system

A voice-activated system converts speech to text, extracts keywords, and searches databases to provide personalized instruction manuals, enhancing work efficiency and accuracy by reducing manual selection time and accounting for user emotions.

JP2026071658APending Publication Date: 2026-04-30SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-17
Publication Date
2026-04-30

AI Technical Summary

Technical Problem

The inefficiency in selecting appropriate instruction manuals and the time-consuming process of determining which manual to use, especially for inexperienced workers, leads to reduced work efficiency and accuracy.

Method used

A system that utilizes voice input through a speech recognition device, converts it to text data, extracts keywords using natural language processing, searches a database for relevant manuals, and presents them to the user, with a generation algorithm to select the most suitable manual.

Benefits of technology

Enables quick and efficient access to necessary instruction manuals, improving work efficiency and accuracy by allowing hands-free operation and personalized guidance based on user emotions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026071658000001_ABST
    Figure 2026071658000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means of receiving work instructions via voice input using a voice recognition device, A means of converting voice input into text data, A method using a natural language processing device to extract keywords from text data, An information processing means that searches for relevant instruction manuals from a database based on extracted keywords, A means having a display device that presents search results to the user, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] At the work site, a large number of instruction manuals are required to handle various tasks, but the process of selecting an appropriate instruction manual from them may take a great deal of time and effort. Also, it is difficult for inexperienced workers to determine which instruction manual to use. Thus, it is an issue to improve the efficiency in the selection and preparation of instruction manuals and achieve a reduction in the overall work time and an improvement in accuracy.

Means for Solving the Problems

[0005] This invention provides a system that receives work instructions via voice input using a speech recognition device. The voice input is converted into text data, and keywords are extracted using natural language processing technology. Based on the extracted keywords, a process is performed to search for relevant instruction manuals from a database, and the results are presented to the user. Furthermore, a generation algorithm is used to select candidate instruction manuals and evaluate their suitability, enabling the selection of the most suitable instruction manual. This system allows users to obtain the necessary instruction manuals quickly and efficiently.

[0006] A "speech recognition device" is a device or software that receives speech input and converts it into digital text.

[0007] "Work instructions" refer to information about the procedures and guidelines necessary to perform a specific task.

[0008] "Text data" refers to the digital representation of character information converted by a speech recognition device.

[0009] A "natural language processing system" is a technology or device used to analyze and extract keywords and meanings from text data.

[0010] A "keyword" refers to an important word or phrase within text data, used to identify its meaning.

[0011] A "database" is a collection of structured information that stores relevant instruction manuals and is searchable based on keywords.

[0012] "Information processing means" refers to a means for implementing a process that retrieves necessary information from a database and manipulates the data according to a specified purpose.

[0013] A "display device" is a screen or display device used to visually provide users with search results and related information.

[0014] A "generation algorithm" is a computational procedure or process for generating a specific output based on input data.

[0015] A "communication device" is a technology of hardware or software for transmitting and receiving data via a network.

Brief Description of Drawings

[0016] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13]It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when combined with an emotion engine. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when combined with an emotion engine.

Modes for Carrying Out the Invention

[0017] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a processor with a reference number (hereinafter simply referred to as "processor") may be one arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be one type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0020] In the following embodiments, a RAM (Random Access Memory) with a reference number is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0021] In the following embodiments, a storage with a reference number is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0024] [First Embodiment]

[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0037] This invention relates to a system for workers to automatically obtain work instructions using voice input. Therefore, embodiments designed to ensure that each device and process functions effectively will be described.

[0038] At the start of a task, the user uses the terminal's voice input function to speak the necessary work instructions into the microphone. The user's voice is converted into digital text data via the terminal's voice recognition device.

[0039] The server receives text data sent from the terminal and uses a natural language processing unit to extract relevant keywords from that text. These keywords serve as important clues in the subsequent database search process.

[0040] Based on the extracted keywords, the server searches the database and selects the most relevant instruction manuals from those held by trading partners and companies. In this context, a generation algorithm is used to provide the optimal choice by matching the keywords with the content of the instruction manuals.

[0041] Subsequently, the terminal retrieves the search results sent from the server and displays the results visually to the user. The user can review these results and select the necessary instructional materials. This selection information is then sent back from the terminal to the server, which provides the instructional materials in an accessible format. Access rights and download links are generated to enable the user to effectively utilize the materials.

[0042] This system allows users to quickly receive the necessary instructions for their work and receive support to make appropriate decisions even in busy work environments. For example, when a work instruction for "daily inspection" is issued, the relevant manuals and checklists for daily inspection can be immediately obtained. In this way, workers can obtain the necessary information in a short time and perform their tasks with high accuracy.

[0043] The following describes the processing flow.

[0044] Step 1:

[0045] The user uses the terminal's voice input function to input work-related requests by voice. The terminal receives this voice in real time and converts it into text data using its built-in speech recognition function. The converted text data is then passed on to the next process.

[0046] Step 2:

[0047] The terminal sends the text data generated by speech recognition to the server. The server utilizes a natural language processing engine to extract keywords relevant to the task from the text data. In this step, important words and phrases are identified and specified through text analysis.

[0048] Step 3:

[0049] The server searches the database for instructional materials based on the extracted keywords. An AI-powered generation algorithm supports this process, matching the keywords with the content of documents in the database to list the most relevant instructional materials.

[0050] Step 4:

[0051] The server sends a list of searched instruction manuals to the terminal. The terminal presents the received list of instruction manuals to the user through a visual interface. It is important here that the user can view previews and summaries of the instruction manuals.

[0052] Step 5:

[0053] The user selects the necessary instruction manual from the displayed options. The terminal sends this selection information to the server, which then prepares the selected manual, sets access permissions, and makes it available for download.

[0054] Step 6:

[0055] The terminal displays instruction manual access information received from the server, allowing the user to download or view the manual. Based on this information, the user can continue with the necessary tasks, enabling efficient work execution.

[0056] (Example 1)

[0057] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0058] Obtaining necessary work instructions quickly on-site is crucial in many work environments. However, conventional systems have the problem of making it difficult for workers to quickly identify and access the specific instructions they need. To address this challenge, a system is needed that improves usability and allows for efficient information acquisition using voice input.

[0059] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0060] In this invention, the server includes means for receiving work instructions via voice input using a voice recognition device, means for converting the voice input into text data, and means for using a natural language processing device to extract keywords from the text data. This makes it possible for workers to efficiently obtain work instructions through voice input.

[0061] A "speech recognition device" is a device that converts speech into digital data and is used to obtain specific commands or information.

[0062] "Voice input" refers to a method of providing voice data to a system using a microphone or other voice receiving device.

[0063] "Text data" refers to digital information that represents audio or other forms of information as written characters.

[0064] A "natural language processing unit" is a device that implements technology to understand and process the language that humans use in everyday life.

[0065] A "keyword" is a word or phrase that is important when searching for or classifying specific information within text data.

[0066] An "information collection" is a database or data store that brings together multiple related documents and data.

[0067] "Information processing means" refers to a device or method for analyzing, retrieving, transforming, or manipulating digital data.

[0068] A "display device" is an electronic device used to present information visually.

[0069] "Communication means" refers to a protocol or device used to send and receive data.

[0070] An "access rights setting mechanism" is a system or process for managing and setting access rights to specific information or resources.

[0071] A "generation algorithm" is a mathematical formula or procedure for generating an output that satisfies specific conditions.

[0072] "User" refers to an individual or organization that uses the system or device.

[0073] This invention relates to a system that allows workers to efficiently obtain work instructions using voice input. In this system, the user provides voice to a terminal, which is then converted into digital text using a speech recognition device. This process is carried out using a common speech recognition technology (e.g., a speech recognition API) integrated into the terminal.

[0074] The server receives text data sent from the terminal and analyzes it using natural language processing techniques. Software such as Python's NLTK and spaCy are often used for this purpose. Through this process, the server extracts necessary keywords from the text.

[0075] Based on the extracted keywords, the server searches the information database and selects the most relevant instruction manuals. A generation algorithm (e.g., a machine learning model) is used to match the specified keywords with the content of the instruction manuals, thereby providing the most relevant information.

[0076] The terminal retrieves search results sent from the server and presents them visually to the user. The user can review the presented results and select the necessary instruction manual. This selection information is sent back to the server, which sets access rights and makes the instruction manual available for download or viewing.

[0077] This system allows users to quickly obtain necessary documents and improve work efficiency. For example, by inputting work instructions related to "daily inspections" via voice, relevant manuals and checklists can be instantly retrieved.

[0078] Examples of prompt messages include the following:

[0079] When given a work instruction to perform a "daily inspection," what kind of manuals or checklists are necessary?

[0080] In this way, users can quickly obtain information from voice input, effectively supporting their actual work.

[0081] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0082] Step 1:

[0083] The user inputs work instructions by voice. When the user says into the microphone, "I want to obtain the daily inspection manual," the voice data is entered into the terminal. This voice data forms the basis for subsequent processing.

[0084] Step 2:

[0085] The terminal uses a speech recognition device to convert voice data into text data. The input voice data is analyzed by a speech recognition API and extracted as text data. For example, the voice saying "I want to obtain the daily inspection manual" is converted into text.

[0086] Step 3:

[0087] The terminal sends the converted text data to the server. This data is securely transmitted to the server using the HTTP protocol, and the server receives this text.

[0088] Step 4:

[0089] The server extracts keywords from the received text data using a natural language processing unit. The natural language processing library identifies keywords (e.g., "daily check," "manual") from the text data. This process clarifies important concepts.

[0090] Step 5:

[0091] The server searches the information database based on the extracted keywords and selects relevant instruction manuals. Using a generative AI model, it matches the specified keywords with data in the digital library and selects relevant documents. For example, manuals related to daily inspections may be obtained as a result.

[0092] Step 6:

[0093] The server sends the search results to the terminal. The information on the selected instruction manuals is sent to the terminal using the HTTP protocol and is ready to be provided to the user.

[0094] Step 7:

[0095] The terminal visually displays the search results received from the server. Users can review the results displayed on the screen and select the necessary instructional materials.

[0096] Step 8:

[0097] The user selects the required instruction manual. The user's selection is recorded by the terminal and resent to the server.

[0098] Step 9:

[0099] The server provides the selected instruction manual in an available format. Based on the selection information, access rights are set, and a link to download or view the instruction manual is generated. The user can then obtain the necessary materials through this link.

[0100] (Application Example 1)

[0101] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0102] There is a problem in that workers have difficulty quickly and efficiently obtaining the necessary instruction manuals when performing tasks on-site. Furthermore, the need to use their hands to refer to the manuals during work reduces work efficiency.

[0103] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0104] In this invention, the server includes means for receiving work instructions via voice input using a voice recognition device, means for converting the voice input into text data, and means for using a natural language processing device to extract keywords from the text data. This makes it possible to quickly obtain necessary instruction manuals via voice instructions and improve work efficiency through visual presentation using visual aids.

[0105] A "speech recognition device" is a device that processes speech as a digital signal and converts it into language data.

[0106] "Voice input" refers to a method of inputting voice commands spoken by the user into the system.

[0107] "Text data" refers to digital character information converted by speech recognition.

[0108] A "natural language processing device" is a device equipped with technology to analyze meaning from text data and extract keywords.

[0109] A "storage medium" is a memory device that records and holds data.

[0110] "Instructional materials" refer to manuals and documents that provide work instructions and procedures.

[0111] An "information processing system" is a mechanism that has the function of processing data and generating and providing necessary information.

[0112] A "display device" refers to a screen or display that visually presents digital information.

[0113] "Visual assistance equipment" refers to devices that visualize information and provide support to help users perform tasks efficiently.

[0114] A "generating algorithm" is a series of steps taken to derive the desired result through the processing and analysis of information.

[0115] A "communication device" is a device that possesses the technology to send and receive data.

[0116] This invention is a system for improving work efficiency by enabling workers to quickly obtain necessary work instructions. The user provides specific work instructions via voice input using a voice recognition device. These voice instructions are converted into digital text data by a terminal. The generated text data is received by a server and analyzed using a device employing natural language processing technology. During the analysis, relevant keywords are extracted, and based on these, appropriate guidance materials are searched from storage media on the server. The server evaluates the suitability of the guidance materials using a generated algorithm. Once appropriate materials are selected, the relevant information is transmitted to the terminal. The terminal then displays the guidance materials on a display device in the user's possession.

[0117] Furthermore, this system integrates with visual assistance devices, allowing users to acquire information visually. This improves work efficiency by enabling users to refer to guide materials without using their hands during work. For example, when assembling a new part in a factory, a voice command such as "Tell me the assembly procedure for part A" will display the relevant work instructions on the glasses display. An example of a prompt message is: "Text to send to the server when the user says 'Tell me the assembly procedure for part A': 'Assembly procedure for part A'". Based on this prompt message, the system can provide guide materials that are appropriate to the workflow in a timely manner.

[0118] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0119] Step 1:

[0120] The user voice-inputs work instructions into a voice-recognition device. This input is received as an audio signal by the terminal. The terminal receives this audio signal and converts it into digital text data using its speech recognition engine. The input is an audio signal, and the output is text data. Through this process, the user's intended message is organized into textual information.

[0121] Step 2:

[0122] The terminal sends the generated text data to the server. The server receives this text data and uses a natural language processing device to extract keywords. The input is text data, and the output is a list of extracted keywords. This data processing gives the server the clues necessary to identify relevant information.

[0123] Step 3:

[0124] The server searches for relevant informational materials from the storage medium based on the extracted keywords. The input is a list of keywords, and the output is the relevant informational materials. This search process evaluates the relevance between the keywords and the content of the informational materials to select the most relevant materials.

[0125] Step 4:

[0126] The server sends the selected informational materials to the terminal. The terminal displays the received information visually on the user's visual aids. The input is the informational materials, and the output is the visual information displayed on the equipment. This operation allows the user to check the necessary information in real time while working.

[0127] Step 5:

[0128] The user checks the guidance materials displayed on the visual aid and proceeds with the work. Because the user can efficiently refer to the instructions without taking their hands off the device, they can continue performing highly accurate work without interrupting the workflow. The input is the information displayed on the visual device, and the output is the user's work actions. This operation enables a fluid workflow.

[0129] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0130] This invention relates to a system that incorporates an emotion engine to recognize user emotions and provide work instructions in a more personalized form. The following is a specific description of how this system is implemented.

[0131] The user provides voice input regarding their task to the terminal. This voice data is collected by the terminal's voice recognition device and converted into digital text data. Simultaneously, the terminal is equipped with an emotion engine that analyzes the tone, tempo, and pitch of the user's voice to identify their emotions.

[0132] The server receives text data sent from the terminal and extracts keywords from it. It also receives sentiment data analyzed by the sentiment engine. Combining the extracted keywords with the user's sentiment information, the server searches the database for the most suitable instructional guide.

[0133] The search process utilizes a generation algorithm to select the most suitable guide, taking into account both keywords and emotions. For example, if the emotion engine detects that the user is stressed, it adjusts the system to prioritize providing simple and easy-to-understand guides.

[0134] The terminal visually displays a list of instruction manuals sent from the server to the user. The content of the instruction manuals presented is tailored to the user's emotions, thus more effectively meeting their needs. The user selects an instruction manual that aligns with their emotions and downloads or views it on screen as needed.

[0135] This system is designed to improve work efficiency and reduce worker stress by providing content tailored to the user's psychological state. For example, users who feel anxious about their work are provided with simplified instructions and encouraging messages, allowing them to continue working with peace of mind.

[0136] The following describes the processing flow.

[0137] Step 1:

[0138] The user inputs the necessary instructions for the task by voice into the terminal. This voice is instantly converted into text data by the terminal's voice recognition device. Simultaneously, the terminal's built-in emotion engine analyzes emotional information from the user's voice. Characteristics such as pitch, speed, and tone of the voice are analyzed to determine what the user is feeling.

[0139] Step 2:

[0140] The device sends the converted text data and analyzed sentiment data to the server. The server then begins the process of extracting necessary keywords from the received text data. Using natural language processing, it efficiently finds keywords from the user's instructions.

[0141] Step 3:

[0142] The server searches the database for instructional materials based on extracted keywords and sentiment data. The generation algorithm then lists the most suitable options, taking sentiment data into consideration. For example, if a user is feeling anxious, instructional materials that provide reassurance can be prioritized.

[0143] Step 4:

[0144] A list of candidate instruction manuals provided by the server is sent to the terminal. The terminal visually displays this list to the user. The displayed instruction manuals are filtered according to the user's emotional state, making them highly likely to appropriately meet the user's needs.

[0145] Step 5:

[0146] The user selects the most suitable instruction manual from those presented. The selection is made directly on the device using actions such as tapping or clicking. Based on the selection, the device sends feedback to the server.

[0147] Step 6:

[0148] The server makes the user's selected instruction manual immediately available and generates a download link as needed. The terminal presents this information to the user, preparing them to download or view the instruction manual. Emotionally conscious instruction manuals enable users to perform their tasks effectively.

[0149] (Example 2)

[0150] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0151] Existing work instruction systems fail to provide individualized instruction materials that take into account the emotional state of users, sometimes resulting in the provision of information that is not most effective for the user. Furthermore, because work instruction is general in nature, it is difficult for users to tailor the instruction to their own specific situation. This can lead to decreased work efficiency and increased stress.

[0152] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0153] In this invention, the server includes means for receiving voice input via an acoustic processing device, means for converting voice data into text information, and means for using a natural language processing device that extracts keywords from the text information and utilizes information for sentiment analysis. This makes it possible to efficiently provide individualized instructional materials that take into account the emotional state of the user.

[0154] An "acoustic processing device" is a device that receives audio input and converts it into a format suitable for analyzing or processing the data.

[0155] "Speech conversion means" refers to a process or device that analyzes speech data and converts it into digital text format.

[0156] A "natural language processing device" is a technology or device that extracts keywords from text data and generates information for sentiment analysis.

[0157] "Emotional analysis information" refers to emotional data obtained from voice and text data, and is used to identify the user's emotional state.

[0158] A "knowledge base" is a database or information source that stores relevant teaching materials and is used to provide the most relevant information through searching.

[0159] "Information retrieval means" refers to a process or apparatus for retrieving specific data or information and providing results tailored to the user.

[0160] A "generative algorithm" is a set of computational steps to create, select, or evaluate the optimal output based on input data.

[0161] "Communication means" refers to the process or technology used to distribute or download instructional materials.

[0162] An "operator" refers to a person or entity that uses this system to receive work instructions.

[0163] This invention is a system that provides user-specific instructional materials based on voice input. The following hardware and software are used to implement this invention.

[0164] The terminal receives work instructions from the user via voice. The terminal is equipped with a voice recognition device that converts voice data into digital text. For example, a common voice recognition API is used for voice conversion.

[0165] Simultaneously, the device runs emotion analysis software. This software analyzes the tone and pitch of the voice to identify the user's emotional state. It then generates emotion analysis information and transmits it to the server via the communication path.

[0166] The server receives text information and sentiment analysis information sent from the terminal. The server uses a text analysis engine to extract keywords from the text data. This process may utilize, for example, a natural language processing library. The server also takes sentiment analysis information into account when searching for the most suitable teaching materials from a knowledge base. Generative AI models are used in the search algorithm.

[0167] For example, if a user inputs a voice command such as "I want to finish this task quickly," the server combines keywords like "quickly" with emotion analysis information such as "anxiety" and processes it. As a result, the server selects simplified instructional materials that are appropriate for the user's situation.

[0168] For example, it's possible to input prompts such as, "Generate a work instruction manual suitable for a user who is feeling anxious." This enables support tailored to the user's emotional state.

[0169] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0170] Step 1:

[0171] The user inputs work instructions by voice into the terminal.

[0172] Input: User's voice instructions

[0173] Output: Unallocated audio data

[0174] Step 2:

[0175] The terminal uses a speech recognition device to convert the user's voice data into digital text. This process utilizes speech recognition software.

[0176] Input: User's voice data

[0177] Output: Digital text data

[0178] Step 3:

[0179] The device acquires voice data and simultaneously runs emotion analysis software, identifying the user's emotions by analyzing the tone and pitch of the voice.

[0180] Input: User's voice data

[0181] Output: Sentiment analysis information

[0182] Step 4:

[0183] The device sends text data and sentiment analysis information to the server.

[0184] Input: Digital text data and sentiment analysis information

[0185] Output: Sending data to the server

[0186] Step 5:

[0187] The server extracts keywords from the received text data using a natural language processing library.

[0188] Input: Digital text data

[0189] Output: Keyword set

[0190] Step 6:

[0191] The server searches for the most suitable teaching materials from its knowledge base based on the extracted keywords and sentiment analysis information. A generative AI model is used in this process.

[0192] Input: Keyword set and sentiment analysis information

[0193] Output: Selection of optimal teaching materials

[0194] Step 7:

[0195] The server sends the retrieved instructional materials to the terminal.

[0196] Input: Optimal teaching materials

[0197] Output: Sending data to the terminal

[0198] Step 8:

[0199] The terminal visually presents the received instructional materials to the user. Specifically, it displays them on the screen through the user interface.

[0200] Input: Instructional materials sent from the server

[0201] Output: Visual presentation to the user

[0202] (Application Example 2)

[0203] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0204] When workers receive work instructions, if those instructions are uniform, efficiency may decrease depending on the worker's psychological state. In particular, if the instructions given to workers experiencing stress or anxiety are inappropriate, work efficiency and safety may decrease.

[0205] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0206] In this invention, the server includes means for analyzing emotions from voice data, means for converting voice input into text data, and means for searching a database for the most suitable instruction manual based on keywords and emotion data. This makes it possible to provide more personalized instruction manuals that are tailored to the emotional state of the worker.

[0207] A "speech recognition device" is a device that receives speech as a digital signal, analyzes its content, and converts it into text data.

[0208] A "natural language processing unit" is a device that processes text data to extract important words and phrases.

[0209] "Information processing means" refers to means that provide a method for retrieving relevant information based on given data and generating or suggesting appropriate results.

[0210] "Emotional data" refers to data about the speaker's psychological or emotional state, obtained by analyzing the characteristics of their voice.

[0211] A "generative algorithm" is a computational method for generating highly relevant results or candidates based on specific input data.

[0212] A "display device" is a device used to visually present processed information to a user.

[0213] A "communication device" is a device that transmits and receives data, enabling data exchange with the outside world.

[0214] The system of this invention is used in a work environment in which an operator wears smart glasses and receives voice instructions. The system mainly consists of a terminal, a server, and smart glasses.

[0215] The device features a function that converts user speech into text data using the Google® Cloud Speech-to-Text API. It also analyzes user emotions based on speech tone and speed using the Microsoft® Azure® Emotion API. Both the emotion data and the text data are then transferred to the server.

[0216] The server uses a generation algorithm based on TENSORFLOW® to search the database for the most suitable instruction manual based on the received emotion and text data. The selected instruction manual takes into account the user's psychological state and is appropriately customized to maximize work efficiency.

[0217] Users view instructions sent from the server on the display of their smart glasses. For example, users who are feeling stressed about a task can be provided with simple and visually easy-to-understand instructions, allowing them to proceed with the task with confidence.

[0218] For example, when a user says, "This task is difficult," stress can be detected from things like mispronunciation or rising intonation. The system then displays a step-by-step, visually guided procedure manual, color-coded accordingly. In this way, it becomes possible to provide task guidance tailored to the user's emotional state.

[0219] Examples of prompt statements include:

[0220] The message is: "We detected audio from a user saying 'the task is difficult' and indicating signs of stress. Please suggest a simplified guide."

[0221] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0222] Step 1:

[0223] The user provides voice input through smart glasses. The input voice data is collected by the device. The collected voice data is then converted into text data using the Google Cloud Speech-to-Text API. This process converts the voice data into a parseable text format.

[0224] Step 2:

[0225] The device simultaneously uses Microsoft Azure's Emotion API to analyze the user's emotions from the tone and speed of their voice. At this stage, it extracts emotional data from the input voice data and interprets the results of the emotion analysis. The output is data indicating the user's emotional state.

[0226] Step 3:

[0227] These text and sentiment data are sent from the terminal to the server. The server extracts keywords from the received text data and then integrates them with the received sentiment data. This data integration creates the dataset necessary for searching.

[0228] Step 4:

[0229] The server executes a generative algorithm using TensorFlow. Here, it searches the database for the most relevant instructional guides based on the integrated dataset. The generative algorithm evaluates both keywords and sentiment data to select the appropriate guide.

[0230] Step 5:

[0231] The selected instruction manuals are sent to the smart glasses in an optimized format to enhance the user's work efficiency. The user reviews the instruction manuals on the smart glasses' display. Here, the system enables a personalized learning experience tailored to the user's psychological state.

[0232] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0233] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0234] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0235] [Second Embodiment]

[0236] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0237] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0238] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0239] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0240] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0241] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0242] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0243] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0244] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0245] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0246] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0247] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0248] This invention relates to a system for workers to automatically obtain work instructions using voice input. Therefore, embodiments designed to ensure that each device and process functions effectively will be described.

[0249] At the start of a task, the user uses the terminal's voice input function to speak the necessary work instructions into the microphone. The user's voice is converted into digital text data via the terminal's voice recognition device.

[0250] The server receives text data sent from the terminal and uses a natural language processing unit to extract relevant keywords from that text. These keywords serve as important clues in the subsequent database search process.

[0251] Based on the extracted keywords, the server searches the database and selects the most relevant instruction manuals from those held by trading partners and companies. In this context, a generation algorithm is used to provide the optimal choice by matching the keywords with the content of the instruction manuals.

[0252] Subsequently, the terminal retrieves the search results sent from the server and displays the results visually to the user. The user can review these results and select the necessary instructional materials. This selection information is then sent back from the terminal to the server, which provides the instructional materials in an accessible format. Access rights and download links are generated to enable the user to effectively utilize the materials.

[0253] This system allows users to quickly receive the necessary instructions for their work and receive support to make appropriate decisions even in busy work environments. For example, when a work instruction for "daily inspection" is issued, the relevant manuals and checklists for daily inspection can be immediately obtained. In this way, workers can obtain the necessary information in a short time and perform their tasks with high accuracy.

[0254] The following describes the processing flow.

[0255] Step 1:

[0256] The user uses the terminal's voice input function to input work-related requests by voice. The terminal receives this voice in real time and converts it into text data using its built-in speech recognition function. The converted text data is then passed on to the next process.

[0257] Step 2:

[0258] The terminal sends the text data generated by speech recognition to the server. The server utilizes a natural language processing engine to extract keywords relevant to the task from the text data. In this step, important words and phrases are identified and specified through text analysis.

[0259] Step 3:

[0260] The server searches the database for instructional materials based on the extracted keywords. An AI-powered generation algorithm supports this process, matching the keywords with the content of documents in the database to list the most relevant instructional materials.

[0261] Step 4:

[0262] The server sends a list of searched instruction manuals to the terminal. The terminal presents the received list of instruction manuals to the user through a visual interface. It is important here that the user can view previews and summaries of the instruction manuals.

[0263] Step 5:

[0264] The user selects the necessary instruction manual from the displayed options. The terminal sends this selection information to the server, which then prepares the selected manual, sets access permissions, and makes it available for download.

[0265] Step 6:

[0266] The terminal displays instruction manual access information received from the server, allowing the user to download or view the manual. Based on this information, the user can continue with the necessary tasks, enabling efficient work execution.

[0267] (Example 1)

[0268] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0269] Obtaining necessary work instructions quickly on-site is crucial in many work environments. However, conventional systems have the problem of making it difficult for workers to quickly identify and access the specific instructions they need. To address this challenge, a system is needed that improves usability and allows for efficient information acquisition using voice input.

[0270] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0271] In this invention, the server includes means for receiving work instructions via voice input using a voice recognition device, means for converting the voice input into text data, and means for using a natural language processing device to extract keywords from the text data. This makes it possible for workers to efficiently obtain work instructions through voice input.

[0272] A "speech recognition device" is a device that converts speech into digital data and is used to obtain specific commands or information.

[0273] "Voice input" refers to a method of providing voice data to a system using a microphone or other voice receiving device.

[0274] "Text data" refers to digital information that represents audio or other forms of information as written characters.

[0275] A "natural language processing unit" is a device that implements technology to understand and process the language that humans use in everyday life.

[0276] A "keyword" is a word or phrase that is important when searching for or classifying specific information within text data.

[0277] "Information collection" refers to a database or data store that aggregates multiple related documents and data.

[0278] "Information processing means" refers to a device or method for analyzing, searching, converting, or manipulating digital data.

[0279] "Display device" refers to an electronic device for visually presenting information.

[0280] "Communication means" refers to a protocol or device used for transmitting and receiving data.

[0281] "Access right setting means" refers to a system or process for managing and setting access rights to specific information or resources.

[0282] "Generation algorithm" refers to a calculation formula or procedure for generating an output that meets specific conditions.

[0283] "User" refers to an individual or organization that uses a system or device.

[0284] This invention relates to a system that enables an operator to efficiently obtain work instructions using voice input. Here, the user provides voice to a terminal, and the voice is converted into digital text using a voice recognition device. This process is carried out using general voice recognition technology (e.g., voice recognition API) integrated into the terminal. <000900>

[0285] The server receives the text data transmitted from the terminal and analyzes it using natural language processing technology. Here, software such as Python's NLTK and spaCy is often used. Through this process, the server extracts the necessary keywords from the text.

[0286] Based on the extracted keywords, the server searches for information collections and selects highly relevant instruction manuals. A generation algorithm (for example, a machine learning model) is used to match the specified keywords with the content of the instruction manuals to provide optimal information.

[0287] The terminal obtains the search results sent by the server and visually presents them to the user. The user can check the presented results and select the necessary instruction manuals. This selection information is sent back to the server, and the server sets the access rights to make the instruction manuals downloadable or viewable.

[0288] With this system, users can quickly obtain the necessary materials, improving work efficiency. As a specific example, by inputting work instructions related to "daily inspection" by voice, relevant manuals and checklists can be obtained immediately.

[0289] Examples of prompt sentences are as follows.

[0290] When receiving the work instruction of "daily inspection", please tell me what manuals and checklists are needed.

[0291] In this way, users can obtain information in a short time from voice input and be effectively supported in actual work.

[0292] The flow of the specific process in Example 1 will be described using FIG. 11.

[0293] Step 1:

[0294] The user inputs the work instruction by voice. When the user says "I want to obtain the daily inspection manual" towards the microphone, voice data is input into the terminal. This voice data serves as the basis for subsequent processing.

[0295] Step 2:

[0296] The terminal uses a speech recognition device to convert voice data into text data. The input voice data is analyzed by a speech recognition API and extracted as text data. For example, the voice saying "I want to obtain the daily inspection manual" is converted into text.

[0297] Step 3:

[0298] The terminal sends the converted text data to the server. This data is securely transmitted to the server using the HTTP protocol, and the server receives this text.

[0299] Step 4:

[0300] The server extracts keywords from the received text data using a natural language processing unit. The natural language processing library identifies keywords (e.g., "daily check," "manual") from the text data. This process clarifies important concepts.

[0301] Step 5:

[0302] The server searches the information database based on the extracted keywords and selects relevant instruction manuals. Using a generative AI model, it matches the specified keywords with data in the digital library and selects relevant documents. For example, manuals related to daily inspections may be obtained as a result.

[0303] Step 6:

[0304] The server sends the search results to the terminal. The information on the selected instruction manuals is sent to the terminal using the HTTP protocol and is ready to be provided to the user.

[0305] Step 7:

[0306] The terminal visually displays the search results received from the server. Users can review the results displayed on the screen and select the necessary instructional materials.

[0307] Step 8:

[0308] The user selects the required instruction manual. The user's selection is recorded by the terminal and resent to the server.

[0309] Step 9:

[0310] The server provides the selected instruction manual in an available form. With the selection information as input, access rights are set, and a link for downloading or viewing the instruction manual is generated. The user can obtain the required materials through this link.

[0311] (Application Example 1)

[0312] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".

[0313] When a worker performs work at the site, there is a problem that it is difficult to quickly and efficiently obtain the required instruction manual. In addition, there is an issue that hands must be used to refer to the instruction manual during work, resulting in a decrease in work efficiency.

[0314] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0315] In this invention, the server includes means for receiving a work instruction by voice input using a voice recognition device, means for converting the voice input into text data, and means using a natural language processing device for extracting keywords from the text data. Thereby, it becomes possible to quickly obtain the required instruction manual by voice instruction and improve work efficiency through visual presentation using visual assistance equipment.

[0316] The "voice recognition device" is a device that processes voice as a digital signal and converts it into language data.

[0317] "Voice input" refers to a method of inputting voice commands spoken by the user into the system.

[0318] "Text data" refers to digital character information converted by speech recognition.

[0319] A "natural language processing device" is a device equipped with technology to analyze meaning from text data and extract keywords.

[0320] A "storage medium" is a memory device that records and holds data.

[0321] "Instructional materials" refer to manuals and documents that provide work instructions and procedures.

[0322] An "information processing system" is a mechanism that has the function of processing data and generating and providing necessary information.

[0323] A "display device" refers to a screen or display that visually presents digital information.

[0324] "Visual assistance equipment" refers to devices that visualize information and provide support to help users perform tasks efficiently.

[0325] A "generating algorithm" is a series of steps taken to derive the desired result through the processing and analysis of information.

[0326] A "communication device" is a device that possesses the technology to send and receive data.

[0327] This invention is a system for improving work efficiency by enabling workers to quickly obtain necessary work instructions. The user provides specific work instructions via voice input using a voice recognition device. These voice instructions are converted into digital text data by a terminal. The generated text data is received by a server and analyzed using a device employing natural language processing technology. During the analysis, relevant keywords are extracted, and based on these, appropriate guidance materials are searched from storage media on the server. The server evaluates the suitability of the guidance materials using a generated algorithm. Once appropriate materials are selected, the relevant information is transmitted to the terminal. The terminal then displays the guidance materials on a display device in the user's possession.

[0328] Furthermore, this system integrates with visual assistance devices, allowing users to acquire information visually. This improves work efficiency by enabling users to refer to guide materials without using their hands during work. For example, when assembling a new part in a factory, a voice command such as "Tell me the assembly procedure for part A" will display the relevant work instructions on the glasses display. An example of a prompt message is: "Text to send to the server when the user says 'Tell me the assembly procedure for part A': 'Assembly procedure for part A'". Based on this prompt message, the system can provide guide materials that are appropriate to the workflow in a timely manner.

[0329] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0330] Step 1:

[0331] The user voice-inputs work instructions into a voice-recognition device. This input is received as an audio signal by the terminal. The terminal receives this audio signal and converts it into digital text data using its speech recognition engine. The input is an audio signal, and the output is text data. Through this process, the user's intended message is organized into textual information.

[0332] Step 2:

[0333] The terminal sends the generated text data to the server. The server receives this text data and uses a natural language processing device to extract keywords. The input is text data, and the output is a list of extracted keywords. This data processing gives the server the clues necessary to identify relevant information.

[0334] Step 3:

[0335] The server searches for relevant informational materials from the storage medium based on the extracted keywords. The input is a list of keywords, and the output is the relevant informational materials. This search process evaluates the relevance between the keywords and the content of the informational materials to select the most relevant materials.

[0336] Step 4:

[0337] The server sends the selected informational materials to the terminal. The terminal displays the received information visually on the user's visual aids. The input is the informational materials, and the output is the visual information displayed on the equipment. This operation allows the user to check the necessary information in real time while working.

[0338] Step 5:

[0339] The user checks the guidance materials displayed on the visual aid and proceeds with the work. Because the user can efficiently refer to the instructions without taking their hands off the device, they can continue performing highly accurate work without interrupting the workflow. The input is the information displayed on the visual device, and the output is the user's work actions. This operation enables a fluid workflow.

[0340] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0341] This invention relates to a system that incorporates an emotion engine to recognize user emotions and provide work instructions in a more personalized form. The following is a specific description of how this system is implemented.

[0342] The user provides voice input regarding their task to the terminal. This voice data is collected by the terminal's voice recognition device and converted into digital text data. Simultaneously, the terminal is equipped with an emotion engine that analyzes the tone, tempo, and pitch of the user's voice to identify their emotions.

[0343] The server receives text data sent from the terminal and extracts keywords from it. It also receives sentiment data analyzed by the sentiment engine. Combining the extracted keywords with the user's sentiment information, the server searches the database for the most suitable instructional guide.

[0344] The search process utilizes a generation algorithm to select the most suitable guide, taking into account both keywords and emotions. For example, if the emotion engine detects that the user is stressed, it adjusts the system to prioritize providing simple and easy-to-understand guides.

[0345] The terminal visually displays a list of instruction manuals sent from the server to the user. The content of the instruction manuals presented is tailored to the user's emotions, thus more effectively meeting their needs. The user selects an instruction manual that aligns with their emotions and downloads or views it on screen as needed.

[0346] This system is designed to improve work efficiency and reduce worker stress by providing content tailored to the user's psychological state. For example, users who feel anxious about their work are provided with simplified instructions and encouraging messages, allowing them to continue working with peace of mind.

[0347] The following describes the processing flow.

[0348] Step 1:

[0349] The user inputs the necessary instructions for the task by voice into the terminal. This voice is instantly converted into text data by the terminal's voice recognition device. Simultaneously, the terminal's built-in emotion engine analyzes emotional information from the user's voice. Characteristics such as pitch, speed, and tone of the voice are analyzed to determine what the user is feeling.

[0350] Step 2:

[0351] The device sends the converted text data and analyzed sentiment data to the server. The server then begins the process of extracting necessary keywords from the received text data. Using natural language processing, it efficiently finds keywords from the user's instructions.

[0352] Step 3:

[0353] The server searches the database for instructional materials based on extracted keywords and sentiment data. The generation algorithm then lists the most suitable options, taking sentiment data into consideration. For example, if a user is feeling anxious, instructional materials that provide reassurance can be prioritized.

[0354] Step 4:

[0355] A list of candidate instruction manuals provided by the server is sent to the terminal. The terminal visually displays this list to the user. The displayed instruction manuals are filtered according to the user's emotional state, making them highly likely to appropriately meet the user's needs.

[0356] Step 5:

[0357] The user selects the most suitable instruction manual from those presented. The selection is made directly on the device using actions such as tapping or clicking. Based on the selection, the device sends feedback to the server.

[0358] Step 6:

[0359] The server makes the user's selected instruction manual immediately available and generates a download link as needed. The terminal presents this information to the user, preparing them to download or view the instruction manual. Emotionally conscious instruction manuals enable users to perform their tasks effectively.

[0360] (Example 2)

[0361] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0362] Existing work instruction systems fail to provide individualized instruction materials that take into account the emotional state of users, sometimes resulting in the provision of information that is not most effective for the user. Furthermore, because work instruction is general in nature, it is difficult for users to tailor the instruction to their own specific situation. This can lead to decreased work efficiency and increased stress.

[0363] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0364] In this invention, the server includes means for receiving voice input via an acoustic processing device, means for converting voice data into text information, and means for using a natural language processing device that extracts keywords from the text information and utilizes information for sentiment analysis. This makes it possible to efficiently provide individualized instructional materials that take into account the emotional state of the user.

[0365] An "acoustic processing device" is a device that receives audio input and converts it into a format suitable for analyzing or processing the data.

[0366] "Speech conversion means" refers to a process or device that analyzes speech data and converts it into digital text format.

[0367] A "natural language processing device" is a technology or device that extracts keywords from text data and generates information for sentiment analysis.

[0368] "Emotional analysis information" refers to emotional data obtained from voice and text data, and is used to identify the user's emotional state.

[0369] A "knowledge base" is a database or information source that stores relevant teaching materials and is used to provide the most relevant information through searching.

[0370] "Information retrieval means" refers to a process or apparatus for retrieving specific data or information and providing results tailored to the user.

[0371] A "generative algorithm" is a set of computational steps to create, select, or evaluate the optimal output based on input data.

[0372] "Communication means" refers to the process or technology used to distribute or download instructional materials.

[0373] An "operator" refers to a person or entity that uses this system to receive work instructions.

[0374] This invention is a system that provides user-specific instructional materials based on voice input. The following hardware and software are used to implement this invention.

[0375] The terminal receives work instructions from the user via voice. The terminal is equipped with a voice recognition device that converts voice data into digital text. For example, a common voice recognition API is used for voice conversion.

[0376] Simultaneously, the device runs emotion analysis software. This software analyzes the tone and pitch of the voice to identify the user's emotional state. It then generates emotion analysis information and transmits it to the server via the communication path.

[0377] The server receives text information and sentiment analysis information sent from the terminal. The server uses a text analysis engine to extract keywords from the text data. This process may utilize, for example, a natural language processing library. The server also takes sentiment analysis information into account when searching for the most suitable teaching materials from a knowledge base. Generative AI models are used in the search algorithm.

[0378] For example, if a user inputs a voice command such as "I want to finish this task quickly," the server combines keywords like "quickly" with emotion analysis information such as "anxiety" and processes it. As a result, the server selects simplified instructional materials that are appropriate for the user's situation.

[0379] For example, it's possible to input prompts such as, "Generate a work instruction manual suitable for a user who is feeling anxious." This enables support tailored to the user's emotional state.

[0380] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0381] Step 1:

[0382] The user inputs work instructions by voice into the terminal.

[0383] Input: User's voice instructions

[0384] Output: Unallocated audio data

[0385] Step 2:

[0386] The terminal uses a speech recognition device to convert the user's voice data into digital text. This process utilizes speech recognition software.

[0387] Input: User's voice data

[0388] Output: Digital text data

[0389] Step 3:

[0390] The device acquires voice data and simultaneously runs emotion analysis software, identifying the user's emotions by analyzing the tone and pitch of the voice.

[0391] Input: User's voice data

[0392] Output: Sentiment analysis information

[0393] Step 4:

[0394] The device sends text data and sentiment analysis information to the server.

[0395] Input: Digital text data and sentiment analysis information

[0396] Output: Sending data to the server

[0397] Step 5:

[0398] The server extracts keywords from the received text data using a natural language processing library.

[0399] Input: Digital text data

[0400] Output: Keyword set

[0401] Step 6:

[0402] The server searches for the most suitable teaching materials from its knowledge base based on the extracted keywords and sentiment analysis information. A generative AI model is used in this process.

[0403] Input: Keyword set and sentiment analysis information

[0404] Output: Selection of optimal teaching materials

[0405] Step 7:

[0406] The server sends the retrieved instructional materials to the terminal.

[0407] Input: Optimal teaching materials

[0408] Output: Sending data to the terminal

[0409] Step 8:

[0410] The terminal visually presents the received instructional materials to the user. Specifically, it displays them on the screen through the user interface.

[0411] Input: Instructional materials sent from the server

[0412] Output: Visual presentation to the user

[0413] (Application Example 2)

[0414] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0415] When workers receive work instructions, if those instructions are uniform, efficiency may decrease depending on the worker's psychological state. In particular, if the instructions given to workers experiencing stress or anxiety are inappropriate, work efficiency and safety may decrease.

[0416] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0417] In this invention, the server includes means for analyzing emotions from voice data, means for converting voice input into text data, and means for searching a database for the most suitable instruction manual based on keywords and emotion data. This makes it possible to provide more personalized instruction manuals that are tailored to the emotional state of the worker.

[0418] A "speech recognition device" is a device that receives speech as a digital signal, analyzes its content, and converts it into text data.

[0419] A "natural language processing unit" is a device that processes text data to extract important words and phrases.

[0420] "Information processing means" refers to means that provide a method for retrieving relevant information based on given data and generating or suggesting appropriate results.

[0421] "Emotional data" refers to data about the speaker's psychological or emotional state, obtained by analyzing the characteristics of their voice.

[0422] A "generative algorithm" is a computational method for generating highly relevant results or candidates based on specific input data.

[0423] A "display device" is a device used to visually present processed information to a user.

[0424] A "communication device" is a device that transmits and receives data, enabling data exchange with the outside world.

[0425] The system of this invention is used in a work environment in which an operator wears smart glasses and receives voice instructions. The system mainly consists of a terminal, a server, and smart glasses.

[0426] The device features a function that converts user speech into text data using the Google Cloud Speech-to-Text API. It also uses the Microsoft Azure Emotion API to analyze user emotions based on speech tone, speed, and other factors. Both the emotion data and the text data are then transferred to a server.

[0427] The server uses a TensorFlow-based generation algorithm to search the database for the most suitable instruction manual based on the received sentiment and text data. The selected instruction manual takes the user's psychological state into consideration and is appropriately customized to maximize work efficiency.

[0428] Users view instructions sent from the server on the display of their smart glasses. For example, users who are feeling stressed about a task can be provided with simple and visually easy-to-understand instructions, allowing them to proceed with the task with confidence.

[0429] For example, when a user says, "This task is difficult," stress can be detected from things like mispronunciation or rising intonation. The system then displays a step-by-step, visually guided procedure manual, color-coded accordingly. In this way, it becomes possible to provide task guidance tailored to the user's emotional state.

[0430] Examples of prompt statements include:

[0431] The message is: "We detected audio from a user saying 'the task is difficult' and indicating signs of stress. Please suggest a simplified guide."

[0432] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0433] Step 1:

[0434] The user provides voice input through smart glasses. The input voice data is collected by the device. The collected voice data is then converted into text data using the Google Cloud Speech-to-Text API. This process converts the voice data into a parseable text format.

[0435] Step 2:

[0436] The device simultaneously uses Microsoft Azure's Emotion API to analyze the user's emotions from the tone and speed of their voice. At this stage, it extracts emotional data from the input voice data and interprets the results of the emotion analysis. The output is data indicating the user's emotional state.

[0437] Step 3:

[0438] These text and sentiment data are sent from the terminal to the server. The server extracts keywords from the received text data and then integrates them with the received sentiment data. This data integration creates the dataset necessary for searching.

[0439] Step 4:

[0440] The server executes a generative algorithm using TensorFlow. Here, it searches the database for the most relevant instructional guides based on the integrated dataset. The generative algorithm evaluates both keywords and sentiment data to select the appropriate guide.

[0441] Step 5:

[0442] The selected instruction manuals are sent to the smart glasses in an optimized format to enhance the user's work efficiency. The user reviews the instruction manuals on the smart glasses' display. Here, the system enables a personalized learning experience tailored to the user's psychological state.

[0443] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0444] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0445] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0446] [Third Embodiment]

[0447] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0448] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0449] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0450] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0451] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0452] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0453] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0454] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0455] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0456] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0457] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0458] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0459] This invention relates to a system for workers to automatically obtain work instructions using voice input. Therefore, embodiments designed to ensure that each device and process functions effectively will be described.

[0460] At the start of a task, the user uses the terminal's voice input function to speak the necessary work instructions into the microphone. The user's voice is converted into digital text data via the terminal's voice recognition device.

[0461] The server receives text data sent from the terminal and uses a natural language processing unit to extract relevant keywords from that text. These keywords serve as important clues in the subsequent database search process.

[0462] Based on the extracted keywords, the server searches the database and selects the most relevant instruction manuals from those held by trading partners and companies. In this context, a generation algorithm is used to provide the optimal choice by matching the keywords with the content of the instruction manuals.

[0463] Subsequently, the terminal retrieves the search results sent from the server and displays the results visually to the user. The user can review these results and select the necessary instructional materials. This selection information is then sent back from the terminal to the server, which provides the instructional materials in an accessible format. Access rights and download links are generated to enable the user to effectively utilize the materials.

[0464] This system allows users to quickly receive the necessary instructions for their work and receive support to make appropriate decisions even in busy work environments. For example, when a work instruction for "daily inspection" is issued, the relevant manuals and checklists for daily inspection can be immediately obtained. In this way, workers can obtain the necessary information in a short time and perform their tasks with high accuracy.

[0465] The following describes the processing flow.

[0466] Step 1:

[0467] The user uses the terminal's voice input function to input work-related requests by voice. The terminal receives this voice in real time and converts it into text data using its built-in speech recognition function. The converted text data is then passed on to the next process.

[0468] Step 2:

[0469] The terminal sends the text data generated by speech recognition to the server. The server utilizes a natural language processing engine to extract keywords relevant to the task from the text data. In this step, important words and phrases are identified and specified through text analysis.

[0470] Step 3:

[0471] The server searches the database for instructional materials based on the extracted keywords. An AI-powered generation algorithm supports this process, matching the keywords with the content of documents in the database to list the most relevant instructional materials.

[0472] Step 4:

[0473] The server sends a list of searched instruction manuals to the terminal. The terminal presents the received list of instruction manuals to the user through a visual interface. It is important here that the user can view previews and summaries of the instruction manuals.

[0474] Step 5:

[0475] The user selects the necessary instruction manual from the displayed options. The terminal sends this selection information to the server, which then prepares the selected manual, sets access permissions, and makes it available for download.

[0476] Step 6:

[0477] The terminal displays instruction manual access information received from the server, allowing the user to download or view the manual. Based on this information, the user can continue with the necessary tasks, enabling efficient work execution.

[0478] (Example 1)

[0479] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0480] Obtaining necessary work instructions quickly on-site is crucial in many work environments. However, conventional systems have the problem of making it difficult for workers to quickly identify and access the specific instructions they need. To address this challenge, a system is needed that improves usability and allows for efficient information acquisition using voice input.

[0481] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0482] In this invention, the server includes means for receiving work instructions via voice input using a voice recognition device, means for converting the voice input into text data, and means for using a natural language processing device to extract keywords from the text data. This makes it possible for workers to efficiently obtain work instructions through voice input.

[0483] A "speech recognition device" is a device that converts speech into digital data and is used to obtain specific commands or information.

[0484] "Voice input" refers to a method of providing voice data to a system using a microphone or other voice receiving device.

[0485] "Text data" refers to digital information that represents audio or other forms of information as written characters.

[0486] A "natural language processing unit" is a device that implements technology to understand and process the language that humans use in everyday life.

[0487] A "keyword" is a word or phrase that is important when searching for or classifying specific information within text data.

[0488] An "information collection" is a database or data store that brings together multiple related documents and data.

[0489] "Information processing means" refers to a device or method for analyzing, retrieving, transforming, or manipulating digital data.

[0490] A "display device" is an electronic device used to present information visually.

[0491] "Communication means" refers to a protocol or device used to send and receive data.

[0492] An "access rights setting mechanism" is a system or process for managing and setting access rights to specific information or resources.

[0493] A "generation algorithm" is a mathematical formula or procedure for generating an output that satisfies specific conditions.

[0494] "User" refers to an individual or organization that uses the system or device.

[0495] This invention relates to a system that allows workers to efficiently obtain work instructions using voice input. In this system, the user provides voice to a terminal, which is then converted into digital text using a speech recognition device. This process is carried out using a common speech recognition technology (e.g., a speech recognition API) integrated into the terminal.

[0496] The server receives text data sent from the terminal and analyzes it using natural language processing techniques. Software such as Python's NLTK and spaCy are often used for this purpose. Through this process, the server extracts necessary keywords from the text.

[0497] Based on the extracted keywords, the server searches the information database and selects the most relevant instruction manuals. A generation algorithm (e.g., a machine learning model) is used to match the specified keywords with the content of the instruction manuals, thereby providing the most relevant information.

[0498] The terminal retrieves search results sent from the server and presents them visually to the user. The user can review the presented results and select the necessary instruction manual. This selection information is sent back to the server, which sets access rights and makes the instruction manual available for download or viewing.

[0499] This system allows users to quickly obtain necessary documents and improve work efficiency. For example, by inputting work instructions related to "daily inspections" via voice, relevant manuals and checklists can be instantly retrieved.

[0500] Examples of prompt messages include the following:

[0501] When given a work instruction to perform a "daily inspection," what kind of manuals or checklists are necessary?

[0502] In this way, users can quickly obtain information from voice input, effectively supporting their actual work.

[0503] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0504] Step 1:

[0505] The user inputs work instructions by voice. When the user says into the microphone, "I want to obtain the daily inspection manual," the voice data is entered into the terminal. This voice data forms the basis for subsequent processing.

[0506] Step 2:

[0507] The terminal uses a speech recognition device to convert voice data into text data. The input voice data is analyzed by a speech recognition API and extracted as text data. For example, the voice saying "I want to obtain the daily inspection manual" is converted into text.

[0508] Step 3:

[0509] The terminal sends the converted text data to the server. This data is securely transmitted to the server using the HTTP protocol, and the server receives this text.

[0510] Step 4:

[0511] The server extracts keywords from the received text data using a natural language processing unit. The natural language processing library identifies keywords (e.g., "daily check," "manual") from the text data. This process clarifies important concepts.

[0512] Step 5:

[0513] The server searches the information database based on the extracted keywords and selects relevant instruction manuals. Using a generative AI model, it matches the specified keywords with data in the digital library and selects relevant documents. For example, manuals related to daily inspections may be obtained as a result.

[0514] Step 6:

[0515] The server sends the search results to the terminal. The information on the selected instruction manuals is sent to the terminal using the HTTP protocol and is ready to be provided to the user.

[0516] Step 7:

[0517] The terminal visually displays the search results received from the server. Users can review the results displayed on the screen and select the necessary instructional materials.

[0518] Step 8:

[0519] The user selects the required instruction manual. The user's selection is recorded by the terminal and resent to the server.

[0520] Step 9:

[0521] The server provides the selected instruction manual in an available format. Based on the selection information, access rights are set, and a link to download or view the instruction manual is generated. The user can then obtain the necessary materials through this link.

[0522] (Application Example 1)

[0523] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0524] There is a problem in that workers have difficulty quickly and efficiently obtaining the necessary instruction manuals when performing tasks on-site. Furthermore, the need to use their hands to refer to the manuals during work reduces work efficiency.

[0525] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0526] In this invention, the server includes means for receiving work instructions via voice input using a voice recognition device, means for converting the voice input into text data, and means for using a natural language processing device to extract keywords from the text data. This makes it possible to quickly obtain necessary instruction manuals via voice instructions and improve work efficiency through visual presentation using visual aids.

[0527] A "speech recognition device" is a device that processes speech as a digital signal and converts it into language data.

[0528] "Voice input" refers to a method of inputting voice commands spoken by the user into the system.

[0529] "Text data" refers to digital character information converted by speech recognition.

[0530] A "natural language processing device" is a device equipped with technology to analyze meaning from text data and extract keywords.

[0531] A "storage medium" is a memory device that records and holds data.

[0532] "Instructional materials" refer to manuals and documents that provide work instructions and procedures.

[0533] An "information processing system" is a mechanism that has the function of processing data and generating and providing necessary information.

[0534] A "display device" refers to a screen or display that visually presents digital information.

[0535] "Visual assistance equipment" refers to devices that visualize information and provide support to help users perform tasks efficiently.

[0536] A "generating algorithm" is a series of steps taken to derive the desired result through the processing and analysis of information.

[0537] A "communication device" is a device that possesses the technology to send and receive data.

[0538] This invention is a system for improving work efficiency by enabling workers to quickly obtain necessary work instructions. The user provides specific work instructions via voice input using a voice recognition device. These voice instructions are converted into digital text data by a terminal. The generated text data is received by a server and analyzed using a device employing natural language processing technology. During the analysis, relevant keywords are extracted, and based on these, appropriate guidance materials are searched from storage media on the server. The server evaluates the suitability of the guidance materials using a generated algorithm. Once appropriate materials are selected, the relevant information is transmitted to the terminal. The terminal then displays the guidance materials on a display device in the user's possession.

[0539] Furthermore, this system integrates with visual assistance devices, allowing users to acquire information visually. This improves work efficiency by enabling users to refer to guide materials without using their hands during work. For example, when assembling a new part in a factory, a voice command such as "Tell me the assembly procedure for part A" will display the relevant work instructions on the glasses display. An example of a prompt message is: "Text to send to the server when the user says 'Tell me the assembly procedure for part A': 'Assembly procedure for part A'". Based on this prompt message, the system can provide guide materials that are appropriate to the workflow in a timely manner.

[0540] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0541] Step 1:

[0542] The user voice-inputs work instructions into a voice-recognition device. This input is received as an audio signal by the terminal. The terminal receives this audio signal and converts it into digital text data using its speech recognition engine. The input is an audio signal, and the output is text data. Through this process, the user's intended message is organized into textual information.

[0543] Step 2:

[0544] The terminal sends the generated text data to the server. The server receives this text data and uses a natural language processing device to extract keywords. The input is text data, and the output is a list of extracted keywords. This data processing gives the server the clues necessary to identify relevant information.

[0545] Step 3:

[0546] The server searches for relevant informational materials from the storage medium based on the extracted keywords. The input is a list of keywords, and the output is the relevant informational materials. This search process evaluates the relevance between the keywords and the content of the informational materials to select the most relevant materials.

[0547] Step 4:

[0548] The server sends the selected informational materials to the terminal. The terminal displays the received information visually on the user's visual aids. The input is the informational materials, and the output is the visual information displayed on the equipment. This operation allows the user to check the necessary information in real time while working.

[0549] Step 5:

[0550] The user checks the guidance materials displayed on the visual aid and proceeds with the work. Because the user can efficiently refer to the instructions without taking their hands off the device, they can continue performing highly accurate work without interrupting the workflow. The input is the information displayed on the visual device, and the output is the user's work actions. This operation enables a fluid workflow.

[0551] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0552] This invention relates to a system that incorporates an emotion engine to recognize user emotions and provide work instructions in a more personalized form. The following is a specific description of how this system is implemented.

[0553] The user provides voice input regarding their task to the terminal. This voice data is collected by the terminal's voice recognition device and converted into digital text data. Simultaneously, the terminal is equipped with an emotion engine that analyzes the tone, tempo, and pitch of the user's voice to identify their emotions.

[0554] The server receives text data sent from the terminal and extracts keywords from it. It also receives sentiment data analyzed by the sentiment engine. Combining the extracted keywords with the user's sentiment information, the server searches the database for the most suitable instructional guide.

[0555] The search process utilizes a generation algorithm to select the most suitable guide, taking into account both keywords and emotions. For example, if the emotion engine detects that the user is stressed, it adjusts the system to prioritize providing simple and easy-to-understand guides.

[0556] The terminal visually displays a list of instruction manuals sent from the server to the user. The content of the instruction manuals presented is tailored to the user's emotions, thus more effectively meeting their needs. The user selects an instruction manual that aligns with their emotions and downloads or views it on screen as needed.

[0557] This system is designed to improve work efficiency and reduce worker stress by providing content tailored to the user's psychological state. For example, users who feel anxious about their work are provided with simplified instructions and encouraging messages, allowing them to continue working with peace of mind.

[0558] The following describes the processing flow.

[0559] Step 1:

[0560] The user inputs the necessary instructions for the task by voice into the terminal. This voice is instantly converted into text data by the terminal's voice recognition device. Simultaneously, the terminal's built-in emotion engine analyzes emotional information from the user's voice. Characteristics such as pitch, speed, and tone of the voice are analyzed to determine what the user is feeling.

[0561] Step 2:

[0562] The device sends the converted text data and analyzed sentiment data to the server. The server then begins the process of extracting necessary keywords from the received text data. Using natural language processing, it efficiently finds keywords from the user's instructions.

[0563] Step 3:

[0564] The server searches the database for instructional materials based on extracted keywords and sentiment data. The generation algorithm then lists the most suitable options, taking sentiment data into consideration. For example, if a user is feeling anxious, instructional materials that provide reassurance can be prioritized.

[0565] Step 4:

[0566] A list of candidate instruction manuals provided by the server is sent to the terminal. The terminal visually displays this list to the user. The displayed instruction manuals are filtered according to the user's emotional state, making them highly likely to appropriately meet the user's needs.

[0567] Step 5:

[0568] The user selects the most suitable instruction manual from those presented. The selection is made directly on the device using actions such as tapping or clicking. Based on the selection, the device sends feedback to the server.

[0569] Step 6:

[0570] The server makes the user's selected instruction manual immediately available and generates a download link as needed. The terminal presents this information to the user, preparing them to download or view the instruction manual. Emotionally conscious instruction manuals enable users to perform their tasks effectively.

[0571] (Example 2)

[0572] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0573] Existing work instruction systems fail to provide individualized instruction materials that take into account the emotional state of users, sometimes resulting in the provision of information that is not most effective for the user. Furthermore, because work instruction is general in nature, it is difficult for users to tailor the instruction to their own specific situation. This can lead to decreased work efficiency and increased stress.

[0574] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0575] In this invention, the server includes means for receiving voice input via an acoustic processing device, means for converting voice data into text information, and means for using a natural language processing device that extracts keywords from the text information and utilizes information for sentiment analysis. This makes it possible to efficiently provide individualized instructional materials that take into account the emotional state of the user.

[0576] An "acoustic processing device" is a device that receives audio input and converts it into a format suitable for analyzing or processing the data.

[0577] "Speech conversion means" refers to a process or device that analyzes speech data and converts it into digital text format.

[0578] A "natural language processing device" is a technology or device that extracts keywords from text data and generates information for sentiment analysis.

[0579] "Emotional analysis information" refers to emotional data obtained from voice and text data, and is used to identify the user's emotional state.

[0580] A "knowledge base" is a database or information source that stores relevant teaching materials and is used to provide the most relevant information through searching.

[0581] "Information retrieval means" refers to a process or apparatus for retrieving specific data or information and providing results tailored to the user.

[0582] A "generative algorithm" is a set of computational steps to create, select, or evaluate the optimal output based on input data.

[0583] "Communication means" refers to the process or technology used to distribute or download instructional materials.

[0584] An "operator" refers to a person or entity that uses this system to receive work instructions.

[0585] This invention is a system that provides user-specific instructional materials based on voice input. The following hardware and software are used to implement this invention.

[0586] The terminal receives work instructions from the user via voice. The terminal is equipped with a voice recognition device that converts voice data into digital text. For example, a common voice recognition API is used for voice conversion.

[0587] Simultaneously, the device runs emotion analysis software. This software analyzes the tone and pitch of the voice to identify the user's emotional state. It then generates emotion analysis information and transmits it to the server via the communication path.

[0588] The server receives text information and sentiment analysis information sent from the terminal. The server uses a text analysis engine to extract keywords from the text data. This process may utilize, for example, a natural language processing library. The server also takes sentiment analysis information into account when searching for the most suitable teaching materials from a knowledge base. Generative AI models are used in the search algorithm.

[0589] For example, if a user inputs a voice command such as "I want to finish this task quickly," the server combines keywords like "quickly" with emotion analysis information such as "anxiety" and processes it. As a result, the server selects simplified instructional materials that are appropriate for the user's situation.

[0590] For example, it's possible to input prompts such as, "Generate a work instruction manual suitable for a user who is feeling anxious." This enables support tailored to the user's emotional state.

[0591] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0592] Step 1:

[0593] The user inputs work instructions by voice into the terminal.

[0594] Input: User's voice instructions

[0595] Output: Unallocated audio data

[0596] Step 2:

[0597] The terminal uses a speech recognition device to convert the user's voice data into digital text. This process utilizes speech recognition software.

[0598] Input: User's voice data

[0599] Output: Digital text data

[0600] Step 3:

[0601] The device acquires voice data and simultaneously runs emotion analysis software, identifying the user's emotions by analyzing the tone and pitch of the voice.

[0602] Input: User's voice data

[0603] Output: Sentiment analysis information

[0604] Step 4:

[0605] The device sends text data and sentiment analysis information to the server.

[0606] Input: Digital text data and sentiment analysis information

[0607] Output: Sending data to the server

[0608] Step 5:

[0609] The server extracts keywords from the received text data using a natural language processing library.

[0610] Input: Digital text data

[0611] Output: Keyword set

[0612] Step 6:

[0613] The server searches for the most suitable teaching materials from its knowledge base based on the extracted keywords and sentiment analysis information. A generative AI model is used in this process.

[0614] Input: Keyword set and sentiment analysis information

[0615] Output: Selection of optimal teaching materials

[0616] Step 7:

[0617] The server sends the retrieved instructional materials to the terminal.

[0618] Input: Optimal teaching materials

[0619] Output: Sending data to the terminal

[0620] Step 8:

[0621] The terminal visually presents the received instructional materials to the user. Specifically, it displays them on the screen through the user interface.

[0622] Input: Instructional materials sent from the server

[0623] Output: Visual presentation to the user

[0624] (Application Example 2)

[0625] Next, we will explain Application Example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0626] When workers receive work instructions, if those instructions are uniform, efficiency may decrease depending on the worker's psychological state. In particular, if the instructions given to workers experiencing stress or anxiety are inappropriate, work efficiency and safety may decrease.

[0627] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0628] In this invention, the server includes means for analyzing emotions from voice data, means for converting voice input into text data, and means for searching a database for the most suitable instruction manual based on keywords and emotion data. This makes it possible to provide more personalized instruction manuals that are tailored to the emotional state of the worker.

[0629] A "speech recognition device" is a device that receives speech as a digital signal, analyzes its content, and converts it into text data.

[0630] A "natural language processing unit" is a device that processes text data to extract important words and phrases.

[0631] "Information processing means" refers to means that provide a method for retrieving relevant information based on given data and generating or suggesting appropriate results.

[0632] "Emotional data" refers to data about the speaker's psychological or emotional state, obtained by analyzing the characteristics of their voice.

[0633] A "generative algorithm" is a computational method for generating highly relevant results or candidates based on specific input data.

[0634] A "display device" is a device used to visually present processed information to a user.

[0635] A "communication device" is a device that transmits and receives data, enabling data exchange with the outside world.

[0636] The system of this invention is used in a work environment in which an operator wears smart glasses and receives voice instructions. The system mainly consists of a terminal, a server, and smart glasses.

[0637] The device features a function that converts user speech into text data using the Google Cloud Speech-to-Text API. It also uses the Microsoft Azure Emotion API to analyze user emotions based on speech tone, speed, and other factors. Both the emotion data and the text data are then transferred to a server.

[0638] The server uses a TensorFlow-based generation algorithm to search the database for the most suitable instruction manual based on the received sentiment and text data. The selected instruction manual takes the user's psychological state into consideration and is appropriately customized to maximize work efficiency.

[0639] Users view instructions sent from the server on the display of their smart glasses. For example, users who are feeling stressed about a task can be provided with simple and visually easy-to-understand instructions, allowing them to proceed with the task with confidence.

[0640] For example, when a user says, "This task is difficult," stress can be detected from things like mispronunciation or rising intonation. The system then displays a step-by-step, visually guided procedure manual, color-coded accordingly. In this way, it becomes possible to provide task guidance tailored to the user's emotional state.

[0641] Examples of prompt statements include:

[0642] The message is: "We detected audio from a user saying 'the task is difficult' and indicating signs of stress. Please suggest a simplified guide."

[0643] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0644] Step 1:

[0645] The user provides voice input through smart glasses. The input voice data is collected by the device. The collected voice data is then converted into text data using the Google Cloud Speech-to-Text API. This process converts the voice data into a parseable text format.

[0646] Step 2:

[0647] The device simultaneously uses Microsoft Azure's Emotion API to analyze the user's emotions from the tone and speed of their voice. At this stage, it extracts emotional data from the input voice data and interprets the results of the emotion analysis. The output is data indicating the user's emotional state.

[0648] Step 3:

[0649] These text and sentiment data are sent from the terminal to the server. The server extracts keywords from the received text data and then integrates them with the received sentiment data. This data integration creates the dataset necessary for searching.

[0650] Step 4:

[0651] The server executes a generative algorithm using TensorFlow. Here, it searches the database for the most relevant instructional guides based on the integrated dataset. The generative algorithm evaluates both keywords and sentiment data to select the appropriate guide.

[0652] Step 5:

[0653] The selected instruction manuals are sent to the smart glasses in an optimized format to enhance the user's work efficiency. The user reviews the instruction manuals on the smart glasses' display. Here, the system enables a personalized learning experience tailored to the user's psychological state.

[0654] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0655] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0656] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0657] [Fourth Embodiment]

[0658] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0659] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0660] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0661] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0662] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0663] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0664] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0665] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0666] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0667] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0668] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0669] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0670] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0671] This invention relates to a system for workers to automatically obtain work instructions using voice input. Therefore, embodiments designed to ensure that each device and process functions effectively will be described.

[0672] At the start of a task, the user uses the terminal's voice input function to speak the necessary work instructions into the microphone. The user's voice is converted into digital text data via the terminal's voice recognition device.

[0673] The server receives text data sent from the terminal and uses a natural language processing unit to extract relevant keywords from that text. These keywords serve as important clues in the subsequent database search process.

[0674] Based on the extracted keywords, the server searches the database and selects the most relevant instruction manuals from those held by trading partners and companies. In this context, a generation algorithm is used to provide the optimal choice by matching the keywords with the content of the instruction manuals.

[0675] Subsequently, the terminal retrieves the search results sent from the server and displays the results visually to the user. The user can review these results and select the necessary instructional materials. This selection information is then sent back from the terminal to the server, which provides the instructional materials in an accessible format. Access rights and download links are generated to enable the user to effectively utilize the materials.

[0676] This system allows users to quickly receive the necessary instructions for their work and receive support to make appropriate decisions even in busy work environments. For example, when a work instruction for "daily inspection" is issued, the relevant manuals and checklists for daily inspection can be immediately obtained. In this way, workers can obtain the necessary information in a short time and perform their tasks with high accuracy.

[0677] The following describes the processing flow.

[0678] Step 1:

[0679] The user uses the terminal's voice input function to input work-related requests by voice. The terminal receives this voice in real time and converts it into text data using its built-in speech recognition function. The converted text data is then passed on to the next process.

[0680] Step 2:

[0681] The terminal sends the text data generated by speech recognition to the server. The server utilizes a natural language processing engine to extract keywords relevant to the task from the text data. In this step, important words and phrases are identified and specified through text analysis.

[0682] Step 3:

[0683] The server searches the database for instructional materials based on the extracted keywords. An AI-powered generation algorithm supports this process, matching the keywords with the content of documents in the database to list the most relevant instructional materials.

[0684] Step 4:

[0685] The server sends a list of searched instruction manuals to the terminal. The terminal presents the received list of instruction manuals to the user through a visual interface. It is important here that the user can view previews and summaries of the instruction manuals.

[0686] Step 5:

[0687] The user selects the necessary instruction manual from the displayed options. The terminal sends this selection information to the server, which then prepares the selected manual, sets access permissions, and makes it available for download.

[0688] Step 6:

[0689] The terminal displays instruction manual access information received from the server, allowing the user to download or view the manual. Based on this information, the user can continue with the necessary tasks, enabling efficient work execution.

[0690] (Example 1)

[0691] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0692] Obtaining necessary work instructions quickly on-site is crucial in many work environments. However, conventional systems have the problem of making it difficult for workers to quickly identify and access the specific instructions they need. To address this challenge, a system is needed that improves usability and allows for efficient information acquisition using voice input.

[0693] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0694] In this invention, the server includes means for receiving work instructions via voice input using a voice recognition device, means for converting the voice input into text data, and means for using a natural language processing device to extract keywords from the text data. This makes it possible for workers to efficiently obtain work instructions through voice input.

[0695] A "speech recognition device" is a device that converts speech into digital data and is used to obtain specific commands or information.

[0696] "Voice input" refers to a method of providing voice data to a system using a microphone or other voice receiving device.

[0697] "Text data" refers to digital information that represents audio or other forms of information as written characters.

[0698] A "natural language processing unit" is a device that implements technology to understand and process the language that humans use in everyday life.

[0699] A "keyword" is a word or phrase that is important when searching for or classifying specific information within text data.

[0700] An "information collection" is a database or data store that brings together multiple related documents and data.

[0701] "Information processing means" refers to a device or method for analyzing, retrieving, transforming, or manipulating digital data.

[0702] A "display device" is an electronic device used to present information visually.

[0703] "Communication means" refers to a protocol or device used to send and receive data.

[0704] An "access rights setting mechanism" is a system or process for managing and setting access rights to specific information or resources.

[0705] A "generation algorithm" is a mathematical formula or procedure for generating an output that satisfies specific conditions.

[0706] "User" refers to an individual or organization that uses the system or device.

[0707] This invention relates to a system that allows workers to efficiently obtain work instructions using voice input. In this system, the user provides voice to a terminal, which is then converted into digital text using a speech recognition device. This process is carried out using a common speech recognition technology (e.g., a speech recognition API) integrated into the terminal.

[0708] The server receives text data sent from the terminal and analyzes it using natural language processing techniques. Software such as Python's NLTK and spaCy are often used for this purpose. Through this process, the server extracts necessary keywords from the text.

[0709] Based on the extracted keywords, the server searches the information database and selects the most relevant instruction manuals. A generation algorithm (e.g., a machine learning model) is used to match the specified keywords with the content of the instruction manuals, thereby providing the most relevant information.

[0710] The terminal retrieves search results sent from the server and presents them visually to the user. The user can review the presented results and select the necessary instruction manual. This selection information is sent back to the server, which sets access rights and makes the instruction manual available for download or viewing.

[0711] This system allows users to quickly obtain necessary documents and improve work efficiency. For example, by inputting work instructions related to "daily inspections" via voice, relevant manuals and checklists can be instantly retrieved.

[0712] Examples of prompt messages include the following:

[0713] When given a work instruction to perform a "daily inspection," what kind of manuals or checklists are necessary?

[0714] In this way, users can quickly obtain information from voice input, effectively supporting their actual work.

[0715] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0716] Step 1:

[0717] The user inputs work instructions by voice. When the user says into the microphone, "I want to obtain the daily inspection manual," the voice data is entered into the terminal. This voice data forms the basis for subsequent processing.

[0718] Step 2:

[0719] The terminal uses a speech recognition device to convert voice data into text data. The input voice data is analyzed by a speech recognition API and extracted as text data. For example, the voice saying "I want to obtain the daily inspection manual" is converted into text.

[0720] Step 3:

[0721] The terminal sends the converted text data to the server. This data is securely transmitted to the server using the HTTP protocol, and the server receives this text.

[0722] Step 4:

[0723] The server extracts keywords from the received text data using a natural language processing unit. The natural language processing library identifies keywords (e.g., "daily check," "manual") from the text data. This process clarifies important concepts.

[0724] Step 5:

[0725] The server searches the information database based on the extracted keywords and selects relevant instruction manuals. Using a generative AI model, it matches the specified keywords with data in the digital library and selects relevant documents. For example, manuals related to daily inspections may be obtained as a result.

[0726] Step 6:

[0727] The server sends the search results to the terminal. The information on the selected instruction manuals is sent to the terminal using the HTTP protocol and is ready to be provided to the user.

[0728] Step 7:

[0729] The terminal visually displays the search results received from the server. Users can review the results displayed on the screen and select the necessary instructional materials.

[0730] Step 8:

[0731] The user selects the required instruction manual. The user's selection is recorded by the terminal and resent to the server.

[0732] Step 9:

[0733] The server provides the selected instruction manual in an available format. Based on the selection information, access rights are set, and a link to download or view the instruction manual is generated. The user can then obtain the necessary materials through this link.

[0734] (Application Example 1)

[0735] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0736] There is a problem in that workers have difficulty quickly and efficiently obtaining the necessary instruction manuals when performing tasks on-site. Furthermore, the need to use their hands to refer to the manuals during work reduces work efficiency.

[0737] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0738] In this invention, the server includes means for receiving work instructions via voice input using a voice recognition device, means for converting the voice input into text data, and means for using a natural language processing device to extract keywords from the text data. This makes it possible to quickly obtain necessary instruction manuals via voice instructions and improve work efficiency through visual presentation using visual aids.

[0739] A "speech recognition device" is a device that processes speech as a digital signal and converts it into language data.

[0740] "Voice input" refers to a method of inputting voice commands spoken by the user into the system.

[0741] "Text data" refers to digital character information converted by speech recognition.

[0742] A "natural language processing device" is a device equipped with technology to analyze meaning from text data and extract keywords.

[0743] A "storage medium" is a memory device that records and holds data.

[0744] "Instructional materials" refer to manuals and documents that provide work instructions and procedures.

[0745] An "information processing system" is a mechanism that has the function of processing data and generating and providing necessary information.

[0746] A "display device" refers to a screen or display that visually presents digital information.

[0747] "Visual assistance equipment" refers to devices that visualize information and provide support to help users perform tasks efficiently.

[0748] A "generating algorithm" is a series of steps taken to derive the desired result through the processing and analysis of information.

[0749] A "communication device" is a device that possesses the technology to send and receive data.

[0750] This invention is a system for improving work efficiency by enabling workers to quickly obtain necessary work instructions. The user provides specific work instructions via voice input using a voice recognition device. These voice instructions are converted into digital text data by a terminal. The generated text data is received by a server and analyzed using a device employing natural language processing technology. During the analysis, relevant keywords are extracted, and based on these, appropriate guidance materials are searched from storage media on the server. The server evaluates the suitability of the guidance materials using a generated algorithm. Once appropriate materials are selected, the relevant information is transmitted to the terminal. The terminal then displays the guidance materials on a display device in the user's possession.

[0751] Furthermore, this system integrates with visual assistance devices, allowing users to acquire information visually. This improves work efficiency by enabling users to refer to guide materials without using their hands during work. For example, when assembling a new part in a factory, a voice command such as "Tell me the assembly procedure for part A" will display the relevant work instructions on the glasses display. An example of a prompt message is: "Text to send to the server when the user says 'Tell me the assembly procedure for part A': 'Assembly procedure for part A'". Based on this prompt message, the system can provide guide materials that are appropriate to the workflow in a timely manner.

[0752] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0753] Step 1:

[0754] The user voice-inputs work instructions into a voice-recognition device. This input is received as an audio signal by the terminal. The terminal receives this audio signal and converts it into digital text data using its speech recognition engine. The input is an audio signal, and the output is text data. Through this process, the user's intended message is organized into textual information.

[0755] Step 2:

[0756] The terminal sends the generated text data to the server. The server receives this text data and uses a natural language processing device to extract keywords. The input is text data, and the output is a list of extracted keywords. This data processing gives the server the clues necessary to identify relevant information.

[0757] Step 3:

[0758] The server searches for relevant informational materials from the storage medium based on the extracted keywords. The input is a list of keywords, and the output is the relevant informational materials. This search process evaluates the relevance between the keywords and the content of the informational materials to select the most relevant materials.

[0759] Step 4:

[0760] The server sends the selected informational materials to the terminal. The terminal displays the received information visually on the user's visual aids. The input is the informational materials, and the output is the visual information displayed on the equipment. This operation allows the user to check the necessary information in real time while working.

[0761] Step 5:

[0762] The user checks the guidance materials displayed on the visual aid and proceeds with the work. Because the user can efficiently refer to the instructions without taking their hands off the device, they can continue performing highly accurate work without interrupting the workflow. The input is the information displayed on the visual device, and the output is the user's work actions. This operation enables a fluid workflow.

[0763] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0764] This invention relates to a system that incorporates an emotion engine to recognize user emotions and provide work instructions in a more personalized form. The following is a specific description of how this system is implemented.

[0765] The user provides voice input regarding their task to the terminal. This voice data is collected by the terminal's voice recognition device and converted into digital text data. Simultaneously, the terminal is equipped with an emotion engine that analyzes the tone, tempo, and pitch of the user's voice to identify their emotions.

[0766] The server receives text data sent from the terminal and extracts keywords from it. It also receives sentiment data analyzed by the sentiment engine. Combining the extracted keywords with the user's sentiment information, the server searches the database for the most suitable instructional guide.

[0767] The search process utilizes a generation algorithm to select the most suitable guide, taking into account both keywords and emotions. For example, if the emotion engine detects that the user is stressed, it adjusts the system to prioritize providing simple and easy-to-understand guides.

[0768] The terminal visually displays a list of instruction manuals sent from the server to the user. The content of the instruction manuals presented is tailored to the user's emotions, thus more effectively meeting their needs. The user selects an instruction manual that aligns with their emotions and downloads or views it on screen as needed.

[0769] This system is designed to improve work efficiency and reduce worker stress by providing content tailored to the user's psychological state. For example, users who feel anxious about their work are provided with simplified instructions and encouraging messages, allowing them to continue working with peace of mind.

[0770] The following describes the processing flow.

[0771] Step 1:

[0772] The user inputs the necessary instructions for the task by voice into the terminal. This voice is instantly converted into text data by the terminal's voice recognition device. Simultaneously, the terminal's built-in emotion engine analyzes emotional information from the user's voice. Characteristics such as pitch, speed, and tone of the voice are analyzed to determine what the user is feeling.

[0773] Step 2:

[0774] The device sends the converted text data and analyzed sentiment data to the server. The server then begins the process of extracting necessary keywords from the received text data. Using natural language processing, it efficiently finds keywords from the user's instructions.

[0775] Step 3:

[0776] The server searches the database for instructional materials based on extracted keywords and sentiment data. The generation algorithm then lists the most suitable options, taking sentiment data into consideration. For example, if a user is feeling anxious, instructional materials that provide reassurance can be prioritized.

[0777] Step 4:

[0778] A list of candidate instruction manuals provided by the server is sent to the terminal. The terminal visually displays this list to the user. The displayed instruction manuals are filtered according to the user's emotional state, making them highly likely to appropriately meet the user's needs.

[0779] Step 5:

[0780] The user selects the most suitable instruction manual from those presented. The selection is made directly on the device using actions such as tapping or clicking. Based on the selection, the device sends feedback to the server.

[0781] Step 6:

[0782] The server makes the user's selected instruction manual immediately available and generates a download link as needed. The terminal presents this information to the user, preparing them to download or view the instruction manual. Emotionally conscious instruction manuals enable users to perform their tasks effectively.

[0783] (Example 2)

[0784] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0785] Existing work instruction systems fail to provide individualized instruction materials that take into account the emotional state of users, sometimes resulting in the provision of information that is not most effective for the user. Furthermore, because work instruction is general in nature, it is difficult for users to tailor the instruction to their own specific situation. This can lead to decreased work efficiency and increased stress.

[0786] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0787] In this invention, the server includes means for receiving voice input via an acoustic processing device, means for converting voice data into text information, and means for using a natural language processing device that extracts keywords from the text information and utilizes information for sentiment analysis. This makes it possible to efficiently provide individualized instructional materials that take into account the emotional state of the user.

[0788] An "acoustic processing device" is a device that receives audio input and converts it into a format suitable for analyzing or processing the data.

[0789] "Speech conversion means" refers to a process or device that analyzes speech data and converts it into digital text format.

[0790] A "natural language processing device" is a technology or device that extracts keywords from text data and generates information for sentiment analysis.

[0791] "Emotional analysis information" refers to emotional data obtained from voice and text data, and is used to identify the user's emotional state.

[0792] A "knowledge base" is a database or information source that stores relevant teaching materials and is used to provide the most relevant information through searching.

[0793] "Information retrieval means" refers to a process or apparatus for retrieving specific data or information and providing results tailored to the user.

[0794] A "generative algorithm" is a set of computational steps to create, select, or evaluate the optimal output based on input data.

[0795] "Communication means" refers to the process or technology used to distribute or download instructional materials.

[0796] An "operator" refers to a person or entity that uses this system to receive work instructions.

[0797] This invention is a system that provides user-specific instructional materials based on voice input. The following hardware and software are used to implement this invention.

[0798] The terminal receives work instructions from the user via voice. The terminal is equipped with a voice recognition device that converts voice data into digital text. For example, a common voice recognition API is used for voice conversion.

[0799] Simultaneously, the device runs emotion analysis software. This software analyzes the tone and pitch of the voice to identify the user's emotional state. It then generates emotion analysis information and transmits it to the server via the communication path.

[0800] The server receives text information and sentiment analysis information sent from the terminal. The server uses a text analysis engine to extract keywords from the text data. This process may utilize, for example, a natural language processing library. The server also takes sentiment analysis information into account when searching for the most suitable teaching materials from a knowledge base. Generative AI models are used in the search algorithm.

[0801] For example, if a user inputs a voice command such as "I want to finish this task quickly," the server combines keywords like "quickly" with emotion analysis information such as "anxiety" and processes it. As a result, the server selects simplified instructional materials that are appropriate for the user's situation.

[0802] For example, it's possible to input prompts such as, "Generate a work instruction manual suitable for a user who is feeling anxious." This enables support tailored to the user's emotional state.

[0803] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0804] Step 1:

[0805] The user inputs work instructions by voice into the terminal.

[0806] Input: User's voice instructions

[0807] Output: Unallocated audio data

[0808] Step 2:

[0809] The terminal uses a speech recognition device to convert the user's voice data into digital text. This process utilizes speech recognition software.

[0810] Input: User's voice data

[0811] Output: Digital text data

[0812] Step 3:

[0813] The device acquires voice data and simultaneously runs emotion analysis software, identifying the user's emotions by analyzing the tone and pitch of the voice.

[0814] Input: User's voice data

[0815] Output: Sentiment analysis information

[0816] Step 4:

[0817] The device sends text data and sentiment analysis information to the server.

[0818] Input: Digital text data and sentiment analysis information

[0819] Output: Sending data to the server

[0820] Step 5:

[0821] The server extracts keywords from the received text data using a natural language processing library.

[0822] Input: Digital text data

[0823] Output: Keyword set

[0824] Step 6:

[0825] The server searches for the most suitable teaching materials from its knowledge base based on the extracted keywords and sentiment analysis information. A generative AI model is used in this process.

[0826] Input: Keyword set and sentiment analysis information

[0827] Output: Selection of optimal teaching materials

[0828] Step 7:

[0829] The server sends the retrieved instructional materials to the terminal.

[0830] Input: Optimal teaching materials

[0831] Output: Sending data to the terminal

[0832] Step 8:

[0833] The terminal visually presents the received instructional materials to the user. Specifically, it displays them on the screen through the user interface.

[0834] Input: Instructional materials sent from the server

[0835] Output: Visual presentation to the user

[0836] (Application Example 2)

[0837] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0838] When workers receive work instructions, if those instructions are uniform, efficiency may decrease depending on the worker's psychological state. In particular, if the instructions given to workers experiencing stress or anxiety are inappropriate, work efficiency and safety may decrease.

[0839] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0840] In this invention, the server includes means for analyzing emotions from voice data, means for converting voice input into text data, and means for searching a database for the most suitable instruction manual based on keywords and emotion data. This makes it possible to provide more personalized instruction manuals that are tailored to the emotional state of the worker.

[0841] A "speech recognition device" is a device that receives speech as a digital signal, analyzes its content, and converts it into text data.

[0842] A "natural language processing unit" is a device that processes text data to extract important words and phrases.

[0843] "Information processing means" refers to means that provide a method for retrieving relevant information based on given data and generating or suggesting appropriate results.

[0844] "Emotional data" refers to data about the speaker's psychological or emotional state, obtained by analyzing the characteristics of their voice.

[0845] A "generative algorithm" is a computational method for generating highly relevant results or candidates based on specific input data.

[0846] A "display device" is a device used to visually present processed information to a user.

[0847] A "communication device" is a device that transmits and receives data, enabling data exchange with the outside world.

[0848] The system of this invention is used in a work environment in which an operator wears smart glasses and receives voice instructions. The system mainly consists of a terminal, a server, and smart glasses.

[0849] The device features a function that converts user speech into text data using the Google Cloud Speech-to-Text API. It also uses the Microsoft Azure Emotion API to analyze user emotions based on speech tone, speed, and other factors. Both the emotion data and the text data are then transferred to a server.

[0850] The server uses a TensorFlow-based generation algorithm to search the database for the most suitable instruction manual based on the received sentiment and text data. The selected instruction manual takes the user's psychological state into consideration and is appropriately customized to maximize work efficiency.

[0851] Users view instructions sent from the server on the display of their smart glasses. For example, users who are feeling stressed about a task can be provided with simple and visually easy-to-understand instructions, allowing them to proceed with the task with confidence.

[0852] For example, when a user says, "This task is difficult," stress can be detected from things like mispronunciation or rising intonation. The system then displays a step-by-step, visually guided procedure manual, color-coded accordingly. In this way, it becomes possible to provide task guidance tailored to the user's emotional state.

[0853] Examples of prompt statements include:

[0854] The message is: "We detected audio from a user saying 'the task is difficult' and indicating signs of stress. Please suggest a simplified guide."

[0855] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0856] Step 1:

[0857] The user provides voice input through smart glasses. The input voice data is collected by the device. The collected voice data is then converted into text data using the Google Cloud Speech-to-Text API. This process converts the voice data into a parseable text format.

[0858] Step 2:

[0859] The device simultaneously uses Microsoft Azure's Emotion API to analyze the user's emotions from the tone and speed of their voice. At this stage, it extracts emotional data from the input voice data and interprets the results of the emotion analysis. The output is data indicating the user's emotional state.

[0860] Step 3:

[0861] These text and sentiment data are sent from the terminal to the server. The server extracts keywords from the received text data and then integrates them with the received sentiment data. This data integration creates the dataset necessary for searching.

[0862] Step 4:

[0863] The server executes a generative algorithm using TensorFlow. Here, it searches the database for the most relevant instructional guides based on the integrated dataset. The generative algorithm evaluates both keywords and sentiment data to select the appropriate guide.

[0864] Step 5:

[0865] The selected instruction manuals are sent to the smart glasses in an optimized format to enhance the user's work efficiency. The user reviews the instruction manuals on the smart glasses' display. Here, the system enables a personalized learning experience tailored to the user's psychological state.

[0866] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0867] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0868] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0869] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0870] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0871] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0872] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0873] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0874] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0875] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0876] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0877] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0878] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0879] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0880] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0881] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0882] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0883] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0884] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0885] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0886] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0887] The following is further disclosed regarding the embodiments described above.

[0888] (Claim 1)

[0889] A means of receiving work instructions via voice input using a voice recognition device,

[0890] A means of converting voice input into text data,

[0891] A method using a natural language processing device to extract keywords from text data,

[0892] An information processing means that searches for relevant instruction manuals from a database based on extracted keywords,

[0893] A means having a display device that presents search results to the user,

[0894] A system that includes this.

[0895] (Claim 2)

[0896] The system according to claim 1, characterized in that it has means for evaluating suitability using a generation algorithm to select candidates for relevant instructional materials for a keyword.

[0897] (Claim 3)

[0898] The system according to claim 1, characterized by having a communication device that distributes or makes downloadable instruction manuals based on user selection.

[0899] "Example 1"

[0900] (Claim 1)

[0901] A means of receiving work instructions via voice input using a voice recognition device,

[0902] A means of converting voice input into text data,

[0903] A method using a natural language processing device to extract keywords from text data,

[0904] An information processing means for searching instruction manuals from an information collection that stores related information based on extracted keywords,

[0905] A means having a display device that visually presents search results to the user,

[0906] A communication method for collecting and transmitting user selection information,

[0907] A means of setting access rights to provide instruction manuals in an accessible format,

[0908] A system that includes this.

[0909] (Claim 2)

[0910] The system according to claim 1, characterized in that it has means for matching candidates of relevant instructional materials for a keyword using a generation algorithm and evaluating their suitability.

[0911] (Claim 3)

[0912] The system according to claim 1, characterized by having means for making instruction manuals downloadable or viewable based on the user's selection.

[0913] "Application Example 1"

[0914] (Claim 1)

[0915] A means of receiving work instructions via voice input using a voice recognition device,

[0916] A means of converting voice input into text data,

[0917] A method using a natural language processing device to extract keywords from text data,

[0918] An information processing means that searches for relevant informational materials from storage media based on extracted keywords,

[0919] A means having a display device that presents search results to the user,

[0920] A means of supporting users to perform tasks efficiently through visual presentation, in conjunction with visual assistance equipment that has a display function for visually conveying information,

[0921] A system that includes this.

[0922] (Claim 2)

[0923] The system according to claim 1, further comprising means for evaluating suitability using a generating algorithm to select candidates for relevant guidance materials for a given keyword.

[0924] (Claim 3)

[0925] The system according to claim 1, comprising a communication device that enables the supply or acquisition of informational materials based on a user's selection.

[0926] "Example 2 of combining an emotion engine"

[0927] (Claim 1)

[0928] A means of receiving audio input via an acoustic processing device,

[0929] A speech conversion means for converting speech data into text information,

[0930] A method using a natural language processing device that extracts keywords from text information and utilizes the information for sentiment analysis,

[0931] An information retrieval method that searches for the most suitable teaching materials from a knowledge base based on extracted keywords and sentiment analysis information,

[0932] A display means for presenting the search results to the operator,

[0933] A system that includes this.

[0934] (Claim 2)

[0935] The system according to claim 1, characterized by having means for selecting relevant instructional materials based on keywords and sentiment analysis information, and for evaluating their optimality using a generation algorithm.

[0936] (Claim 3)

[0937] The system according to claim 1, characterized by having communication means that enable the transmission or acquisition of instructional materials based on the operator's selection.

[0938] "Application example 2 when combining with an emotional engine"

[0939] (Claim 1)

[0940] A means of receiving work instructions via voice input using a voice recognition device,

[0941] A means of converting voice input into text data,

[0942] A method using a natural language processing device to extract keywords from text data,

[0943] An information processing method that searches a database for relevant instructional materials based on extracted keywords and emotional data analyzed from audio,

[0944] A means for displaying instruction manuals optimized according to text data and sentiment data on a display device,

[0945] A system that includes this.

[0946] (Claim 2)

[0947] The system according to claim 1, further comprising means for evaluating suitability using a generation algorithm to select relevant instructional guide candidates based on keywords and sentiment data.

[0948] (Claim 3)

[0949] The system according to claim 1, comprising a communication device that selects and distributes or makes downloadable instruction manuals optimized according to the user's emotions. [Explanation of symbols]

[0950] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of receiving work instructions via voice input using a voice recognition device, A means of converting voice input into text data, A method using a natural language processing device to extract keywords from text data, An information processing means that searches for relevant instruction manuals from a database based on extracted keywords, A means having a display device that presents search results to the user, A system that includes this.

2. The system according to claim 1, characterized in that it has means for evaluating suitability using a generation algorithm to select candidates for relevant instructional materials for a keyword.

3. The system according to claim 1, characterized by having a communication device that distributes or makes downloadable instruction manuals based on user selection.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A