System
The system addresses the limitations of traditional manuals by using AR and audio guidance to intuitively assist users and offer emergency support.
Patent Information
- Application Number
- JP2024137133
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-16
- Publication Date
- 2026-02-27
AI Technical Summary
Current paper manuals and conventional text-based interactive AI manuals are difficult for users to understand intuitively, lack visual and audio support, and have limited emergency response capabilities.
A system that receives user instructions, acquires camera footage, analyzes it to identify appropriate manual information, and provides this information through AR display and audio guidance, while also generating answers to questions and notifying emergency situations.
Enables users to quickly and intuitively perform tasks with visual and audio support, and provides emergency assistance when needed.
Smart Images

Figure 2026034012000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Current paper manuals and conventional text-based interactive AI manuals make it difficult for users to quickly and intuitively obtain the information they need. Furthermore, these methods lack visual support and audio guidance, making them difficult for users to understand, especially complex tasks and operations. Furthermore, there are limited support methods available to quickly respond when users encounter difficulties. There is a need for a new manual system that can solve these issues and enable users to understand intuitively. [Means for solving the problem]
[0005] The present invention is a system that includes a means for receiving user instructions, a means for acquiring camera footage and transmitting it to a server, a means for analyzing the received camera footage and identifying appropriate manual information, a means for searching for the identified manual information and providing it to a terminal, and a means for providing information to the user on the terminal through AR display and audio guidance. This allows the user to intuitively understand information visually and audibly, enabling them to quickly perform the necessary operations and tasks. The system also includes a means for generating answers to user questions based on the analyzed camera footage, and a means for notifying the user of an emergency situation and contacting a manned operator if the user is in trouble, thereby strengthening the user support system.
[0006] "Means for receiving user instructions" refers to devices or software that receive input from the user, such as voice or touch operations, and transmit them to the system.
[0007] "Means for acquiring camera images and transmitting them to a server" refers to a device or program that uses a camera to capture the object or situation the user is looking at and transmits the images to a server in real time.
[0008] "Means for analyzing received camera footage and identifying appropriate manual information" refers to software or algorithms that allow the server to analyze camera footage, recognize structures and objects, and identify the most appropriate manual information based on that.
[0009] "Means for searching for identified manual information and providing it to the terminal" refers to the function or process for searching for related manual information from a database based on the analysis results and sending it to the terminal.
[0010] "Means for providing information to users through AR displays and voice guidance" refers to devices and software that allow a terminal to visually display information using augmented reality (AR) technology and also provide voice instructions using voice synthesis technology.
[0011] "Means for generating answers to user questions based on analyzed camera footage" refers to interactive AI or algorithms that automatically generate appropriate answers to user questions based on information obtained from camera footage.
[0012] "Means for notifying an emergency situation and contacting a live operator when a user is in trouble" refers to a system or process that allows a user to notify a server of their situation by operating an interface in an emergency, and then connect to a live operator in real time to receive support. [Brief explanation of the drawings]
[0013] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0014] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0015] First, the terms used in the following description will be explained.
[0016] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0017] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0018] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0019] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0021] [First embodiment]
[0022] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0023] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0024] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0025] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0026] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0028] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0029] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0030] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0031] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0032] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0033] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0034] The present invention is a system for receiving user instructions, analyzing camera footage based on those instructions, and providing appropriate manual information to the user. This system is composed of a server and a terminal, and aims to provide the user with visual and audio information that can be intuitively understood. Specific embodiments of this system are described below.
[0035] Receiving user instructions
[0036] The device receives instructions from the user through voice or touch operations. For example, if the user says, "Tell me how to connect to my smart TV," the device picks up the voice through a microphone and converts it into text using voice recognition technology. This textual instruction is then sent to the server.
[0037] Video acquisition and transmission
[0038] When a user points the back of a smart TV at the device's camera, the device captures the camera image in real time and sends it to the server. The device is equipped with a camera and a network connection to reliably capture the image.
[0039] Video analysis and information identification
[0040] The server receives the camera footage sent from the device and analyzes it using an image recognition model. The server identifies each port and cable connection in the footage, and based on the results, identifies the appropriate connection procedures and manual information. This information is retrieved from a database.
[0041] Providing manual information
[0042] The server sends the identified manual information to the device. The device receives this information and provides it to the user through AR display and voice guidance. For example, the device display may overlay the corresponding cable connection point on the back of the smart TV, and voice guidance may say, "Plug the HDMI cable into the third port from the left."
[0043] Response to additional questions
[0044] If the user asks a follow-up question, such as "Where is the power button?", the device sends this question to the server. The server uses conversational AI to analyze the question and generate an answer. For example, it might determine that "The round button on the bottom right is the power button" and send this information to the device. The device then presents this information in an AR display and voice guidance.
[0045] Emergency response
[0046] If a user encounters a problem, such as difficulty operating the device or making a mistake, they can press the emergency button on the device, which will then notify the server of the emergency situation. The server receives this information and connects to a live operator in real time. The emergency operator will monitor the user's current situation in real time and provide appropriate support.
[0047] Specific examples
[0048] Here is a concrete example: If a user purchases a new smart TV and does not know how to connect it, they can use this system to receive the following support:
[0049] 1. The user points the camera at the back of the smart TV and gives voice instructions from the device saying, "Tell me how to connect to the smart TV."
[0050] 2. The device analyzes the voice, converts the instructions into text, and sends it to the server, along with the camera footage.
[0051] 3. The server analyzes the camera footage, identifies each port on the smart TV, determines the appropriate connection procedure, and sends the manual information to the device.
[0052] 4. The device will use AR to display the appropriate cable connection location on the back of the TV and will provide voice guidance such as "Plug the HDMI cable into the third port from the left."
[0053] 5. While connected, if the user asks, "Where is the power button?", the device sends this question to the server, which generates an answer and sends it to the device.
[0054] 6. The device will display AR and provide voice guidance, saying, "The round button on the bottom right is the power button."
[0055] In this way, the present invention allows the user to quickly perform operations and tasks while receiving intuitive visual and audio instructions.
[0056] The processing flow will be explained below.
[0057] Step 1:
[0058] The server initializes the system, specifically loading the AI models (conversational AI, image recognition model, and RAG model) and databases, and securing the necessary resources.
[0059] Step 2:
[0060] The device receives user instructions and uses speech recognition to convert user voice instructions, such as "Tell me how to connect my smart TV," into text.
[0061] Step 3:
[0062] The device captures images of the user's surroundings with a camera, which then captures the images in real time and transmits them to a server via a network.
[0063] Step 4:
[0064] The server analyzes the received camera footage and uses image recognition models to identify objects in the footage, as well as each port and connection on the back of the smart TV.
[0065] Step 5:
[0066] Based on the analysis results, the server searches the database for the appropriate connection manual information, formats the found manual information, and sends it to the terminal.
[0067] Step 6:
[0068] The device then provides the received manual information to the user in AR format, overlaying images of the cables corresponding to each port on the smart TV and the connection instructions on the screen.
[0069] Step 7:
[0070] The device uses a voice synthesis function to guide the user, saying, "Please insert the HDMI cable into the third port from the left." The user then follows the instructions.
[0071] Step 8:
[0072] If the user has an additional question, they can ask aloud, for example, "Where is the power button?"
[0073] Step 9:
[0074] The device uses voice recognition to convert the user's follow-up questions into text and sends them to the server, which then uses conversational AI to analyze the questions and generate appropriate answers.
[0075] Step 10:
[0076] The server sends the generated answer to the device, and the device provides the answer to the user in AR display and voice guidance, such as "The round button at the bottom right is the power button."
[0077] Step 11:
[0078] If the user experiences difficulty in operation, they can press the emergency button, and the device will immediately notify the server of the emergency situation.
[0079] Step 12:
[0080] The server receives emergency notifications and connects to live operators, who provide the user's current camera footage and status to provide real-time support.
[0081] This series of processing steps allows the user to smoothly complete operations and learning while receiving visual and audio guidance.
[0082] Example 1
[0083] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0084] Conventional manual information provision systems have the problem that they are difficult for users to understand intuitively and are unable to provide information efficiently. Another problem is that users cannot solve problems themselves, making it difficult to respond quickly in the event of an emergency. To solve these problems, a system is needed that can understand the user's voice instructions, analyze camera footage, provide appropriate manual information, and respond quickly if the user has additional questions or is in trouble.
[0085] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0086] In this invention, the server includes a means for receiving user instructions, a means for converting voice data into text data, a means for acquiring camera images and transmitting them to the server, a means for analyzing the received camera images and identifying appropriate manual information, a means for searching for the identified manual information and providing it to the terminal, and a means for providing information to the user through an augmented reality display and voice guidance on the terminal. This allows for intuitive information provision via visual and audio, making it easier for the user to quickly solve problems. Furthermore, if the user encounters a difficult situation, the server can notify the user of an emergency situation and quickly contact a manned operator.
[0087] A "user" is an entity that uses this system to provide camera images and give voice instructions and perform touch operations.
[0088] A "means for receiving instructions" is a device or software that has the function of recognizing voice or touch operations made by the user and converting that information into text data.
[0089] The "means for converting voice data into text data" is a technology for converting a user's voice into text, and is a function realized by using voice recognition technology.
[0090] The "means for acquiring camera images and transmitting them to a server" refers to a device or software that can acquire real-time video data captured by a user and transmit it to a server.
[0091] The "means for analyzing the received camera footage and identifying appropriate manual information" refers to a technology in which the server uses video analysis technology to analyze the camera footage and search for and identify the manual information based on the analysis results.
[0092] The "means for searching for specified manual information and providing it to the terminal" is a function for searching for specific information from the manual information stored in the database and transmitting that information to the terminal.
[0093] "Means for providing information to users through augmented reality display and voice guidance on a terminal" refers to a function that uses AR technology to visually display information on a terminal and voice guidance to users using voice synthesis technology.
[0094] "Means for generating answers to questions" refers to a function that analyzes questions received from users and generates appropriate answers using conversational AI.
[0095] The "means for notifying an emergency situation and contacting a manned operator" is a function that, when a user faces a difficult situation, notifies the server of an emergency situation from the terminal and the server contacts a manned operator.
[0096] This invention is a system that receives user instructions, analyzes camera images based on those instructions, and provides appropriate manual information to the user. This system is composed of a server and a terminal, and aims to provide information that the user can intuitively understand visually and audibly.
[0097] Receiving user instructions
[0098] The device receives instructions from the user through voice or touch operations. For example, the user may say, "Tell me how to connect to my smart TV." At this time, the device's microphone picks up the voice and converts it into text using voice recognition technology (e.g., Google (registered trademark) Speech-to-Text API). This textual instruction is then sent to the server.
[0099] Video acquisition and transmission
[0100] When a user points the rear of a smart TV at the device's camera, the device captures the camera image in real time and sends it to the server. The device requires a high-resolution camera and a stable network connection to ensure reliable transmission of the image data to the server.
[0101] Video analysis and information identification
[0102] The server receives the camera images sent from the device and analyzes them using an image recognition model (e.g., the image recognition model of TENSORFLOW (registered trademark)). The server identifies each port and cable connection in the image, and based on the results, retrieves the appropriate connection procedures and manual information from a database (e.g., MySQL (registered trademark)).
[0103] Providing manual information
[0104] The server sends the identified manual information to the device. The device receives this information and uses augmented reality (AR) display technology (e.g., Vuforia SDK) to display the corresponding cable connection location on the back of the TV. It also uses voice guidance technology (e.g., Amazon Polly) to guide the user, saying, "Plug the HDMI cable into the third port from the left."
[0105] Response to additional questions
[0106] If the user asks a follow-up question, such as "Where is the power button?", the device sends this question to the server. The server uses conversational AI (e.g., OpenAI's GPT-3) to analyze the question, generate an answer, and send it to the device. The device then presents this information again in AR displays and voice guidance.
[0107] Emergency response
[0108] If a user experiences difficulty operating the device, for example if they are having difficulty or make a mistake, they can press the emergency button on the device, which will then notify the server of the emergency situation. The server receives this information and connects the user to a live operator, who then monitors the user's current situation and provides appropriate support.
[0109] Specific examples
[0110] Here is a concrete example: If a user purchases a new smart TV and does not know how to connect it, they can use this system to receive the following support:
[0111] 1. The user points the camera at the back of the smart TV and gives a voice command saying, "Tell me how to connect my smart TV."
[0112] 2. The device analyzes the voice, converts the instructions into text, and sends it to the server, while also sending the camera footage.
[0113] 3. The server analyzes the camera footage, identifies each port on the smart TV, retrieves the appropriate connection instructions from the database, and sends the manual information to the device.
[0114] 4. The device will use AR to display the appropriate cable connection location on the back of the TV and will provide voice guidance such as "Plug the HDMI cable into the third port from the left."
[0115] 5. While connected, if the user asks, "Where is the power button?", the device sends this question to the server, which generates an answer and sends it to the device.
[0116] 6. The device will display AR and provide voice guidance, saying, "The round button on the bottom right is the power button."
[0117] Prompt Sentence Examples
[0118] Here are some example prompts to input to a generative AI model:
[0119] "The user gives voice instructions such as, 'Tell me how to connect my smart TV.' The device converts the voice to text using speech recognition technology and sends this information to the server along with the camera footage. The server uses an image recognition model to analyze the camera footage and generate appropriate manual information. The device then provides instructions to the user using AR displays and voice guidance. If the user asks a follow-up question such as, 'Where is the power button?' the device sends this question to the server, and the server uses conversational AI to generate an answer, which is again conveyed to the user using AR displays and voice guidance."
[0120] This prompt allows the generative AI model to understand the system's behavior and generate appropriate instructions.
[0121] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0122] Step 1:
[0123] The user points the camera at the back of the smart TV and gives a voice command such as "Tell me how to connect my smart TV." The user's voice data is acquired as input. The device uses a microphone to capture the user's voice. Next, it calls voice recognition technology (e.g., Google Speech-to-Text API) to convert the voice data into text data. This outputs text data from the voice data. This text data is then transferred to the server, representing the user's command.
[0124] Step 2:
[0125] The device uses a camera to capture real-time video data. As input, video data from the back of the smart TV is captured from the camera. This video data is compressed, encoded, and sent to the server using a network protocol. The output is the video data sent from the device to the server.
[0126] Step 3:
[0127] The server receives video data sent from the device. The input is the image data sent from the device. The server invokes an image recognition model (e.g., TensorFlow's image recognition model) to analyze the video data. This analysis identifies each port or cable connection in the video. The output is the location data of the identified ports or cable connections.
[0128] Step 4:
[0129] The server searches a database (e.g., MySQL) based on the analysis results to obtain the appropriate connection procedures and manual information. The input is the video analysis results. A database query is executed to obtain the connection procedures and manual information. This information is formatted in JSON format or similar and sent to the device. The output is the obtained manual information.
[0130] Step 5:
[0131] The device uses the received manual information to display an augmented reality (AR) image. The input is the manual information sent from the server. The device uses AR display technology (e.g., Vuforia SDK) to display the corresponding cable connection location on the back of the smart TV. It also uses voice guidance technology (e.g., Amazon Polly) to guide the user, saying, "Plug the HDMI cable into the third port from the left." The output is the displayed AR instructions and the voice guidance.
[0132] Step 6:
[0133] The user asks a follow-up question: "Where is the power button?" The user's voice data is again taken as input. The device converts the voice to text and sends the question to the server.
[0134] Step 7:
[0135] The server uses a conversational AI (e.g., OpenAI's GPT-3) to analyze the user's question. The input is the textual question. The AI model analyzes the question and generates an appropriate answer. This answer is sent back to the device. The output is the generated answer.
[0136] Step 8:
[0137] The device provides the received answer to the user through AR display and voice guidance. The input is the answer data sent from the server. The device uses AR display technology to display the location of the power button. It also provides voice guidance saying, "The round button on the bottom right is the power button." The output is the displayed AR instructions and voice guidance.
[0138] Step 9:
[0139] When a user finds themselves in a difficult situation, they press the emergency button on their device. The input is the act of pressing the emergency button. The device notifies the server of the emergency situation. The server receives this information and sends an emergency notification to a manned operator. The operator monitors the user's situation in real time and gives instructions via video or voice call. The output is the operator providing support.
[0140] (Application example 1)
[0141] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0142] In factory maintenance work, there is a challenge for workers to quickly understand the appropriate procedures and perform the work accurately. Also, if a problem occurs during work or if they do not understand the procedures, there are limited means to smoothly obtain support. This reduces work efficiency and, in some cases, can lead to serious mistakes or accidents.
[0143] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0144] In this invention, the server includes a means for receiving user instructions via voice or touch operation, a means for acquiring camera images and transmitting them to the server, a means for analyzing the received camera images and identifying appropriate manual information, a means for searching for the identified manual information and providing it to the terminal, a means for providing information to the user through augmented reality display and audio guidance on the terminal, and a means for providing the user with audio guidance on appropriate procedures and operating methods based on the analysis results. This allows workers to quickly understand and perform appropriate maintenance procedures while receiving intuitive visual and audio instructions. Furthermore, a function for generating answers to user questions using a generative AI model allows for immediate resolution of questions that arise during work.
[0145] "Means for receiving user instructions by voice or touch operation" refers to a system that uses voice recognition technology or a touch sensor to obtain operations or questions from the user as input data.
[0146] The "means for acquiring camera images and transmitting them to a server" is a function for capturing real-time video data using a camera device and transmitting the video data to a remote server via a network.
[0147] "Means for analyzing received camera footage and identifying appropriate manual information" refers to a system that analyzes received video data using image recognition technology and machine learning models, and then searches and retrieves the necessary manuals and procedural information from a database based on the analysis results.
[0148] The "means for searching for specified manual information and providing it to the terminal" is a function for quickly searching for related information in the database based on the analysis results and providing that information to the user's terminal.
[0149] "Means for providing information to users through augmented reality display and voice guidance on a terminal" refers to a system that uses augmented reality technology to overlay visual information on the terminal display and also uses voice synthesis technology to provide voice guidance to users.
[0150] "Means for providing the user with voice guidance on the appropriate procedures and operation methods based on the analysis results" is a function that uses voice synthesis technology to provide instructions on the next operation or procedure that the user should perform based on the analysis results.
[0151] "Means for generating answers to user questions using a generative AI model" refers to a system that uses natural language processing technology and machine learning models to automatically generate and provide appropriate answers to user questions.
[0152] "Means for notifying an emergency situation and contacting a manned operator" is a function that receives an alert from the user in the event of an emergency, notifies the information to a manned support person in real time, and allows direct contact.
[0153] This invention is a system that aims to improve the efficiency of factory maintenance work and support workers. This system consists of a server and a terminal, and allows users to intuitively give instructions by voice or touch operation, and analyzes camera images to provide appropriate manual information.
[0154] System configuration
[0155] The system consists of the following main components:
[0156] 1. Terminal: Equipped with voice recognition and touch input means, camera image acquisition means, augmented reality display and voice guidance functions. Specifically, the terminal is equipped with a microphone, camera, display, speaker, and network connection function.
[0157] 2. Server: Has the means to receive and analyze camera footage. It uses a generative AI model to generate answers to user questions, searches a database for appropriate manual information, and provides it to the device.
[0158] Program processing overview
[0159] 1. Speech recognition and text conversion: The device uses speech recognition technology (specifically, the speech_recognition library) to convert the user's voice instructions into text data. For touch operations, the device uses the touch sensor.
[0160] 2. Acquiring and transmitting camera images: The device's camera is used to acquire real-time images, and the image data is transmitted to the server via the network.
[0161] 3. Video analysis: The received camera footage is analyzed on the server. Image recognition technology and machine learning models (specifically, the BERT-based pre-trained model bert-base-japanese) are used to recognize objects and text in the footage and identify appropriate manual information.
[0162] 4. Providing manual information: The server sends the identified manual information to the terminal, which then provides the information to the user through augmented reality display and voice guidance (using the pyttsx3 library).
[0163] 5. Answer generation using the generative AI model: The server uses the generative AI model to generate answers to the user's follow-up questions and sends the answers to the device.
[0164] Specific examples
[0165] For example, if a worker says, "Tell me how to remove the bearing," the device analyzes the voice and sends the camera footage to the server. The server analyzes the footage, identifies the appropriate procedure, and sends it to the device. The device then displays an augmented reality display and provides voice guidance, saying, "Remove the bolt, then remove the bearing."
[0166] Example prompt sentence:
[0167] "Please explain the proper procedure to remove the bearing from the machine, using the captured image as a reference."
[0168] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0169] Step 1:
[0170] User instruction input
[0171] The user inputs instructions by voice or touch. In the case of voice instructions, voice data is acquired through the device's microphone. In the case of touch operations, tactile information is acquired by the touch sensor.
[0172] Input: Voice or touch data
[0173] Output: Acquired voice data or touch operation data
[0174] Step 2:
[0175] Speech recognition and text conversion
[0176] The device uses speech recognition technology to convert voice data into text data. Specifically, it uses the speech_recognition library to recognize speech and convert it into text.
[0177] Input: Audio data
[0178] Output: Text data
[0179] Specific operation: The device analyzes the acquired voice data, performs voice recognition, and generates text data such as "Please tell me how to remove the bearing."
[0180] Step 3:
[0181] Acquiring and transmitting camera images
[0182] The user captures the object being handled on the camera, and the terminal uses the camera to capture the image in real time and transmits it to the server via the network.
[0183] Input: Camera video data
[0184] Output: Video data sent to the server
[0185] Specific operation: The device operates the camera and streams the captured video data to the server in real time.
[0186] Step 4:
[0187] Video analysis and information identification
[0188] The server analyzes the received video data and identifies appropriate manual information based on the object's condition using image recognition technology and machine learning models (e.g., bert-base-japanese).
[0189] Input: Video data
[0190] Output: Identified manual information
[0191] Specific operation: The server recognizes each object in the video, identifies elements such as "bolt" and "bearing," and then searches for related manual information in a database.
[0192] Step 5:
[0193] Providing manual information
[0194] The server transmits the identified manual information to the terminal, which provides the information to the user through an augmented reality display and voice guidance.
[0195] Input: Identified manual information
[0196] Output: Visual and audio information provided to the user
[0197] Specific operation: The device uses AR technology to superimpose guidelines on the user's display and provides voice guidance such as, "Remove the bolt, then remove the bearing."
[0198] Step 6:
[0199] Responding to additional questions
[0200] If the user asks a follow-up question, the device converts the question into text and sends it to the server, which uses a generative AI model to generate an answer and sends it to the device, which then provides it to the user.
[0201] Input: User's additional question (voice data or touch operation data)
[0202] Output: The answer generated by the server
[0203] Specific behavior: For example, when a user asks, "How do I insert a new bearing?", the server uses a generative AI model to answer, "Insert the bearing in the correct direction and secure it with a bolt."
[0204] Step 7:
[0205] Emergency response
[0206] When a user is in trouble, they can press the emergency button on their device, which will immediately notify the server of the emergency situation, and the server will contact a live operator to provide support.
[0207] Input: Emergency alert signal
[0208] Output: Support by manned operators begins
[0209] Specific operation: When an emergency occurs, the server notifies the monitoring operator, who checks the user's situation in real time and provides appropriate support.
[0210] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0211] This invention is a system that receives user instructions, analyzes camera footage based on those instructions, and provides appropriate manual information to the user.By combining a server, a terminal, and an emotion engine, this system provides more flexible and effective support according to the user's emotional state.
[0212] Receiving user instructions
[0213] The device receives instructions from the user through voice or touch operations. For example, if the user says, "Tell me how to connect to my smart TV," the device picks up the voice through a microphone and converts it into text using voice recognition technology. This textual instruction is then sent to the server.
[0214] Video acquisition and transmission
[0215] When a user points the back of a smart TV at the device's camera, the device captures the camera image in real time and sends it to the server, which uses the captured image to understand what the user is looking at and the input status to the system.
[0216] emotion recognition
[0217] The device is equipped with an emotion engine that recognizes emotions in real time from the user's voice and facial expressions. For example, it can analyze the user's emotions, such as irritation or distress, from the tone of their voice and facial expressions.
[0218] Video analysis and information identification
[0219] The server analyzes the received camera footage and uses image recognition models to identify objects within the footage, recognizes each port and connection on the back of the smart TV, and uses the results to identify the appropriate connection instructions and manual information. This information is retrieved from a database.
[0220] Providing manual information
[0221] The server sends the identified manual information to the device. The device receives this information and provides it to the user through AR display and voice guidance. For example, the device display may overlay the corresponding cable connection point on the back of the smart TV, and voice guidance may say, "Plug the HDMI cable into the third port from the left."
[0222] Using an emotion engine, the content and tone of the instructions can be adjusted depending on the user's emotional state. For example, if the user is confused, the instructions will be slowed down and explained in a gentler tone.
[0223] Response to additional questions
[0224] If the user has an additional question, they can ask aloud, for example, "Where is the power button?"
[0225] The device uses voice recognition to convert the user's follow-up questions into text and sends them to the server, which then uses conversational AI to analyze the questions and generate appropriate answers.
[0226] The server sends the generated answer to the device, and the device provides the answer to the user in AR display and voice, for example, "The round button on the bottom right is the power button."
[0227] Emergency response
[0228] If a user encounters a problem, such as difficulty operating the device or making a mistake, they can press the emergency button on their device, which will then notify the server of the emergency situation. The server receives this information and connects them to a live operator in real time. The emergency operator will monitor the user's current situation in real time and provide appropriate support.
[0229] Specific examples
[0230] Here is a concrete example: If a user purchases a new smart TV and does not know how to connect it, they can use this system to receive the following support:
[0231] 1. The user points the camera at the back of the smart TV and gives voice instructions from the device saying, "Tell me how to connect to the smart TV."
[0232] 2. The device analyzes the voice, converts the instructions into text, and sends it to the server, along with the camera footage.
[0233] 3. The server analyzes the camera footage, identifies each port on the smart TV, determines the appropriate connection procedure, and sends the manual information to the device.
[0234] 4. The device will use AR to display the appropriate cable connection location on the back of the TV and will provide voice guidance such as "Plug the HDMI cable into the third port from the left."
[0235] 5. While connected, if the user asks, "Where is the power button?", the device sends this question to the server, which generates an answer and sends it to the device.
[0236] 6. The device will display AR and provide voice guidance, saying, "The round button on the bottom right is the power button."
[0237] 7. If the emotion engine recognizes that the user is in trouble, it will adjust the guidance and provide a more friendly tone.
[0238] 8. Finally, if users experience difficulties, they can press an emergency button to connect to a live operator for real-time support.
[0239] In this way, the present invention allows users to quickly perform operations and tasks while receiving intuitive visual and audio instructions. In addition, the emotion recognition function can provide more appropriate support for the user's condition, improving the user experience.
[0240] The processing flow will be explained below.
[0241] Step 1:
[0242] The server initializes the system, specifically loading the AI models (conversational AI, image recognition model, and emotion engine) and databases, and securing the necessary resources.
[0243] Step 2:
[0244] The device receives instructions from the user. For example, if the user issues a voice command such as "Tell me how to connect to my smart TV," the device captures the voice through a microphone and converts it into text using voice recognition technology. The converted text data is then sent to the server.
[0245] Step 3:
[0246] The device captures images of the user's surroundings with a camera, and the captured images are sent to a server via a network in real time.
[0247] Step 4:
[0248] The device's built-in emotion engine analyzes the user's voice and facial expressions to recognize their emotional state. For example, it can recognize that the user is irritated based on their voice tone and facial expression.
[0249] Step 5:
[0250] The server analyzes the received camera footage and uses image recognition models to identify objects in the footage, such as the individual ports of a smart TV, and then identifies the appropriate connection procedures and manual information.
[0251] Step 6:
[0252] The server searches the database for the appropriate manual information and sends the identified information, including cable connection procedures and reference images, to the terminal.
[0253] Step 7:
[0254] The device then provides the received manual information to the user through AR displays and voice guidance. For example, the device might overlay the corresponding cable connection on the back of a smart TV and provide voice guidance such as, "Plug the HDMI cable into the third port from the left."
[0255] Step 8:
[0256] The device's emotion engine monitors the user's emotional state and, if it detects that the user is confused, adjusts the content and tone of the instructions, for example, slowing down the pace of the instructions and using a gentler tone.
[0257] Step 9:
[0258] If the user has a follow-up question, they may provide a voice prompt, such as "Where is the power button?"
[0259] Step 10:
[0260] The device uses voice recognition technology to convert follow-up questions into text and send it to the server, which uses conversational AI to analyze the questions and generate appropriate answers.
[0261] Step 11:
[0262] The server sends the generated answer to the device, which then provides it to the user in the form of an AR display and voice guidance, such as "The round button at the bottom right is the power button."
[0263] Step 12:
[0264] If a user experiences difficulty operating the device or needs additional support, they can press the emergency button on the device, which will immediately notify the server of the emergency situation.
[0265] Step 13:
[0266] When the server receives an emergency notification, it connects to a live operator, who is provided with the user's current camera footage and situation in real time.
[0267] Step 14:
[0268] A human operator checks the user's situation in real time and provides appropriate support, providing the instructions necessary to help the user solve the problem in real time.
[0269] Through the above processing steps, the system of the present invention provides visual and audio support to the user, and recognizes the user's emotional state to enable optimal assistance.
[0270] Example 2
[0271] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0272] Conventional manual provision systems often do not provide sufficient support when users have difficulty operating the system, and tasks often stall due to emotional confusion or a lack of understanding of the operating procedures. Furthermore, conventional systems do not support real-time dialogue or emotion recognition, which hinders the user experience.
[0273] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0274] In this invention, the server includes means for receiving user instructions, means for acquiring camera images and transmitting them to the server, means for analyzing the received camera images and identifying appropriate manual information, means for retrieving the identified manual information and providing it to the terminal, means for providing information to the user through augmented reality display and voice guidance at the terminal, means for recognizing emotions from the user's tone of voice and facial expressions, and means for adjusting the content and tone of guidance according to the emotions. This allows the server to provide flexible support according to the user's emotions even when the user is confused, helping the user understand the operating procedures and quickly solving the problem.
[0275] The "means for receiving user instructions" refers to a mechanism for recognizing instructions given by the user through voice or touch operation, converting them into a format such as text data, and processing them.
[0276] "Means for acquiring camera images and transmitting them to a server" refers to a technology for acquiring video data in real time using a camera and transmitting that data to a server.
[0277] "Means for analyzing received camera footage and identifying appropriate manual information" refers to algorithms or programs that analyze the video data received by the server and identify the most appropriate instructions or manual information based on its content.
[0278] The "means for searching for the specified manual information and providing it to the terminal" is a mechanism by which the server searches the database for the specified manual information and transmits it to the user's terminal for provision.
[0279] "Means for providing information to users through augmented reality display and voice guidance on a terminal" refers to a technology that uses the terminal's display to visually display information in augmented reality (AR) and simultaneously conveys the information to users using voice guidance.
[0280] "Means for recognizing emotions from a user's tone of voice and facial expressions" refers to a program or algorithm that analyzes a user's tone of voice and facial expressions in real time to determine their emotional state.
[0281] "Means for adjusting the content and tone of guidance according to emotions" refers to a mechanism for appropriately adjusting the content and tone of the guidance or explanation provided based on the recognized emotional state of the user.
[0282] "Means for generating answers to user questions based on analyzed camera footage and emotional data" refers to a technology in which the server analyzes camera footage and recognized emotional data, and generates appropriate answers to user questions based on that information.
[0283] "Means for notifying an emergency situation and contacting a manned operator" refers to a mechanism that, when a user operates an emergency button or the like, notifies the server of the situation and allows for real-time contact with a manned operator.
[0284] This invention is a system that receives user instructions, analyzes camera images based on those instructions, and provides appropriate manual information to the user. In addition, by combining it with an emotion engine, it provides flexible and effective support according to the user's emotional state.
[0285] Hardware Configuration
[0286] This system mainly consists of a terminal and a server. The terminal is equipped with a camera, microphone, display, and emotion engine. The server analyzes the camera footage and identifies and provides manual information.
[0287] Software Configuration
[0288] The system integrates Google's speech recognition API, an image recognition model using TensorFlow, and an emotion recognition engine using the Affectiva SDK, enabling multiple functions such as voice recognition to receive user instructions, analyzing camera footage, and recognizing the user's emotional state.
[0289] System Operation
[0290] When a user speaks to the device, the device picks up the voice through its built-in microphone and converts it into text using Google's speech recognition API. This textual instruction is then sent to the server. At the same time, a video of the object the user specified through the camera (for example, the back of a smart TV) is captured in real time and sent to the server. The server then analyzes the received video using TensorFlow and identifies the object within the video. Based on the results, it identifies the appropriate connection procedures and manual information from a database and sends them to the device.
[0291] emotion recognition
[0292] The device is also equipped with an emotion engine using the Affectiva SDK, which recognizes the user's emotions in real time from their tone of voice and facial expressions. For example, if the user is confused, the device will provide appropriate support based on their emotions, such as slowing down the pace of the instructions and using a gentler tone of voice.
[0293] Specific examples
[0294] For example, if a user purchases a new smart TV and doesn't know how to connect it, they can use this system to receive the following support:
[0295] 1. The user points the camera at the back of the smart TV and speaks to the device, saying, "Tell me how to connect to my smart TV."
[0296] 2. The device analyzes the voice, converts the instructions into text, and sends it to the server, along with the camera footage.
[0297] 3. The server analyzes the camera footage, identifies each port on the smart TV, determines the appropriate connection procedure, and sends the manual information to the device.
[0298] 4. The device will use AR to display the appropriate cable connection location on the back of the TV and will provide voice guidance such as "Plug the HDMI cable into the third port from the left."
[0299] 5. While connected, if the user asks, "Where is the power button?", the device sends this question to the server, which uses conversational AI to generate an answer and sends it to the device.
[0300] 6. The device will display AR and provide voice guidance, saying, "The round button on the bottom right is the power button."
[0301] 7. If the emotion engine recognizes that the user is in trouble, it will adjust the guidance and provide a more friendly tone.
[0302] 8. Finally, if users experience difficulties, they can press an emergency button to connect to a live operator for real-time support.
[0303] Prompt Sentence Examples
[0304] "The user gives voice instructions such as, 'Tell me how to connect my smart TV,' and sends the video to the server. The server analyzes the video, identifies the appropriate connection procedure, and provides instructions on how to connect using AR displays and voice guidance. If the user asks, 'Where is the power button?' along the way, the server uses conversational AI to generate an answer, and the device provides guidance to the user. If the user has trouble, they can press the emergency button to connect to an operator. How does the system explain this series of steps?"
[0305] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0306] Step 1: Receive user instructions
[0307] The device receives the user's voice instructions and acquires voice data using the built-in microphone. The input is the user's voice instructions and the output is voice data. The device sends this voice data to Google's speech recognition API, which converts the voice into text data. The input is voice data and the output is text data. This textual instruction is sent to the server. The input is text data and the output is data sent to the server.
[0308] Specific behavior:
[0309] 1. A user says, "How do I connect my smart TV?"
[0310] 2. The device's microphone picks up audio data.
[0311] 3. The voice data acquired by the device is sent to Google's speech recognition API and converted into text data.
[0312] 4. The device sends the text data to the server.
[0313] Step 2: Acquire and transmit video
[0314] The device uses a camera to capture real-time video of an object specified by the user (e.g., the back of a smart TV). The input is the camera image, and the output is real-time video data. This video data is compressed and sent to the server. The input is real-time video data, and the output is data sent to the server.
[0315] Specific behavior:
[0316] 1. The user points the back of the smart TV at the device's camera.
[0317] 2. The device's camera captures video in real time.
[0318] 3. The device compresses the video data and sends it to the server.
[0319] Step 3: Emotion Recognition
[0320] The emotion engine installed on the device recognizes emotions from the user's voice tone and facial expressions. The input is the user's voice tone and facial expression data, and the output is the recognized emotion data. Emotions are analyzed in real time using the Affectiva SDK. The input is voice tone and facial expression data, and the output is emotion data.
[0321] Specific behavior:
[0322] 1. The device's camera and microphone capture the user's facial expressions and voice.
[0323] 2. The Affectiva SDK installed on the device analyzes this data and recognizes emotions.
[0324] Step 4: Video analysis and information identification
[0325] The server analyzes the received camera footage and identifies the appropriate manual information. The input is the camera video data, and the output is the identified manual information. An image recognition model is deployed using TensorFlow to identify objects in the video. The input is the camera video data, and the output is the object recognition results. Based on this result, the appropriate manual information is retrieved from the database. The input is the object recognition results, and the output is the manual information.
[0326] Specific behavior:
[0327] 1. The server analyzes the received video data using a TensorFlow model.
[0328] 2. The server identifies each port on the smart TV from the analysis results.
[0329] 3. The server retrieves the manual information from the database based on the identified information.
[0330] Step 5: Provide manual information
[0331] The server sends the identified manual information to the device. The input is the manual information, and the output is data transmission to the device. The device receives this information and uses AR display technology (e.g., Vuforia) to display the appropriate connection location on the back of the TV and provide voice guidance. The input is the manual information, and the output is the AR display and voice guidance.
[0332] Specific behavior:
[0333] 1. The server sends the specified manual information to the terminal.
[0334] 2. The device uses Vuforia to display an AR display on the back of the TV showing the appropriate connection location.
[0335] 3. The device will say "Please insert the HDMI cable into the third port from the left."
[0336] Step 6: Addressing additional questions
[0337] The server uses conversational AI to generate answers to the user's follow-up questions and sends them to the device. The input is the user's question text and the output is the generated answer. The device converts the question into text using voice recognition and sends it to the server. The input is voice recognition data and the output is text data. The server uses conversational AI to generate an appropriate answer. The input is text data and the output is answer data. The device provides the answer as voice guidance and AR display. The input is answer data and the output is voice guidance and AR display.
[0338] Specific behavior:
[0339] 1. The user asks, "Where is the power button?"
[0340] 2. The device converts the question into text and sends it to the server.
[0341] 3. The server uses OpenAI GPT-3 to generate an appropriate answer and send it to the device.
[0342] 4. The device will announce, "The round button on the bottom right is the power button," and will also show this in AR.
[0343] Step 7: Emergency response
[0344] If the user experiences difficulty, the terminal will connect to a manned operator in real time. The input is the operation of the emergency button, and the output is an emergency notification to the server. When the emergency button on the terminal is pressed, an emergency notification is sent to the server. The input is emergency status data, and the output is data sent to the server. The server receives this information and connects to a manned operator in real time. The input is emergency notification data, and the output is operator connection data.
[0345] Specific behavior:
[0346] 1. The user presses the emergency button.
[0347] 2. The device sends an emergency notification to the server.
[0348] 3. The server connects to a live operator.
[0349] (Application example 2)
[0350] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0351] In conventional work support systems and shopping assistant systems, users often have difficulty understanding specific work procedures or detailed product usage. Furthermore, because they provide uniform support without considering the user's emotional state, they often fail to provide effective support. This can lead to operational errors and lack of understanding, resulting in a poor user experience.
[0352] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0353] In this invention, the server includes means for receiving user instructions, means for acquiring camera images and transmitting them to the server, means for analyzing the received camera images and identifying appropriate manual information, means for searching for the identified manual information and providing it to the terminal, means for providing information to the user through augmented reality display and audio guidance on the terminal, means for providing product usage instructions and features based on product information detected from the camera images, and means for analyzing the user's emotional state and providing flexible and effective support according to the user's emotions. This allows the user to quickly perform operations and tasks while receiving intuitive visual and audio instructions. Furthermore, the emotion recognition function can more appropriately support the user's state, improving the user experience.
[0354] "Means for receiving user instructions" refers to the means by which a user inputs instructions to the system through voice or touch operations.
[0355] "Means for acquiring camera images and transmitting them to a server" refers to means that has the function of capturing images using a camera and transferring the images to a server.
[0356] "Means for analyzing received camera footage and identifying appropriate manual information" refers to the means by which the server performs object recognition and information analysis based on the video data received, and extracts the corresponding manual information.
[0357] The "means for retrieving the specified manual information and providing it to the terminal" refers to a means for retrieving the extracted manual information from the database and transmitting the information to the user's terminal.
[0358] "Means for providing information to users through augmented reality displays and audio guidance on a device" refers to means for providing information to users visually and audibly using AR technology and audio guidance on a device.
[0359] "Means for providing product usage and features based on product information detected from camera footage" refers to means for providing users with information about detailed usage and features of products identified by analyzing camera footage.
[0360] "Means for analyzing the emotional state of a user and providing flexible and effective support according to the emotion" refers to means for analyzing the emotional state of a user using emotion recognition technology and providing appropriate support according to that state.
[0361] "Means for generating an answer to a user's question based on analyzed camera footage" refers to means for generating an answer to a follow-up question from a user based on analyzed video data.
[0362] "Means for notifying an emergency situation and contacting a manned operator" refers to a means for a user to notify the system of a difficult situation and immediately connect to a manned operator.
[0363] This invention is a system that receives user instructions, analyzes camera footage based on those instructions, and provides appropriate manual information to the user. This system links the terminal and server, and further combines it with an emotion engine to provide flexible and effective support according to the user's emotional state.
[0364] 1. Receiving user instructions
[0365] The device receives instructions from the user through voice or touch operations. For example, if the user says, "Tell me how to use this product," the device picks up the voice through a microphone and converts it into text using voice recognition technology. This text instruction is then sent to the server.
[0366] 2. Image acquisition and transmission
[0367] When a user photographs a product with the device's camera, the device captures the camera image in real time and sends it to the server. The captured image is used to understand what the user is looking at and the input status to the system.
[0368] 3. Video analysis and information identification
[0369] The server analyzes the received camera footage and uses image recognition technology to identify objects in the footage, recognizes each detail about the product, and determines the appropriate usage and characteristic information based on the results. This information is obtained from a database.
[0370] 4. Providing manual information
[0371] The server sends the identified manual information to the terminal. The terminal receives this information and provides it to the user through an augmented reality display and voice guidance. For example, the terminal may overlay an animation of the product or operating instructions on the display, and provide voice guidance such as "Please do this part like this."
[0372] 5. Emotion recognition
[0373] The device is equipped with an emotion engine that recognizes emotions in real time from the user's voice and facial expressions. For example, it can analyze the user's emotions, such as whether they are in trouble, from the tone of their voice and facial expression. Based on the emotion recognition results, the tone and pace of the guidance are adjusted.
[0374] 6. Response to additional questions
[0375] If the user has an additional question, they can ask it by voice, for example, "How do I use this button?" The device converts the user's additional question into text using its voice recognition function and sends it to the server. The server uses conversational AI to analyze the question and generate an appropriate answer. The server then sends the generated answer to the device, which then provides the answer to the user in AR display and voice.
[0376] 7. Emergency Response
[0377] If a user encounters a problem, for example, if they are having difficulty operating the device or make a mistake, they can press the emergency button on their device, which will then notify the server of the emergency. The server will receive this information and connect them to a live operator in real time. The operator will then monitor the user's current situation in real time and provide appropriate support.
[0378] Hardware and software used
[0379] 1. Camera: For example, a smartphone camera.
[0380] 2. Microphone: For example, the microphone on your smartphone.
[0381] 3. Displays: smartphone displays, AR glasses, etc.
[0382] 4. Emotion recognition models: EmotionRecognizer, etc.
[0383] 5. Image recognition technology: Image analysis software such as OpenCV.
[0384] 6. Speech recognition technology: SpeechRecognition, etc.
[0385] Specific examples
[0386] If a user purchases a new kitchen gadget at a brick-and-mortar store and doesn't know how to use it, the system can provide the following assistance:
[0387] 1. The user points the camera at a kitchen gadget and gives voice instructions on the device, such as "Tell me how to use this kitchen gadget."
[0388] 2. The device analyzes the voice, converts the instructions into text, and sends it to the server, along with the camera footage.
[0389] 3. The server analyzes the camera image, recognizes the product, determines the appropriate usage procedure, and sends the manual information to the terminal.
[0390] 4. The device uses AR to display how to use the product and provides voice guidance such as "First, do this."
[0391] 5. If the emotion engine recognizes that the user is in trouble, it will adjust the guidance and provide a more friendly tone.
[0392] 6. 24-hour operator connection button for real-time support.
[0393] Prompt Sentence Examples
[0394] "Please identify the product in this image and provide a detailed description of how to use it."
[0395] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0396] Step 1:
[0397] The terminal receives voice instructions from the user. When the user gives a voice instruction such as "Tell me how to use the product," the terminal uses a microphone to capture the voice data and converts it into text data using voice recognition technology. The input is the user's voice data, and the output is text data.
[0398] Step 2:
[0399] The device sends the converted text data to the server. At the same time, it also sends the product video captured by the device's camera to the server. The input is the text data and the video data, and the output is the data sent to the server.
[0400] Step 3:
[0401] The server analyzes the received video data and uses image recognition technology (e.g., OpenCV) to identify objects within the video. Specifically, it identifies each product or part within the video and extracts information such as its position and shape. The input is the video data, and the output is the product information that is the result of the analysis.
[0402] Step 4:
[0403] The server searches the database for corresponding manual information based on the product information. The searched manual information includes product usage instructions and features. The input is product information, and the output is manual information.
[0404] Step 5:
[0405] The server sends the retrieved manual information to the terminal. The input is the manual information, and the output is the data sent to the terminal.
[0406] Step 6:
[0407] The terminal then provides the received manual information to the user. Specifically, it uses augmented reality technology to overlay product usage and features onto the display. It also provides specific instructions such as "Do this part like this" through voice guidance. The input is the manual information, and the output is the information provided to the user.
[0408] Step 7:
[0409] The device analyzes the user's emotional state. Using an emotion engine, it recognizes emotions in real time from the user's tone of voice and facial expressions. The input is audio and video data, and the output is the user's emotional state.
[0410] Step 8:
[0411] The device adjusts the tone and pace of the guidance based on the user's emotional state. For example, if it recognizes that the user is in a difficult situation, it will explain the guidance in a more friendly tone and at a slower pace. The input is the user's emotional state, and the output is the adjusted guidance.
[0412] Step 9:
[0413] If the user asks a follow-up question, the device converts the question into text using voice recognition and sends it back to the server. The server analyzes the question, generates an appropriate answer, and sends it to the device. The device then provides the answer to the user in AR display and voice. The input is the follow-up question and its analysis results, and the output is the generated answer.
[0414] Step 10:
[0415] If a user encounters a problem, for example, if they are having difficulty operating the device or make a mistake, they press the emergency button on their device. The device then notifies the server of the emergency situation, and the server connects to a live operator in real time. The input is the notification of the emergency situation, and the output is a connection to a live operator.
[0416] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0417] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0418] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0419] [Second embodiment]
[0420] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0421] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0422] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0423] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0424] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0425] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0426] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0427] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0428] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0429] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0430] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0431] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0432] The present invention is a system for receiving user instructions, analyzing camera footage based on those instructions, and providing appropriate manual information to the user. This system is composed of a server and a terminal, and aims to provide the user with visual and audio information that can be intuitively understood. Specific embodiments of this system are described below.
[0433] Receiving user instructions
[0434] The device receives instructions from the user through voice or touch operations. For example, if the user says, "Tell me how to connect to my smart TV," the device picks up the voice through a microphone and converts it into text using voice recognition technology. This textual instruction is then sent to the server.
[0435] Video acquisition and transmission
[0436] When a user points the back of a smart TV at the device's camera, the device captures the camera image in real time and sends it to the server. The device is equipped with a camera and a network connection to reliably capture the image.
[0437] Video analysis and information identification
[0438] The server receives the camera footage sent from the device and analyzes it using an image recognition model. The server identifies each port and cable connection in the footage, and based on the results, identifies the appropriate connection procedures and manual information. This information is retrieved from a database.
[0439] Providing manual information
[0440] The server sends the identified manual information to the device. The device receives this information and provides it to the user through AR display and voice guidance. For example, the device display may overlay the corresponding cable connection point on the back of the smart TV, and voice guidance may say, "Plug the HDMI cable into the third port from the left."
[0441] Response to additional questions
[0442] If the user asks a follow-up question, such as "Where is the power button?", the device sends this question to the server. The server uses conversational AI to analyze the question and generate an answer. For example, it might determine that "The round button on the bottom right is the power button" and send this information to the device. The device then presents this information in an AR display and voice guidance.
[0443] Emergency response
[0444] If a user encounters a problem, such as difficulty operating the device or making a mistake, they can press the emergency button on the device, which will then notify the server of the emergency situation. The server receives this information and connects to a live operator in real time. The emergency operator will monitor the user's current situation in real time and provide appropriate support.
[0445] Specific examples
[0446] Here is a concrete example: If a user purchases a new smart TV and does not know how to connect it, they can use this system to receive the following support:
[0447] 1. The user points the camera at the back of the smart TV and gives voice instructions from the device saying, "Tell me how to connect to the smart TV."
[0448] 2. The device analyzes the voice, converts the instructions into text, and sends it to the server, along with the camera footage.
[0449] 3. The server analyzes the camera footage, identifies each port on the smart TV, determines the appropriate connection procedure, and sends the manual information to the device.
[0450] 4. The device will use AR to display the appropriate cable connection location on the back of the TV and will provide voice guidance such as "Plug the HDMI cable into the third port from the left."
[0451] 5. While connected, if the user asks, "Where is the power button?", the device sends this question to the server, which generates an answer and sends it to the device.
[0452] 6. The device will display AR and provide voice guidance, saying, "The round button on the bottom right is the power button."
[0453] In this way, the present invention allows the user to quickly perform operations and tasks while receiving intuitive visual and audio instructions.
[0454] The processing flow will be explained below.
[0455] Step 1:
[0456] The server initializes the system, specifically loading the AI models (conversational AI, image recognition model, and RAG model) and databases, and securing the necessary resources.
[0457] Step 2:
[0458] The device receives user instructions and uses speech recognition to convert user voice instructions, such as "Tell me how to connect my smart TV," into text.
[0459] Step 3:
[0460] The device captures images of the user's surroundings with a camera, which then captures the images in real time and transmits them to a server via a network.
[0461] Step 4:
[0462] The server analyzes the received camera footage and uses image recognition models to identify objects in the footage, as well as each port and connection on the back of the smart TV.
[0463] Step 5:
[0464] Based on the analysis results, the server searches the database for the appropriate connection manual information, formats the found manual information, and sends it to the terminal.
[0465] Step 6:
[0466] The device then provides the received manual information to the user in AR format, overlaying images of the cables corresponding to each port on the smart TV and the connection instructions on the screen.
[0467] Step 7:
[0468] The device uses a voice synthesis function to guide the user, saying, "Please insert the HDMI cable into the third port from the left." The user then follows the instructions.
[0469] Step 8:
[0470] If the user has an additional question, they can ask aloud, for example, "Where is the power button?"
[0471] Step 9:
[0472] The device uses voice recognition to convert the user's follow-up questions into text and sends them to the server, which then uses conversational AI to analyze the questions and generate appropriate answers.
[0473] Step 10:
[0474] The server sends the generated answer to the device, and the device provides the answer to the user in AR display and voice guidance, such as "The round button at the bottom right is the power button."
[0475] Step 11:
[0476] If the user experiences difficulty in operation, they can press the emergency button, and the device will immediately notify the server of the emergency situation.
[0477] Step 12:
[0478] The server receives emergency notifications and connects to live operators, who provide the user's current camera footage and status to provide real-time support.
[0479] This series of processing steps allows the user to smoothly complete operations and learning while receiving visual and audio guidance.
[0480] Example 1
[0481] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0482] Conventional manual information provision systems have the problem that they are difficult for users to understand intuitively and are unable to provide information efficiently. Another problem is that users cannot solve problems themselves, making it difficult to respond quickly in the event of an emergency. To solve these problems, a system is needed that can understand the user's voice instructions, analyze camera footage, provide appropriate manual information, and respond quickly if the user has additional questions or is in trouble.
[0483] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0484] In this invention, the server includes a means for receiving user instructions, a means for converting voice data into text data, a means for acquiring camera images and transmitting them to the server, a means for analyzing the received camera images and identifying appropriate manual information, a means for searching for the identified manual information and providing it to the terminal, and a means for providing information to the user through an augmented reality display and voice guidance on the terminal. This allows for intuitive information provision via visual and audio, making it easier for the user to quickly solve problems. Furthermore, if the user encounters a difficult situation, the server can notify the user of an emergency situation and quickly contact a manned operator.
[0485] A "user" is an entity that uses this system to provide camera images and give voice instructions and perform touch operations.
[0486] A "means for receiving instructions" is a device or software that has the function of recognizing voice or touch operations made by the user and converting that information into text data.
[0487] The "means for converting voice data into text data" is a technology for converting a user's voice into text, and is a function realized by using voice recognition technology.
[0488] The "means for acquiring camera images and transmitting them to a server" refers to a device or software that can acquire real-time video data captured by a user and transmit it to a server.
[0489] The "means for analyzing the received camera footage and identifying appropriate manual information" refers to a technology in which the server uses video analysis technology to analyze the camera footage and search for and identify the manual information based on the analysis results.
[0490] The "means for searching for specified manual information and providing it to the terminal" is a function for searching for specific information from the manual information stored in the database and transmitting that information to the terminal.
[0491] "Means for providing information to users through augmented reality display and voice guidance on a terminal" refers to a function that uses AR technology to visually display information on a terminal and voice guidance to users using voice synthesis technology.
[0492] "Means for generating answers to questions" refers to a function that analyzes questions received from users and generates appropriate answers using conversational AI.
[0493] The "means for notifying an emergency situation and contacting a manned operator" is a function that, when a user faces a difficult situation, notifies the server of an emergency situation from the terminal and the server contacts a manned operator.
[0494] This invention is a system that receives user instructions, analyzes camera images based on those instructions, and provides appropriate manual information to the user. This system is composed of a server and a terminal, and aims to provide information that the user can intuitively understand visually and audibly.
[0495] Receiving user instructions
[0496] The device receives instructions from the user through voice or touch operations. For example, the user may say, "Tell me how to connect to my smart TV." The device's microphone picks up the voice and converts it into text using voice recognition technology (e.g., Google Speech-to-Text API). This textual instruction is then sent to the server.
[0497] Video acquisition and transmission
[0498] When a user points the rear of a smart TV at the device's camera, the device captures the camera image in real time and sends it to the server. The device requires a high-resolution camera and a stable network connection to ensure reliable transmission of the image data to the server.
[0499] Video analysis and information identification
[0500] The server receives the camera images sent from the device and analyzes them using an image recognition model (e.g., TensorFlow image recognition model).The server identifies each port and cable connection in the image and, based on the results, retrieves the appropriate connection procedures and manual information from a database (e.g., MySQL).
[0501] Providing manual information
[0502] The server sends the identified manual information to the device. The device receives this information and uses augmented reality (AR) display technology (e.g., Vuforia SDK) to display the corresponding cable connection location on the back of the TV. It also uses voice guidance technology (e.g., Amazon Polly) to guide the user, saying, "Plug the HDMI cable into the third port from the left."
[0503] Response to additional questions
[0504] If the user asks a follow-up question, such as "Where is the power button?", the device sends this question to the server. The server uses conversational AI (e.g., OpenAI's GPT-3) to analyze the question, generate an answer, and send it to the device. The device then presents this information again in AR displays and voice guidance.
[0505] Emergency response
[0506] If a user experiences difficulty operating the device, for example if they are having difficulty or make a mistake, they can press the emergency button on the device, which will then notify the server of the emergency situation. The server receives this information and connects the user to a live operator, who then monitors the user's current situation and provides appropriate support.
[0507] Specific examples
[0508] Here is a concrete example: If a user purchases a new smart TV and does not know how to connect it, they can use this system to receive the following support:
[0509] 1. The user points the camera at the back of the smart TV and gives a voice command saying, "Tell me how to connect my smart TV."
[0510] 2. The device analyzes the voice, converts the instructions into text, and sends it to the server, while also sending the camera footage.
[0511] 3. The server analyzes the camera footage, identifies each port on the smart TV, retrieves the appropriate connection instructions from the database, and sends the manual information to the device.
[0512] 4. The device will use AR to display the appropriate cable connection location on the back of the TV and will provide voice guidance such as "Plug the HDMI cable into the third port from the left."
[0513] 5. While connected, if the user asks, "Where is the power button?", the device sends this question to the server, which generates an answer and sends it to the device.
[0514] 6. The device will display AR and provide voice guidance, saying, "The round button on the bottom right is the power button."
[0515] Prompt Sentence Examples
[0516] Here are some example prompts to input to a generative AI model:
[0517] "The user gives voice instructions such as, 'Tell me how to connect my smart TV.' The device converts the voice to text using speech recognition technology and sends this information to the server along with the camera footage. The server uses an image recognition model to analyze the camera footage and generate appropriate manual information. The device then provides instructions to the user using AR displays and voice guidance. If the user asks a follow-up question such as, 'Where is the power button?' the device sends this question to the server, and the server uses conversational AI to generate an answer, which is again conveyed to the user using AR displays and voice guidance."
[0518] This prompt allows the generative AI model to understand the system's behavior and generate appropriate instructions.
[0519] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0520] Step 1:
[0521] The user points the camera at the back of the smart TV and gives a voice command such as "Tell me how to connect my smart TV." The user's voice data is acquired as input. The device uses a microphone to capture the user's voice. Next, it calls voice recognition technology (e.g., Google Speech-to-Text API) to convert the voice data into text data. This outputs text data from the voice data. This text data is then transferred to the server, representing the user's command.
[0522] Step 2:
[0523] The device uses a camera to capture real-time video data. As input, video data from the back of the smart TV is captured from the camera. This video data is compressed, encoded, and sent to the server using a network protocol. The output is the video data sent from the device to the server.
[0524] Step 3:
[0525] The server receives video data sent from the device. The input is the image data sent from the device. The server invokes an image recognition model (e.g., TensorFlow's image recognition model) to analyze the video data. This analysis identifies each port or cable connection in the video. The output is the location data of the identified ports or cable connections.
[0526] Step 4:
[0527] The server searches a database (e.g., MySQL) based on the analysis results to obtain the appropriate connection procedures and manual information. The input is the video analysis results. A database query is executed to obtain the connection procedures and manual information. This information is formatted in JSON format or similar and sent to the device. The output is the obtained manual information.
[0528] Step 5:
[0529] The device uses the received manual information to display an augmented reality (AR) image. The input is the manual information sent from the server. The device uses AR display technology (e.g., Vuforia SDK) to display the corresponding cable connection location on the back of the smart TV. It also uses voice guidance technology (e.g., Amazon Polly) to guide the user, saying, "Plug the HDMI cable into the third port from the left." The output is the displayed AR instructions and the voice guidance.
[0530] Step 6:
[0531] The user asks a follow-up question: "Where is the power button?" The user's voice data is again taken as input. The device converts the voice to text and sends the question to the server.
[0532] Step 7:
[0533] The server uses a conversational AI (e.g., OpenAI's GPT-3) to analyze the user's question. The input is the textual question. The AI model analyzes the question and generates an appropriate answer. This answer is sent back to the device. The output is the generated answer.
[0534] Step 8:
[0535] The device provides the received answer to the user through AR display and voice guidance. The input is the answer data sent from the server. The device uses AR display technology to display the location of the power button. It also provides voice guidance saying, "The round button on the bottom right is the power button." The output is the displayed AR instructions and voice guidance.
[0536] Step 9:
[0537] When a user finds themselves in a difficult situation, they press the emergency button on their device. The input is the act of pressing the emergency button. The device notifies the server of the emergency situation. The server receives this information and sends an emergency notification to a manned operator. The operator monitors the user's situation in real time and gives instructions via video or voice call. The output is the operator providing support.
[0538] (Application example 1)
[0539] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0540] In factory maintenance work, there is a challenge for workers to quickly understand the appropriate procedures and perform the work accurately. Also, if a problem occurs during work or if they do not understand the procedures, there are limited means to smoothly obtain support. This reduces work efficiency and, in some cases, can lead to serious mistakes or accidents.
[0541] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0542] In this invention, the server includes a means for receiving user instructions via voice or touch operation, a means for acquiring camera images and transmitting them to the server, a means for analyzing the received camera images and identifying appropriate manual information, a means for searching for the identified manual information and providing it to the terminal, a means for providing information to the user through augmented reality display and audio guidance on the terminal, and a means for providing the user with audio guidance on appropriate procedures and operating methods based on the analysis results. This allows workers to quickly understand and perform appropriate maintenance procedures while receiving intuitive visual and audio instructions. Furthermore, a function for generating answers to user questions using a generative AI model allows for immediate resolution of questions that arise during work.
[0543] "Means for receiving user instructions by voice or touch operation" refers to a system that uses voice recognition technology or a touch sensor to obtain operations or questions from the user as input data.
[0544] The "means for acquiring camera images and transmitting them to a server" is a function for capturing real-time video data using a camera device and transmitting the video data to a remote server via a network.
[0545] "Means for analyzing received camera footage and identifying appropriate manual information" refers to a system that analyzes received video data using image recognition technology and machine learning models, and then searches and retrieves the necessary manuals and procedural information from a database based on the analysis results.
[0546] The "means for searching for specified manual information and providing it to the terminal" is a function for quickly searching for related information in the database based on the analysis results and providing that information to the user's terminal.
[0547] "Means for providing information to users through augmented reality display and voice guidance on a terminal" refers to a system that uses augmented reality technology to overlay visual information on the terminal display and also uses voice synthesis technology to provide voice guidance to users.
[0548] "Means for providing the user with voice guidance on the appropriate procedures and operation methods based on the analysis results" is a function that uses voice synthesis technology to provide instructions on the next operation or procedure that the user should perform based on the analysis results.
[0549] "Means for generating answers to user questions using a generative AI model" refers to a system that uses natural language processing technology and machine learning models to automatically generate and provide appropriate answers to user questions.
[0550] "Means for notifying an emergency situation and contacting a manned operator" is a function that receives an alert from the user in the event of an emergency, notifies the information to a manned support person in real time, and allows direct contact.
[0551] This invention is a system that aims to improve the efficiency of factory maintenance work and support workers. This system consists of a server and a terminal, and allows users to intuitively give instructions by voice or touch operation, and analyzes camera images to provide appropriate manual information.
[0552] System configuration
[0553] The system consists of the following main components:
[0554] 1. Terminal: Equipped with voice recognition and touch input means, camera image acquisition means, augmented reality display and voice guidance functions. Specifically, the terminal is equipped with a microphone, camera, display, speaker, and network connection function.
[0555] 2. Server: Has the means to receive and analyze camera footage. It uses a generative AI model to generate answers to user questions, searches a database for appropriate manual information, and provides it to the device.
[0556] Program processing overview
[0557] 1. Speech recognition and text conversion: The device uses speech recognition technology (specifically, the speech_recognition library) to convert the user's voice instructions into text data. For touch operations, the device uses the touch sensor.
[0558] 2. Acquiring and transmitting camera images: The device's camera is used to acquire real-time images, and the image data is transmitted to the server via the network.
[0559] 3. Video analysis: The received camera footage is analyzed on the server. Image recognition technology and machine learning models (specifically, the BERT-based pre-trained model bert-base-japanese) are used to recognize objects and text in the footage and identify appropriate manual information.
[0560] 4. Providing manual information: The server sends the identified manual information to the terminal, which then provides the information to the user through augmented reality display and voice guidance (using the pyttsx3 library).
[0561] 5. Answer generation using the generative AI model: The server uses the generative AI model to generate answers to the user's follow-up questions and sends the answers to the device.
[0562] Specific examples
[0563] For example, if a worker says, "Tell me how to remove the bearing," the device analyzes the voice and sends the camera footage to the server. The server analyzes the footage, identifies the appropriate procedure, and sends it to the device. The device then displays an augmented reality display and provides voice guidance, saying, "Remove the bolt, then remove the bearing."
[0564] Example prompt sentence:
[0565] "Please explain the proper procedure to remove the bearing from the machine, using the captured image as a reference."
[0566] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0567] Step 1:
[0568] User instruction input
[0569] The user inputs instructions by voice or touch. In the case of voice instructions, voice data is acquired through the device's microphone. In the case of touch operations, tactile information is acquired by the touch sensor.
[0570] Input: Voice or touch data
[0571] Output: Acquired voice data or touch operation data
[0572] Step 2:
[0573] Speech recognition and text conversion
[0574] The device uses speech recognition technology to convert voice data into text data. Specifically, it uses the speech_recognition library to recognize speech and convert it into text.
[0575] Input: Audio data
[0576] Output: Text data
[0577] Specific operation: The device analyzes the acquired voice data, performs voice recognition, and generates text data such as "Please tell me how to remove the bearing."
[0578] Step 3:
[0579] Acquiring and transmitting camera images
[0580] The user captures the object being handled on the camera, and the terminal uses the camera to capture the image in real time and transmits it to the server via the network.
[0581] Input: Camera video data
[0582] Output: Video data sent to the server
[0583] Specific operation: The device operates the camera and streams the captured video data to the server in real time.
[0584] Step 4:
[0585] Video analysis and information identification
[0586] The server analyzes the received video data and identifies appropriate manual information based on the object's condition using image recognition technology and machine learning models (e.g., bert-base-japanese).
[0587] Input: Video data
[0588] Output: Identified manual information
[0589] Specific operation: The server recognizes each object in the video, identifies elements such as "bolt" and "bearing," and then searches for related manual information in a database.
[0590] Step 5:
[0591] Providing manual information
[0592] The server transmits the identified manual information to the terminal, which provides the information to the user through an augmented reality display and voice guidance.
[0593] Input: Identified manual information
[0594] Output: Visual and audio information provided to the user
[0595] Specific operation: The device uses AR technology to superimpose guidelines on the user's display and provides voice guidance such as, "Remove the bolt, then remove the bearing."
[0596] Step 6:
[0597] Responding to additional questions
[0598] If the user asks a follow-up question, the device converts the question into text and sends it to the server, which uses a generative AI model to generate an answer and sends it to the device, which then provides it to the user.
[0599] Input: User's additional question (voice data or touch operation data)
[0600] Output: The answer generated by the server
[0601] Specific behavior: For example, when a user asks, "How do I insert a new bearing?", the server uses a generative AI model to answer, "Insert the bearing in the correct direction and secure it with a bolt."
[0602] Step 7:
[0603] Emergency response
[0604] When a user is in trouble, they can press the emergency button on their device, which will immediately notify the server of the emergency situation, and the server will contact a live operator to provide support.
[0605] Input: Emergency alert signal
[0606] Output: Support by manned operators begins
[0607] Specific operation: When an emergency occurs, the server notifies the monitoring operator, who checks the user's situation in real time and provides appropriate support.
[0608] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0609] This invention is a system that receives user instructions, analyzes camera footage based on those instructions, and provides appropriate manual information to the user.By combining a server, a terminal, and an emotion engine, this system provides more flexible and effective support according to the user's emotional state.
[0610] Receiving user instructions
[0611] The device receives instructions from the user through voice or touch operations. For example, if the user says, "Tell me how to connect to my smart TV," the device picks up the voice through a microphone and converts it into text using voice recognition technology. This textual instruction is then sent to the server.
[0612] Video acquisition and transmission
[0613] When a user points the back of a smart TV at the device's camera, the device captures the camera image in real time and sends it to the server, which uses the captured image to understand what the user is looking at and the input status to the system.
[0614] emotion recognition
[0615] The device is equipped with an emotion engine that recognizes emotions in real time from the user's voice and facial expressions. For example, it can analyze the user's emotions, such as irritation or distress, from the tone of their voice and facial expressions.
[0616] Video analysis and information identification
[0617] The server analyzes the received camera footage and uses image recognition models to identify objects within the footage, recognizes each port and connection on the back of the smart TV, and uses the results to identify the appropriate connection instructions and manual information. This information is retrieved from a database.
[0618] Providing manual information
[0619] The server sends the identified manual information to the device. The device receives this information and provides it to the user through AR display and voice guidance. For example, the device display may overlay the corresponding cable connection point on the back of the smart TV, and voice guidance may say, "Plug the HDMI cable into the third port from the left."
[0620] Using an emotion engine, the content and tone of the instructions can be adjusted depending on the user's emotional state. For example, if the user is confused, the instructions will be slowed down and explained in a gentler tone.
[0621] Response to additional questions
[0622] If the user has an additional question, they can ask aloud, for example, "Where is the power button?"
[0623] The device uses voice recognition to convert the user's follow-up questions into text and sends them to the server, which then uses conversational AI to analyze the questions and generate appropriate answers.
[0624] The server sends the generated answer to the device, and the device provides the answer to the user in AR display and voice, for example, "The round button on the bottom right is the power button."
[0625] Emergency response
[0626] If a user encounters a problem, such as difficulty operating the device or making a mistake, they can press the emergency button on their device, which will then notify the server of the emergency situation. The server receives this information and connects them to a live operator in real time. The emergency operator will monitor the user's current situation in real time and provide appropriate support.
[0627] Specific examples
[0628] Here is a concrete example: If a user purchases a new smart TV and does not know how to connect it, they can use this system to receive the following support:
[0629] 1. The user points the camera at the back of the smart TV and gives voice instructions from the device saying, "Tell me how to connect to the smart TV."
[0630] 2. The device analyzes the voice, converts the instructions into text, and sends it to the server, along with the camera footage.
[0631] 3. The server analyzes the camera footage, identifies each port on the smart TV, determines the appropriate connection procedure, and sends the manual information to the device.
[0632] 4. The device will use AR to display the appropriate cable connection location on the back of the TV and will provide voice guidance such as "Plug the HDMI cable into the third port from the left."
[0633] 5. While connected, if the user asks, "Where is the power button?", the device sends this question to the server, which generates an answer and sends it to the device.
[0634] 6. The device will display AR and provide voice guidance, saying, "The round button on the bottom right is the power button."
[0635] 7. If the emotion engine recognizes that the user is in trouble, it will adjust the guidance and provide a more friendly tone.
[0636] 8. Finally, if users experience difficulties, they can press an emergency button to connect to a live operator for real-time support.
[0637] In this way, the present invention allows users to quickly perform operations and tasks while receiving intuitive visual and audio instructions. In addition, the emotion recognition function can provide more appropriate support for the user's condition, improving the user experience.
[0638] The processing flow will be explained below.
[0639] Step 1:
[0640] The server initializes the system, specifically loading the AI models (conversational AI, image recognition model, and emotion engine) and databases, and securing the necessary resources.
[0641] Step 2:
[0642] The device receives instructions from the user. For example, if the user issues a voice command such as "Tell me how to connect to my smart TV," the device captures the voice through a microphone and converts it into text using voice recognition technology. The converted text data is then sent to the server.
[0643] Step 3:
[0644] The device captures images of the user's surroundings with a camera, and the captured images are sent to a server via a network in real time.
[0645] Step 4:
[0646] The device's built-in emotion engine analyzes the user's voice and facial expressions to recognize their emotional state. For example, it can recognize that the user is irritated based on their voice tone and facial expression.
[0647] Step 5:
[0648] The server analyzes the received camera footage and uses image recognition models to identify objects in the footage, such as the individual ports of a smart TV, and then identifies the appropriate connection procedures and manual information.
[0649] Step 6:
[0650] The server searches the database for the appropriate manual information and sends the identified information, including cable connection procedures and reference images, to the terminal.
[0651] Step 7:
[0652] The device then provides the received manual information to the user through AR displays and voice guidance. For example, the device might overlay the corresponding cable connection on the back of a smart TV and provide voice guidance such as, "Plug the HDMI cable into the third port from the left."
[0653] Step 8:
[0654] The device's emotion engine monitors the user's emotional state and, if it detects that the user is confused, adjusts the content and tone of the instructions, for example, slowing down the pace of the instructions and using a gentler tone.
[0655] Step 9:
[0656] If the user has a follow-up question, they may provide a voice prompt, such as "Where is the power button?"
[0657] Step 10:
[0658] The device uses voice recognition technology to convert follow-up questions into text and send it to the server, which uses conversational AI to analyze the questions and generate appropriate answers.
[0659] Step 11:
[0660] The server sends the generated answer to the device, which then provides it to the user in the form of an AR display and voice guidance, such as "The round button at the bottom right is the power button."
[0661] Step 12:
[0662] If a user experiences difficulty operating the device or needs additional support, they can press the emergency button on the device, which will immediately notify the server of the emergency situation.
[0663] Step 13:
[0664] When the server receives an emergency notification, it connects to a live operator, who is provided with the user's current camera footage and situation in real time.
[0665] Step 14:
[0666] A human operator checks the user's situation in real time and provides appropriate support, providing the instructions necessary to help the user solve the problem in real time.
[0667] Through the above processing steps, the system of the present invention provides visual and audio support to the user, and recognizes the user's emotional state to enable optimal assistance.
[0668] Example 2
[0669] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0670] Conventional manual provision systems often do not provide sufficient support when users have difficulty operating the system, and tasks often stall due to emotional confusion or a lack of understanding of the operating procedures. Furthermore, conventional systems do not support real-time dialogue or emotion recognition, which hinders the user experience.
[0671] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0672] In this invention, the server includes means for receiving user instructions, means for acquiring camera images and transmitting them to the server, means for analyzing the received camera images and identifying appropriate manual information, means for retrieving the identified manual information and providing it to the terminal, means for providing information to the user through augmented reality display and voice guidance at the terminal, means for recognizing emotions from the user's tone of voice and facial expressions, and means for adjusting the content and tone of guidance according to the emotions. This allows the server to provide flexible support according to the user's emotions even when the user is confused, helping the user understand the operating procedures and quickly solving the problem.
[0673] The "means for receiving user instructions" refers to a mechanism for recognizing instructions given by the user through voice or touch operation, converting them into a format such as text data, and processing them.
[0674] "Means for acquiring camera images and transmitting them to a server" refers to a technology for acquiring video data in real time using a camera and transmitting that data to a server.
[0675] "Means for analyzing received camera footage and identifying appropriate manual information" refers to algorithms or programs that analyze the video data received by the server and identify the most appropriate instructions or manual information based on its content.
[0676] The "means for searching for the specified manual information and providing it to the terminal" is a mechanism by which the server searches the database for the specified manual information and transmits it to the user's terminal for provision.
[0677] "Means for providing information to users through augmented reality display and voice guidance on a terminal" refers to a technology that uses the terminal's display to visually display information in augmented reality (AR) and simultaneously conveys the information to users using voice guidance.
[0678] "Means for recognizing emotions from a user's tone of voice and facial expressions" refers to a program or algorithm that analyzes a user's tone of voice and facial expressions in real time to determine their emotional state.
[0679] "Means for adjusting the content and tone of guidance according to emotions" refers to a mechanism for appropriately adjusting the content and tone of the guidance or explanation provided based on the recognized emotional state of the user.
[0680] "Means for generating answers to user questions based on analyzed camera footage and emotional data" refers to a technology in which the server analyzes camera footage and recognized emotional data, and generates appropriate answers to user questions based on that information.
[0681] "Means for notifying an emergency situation and contacting a manned operator" refers to a mechanism that, when a user operates an emergency button or the like, notifies the server of the situation and allows for real-time contact with a manned operator.
[0682] This invention is a system that receives user instructions, analyzes camera images based on those instructions, and provides appropriate manual information to the user. In addition, by combining it with an emotion engine, it provides flexible and effective support according to the user's emotional state.
[0683] Hardware Configuration
[0684] This system mainly consists of a terminal and a server. The terminal is equipped with a camera, microphone, display, and emotion engine. The server analyzes the camera footage and identifies and provides manual information.
[0685] Software Configuration
[0686] The system integrates Google's speech recognition API, an image recognition model using TensorFlow, and an emotion recognition engine using the Affectiva SDK, enabling multiple functions such as voice recognition to receive user instructions, analyzing camera footage, and recognizing the user's emotional state.
[0687] System Operation
[0688] When a user speaks to the device, the device picks up the voice through its built-in microphone and converts it into text using Google's speech recognition API. This textual instruction is then sent to the server. At the same time, a video of the object the user specified through the camera (for example, the back of a smart TV) is captured in real time and sent to the server. The server then analyzes the received video using TensorFlow and identifies the object within the video. Based on the results, it identifies the appropriate connection procedures and manual information from a database and sends them to the device.
[0689] emotion recognition
[0690] The device is also equipped with an emotion engine using the Affectiva SDK, which recognizes the user's emotions in real time from their tone of voice and facial expressions. For example, if the user is confused, the device will provide appropriate support based on their emotions, such as slowing down the pace of the instructions and using a gentler tone of voice.
[0691] Specific examples
[0692] For example, if a user purchases a new smart TV and doesn't know how to connect it, they can use this system to receive the following support:
[0693] 1. The user points the camera at the back of the smart TV and speaks to the device, saying, "Tell me how to connect to my smart TV."
[0694] 2. The device analyzes the voice, converts the instructions into text, and sends it to the server, along with the camera footage.
[0695] 3. The server analyzes the camera footage, identifies each port on the smart TV, determines the appropriate connection procedure, and sends the manual information to the device.
[0696] 4. The device will use AR to display the appropriate cable connection location on the back of the TV and will provide voice guidance such as "Plug the HDMI cable into the third port from the left."
[0697] 5. While connected, if the user asks, "Where is the power button?", the device sends this question to the server, which uses conversational AI to generate an answer and sends it to the device.
[0698] 6. The device will display AR and provide voice guidance, saying, "The round button on the bottom right is the power button."
[0699] 7. If the emotion engine recognizes that the user is in trouble, it will adjust the guidance and provide a more friendly tone.
[0700] 8. Finally, if users experience difficulties, they can press an emergency button to connect to a live operator for real-time support.
[0701] Prompt Sentence Examples
[0702] "The user gives voice instructions such as, 'Tell me how to connect my smart TV,' and sends the video to the server. The server analyzes the video, identifies the appropriate connection procedure, and provides instructions on how to connect using AR displays and voice guidance. If the user asks, 'Where is the power button?' along the way, the server uses conversational AI to generate an answer, and the device provides guidance to the user. If the user has trouble, they can press the emergency button to connect to an operator. How does the system explain this series of steps?"
[0703] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0704] Step 1: Receive user instructions
[0705] The device receives the user's voice instructions and acquires voice data using the built-in microphone. The input is the user's voice instructions and the output is voice data. The device sends this voice data to Google's speech recognition API, which converts the voice into text data. The input is voice data and the output is text data. This textual instruction is sent to the server. The input is text data and the output is data sent to the server.
[0706] Specific behavior:
[0707] 1. A user says, "How do I connect my smart TV?"
[0708] 2. The device's microphone picks up audio data.
[0709] 3. The voice data acquired by the device is sent to Google's speech recognition API and converted into text data.
[0710] 4. The device sends the text data to the server.
[0711] Step 2: Acquire and transmit video
[0712] The device uses a camera to capture real-time video of an object specified by the user (e.g., the back of a smart TV). The input is the camera image, and the output is real-time video data. This video data is compressed and sent to the server. The input is real-time video data, and the output is data sent to the server.
[0713] Specific behavior:
[0714] 1. The user points the back of the smart TV at the device's camera.
[0715] 2. The device's camera captures video in real time.
[0716] 3. The device compresses the video data and sends it to the server.
[0717] Step 3: Emotion Recognition
[0718] The emotion engine installed on the device recognizes emotions from the user's voice tone and facial expressions. The input is the user's voice tone and facial expression data, and the output is the recognized emotion data. Emotions are analyzed in real time using the Affectiva SDK. The input is voice tone and facial expression data, and the output is emotion data.
[0719] Specific behavior:
[0720] 1. The device's camera and microphone capture the user's facial expressions and voice.
[0721] 2. The Affectiva SDK installed on the device analyzes this data and recognizes emotions.
[0722] Step 4: Video analysis and information identification
[0723] The server analyzes the received camera footage and identifies the appropriate manual information. The input is the camera video data, and the output is the identified manual information. An image recognition model is deployed using TensorFlow to identify objects in the video. The input is the camera video data, and the output is the object recognition results. Based on this result, the appropriate manual information is retrieved from the database. The input is the object recognition results, and the output is the manual information.
[0724] Specific behavior:
[0725] 1. The server analyzes the received video data using a TensorFlow model.
[0726] 2. The server identifies each port on the smart TV from the analysis results.
[0727] 3. The server retrieves the manual information from the database based on the identified information.
[0728] Step 5: Provide manual information
[0729] The server sends the identified manual information to the device. The input is the manual information, and the output is data transmission to the device. The device receives this information and uses AR display technology (e.g., Vuforia) to display the appropriate connection location on the back of the TV and provide voice guidance. The input is the manual information, and the output is the AR display and voice guidance.
[0730] Specific behavior:
[0731] 1. The server sends the specified manual information to the terminal.
[0732] 2. The device uses Vuforia to display an AR display on the back of the TV showing the appropriate connection location.
[0733] 3. The device will say "Please insert the HDMI cable into the third port from the left."
[0734] Step 6: Addressing additional questions
[0735] The server uses conversational AI to generate answers to the user's follow-up questions and sends them to the device. The input is the user's question text and the output is the generated answer. The device converts the question into text using voice recognition and sends it to the server. The input is voice recognition data and the output is text data. The server uses conversational AI to generate an appropriate answer. The input is text data and the output is answer data. The device provides the answer as voice guidance and AR display. The input is answer data and the output is voice guidance and AR display.
[0736] Specific behavior:
[0737] 1. The user asks, "Where is the power button?"
[0738] 2. The device converts the question into text and sends it to the server.
[0739] 3. The server uses OpenAI GPT-3 to generate an appropriate answer and send it to the device.
[0740] 4. The device will announce, "The round button on the bottom right is the power button," and will also show this in AR.
[0741] Step 7: Emergency response
[0742] If the user experiences difficulty, the terminal will connect to a manned operator in real time. The input is the operation of the emergency button, and the output is an emergency notification to the server. When the emergency button on the terminal is pressed, an emergency notification is sent to the server. The input is emergency status data, and the output is data sent to the server. The server receives this information and connects to a manned operator in real time. The input is emergency notification data, and the output is operator connection data.
[0743] Specific behavior:
[0744] 1. The user presses the emergency button.
[0745] 2. The device sends an emergency notification to the server.
[0746] 3. The server connects to a live operator.
[0747] (Application example 2)
[0748] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0749] In conventional work support systems and shopping assistant systems, users often have difficulty understanding specific work procedures or detailed product usage. Furthermore, because they provide uniform support without considering the user's emotional state, they often fail to provide effective support. This can lead to operational errors and lack of understanding, resulting in a poor user experience.
[0750] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0751] In this invention, the server includes means for receiving user instructions, means for acquiring camera images and transmitting them to the server, means for analyzing the received camera images and identifying appropriate manual information, means for searching for the identified manual information and providing it to the terminal, means for providing information to the user through augmented reality display and audio guidance on the terminal, means for providing product usage instructions and features based on product information detected from the camera images, and means for analyzing the user's emotional state and providing flexible and effective support according to the user's emotions. This allows the user to quickly perform operations and tasks while receiving intuitive visual and audio instructions. Furthermore, the emotion recognition function can more appropriately support the user's state, improving the user experience.
[0752] "Means for receiving user instructions" refers to the means by which a user inputs instructions to the system through voice or touch operations.
[0753] "Means for acquiring camera images and transmitting them to a server" refers to means that has the function of capturing images using a camera and transferring the images to a server.
[0754] "Means for analyzing received camera footage and identifying appropriate manual information" refers to the means by which the server performs object recognition and information analysis based on the video data received, and extracts the corresponding manual information.
[0755] The "means for retrieving the specified manual information and providing it to the terminal" refers to a means for retrieving the extracted manual information from the database and transmitting the information to the user's terminal.
[0756] "Means for providing information to users through augmented reality displays and audio guidance on a device" refers to means for providing information to users visually and audibly using AR technology and audio guidance on a device.
[0757] "Means for providing product usage and features based on product information detected from camera footage" refers to means for providing users with information about detailed usage and features of products identified by analyzing camera footage.
[0758] "Means for analyzing the emotional state of a user and providing flexible and effective support according to the emotion" refers to means for analyzing the emotional state of a user using emotion recognition technology and providing appropriate support according to that state.
[0759] "Means for generating an answer to a user's question based on analyzed camera footage" refers to means for generating an answer to a follow-up question from a user based on analyzed video data.
[0760] "Means for notifying an emergency situation and contacting a manned operator" refers to a means for a user to notify the system of a difficult situation and immediately connect to a manned operator.
[0761] This invention is a system that receives user instructions, analyzes camera footage based on those instructions, and provides appropriate manual information to the user. This system links the terminal and server, and further combines it with an emotion engine to provide flexible and effective support according to the user's emotional state.
[0762] 1. Receiving user instructions
[0763] The device receives instructions from the user through voice or touch operations. For example, if the user says, "Tell me how to use this product," the device picks up the voice through a microphone and converts it into text using voice recognition technology. This text instruction is then sent to the server.
[0764] 2. Image acquisition and transmission
[0765] When a user photographs a product with the device's camera, the device captures the camera image in real time and sends it to the server. The captured image is used to understand what the user is looking at and the input status to the system.
[0766] 3. Video analysis and information identification
[0767] The server analyzes the received camera footage and uses image recognition technology to identify objects in the footage, recognizes each detail about the product, and determines the appropriate usage and characteristic information based on the results. This information is obtained from a database.
[0768] 4. Providing manual information
[0769] The server sends the identified manual information to the terminal. The terminal receives this information and provides it to the user through an augmented reality display and voice guidance. For example, the terminal may overlay an animation of the product or operating instructions on the display, and provide voice guidance such as "Please do this part like this."
[0770] 5. Emotion recognition
[0771] The device is equipped with an emotion engine that recognizes emotions in real time from the user's voice and facial expressions. For example, it can analyze the user's emotions, such as whether they are in trouble, from the tone of their voice and facial expression. Based on the emotion recognition results, the tone and pace of the guidance are adjusted.
[0772] 6. Response to additional questions
[0773] If the user has an additional question, they can ask it by voice, for example, "How do I use this button?" The device converts the user's additional question into text using its voice recognition function and sends it to the server. The server uses conversational AI to analyze the question and generate an appropriate answer. The server then sends the generated answer to the device, which then provides the answer to the user in AR display and voice.
[0774] 7. Emergency Response
[0775] If a user encounters a problem, for example, if they are having difficulty operating the device or make a mistake, they can press the emergency button on their device, which will then notify the server of the emergency. The server will receive this information and connect them to a live operator in real time. The operator will then monitor the user's current situation in real time and provide appropriate support.
[0776] Hardware and software used
[0777] 1. Camera: For example, a smartphone camera.
[0778] 2. Microphone: For example, the microphone on your smartphone.
[0779] 3. Displays: smartphone displays, AR glasses, etc.
[0780] 4. Emotion recognition models: EmotionRecognizer, etc.
[0781] 5. Image recognition technology: Image analysis software such as OpenCV.
[0782] 6. Speech recognition technology: SpeechRecognition, etc.
[0783] Specific examples
[0784] If a user purchases a new kitchen gadget at a brick-and-mortar store and doesn't know how to use it, the system can provide the following assistance:
[0785] 1. The user points the camera at a kitchen gadget and gives voice instructions on the device, such as "Tell me how to use this kitchen gadget."
[0786] 2. The device analyzes the voice, converts the instructions into text, and sends it to the server, along with the camera footage.
[0787] 3. The server analyzes the camera image, recognizes the product, determines the appropriate usage procedure, and sends the manual information to the terminal.
[0788] 4. The device uses AR to display how to use the product and provides voice guidance such as "First, do this."
[0789] 5. If the emotion engine recognizes that the user is in trouble, it will adjust the guidance and provide a more friendly tone.
[0790] 6. 24-hour operator connection button for real-time support.
[0791] Prompt Sentence Examples
[0792] "Please identify the product in this image and provide a detailed description of how to use it."
[0793] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0794] Step 1:
[0795] The terminal receives voice instructions from the user. When the user gives a voice instruction such as "Tell me how to use the product," the terminal uses a microphone to capture the voice data and converts it into text data using voice recognition technology. The input is the user's voice data, and the output is text data.
[0796] Step 2:
[0797] The device sends the converted text data to the server. At the same time, it also sends the product video captured by the device's camera to the server. The input is the text data and the video data, and the output is the data sent to the server.
[0798] Step 3:
[0799] The server analyzes the received video data and uses image recognition technology (e.g., OpenCV) to identify objects within the video. Specifically, it identifies each product or part within the video and extracts information such as its position and shape. The input is the video data, and the output is the product information that is the result of the analysis.
[0800] Step 4:
[0801] The server searches the database for corresponding manual information based on the product information. The searched manual information includes product usage instructions and features. The input is product information, and the output is manual information.
[0802] Step 5:
[0803] The server sends the retrieved manual information to the terminal. The input is the manual information, and the output is the data sent to the terminal.
[0804] Step 6:
[0805] The terminal then provides the received manual information to the user. Specifically, it uses augmented reality technology to overlay product usage and features onto the display. It also provides specific instructions such as "Do this part like this" through voice guidance. The input is the manual information, and the output is the information provided to the user.
[0806] Step 7:
[0807] The device analyzes the user's emotional state. Using an emotion engine, it recognizes emotions in real time from the user's tone of voice and facial expressions. The input is audio and video data, and the output is the user's emotional state.
[0808] Step 8:
[0809] The device adjusts the tone and pace of the guidance based on the user's emotional state. For example, if it recognizes that the user is in a difficult situation, it will explain the guidance in a more friendly tone and at a slower pace. The input is the user's emotional state, and the output is the adjusted guidance.
[0810] Step 9:
[0811] If the user asks a follow-up question, the device converts the question into text using voice recognition and sends it back to the server. The server analyzes the question, generates an appropriate answer, and sends it to the device. The device then provides the answer to the user in AR display and voice. The input is the follow-up question and its analysis results, and the output is the generated answer.
[0812] Step 10:
[0813] If a user encounters a problem, for example, if they are having difficulty operating the device or make a mistake, they press the emergency button on their device. The device then notifies the server of the emergency situation, and the server connects to a live operator in real time. The input is the notification of the emergency situation, and the output is a connection to a live operator.
[0814] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0815] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0816] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0817] [Third embodiment]
[0818] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0819] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0820] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0821] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0822] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0823] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0824] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0825] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0826] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0827] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0828] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0829] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0830] The present invention is a system for receiving user instructions, analyzing camera footage based on those instructions, and providing appropriate manual information to the user. This system is composed of a server and a terminal, and aims to provide the user with visual and audio information that can be intuitively understood. Specific embodiments of this system are described below.
[0831] Receiving user instructions
[0832] The device receives instructions from the user through voice or touch operations. For example, if the user says, "Tell me how to connect to my smart TV," the device picks up the voice through a microphone and converts it into text using voice recognition technology. This textual instruction is then sent to the server.
[0833] Video acquisition and transmission
[0834] When a user points the back of a smart TV at the device's camera, the device captures the camera image in real time and sends it to the server. The device is equipped with a camera and a network connection to reliably capture the image.
[0835] Video analysis and information identification
[0836] The server receives the camera footage sent from the device and analyzes it using an image recognition model. The server identifies each port and cable connection in the footage, and based on the results, identifies the appropriate connection procedures and manual information. This information is retrieved from a database.
[0837] Providing manual information
[0838] The server sends the identified manual information to the device. The device receives this information and provides it to the user through AR display and voice guidance. For example, the device display may overlay the corresponding cable connection point on the back of the smart TV, and voice guidance may say, "Plug the HDMI cable into the third port from the left."
[0839] Response to additional questions
[0840] If the user asks a follow-up question, such as "Where is the power button?", the device sends this question to the server. The server uses conversational AI to analyze the question and generate an answer. For example, it might determine that "The round button on the bottom right is the power button" and send this information to the device. The device then presents this information in an AR display and voice guidance.
[0841] Emergency response
[0842] If a user encounters a problem, such as difficulty operating the device or making a mistake, they can press the emergency button on the device, which will then notify the server of the emergency situation. The server receives this information and connects to a live operator in real time. The emergency operator will monitor the user's current situation in real time and provide appropriate support.
[0843] Specific examples
[0844] Here is a concrete example: If a user purchases a new smart TV and does not know how to connect it, they can use this system to receive the following support:
[0845] 1. The user points the camera at the back of the smart TV and gives voice instructions from the device saying, "Tell me how to connect to the smart TV."
[0846] 2. The device analyzes the voice, converts the instructions into text, and sends it to the server, along with the camera footage.
[0847] 3. The server analyzes the camera footage, identifies each port on the smart TV, determines the appropriate connection procedure, and sends the manual information to the device.
[0848] 4. The device will use AR to display the appropriate cable connection location on the back of the TV and will provide voice guidance such as "Plug the HDMI cable into the third port from the left."
[0849] 5. While connected, if the user asks, "Where is the power button?", the device sends this question to the server, which generates an answer and sends it to the device.
[0850] 6. The device will display AR and provide voice guidance, saying, "The round button on the bottom right is the power button."
[0851] In this way, the present invention allows the user to quickly perform operations and tasks while receiving intuitive visual and audio instructions.
[0852] The processing flow will be explained below.
[0853] Step 1:
[0854] The server initializes the system, specifically loading the AI models (conversational AI, image recognition model, and RAG model) and databases, and securing the necessary resources.
[0855] Step 2:
[0856] The device receives user instructions and uses speech recognition to convert user voice instructions, such as "Tell me how to connect my smart TV," into text.
[0857] Step 3:
[0858] The device captures images of the user's surroundings with a camera, which then captures the images in real time and transmits them to a server via a network.
[0859] Step 4:
[0860] The server analyzes the received camera footage and uses image recognition models to identify objects in the footage, as well as each port and connection on the back of the smart TV.
[0861] Step 5:
[0862] Based on the analysis results, the server searches the database for the appropriate connection manual information, formats the found manual information, and sends it to the terminal.
[0863] Step 6:
[0864] The device then provides the received manual information to the user in AR format, overlaying images of the cables corresponding to each port on the smart TV and the connection instructions on the screen.
[0865] Step 7:
[0866] The device uses a voice synthesis function to guide the user, saying, "Please insert the HDMI cable into the third port from the left." The user then follows the instructions.
[0867] Step 8:
[0868] If the user has an additional question, they can ask aloud, for example, "Where is the power button?"
[0869] Step 9:
[0870] The device uses voice recognition to convert the user's follow-up questions into text and sends them to the server, which then uses conversational AI to analyze the questions and generate appropriate answers.
[0871] Step 10:
[0872] The server sends the generated answer to the device, and the device provides the answer to the user in AR display and voice guidance, such as "The round button at the bottom right is the power button."
[0873] Step 11:
[0874] If the user experiences difficulty in operation, they can press the emergency button, and the device will immediately notify the server of the emergency situation.
[0875] Step 12:
[0876] The server receives emergency notifications and connects to live operators, who provide the user's current camera footage and status to provide real-time support.
[0877] This series of processing steps allows the user to smoothly complete operations and learning while receiving visual and audio guidance.
[0878] Example 1
[0879] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0880] Conventional manual information provision systems have the problem that they are difficult for users to understand intuitively and are unable to provide information efficiently. Another problem is that users cannot solve problems themselves, making it difficult to respond quickly in the event of an emergency. To solve these problems, a system is needed that can understand the user's voice instructions, analyze camera footage, provide appropriate manual information, and respond quickly if the user has additional questions or is in trouble.
[0881] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0882] In this invention, the server includes a means for receiving user instructions, a means for converting voice data into text data, a means for acquiring camera images and transmitting them to the server, a means for analyzing the received camera images and identifying appropriate manual information, a means for searching for the identified manual information and providing it to the terminal, and a means for providing information to the user through an augmented reality display and voice guidance on the terminal. This allows for intuitive information provision via visual and audio, making it easier for the user to quickly solve problems. Furthermore, if the user encounters a difficult situation, the server can notify the user of an emergency situation and quickly contact a manned operator.
[0883] A "user" is an entity that uses this system to provide camera images and give voice instructions and perform touch operations.
[0884] A "means for receiving instructions" is a device or software that has the function of recognizing voice or touch operations made by the user and converting that information into text data.
[0885] The "means for converting voice data into text data" is a technology for converting a user's voice into text, and is a function realized by using voice recognition technology.
[0886] The "means for acquiring camera images and transmitting them to a server" refers to a device or software that can acquire real-time video data captured by a user and transmit it to a server.
[0887] The "means for analyzing the received camera footage and identifying appropriate manual information" refers to a technology in which the server uses video analysis technology to analyze the camera footage and search for and identify the manual information based on the analysis results.
[0888] The "means for searching for specified manual information and providing it to the terminal" is a function for searching for specific information from the manual information stored in the database and transmitting that information to the terminal.
[0889] "Means for providing information to users through augmented reality display and voice guidance on a terminal" refers to a function that uses AR technology to visually display information on a terminal and voice guidance to users using voice synthesis technology.
[0890] "Means for generating answers to questions" refers to a function that analyzes questions received from users and generates appropriate answers using conversational AI.
[0891] The "means for notifying an emergency situation and contacting a manned operator" is a function that, when a user faces a difficult situation, notifies the server of an emergency situation from the terminal and the server contacts a manned operator.
[0892] This invention is a system that receives user instructions, analyzes camera images based on those instructions, and provides appropriate manual information to the user. This system is composed of a server and a terminal, and aims to provide information that the user can intuitively understand visually and audibly.
[0893] Receiving user instructions
[0894] The device receives instructions from the user through voice or touch operations. For example, the user may say, "Tell me how to connect to my smart TV." The device's microphone picks up the voice and converts it into text using voice recognition technology (e.g., Google Speech-to-Text API). This textual instruction is then sent to the server.
[0895] Video acquisition and transmission
[0896] When a user points the rear of a smart TV at the device's camera, the device captures the camera image in real time and sends it to the server. The device requires a high-resolution camera and a stable network connection to ensure reliable transmission of the image data to the server.
[0897] Video analysis and information identification
[0898] The server receives the camera images sent from the device and analyzes them using an image recognition model (e.g., TensorFlow image recognition model).The server identifies each port and cable connection in the image and, based on the results, retrieves the appropriate connection procedures and manual information from a database (e.g., MySQL).
[0899] Providing manual information
[0900] The server sends the identified manual information to the device. The device receives this information and uses augmented reality (AR) display technology (e.g., Vuforia SDK) to display the corresponding cable connection location on the back of the TV. It also uses voice guidance technology (e.g., Amazon Polly) to guide the user, saying, "Plug the HDMI cable into the third port from the left."
[0901] Response to additional questions
[0902] If the user asks a follow-up question, such as "Where is the power button?", the device sends this question to the server. The server uses conversational AI (e.g., OpenAI's GPT-3) to analyze the question, generate an answer, and send it to the device. The device then presents this information again in AR displays and voice guidance.
[0903] Emergency response
[0904] If a user experiences difficulty operating the device, for example if they are having difficulty or make a mistake, they can press the emergency button on the device, which will then notify the server of the emergency situation. The server receives this information and connects the user to a live operator, who then monitors the user's current situation and provides appropriate support.
[0905] Specific examples
[0906] Here is a concrete example: If a user purchases a new smart TV and does not know how to connect it, they can use this system to receive the following support:
[0907] 1. The user points the camera at the back of the smart TV and gives a voice command saying, "Tell me how to connect my smart TV."
[0908] 2. The device analyzes the voice, converts the instructions into text, and sends it to the server, while also sending the camera footage.
[0909] 3. The server analyzes the camera footage, identifies each port on the smart TV, retrieves the appropriate connection instructions from the database, and sends the manual information to the device.
[0910] 4. The device will use AR to display the appropriate cable connection location on the back of the TV and will provide voice guidance such as "Plug the HDMI cable into the third port from the left."
[0911] 5. While connected, if the user asks, "Where is the power button?", the device sends this question to the server, which generates an answer and sends it to the device.
[0912] 6. The device will display AR and provide voice guidance, saying, "The round button on the bottom right is the power button."
[0913] Prompt Sentence Examples
[0914] Here are some example prompts to input to a generative AI model:
[0915] "The user gives voice instructions such as, 'Tell me how to connect my smart TV.' The device converts the voice to text using speech recognition technology and sends this information to the server along with the camera footage. The server uses an image recognition model to analyze the camera footage and generate appropriate manual information. The device then provides instructions to the user using AR displays and voice guidance. If the user asks a follow-up question such as, 'Where is the power button?' the device sends this question to the server, and the server uses conversational AI to generate an answer, which is again conveyed to the user using AR displays and voice guidance."
[0916] This prompt allows the generative AI model to understand the system's behavior and generate appropriate instructions.
[0917] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0918] Step 1:
[0919] The user points the camera at the back of the smart TV and gives a voice command such as "Tell me how to connect my smart TV." The user's voice data is acquired as input. The device uses a microphone to capture the user's voice. Next, it calls voice recognition technology (e.g., Google Speech-to-Text API) to convert the voice data into text data. This outputs text data from the voice data. This text data is then transferred to the server, representing the user's command.
[0920] Step 2:
[0921] The device uses a camera to capture real-time video data. As input, video data from the back of the smart TV is captured from the camera. This video data is compressed, encoded, and sent to the server using a network protocol. The output is the video data sent from the device to the server.
[0922] Step 3:
[0923] The server receives video data sent from the device. The input is the image data sent from the device. The server invokes an image recognition model (e.g., TensorFlow's image recognition model) to analyze the video data. This analysis identifies each port or cable connection in the video. The output is the location data of the identified ports or cable connections.
[0924] Step 4:
[0925] The server searches a database (e.g., MySQL) based on the analysis results to obtain the appropriate connection procedures and manual information. The input is the video analysis results. A database query is executed to obtain the connection procedures and manual information. This information is formatted in JSON format or similar and sent to the device. The output is the obtained manual information.
[0926] Step 5:
[0927] The device uses the received manual information to display an augmented reality (AR) image. The input is the manual information sent from the server. The device uses AR display technology (e.g., Vuforia SDK) to display the corresponding cable connection location on the back of the smart TV. It also uses voice guidance technology (e.g., Amazon Polly) to guide the user, saying, "Plug the HDMI cable into the third port from the left." The output is the displayed AR instructions and the voice guidance.
[0928] Step 6:
[0929] The user asks a follow-up question: "Where is the power button?" The user's voice data is again taken as input. The device converts the voice to text and sends the question to the server.
[0930] Step 7:
[0931] The server uses a conversational AI (e.g., OpenAI's GPT-3) to analyze the user's question. The input is the textual question. The AI model analyzes the question and generates an appropriate answer. This answer is sent back to the device. The output is the generated answer.
[0932] Step 8:
[0933] The device provides the received answer to the user through AR display and voice guidance. The input is the answer data sent from the server. The device uses AR display technology to display the location of the power button. It also provides voice guidance saying, "The round button on the bottom right is the power button." The output is the displayed AR instructions and voice guidance.
[0934] Step 9:
[0935] When a user finds themselves in a difficult situation, they press the emergency button on their device. The input is the act of pressing the emergency button. The device notifies the server of the emergency situation. The server receives this information and sends an emergency notification to a manned operator. The operator monitors the user's situation in real time and gives instructions via video or voice call. The output is the operator providing support.
[0936] (Application example 1)
[0937] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0938] In factory maintenance work, there is a challenge for workers to quickly understand the appropriate procedures and perform the work accurately. Also, if a problem occurs during work or if they do not understand the procedures, there are limited means to smoothly obtain support. This reduces work efficiency and, in some cases, can lead to serious mistakes or accidents.
[0939] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0940] In this invention, the server includes a means for receiving user instructions via voice or touch operation, a means for acquiring camera images and transmitting them to the server, a means for analyzing the received camera images and identifying appropriate manual information, a means for searching for the identified manual information and providing it to the terminal, a means for providing information to the user through augmented reality display and audio guidance on the terminal, and a means for providing the user with audio guidance on appropriate procedures and operating methods based on the analysis results. This allows workers to quickly understand and perform appropriate maintenance procedures while receiving intuitive visual and audio instructions. Furthermore, a function for generating answers to user questions using a generative AI model allows for immediate resolution of questions that arise during work.
[0941] "Means for receiving user instructions by voice or touch operation" refers to a system that uses voice recognition technology or a touch sensor to obtain operations or questions from the user as input data.
[0942] The "means for acquiring camera images and transmitting them to a server" is a function for capturing real-time video data using a camera device and transmitting the video data to a remote server via a network.
[0943] "Means for analyzing received camera footage and identifying appropriate manual information" refers to a system that analyzes received video data using image recognition technology and machine learning models, and then searches and retrieves the necessary manuals and procedural information from a database based on the analysis results.
[0944] The "means for searching for specified manual information and providing it to the terminal" is a function for quickly searching for related information in the database based on the analysis results and providing that information to the user's terminal.
[0945] "Means for providing information to users through augmented reality display and voice guidance on a terminal" refers to a system that uses augmented reality technology to overlay visual information on the terminal display and also uses voice synthesis technology to provide voice guidance to users.
[0946] "Means for providing the user with voice guidance on the appropriate procedures and operation methods based on the analysis results" is a function that uses voice synthesis technology to provide instructions on the next operation or procedure that the user should perform based on the analysis results.
[0947] "Means for generating answers to user questions using a generative AI model" refers to a system that uses natural language processing technology and machine learning models to automatically generate and provide appropriate answers to user questions.
[0948] "Means for notifying an emergency situation and contacting a manned operator" is a function that receives an alert from the user in the event of an emergency, notifies the information to a manned support person in real time, and allows direct contact.
[0949] This invention is a system that aims to improve the efficiency of factory maintenance work and support workers. This system consists of a server and a terminal, and allows users to intuitively give instructions by voice or touch operation, and analyzes camera images to provide appropriate manual information.
[0950] System configuration
[0951] The system consists of the following main components:
[0952] 1. Terminal: Equipped with voice recognition and touch input means, camera image acquisition means, augmented reality display and voice guidance functions. Specifically, the terminal is equipped with a microphone, camera, display, speaker, and network connection function.
[0953] 2. Server: Has the means to receive and analyze camera footage. It uses a generative AI model to generate answers to user questions, searches a database for appropriate manual information, and provides it to the device.
[0954] Program processing overview
[0955] 1. Speech recognition and text conversion: The device uses speech recognition technology (specifically, the speech_recognition library) to convert the user's voice instructions into text data. For touch operations, the device uses the touch sensor.
[0956] 2. Acquiring and transmitting camera images: The device's camera is used to acquire real-time images, and the image data is transmitted to the server via the network.
[0957] 3. Video analysis: The received camera footage is analyzed on the server. Image recognition technology and machine learning models (specifically, the BERT-based pre-trained model bert-base-japanese) are used to recognize objects and text in the footage and identify appropriate manual information.
[0958] 4. Providing manual information: The server sends the identified manual information to the terminal, which then provides the information to the user through augmented reality display and voice guidance (using the pyttsx3 library).
[0959] 5. Answer generation using the generative AI model: The server uses the generative AI model to generate answers to the user's follow-up questions and sends the answers to the device.
[0960] Specific examples
[0961] For example, if a worker says, "Tell me how to remove the bearing," the device analyzes the voice and sends the camera footage to the server. The server analyzes the footage, identifies the appropriate procedure, and sends it to the device. The device then displays an augmented reality display and provides voice guidance, saying, "Remove the bolt, then remove the bearing."
[0962] Example prompt sentence:
[0963] "Please explain the proper procedure to remove the bearing from the machine, using the captured image as a reference."
[0964] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0965] Step 1:
[0966] User instruction input
[0967] The user inputs instructions by voice or touch. In the case of voice instructions, voice data is acquired through the device's microphone. In the case of touch operations, tactile information is acquired by the touch sensor.
[0968] Input: Voice or touch data
[0969] Output: Acquired voice data or touch operation data
[0970] Step 2:
[0971] Speech recognition and text conversion
[0972] The device uses speech recognition technology to convert voice data into text data. Specifically, it uses the speech_recognition library to recognize speech and convert it into text.
[0973] Input: Audio data
[0974] Output: Text data
[0975] Specific operation: The device analyzes the acquired voice data, performs voice recognition, and generates text data such as "Please tell me how to remove the bearing."
[0976] Step 3:
[0977] Acquiring and transmitting camera images
[0978] The user captures the object being handled on the camera, and the terminal uses the camera to capture the image in real time and transmits it to the server via the network.
[0979] Input: Camera video data
[0980] Output: Video data sent to the server
[0981] Specific operation: The device operates the camera and streams the captured video data to the server in real time.
[0982] Step 4:
[0983] Video analysis and information identification
[0984] The server analyzes the received video data and identifies appropriate manual information based on the object's condition using image recognition technology and machine learning models (e.g., bert-base-japanese).
[0985] Input: Video data
[0986] Output: Identified manual information
[0987] Specific operation: The server recognizes each object in the video, identifies elements such as "bolt" and "bearing," and then searches for related manual information in a database.
[0988] Step 5:
[0989] Providing manual information
[0990] The server transmits the identified manual information to the terminal, which provides the information to the user through an augmented reality display and voice guidance.
[0991] Input: Identified manual information
[0992] Output: Visual and audio information provided to the user
[0993] Specific operation: The device uses AR technology to superimpose guidelines on the user's display and provides voice guidance such as, "Remove the bolt, then remove the bearing."
[0994] Step 6:
[0995] Responding to additional questions
[0996] If the user asks a follow-up question, the device converts the question into text and sends it to the server, which uses a generative AI model to generate an answer and sends it to the device, which then provides it to the user.
[0997] Input: User's additional question (voice data or touch operation data)
[0998] Output: The answer generated by the server
[0999] Specific behavior: For example, when a user asks, "How do I insert a new bearing?", the server uses a generative AI model to answer, "Insert the bearing in the correct direction and secure it with a bolt."
[1000] Step 7:
[1001] Emergency response
[1002] When a user is in trouble, they can press the emergency button on their device, which will immediately notify the server of the emergency situation, and the server will contact a live operator to provide support.
[1003] Input: Emergency alert signal
[1004] Output: Support by manned operators begins
[1005] Specific operation: When an emergency occurs, the server notifies the monitoring operator, who checks the user's situation in real time and provides appropriate support.
[1006] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1007] This invention is a system that receives user instructions, analyzes camera footage based on those instructions, and provides appropriate manual information to the user.By combining a server, a terminal, and an emotion engine, this system provides more flexible and effective support according to the user's emotional state.
[1008] Receiving user instructions
[1009] The device receives instructions from the user through voice or touch operations. For example, if the user says, "Tell me how to connect to my smart TV," the device picks up the voice through a microphone and converts it into text using voice recognition technology. This textual instruction is then sent to the server.
[1010] Video acquisition and transmission
[1011] When a user points the back of a smart TV at the device's camera, the device captures the camera image in real time and sends it to the server, which uses the captured image to understand what the user is looking at and the input status to the system.
[1012] emotion recognition
[1013] The device is equipped with an emotion engine that recognizes emotions in real time from the user's voice and facial expressions. For example, it can analyze the user's emotions, such as irritation or distress, from the tone of their voice and facial expressions.
[1014] Video analysis and information identification
[1015] The server analyzes the received camera footage and uses image recognition models to identify objects within the footage, recognizes each port and connection on the back of the smart TV, and uses the results to identify the appropriate connection instructions and manual information. This information is retrieved from a database.
[1016] Providing manual information
[1017] The server sends the identified manual information to the device. The device receives this information and provides it to the user through AR display and voice guidance. For example, the device display may overlay the corresponding cable connection point on the back of the smart TV, and voice guidance may say, "Plug the HDMI cable into the third port from the left."
[1018] Using an emotion engine, the content and tone of the instructions can be adjusted depending on the user's emotional state. For example, if the user is confused, the instructions will be slowed down and explained in a gentler tone.
[1019] Response to additional questions
[1020] If the user has an additional question, they can ask aloud, for example, "Where is the power button?"
[1021] The device uses voice recognition to convert the user's follow-up questions into text and sends them to the server, which then uses conversational AI to analyze the questions and generate appropriate answers.
[1022] The server sends the generated answer to the device, and the device provides the answer to the user in AR display and voice, for example, "The round button on the bottom right is the power button."
[1023] Emergency response
[1024] If a user encounters a problem, such as difficulty operating the device or making a mistake, they can press the emergency button on their device, which will then notify the server of the emergency situation. The server receives this information and connects them to a live operator in real time. The emergency operator will monitor the user's current situation in real time and provide appropriate support.
[1025] Specific examples
[1026] Here is a concrete example: If a user purchases a new smart TV and does not know how to connect it, they can use this system to receive the following support:
[1027] 1. The user points the camera at the back of the smart TV and gives voice instructions from the device saying, "Tell me how to connect to the smart TV."
[1028] 2. The device analyzes the voice, converts the instructions into text, and sends it to the server, along with the camera footage.
[1029] 3. The server analyzes the camera footage, identifies each port on the smart TV, determines the appropriate connection procedure, and sends the manual information to the device.
[1030] 4. The device will use AR to display the appropriate cable connection location on the back of the TV and will provide voice guidance such as "Plug the HDMI cable into the third port from the left."
[1031] 5. While connected, if the user asks, "Where is the power button?", the device sends this question to the server, which generates an answer and sends it to the device.
[1032] 6. The device will display AR and provide voice guidance, saying, "The round button on the bottom right is the power button."
[1033] 7. If the emotion engine recognizes that the user is in trouble, it will adjust the guidance and provide a more friendly tone.
[1034] 8. Finally, if users experience difficulties, they can press an emergency button to connect to a live operator for real-time support.
[1035] In this way, the present invention allows users to quickly perform operations and tasks while receiving intuitive visual and audio instructions. In addition, the emotion recognition function can provide more appropriate support for the user's condition, improving the user experience.
[1036] The processing flow will be explained below.
[1037] Step 1:
[1038] The server initializes the system, specifically loading the AI models (conversational AI, image recognition model, and emotion engine) and databases, and securing the necessary resources.
[1039] Step 2:
[1040] The device receives instructions from the user. For example, if the user issues a voice command such as "Tell me how to connect to my smart TV," the device captures the voice through a microphone and converts it into text using voice recognition technology. The converted text data is then sent to the server.
[1041] Step 3:
[1042] The device captures images of the user's surroundings with a camera, and the captured images are sent to a server via a network in real time.
[1043] Step 4:
[1044] The device's built-in emotion engine analyzes the user's voice and facial expressions to recognize their emotional state. For example, it can recognize that the user is irritated based on their voice tone and facial expression.
[1045] Step 5:
[1046] The server analyzes the received camera footage and uses image recognition models to identify objects in the footage, such as the individual ports of a smart TV, and then identifies the appropriate connection procedures and manual information.
[1047] Step 6:
[1048] The server searches the database for the appropriate manual information and sends the identified information, including cable connection procedures and reference images, to the terminal.
[1049] Step 7:
[1050] The device then provides the received manual information to the user through AR displays and voice guidance. For example, the device might overlay the corresponding cable connection on the back of a smart TV and provide voice guidance such as, "Plug the HDMI cable into the third port from the left."
[1051] Step 8:
[1052] The device's emotion engine monitors the user's emotional state and, if it detects that the user is confused, adjusts the content and tone of the instructions, for example, slowing down the pace of the instructions and using a gentler tone.
[1053] Step 9:
[1054] If the user has a follow-up question, they may provide a voice prompt, such as "Where is the power button?"
[1055] Step 10:
[1056] The device uses voice recognition technology to convert follow-up questions into text and send it to the server, which uses conversational AI to analyze the questions and generate appropriate answers.
[1057] Step 11:
[1058] The server sends the generated answer to the device, which then provides it to the user in the form of an AR display and voice guidance, such as "The round button at the bottom right is the power button."
[1059] Step 12:
[1060] If a user experiences difficulty operating the device or needs additional support, they can press the emergency button on the device, which will immediately notify the server of the emergency situation.
[1061] Step 13:
[1062] When the server receives an emergency notification, it connects to a live operator, who is provided with the user's current camera footage and situation in real time.
[1063] Step 14:
[1064] A human operator checks the user's situation in real time and provides appropriate support, providing the instructions necessary to help the user solve the problem in real time.
[1065] Through the above processing steps, the system of the present invention provides visual and audio support to the user, and recognizes the user's emotional state to enable optimal assistance.
[1066] Example 2
[1067] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1068] Conventional manual provision systems often do not provide sufficient support when users have difficulty operating the system, and tasks often stall due to emotional confusion or a lack of understanding of the operating procedures. Furthermore, conventional systems do not support real-time dialogue or emotion recognition, which hinders the user experience.
[1069] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1070] In this invention, the server includes means for receiving user instructions, means for acquiring camera images and transmitting them to the server, means for analyzing the received camera images and identifying appropriate manual information, means for retrieving the identified manual information and providing it to the terminal, means for providing information to the user through augmented reality display and voice guidance at the terminal, means for recognizing emotions from the user's tone of voice and facial expressions, and means for adjusting the content and tone of guidance according to the emotions. This allows the server to provide flexible support according to the user's emotions even when the user is confused, helping the user understand the operating procedures and quickly solving the problem.
[1071] The "means for receiving user instructions" refers to a mechanism for recognizing instructions given by the user through voice or touch operation, converting them into a format such as text data, and processing them.
[1072] "Means for acquiring camera images and transmitting them to a server" refers to a technology for acquiring video data in real time using a camera and transmitting that data to a server.
[1073] "Means for analyzing received camera footage and identifying appropriate manual information" refers to algorithms or programs that analyze the video data received by the server and identify the most appropriate instructions or manual information based on its content.
[1074] The "means for searching for the specified manual information and providing it to the terminal" is a mechanism by which the server searches the database for the specified manual information and transmits it to the user's terminal for provision.
[1075] "Means for providing information to users through augmented reality display and voice guidance on a terminal" refers to a technology that uses the terminal's display to visually display information in augmented reality (AR) and simultaneously conveys the information to users using voice guidance.
[1076] "Means for recognizing emotions from a user's tone of voice and facial expressions" refers to a program or algorithm that analyzes a user's tone of voice and facial expressions in real time to determine their emotional state.
[1077] "Means for adjusting the content and tone of guidance according to emotions" refers to a mechanism for appropriately adjusting the content and tone of the guidance or explanation provided based on the recognized emotional state of the user.
[1078] "Means for generating answers to user questions based on analyzed camera footage and emotional data" refers to a technology in which the server analyzes camera footage and recognized emotional data, and generates appropriate answers to user questions based on that information.
[1079] "Means for notifying an emergency situation and contacting a manned operator" refers to a mechanism that, when a user operates an emergency button or the like, notifies the server of the situation and allows for real-time contact with a manned operator.
[1080] This invention is a system that receives user instructions, analyzes camera images based on those instructions, and provides appropriate manual information to the user. In addition, by combining it with an emotion engine, it provides flexible and effective support according to the user's emotional state.
[1081] Hardware Configuration
[1082] This system mainly consists of a terminal and a server. The terminal is equipped with a camera, microphone, display, and emotion engine. The server analyzes the camera footage and identifies and provides manual information.
[1083] Software Configuration
[1084] The system integrates Google's speech recognition API, an image recognition model using TensorFlow, and an emotion recognition engine using the Affectiva SDK, enabling multiple functions such as voice recognition to receive user instructions, analyzing camera footage, and recognizing the user's emotional state.
[1085] System Operation
[1086] When a user speaks to the device, the device picks up the voice through its built-in microphone and converts it into text using Google's speech recognition API. This textual instruction is then sent to the server. At the same time, a video of the object the user specified through the camera (for example, the back of a smart TV) is captured in real time and sent to the server. The server then analyzes the received video using TensorFlow and identifies the object within the video. Based on the results, it identifies the appropriate connection procedures and manual information from a database and sends them to the device.
[1087] emotion recognition
[1088] The device is also equipped with an emotion engine using the Affectiva SDK, which recognizes the user's emotions in real time from their tone of voice and facial expressions. For example, if the user is confused, the device will provide appropriate support based on their emotions, such as slowing down the pace of the instructions and using a gentler tone of voice.
[1089] Specific examples
[1090] For example, if a user purchases a new smart TV and doesn't know how to connect it, they can use this system to receive the following support:
[1091] 1. The user points the camera at the back of the smart TV and speaks to the device, saying, "Tell me how to connect to my smart TV."
[1092] 2. The device analyzes the voice, converts the instructions into text, and sends it to the server, along with the camera footage.
[1093] 3. The server analyzes the camera footage, identifies each port on the smart TV, determines the appropriate connection procedure, and sends the manual information to the device.
[1094] 4. The device will use AR to display the appropriate cable connection location on the back of the TV and will provide voice guidance such as "Plug the HDMI cable into the third port from the left."
[1095] 5. While connected, if the user asks, "Where is the power button?", the device sends this question to the server, which uses conversational AI to generate an answer and sends it to the device.
[1096] 6. The device will display AR and provide voice guidance, saying, "The round button on the bottom right is the power button."
[1097] 7. If the emotion engine recognizes that the user is in trouble, it will adjust the guidance and provide a more friendly tone.
[1098] 8. Finally, if users experience difficulties, they can press an emergency button to connect to a live operator for real-time support.
[1099] Prompt Sentence Examples
[1100] "The user gives voice instructions such as, 'Tell me how to connect my smart TV,' and sends the video to the server. The server analyzes the video, identifies the appropriate connection procedure, and provides instructions on how to connect using AR displays and voice guidance. If the user asks, 'Where is the power button?' along the way, the server uses conversational AI to generate an answer, and the device provides guidance to the user. If the user has trouble, they can press the emergency button to connect to an operator. How does the system explain this series of steps?"
[1101] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1102] Step 1: Receive user instructions
[1103] The device receives the user's voice instructions and acquires voice data using the built-in microphone. The input is the user's voice instructions and the output is voice data. The device sends this voice data to Google's speech recognition API, which converts the voice into text data. The input is voice data and the output is text data. This textual instruction is sent to the server. The input is text data and the output is data sent to the server.
[1104] Specific behavior:
[1105] 1. A user says, "How do I connect my smart TV?"
[1106] 2. The device's microphone picks up audio data.
[1107] 3. The voice data acquired by the device is sent to Google's speech recognition API and converted into text data.
[1108] 4. The device sends the text data to the server.
[1109] Step 2: Acquire and transmit video
[1110] The device uses a camera to capture real-time video of an object specified by the user (e.g., the back of a smart TV). The input is the camera image, and the output is real-time video data. This video data is compressed and sent to the server. The input is real-time video data, and the output is data sent to the server.
[1111] Specific behavior:
[1112] 1. The user points the back of the smart TV at the device's camera.
[1113] 2. The device's camera captures video in real time.
[1114] 3. The device compresses the video data and sends it to the server.
[1115] Step 3: Emotion Recognition
[1116] The emotion engine installed on the device recognizes emotions from the user's voice tone and facial expressions. The input is the user's voice tone and facial expression data, and the output is the recognized emotion data. Emotions are analyzed in real time using the Affectiva SDK. The input is voice tone and facial expression data, and the output is emotion data.
[1117] Specific behavior:
[1118] 1. The device's camera and microphone capture the user's facial expressions and voice.
[1119] 2. The Affectiva SDK installed on the device analyzes this data and recognizes emotions.
[1120] Step 4: Video analysis and information identification
[1121] The server analyzes the received camera footage and identifies the appropriate manual information. The input is the camera video data, and the output is the identified manual information. An image recognition model is deployed using TensorFlow to identify objects in the video. The input is the camera video data, and the output is the object recognition results. Based on this result, the appropriate manual information is retrieved from the database. The input is the object recognition results, and the output is the manual information.
[1122] Specific behavior:
[1123] 1. The server analyzes the received video data using a TensorFlow model.
[1124] 2. The server identifies each port on the smart TV from the analysis results.
[1125] 3. The server retrieves the manual information from the database based on the identified information.
[1126] Step 5: Provide manual information
[1127] The server sends the identified manual information to the device. The input is the manual information, and the output is data transmission to the device. The device receives this information and uses AR display technology (e.g., Vuforia) to display the appropriate connection location on the back of the TV and provide voice guidance. The input is the manual information, and the output is the AR display and voice guidance.
[1128] Specific behavior:
[1129] 1. The server sends the specified manual information to the terminal.
[1130] 2. The device uses Vuforia to display an AR display on the back of the TV showing the appropriate connection location.
[1131] 3. The device will say "Please insert the HDMI cable into the third port from the left."
[1132] Step 6: Addressing additional questions
[1133] The server uses conversational AI to generate answers to the user's follow-up questions and sends them to the device. The input is the user's question text and the output is the generated answer. The device converts the question into text using voice recognition and sends it to the server. The input is voice recognition data and the output is text data. The server uses conversational AI to generate an appropriate answer. The input is text data and the output is answer data. The device provides the answer as voice guidance and AR display. The input is answer data and the output is voice guidance and AR display.
[1134] Specific behavior:
[1135] 1. The user asks, "Where is the power button?"
[1136] 2. The device converts the question into text and sends it to the server.
[1137] 3. The server uses OpenAI GPT-3 to generate an appropriate answer and send it to the device.
[1138] 4. The device will announce, "The round button on the bottom right is the power button," and will also show this in AR.
[1139] Step 7: Emergency response
[1140] If the user experiences difficulty, the terminal will connect to a manned operator in real time. The input is the operation of the emergency button, and the output is an emergency notification to the server. When the emergency button on the terminal is pressed, an emergency notification is sent to the server. The input is emergency status data, and the output is data sent to the server. The server receives this information and connects to a manned operator in real time. The input is emergency notification data, and the output is operator connection data.
[1141] Specific behavior:
[1142] 1. The user presses the emergency button.
[1143] 2. The device sends an emergency notification to the server.
[1144] 3. The server connects to a live operator.
[1145] (Application example 2)
[1146] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1147] In conventional work support systems and shopping assistant systems, users often have difficulty understanding specific work procedures or detailed product usage. Furthermore, because they provide uniform support without considering the user's emotional state, they often fail to provide effective support. This can lead to operational errors and lack of understanding, resulting in a poor user experience.
[1148] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1149] In this invention, the server includes means for receiving user instructions, means for acquiring camera images and transmitting them to the server, means for analyzing the received camera images and identifying appropriate manual information, means for searching for the identified manual information and providing it to the terminal, means for providing information to the user through augmented reality display and audio guidance on the terminal, means for providing product usage instructions and features based on product information detected from the camera images, and means for analyzing the user's emotional state and providing flexible and effective support according to the user's emotions. This allows the user to quickly perform operations and tasks while receiving intuitive visual and audio instructions. Furthermore, the emotion recognition function can more appropriately support the user's state, improving the user experience.
[1150] "Means for receiving user instructions" refers to the means by which a user inputs instructions to the system through voice or touch operations.
[1151] "Means for acquiring camera images and transmitting them to a server" refers to means that has the function of capturing images using a camera and transferring the images to a server.
[1152] "Means for analyzing received camera footage and identifying appropriate manual information" refers to the means by which the server performs object recognition and information analysis based on the video data received, and extracts the corresponding manual information.
[1153] The "means for retrieving the specified manual information and providing it to the terminal" refers to a means for retrieving the extracted manual information from the database and transmitting the information to the user's terminal.
[1154] "Means for providing information to users through augmented reality displays and audio guidance on a device" refers to means for providing information to users visually and audibly using AR technology and audio guidance on a device.
[1155] "Means for providing product usage and features based on product information detected from camera footage" refers to means for providing users with information about detailed usage and features of products identified by analyzing camera footage.
[1156] "Means for analyzing the emotional state of a user and providing flexible and effective support according to the emotion" refers to means for analyzing the emotional state of a user using emotion recognition technology and providing appropriate support according to that state.
[1157] "Means for generating an answer to a user's question based on analyzed camera footage" refers to means for generating an answer to a follow-up question from a user based on analyzed video data.
[1158] "Means for notifying an emergency situation and contacting a manned operator" refers to a means for a user to notify the system of a difficult situation and immediately connect to a manned operator.
[1159] This invention is a system that receives user instructions, analyzes camera footage based on those instructions, and provides appropriate manual information to the user. This system links the terminal and server, and further combines it with an emotion engine to provide flexible and effective support according to the user's emotional state.
[1160] 1. Receiving user instructions
[1161] The device receives instructions from the user through voice or touch operations. For example, if the user says, "Tell me how to use this product," the device picks up the voice through a microphone and converts it into text using voice recognition technology. This text instruction is then sent to the server.
[1162] 2. Image acquisition and transmission
[1163] When a user photographs a product with the device's camera, the device captures the camera image in real time and sends it to the server. The captured image is used to understand what the user is looking at and the input status to the system.
[1164] 3. Video analysis and information identification
[1165] The server analyzes the received camera footage and uses image recognition technology to identify objects in the footage, recognizes each detail about the product, and determines the appropriate usage and characteristic information based on the results. This information is obtained from a database.
[1166] 4. Providing manual information
[1167] The server sends the identified manual information to the terminal. The terminal receives this information and provides it to the user through an augmented reality display and voice guidance. For example, the terminal may overlay an animation of the product or operating instructions on the display, and provide voice guidance such as "Please do this part like this."
[1168] 5. Emotion recognition
[1169] The device is equipped with an emotion engine that recognizes emotions in real time from the user's voice and facial expressions. For example, it can analyze the user's emotions, such as whether they are in trouble, from the tone of their voice and facial expression. Based on the emotion recognition results, the tone and pace of the guidance are adjusted.
[1170] 6. Response to additional questions
[1171] If the user has an additional question, they can ask it by voice, for example, "How do I use this button?" The device converts the user's additional question into text using its voice recognition function and sends it to the server. The server uses conversational AI to analyze the question and generate an appropriate answer. The server then sends the generated answer to the device, which then provides the answer to the user in AR display and voice.
[1172] 7. Emergency Response
[1173] If a user encounters a problem, for example, if they are having difficulty operating the device or make a mistake, they can press the emergency button on their device, which will then notify the server of the emergency. The server will receive this information and connect them to a live operator in real time. The operator will then monitor the user's current situation in real time and provide appropriate support.
[1174] Hardware and software used
[1175] 1. Camera: For example, a smartphone camera.
[1176] 2. Microphone: For example, the microphone on your smartphone.
[1177] 3. Displays: smartphone displays, AR glasses, etc.
[1178] 4. Emotion recognition models: EmotionRecognizer, etc.
[1179] 5. Image recognition technology: Image analysis software such as OpenCV.
[1180] 6. Speech recognition technology: SpeechRecognition, etc.
[1181] Specific examples
[1182] If a user purchases a new kitchen gadget at a brick-and-mortar store and doesn't know how to use it, the system can provide the following assistance:
[1183] 1. The user points the camera at a kitchen gadget and gives voice instructions on the device, such as "Tell me how to use this kitchen gadget."
[1184] 2. The device analyzes the voice, converts the instructions into text, and sends it to the server, along with the camera footage.
[1185] 3. The server analyzes the camera image, recognizes the product, determines the appropriate usage procedure, and sends the manual information to the terminal.
[1186] 4. The device uses AR to display how to use the product and provides voice guidance such as "First, do this."
[1187] 5. If the emotion engine recognizes that the user is in trouble, it will adjust the guidance and provide a more friendly tone.
[1188] 6. 24-hour operator connection button for real-time support.
[1189] Prompt Sentence Examples
[1190] "Please identify the product in this image and provide a detailed description of how to use it."
[1191] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1192] Step 1:
[1193] The terminal receives voice instructions from the user. When the user gives a voice instruction such as "Tell me how to use the product," the terminal uses a microphone to capture the voice data and converts it into text data using voice recognition technology. The input is the user's voice data, and the output is text data.
[1194] Step 2:
[1195] The device sends the converted text data to the server. At the same time, it also sends the product video captured by the device's camera to the server. The input is the text data and the video data, and the output is the data sent to the server.
[1196] Step 3:
[1197] The server analyzes the received video data and uses image recognition technology (e.g., OpenCV) to identify objects within the video. Specifically, it identifies each product or part within the video and extracts information such as its position and shape. The input is the video data, and the output is the product information that is the result of the analysis.
[1198] Step 4:
[1199] The server searches the database for corresponding manual information based on the product information. The searched manual information includes product usage instructions and features. The input is product information, and the output is manual information.
[1200] Step 5:
[1201] The server sends the retrieved manual information to the terminal. The input is the manual information, and the output is the data sent to the terminal.
[1202] Step 6:
[1203] The terminal then provides the received manual information to the user. Specifically, it uses augmented reality technology to overlay product usage and features onto the display. It also provides specific instructions such as "Do this part like this" through voice guidance. The input is the manual information, and the output is the information provided to the user.
[1204] Step 7:
[1205] The device analyzes the user's emotional state. Using an emotion engine, it recognizes emotions in real time from the user's tone of voice and facial expressions. The input is audio and video data, and the output is the user's emotional state.
[1206] Step 8:
[1207] The device adjusts the tone and pace of the guidance based on the user's emotional state. For example, if it recognizes that the user is in a difficult situation, it will explain the guidance in a more friendly tone and at a slower pace. The input is the user's emotional state, and the output is the adjusted guidance.
[1208] Step 9:
[1209] If the user asks a follow-up question, the device converts the question into text using voice recognition and sends it back to the server. The server analyzes the question, generates an appropriate answer, and sends it to the device. The device then provides the answer to the user in AR display and voice. The input is the follow-up question and its analysis results, and the output is the generated answer.
[1210] Step 10:
[1211] If a user encounters a problem, for example, if they are having difficulty operating the device or make a mistake, they press the emergency button on their device. The device then notifies the server of the emergency situation, and the server connects to a live operator in real time. The input is the notification of the emergency situation, and the output is a connection to a live operator.
[1212] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1213] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1214] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1215] [Fourth embodiment]
[1216] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1217] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1218] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1219] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1220] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1221] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1222] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1223] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1224] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1225] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1226] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1227] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1228] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1229] The present invention is a system for receiving user instructions, analyzing camera footage based on those instructions, and providing appropriate manual information to the user. This system is composed of a server and a terminal, and aims to provide the user with visual and audio information that can be intuitively understood. Specific embodiments of this system are described below.
[1230] Receiving user instructions
[1231] The device receives instructions from the user through voice or touch operations. For example, if the user says, "Tell me how to connect to my smart TV," the device picks up the voice through a microphone and converts it into text using voice recognition technology. This textual instruction is then sent to the server.
[1232] Video acquisition and transmission
[1233] When a user points the back of a smart TV at the device's camera, the device captures the camera image in real time and sends it to the server. The device is equipped with a camera and a network connection to reliably capture the image.
[1234] Video analysis and information identification
[1235] The server receives the camera footage sent from the device and analyzes it using an image recognition model. The server identifies each port and cable connection in the footage, and based on the results, identifies the appropriate connection procedures and manual information. This information is retrieved from a database.
[1236] Providing manual information
[1237] The server sends the identified manual information to the device. The device receives this information and provides it to the user through AR display and voice guidance. For example, the device display may overlay the corresponding cable connection point on the back of the smart TV, and voice guidance may say, "Plug the HDMI cable into the third port from the left."
[1238] Response to additional questions
[1239] If the user asks a follow-up question, such as "Where is the power button?", the device sends this question to the server. The server uses conversational AI to analyze the question and generate an answer. For example, it might determine that "The round button on the bottom right is the power button" and send this information to the device. The device then presents this information in an AR display and voice guidance.
[1240] Emergency response
[1241] If a user encounters a problem, such as difficulty operating the device or making a mistake, they can press the emergency button on the device, which will then notify the server of the emergency situation. The server receives this information and connects to a live operator in real time. The emergency operator will monitor the user's current situation in real time and provide appropriate support.
[1242] Specific examples
[1243] Here is a concrete example: If a user purchases a new smart TV and does not know how to connect it, they can use this system to receive the following support:
[1244] 1. The user points the camera at the back of the smart TV and gives voice instructions from the device saying, "Tell me how to connect to the smart TV."
[1245] 2. The device analyzes the voice, converts the instructions into text, and sends it to the server, along with the camera footage.
[1246] 3. The server analyzes the camera footage, identifies each port on the smart TV, determines the appropriate connection procedure, and sends the manual information to the device.
[1247] 4. The device will use AR to display the appropriate cable connection location on the back of the TV and will provide voice guidance such as "Plug the HDMI cable into the third port from the left."
[1248] 5. While connected, if the user asks, "Where is the power button?", the device sends this question to the server, which generates an answer and sends it to the device.
[1249] 6. The device will display AR and provide voice guidance, saying, "The round button on the bottom right is the power button."
[1250] In this way, the present invention allows the user to quickly perform operations and tasks while receiving intuitive visual and audio instructions.
[1251] The processing flow will be explained below.
[1252] Step 1:
[1253] The server initializes the system, specifically loading the AI models (conversational AI, image recognition model, and RAG model) and databases, and securing the necessary resources.
[1254] Step 2:
[1255] The device receives user instructions and uses speech recognition to convert user voice instructions, such as "Tell me how to connect my smart TV," into text.
[1256] Step 3:
[1257] The device captures images of the user's surroundings with a camera, which then captures the images in real time and transmits them to a server via a network.
[1258] Step 4:
[1259] The server analyzes the received camera footage and uses image recognition models to identify objects in the footage, as well as each port and connection on the back of the smart TV.
[1260] Step 5:
[1261] Based on the analysis results, the server searches the database for the appropriate connection manual information, formats the found manual information, and sends it to the terminal.
[1262] Step 6:
[1263] The device then provides the received manual information to the user in AR format, overlaying images of the cables corresponding to each port on the smart TV and the connection instructions on the screen.
[1264] Step 7:
[1265] The device uses a voice synthesis function to guide the user, saying, "Please insert the HDMI cable into the third port from the left." The user then follows the instructions.
[1266] Step 8:
[1267] If the user has an additional question, they can ask aloud, for example, "Where is the power button?"
[1268] Step 9:
[1269] The device uses voice recognition to convert the user's follow-up questions into text and sends them to the server, which then uses conversational AI to analyze the questions and generate appropriate answers.
[1270] Step 10:
[1271] The server sends the generated answer to the device, and the device provides the answer to the user in AR display and voice guidance, such as "The round button at the bottom right is the power button."
[1272] Step 11:
[1273] If the user experiences difficulty in operation, they can press the emergency button, and the device will immediately notify the server of the emergency situation.
[1274] Step 12:
[1275] The server receives emergency notifications and connects to live operators, who provide the user's current camera footage and status to provide real-time support.
[1276] This series of processing steps allows the user to smoothly complete operations and learning while receiving visual and audio guidance.
[1277] Example 1
[1278] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1279] Conventional manual information provision systems have the problem that they are difficult for users to understand intuitively and are unable to provide information efficiently. Another problem is that users cannot solve problems themselves, making it difficult to respond quickly in the event of an emergency. To solve these problems, a system is needed that can understand the user's voice instructions, analyze camera footage, provide appropriate manual information, and respond quickly if the user has additional questions or is in trouble.
[1280] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1281] In this invention, the server includes a means for receiving user instructions, a means for converting voice data into text data, a means for acquiring camera images and transmitting them to the server, a means for analyzing the received camera images and identifying appropriate manual information, a means for searching for the identified manual information and providing it to the terminal, and a means for providing information to the user through an augmented reality display and voice guidance on the terminal. This allows for intuitive information provision via visual and audio, making it easier for the user to quickly solve problems. Furthermore, if the user encounters a difficult situation, the server can notify the user of an emergency situation and quickly contact a manned operator.
[1282] A "user" is an entity that uses this system to provide camera images and give voice instructions and perform touch operations.
[1283] A "means for receiving instructions" is a device or software that has the function of recognizing voice or touch operations made by the user and converting that information into text data.
[1284] The "means for converting voice data into text data" is a technology for converting a user's voice into text, and is a function realized by using voice recognition technology.
[1285] The "means for acquiring camera images and transmitting them to a server" refers to a device or software that can acquire real-time video data captured by a user and transmit it to a server.
[1286] The "means for analyzing the received camera footage and identifying appropriate manual information" refers to a technology in which the server uses video analysis technology to analyze the camera footage and search for and identify the manual information based on the analysis results.
[1287] The "means for searching for specified manual information and providing it to the terminal" is a function for searching for specific information from the manual information stored in the database and transmitting that information to the terminal.
[1288] "Means for providing information to users through augmented reality display and voice guidance on a terminal" refers to a function that uses AR technology to visually display information on a terminal and voice guidance to users using voice synthesis technology.
[1289] "Means for generating answers to questions" refers to a function that analyzes questions received from users and generates appropriate answers using conversational AI.
[1290] The "means for notifying an emergency situation and contacting a manned operator" is a function that, when a user faces a difficult situation, notifies the server of an emergency situation from the terminal and the server contacts a manned operator.
[1291] This invention is a system that receives user instructions, analyzes camera images based on those instructions, and provides appropriate manual information to the user. This system is composed of a server and a terminal, and aims to provide information that the user can intuitively understand visually and audibly.
[1292] Receiving user instructions
[1293] The device receives instructions from the user through voice or touch operations. For example, the user may say, "Tell me how to connect to my smart TV." The device's microphone picks up the voice and converts it into text using voice recognition technology (e.g., Google Speech-to-Text API). This textual instruction is then sent to the server.
[1294] Video acquisition and transmission
[1295] When a user points the rear of a smart TV at the device's camera, the device captures the camera image in real time and sends it to the server. The device requires a high-resolution camera and a stable network connection to ensure reliable transmission of the image data to the server.
[1296] Video analysis and information identification
[1297] The server receives the camera images sent from the device and analyzes them using an image recognition model (e.g., TensorFlow image recognition model).The server identifies each port and cable connection in the image and, based on the results, retrieves the appropriate connection procedures and manual information from a database (e.g., MySQL).
[1298] Providing manual information
[1299] The server sends the identified manual information to the device. The device receives this information and uses augmented reality (AR) display technology (e.g., Vuforia SDK) to display the corresponding cable connection location on the back of the TV. It also uses voice guidance technology (e.g., Amazon Polly) to guide the user, saying, "Plug the HDMI cable into the third port from the left."
[1300] Response to additional questions
[1301] If the user asks a follow-up question, such as "Where is the power button?", the device sends this question to the server. The server uses conversational AI (e.g., OpenAI's GPT-3) to analyze the question, generate an answer, and send it to the device. The device then presents this information again in AR displays and voice guidance.
[1302] Emergency response
[1303] If a user experiences difficulty operating the device, for example if they are having difficulty or make a mistake, they can press the emergency button on the device, which will then notify the server of the emergency situation. The server receives this information and connects the user to a live operator, who then monitors the user's current situation and provides appropriate support.
[1304] Specific examples
[1305] Here is a concrete example: If a user purchases a new smart TV and does not know how to connect it, they can use this system to receive the following support:
[1306] 1. The user points the camera at the back of the smart TV and gives a voice command saying, "Tell me how to connect my smart TV."
[1307] 2. The device analyzes the voice, converts the instructions into text, and sends it to the server, while also sending the camera footage.
[1308] 3. The server analyzes the camera footage, identifies each port on the smart TV, retrieves the appropriate connection instructions from the database, and sends the manual information to the device.
[1309] 4. The device will use AR to display the appropriate cable connection location on the back of the TV and will provide voice guidance such as "Plug the HDMI cable into the third port from the left."
[1310] 5. While connected, if the user asks, "Where is the power button?", the device sends this question to the server, which generates an answer and sends it to the device.
[1311] 6. The device will display AR and provide voice guidance, saying, "The round button on the bottom right is the power button."
[1312] Prompt Sentence Examples
[1313] Here are some example prompts to input to a generative AI model:
[1314] "The user gives voice instructions such as, 'Tell me how to connect my smart TV.' The device converts the voice to text using speech recognition technology and sends this information to the server along with the camera footage. The server uses an image recognition model to analyze the camera footage and generate appropriate manual information. The device then provides instructions to the user using AR displays and voice guidance. If the user asks a follow-up question such as, 'Where is the power button?' the device sends this question to the server, and the server uses conversational AI to generate an answer, which is again conveyed to the user using AR displays and voice guidance."
[1315] This prompt allows the generative AI model to understand the system's behavior and generate appropriate instructions.
[1316] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1317] Step 1:
[1318] The user points the camera at the back of the smart TV and gives a voice command such as "Tell me how to connect my smart TV." The user's voice data is acquired as input. The device uses a microphone to capture the user's voice. Next, it calls voice recognition technology (e.g., Google Speech-to-Text API) to convert the voice data into text data. This outputs text data from the voice data. This text data is then transferred to the server, representing the user's command.
[1319] Step 2:
[1320] The device uses a camera to capture real-time video data. As input, video data from the back of the smart TV is captured from the camera. This video data is compressed, encoded, and sent to the server using a network protocol. The output is the video data sent from the device to the server.
[1321] Step 3:
[1322] The server receives video data sent from the device. The input is the image data sent from the device. The server invokes an image recognition model (e.g., TensorFlow's image recognition model) to analyze the video data. This analysis identifies each port or cable connection in the video. The output is the location data of the identified ports or cable connections.
[1323] Step 4:
[1324] The server searches a database (e.g., MySQL) based on the analysis results to obtain the appropriate connection procedures and manual information. The input is the video analysis results. A database query is executed to obtain the connection procedures and manual information. This information is formatted in JSON format or similar and sent to the device. The output is the obtained manual information.
[1325] Step 5:
[1326] The device uses the received manual information to display an augmented reality (AR) image. The input is the manual information sent from the server. The device uses AR display technology (e.g., Vuforia SDK) to display the corresponding cable connection location on the back of the smart TV. It also uses voice guidance technology (e.g., Amazon Polly) to guide the user, saying, "Plug the HDMI cable into the third port from the left." The output is the displayed AR instructions and the voice guidance.
[1327] Step 6:
[1328] The user asks a follow-up question: "Where is the power button?" The user's voice data is again taken as input. The device converts the voice to text and sends the question to the server.
[1329] Step 7:
[1330] The server uses a conversational AI (e.g., OpenAI's GPT-3) to analyze the user's question. The input is the textual question. The AI model analyzes the question and generates an appropriate answer. This answer is sent back to the device. The output is the generated answer.
[1331] Step 8:
[1332] The device provides the received answer to the user through AR display and voice guidance. The input is the answer data sent from the server. The device uses AR display technology to display the location of the power button. It also provides voice guidance saying, "The round button on the bottom right is the power button." The output is the displayed AR instructions and voice guidance.
[1333] Step 9:
[1334] When a user finds themselves in a difficult situation, they press the emergency button on their device. The input is the act of pressing the emergency button. The device notifies the server of the emergency situation. The server receives this information and sends an emergency notification to a manned operator. The operator monitors the user's situation in real time and gives instructions via video or voice call. The output is the operator providing support.
[1335] (Application example 1)
[1336] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1337] In factory maintenance work, there is a challenge for workers to quickly understand the appropriate procedures and perform the work accurately. Also, if a problem occurs during work or if they do not understand the procedures, there are limited means to smoothly obtain support. This reduces work efficiency and, in some cases, can lead to serious mistakes or accidents.
[1338] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1339] In this invention, the server includes a means for receiving user instructions via voice or touch operation, a means for acquiring camera images and transmitting them to the server, a means for analyzing the received camera images and identifying appropriate manual information, a means for searching for the identified manual information and providing it to the terminal, a means for providing information to the user through augmented reality display and audio guidance on the terminal, and a means for providing the user with audio guidance on appropriate procedures and operating methods based on the analysis results. This allows workers to quickly understand and perform appropriate maintenance procedures while receiving intuitive visual and audio instructions. Furthermore, a function for generating answers to user questions using a generative AI model allows for immediate resolution of questions that arise during work.
[1340] "Means for receiving user instructions by voice or touch operation" refers to a system that uses voice recognition technology or a touch sensor to obtain operations or questions from the user as input data.
[1341] The "means for acquiring camera images and transmitting them to a server" is a function for capturing real-time video data using a camera device and transmitting the video data to a remote server via a network.
[1342] "Means for analyzing received camera footage and identifying appropriate manual information" refers to a system that analyzes received video data using image recognition technology and machine learning models, and then searches and retrieves the necessary manuals and procedural information from a database based on the analysis results.
[1343] The "means for searching for specified manual information and providing it to the terminal" is a function for quickly searching for related information in the database based on the analysis results and providing that information to the user's terminal.
[1344] "Means for providing information to users through augmented reality display and voice guidance on a terminal" refers to a system that uses augmented reality technology to overlay visual information on the terminal display and also uses voice synthesis technology to provide voice guidance to users.
[1345] "Means for providing the user with voice guidance on the appropriate procedures and operation methods based on the analysis results" is a function that uses voice synthesis technology to provide instructions on the next operation or procedure that the user should perform based on the analysis results.
[1346] "Means for generating answers to user questions using a generative AI model" refers to a system that uses natural language processing technology and machine learning models to automatically generate and provide appropriate answers to user questions.
[1347] "Means for notifying an emergency situation and contacting a manned operator" is a function that receives an alert from the user in the event of an emergency, notifies the information to a manned support person in real time, and allows direct contact.
[1348] This invention is a system that aims to improve the efficiency of factory maintenance work and support workers. This system consists of a server and a terminal, and allows users to intuitively give instructions by voice or touch operation, and analyzes camera images to provide appropriate manual information.
[1349] System configuration
[1350] The system consists of the following main components:
[1351] 1. Terminal: Equipped with voice recognition and touch input means, camera image acquisition means, augmented reality display and voice guidance functions. Specifically, the terminal is equipped with a microphone, camera, display, speaker, and network connection function.
[1352] 2. Server: Has the means to receive and analyze camera footage. It uses a generative AI model to generate answers to user questions, searches a database for appropriate manual information, and provides it to the device.
[1353] Program processing overview
[1354] 1. Speech recognition and text conversion: The device uses speech recognition technology (specifically, the speech_recognition library) to convert the user's voice instructions into text data. For touch operations, the device uses the touch sensor.
[1355] 2. Acquiring and transmitting camera images: The device's camera is used to acquire real-time images, and the image data is transmitted to the server via the network.
[1356] 3. Video analysis: The received camera footage is analyzed on the server. Image recognition technology and machine learning models (specifically, the BERT-based pre-trained model bert-base-japanese) are used to recognize objects and text in the footage and identify appropriate manual information.
[1357] 4. Providing manual information: The server sends the identified manual information to the terminal, which then provides the information to the user through augmented reality display and voice guidance (using the pyttsx3 library).
[1358] 5. Answer generation using the generative AI model: The server uses the generative AI model to generate answers to the user's follow-up questions and sends the answers to the device.
[1359] Specific examples
[1360] For example, if a worker says, "Tell me how to remove the bearing," the device analyzes the voice and sends the camera footage to the server. The server analyzes the footage, identifies the appropriate procedure, and sends it to the device. The device then displays an augmented reality display and provides voice guidance, saying, "Remove the bolt, then remove the bearing."
[1361] Example prompt sentence:
[1362] "Please explain the proper procedure to remove the bearing from the machine, using the captured image as a reference."
[1363] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1364] Step 1:
[1365] User instruction input
[1366] The user inputs instructions by voice or touch. In the case of voice instructions, voice data is acquired through the device's microphone. In the case of touch operations, tactile information is acquired by the touch sensor.
[1367] Input: Voice or touch data
[1368] Output: Acquired voice data or touch operation data
[1369] Step 2:
[1370] Speech recognition and text conversion
[1371] The device uses speech recognition technology to convert voice data into text data. Specifically, it uses the speech_recognition library to recognize speech and convert it into text.
[1372] Input: Audio data
[1373] Output: Text data
[1374] Specific operation: The device analyzes the acquired voice data, performs voice recognition, and generates text data such as "Please tell me how to remove the bearing."
[1375] Step 3:
[1376] Acquiring and transmitting camera images
[1377] The user captures the object being handled on the camera, and the terminal uses the camera to capture the image in real time and transmits it to the server via the network.
[1378] Input: Camera video data
[1379] Output: Video data sent to the server
[1380] Specific operation: The device operates the camera and streams the captured video data to the server in real time.
[1381] Step 4:
[1382] Video analysis and information identification
[1383] The server analyzes the received video data and identifies appropriate manual information based on the object's condition using image recognition technology and machine learning models (e.g., bert-base-japanese).
[1384] Input: Video data
[1385] Output: Identified manual information
[1386] Specific operation: The server recognizes each object in the video, identifies elements such as "bolt" and "bearing," and then searches for related manual information in a database.
[1387] Step 5:
[1388] Providing manual information
[1389] The server transmits the identified manual information to the terminal, which provides the information to the user through an augmented reality display and voice guidance.
[1390] Input: Identified manual information
[1391] Output: Visual and audio information provided to the user
[1392] Specific operation: The device uses AR technology to superimpose guidelines on the user's display and provides voice guidance such as, "Remove the bolt, then remove the bearing."
[1393] Step 6:
[1394] Responding to additional questions
[1395] If the user asks a follow-up question, the device converts the question into text and sends it to the server, which uses a generative AI model to generate an answer and sends it to the device, which then provides it to the user.
[1396] Input: User's additional question (voice data or touch operation data)
[1397] Output: The answer generated by the server
[1398] Specific behavior: For example, when a user asks, "How do I insert a new bearing?", the server uses a generative AI model to answer, "Insert the bearing in the correct direction and secure it with a bolt."
[1399] Step 7:
[1400] Emergency response
[1401] When a user is in trouble, they can press the emergency button on their device, which will immediately notify the server of the emergency situation, and the server will contact a live operator to provide support.
[1402] Input: Emergency alert signal
[1403] Output: Support by manned operators begins
[1404] Specific operation: When an emergency occurs, the server notifies the monitoring operator, who checks the user's situation in real time and provides appropriate support.
[1405] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1406] This invention is a system that receives user instructions, analyzes camera footage based on those instructions, and provides appropriate manual information to the user.By combining a server, a terminal, and an emotion engine, this system provides more flexible and effective support according to the user's emotional state.
[1407] Receiving user instructions
[1408] The device receives instructions from the user through voice or touch operations. For example, if the user says, "Tell me how to connect to my smart TV," the device picks up the voice through a microphone and converts it into text using voice recognition technology. This textual instruction is then sent to the server.
[1409] Video acquisition and transmission
[1410] When a user points the back of a smart TV at the device's camera, the device captures the camera image in real time and sends it to the server, which uses the captured image to understand what the user is looking at and the input status to the system.
[1411] emotion recognition
[1412] The device is equipped with an emotion engine that recognizes emotions in real time from the user's voice and facial expressions. For example, it can analyze the user's emotions, such as irritation or distress, from the tone of their voice and facial expressions.
[1413] Video analysis and information identification
[1414] The server analyzes the received camera footage and uses image recognition models to identify objects within the footage, recognizes each port and connection on the back of the smart TV, and uses the results to identify the appropriate connection instructions and manual information. This information is retrieved from a database.
[1415] Providing manual information
[1416] The server sends the identified manual information to the device. The device receives this information and provides it to the user through AR display and voice guidance. For example, the device display may overlay the corresponding cable connection point on the back of the smart TV, and voice guidance may say, "Plug the HDMI cable into the third port from the left."
[1417] Using an emotion engine, the content and tone of the instructions can be adjusted depending on the user's emotional state. For example, if the user is confused, the instructions will be slowed down and explained in a gentler tone.
[1418] Response to additional questions
[1419] If the user has an additional question, they can ask aloud, for example, "Where is the power button?"
[1420] The device uses voice recognition to convert the user's follow-up questions into text and sends them to the server, which then uses conversational AI to analyze the questions and generate appropriate answers.
[1421] The server sends the generated answer to the device, and the device provides the answer to the user in AR display and voice, for example, "The round button on the bottom right is the power button."
[1422] Emergency response
[1423] If a user encounters a problem, such as difficulty operating the device or making a mistake, they can press the emergency button on their device, which will then notify the server of the emergency situation. The server receives this information and connects them to a live operator in real time. The emergency operator will monitor the user's current situation in real time and provide appropriate support.
[1424] Specific examples
[1425] Here is a concrete example: If a user purchases a new smart TV and does not know how to connect it, they can use this system to receive the following support:
[1426] 1. The user points the camera at the back of the smart TV and gives voice instructions from the device saying, "Tell me how to connect to the smart TV."
[1427] 2. The device analyzes the voice, converts the instructions into text, and sends it to the server, along with the camera footage.
[1428] 3. The server analyzes the camera footage, identifies each port on the smart TV, determines the appropriate connection procedure, and sends the manual information to the device.
[1429] 4. The device will use AR to display the appropriate cable connection location on the back of the TV and will provide voice guidance such as "Plug the HDMI cable into the third port from the left."
[1430] 5. While connected, if the user asks, "Where is the power button?", the device sends this question to the server, which generates an answer and sends it to the device.
[1431] 6. The device will display AR and provide voice guidance, saying, "The round button on the bottom right is the power button."
[1432] 7. If the emotion engine recognizes that the user is in trouble, it will adjust the guidance and provide a more friendly tone.
[1433] 8. Finally, if users experience difficulties, they can press an emergency button to connect to a live operator for real-time support.
[1434] In this way, the present invention allows users to quickly perform operations and tasks while receiving intuitive visual and audio instructions. In addition, the emotion recognition function can provide more appropriate support for the user's condition, improving the user experience.
[1435] The processing flow will be explained below.
[1436] Step 1:
[1437] The server initializes the system, specifically loading the AI models (conversational AI, image recognition model, and emotion engine) and databases, and securing the necessary resources.
[1438] Step 2:
[1439] The device receives instructions from the user. For example, if the user issues a voice command such as "Tell me how to connect to my smart TV," the device captures the voice through a microphone and converts it into text using voice recognition technology. The converted text data is then sent to the server.
[1440] Step 3:
[1441] The device captures images of the user's surroundings with a camera, and the captured images are sent to a server via a network in real time.
[1442] Step 4:
[1443] The device's built-in emotion engine analyzes the user's voice and facial expressions to recognize their emotional state. For example, it can recognize that the user is irritated based on their voice tone and facial expression.
[1444] Step 5:
[1445] The server analyzes the received camera footage and uses image recognition models to identify objects in the footage, such as the individual ports of a smart TV, and then identifies the appropriate connection procedures and manual information.
[1446] Step 6:
[1447] The server searches the database for the appropriate manual information and sends the identified information, including cable connection procedures and reference images, to the terminal.
[1448] Step 7:
[1449] The device then provides the received manual information to the user through AR displays and voice guidance. For example, the device might overlay the corresponding cable connection on the back of a smart TV and provide voice guidance such as, "Plug the HDMI cable into the third port from the left."
[1450] Step 8:
[1451] The device's emotion engine monitors the user's emotional state and, if it detects that the user is confused, adjusts the content and tone of the instructions, for example, slowing down the pace of the instructions and using a gentler tone.
[1452] Step 9:
[1453] If the user has a follow-up question, they may provide a voice prompt, such as "Where is the power button?"
[1454] Step 10:
[1455] The device uses voice recognition technology to convert follow-up questions into text and send it to the server, which uses conversational AI to analyze the questions and generate appropriate answers.
[1456] Step 11:
[1457] The server sends the generated answer to the device, which then provides it to the user in the form of an AR display and voice guidance, such as "The round button at the bottom right is the power button."
[1458] Step 12:
[1459] If a user experiences difficulty operating the device or needs additional support, they can press the emergency button on the device, which will immediately notify the server of the emergency situation.
[1460] Step 13:
[1461] When the server receives an emergency notification, it connects to a live operator, who is provided with the user's current camera footage and situation in real time.
[1462] Step 14:
[1463] A human operator checks the user's situation in real time and provides appropriate support, providing the instructions necessary to help the user solve the problem in real time.
[1464] Through the above processing steps, the system of the present invention provides visual and audio support to the user, and recognizes the user's emotional state to enable optimal assistance.
[1465] Example 2
[1466] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1467] Conventional manual provision systems often do not provide sufficient support when users have difficulty operating the system, and tasks often stall due to emotional confusion or a lack of understanding of the operating procedures. Furthermore, conventional systems do not support real-time dialogue or emotion recognition, which hinders the user experience.
[1468] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1469] In this invention, the server includes means for receiving user instructions, means for acquiring camera images and transmitting them to the server, means for analyzing the received camera images and identifying appropriate manual information, means for retrieving the identified manual information and providing it to the terminal, means for providing information to the user through augmented reality display and voice guidance at the terminal, means for recognizing emotions from the user's tone of voice and facial expressions, and means for adjusting the content and tone of guidance according to the emotions. This allows the server to provide flexible support according to the user's emotions even when the user is confused, helping the user understand the operating procedures and quickly solving the problem.
[1470] The "means for receiving user instructions" refers to a mechanism for recognizing instructions given by the user through voice or touch operation, converting them into a format such as text data, and processing them.
[1471] "Means for acquiring camera images and transmitting them to a server" refers to a technology for acquiring video data in real time using a camera and transmitting that data to a server.
[1472] "Means for analyzing received camera footage and identifying appropriate manual information" refers to algorithms or programs that analyze the video data received by the server and identify the most appropriate instructions or manual information based on its content.
[1473] The "means for searching for the specified manual information and providing it to the terminal" is a mechanism by which the server searches the database for the specified manual information and transmits it to the user's terminal for provision.
[1474] "Means for providing information to users through augmented reality display and voice guidance on a terminal" refers to a technology that uses the terminal's display to visually display information in augmented reality (AR) and simultaneously conveys the information to users using voice guidance.
[1475] "Means for recognizing emotions from a user's tone of voice and facial expressions" refers to a program or algorithm that analyzes a user's tone of voice and facial expressions in real time to determine their emotional state.
[1476] "Means for adjusting the content and tone of guidance according to emotions" refers to a mechanism for appropriately adjusting the content and tone of the guidance or explanation provided based on the recognized emotional state of the user.
[1477] "Means for generating answers to user questions based on analyzed camera footage and emotional data" refers to a technology in which the server analyzes camera footage and recognized emotional data, and generates appropriate answers to user questions based on that information.
[1478] "Means for notifying an emergency situation and contacting a manned operator" refers to a mechanism that, when a user operates an emergency button or the like, notifies the server of the situation and allows for real-time contact with a manned operator.
[1479] This invention is a system that receives user instructions, analyzes camera images based on those instructions, and provides appropriate manual information to the user. In addition, by combining it with an emotion engine, it provides flexible and effective support according to the user's emotional state.
[1480] Hardware Configuration
[1481] This system mainly consists of a terminal and a server. The terminal is equipped with a camera, microphone, display, and emotion engine. The server analyzes the camera footage and identifies and provides manual information.
[1482] Software Configuration
[1483] The system integrates Google's speech recognition API, an image recognition model using TensorFlow, and an emotion recognition engine using the Affectiva SDK, enabling multiple functions such as voice recognition to receive user instructions, analyzing camera footage, and recognizing the user's emotional state.
[1484] System Operation
[1485] When a user speaks to the device, the device picks up the voice through its built-in microphone and converts it into text using Google's speech recognition API. This textual instruction is then sent to the server. At the same time, a video of the object the user specified through the camera (for example, the back of a smart TV) is captured in real time and sent to the server. The server then analyzes the received video using TensorFlow and identifies the object within the video. Based on the results, it identifies the appropriate connection procedures and manual information from a database and sends them to the device.
[1486] emotion recognition
[1487] The device is also equipped with an emotion engine using the Affectiva SDK, which recognizes the user's emotions in real time from their tone of voice and facial expressions. For example, if the user is confused, the device will provide appropriate support based on their emotions, such as slowing down the pace of the instructions and using a gentler tone of voice.
[1488] Specific examples
[1489] For example, if a user purchases a new smart TV and doesn't know how to connect it, they can use this system to receive the following support:
[1490] 1. The user points the camera at the back of the smart TV and speaks to the device, saying, "Tell me how to connect to my smart TV."
[1491] 2. The device analyzes the voice, converts the instructions into text, and sends it to the server, along with the camera footage.
[1492] 3. The server analyzes the camera footage, identifies each port on the smart TV, determines the appropriate connection procedure, and sends the manual information to the device.
[1493] 4. The device will use AR to display the appropriate cable connection location on the back of the TV and will provide voice guidance such as "Plug the HDMI cable into the third port from the left."
[1494] 5. While connected, if the user asks, "Where is the power button?", the device sends this question to the server, which uses conversational AI to generate an answer and sends it to the device.
[1495] 6. The device will display AR and provide voice guidance, saying, "The round button on the bottom right is the power button."
[1496] 7. If the emotion engine recognizes that the user is in trouble, it will adjust the guidance and provide a more friendly tone.
[1497] 8. Finally, if users experience difficulties, they can press an emergency button to connect to a live operator for real-time support.
[1498] Prompt Sentence Examples
[1499] "The user gives voice instructions such as, 'Tell me how to connect my smart TV,' and sends the video to the server. The server analyzes the video, identifies the appropriate connection procedure, and provides instructions on how to connect using AR displays and voice guidance. If the user asks, 'Where is the power button?' along the way, the server uses conversational AI to generate an answer, and the device provides guidance to the user. If the user has trouble, they can press the emergency button to connect to an operator. How does the system explain this series of steps?"
[1500] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1501] Step 1: Receive user instructions
[1502] The device receives the user's voice instructions and acquires voice data using the built-in microphone. The input is the user's voice instructions and the output is voice data. The device sends this voice data to Google's speech recognition API, which converts the voice into text data. The input is voice data and the output is text data. This textual instruction is sent to the server. The input is text data and the output is data sent to the server.
[1503] Specific behavior:
[1504] 1. A user says, "How do I connect my smart TV?"
[1505] 2. The device's microphone picks up audio data.
[1506] 3. The voice data acquired by the device is sent to Google's speech recognition API and converted into text data.
[1507] 4. The device sends the text data to the server.
[1508] Step 2: Acquire and transmit video
[1509] The device uses a camera to capture real-time video of an object specified by the user (e.g., the back of a smart TV). The input is the camera image, and the output is real-time video data. This video data is compressed and sent to the server. The input is real-time video data, and the output is data sent to the server.
[1510] Specific behavior:
[1511] 1. The user points the back of the smart TV at the device's camera.
[1512] 2. The device's camera captures video in real time.
[1513] 3. The device compresses the video data and sends it to the server.
[1514] Step 3: Emotion Recognition
[1515] The emotion engine installed on the device recognizes emotions from the user's voice tone and facial expressions. The input is the user's voice tone and facial expression data, and the output is the recognized emotion data. Emotions are analyzed in real time using the Affectiva SDK. The input is voice tone and facial expression data, and the output is emotion data.
[1516] Specific behavior:
[1517] 1. The device's camera and microphone capture the user's facial expressions and voice.
[1518] 2. The Affectiva SDK installed on the device analyzes this data and recognizes emotions.
[1519] Step 4: Video analysis and information identification
[1520] The server analyzes the received camera footage and identifies the appropriate manual information. The input is the camera video data, and the output is the identified manual information. An image recognition model is deployed using TensorFlow to identify objects in the video. The input is the camera video data, and the output is the object recognition results. Based on this result, the appropriate manual information is retrieved from the database. The input is the object recognition results, and the output is the manual information.
[1521] Specific behavior:
[1522] 1. The server analyzes the received video data using a TensorFlow model.
[1523] 2. The server identifies each port on the smart TV from the analysis results.
[1524] 3. The server retrieves the manual information from the database based on the identified information.
[1525] Step 5: Provide manual information
[1526] The server sends the identified manual information to the device. The input is the manual information, and the output is data transmission to the device. The device receives this information and uses AR display technology (e.g., Vuforia) to display the appropriate connection location on the back of the TV and provide voice guidance. The input is the manual information, and the output is the AR display and voice guidance.
[1527] Specific behavior:
[1528] 1. The server sends the specified manual information to the terminal.
[1529] 2. The device uses Vuforia to display an AR display on the back of the TV showing the appropriate connection location.
[1530] 3. The device will say "Please insert the HDMI cable into the third port from the left."
[1531] Step 6: Addressing additional questions
[1532] The server uses conversational AI to generate answers to the user's follow-up questions and sends them to the device. The input is the user's question text and the output is the generated answer. The device converts the question into text using voice recognition and sends it to the server. The input is voice recognition data and the output is text data. The server uses conversational AI to generate an appropriate answer. The input is text data and the output is answer data. The device provides the answer as voice guidance and AR display. The input is answer data and the output is voice guidance and AR display.
[1533] Specific behavior:
[1534] 1. The user asks, "Where is the power button?"
[1535] 2. The device converts the question into text and sends it to the server.
[1536] 3. The server uses OpenAI GPT-3 to generate an appropriate answer and send it to the device.
[1537] 4. The device will announce, "The round button on the bottom right is the power button," and will also show this in AR.
[1538] Step 7: Emergency response
[1539] If the user experiences difficulty, the terminal will connect to a manned operator in real time. The input is the operation of the emergency button, and the output is an emergency notification to the server. When the emergency button on the terminal is pressed, an emergency notification is sent to the server. The input is emergency status data, and the output is data sent to the server. The server receives this information and connects to a manned operator in real time. The input is emergency notification data, and the output is operator connection data.
[1540] Specific behavior:
[1541] 1. The user presses the emergency button.
[1542] 2. The device sends an emergency notification to the server.
[1543] 3. The server connects to a live operator.
[1544] (Application example 2)
[1545] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1546] In conventional work support systems and shopping assistant systems, users often have difficulty understanding specific work procedures or detailed product usage. Furthermore, because they provide uniform support without considering the user's emotional state, they often fail to provide effective support. This can lead to operational errors and lack of understanding, resulting in a poor user experience.
[1547] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1548] In this invention, the server includes means for receiving user instructions, means for acquiring camera images and transmitting them to the server, means for analyzing the received camera images and identifying appropriate manual information, means for searching for the identified manual information and providing it to the terminal, means for providing information to the user through augmented reality display and audio guidance on the terminal, means for providing product usage instructions and features based on product information detected from the camera images, and means for analyzing the user's emotional state and providing flexible and effective support according to the user's emotions. This allows the user to quickly perform operations and tasks while receiving intuitive visual and audio instructions. Furthermore, the emotion recognition function can more appropriately support the user's state, improving the user experience.
[1549] "Means for receiving user instructions" refers to the means by which a user inputs instructions to the system through voice or touch operations.
[1550] "Means for acquiring camera images and transmitting them to a server" refers to means that has the function of capturing images using a camera and transferring the images to a server.
[1551] "Means for analyzing received camera footage and identifying appropriate manual information" refers to the means by which the server performs object recognition and information analysis based on the video data received, and extracts the corresponding manual information.
[1552] The "means for retrieving the specified manual information and providing it to the terminal" refers to a means for retrieving the extracted manual information from the database and transmitting the information to the user's terminal.
[1553] "Means for providing information to users through augmented reality displays and audio guidance on a device" refers to means for providing information to users visually and audibly using AR technology and audio guidance on a device.
[1554] "Means for providing product usage and features based on product information detected from camera footage" refers to means for providing users with information about detailed usage and features of products identified by analyzing camera footage.
[1555] "Means for analyzing the emotional state of a user and providing flexible and effective support according to the emotion" refers to means for analyzing the emotional state of a user using emotion recognition technology and providing appropriate support according to that state.
[1556] "Means for generating an answer to a user's question based on analyzed camera footage" refers to means for generating an answer to a follow-up question from a user based on analyzed video data.
[1557] "Means for notifying an emergency situation and contacting a manned operator" refers to a means for a user to notify the system of a difficult situation and immediately connect to a manned operator.
[1558] This invention is a system that receives user instructions, analyzes camera footage based on those instructions, and provides appropriate manual information to the user. This system links the terminal and server, and further combines it with an emotion engine to provide flexible and effective support according to the user's emotional state.
[1559] 1. Receiving user instructions
[1560] The device receives instructions from the user through voice or touch operations. For example, if the user says, "Tell me how to use this product," the device picks up the voice through a microphone and converts it into text using voice recognition technology. This text instruction is then sent to the server.
[1561] 2. Image acquisition and transmission
[1562] When a user photographs a product with the device's camera, the device captures the camera image in real time and sends it to the server. The captured image is used to understand what the user is looking at and the input status to the system.
[1563] 3. Video analysis and information identification
[1564] The server analyzes the received camera footage and uses image recognition technology to identify objects in the footage, recognizes each detail about the product, and determines the appropriate usage and characteristic information based on the results. This information is obtained from a database.
[1565] 4. Providing manual information
[1566] The server sends the identified manual information to the terminal. The terminal receives this information and provides it to the user through an augmented reality display and voice guidance. For example, the terminal may overlay an animation of the product or operating instructions on the display, and provide voice guidance such as "Please do this part like this."
[1567] 5. Emotion recognition
[1568] The device is equipped with an emotion engine that recognizes emotions in real time from the user's voice and facial expressions. For example, it can analyze the user's emotions, such as whether they are in trouble, from the tone of their voice and facial expression. Based on the emotion recognition results, the tone and pace of the guidance are adjusted.
[1569] 6. Response to additional questions
[1570] If the user has an additional question, they can ask it by voice, for example, "How do I use this button?" The device converts the user's additional question into text using its voice recognition function and sends it to the server. The server uses conversational AI to analyze the question and generate an appropriate answer. The server then sends the generated answer to the device, which then provides the answer to the user in AR display and voice.
[1571] 7. Emergency Response
[1572] If a user encounters a problem, for example, if they are having difficulty operating the device or make a mistake, they can press the emergency button on their device, which will then notify the server of the emergency. The server will receive this information and connect them to a live operator in real time. The operator will then monitor the user's current situation in real time and provide appropriate support.
[1573] Hardware and software used
[1574] 1. Camera: For example, a smartphone camera.
[1575] 2. Microphone: For example, the microphone on your smartphone.
[1576] 3. Displays: smartphone displays, AR glasses, etc.
[1577] 4. Emotion recognition models: EmotionRecognizer, etc.
[1578] 5. Image recognition technology: Image analysis software such as OpenCV.
[1579] 6. Speech recognition technology: SpeechRecognition, etc.
[1580] Specific examples
[1581] If a user purchases a new kitchen gadget at a brick-and-mortar store and doesn't know how to use it, the system can provide the following assistance:
[1582] 1. The user points the camera at a kitchen gadget and gives voice instructions on the device, such as "Tell me how to use this kitchen gadget."
[1583] 2. The device analyzes the voice, converts the instructions into text, and sends it to the server, along with the camera footage.
[1584] 3. The server analyzes the camera image, recognizes the product, determines the appropriate usage procedure, and sends the manual information to the terminal.
[1585] 4. The device uses AR to display how to use the product and provides voice guidance such as "First, do this."
[1586] 5. If the emotion engine recognizes that the user is in trouble, it will adjust the guidance and provide a more friendly tone.
[1587] 6. 24-hour operator connection button for real-time support.
[1588] Prompt Sentence Examples
[1589] "Please identify the product in this image and provide a detailed description of how to use it."
[1590] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1591] Step 1:
[1592] The terminal receives voice instructions from the user. When the user gives a voice instruction such as "Tell me how to use the product," the terminal uses a microphone to capture the voice data and converts it into text data using voice recognition technology. The input is the user's voice data, and the output is text data.
[1593] Step 2:
[1594] The device sends the converted text data to the server. At the same time, it also sends the product video captured by the device's camera to the server. The input is the text data and the video data, and the output is the data sent to the server.
[1595] Step 3:
[1596] The server analyzes the received video data and uses image recognition technology (e.g., OpenCV) to identify objects within the video. Specifically, it identifies each product or part within the video and extracts information such as its position and shape. The input is the video data, and the output is the product information that is the result of the analysis.
[1597] Step 4:
[1598] The server searches the database for corresponding manual information based on the product information. The searched manual information includes product usage instructions and features. The input is product information, and the output is manual information.
[1599] Step 5:
[1600] The server sends the retrieved manual information to the terminal. The input is the manual information, and the output is the data sent to the terminal.
[1601] Step 6:
[1602] The terminal then provides the received manual information to the user. Specifically, it uses augmented reality technology to overlay product usage and features onto the display. It also provides specific instructions such as "Do this part like this" through voice guidance. The input is the manual information, and the output is the information provided to the user.
[1603] Step 7:
[1604] The device analyzes the user's emotional state. Using an emotion engine, it recognizes emotions in real time from the user's tone of voice and facial expressions. The input is audio and video data, and the output is the user's emotional state.
[1605] Step 8:
[1606] The device adjusts the tone and pace of the guidance based on the user's emotional state. For example, if it recognizes that the user is in a difficult situation, it will explain the guidance in a more friendly tone and at a slower pace. The input is the user's emotional state, and the output is the adjusted guidance.
[1607] Step 9:
[1608] If the user asks a follow-up question, the device converts the question into text using voice recognition and sends it back to the server. The server analyzes the question, generates an appropriate answer, and sends it to the device. The device then provides the answer to the user in AR display and voice. The input is the follow-up question and its analysis results, and the output is the generated answer.
[1609] Step 10:
[1610] If a user encounters a problem, for example, if they are having difficulty operating the device or make a mistake, they press the emergency button on their device. The device then notifies the server of the emergency situation, and the server connects to a live operator in real time. The input is the notification of the emergency situation, and the output is a connection to a live operator.
[1611] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1612] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1613] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1614] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1615] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1616] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm. 【16...
Claims
1. means for receiving user instructions; A means for acquiring camera images and transmitting them to a server; means for analyzing the received camera footage and identifying appropriate manual information; A means for searching for the specified manual information and providing it to the terminal; a means for providing information to a user through AR display and voice guidance in the terminal; A system including:
2. The system of claim 1 further comprising means for generating an answer to a user's question based on the analyzed camera footage.
3. 2. The system of claim 1, further comprising means for notifying an emergency situation and contacting a live operator when a user is in distress.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A