System

The system addresses the challenge of understanding complex home appliance manuals by digitizing and summarizing content, using augmented reality for visual guidance, and personalizing advice based on user interactions, improving operational efficiency and user satisfaction.

JP2026028028APending Publication Date: 2026-02-19SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024130326
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-06
Publication Date
2026-02-19

AI Technical Summary

Technical Problem

Users find it difficult to understand instruction manuals for sophisticated home appliances, and there are limited ways to quickly obtain operating instructions and troubleshooting information, leading to inefficiencies and a lack of personalized operation and setting suggestions.

Method used

A system that digitizes instruction manuals using optical character recognition, summarizes text data with natural language processing, provides visual guidance through augmented reality, and uses a support chatbot to answer questions, while recording user interactions to learn preferences and habits for personalized advice.

Benefits of technology

Enables users to intuitively and efficiently operate home appliances by providing personalized operation advice and troubleshooting information, enhancing user experience and utilization of appliance features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026028028000001_ABST
    Figure 2026028028000001_ABST
Patent Text Reader

Abstract

To provide a system capable of digitizing an instruction manual of a home electric appliance, allowing a user to intuitively operate it, and quickly acquiring necessary information.SOLUTION: A means for acquiring instruction book data and extracting text data by optical character recognition, a means for applying a natural language processing model to the extracted text data and summarizing the text data, a means for selecting a model of a home appliance through a user interface and providing a related operation method and troubleshooting information, a means for scanning the home appliance using a camera and visualizing an operation procedure by an augmented reality technology, a means for answering a question from a user by a support chatbot using the natural language processing model, a means for recording an operation history and a setting change of the user and learning a preference and a use habit of the user using a machine learning algorithm, and a means for proposing individual operation advice and a setting change on the basis of a learning result.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] As home appliances become increasingly sophisticated, many users find it difficult to understand instruction manuals. Furthermore, there are limited ways to quickly obtain operating instructions and troubleshooting information, forcing users to take the time to contact support centers. Furthermore, there are few cases where optimal operation and settings tailored to individual users' preferences and usage habits are provided, resulting in the issue of not fully utilizing the convenience of home appliances. [Means for solving the problem]

[0005] To solve the above-mentioned problems, the present invention provides a system as follows. First, a home appliance instruction manual is acquired in digital format and text data is extracted using optical character recognition. Next, the extracted text data is summarized using a natural language processing model. The user selects the home appliance model through a user interface and is provided with related operating instructions and troubleshooting information. Furthermore, the home appliance is scanned using a camera, and operating procedures are visualized using augmented reality technology. A support chatbot using a natural language processing model quickly and accurately answers user questions. Furthermore, the system records the user's operation history and setting changes, and uses a machine learning algorithm to learn the user's preferences and usage habits. It then provides personalized operating advice and setting change suggestions, thereby providing the user with the optimal operating method. This allows the user to intuitively operate the home appliance and quickly obtain the necessary information.

[0006] "Home appliances" refers to all appliances and devices used in the home that operate using electricity.

[0007] "Instruction manual" refers to a document that includes information on how to use the product, an explanation of each part, troubleshooting information, etc.

[0008] "Digitalization" refers to the conversion of information into electronic form.

[0009] "Optical character recognition" refers to the technology that converts printed text or handwritten characters into digital data.

[0010] A "natural language processing model" refers to algorithms and technologies that allow computers to understand and generate human language.

[0011] "User interface" refers to the screens and methods of operation that allow a user to interact with a computer or application.

[0012] "Troubleshooting Information" refers to information that describes how to solve product problems and how to deal with errors.

[0013] "Augmented reality technology (AR)" refers to the technology of overlaying digital information onto images of the real world.

[0014] A "support chatbot" is a software program that automatically answers users' questions.

[0015] "Operation history" refers to a record of operations performed by a user.

[0016] "Change settings" refers to changing the product's operating conditions and display content according to the user's wishes.

[0017] "Machine learning algorithms" refer to algorithms that self-improve and make predictions and decisions based on data.

[0018] "Preferences" refer to the tendencies and patterns that users prefer.

[0019] "Usage habits" refers to consistent patterns of behavior when a user uses a product.

[0020] "Operation advice" refers to suggestions for optimal operation methods and settings. [Brief explanation of the drawings]

[0021] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0022] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0023] First, the terms used in the following description will be explained.

[0024] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0025] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0026] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0027] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0028] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0029] [First embodiment]

[0030] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0031] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0032] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0033] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0034] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0035] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0036] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0037] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0038] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0039] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0040] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0041] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0042] This invention is a home appliance instruction manual system that utilizes generative AI and augmented reality (AR) technology, allowing users to intuitively and quickly obtain operation methods and troubleshooting information for home appliances. The invention can be implemented as follows through the roles of a server, a terminal, and a user.

[0043] Overall system overview

[0044] This system acquires digital versions of instruction manuals for home appliances and extracts text data using optical character recognition (OCR) technology. The extracted text data is summarized using a natural language processing (NLP) model and provided as information to assist users in their operations. It also utilizes AR technology to enable users to visually understand product operation procedures using their smartphones. Furthermore, a support chatbot responds to user inquiries, providing a means for users to instantly obtain the information they need.

[0045] Program processing overview

[0046] 1. The server retrieves the instruction manual data for the home appliance from the database. The retrieved data is converted into text data using OCR technology and summarized using an NLP model. This converts the long instructions into a short, easy-to-understand format.

[0047] 2. The device launches the application and prompts the user to select the model of the home appliance through the user interface. Based on the selected model, the device communicates with the server to obtain relevant operating instructions and troubleshooting information.

[0048] 3. The device allows users to scan home appliances using their smartphone camera, and uses AR technology to overlay operating instructions on the camera image, providing users with a visual guide.

[0049] 4. The user enters a question using the chatbot function within the application. The chatbot uses an NLP model to analyze the question and generate an appropriate answer.

[0050] 5. The server records the user's operation history and setting changes, and uses a machine learning algorithm to learn the user's preferences and usage habits. Based on this learning, it generates individualized operation advice and setting change suggestions and sends them to the device.

[0051] Specific Examples

[0052] The following is a specific scenario.

[0053] Changing the air conditioner remote control settings

[0054] 1. The user opens the application and selects the air conditioner model.

[0055] 2. The device requests the selected model information from the server, and the server returns the operation instructions and setting data for the corresponding model.

[0056] 3. The device instructs the user to scan the air conditioner remote control with the camera, and overlays operating instructions on the scanned image of the remote control, allowing the user to visually confirm the specific button operations.

[0057] 4. When a user asks the chatbot, "I don't know how to set it up," the server analyzes the question, generates detailed instructions, and sends them to the device.

[0058] 5. The device displays the generated instructions on the chatbot screen, and the user follows them to complete the setup.

[0059] In this way, the system intuitively and efficiently provides users with operation and troubleshooting information for home appliances, making it easier to understand instruction manuals and providing optimal operation advice tailored to the user's individual preferences and usage habits.

[0060] The processing flow will be explained below.

[0061] Step 1:

[0062] The server queries the database to retrieve the instruction manual (PDF format) for the home appliance, and stores the retrieved PDF file in local storage or temporarily in memory.

[0063] Step 2:

[0064] The server uses OCR (Optical Character Recognition) technology to extract text data from the PDF. Using an OCR library such as Tesseract, it analyzes all pages in the PDF and converts them into text data.

[0065] Step 3:

[0066] The server then runs the extracted text data through a natural language processing (NLP) model to summarize it, using models like BERT and GPT to reduce redundant descriptions and extract key information.

[0067] Step 4:

[0068] The device launches the application and displays the user interface, initially displaying a search bar and drop-down lists for selecting appliance categories and models.

[0069] Step 5:

[0070] The user selects the model of the home appliance they are using on the interface, and the selected model information is sent to the server, requesting related operating instructions and troubleshooting information.

[0071] Step 6:

[0072] Based on the received model information, the server retrieves relevant operating instructions and troubleshooting information and sends it to the device. The information is retrieved from a database and appropriately formatted before being sent.

[0073] Step 7:

[0074] The device responds to user requests and displays the received instructional and troubleshooting information, appropriately laid out in the app's interface.

[0075] Step 8:

[0076] The device will activate the camera, allowing the user to scan home appliances, capture camera footage in real time, and provide guidance to the user.

[0077] Step 9:

[0078] The device uses augmented reality (AR) technology to overlay operation instructions on the camera image, using libraries such as ARKit and ARCore to overlay instruction icons and text on the operation panel.

[0079] Step 10:

[0080] Users operate home appliances by following the instructions on the camera footage, and follow the on-screen instructions and guides to complete the required operations.

[0081] Step 11:

[0082] Users enter questions into the in-app chatbot, using natural language, and submit the question.

[0083] Step 12:

[0084] The server analyzes the questions received from users through a natural language processing model, interprets the input text to understand its meaning, and generates an appropriate answer.

[0085] Step 13:

[0086] The server sends the generated answer to the terminal, where it is formatted appropriately and displayed on the chatbot screen.

[0087] Step 14:

[0088] The device displays the received response on the chatbot's conversation screen, and the user can confirm the displayed information and perform any necessary operations or settings.

[0089] Step 15:

[0090] The device records the user's operation history and setting changes, and sends the operation details and timestamps to the server as log data.

[0091] Step 16:

[0092] The server compiles the received operation history and setting change data and uses machine learning algorithms to learn the user's preferences and usage habits.

[0093] Step 17:

[0094] The server generates individualized operation advice and setting change suggestions based on the learning results, and sends the suggestions to the device as a notification.

[0095] Step 18:

[0096] The device will notify the user of the received suggestions, either as a push notification or a pop-up message, suggesting new settings or operation methods to the user.

[0097] Example 1

[0098] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0099] Conventional paper manuals for home appliance operation and troubleshooting information are often difficult for users to understand and take a long time to understand. Furthermore, manuals are often lost, making it difficult to quickly obtain the necessary information. Furthermore, when operation methods or solutions are complex, users can easily become lost. There is a need for a system that can solve these problems and enable users to quickly and intuitively understand and operate home appliances.

[0100] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0101] In this invention, the server includes means for acquiring instruction manual data for the home appliance and extracting text data by optical character recognition, means for summarizing the extracted text data through a natural language processing model, means for providing the contents of the instruction manual in a format that is easy for the user to understand using a generative AI model, and means for prompting the user to ask specific questions using prompt sentences and generating detailed answers in response to the questions, thereby enabling the user to quickly and intuitively obtain complex operating instructions and troubleshooting information.

[0102] "Home appliances" refers to all household electronic devices that users use on a daily basis.

[0103] An "instruction manual" refers to a document that describes how to operate, use, maintain, and troubleshoot a home appliance.

[0104] "Digitalization" refers to the process of converting data in paper or physical form into electronic form.

[0105] "User interface" refers to the screen and input devices that users use to operate the device.

[0106] "Optical character recognition (OCR)" refers to the technology that identifies character information from image data and converts it into digital text.

[0107] "Text data" refers to data that is stored as character information and can be processed electronically.

[0108] A "natural language processing (NLP) model" refers to an artificial intelligence model for understanding, analyzing, and generating human language.

[0109] An "abstract" is a concise summary of the contents of a longer document.

[0110] "Augmented reality (AR) technology" refers to the technology of overlaying digital information onto images of the real world.

[0111] A "support chatbot" is a program that automatically responds to questions from users.

[0112] A "machine learning algorithm" refers to a method that allows a computer to learn on its own based on data and make predictions and classifications.

[0113] A "generative AI model" refers to an artificial intelligence model that generates new sentences and answers based on human language.

[0114] A "prompt sentence" refers to a question-type sentence that guides the user's input.

[0115] This invention is a home appliance instruction manual system that utilizes generative AI models and augmented reality (AR) technology, allowing users to intuitively and quickly obtain operating instructions and troubleshooting information for home appliances. This invention is realized through the roles of a server, a terminal, and a user.

[0116] Overall system overview

[0117] This system acquires the digital version of the appliance's instruction manual and extracts text data using optical character recognition (OCR) technology. The extracted text data is summarized using a natural language processing (NLP) model. Furthermore, a generative AI model is used to present the contents of the instruction manual in a format that is easy for users to understand. The user selects the appliance model through the user interface, and related operating instructions and troubleshooting information are provided. When the user scans the appliance using a smartphone, AR technology is used to overlay operating instructions on the camera image. Furthermore, a support chatbot answers questions from the user, quickly providing the necessary information. Through this process, the user's operation history and setting changes are recorded, and a machine learning algorithm learns the user's preferences and usage habits. Based on the learning results, personalized operating advice and setting change suggestions are provided.

[0118] Hardware and software used

[0119] Server: A common RDBMS (e.g., MySQL, PostgreSQL) is used as the database, Tesseract OCR is used as the OCR technology, and BERT or GPT-4 is used as the NLP model.

[0120] Device: A mobile device such as a smartphone or tablet that uses ARCore (for Android) or ARKit (for iOS) as AR technology.

[0121] Users: Use the user interface and chatbot functions within the application.

[0122] Specific examples

[0123] Changing the air conditioner remote control settings

[0124] 1. The user opens the application and selects the model of the air conditioner.

[0125] 2. The device requests the selected model information from the server, and the server returns the operation instructions and setting data for the corresponding model.

[0126] 3. The device instructs the user to scan the air conditioner remote control with the camera, and overlays operating instructions on the scanned image of the remote control, allowing the user to visually confirm the specific button operations.

[0127] 4. When a user asks the chatbot, "I don't know how to set it up," the server analyzes the question, generates detailed instructions, and sends them to the device.

[0128] 5. The device displays the generated instructions on the chatbot screen, and the user follows them to complete the setup.

[0129] Prompt Sentence Examples

[0130] Example question: "How do I switch my air conditioner to cooling mode?"

[0131] Example prompt: "My question is about how to operate my air conditioner. What are the specific steps to switch it to cooling mode?"

[0132] Based on these prompts, the generative AI model generates easy-to-understand answers for users, allowing them to intuitively and efficiently obtain instructions and troubleshooting information for their home appliances.

[0133] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0134] Step 1:

[0135] The server retrieves instruction manual data for the home appliance from the database.

[0136] Input: Product Model ID

[0137] Data processing: Search instruction manual data using database queries

[0138] Output: Instruction manual data

[0139] Specific operation: The server executes an SQL query based on the product model ID to retrieve the instruction manual data. For example, it uses a query like "SELECT FROM manuals WHERE product_id = 'AC1234';".

[0140] Step 2:

[0141] The instruction manual data acquired by the server is converted into text data using OCR technology.

[0142] Input: Instruction manual data

[0143] Data processing: Applying OCR technology to extract text data from image data

[0144] Output: Extracted text data

[0145] Specific operation: Convert image-format instruction manual data into text using Tesseract OCR. For example, read the image using Python's PIL library and extract the text using pytesseract.

[0146] Step 3:

[0147] The server then applies the extracted text data to an NLP model to summarize it.

[0148] Input: Text data

[0149] Data processing: Summarizing text using NLP models

[0150] Output: Summarized text data

[0151] How it works: Summarize long text data concisely using NLP models such as BERT and GPT-4, for example using the summarization pipeline in the transformers library.

[0152] Step 4:

[0153] The terminal allows the user to select the model of the home appliance through a user interface.

[0154] Input: User selected model information

[0155] Data processing: Generate API requests and send them to the server

[0156] Output: Relevant information received from the server

[0157] Specific operation: The user selects a product model on the initial screen of the application, and an API request is sent to the server based on that information. For example, an HTTP request such as "GET / manuals / AC1234".

[0158] Step 5:

[0159] The server transmits the acquired related information to the terminal.

[0160] Input: User selected model information

[0161] Data processing: Generate related information after search processing

[0162] Output: Related operating instructions and troubleshooting information

[0163] Specific operation: The server receives the request, retrieves the relevant operation instructions and troubleshooting information from the database, and sends them to the terminal.

[0164] Step 6:

[0165] The device uses its camera to scan home appliances and uses AR technology to overlay operating instructions.

[0166] Input: Acquired related information, camera footage

[0167] Data processing: AR technology overlays operation procedures onto camera images

[0168] Output: AR displayed operation procedure

[0169] Specific operation: When a user points the camera at a home appliance, ARCore or ARKit is used to overlay operating instructions on the real-world image.

[0170] Step 7:

[0171] A user enters a question using the chatbot function within the application.

[0172] Input: User's question text

[0173] Data processing: Parse the question with an NLP model

[0174] Output: Parsed question

[0175] Specific operation: A question is entered into the chatbot, and the question is analyzed using an NLP model.

[0176] Step 8:

[0177] The server generates an appropriate answer to the analyzed question and sends it to the terminal.

[0178] Input: Parsed question content

[0179] Data processing: Generative AI models generate answers

[0180] Output: The generated answer

[0181] Specific operation: The server inputs the received question into a generative AI model to generate an appropriate answer. For example, it uses GPT-4 to generate an answer and sends it to the device.

[0182] Step 9:

[0183] The device displays the generated answer on the chatbot screen.

[0184] Input: Generated Answer

[0185] Data processing: None

[0186] Output: The answer displayed on the chatbot screen

[0187] Specific operation: The generated answer text is displayed on the chatbot screen for the user to see.

[0188] Step 10:

[0189] The server records user operation history and setting changes and learns from them using machine learning algorithms.

[0190] Input: User operation history and setting change information

[0191] Data processing: Learning processing using machine learning algorithms

[0192] Output: Learning results

[0193] Specific operation: Collect operation history and setting change data, and cluster behavioral patterns using, for example, Scikit-learn.

[0194] Step 11:

[0195] The server generates individual operation advice and suggestions for setting changes based on the learning results and sends them to the device.

[0196] Input: Learning results

[0197] Data processing: generating personalized advice and suggestions for changing settings

[0198] Output: Generated advice and suggestions

[0199] Specific operation: Generates personalized advice based on user behavior data and sends it to the device.

[0200] (Application example 1)

[0201] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0202] Instruction manuals for home appliances are often provided in paper format, and understanding them takes time and effort. Furthermore, when users encounter difficulties setting up or operating the device, it is difficult to intuitively understand the specific operation method. Furthermore, with the recent spread of smart security systems, there is a demand for intuitive guidance for their installation and configuration. Conventional technologies lack systems that provide comprehensive, immediate, and appropriate support, so there is a need to solve these issues.

[0203] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0204] In this invention, the server includes means for acquiring instruction manual data and extracting text data using optical character recognition, means for summarizing the extracted text data using a natural language processing model, means for selecting a home appliance model through a user interface and providing related operating instructions and troubleshooting information, means for scanning the home appliance using a camera and visualizing operating procedures using augmented reality technology, means for answering questions from the user using a support chatbot using the natural language processing model, means for recording user operation history and setting changes and learning user preferences and usage habits using a machine learning algorithm, means for providing individual operating advice and suggesting setting changes based on the learning results, and means for visualizing installation and configuration methods for the smart security system using augmented reality technology, thereby enabling intuitive and efficient operation and configuration of home appliances and smart security systems.

[0205] An "instruction manual" is a document that provides users with information on how to operate, set up, and maintain home appliances and other devices.

[0206] "Optical character recognition" is a technology that extracts text data from images and documents.

[0207] A "natural language processing model" is an algorithm or system for understanding and generating human language.

[0208] "User interface" refers to the means or screen through which a user interacts with a system.

[0209] A "camera" is a device that receives light and captures image information.

[0210] "Augmented reality technology" is a technology that displays digital information superimposed on images of the real world.

[0211] A "support chatbot" is software that automatically answers users' questions via text or voice.

[0212] A "machine learning algorithm" is a computational method for learning patterns from data and making predictions and classifications.

[0213] "Preferences" refer to the preferences and tendencies of an individual.

[0214] A "smart security system" is a system that automatically monitors and manages safety using network-connected cameras and sensors.

[0215] An "operating procedure" refers to the steps or methods for using a piece of equipment or system.

[0216] "Troubleshooting Information" means information containing guidance or advice for resolving equipment malfunctions or errors.

[0217] "Operation history" is a record of when a user uses a system or device.

[0218] "Configuration change" refers to the change made to adjust the operation of a system or device.

[0219] This invention is an instruction manual system for a smart security system that utilizes generative AI and augmented reality (AR) technology, allowing users to intuitively and quickly obtain operating instructions and troubleshooting information for security devices. The invention can be implemented as follows through the roles of a server, a terminal, and a user.

[0220] Overall system overview

[0221] This system digitally acquires the instruction manual for the smart security system and extracts the text data using optical character recognition (OCR) technology. The extracted text data is summarized using a natural language processing (NLP) model and provided as information to assist the user in operation. It also utilizes AR technology to allow users to visually understand the product's operation procedures and installation methods using their smartphone. Furthermore, a support chatbot responds to user inquiries, providing a means for users to instantly obtain the information they need.

[0222] Program processing overview

[0223] 1. The server retrieves the instruction manual data for the smart security system from the database. The retrieved data is converted into text data using OCR technology (e.g., Google Cloud Vision API) and then summarized using an NLP model (e.g., OpenAI GPT-3). This converts the long instruction manual into a short, easy-to-understand format.

[0224] 2. The device (smartphone, smart glasses, head-mounted display) launches the application and prompts the user to select the security device model through the user interface. Based on the selected model, it communicates with the server to obtain relevant operating instructions and troubleshooting information.

[0225] 3. The device allows users to use the camera to scan for security devices, and uses AR technology (e.g., Apple ARKit, Google ARCore) to overlay operation instructions and installation locations on the camera image, providing users with a visual guide.

[0226] 4. The user enters a question using the chatbot function within the application. The chatbot uses an NLP model to analyze the question and generate an appropriate answer.

[0227] 5. The server records the user's operation history and setting changes, and uses machine learning algorithms to learn the user's preferences and usage habits. Based on this learning, it generates personalized operation advice and setting change suggestions and sends them to the device.

[0228] Specific Examples

[0229] Security camera Wi-Fi settings

[0230] 1. The user opens the application and selects the security camera model.

[0231] 2. The device requests the selected model information from the server, and the server returns the operation instructions and setting data for the corresponding model.

[0232] 3. The device prompts the user to scan the security camera and overlays Wi-Fi setup instructions on the scanned camera image, allowing the user to visually confirm the specific steps.

[0233] 4. When a user asks the chatbot, "I don't know how to set up Wi-Fi," the server analyzes the question, generates detailed instructions, and sends them to the device.

[0234] 5. The device displays the generated instructions on the chatbot screen, and the user follows them to complete the setup.

[0235] Prompt Sentence Examples

[0236] "Please tell me about the Wi-Fi settings for the security camera."

[0237] "I'd like to set up the initial settings for my security system. Please tell me the procedure."

[0238] "Please show me in AR where to place the motion sensor."

[0239] In this way, the system provides users with intuitive and efficient operation and troubleshooting information for their smart security system, making it easier to understand the instruction manual and providing optimal operation advice tailored to the user's individual preferences and usage habits.

[0240] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0241] Step 1:

[0242] The server retrieves the instruction manual data for the smart security system from the database. The input is a request for instruction manual data, and the output is the retrieved instruction manual data. Using OCR technology (e.g., Google Cloud Vision API), the image data is converted into text data. Specifically, the server accesses the database upon receiving the request, retrieves the image data of the specified instruction manual, and applies OCR technology to extract the text data.

[0243] Step 2:

[0244] The server summarizes the extracted text data by running it through a natural language processing (NLP) model (e.g., OpenAI GPT-3). The input is the OCR-processed text data, and the output is the summarized text data. Specifically, the server inputs the text data into the NLP model and receives the generated summary data.

[0245] Step 3:

[0246] The terminal (smartphone, smart glasses, head-mounted display) allows the user to select a security device model through a user interface. The input is the user's model selection operation, and the output is the selected model information. Specifically, the terminal presents multiple models through the display interface and collects the information selected by the user.

[0247] Step 4:

[0248] The terminal requests the selected model information from the server, and the server provides the related operation method and troubleshooting information. The input is the model information request, and the output is the operation method and troubleshooting information. In concrete terms, the terminal sends a request to the server, and the server searches for the corresponding model information and returns it to the terminal.

[0249] Step 5:

[0250] The device allows users to scan security devices using a camera. The input is the camera image, and the output is an overlay display of location information and operation procedures. The display is achieved using augmented reality technology (e.g., Apple ARKit, Google ARCore). Specifically, the device acquires real-time images from the camera and uses AR technology to overlay the necessary guide information on the image.

[0251] Step 6:

[0252] A user inputs a question using the chatbot function within the application. The chatbot uses an NLP model to analyze the input question and generate an appropriate answer. The input is the user's question, and the output is the generated answer. Specifically, the user inputs a question as text, the chatbot sends the question to the NLP model for analysis, and then displays the generated answer to the user.

[0253] Step 7:

[0254] The server records the user's operation history and setting changes, and uses a machine learning algorithm to learn the user's preferences and usage habits. The input is the operation history and setting change data, and the output is individual operation advice and setting change suggestions based on the learning results. Specifically, the server records the operation history in a database, analyzes the data using a machine learning algorithm, and generates individually optimized advice and setting change suggestions.

[0255] Step 8:

[0256] The device provides the user with personalized operation advice and suggestions for setting changes based on the learning results. The input is the suggestion data sent from the server, and the output is a presentation to the user. Specifically, the device receives the suggestion data and displays it on the user interface.

[0257] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0258] The present invention is a home appliance instruction manual system that utilizes generative AI, augmented reality (AR) technology, and an emotion engine, allowing users to intuitively and quickly obtain operation instructions and troubleshooting information for home appliances. The present invention can be implemented as follows through the roles of a server, a terminal, and a user.

[0259] Overall system overview

[0260] This system acquires digital versions of instruction manuals for home appliances, extracts text data using optical character recognition (OCR), summarizes it using natural language processing (NLP) models, and provides it as information to assist users in their operations. Augmented reality (AR) technology also allows users to visually understand product operation procedures using their smartphones. A support chatbot responds to user inquiries, and an emotion engine recognizes the user's emotions and provides optimal advice.

[0261] Program processing overview

[0262] 1. The server retrieves the instruction manual data for the home appliance from the database. The retrieved data is converted into text data using OCR technology and summarized using an NLP model, converting the long instructions into a short, easy-to-understand format.

[0263] 2. The device launches the application and prompts the user to select the model of the home appliance through the user interface. Based on the selected model, the device communicates with the server to obtain relevant operating instructions and troubleshooting information.

[0264] 3. The device allows users to scan home appliances using their smartphone camera, and uses AR technology to overlay operating instructions on the camera image, providing users with a visual guide.

[0265] 4. The user enters a question using the chatbot function within the application. The chatbot uses an NLP model to analyze the question and generate an appropriate answer.

[0266] 5. The server records the user's operation history and setting changes, and uses a machine learning algorithm to learn the user's preferences and usage habits. Based on this learning, it generates individualized operation advice and setting change suggestions and sends them to the device.

[0267] 6. Recognize user emotions using an emotion engine. Analyze facial expressions and voice recorded through the camera while the user is operating the device, and evaluate the user's emotional state in real time.

[0268] 7. The server adjusts its operational advice and suggestions for setting changes based on the emotional data obtained by the emotion engine. For example, if the user is confused, it will provide more detailed and gentler explanations.

[0269] Specific Examples

[0270] The following is a specific scenario.

[0271] Changing the air conditioner remote control settings

[0272] 1. The user opens the application and selects the air conditioner model.

[0273] 2. The device requests the selected model information from the server, and the server returns the operation instructions and setting data for the corresponding model.

[0274] 3. The device instructs the user to scan the air conditioner remote control with the camera, and overlays operating instructions on the scanned image of the remote control, allowing the user to visually confirm the specific button operations.

[0275] 4. When a user asks the chatbot, "I don't know how to set it up," the server analyzes the question, generates detailed instructions, and sends them to the device.

[0276] 5. The device displays the generated instructions on the chatbot screen, and the user follows them to complete the setup.

[0277] 6. If the emotion engine detects confusion or stress in the user's facial expression, the server will adjust the advice accordingly and send a more understandable explanation to the device.

[0278] In this way, the system not only provides users with intuitive and efficient instructions on how to operate home appliances and troubleshooting information, making it easier to understand instruction manuals, but also provides appropriate support according to the user's emotional state.

[0279] The processing flow will be explained below.

[0280] Step 1:

[0281] The server queries the database to retrieve the instruction manual (PDF format) for the home appliance, and stores the retrieved PDF file in local storage or temporarily in memory.

[0282] Step 2:

[0283] The server extracts text data from the PDF using OCR (Optical Character Recognition) technology. Using an OCR library such as Tesseract, it analyzes all pages in the PDF and converts them into text data.

[0284] Step 3:

[0285] The server then runs the extracted text data through a natural language processing (NLP) model to summarize it, using models like BERT and GPT to reduce redundant descriptions and extract key information.

[0286] Step 4:

[0287] The device launches the application and displays the user interface, initially displaying a search bar and drop-down lists for selecting appliance categories and models.

[0288] Step 5:

[0289] The user selects the model of the home appliance they are using on the interface, and the selected model information is sent to the server, requesting related operating instructions and troubleshooting information.

[0290] Step 6:

[0291] Based on the received model information, the server retrieves relevant operating instructions and troubleshooting information and sends it to the device. The information is retrieved from a database and appropriately formatted before being sent.

[0292] Step 7:

[0293] The device responds to user requests and displays the received instructional and troubleshooting information, appropriately laid out in the app's interface.

[0294] Step 8:

[0295] The device will activate the camera, allowing the user to scan home appliances, capture camera footage in real time, and provide guidance to the user.

[0296] Step 9:

[0297] The device uses augmented reality (AR) technology to overlay operation instructions on the camera image, using libraries such as ARKit and ARCore to overlay instruction icons and text on the operation panel.

[0298] Step 10:

[0299] Users operate home appliances by following the instructions on the camera footage, and follow the on-screen instructions and guides to complete the required operations.

[0300] Step 11:

[0301] Users enter questions into the in-app chatbot, using natural language, and submit the question.

[0302] Step 12:

[0303] The server analyzes the questions received from users through a natural language processing model, interprets the input text to understand its meaning, and generates an appropriate answer.

[0304] Step 13:

[0305] The server sends the generated answer to the terminal, where it is formatted appropriately and displayed on the chatbot screen.

[0306] Step 14:

[0307] The device displays the received response on the chatbot's conversation screen, and the user can confirm the displayed information and perform any necessary operations or settings.

[0308] Step 15:

[0309] The device records the user's operation history and setting changes, and sends the operation details and timestamps to the server as log data.

[0310] Step 16:

[0311] The server compiles the received operation history and setting change data and uses machine learning algorithms to learn the user's preferences and usage habits.

[0312] Step 17:

[0313] The server generates individualized operation advice and setting change suggestions based on the learning results, and sends the suggestions to the device as a notification.

[0314] Step 18:

[0315] The device will notify the user of the received suggestions, either as a push notification or a pop-up message, suggesting new settings or operation methods to the user.

[0316] Step 19:

[0317] The device captures facial expression data through the camera while the user is operating the device, and transmits it to the emotion engine, which analyzes facial expressions and voice to evaluate the user's emotional state in real time.

[0318] Step 20:

[0319] The emotion data recognized by the emotion engine is sent to the server, which then adjusts the advice and suggested settings changes accordingly.

[0320] Step 21:

[0321] If the user's emotional state is confused or stressed, the server generates a more detailed and gentle explanation and sends it to the device, improving the quality of support according to the user's emotional state.

[0322] Example 2

[0323] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0324] Modern home appliances are becoming increasingly multifunctional, making it difficult to accurately understand their operation methods and troubleshooting information. Traditional instruction manuals are provided in paper or digital format, but the volume of information makes it difficult for users to quickly find the information they need. Furthermore, the lack of flexible responses or personalized advice based on the user's emotional state leaves the user with a poor user experience. Furthermore, there is a need for systems that can learn users' preferences and usage habits and suggest optimal operations and settings.

[0325] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0326] In this invention, the server includes means for acquiring instruction manual data and extracting character data using optical character recognition, means for summarizing the extracted character data through a natural language processing model, means for recording a user's operation history and setting changes and learning the user's preferences and usage habits using a machine learning algorithm, means for recognizing the user's emotional state using an emotion analysis engine, and means for adjusting operation advice and setting change suggestions based on the results of the emotion analysis engine. This allows the user to quickly obtain the necessary information and operate the home appliance intuitively and efficiently. In addition, to improve the user experience, the user can receive optimal advice and setting suggestions based on the user's emotional state.

[0327] "Home appliances" refer to electronic devices and electrically powered devices that are generally used in the home, including, for example, air conditioners, refrigerators, washing machines, and the like.

[0328] An "instruction manual" is a document that describes how to use, install, and troubleshoot a home appliance device.

[0329] "Digitization" is the process of converting paper or analog information into digital form, making it possible to view, store, and process the information on computers and other digital devices.

[0330] Optical character recognition (OCR) is a technology that extracts character data from images and scanned documents, allowing handwritten or printed characters to be recognized as digital text.

[0331] A "natural language processing (NLP) model" is an artificial intelligence technique for understanding, analyzing, and processing natural language, and can perform tasks such as text summarization, translation, and sentiment analysis.

[0332] A "user interface" is a medium through which a user interacts with a system, and includes a graphical user interface (GUI) and a voice user interface (VUI).

[0333] An "image acquisition device" is a device for acquiring image data, such as a camera or scanner, and particularly includes cameras installed in smartphones.

[0334] Augmented reality (AR) technology is a technology that displays digital information overlaid on images of the real world, allowing users to visually recognize the real world and virtual information in an integrated manner.

[0335] A "support chatbot" is a program that automatically responds to questions and requests from users, and uses natural language processing technology to provide appropriate answers.

[0336] A "machine learning algorithm" is a computational method for analyzing data, learning patterns, and making predictions. It is particularly used to analyze user behavior patterns and suggest optimal operations and settings.

[0337] An "emotion analysis engine" is an artificial intelligence technology that analyzes a user's facial expressions and voice data to evaluate their emotional state, and can provide feedback according to the user's psychological state.

[0338] "Operation history" is a record of operations and setting changes performed by a user using the system, and this information is used to understand the user's preferences and behavioral patterns.

[0339] "Adjusting suggestions" means dynamically changing the appropriate advice and recommendations for setting changes to the user based on collected data and analysis results.

[0340] The present invention is a user manual system for home appliances that utilizes generative AI models, augmented reality (AR) technology, and an emotion engine, allowing users to intuitively and quickly obtain operating instructions and troubleshooting information for home appliances. The present invention is implemented through the roles of a server, a terminal, and a user. This improves user convenience and satisfaction.

[0341] Hardware and Software Use

[0342] Server: Database, Optical Character Recognition (OCR) technology, Natural Language Processing (NLP) models, Sentiment Analysis Engine, Machine Learning Algorithms

[0343] Devices: Smartphones, cameras, and applications using augmented reality (AR) technology

[0344] User: Operating the application, using the camera, using the chat function

[0345] Data acquisition and preprocessing

[0346] The server retrieves instruction manual data for a home appliance from the database. For example, it retrieves instruction manual data for an air conditioner. The retrieved data is often in PDF or image file format.

[0347] Text extraction and summarization

[0348] The server converts the acquired instruction manual data into text data using optical character recognition (OCR) technology. For example, it uses Tesseract OCR to extract text from PDFs and images. The extracted text data is then input into a natural language processing (NLP) model (e.g., BERT, GPT-3) to generate a summary. For example, it can extract only important setup steps and necessary information from long instructions and create a short summary.

[0349] Product Scanning and Display

[0350] The terminal provides an application that allows the user to select the model of a home appliance. When the user selects "air conditioner," the terminal requests the model information from the server, and the server sends related operating instructions and troubleshooting information to the terminal.

[0351] Inquiry response

[0352] The user uses the app's chatbot function to input a question, such as "How do I set a timer?" The server analyzes the question using a natural language processing (NLP) model, generates an appropriate answer, and sends it to the device.

[0353] Operation history recording and learning

[0354] The server records the user's operation history and setting changes, and uses a machine learning algorithm to learn the user's preferences and usage habits. For example, it records a history such as "Timer settings were changed on October 1, 2023," and based on this, suggests optimal settings for users who often use the device at night.

[0355] Recognition of emotional states

[0356] The emotion analysis engine analyzes the user's facial expressions and voice via a camera and microphone to assess their emotional state. For example, it may recognize that the user is confused. The server records this emotional data and generates appropriate feedback.

[0357] Providing advice and coordination

[0358] The server tailors operational advice and setting change suggestions based on the results of the emotion analysis engine. For example, if the user is confused, it generates detailed and easy-to-understand instructions and sends them to the device.

[0359] Examples of concrete examples and prompts

[0360] For example, when a user wants to change the remote control settings of an air conditioner, the specific operating procedure is as follows.

[0361] 1. The user opens the application and selects the air conditioner model.

[0362] 2. The device requests the selected model information from the server, and the server returns the operation instructions and setting data for the corresponding model.

[0363] 3. The device instructs the user to scan the air conditioner remote control with the camera and overlays operating instructions on the scanned image.

[0364] 4. When a user asks the chatbot, "I don't know how to set it up," the server analyzes the question, generates detailed instructions, and sends them to the device.

[0365] 5. The device displays the generated instructions on the chatbot screen, and the user follows them to complete the setup.

[0366] 6. The emotion engine detects confusion or stress from the user's facial expressions, and the server adjusts the advice based on this and sends a more understandable explanation to the device.

[0367] As an example of a prompt sentence, you could feed the generative AI model the following:

[0368] "How do I set the timer on the air conditioner remote control?"

[0369] "What should I do if a user is confused?"

[0370] In this way, the system not only provides users with intuitive and efficient operating instructions and troubleshooting information for home appliances, making it easier to understand instruction manuals, but also provides appropriate support according to the user's emotional state.

[0371] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0372] Step 1:

[0373] The server retrieves instruction manual data for home appliances from a database. The retrieved data may be in PDF or image format. The input data may be, for example, "air conditioner instruction manual data," which is then subjected to OCR processing on the server side. The server then converts this data into text data using optical character recognition (OCR) technology. The output is the text data for the instruction manual. For example, the text may include "How to operate the air conditioner" or "Troubleshooting information."

[0374] Step 2:

[0375] The server inputs the character data extracted by OCR into a natural language processing (NLP) model to generate a summary. The character data is used as input data and passed to an NLP model (e.g., BERT or GPT-3). Specifically, the text is summarized using an application with a summarization function. As output, a summary text is generated that extracts only the important steps from the detailed description. For example, "Description of all buttons on a remote control" is summarized as "Description of the main buttons."

[0376] Step 3:

[0377] The terminal launches an application to allow the user to select a model of a home appliance. The user operates the terminal's user interface to select a model, such as "air conditioner." The user's model selection information is used as input. Based on this, the terminal sends a request for model information to the system. The model selection information is sent to the server as output.

[0378] Step 4:

[0379] Based on the request from the terminal, the server returns the operation method and troubleshooting information for the relevant model. The input is the model information selected by the user, and based on that, the server retrieves the relevant information from the database. Specifically, this includes detailed data such as operation procedures and how to deal with errors. As output, the server sends the related operation method and troubleshooting information to the terminal.

[0380] Step 5:

[0381] The device instructs the user to scan the home appliance with a camera and visualizes the operation procedure using augmented reality (AR) technology. The input is the image data scanned by the user with the camera. Specifically, AR technology is used to overlay instructions such as "Press the POWER button" on the camera image. The output is the visualized operation procedure provided to the user.

[0382] Step 6:

[0383] Users can use the chatbot function within the app to input specific questions, such as, "I don't know how to set the timer on my air conditioner." This question is then sent to the chatbot.

[0384] Step 7:

[0385] The server analyzes questions received via the chatbot using a natural language processing (NLP) model, generates appropriate answers, and sends them to the device. The input is the user's question text. Based on this, the NLP model operates and generates text with specific setup procedures and operation instructions. As output, the generated answer text is sent to the device and displayed on the chatbot screen.

[0386] Step 8:

[0387] The server records the user's operation history and setting changes, and uses a machine learning algorithm to learn the user's preferences and usage habits. The input is the user's operation history data. Specifically, a history such as "Change timer settings on October 1, 2023" is saved. The output is recommended settings and operation methods based on the learning results.

[0388] Step 9:

[0389] The emotion analysis engine analyzes the user's facial expressions and voice via a camera and microphone to evaluate their emotional state. The inputs include facial expression data captured by the camera and voice data captured by the microphone. Specifically, it analyzes in real time whether the user is confused or not. The output is the emotion analysis result.

[0390] Step 10:

[0391] The server adjusts operational advice and setting change suggestions based on the results of the sentiment analysis engine. The input is the sentiment analysis data. Specifically, it prepares detailed and friendly explanations for confused users. The output is the adjusted advice and setting change suggestions sent to the device.

[0392] The above is the specific processing flow of the program of this system.

[0393] (Application example 2)

[0394] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0395] Conventional instruction manual systems for home appliances often make it difficult for users to intuitively understand how to operate them, especially when it comes to complex operations and troubleshooting. Furthermore, due to a lack of systems utilizing emotion engines and generative AI, users' emotional state and individual operational support are insufficient. This results in reduced user operational efficiency and longer troubleshooting times. Furthermore, there is a demand for more efficient picking operations in warehouses, but intuitive support using AR technology is lacking.

[0396] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0397] In this invention, the server includes means for acquiring instruction manual data and extracting text data by optical character recognition, means for summarizing the extracted text data by applying a natural language processing model, and means for selecting a device model through a user interface and providing related operating instructions and troubleshooting information, thereby enabling intuitive and rapid provision of operating instructions and troubleshooting information to the user.

[0398] The server also includes a means for scanning devices using a camera and visualizing operation procedures using augmented reality technology, a means for answering user questions using a support chatbot using a natural language processing model, a means for recording user operation history and setting changes and learning user preferences and usage habits using a machine learning algorithm, a means for providing individual operation advice and suggesting setting changes based on the learning results, a means for proposing work procedures using a generative AI model, a means for analyzing the user's emotional state using an emotion engine and adjusting the operation procedures, and a means for displaying item picking information and supporting the user using augmented reality technology. This makes it possible to provide optimal support according to the user's emotional state and behavioral patterns, thereby improving operation efficiency and work accuracy.

[0399] "Instruction Data" means digital instruction manual information including operating instructions and troubleshooting information for home appliances and other devices.

[0400] Optical character recognition is a technology that automatically reads letters and numbers from images or handwritten characters and converts them into digital text.

[0401] A "natural language processing model" is a machine learning model for understanding, analyzing, and generating human language, allowing it to summarize text data and answer questions.

[0402] A "user interface" is an interaction mechanism that provides a screen and operating methods for users to interact with a system.

[0403] "Augmented reality technology" is a technology that displays digital information overlaid on the real environment, and is used to overlay operating procedures and instructions on camera images.

[0404] A "support chatbot" is a program that uses a natural language processing model to automatically respond to questions from users.

[0405] A "machine learning algorithm" is an algorithm that learns patterns using large amounts of data and makes predictions and classifications for new data.

[0406] A "generative AI model" is an artificial intelligence model that generates new content based on existing data, and is particularly used to generate text, audio, images, etc.

[0407] An "emotion engine" is a system that analyzes a person's emotional state from facial expressions and voice, and responds adaptively based on the results.

[0408] "Item picking information" is information about the work of picking items from warehouses and logistics centers, and includes a picking list and the like.

[0409] This invention is a logistics center support system that utilizes generative AI models, augmented reality (AR) technology, and an emotion engine to enable users to intuitively and quickly perform item picking tasks. This system can be implemented as follows through the roles of the server, terminal, and user.

[0410] Overall system overview

[0411] The system acquires item picking information in digital format and extracts text data using optical character recognition (OCR) technology. The extracted data is summarized using a natural language processing (NLP) model and provided as information to assist the user in their work. AR technology also allows users to visually understand the picking process using smart glasses. A support chatbot answers user inquiries, and an emotion engine recognizes the user's emotions to provide optimal assistance.

[0412] Program processing overview

[0413] The server retrieves item picking information from a database. The retrieved data is converted into text data using OCR technology and summarized using an NLP model. The server then uses a generative AI model to suggest work procedures. The user obtains relevant picking information through a device equipped with a user interface. The device scans the item using smart glasses and overlays the picking list information using AR technology. The user can intuitively perform the work using the smart glasses, following the visual guide. When the user inputs a question using the smart glasses' chatbot function, the server uses an NLP model to generate an appropriate answer and sends it to the device. The server records the user's work history and setting changes and uses a machine learning algorithm to learn the user's preferences and usage habits. Based on this learning result, it generates individual work procedures and setting change suggestions and sends them to the device. The emotion engine analyzes the user's emotional state from their facial expressions and voice and adjusts the operating procedures. The server adjusts assistance advice and provides more appropriate explanations based on the emotional data obtained by the emotion engine.

[0414] Specific Examples

[0415] The following is a specific scenario.

[0416] Support for picking work in the warehouse

[0417] 1. The server retrieves item picking information from the database.

[0418] 2. The server converts the acquired picking information into text data using OCR technology.

[0419] 3. The server summarizes the converted text data using an NLP model.

[0420] 4. The server uses the generative AI model to generate the optimal picking procedure.

[0421] 5. The terminal overlays the picking list information to the user through the smart glasses.

[0422] 6. The user wears the smart glasses and follows visual guidance to pick the item.

[0423] 7. When a user uses the chatbot function to ask a question, the server uses an NLP model to generate an appropriate answer and sends it to the device.

[0424] 8. The emotion engine assesses the user's emotional state from their facial expressions and voice, and adjusts the assistance provided depending on the problem that has occurred.

[0425] Example prompts for generative AI models

[0426] "Picking List Information: Box 123, Section A1, Item 345

[0427] Please suggest the best warehouse operation procedure based on the information below:

[0428] In this way, the system not only provides users with intuitive and efficient support for picking tasks, improving work efficiency within the warehouse, but also provides appropriate support according to the user's emotional state.

[0429] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0430] Step 1:

[0431] The server retrieves item picking information from the database. It uses the picking list information stored in the database as input to obtain the necessary picking data. It obtains raw data for OCR processing as output. Specific operations include issuing a database query to obtain the list information.

[0432] Step 2:

[0433] The server converts the acquired picking information into text data using OCR technology. It uses the raw data obtained in step 1 as input and uses OCR software to obtain the output as digital text. Specifically, it calls an OCR module (e.g., Tesseract) to extract text information from images and documents.

[0434] Step 3:

[0435] The server summarizes the converted text data using a natural language processing (NLP) model. The text data is provided as input to the NLP model, which performs the summarization process and outputs a concise operating procedure. Specifically, an NLP library (e.g., SpaCy or OpenAI GPT-3) is used to generate the summary text.

[0436] Step 4:

[0437] The server uses a generative AI model to propose the optimal picking procedure. The generative AI model receives the summarized text data and the prompt as input. Based on the prompt, the server generates the optimal picking procedure and obtains the procedure as output. Specifically, the server invokes the generative AI (e.g., OpenAI GPT-3), inputs the prompt, and generates the recommended procedure.

[0438] Step 5:

[0439] The terminal overlays the picking list information to the user through the smart glasses. It uses the picking procedure sent from the server as input. It uses AR technology to overlay a visual guide on the smart glasses display and provides visual assistance as output. Specific operations include using AR software (e.g., Unity or ARKit) to display the information on the glasses display.

[0440] Step 6:

[0441] The user wears smart glasses and follows visual guidance to pick items. The input is the information displayed on the smart glasses. The output is to accurately pick the specified items. Specific actions involve looking at the display on the glasses, picking out the items as instructed, and placing them on a cart or other device for inspection.

[0442] Step 7:

[0443] When a user uses the chatbot function to ask a question, the server uses an NLP model to generate an appropriate answer and sends it to the device. The user's question text is used as input. The NLP model analyzes it and outputs the answer. Specifically, the natural language processing model analyzes the user's question, generates an appropriate answer text, and displays it on the chat screen.

[0444] Step 8:

[0445] The emotion engine evaluates the user's emotional state from their facial expressions and voice, and adjusts the assistance provided depending on the problem that has occurred. The input is the user's video and audio data. The emotion engine analyzes the data and outputs the user's emotional state. Specifically, it uses facial recognition technology (e.g., DeepFace or face_recognition) and voice analysis technology to analyze the user's emotions and adjust appropriate advice and procedures.

[0446] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0447] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0448] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0449] [Second embodiment]

[0450] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0451] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0452] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0453] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0454] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0455] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0456] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0457] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0458] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0459] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0460] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0461] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0462] This invention is a home appliance instruction manual system that utilizes generative AI and augmented reality (AR) technology, allowing users to intuitively and quickly obtain operation methods and troubleshooting information for home appliances. The invention can be implemented as follows through the roles of a server, a terminal, and a user.

[0463] Overall system overview

[0464] This system acquires digital versions of instruction manuals for home appliances and extracts text data using optical character recognition (OCR) technology. The extracted text data is summarized using a natural language processing (NLP) model and provided as information to assist users in their operations. It also utilizes AR technology to enable users to visually understand product operation procedures using their smartphones. Furthermore, a support chatbot responds to user inquiries, providing a means for users to instantly obtain the information they need.

[0465] Program processing overview

[0466] 1. The server retrieves the instruction manual data for the home appliance from the database. The retrieved data is converted into text data using OCR technology and summarized using an NLP model. This converts the long instructions into a short, easy-to-understand format.

[0467] 2. The device launches the application and prompts the user to select the model of the home appliance through the user interface. Based on the selected model, the device communicates with the server to obtain relevant operating instructions and troubleshooting information.

[0468] 3. The device allows users to scan home appliances using their smartphone camera, and uses AR technology to overlay operating instructions on the camera image, providing users with a visual guide.

[0469] 4. The user enters a question using the chatbot function within the application. The chatbot uses an NLP model to analyze the question and generate an appropriate answer.

[0470] 5. The server records the user's operation history and setting changes, and uses a machine learning algorithm to learn the user's preferences and usage habits. Based on this learning, it generates individualized operation advice and setting change suggestions and sends them to the device.

[0471] Specific Examples

[0472] The following is a specific scenario.

[0473] Changing the air conditioner remote control settings

[0474] 1. The user opens the application and selects the air conditioner model.

[0475] 2. The device requests the selected model information from the server, and the server returns the operation instructions and setting data for the corresponding model.

[0476] 3. The device instructs the user to scan the air conditioner remote control with the camera, and overlays operating instructions on the scanned image of the remote control, allowing the user to visually confirm the specific button operations.

[0477] 4. When a user asks the chatbot, "I don't know how to set it up," the server analyzes the question, generates detailed instructions, and sends them to the device.

[0478] 5. The device displays the generated instructions on the chatbot screen, and the user follows them to complete the setup.

[0479] In this way, the system intuitively and efficiently provides users with operation and troubleshooting information for home appliances, making it easier to understand instruction manuals and providing optimal operation advice tailored to the user's individual preferences and usage habits.

[0480] The processing flow will be explained below.

[0481] Step 1:

[0482] The server queries the database to retrieve the instruction manual (PDF format) for the home appliance, and stores the retrieved PDF file in local storage or temporarily in memory.

[0483] Step 2:

[0484] The server uses OCR (Optical Character Recognition) technology to extract text data from the PDF. Using an OCR library such as Tesseract, it analyzes all pages in the PDF and converts them into text data.

[0485] Step 3:

[0486] The server then runs the extracted text data through a natural language processing (NLP) model to summarize it, using models like BERT and GPT to reduce redundant descriptions and extract key information.

[0487] Step 4:

[0488] The device launches the application and displays the user interface, initially displaying a search bar and drop-down lists for selecting appliance categories and models.

[0489] Step 5:

[0490] The user selects the model of the home appliance they are using on the interface, and the selected model information is sent to the server, requesting related operating instructions and troubleshooting information.

[0491] Step 6:

[0492] Based on the received model information, the server retrieves relevant operating instructions and troubleshooting information and sends it to the device. The information is retrieved from a database and appropriately formatted before being sent.

[0493] Step 7:

[0494] The device responds to user requests and displays the received instructional and troubleshooting information, appropriately laid out in the app's interface.

[0495] Step 8:

[0496] The device will activate the camera, allowing the user to scan home appliances, capture camera footage in real time, and provide guidance to the user.

[0497] Step 9:

[0498] The device uses augmented reality (AR) technology to overlay operation instructions on the camera image, using libraries such as ARKit and ARCore to overlay instruction icons and text on the operation panel.

[0499] Step 10:

[0500] Users operate home appliances by following the instructions on the camera footage, and follow the on-screen instructions and guides to complete the required operations.

[0501] Step 11:

[0502] Users enter questions into the in-app chatbot, using natural language, and submit the question.

[0503] Step 12:

[0504] The server analyzes the questions received from users through a natural language processing model, interprets the input text to understand its meaning, and generates an appropriate answer.

[0505] Step 13:

[0506] The server sends the generated answer to the terminal, where it is formatted appropriately and displayed on the chatbot screen.

[0507] Step 14:

[0508] The device displays the received response on the chatbot's conversation screen, and the user can confirm the displayed information and perform any necessary operations or settings.

[0509] Step 15:

[0510] The device records the user's operation history and setting changes, and sends the operation details and timestamps to the server as log data.

[0511] Step 16:

[0512] The server compiles the received operation history and setting change data and uses machine learning algorithms to learn the user's preferences and usage habits.

[0513] Step 17:

[0514] The server generates individualized operation advice and setting change suggestions based on the learning results, and sends the suggestions to the device as a notification.

[0515] Step 18:

[0516] The device will notify the user of the received suggestions, either as a push notification or a pop-up message, suggesting new settings or operation methods to the user.

[0517] Example 1

[0518] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0519] Conventional paper manuals for home appliance operation and troubleshooting information are often difficult for users to understand and take a long time to understand. Furthermore, manuals are often lost, making it difficult to quickly obtain the necessary information. Furthermore, when operation methods or solutions are complex, users can easily become lost. There is a need for a system that can solve these problems and enable users to quickly and intuitively understand and operate home appliances.

[0520] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0521] In this invention, the server includes means for acquiring instruction manual data for the home appliance and extracting text data by optical character recognition, means for summarizing the extracted text data through a natural language processing model, means for providing the contents of the instruction manual in a format that is easy for the user to understand using a generative AI model, and means for prompting the user to ask specific questions using prompt sentences and generating detailed answers in response to the questions, thereby enabling the user to quickly and intuitively obtain complex operating instructions and troubleshooting information.

[0522] "Home appliances" refers to all household electronic devices that users use on a daily basis.

[0523] An "instruction manual" refers to a document that describes how to operate, use, maintain, and troubleshoot a home appliance.

[0524] "Digitalization" refers to the process of converting data in paper or physical form into electronic form.

[0525] "User interface" refers to the screen and input devices that users use to operate the device.

[0526] "Optical character recognition (OCR)" refers to the technology that identifies character information from image data and converts it into digital text.

[0527] "Text data" refers to data that is stored as character information and can be processed electronically.

[0528] A "natural language processing (NLP) model" refers to an artificial intelligence model for understanding, analyzing, and generating human language.

[0529] An "abstract" is a concise summary of the contents of a longer document.

[0530] "Augmented reality (AR) technology" refers to the technology of overlaying digital information onto images of the real world.

[0531] A "support chatbot" is a program that automatically responds to questions from users.

[0532] A "machine learning algorithm" refers to a method that allows a computer to learn on its own based on data and make predictions and classifications.

[0533] A "generative AI model" refers to an artificial intelligence model that generates new sentences and answers based on human language.

[0534] A "prompt sentence" refers to a question-type sentence that guides the user's input.

[0535] This invention is a home appliance instruction manual system that utilizes generative AI models and augmented reality (AR) technology, allowing users to intuitively and quickly obtain operating instructions and troubleshooting information for home appliances. This invention is realized through the roles of a server, a terminal, and a user.

[0536] Overall system overview

[0537] This system acquires the digital version of the appliance's instruction manual and extracts text data using optical character recognition (OCR) technology. The extracted text data is summarized using a natural language processing (NLP) model. Furthermore, a generative AI model is used to present the contents of the instruction manual in a format that is easy for users to understand. The user selects the appliance model through the user interface, and related operating instructions and troubleshooting information are provided. When the user scans the appliance using a smartphone, AR technology is used to overlay operating instructions on the camera image. Furthermore, a support chatbot answers questions from the user, quickly providing the necessary information. Through this process, the user's operation history and setting changes are recorded, and a machine learning algorithm learns the user's preferences and usage habits. Based on the learning results, personalized operating advice and setting change suggestions are provided.

[0538] Hardware and software used

[0539] Server: A common RDBMS (e.g., MySQL, PostgreSQL) is used as the database, Tesseract OCR is used as the OCR technology, and BERT or GPT-4 is used as the NLP model.

[0540] Device: A mobile device such as a smartphone or tablet that uses ARCore (for Android) or ARKit (for iOS) as AR technology.

[0541] Users: Use the user interface and chatbot functions within the application.

[0542] Specific examples

[0543] Changing the air conditioner remote control settings

[0544] 1. The user opens the application and selects the model of the air conditioner.

[0545] 2. The device requests the selected model information from the server, and the server returns the operation instructions and setting data for the corresponding model.

[0546] 3. The device instructs the user to scan the air conditioner remote control with the camera, and overlays operating instructions on the scanned image of the remote control, allowing the user to visually confirm the specific button operations.

[0547] 4. When a user asks the chatbot, "I don't know how to set it up," the server analyzes the question, generates detailed instructions, and sends them to the device.

[0548] 5. The device displays the generated instructions on the chatbot screen, and the user follows them to complete the setup.

[0549] Prompt Sentence Examples

[0550] Example question: "How do I switch my air conditioner to cooling mode?"

[0551] Example prompt: "My question is about how to operate my air conditioner. What are the specific steps to switch it to cooling mode?"

[0552] Based on these prompts, the generative AI model generates easy-to-understand answers for users, allowing them to intuitively and efficiently obtain instructions and troubleshooting information for their home appliances.

[0553] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0554] Step 1:

[0555] The server retrieves instruction manual data for the home appliance from the database.

[0556] Input: Product Model ID

[0557] Data processing: Search instruction manual data using database queries

[0558] Output: Instruction manual data

[0559] Specific operation: The server executes an SQL query based on the product model ID to retrieve the instruction manual data. For example, it uses a query like "SELECT FROM manuals WHERE product_id = 'AC1234';".

[0560] Step 2:

[0561] The instruction manual data acquired by the server is converted into text data using OCR technology.

[0562] Input: Instruction manual data

[0563] Data processing: Applying OCR technology to extract text data from image data

[0564] Output: Extracted text data

[0565] Specific operation: Convert image-format instruction manual data into text using Tesseract OCR. For example, read the image using Python's PIL library and extract the text using pytesseract.

[0566] Step 3:

[0567] The server then applies the extracted text data to an NLP model to summarize it.

[0568] Input: Text data

[0569] Data processing: Summarizing text using NLP models

[0570] Output: Summarized text data

[0571] How it works: Summarize long text data concisely using NLP models such as BERT and GPT-4, for example using the summarization pipeline in the transformers library.

[0572] Step 4:

[0573] The terminal allows the user to select the model of the home appliance through a user interface.

[0574] Input: User selected model information

[0575] Data processing: Generate API requests and send them to the server

[0576] Output: Relevant information received from the server

[0577] Specific operation: The user selects a product model on the initial screen of the application, and an API request is sent to the server based on that information. For example, an HTTP request such as "GET / manuals / AC1234".

[0578] Step 5:

[0579] The server transmits the acquired related information to the terminal.

[0580] Input: User selected model information

[0581] Data processing: Generate related information after search processing

[0582] Output: Related operating instructions and troubleshooting information

[0583] Specific operation: The server receives the request, retrieves the relevant operation instructions and troubleshooting information from the database, and sends them to the terminal.

[0584] Step 6:

[0585] The device uses its camera to scan home appliances and uses AR technology to overlay operating instructions.

[0586] Input: Acquired related information, camera footage

[0587] Data processing: AR technology overlays operation procedures onto camera images

[0588] Output: AR displayed operation procedure

[0589] Specific operation: When a user points the camera at a home appliance, ARCore or ARKit is used to overlay operating instructions on the real-world image.

[0590] Step 7:

[0591] A user enters a question using the chatbot function within the application.

[0592] Input: User's question text

[0593] Data processing: Parse the question with an NLP model

[0594] Output: Parsed question

[0595] Specific operation: A question is entered into the chatbot, and the question is analyzed using an NLP model.

[0596] Step 8:

[0597] The server generates an appropriate answer to the analyzed question and sends it to the terminal.

[0598] Input: Parsed question content

[0599] Data processing: Generative AI models generate answers

[0600] Output: The generated answer

[0601] Specific operation: The server inputs the received question into a generative AI model to generate an appropriate answer. For example, it uses GPT-4 to generate an answer and sends it to the device.

[0602] Step 9:

[0603] The device displays the generated answer on the chatbot screen.

[0604] Input: Generated Answer

[0605] Data processing: None

[0606] Output: The answer displayed on the chatbot screen

[0607] Specific operation: The generated answer text is displayed on the chatbot screen for the user to see.

[0608] Step 10:

[0609] The server records user operation history and setting changes and learns from them using machine learning algorithms.

[0610] Input: User operation history and setting change information

[0611] Data processing: Learning processing using machine learning algorithms

[0612] Output: Learning results

[0613] Specific operation: Collect operation history and setting change data, and cluster behavioral patterns using, for example, Scikit-learn.

[0614] Step 11:

[0615] The server generates individual operation advice and suggestions for setting changes based on the learning results and sends them to the device.

[0616] Input: Learning results

[0617] Data processing: generating personalized advice and suggestions for changing settings

[0618] Output: Generated advice and suggestions

[0619] Specific operation: Generates personalized advice based on user behavior data and sends it to the device.

[0620] (Application example 1)

[0621] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0622] Instruction manuals for home appliances are often provided in paper format, and understanding them takes time and effort. Furthermore, when users encounter difficulties setting up or operating the device, it is difficult to intuitively understand the specific operation method. Furthermore, with the recent spread of smart security systems, there is a demand for intuitive guidance for their installation and configuration. Conventional technologies lack systems that provide comprehensive, immediate, and appropriate support, so there is a need to solve these issues.

[0623] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0624] In this invention, the server includes means for acquiring instruction manual data and extracting text data using optical character recognition, means for summarizing the extracted text data using a natural language processing model, means for selecting a home appliance model through a user interface and providing related operating instructions and troubleshooting information, means for scanning the home appliance using a camera and visualizing operating procedures using augmented reality technology, means for answering questions from the user using a support chatbot using the natural language processing model, means for recording user operation history and setting changes and learning user preferences and usage habits using a machine learning algorithm, means for providing individual operating advice and suggesting setting changes based on the learning results, and means for visualizing installation and configuration methods for the smart security system using augmented reality technology, thereby enabling intuitive and efficient operation and configuration of home appliances and smart security systems.

[0625] An "instruction manual" is a document that provides users with information on how to operate, set up, and maintain home appliances and other devices.

[0626] "Optical character recognition" is a technology that extracts text data from images and documents.

[0627] A "natural language processing model" is an algorithm or system for understanding and generating human language.

[0628] "User interface" refers to the means or screen through which a user interacts with a system.

[0629] A "camera" is a device that receives light and captures image information.

[0630] "Augmented reality technology" is a technology that displays digital information superimposed on images of the real world.

[0631] A "support chatbot" is software that automatically answers users' questions via text or voice.

[0632] A "machine learning algorithm" is a computational method for learning patterns from data and making predictions and classifications.

[0633] "Preferences" refer to the preferences and tendencies of an individual.

[0634] A "smart security system" is a system that automatically monitors and manages safety using network-connected cameras and sensors.

[0635] An "operating procedure" refers to the steps or methods for using a piece of equipment or system.

[0636] "Troubleshooting Information" means information containing guidance or advice for resolving equipment malfunctions or errors.

[0637] "Operation history" is a record of when a user uses a system or device.

[0638] "Configuration change" refers to the change made to adjust the operation of a system or device.

[0639] This invention is an instruction manual system for a smart security system that utilizes generative AI and augmented reality (AR) technology, allowing users to intuitively and quickly obtain operating instructions and troubleshooting information for security devices. The invention can be implemented as follows through the roles of a server, a terminal, and a user.

[0640] Overall system overview

[0641] This system digitally acquires the instruction manual for the smart security system and extracts the text data using optical character recognition (OCR) technology. The extracted text data is summarized using a natural language processing (NLP) model and provided as information to assist the user in operation. It also utilizes AR technology to allow users to visually understand the product's operation procedures and installation methods using their smartphone. Furthermore, a support chatbot responds to user inquiries, providing a means for users to instantly obtain the information they need.

[0642] Program processing overview

[0643] 1. The server retrieves the instruction manual data for the smart security system from the database. The retrieved data is converted into text data using OCR technology (e.g., Google Cloud Vision API) and then summarized using an NLP model (e.g., OpenAI GPT-3). This converts the long instruction manual into a short, easy-to-understand format.

[0644] 2. The device (smartphone, smart glasses, head-mounted display) launches the application and prompts the user to select the security device model through the user interface. Based on the selected model, it communicates with the server to obtain relevant operating instructions and troubleshooting information.

[0645] 3. The device allows users to use the camera to scan for security devices, and uses AR technology (e.g., Apple ARKit, Google ARCore) to overlay operation instructions and installation locations on the camera image, providing users with a visual guide.

[0646] 4. The user enters a question using the chatbot function within the application. The chatbot uses an NLP model to analyze the question and generate an appropriate answer.

[0647] 5. The server records the user's operation history and setting changes, and uses machine learning algorithms to learn the user's preferences and usage habits. Based on this learning, it generates personalized operation advice and setting change suggestions and sends them to the device.

[0648] Specific Examples

[0649] Security camera Wi-Fi settings

[0650] 1. The user opens the application and selects the security camera model.

[0651] 2. The device requests the selected model information from the server, and the server returns the operation instructions and setting data for the corresponding model.

[0652] 3. The device prompts the user to scan the security camera and overlays Wi-Fi setup instructions on the scanned camera image, allowing the user to visually confirm the specific steps.

[0653] 4. When a user asks the chatbot, "I don't know how to set up Wi-Fi," the server analyzes the question, generates detailed instructions, and sends them to the device.

[0654] 5. The device displays the generated instructions on the chatbot screen, and the user follows them to complete the setup.

[0655] Prompt Sentence Examples

[0656] "Please tell me about the Wi-Fi settings for the security camera."

[0657] "I'd like to set up the initial settings for my security system. Please tell me the procedure."

[0658] "Please show me in AR where to place the motion sensor."

[0659] In this way, the system provides users with intuitive and efficient operation and troubleshooting information for their smart security system, making it easier to understand the instruction manual and providing optimal operation advice tailored to the user's individual preferences and usage habits.

[0660] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0661] Step 1:

[0662] The server retrieves the instruction manual data for the smart security system from the database. The input is a request for instruction manual data, and the output is the retrieved instruction manual data. Using OCR technology (e.g., Google Cloud Vision API), the image data is converted into text data. Specifically, the server accesses the database upon receiving the request, retrieves the image data of the specified instruction manual, and applies OCR technology to extract the text data.

[0663] Step 2:

[0664] The server summarizes the extracted text data by running it through a natural language processing (NLP) model (e.g., OpenAI GPT-3). The input is the OCR-processed text data, and the output is the summarized text data. Specifically, the server inputs the text data into the NLP model and receives the generated summary data.

[0665] Step 3:

[0666] The terminal (smartphone, smart glasses, head-mounted display) allows the user to select a security device model through a user interface. The input is the user's model selection operation, and the output is the selected model information. Specifically, the terminal presents multiple models through the display interface and collects the information selected by the user.

[0667] Step 4:

[0668] The terminal requests the selected model information from the server, and the server provides the related operation method and troubleshooting information. The input is the model information request, and the output is the operation method and troubleshooting information. In concrete terms, the terminal sends a request to the server, and the server searches for the corresponding model information and returns it to the terminal.

[0669] Step 5:

[0670] The device allows users to scan security devices using a camera. The input is the camera image, and the output is an overlay display of location information and operation procedures. The display is achieved using augmented reality technology (e.g., Apple ARKit, Google ARCore). Specifically, the device acquires real-time images from the camera and uses AR technology to overlay the necessary guide information on the image.

[0671] Step 6:

[0672] A user inputs a question using the chatbot function within the application. The chatbot uses an NLP model to analyze the input question and generate an appropriate answer. The input is the user's question, and the output is the generated answer. Specifically, the user inputs a question as text, the chatbot sends the question to the NLP model for analysis, and then displays the generated answer to the user.

[0673] Step 7:

[0674] The server records the user's operation history and setting changes, and uses a machine learning algorithm to learn the user's preferences and usage habits. The input is the operation history and setting change data, and the output is individual operation advice and setting change suggestions based on the learning results. Specifically, the server records the operation history in a database, analyzes the data using a machine learning algorithm, and generates individually optimized advice and setting change suggestions.

[0675] Step 8:

[0676] The device provides the user with personalized operation advice and suggestions for setting changes based on the learning results. The input is the suggestion data sent from the server, and the output is a presentation to the user. Specifically, the device receives the suggestion data and displays it on the user interface.

[0677] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0678] The present invention is a home appliance instruction manual system that utilizes generative AI, augmented reality (AR) technology, and an emotion engine, allowing users to intuitively and quickly obtain operation instructions and troubleshooting information for home appliances. The present invention can be implemented as follows through the roles of a server, a terminal, and a user.

[0679] Overall system overview

[0680] This system acquires digital versions of instruction manuals for home appliances, extracts text data using optical character recognition (OCR), summarizes it using natural language processing (NLP) models, and provides it as information to assist users in their operations. Augmented reality (AR) technology also allows users to visually understand product operation procedures using their smartphones. A support chatbot responds to user inquiries, and an emotion engine recognizes the user's emotions and provides optimal advice.

[0681] Program processing overview

[0682] 1. The server retrieves the instruction manual data for the home appliance from the database. The retrieved data is converted into text data using OCR technology and summarized using an NLP model, converting the long instructions into a short, easy-to-understand format.

[0683] 2. The device launches the application and prompts the user to select the model of the home appliance through the user interface. Based on the selected model, the device communicates with the server to obtain relevant operating instructions and troubleshooting information.

[0684] 3. The device allows users to scan home appliances using their smartphone camera, and uses AR technology to overlay operating instructions on the camera image, providing users with a visual guide.

[0685] 4. The user enters a question using the chatbot function within the application. The chatbot uses an NLP model to analyze the question and generate an appropriate answer.

[0686] 5. The server records the user's operation history and setting changes, and uses a machine learning algorithm to learn the user's preferences and usage habits. Based on this learning, it generates individualized operation advice and setting change suggestions and sends them to the device.

[0687] 6. Recognize user emotions using an emotion engine. Analyze facial expressions and voice recorded through the camera while the user is operating the device, and evaluate the user's emotional state in real time.

[0688] 7. The server adjusts its operational advice and suggestions for setting changes based on the emotional data obtained by the emotion engine. For example, if the user is confused, it will provide more detailed and gentler explanations.

[0689] Specific Examples

[0690] The following is a specific scenario.

[0691] Changing the air conditioner remote control settings

[0692] 1. The user opens the application and selects the air conditioner model.

[0693] 2. The device requests the selected model information from the server, and the server returns the operation instructions and setting data for the corresponding model.

[0694] 3. The device instructs the user to scan the air conditioner remote control with the camera, and overlays operating instructions on the scanned image of the remote control, allowing the user to visually confirm the specific button operations.

[0695] 4. When a user asks the chatbot, "I don't know how to set it up," the server analyzes the question, generates detailed instructions, and sends them to the device.

[0696] 5. The device displays the generated instructions on the chatbot screen, and the user follows them to complete the setup.

[0697] 6. If the emotion engine detects confusion or stress in the user's facial expression, the server will adjust the advice accordingly and send a more understandable explanation to the device.

[0698] In this way, the system not only provides users with intuitive and efficient instructions on how to operate home appliances and troubleshooting information, making it easier to understand instruction manuals, but also provides appropriate support according to the user's emotional state.

[0699] The processing flow will be explained below.

[0700] Step 1:

[0701] The server queries the database to retrieve the instruction manual (PDF format) for the home appliance, and stores the retrieved PDF file in local storage or temporarily in memory.

[0702] Step 2:

[0703] The server extracts text data from the PDF using OCR (Optical Character Recognition) technology. Using an OCR library such as Tesseract, it analyzes all pages in the PDF and converts them into text data.

[0704] Step 3:

[0705] The server then runs the extracted text data through a natural language processing (NLP) model to summarize it, using models like BERT and GPT to reduce redundant descriptions and extract key information.

[0706] Step 4:

[0707] The device launches the application and displays the user interface, initially displaying a search bar and drop-down lists for selecting appliance categories and models.

[0708] Step 5:

[0709] The user selects the model of the home appliance they are using on the interface, and the selected model information is sent to the server, requesting related operating instructions and troubleshooting information.

[0710] Step 6:

[0711] Based on the received model information, the server retrieves relevant operating instructions and troubleshooting information and sends it to the device. The information is retrieved from a database and appropriately formatted before being sent.

[0712] Step 7:

[0713] The device responds to user requests and displays the received instructional and troubleshooting information, appropriately laid out in the app's interface.

[0714] Step 8:

[0715] The device will activate the camera, allowing the user to scan home appliances, capture camera footage in real time, and provide guidance to the user.

[0716] Step 9:

[0717] The device uses augmented reality (AR) technology to overlay operation instructions on the camera image, using libraries such as ARKit and ARCore to overlay instruction icons and text on the operation panel.

[0718] Step 10:

[0719] Users operate home appliances by following the instructions on the camera footage, and follow the on-screen instructions and guides to complete the required operations.

[0720] Step 11:

[0721] Users enter questions into the in-app chatbot, using natural language, and submit the question.

[0722] Step 12:

[0723] The server analyzes the questions received from users through a natural language processing model, interprets the input text to understand its meaning, and generates an appropriate answer.

[0724] Step 13:

[0725] The server sends the generated answer to the terminal, where it is formatted appropriately and displayed on the chatbot screen.

[0726] Step 14:

[0727] The device displays the received response on the chatbot's conversation screen, and the user can confirm the displayed information and perform any necessary operations or settings.

[0728] Step 15:

[0729] The device records the user's operation history and setting changes, and sends the operation details and timestamps to the server as log data.

[0730] Step 16:

[0731] The server compiles the received operation history and setting change data and uses machine learning algorithms to learn the user's preferences and usage habits.

[0732] Step 17:

[0733] The server generates individualized operation advice and setting change suggestions based on the learning results, and sends the suggestions to the device as a notification.

[0734] Step 18:

[0735] The device will notify the user of the received suggestions, either as a push notification or a pop-up message, suggesting new settings or operation methods to the user.

[0736] Step 19:

[0737] The device captures facial expression data through the camera while the user is operating the device, and transmits it to the emotion engine, which analyzes facial expressions and voice to evaluate the user's emotional state in real time.

[0738] Step 20:

[0739] The emotion data recognized by the emotion engine is sent to the server, which then adjusts the advice and suggested settings changes accordingly.

[0740] Step 21:

[0741] If the user's emotional state is confused or stressed, the server generates a more detailed and gentle explanation and sends it to the device, improving the quality of support according to the user's emotional state.

[0742] Example 2

[0743] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0744] Modern home appliances are becoming increasingly multifunctional, making it difficult to accurately understand their operation methods and troubleshooting information. Traditional instruction manuals are provided in paper or digital format, but the volume of information makes it difficult for users to quickly find the information they need. Furthermore, the lack of flexible responses or personalized advice based on the user's emotional state leaves the user with a poor user experience. Furthermore, there is a need for systems that can learn users' preferences and usage habits and suggest optimal operations and settings.

[0745] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0746] In this invention, the server includes means for acquiring instruction manual data and extracting character data using optical character recognition, means for summarizing the extracted character data through a natural language processing model, means for recording a user's operation history and setting changes and learning the user's preferences and usage habits using a machine learning algorithm, means for recognizing the user's emotional state using an emotion analysis engine, and means for adjusting operation advice and setting change suggestions based on the results of the emotion analysis engine. This allows the user to quickly obtain the necessary information and operate the home appliance intuitively and efficiently. In addition, to improve the user experience, the user can receive optimal advice and setting suggestions based on the user's emotional state.

[0747] "Home appliances" refer to electronic devices and electrically powered devices that are generally used in the home, including, for example, air conditioners, refrigerators, washing machines, and the like.

[0748] An "instruction manual" is a document that describes how to use, install, and troubleshoot a home appliance device.

[0749] "Digitization" is the process of converting paper or analog information into digital form, making it possible to view, store, and process the information on computers and other digital devices.

[0750] Optical character recognition (OCR) is a technology that extracts character data from images and scanned documents, allowing handwritten or printed characters to be recognized as digital text.

[0751] A "natural language processing (NLP) model" is an artificial intelligence technique for understanding, analyzing, and processing natural language, and can perform tasks such as text summarization, translation, and sentiment analysis.

[0752] A "user interface" is a medium through which a user interacts with a system, and includes a graphical user interface (GUI) and a voice user interface (VUI).

[0753] An "image acquisition device" is a device for acquiring image data, such as a camera or scanner, and particularly includes cameras installed in smartphones.

[0754] Augmented reality (AR) technology is a technology that displays digital information overlaid on images of the real world, allowing users to visually recognize the real world and virtual information in an integrated manner.

[0755] A "support chatbot" is a program that automatically responds to questions and requests from users, and uses natural language processing technology to provide appropriate answers.

[0756] A "machine learning algorithm" is a computational method for analyzing data, learning patterns, and making predictions. It is particularly used to analyze user behavior patterns and suggest optimal operations and settings.

[0757] An "emotion analysis engine" is an artificial intelligence technology that analyzes a user's facial expressions and voice data to evaluate their emotional state, and can provide feedback according to the user's psychological state.

[0758] "Operation history" is a record of operations and setting changes performed by a user using the system, and this information is used to understand the user's preferences and behavioral patterns.

[0759] "Adjusting suggestions" means dynamically changing the appropriate advice and recommendations for setting changes to the user based on collected data and analysis results.

[0760] The present invention is a user manual system for home appliances that utilizes generative AI models, augmented reality (AR) technology, and an emotion engine, allowing users to intuitively and quickly obtain operating instructions and troubleshooting information for home appliances. The present invention is implemented through the roles of a server, a terminal, and a user. This improves user convenience and satisfaction.

[0761] Hardware and Software Use

[0762] Server: Database, Optical Character Recognition (OCR) technology, Natural Language Processing (NLP) models, Sentiment Analysis Engine, Machine Learning Algorithms

[0763] Devices: Smartphones, cameras, and applications using augmented reality (AR) technology

[0764] User: Operating the application, using the camera, using the chat function

[0765] Data acquisition and preprocessing

[0766] The server retrieves instruction manual data for a home appliance from the database. For example, it retrieves instruction manual data for an air conditioner. The retrieved data is often in PDF or image file format.

[0767] Text extraction and summarization

[0768] The server converts the acquired instruction manual data into text data using optical character recognition (OCR) technology. For example, it uses Tesseract OCR to extract text from PDFs and images. The extracted text data is then input into a natural language processing (NLP) model (e.g., BERT, GPT-3) to generate a summary. For example, it can extract only important setup steps and necessary information from long instructions and create a short summary.

[0769] Product Scanning and Display

[0770] The terminal provides an application that allows the user to select the model of a home appliance. When the user selects "air conditioner," the terminal requests the model information from the server, and the server sends related operating instructions and troubleshooting information to the terminal.

[0771] Inquiry response

[0772] The user uses the app's chatbot function to input a question, such as "How do I set a timer?" The server analyzes the question using a natural language processing (NLP) model, generates an appropriate answer, and sends it to the device.

[0773] Operation history recording and learning

[0774] The server records the user's operation history and setting changes, and uses a machine learning algorithm to learn the user's preferences and usage habits. For example, it records a history such as "Timer settings were changed on October 1, 2023," and based on this, suggests optimal settings for users who often use the device at night.

[0775] Recognition of emotional states

[0776] The emotion analysis engine analyzes the user's facial expressions and voice via a camera and microphone to assess their emotional state. For example, it may recognize that the user is confused. The server records this emotional data and generates appropriate feedback.

[0777] Providing advice and coordination

[0778] The server tailors operational advice and setting change suggestions based on the results of the emotion analysis engine. For example, if the user is confused, it generates detailed and easy-to-understand instructions and sends them to the device.

[0779] Examples of concrete examples and prompts

[0780] For example, when a user wants to change the remote control settings of an air conditioner, the specific operating procedure is as follows.

[0781] 1. The user opens the application and selects the air conditioner model.

[0782] 2. The device requests the selected model information from the server, and the server returns the operation instructions and setting data for the corresponding model.

[0783] 3. The device instructs the user to scan the air conditioner remote control with the camera and overlays operating instructions on the scanned image.

[0784] 4. When a user asks the chatbot, "I don't know how to set it up," the server analyzes the question, generates detailed instructions, and sends them to the device.

[0785] 5. The device displays the generated instructions on the chatbot screen, and the user follows them to complete the setup.

[0786] 6. The emotion engine detects confusion or stress from the user's facial expressions, and the server adjusts the advice based on this and sends a more understandable explanation to the device.

[0787] As an example of a prompt sentence, you could feed the generative AI model the following:

[0788] "How do I set the timer on the air conditioner remote control?"

[0789] "What should I do if a user is confused?"

[0790] In this way, the system not only provides users with intuitive and efficient operating instructions and troubleshooting information for home appliances, making it easier to understand instruction manuals, but also provides appropriate support according to the user's emotional state.

[0791] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0792] Step 1:

[0793] The server retrieves instruction manual data for home appliances from a database. The retrieved data may be in PDF or image format. The input data may be, for example, "air conditioner instruction manual data," which is then subjected to OCR processing on the server side. The server then converts this data into text data using optical character recognition (OCR) technology. The output is the text data for the instruction manual. For example, the text may include "How to operate the air conditioner" or "Troubleshooting information."

[0794] Step 2:

[0795] The server inputs the character data extracted by OCR into a natural language processing (NLP) model to generate a summary. The character data is used as input data and passed to an NLP model (e.g., BERT or GPT-3). Specifically, the text is summarized using an application with a summarization function. As output, a summary text is generated that extracts only the important steps from the detailed description. For example, "Description of all buttons on a remote control" is summarized as "Description of the main buttons."

[0796] Step 3:

[0797] The terminal launches an application to allow the user to select a model of a home appliance. The user operates the terminal's user interface to select a model, such as "air conditioner." The user's model selection information is used as input. Based on this, the terminal sends a request for model information to the system. The model selection information is sent to the server as output.

[0798] Step 4:

[0799] Based on the request from the terminal, the server returns the operation method and troubleshooting information for the relevant model. The input is the model information selected by the user, and based on that, the server retrieves the relevant information from the database. Specifically, this includes detailed data such as operation procedures and how to deal with errors. As output, the server sends the related operation method and troubleshooting information to the terminal.

[0800] Step 5:

[0801] The device instructs the user to scan the home appliance with a camera and visualizes the operation procedure using augmented reality (AR) technology. The input is the image data scanned by the user with the camera. Specifically, AR technology is used to overlay instructions such as "Press the POWER button" on the camera image. The output is the visualized operation procedure provided to the user.

[0802] Step 6:

[0803] Users can use the chatbot function within the app to input specific questions, such as, "I don't know how to set the timer on my air conditioner." This question is then sent to the chatbot.

[0804] Step 7:

[0805] The server analyzes questions received via the chatbot using a natural language processing (NLP) model, generates appropriate answers, and sends them to the device. The input is the user's question text. Based on this, the NLP model operates and generates text with specific setup procedures and operation instructions. As output, the generated answer text is sent to the device and displayed on the chatbot screen.

[0806] Step 8:

[0807] The server records the user's operation history and setting changes, and uses a machine learning algorithm to learn the user's preferences and usage habits. The input is the user's operation history data. Specifically, a history such as "Change timer settings on October 1, 2023" is saved. The output is recommended settings and operation methods based on the learning results.

[0808] Step 9:

[0809] The emotion analysis engine analyzes the user's facial expressions and voice via a camera and microphone to evaluate their emotional state. The inputs include facial expression data captured by the camera and voice data captured by the microphone. Specifically, it analyzes in real time whether the user is confused or not. The output is the emotion analysis result.

[0810] Step 10:

[0811] The server adjusts operational advice and setting change suggestions based on the results of the sentiment analysis engine. The input is the sentiment analysis data. Specifically, it prepares detailed and friendly explanations for confused users. The output is the adjusted advice and setting change suggestions sent to the device.

[0812] The above is the specific processing flow of the program of this system.

[0813] (Application example 2)

[0814] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0815] Conventional instruction manual systems for home appliances often make it difficult for users to intuitively understand how to operate them, especially when it comes to complex operations and troubleshooting. Furthermore, due to a lack of systems utilizing emotion engines and generative AI, users' emotional state and individual operational support are insufficient. This results in reduced user operational efficiency and longer troubleshooting times. Furthermore, there is a demand for more efficient picking operations in warehouses, but intuitive support using AR technology is lacking.

[0816] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0817] In this invention, the server includes means for acquiring instruction manual data and extracting text data by optical character recognition, means for summarizing the extracted text data by applying a natural language processing model, and means for selecting a device model through a user interface and providing related operating instructions and troubleshooting information, thereby enabling intuitive and rapid provision of operating instructions and troubleshooting information to the user.

[0818] The server also includes a means for scanning devices using a camera and visualizing operation procedures using augmented reality technology, a means for answering user questions using a support chatbot using a natural language processing model, a means for recording user operation history and setting changes and learning user preferences and usage habits using a machine learning algorithm, a means for providing individual operation advice and suggesting setting changes based on the learning results, a means for proposing work procedures using a generative AI model, a means for analyzing the user's emotional state using an emotion engine and adjusting the operation procedures, and a means for displaying item picking information and supporting the user using augmented reality technology. This makes it possible to provide optimal support according to the user's emotional state and behavioral patterns, thereby improving operation efficiency and work accuracy.

[0819] "Instruction Data" means digital instruction manual information including operating instructions and troubleshooting information for home appliances and other devices.

[0820] Optical character recognition is a technology that automatically reads letters and numbers from images or handwritten characters and converts them into digital text.

[0821] A "natural language processing model" is a machine learning model for understanding, analyzing, and generating human language, allowing it to summarize text data and answer questions.

[0822] A "user interface" is an interaction mechanism that provides a screen and operating methods for users to interact with a system.

[0823] "Augmented reality technology" is a technology that displays digital information overlaid on the real environment, and is used to overlay operating procedures and instructions on camera images.

[0824] A "support chatbot" is a program that uses a natural language processing model to automatically respond to questions from users.

[0825] A "machine learning algorithm" is an algorithm that learns patterns using large amounts of data and makes predictions and classifications for new data.

[0826] A "generative AI model" is an artificial intelligence model that generates new content based on existing data, and is particularly used to generate text, audio, images, etc.

[0827] An "emotion engine" is a system that analyzes a person's emotional state from facial expressions and voice, and responds adaptively based on the results.

[0828] "Item picking information" is information about the work of picking items from warehouses and logistics centers, and includes a picking list and the like.

[0829] This invention is a logistics center support system that utilizes generative AI models, augmented reality (AR) technology, and an emotion engine to enable users to intuitively and quickly perform item picking tasks. This system can be implemented as follows through the roles of the server, terminal, and user.

[0830] Overall system overview

[0831] The system acquires item picking information in digital format and extracts text data using optical character recognition (OCR) technology. The extracted data is summarized using a natural language processing (NLP) model and provided as information to assist the user in their work. AR technology also allows users to visually understand the picking process using smart glasses. A support chatbot answers user inquiries, and an emotion engine recognizes the user's emotions to provide optimal assistance.

[0832] Program processing overview

[0833] The server retrieves item picking information from a database. The retrieved data is converted into text data using OCR technology and summarized using an NLP model. The server then uses a generative AI model to suggest work procedures. The user obtains relevant picking information through a device equipped with a user interface. The device scans the item using smart glasses and overlays the picking list information using AR technology. The user can intuitively perform the work using the smart glasses, following the visual guide. When the user inputs a question using the smart glasses' chatbot function, the server uses an NLP model to generate an appropriate answer and sends it to the device. The server records the user's work history and setting changes and uses a machine learning algorithm to learn the user's preferences and usage habits. Based on this learning result, it generates individual work procedures and setting change suggestions and sends them to the device. The emotion engine analyzes the user's emotional state from their facial expressions and voice and adjusts the operating procedures. The server adjusts assistance advice and provides more appropriate explanations based on the emotional data obtained by the emotion engine.

[0834] Specific Examples

[0835] The following is a specific scenario.

[0836] Support for picking work in the warehouse

[0837] 1. The server retrieves item picking information from the database.

[0838] 2. The server converts the acquired picking information into text data using OCR technology.

[0839] 3. The server summarizes the converted text data using an NLP model.

[0840] 4. The server uses the generative AI model to generate the optimal picking procedure.

[0841] 5. The terminal overlays the picking list information to the user through the smart glasses.

[0842] 6. The user wears the smart glasses and follows visual guidance to pick the item.

[0843] 7. When a user uses the chatbot function to ask a question, the server uses an NLP model to generate an appropriate answer and sends it to the device.

[0844] 8. The emotion engine assesses the user's emotional state from their facial expressions and voice, and adjusts the assistance provided depending on the problem that has occurred.

[0845] Example prompts for generative AI models

[0846] "Picking List Information: Box 123, Section A1, Item 345

[0847] Please suggest the best warehouse operation procedure based on the information below:

[0848] In this way, the system not only provides users with intuitive and efficient support for picking tasks, improving work efficiency within the warehouse, but also provides appropriate support according to the user's emotional state.

[0849] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0850] Step 1:

[0851] The server retrieves item picking information from the database. It uses the picking list information stored in the database as input to obtain the necessary picking data. It obtains raw data for OCR processing as output. Specific operations include issuing a database query to obtain the list information.

[0852] Step 2:

[0853] The server converts the acquired picking information into text data using OCR technology. It uses the raw data obtained in step 1 as input and uses OCR software to obtain the output as digital text. Specifically, it calls an OCR module (e.g., Tesseract) to extract text information from images and documents.

[0854] Step 3:

[0855] The server summarizes the converted text data using a natural language processing (NLP) model. The text data is provided as input to the NLP model, which performs the summarization process and outputs a concise operating procedure. Specifically, an NLP library (e.g., SpaCy or OpenAI GPT-3) is used to generate the summary text.

[0856] Step 4:

[0857] The server uses a generative AI model to propose the optimal picking procedure. The generative AI model receives the summarized text data and the prompt as input. Based on the prompt, the server generates the optimal picking procedure and obtains the procedure as output. Specifically, the server invokes the generative AI (e.g., OpenAI GPT-3), inputs the prompt, and generates the recommended procedure.

[0858] Step 5:

[0859] The terminal overlays the picking list information to the user through the smart glasses. It uses the picking procedure sent from the server as input. It uses AR technology to overlay a visual guide on the smart glasses display and provides visual assistance as output. Specific operations include using AR software (e.g., Unity or ARKit) to display the information on the glasses display.

[0860] Step 6:

[0861] The user wears smart glasses and follows visual guidance to pick items. The input is the information displayed on the smart glasses. The output is to accurately pick the specified items. Specific actions involve looking at the display on the glasses, picking out the items as instructed, and placing them on a cart or other device for inspection.

[0862] Step 7:

[0863] When a user uses the chatbot function to ask a question, the server uses an NLP model to generate an appropriate answer and sends it to the device. The user's question text is used as input. The NLP model analyzes it and outputs the answer. Specifically, the natural language processing model analyzes the user's question, generates an appropriate answer text, and displays it on the chat screen.

[0864] Step 8:

[0865] The emotion engine evaluates the user's emotional state from their facial expressions and voice, and adjusts the assistance provided depending on the problem that has occurred. The input is the user's video and audio data. The emotion engine analyzes the data and outputs the user's emotional state. Specifically, it uses facial recognition technology (e.g., DeepFace or face_recognition) and voice analysis technology to analyze the user's emotions and adjust appropriate advice and procedures.

[0866] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0867] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0868] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0869] [Third embodiment]

[0870] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0871] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0872] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0873] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0874] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0875] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0876] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0877] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0878] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0879] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0880] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0881] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0882] This invention is a home appliance instruction manual system that utilizes generative AI and augmented reality (AR) technology, allowing users to intuitively and quickly obtain operation methods and troubleshooting information for home appliances. The invention can be implemented as follows through the roles of a server, a terminal, and a user.

[0883] Overall system overview

[0884] This system acquires digital versions of instruction manuals for home appliances and extracts text data using optical character recognition (OCR) technology. The extracted text data is summarized using a natural language processing (NLP) model and provided as information to assist users in their operations. It also utilizes AR technology to enable users to visually understand product operation procedures using their smartphones. Furthermore, a support chatbot responds to user inquiries, providing a means for users to instantly obtain the information they need.

[0885] Program processing overview

[0886] 1. The server retrieves the instruction manual data for the home appliance from the database. The retrieved data is converted into text data using OCR technology and summarized using an NLP model. This converts the long instructions into a short, easy-to-understand format.

[0887] 2. The device launches the application and prompts the user to select the model of the home appliance through the user interface. Based on the selected model, the device communicates with the server to obtain relevant operating instructions and troubleshooting information.

[0888] 3. The device allows users to scan home appliances using their smartphone camera, and uses AR technology to overlay operating instructions on the camera image, providing users with a visual guide.

[0889] 4. The user enters a question using the chatbot function within the application. The chatbot uses an NLP model to analyze the question and generate an appropriate answer.

[0890] 5. The server records the user's operation history and setting changes, and uses a machine learning algorithm to learn the user's preferences and usage habits. Based on this learning, it generates individualized operation advice and setting change suggestions and sends them to the device.

[0891] Specific Examples

[0892] The following is a specific scenario.

[0893] Changing the air conditioner remote control settings

[0894] 1. The user opens the application and selects the air conditioner model.

[0895] 2. The device requests the selected model information from the server, and the server returns the operation instructions and setting data for the corresponding model.

[0896] 3. The device instructs the user to scan the air conditioner remote control with the camera, and overlays operating instructions on the scanned image of the remote control, allowing the user to visually confirm the specific button operations.

[0897] 4. When a user asks the chatbot, "I don't know how to set it up," the server analyzes the question, generates detailed instructions, and sends them to the device.

[0898] 5. The device displays the generated instructions on the chatbot screen, and the user follows them to complete the setup.

[0899] In this way, the system intuitively and efficiently provides users with operation and troubleshooting information for home appliances, making it easier to understand instruction manuals and providing optimal operation advice tailored to the user's individual preferences and usage habits.

[0900] The processing flow will be explained below.

[0901] Step 1:

[0902] The server queries the database to retrieve the instruction manual (PDF format) for the home appliance, and stores the retrieved PDF file in local storage or temporarily in memory.

[0903] Step 2:

[0904] The server uses OCR (Optical Character Recognition) technology to extract text data from the PDF. Using an OCR library such as Tesseract, it analyzes all pages in the PDF and converts them into text data.

[0905] Step 3:

[0906] The server then runs the extracted text data through a natural language processing (NLP) model to summarize it, using models like BERT and GPT to reduce redundant descriptions and extract key information.

[0907] Step 4:

[0908] The device launches the application and displays the user interface, initially displaying a search bar and drop-down lists for selecting appliance categories and models.

[0909] Step 5:

[0910] The user selects the model of the home appliance they are using on the interface, and the selected model information is sent to the server, requesting related operating instructions and troubleshooting information.

[0911] Step 6:

[0912] Based on the received model information, the server retrieves relevant operating instructions and troubleshooting information and sends it to the device. The information is retrieved from a database and appropriately formatted before being sent.

[0913] Step 7:

[0914] The device responds to user requests and displays the received instructional and troubleshooting information, appropriately laid out in the app's interface.

[0915] Step 8:

[0916] The device will activate the camera, allowing the user to scan home appliances, capture camera footage in real time, and provide guidance to the user.

[0917] Step 9:

[0918] The device uses augmented reality (AR) technology to overlay operation instructions on the camera image, using libraries such as ARKit and ARCore to overlay instruction icons and text on the operation panel.

[0919] Step 10:

[0920] Users operate home appliances by following the instructions on the camera footage, and follow the on-screen instructions and guides to complete the required operations.

[0921] Step 11:

[0922] Users enter questions into the in-app chatbot, using natural language, and submit the question.

[0923] Step 12:

[0924] The server analyzes the questions received from users through a natural language processing model, interprets the input text to understand its meaning, and generates an appropriate answer.

[0925] Step 13:

[0926] The server sends the generated answer to the terminal, where it is formatted appropriately and displayed on the chatbot screen.

[0927] Step 14:

[0928] The device displays the received response on the chatbot's conversation screen, and the user can confirm the displayed information and perform any necessary operations or settings.

[0929] Step 15:

[0930] The device records the user's operation history and setting changes, and sends the operation details and timestamps to the server as log data.

[0931] Step 16:

[0932] The server compiles the received operation history and setting change data and uses machine learning algorithms to learn the user's preferences and usage habits.

[0933] Step 17:

[0934] The server generates individualized operation advice and setting change suggestions based on the learning results, and sends the suggestions to the device as a notification.

[0935] Step 18:

[0936] The device will notify the user of the received suggestions, either as a push notification or a pop-up message, suggesting new settings or operation methods to the user.

[0937] Example 1

[0938] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0939] Conventional paper manuals for home appliance operation and troubleshooting information are often difficult for users to understand and take a long time to understand. Furthermore, manuals are often lost, making it difficult to quickly obtain the necessary information. Furthermore, when operation methods or solutions are complex, users can easily become lost. There is a need for a system that can solve these problems and enable users to quickly and intuitively understand and operate home appliances.

[0940] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0941] In this invention, the server includes means for acquiring instruction manual data for the home appliance and extracting text data by optical character recognition, means for summarizing the extracted text data through a natural language processing model, means for providing the contents of the instruction manual in a format that is easy for the user to understand using a generative AI model, and means for prompting the user to ask specific questions using prompt sentences and generating detailed answers in response to the questions, thereby enabling the user to quickly and intuitively obtain complex operating instructions and troubleshooting information.

[0942] "Home appliances" refers to all household electronic devices that users use on a daily basis.

[0943] An "instruction manual" refers to a document that describes how to operate, use, maintain, and troubleshoot a home appliance.

[0944] "Digitalization" refers to the process of converting data in paper or physical form into electronic form.

[0945] "User interface" refers to the screen and input devices that users use to operate the device.

[0946] "Optical character recognition (OCR)" refers to the technology that identifies character information from image data and converts it into digital text.

[0947] "Text data" refers to data that is stored as character information and can be processed electronically.

[0948] A "natural language processing (NLP) model" refers to an artificial intelligence model for understanding, analyzing, and generating human language.

[0949] An "abstract" is a concise summary of the contents of a longer document.

[0950] "Augmented reality (AR) technology" refers to the technology of overlaying digital information onto images of the real world.

[0951] A "support chatbot" is a program that automatically responds to questions from users.

[0952] A "machine learning algorithm" refers to a method that allows a computer to learn on its own based on data and make predictions and classifications.

[0953] A "generative AI model" refers to an artificial intelligence model that generates new sentences and answers based on human language.

[0954] A "prompt sentence" refers to a question-type sentence that guides the user's input.

[0955] This invention is a home appliance instruction manual system that utilizes generative AI models and augmented reality (AR) technology, allowing users to intuitively and quickly obtain operating instructions and troubleshooting information for home appliances. This invention is realized through the roles of a server, a terminal, and a user.

[0956] Overall system overview

[0957] This system acquires the digital version of the appliance's instruction manual and extracts text data using optical character recognition (OCR) technology. The extracted text data is summarized using a natural language processing (NLP) model. Furthermore, a generative AI model is used to present the contents of the instruction manual in a format that is easy for users to understand. The user selects the appliance model through the user interface, and related operating instructions and troubleshooting information are provided. When the user scans the appliance using a smartphone, AR technology is used to overlay operating instructions on the camera image. Furthermore, a support chatbot answers questions from the user, quickly providing the necessary information. Through this process, the user's operation history and setting changes are recorded, and a machine learning algorithm learns the user's preferences and usage habits. Based on the learning results, personalized operating advice and setting change suggestions are provided.

[0958] Hardware and software used

[0959] Server: A common RDBMS (e.g., MySQL, PostgreSQL) is used as the database, Tesseract OCR is used as the OCR technology, and BERT or GPT-4 is used as the NLP model.

[0960] Device: A mobile device such as a smartphone or tablet that uses ARCore (for Android) or ARKit (for iOS) as AR technology.

[0961] Users: Use the user interface and chatbot functions within the application.

[0962] Specific examples

[0963] Changing the air conditioner remote control settings

[0964] 1. The user opens the application and selects the model of the air conditioner.

[0965] 2. The device requests the selected model information from the server, and the server returns the operation instructions and setting data for the corresponding model.

[0966] 3. The device instructs the user to scan the air conditioner remote control with the camera, and overlays operating instructions on the scanned image of the remote control, allowing the user to visually confirm the specific button operations.

[0967] 4. When a user asks the chatbot, "I don't know how to set it up," the server analyzes the question, generates detailed instructions, and sends them to the device.

[0968] 5. The device displays the generated instructions on the chatbot screen, and the user follows them to complete the setup.

[0969] Prompt Sentence Examples

[0970] Example question: "How do I switch my air conditioner to cooling mode?"

[0971] Example prompt: "My question is about how to operate my air conditioner. What are the specific steps to switch it to cooling mode?"

[0972] Based on these prompts, the generative AI model generates easy-to-understand answers for users, allowing them to intuitively and efficiently obtain instructions and troubleshooting information for their home appliances.

[0973] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0974] Step 1:

[0975] The server retrieves instruction manual data for the home appliance from the database.

[0976] Input: Product Model ID

[0977] Data processing: Search instruction manual data using database queries

[0978] Output: Instruction manual data

[0979] Specific operation: The server executes an SQL query based on the product model ID to retrieve the instruction manual data. For example, it uses a query like "SELECT FROM manuals WHERE product_id = 'AC1234';".

[0980] Step 2:

[0981] The instruction manual data acquired by the server is converted into text data using OCR technology.

[0982] Input: Instruction manual data

[0983] Data processing: Applying OCR technology to extract text data from image data

[0984] Output: Extracted text data

[0985] Specific operation: Convert image-format instruction manual data into text using Tesseract OCR. For example, read the image using Python's PIL library and extract the text using pytesseract.

[0986] Step 3:

[0987] The server then applies the extracted text data to an NLP model to summarize it.

[0988] Input: Text data

[0989] Data processing: Summarizing text using NLP models

[0990] Output: Summarized text data

[0991] How it works: Summarize long text data concisely using NLP models such as BERT and GPT-4, for example using the summarization pipeline in the transformers library.

[0992] Step 4:

[0993] The terminal allows the user to select the model of the home appliance through a user interface.

[0994] Input: User selected model information

[0995] Data processing: Generate API requests and send them to the server

[0996] Output: Relevant information received from the server

[0997] Specific operation: The user selects a product model on the initial screen of the application, and an API request is sent to the server based on that information. For example, an HTTP request such as "GET / manuals / AC1234".

[0998] Step 5:

[0999] The server transmits the acquired related information to the terminal.

[1000] Input: User selected model information

[1001] Data processing: Generate related information after search processing

[1002] Output: Related operating instructions and troubleshooting information

[1003] Specific operation: The server receives the request, retrieves the relevant operation instructions and troubleshooting information from the database, and sends them to the terminal.

[1004] Step 6:

[1005] The device uses its camera to scan home appliances and uses AR technology to overlay operating instructions.

[1006] Input: Acquired related information, camera footage

[1007] Data processing: AR technology overlays operation procedures onto camera images

[1008] Output: AR displayed operation procedure

[1009] Specific operation: When a user points the camera at a home appliance, ARCore or ARKit is used to overlay operating instructions on the real-world image.

[1010] Step 7:

[1011] A user enters a question using the chatbot function within the application.

[1012] Input: User's question text

[1013] Data processing: Parse the question with an NLP model

[1014] Output: Parsed question

[1015] Specific operation: A question is entered into the chatbot, and the question is analyzed using an NLP model.

[1016] Step 8:

[1017] The server generates an appropriate answer to the analyzed question and sends it to the terminal.

[1018] Input: Parsed question content

[1019] Data processing: Generative AI models generate answers

[1020] Output: The generated answer

[1021] Specific operation: The server inputs the received question into a generative AI model to generate an appropriate answer. For example, it uses GPT-4 to generate an answer and sends it to the device.

[1022] Step 9:

[1023] The device displays the generated answer on the chatbot screen.

[1024] Input: Generated Answer

[1025] Data processing: None

[1026] Output: The answer displayed on the chatbot screen

[1027] Specific operation: The generated answer text is displayed on the chatbot screen for the user to see.

[1028] Step 10:

[1029] The server records user operation history and setting changes and learns from them using machine learning algorithms.

[1030] Input: User operation history and setting change information

[1031] Data processing: Learning processing using machine learning algorithms

[1032] Output: Learning results

[1033] Specific operation: Collect operation history and setting change data, and cluster behavioral patterns using, for example, Scikit-learn.

[1034] Step 11:

[1035] The server generates individual operation advice and suggestions for setting changes based on the learning results and sends them to the device.

[1036] Input: Learning results

[1037] Data processing: generating personalized advice and suggestions for changing settings

[1038] Output: Generated advice and suggestions

[1039] Specific operation: Generates personalized advice based on user behavior data and sends it to the device.

[1040] (Application example 1)

[1041] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1042] Instruction manuals for home appliances are often provided in paper format, and understanding them takes time and effort. Furthermore, when users encounter difficulties setting up or operating the device, it is difficult to intuitively understand the specific operation method. Furthermore, with the recent spread of smart security systems, there is a demand for intuitive guidance for their installation and configuration. Conventional technologies lack systems that provide comprehensive, immediate, and appropriate support, so there is a need to solve these issues.

[1043] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1044] In this invention, the server includes means for acquiring instruction manual data and extracting text data using optical character recognition, means for summarizing the extracted text data using a natural language processing model, means for selecting a home appliance model through a user interface and providing related operating instructions and troubleshooting information, means for scanning the home appliance using a camera and visualizing operating procedures using augmented reality technology, means for answering questions from the user using a support chatbot using the natural language processing model, means for recording user operation history and setting changes and learning user preferences and usage habits using a machine learning algorithm, means for providing individual operating advice and suggesting setting changes based on the learning results, and means for visualizing installation and configuration methods for the smart security system using augmented reality technology, thereby enabling intuitive and efficient operation and configuration of home appliances and smart security systems.

[1045] An "instruction manual" is a document that provides users with information on how to operate, set up, and maintain home appliances and other devices.

[1046] "Optical character recognition" is a technology that extracts text data from images and documents.

[1047] A "natural language processing model" is an algorithm or system for understanding and generating human language.

[1048] "User interface" refers to the means or screen through which a user interacts with a system.

[1049] A "camera" is a device that receives light and captures image information.

[1050] "Augmented reality technology" is a technology that displays digital information superimposed on images of the real world.

[1051] A "support chatbot" is software that automatically answers users' questions via text or voice.

[1052] A "machine learning algorithm" is a computational method for learning patterns from data and making predictions and classifications.

[1053] "Preferences" refer to the preferences and tendencies of an individual.

[1054] A "smart security system" is a system that automatically monitors and manages safety using network-connected cameras and sensors.

[1055] An "operating procedure" refers to the steps or methods for using a piece of equipment or system.

[1056] "Troubleshooting Information" means information containing guidance or advice for resolving equipment malfunctions or errors.

[1057] "Operation history" is a record of when a user uses a system or device.

[1058] "Configuration change" refers to the change made to adjust the operation of a system or device.

[1059] This invention is an instruction manual system for a smart security system that utilizes generative AI and augmented reality (AR) technology, allowing users to intuitively and quickly obtain operating instructions and troubleshooting information for security devices. The invention can be implemented as follows through the roles of a server, a terminal, and a user.

[1060] Overall system overview

[1061] This system digitally acquires the instruction manual for the smart security system and extracts the text data using optical character recognition (OCR) technology. The extracted text data is summarized using a natural language processing (NLP) model and provided as information to assist the user in operation. It also utilizes AR technology to allow users to visually understand the product's operation procedures and installation methods using their smartphone. Furthermore, a support chatbot responds to user inquiries, providing a means for users to instantly obtain the information they need.

[1062] Program processing overview

[1063] 1. The server retrieves the instruction manual data for the smart security system from the database. The retrieved data is converted into text data using OCR technology (e.g., Google Cloud Vision API) and then summarized using an NLP model (e.g., OpenAI GPT-3). This converts the long instruction manual into a short, easy-to-understand format.

[1064] 2. The device (smartphone, smart glasses, head-mounted display) launches the application and prompts the user to select the security device model through the user interface. Based on the selected model, it communicates with the server to obtain relevant operating instructions and troubleshooting information.

[1065] 3. The device allows users to use the camera to scan for security devices, and uses AR technology (e.g., Apple ARKit, Google ARCore) to overlay operation instructions and installation locations on the camera image, providing users with a visual guide.

[1066] 4. The user enters a question using the chatbot function within the application. The chatbot uses an NLP model to analyze the question and generate an appropriate answer.

[1067] 5. The server records the user's operation history and setting changes, and uses machine learning algorithms to learn the user's preferences and usage habits. Based on this learning, it generates personalized operation advice and setting change suggestions and sends them to the device.

[1068] Specific Examples

[1069] Security camera Wi-Fi settings

[1070] 1. The user opens the application and selects the security camera model.

[1071] 2. The device requests the selected model information from the server, and the server returns the operation instructions and setting data for the corresponding model.

[1072] 3. The device prompts the user to scan the security camera and overlays Wi-Fi setup instructions on the scanned camera image, allowing the user to visually confirm the specific steps.

[1073] 4. When a user asks the chatbot, "I don't know how to set up Wi-Fi," the server analyzes the question, generates detailed instructions, and sends them to the device.

[1074] 5. The device displays the generated instructions on the chatbot screen, and the user follows them to complete the setup.

[1075] Prompt Sentence Examples

[1076] "Please tell me about the Wi-Fi settings for the security camera."

[1077] "I'd like to set up the initial settings for my security system. Please tell me the procedure."

[1078] "Please show me in AR where to place the motion sensor."

[1079] In this way, the system provides users with intuitive and efficient operation and troubleshooting information for their smart security system, making it easier to understand the instruction manual and providing optimal operation advice tailored to the user's individual preferences and usage habits.

[1080] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1081] Step 1:

[1082] The server retrieves the instruction manual data for the smart security system from the database. The input is a request for instruction manual data, and the output is the retrieved instruction manual data. Using OCR technology (e.g., Google Cloud Vision API), the image data is converted into text data. Specifically, the server accesses the database upon receiving the request, retrieves the image data of the specified instruction manual, and applies OCR technology to extract the text data.

[1083] Step 2:

[1084] The server summarizes the extracted text data by running it through a natural language processing (NLP) model (e.g., OpenAI GPT-3). The input is the OCR-processed text data, and the output is the summarized text data. Specifically, the server inputs the text data into the NLP model and receives the generated summary data.

[1085] Step 3:

[1086] The terminal (smartphone, smart glasses, head-mounted display) allows the user to select a security device model through a user interface. The input is the user's model selection operation, and the output is the selected model information. Specifically, the terminal presents multiple models through the display interface and collects the information selected by the user.

[1087] Step 4:

[1088] The terminal requests the selected model information from the server, and the server provides the related operation method and troubleshooting information. The input is the model information request, and the output is the operation method and troubleshooting information. In concrete terms, the terminal sends a request to the server, and the server searches for the corresponding model information and returns it to the terminal.

[1089] Step 5:

[1090] The device allows users to scan security devices using a camera. The input is the camera image, and the output is an overlay display of location information and operation procedures. The display is achieved using augmented reality technology (e.g., Apple ARKit, Google ARCore). Specifically, the device acquires real-time images from the camera and uses AR technology to overlay the necessary guide information on the image.

[1091] Step 6:

[1092] A user inputs a question using the chatbot function within the application. The chatbot uses an NLP model to analyze the input question and generate an appropriate answer. The input is the user's question, and the output is the generated answer. Specifically, the user inputs a question as text, the chatbot sends the question to the NLP model for analysis, and then displays the generated answer to the user.

[1093] Step 7:

[1094] The server records the user's operation history and setting changes, and uses a machine learning algorithm to learn the user's preferences and usage habits. The input is the operation history and setting change data, and the output is individual operation advice and setting change suggestions based on the learning results. Specifically, the server records the operation history in a database, analyzes the data using a machine learning algorithm, and generates individually optimized advice and setting change suggestions.

[1095] Step 8:

[1096] The device provides the user with personalized operation advice and suggestions for setting changes based on the learning results. The input is the suggestion data sent from the server, and the output is a presentation to the user. Specifically, the device receives the suggestion data and displays it on the user interface.

[1097] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1098] The present invention is a home appliance instruction manual system that utilizes generative AI, augmented reality (AR) technology, and an emotion engine, allowing users to intuitively and quickly obtain operation instructions and troubleshooting information for home appliances. The present invention can be implemented as follows through the roles of a server, a terminal, and a user.

[1099] Overall system overview

[1100] This system acquires digital versions of instruction manuals for home appliances, extracts text data using optical character recognition (OCR), summarizes it using natural language processing (NLP) models, and provides it as information to assist users in their operations. Augmented reality (AR) technology also allows users to visually understand product operation procedures using their smartphones. A support chatbot responds to user inquiries, and an emotion engine recognizes the user's emotions and provides optimal advice.

[1101] Program processing overview

[1102] 1. The server retrieves the instruction manual data for the home appliance from the database. The retrieved data is converted into text data using OCR technology and summarized using an NLP model, converting the long instructions into a short, easy-to-understand format.

[1103] 2. The device launches the application and prompts the user to select the model of the home appliance through the user interface. Based on the selected model, the device communicates with the server to obtain relevant operating instructions and troubleshooting information.

[1104] 3. The device allows users to scan home appliances using their smartphone camera, and uses AR technology to overlay operating instructions on the camera image, providing users with a visual guide.

[1105] 4. The user enters a question using the chatbot function within the application. The chatbot uses an NLP model to analyze the question and generate an appropriate answer.

[1106] 5. The server records the user's operation history and setting changes, and uses a machine learning algorithm to learn the user's preferences and usage habits. Based on this learning, it generates individualized operation advice and setting change suggestions and sends them to the device.

[1107] 6. Recognize user emotions using an emotion engine. Analyze facial expressions and voice recorded through the camera while the user is operating the device, and evaluate the user's emotional state in real time.

[1108] 7. The server adjusts its operational advice and suggestions for setting changes based on the emotional data obtained by the emotion engine. For example, if the user is confused, it will provide more detailed and gentler explanations.

[1109] Specific Examples

[1110] The following is a specific scenario.

[1111] Changing the air conditioner remote control settings

[1112] 1. The user opens the application and selects the air conditioner model.

[1113] 2. The device requests the selected model information from the server, and the server returns the operation instructions and setting data for the corresponding model.

[1114] 3. The device instructs the user to scan the air conditioner remote control with the camera, and overlays operating instructions on the scanned image of the remote control, allowing the user to visually confirm the specific button operations.

[1115] 4. When a user asks the chatbot, "I don't know how to set it up," the server analyzes the question, generates detailed instructions, and sends them to the device.

[1116] 5. The device displays the generated instructions on the chatbot screen, and the user follows them to complete the setup.

[1117] 6. If the emotion engine detects confusion or stress in the user's facial expression, the server will adjust the advice accordingly and send a more understandable explanation to the device.

[1118] In this way, the system not only provides users with intuitive and efficient instructions on how to operate home appliances and troubleshooting information, making it easier to understand instruction manuals, but also provides appropriate support according to the user's emotional state.

[1119] The processing flow will be explained below.

[1120] Step 1:

[1121] The server queries the database to retrieve the instruction manual (PDF format) for the home appliance, and stores the retrieved PDF file in local storage or temporarily in memory.

[1122] Step 2:

[1123] The server extracts text data from the PDF using OCR (Optical Character Recognition) technology. Using an OCR library such as Tesseract, it analyzes all pages in the PDF and converts them into text data.

[1124] Step 3:

[1125] The server then runs the extracted text data through a natural language processing (NLP) model to summarize it, using models like BERT and GPT to reduce redundant descriptions and extract key information.

[1126] Step 4:

[1127] The device launches the application and displays the user interface, initially displaying a search bar and drop-down lists for selecting appliance categories and models.

[1128] Step 5:

[1129] The user selects the model of the home appliance they are using on the interface, and the selected model information is sent to the server, requesting related operating instructions and troubleshooting information.

[1130] Step 6:

[1131] Based on the received model information, the server retrieves relevant operating instructions and troubleshooting information and sends it to the device. The information is retrieved from a database and appropriately formatted before being sent.

[1132] Step 7:

[1133] The device responds to user requests and displays the received instructional and troubleshooting information, appropriately laid out in the app's interface.

[1134] Step 8:

[1135] The device will activate the camera, allowing the user to scan home appliances, capture camera footage in real time, and provide guidance to the user.

[1136] Step 9:

[1137] The device uses augmented reality (AR) technology to overlay operation instructions on the camera image, using libraries such as ARKit and ARCore to overlay instruction icons and text on the operation panel.

[1138] Step 10:

[1139] Users operate home appliances by following the instructions on the camera footage, and follow the on-screen instructions and guides to complete the required operations.

[1140] Step 11:

[1141] Users enter questions into the in-app chatbot, using natural language, and submit the question.

[1142] Step 12:

[1143] The server analyzes the questions received from users through a natural language processing model, interprets the input text to understand its meaning, and generates an appropriate answer.

[1144] Step 13:

[1145] The server sends the generated answer to the terminal, where it is formatted appropriately and displayed on the chatbot screen.

[1146] Step 14:

[1147] The device displays the received response on the chatbot's conversation screen, and the user can confirm the displayed information and perform any necessary operations or settings.

[1148] Step 15:

[1149] The device records the user's operation history and setting changes, and sends the operation details and timestamps to the server as log data.

[1150] Step 16:

[1151] The server compiles the received operation history and setting change data and uses machine learning algorithms to learn the user's preferences and usage habits.

[1152] Step 17:

[1153] The server generates individualized operation advice and setting change suggestions based on the learning results, and sends the suggestions to the device as a notification.

[1154] Step 18:

[1155] The device will notify the user of the received suggestions, either as a push notification or a pop-up message, suggesting new settings or operation methods to the user.

[1156] Step 19:

[1157] The device captures facial expression data through the camera while the user is operating the device, and transmits it to the emotion engine, which analyzes facial expressions and voice to evaluate the user's emotional state in real time.

[1158] Step 20:

[1159] The emotion data recognized by the emotion engine is sent to the server, which then adjusts the advice and suggested settings changes accordingly.

[1160] Step 21:

[1161] If the user's emotional state is confused or stressed, the server generates a more detailed and gentle explanation and sends it to the device, improving the quality of support according to the user's emotional state.

[1162] Example 2

[1163] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1164] Modern home appliances are becoming increasingly multifunctional, making it difficult to accurately understand their operation methods and troubleshooting information. Traditional instruction manuals are provided in paper or digital format, but the volume of information makes it difficult for users to quickly find the information they need. Furthermore, the lack of flexible responses or personalized advice based on the user's emotional state leaves the user with a poor user experience. Furthermore, there is a need for systems that can learn users' preferences and usage habits and suggest optimal operations and settings.

[1165] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1166] In this invention, the server includes means for acquiring instruction manual data and extracting character data using optical character recognition, means for summarizing the extracted character data through a natural language processing model, means for recording a user's operation history and setting changes and learning the user's preferences and usage habits using a machine learning algorithm, means for recognizing the user's emotional state using an emotion analysis engine, and means for adjusting operation advice and setting change suggestions based on the results of the emotion analysis engine. This allows the user to quickly obtain the necessary information and operate the home appliance intuitively and efficiently. In addition, to improve the user experience, the user can receive optimal advice and setting suggestions based on the user's emotional state.

[1167] "Home appliances" refer to electronic devices and electrically powered devices that are generally used in the home, including, for example, air conditioners, refrigerators, washing machines, and the like.

[1168] An "instruction manual" is a document that describes how to use, install, and troubleshoot a home appliance device.

[1169] "Digitization" is the process of converting paper or analog information into digital form, making it possible to view, store, and process the information on computers and other digital devices.

[1170] Optical character recognition (OCR) is a technology that extracts character data from images and scanned documents, allowing handwritten or printed characters to be recognized as digital text.

[1171] A "natural language processing (NLP) model" is an artificial intelligence technique for understanding, analyzing, and processing natural language, and can perform tasks such as text summarization, translation, and sentiment analysis.

[1172] A "user interface" is a medium through which a user interacts with a system, and includes a graphical user interface (GUI) and a voice user interface (VUI).

[1173] An "image acquisition device" is a device for acquiring image data, such as a camera or scanner, and particularly includes cameras installed in smartphones.

[1174] Augmented reality (AR) technology is a technology that displays digital information overlaid on images of the real world, allowing users to visually recognize the real world and virtual information in an integrated manner.

[1175] A "support chatbot" is a program that automatically responds to questions and requests from users, and uses natural language processing technology to provide appropriate answers.

[1176] A "machine learning algorithm" is a computational method for analyzing data, learning patterns, and making predictions. It is particularly used to analyze user behavior patterns and suggest optimal operations and settings.

[1177] An "emotion analysis engine" is an artificial intelligence technology that analyzes a user's facial expressions and voice data to evaluate their emotional state, and can provide feedback according to the user's psychological state.

[1178] "Operation history" is a record of operations and setting changes performed by a user using the system, and this information is used to understand the user's preferences and behavioral patterns.

[1179] "Adjusting suggestions" means dynamically changing the appropriate advice and recommendations for setting changes to the user based on collected data and analysis results.

[1180] The present invention is a user manual system for home appliances that utilizes generative AI models, augmented reality (AR) technology, and an emotion engine, allowing users to intuitively and quickly obtain operating instructions and troubleshooting information for home appliances. The present invention is implemented through the roles of a server, a terminal, and a user. This improves user convenience and satisfaction.

[1181] Hardware and Software Use

[1182] Server: Database, Optical Character Recognition (OCR) technology, Natural Language Processing (NLP) models, Sentiment Analysis Engine, Machine Learning Algorithms

[1183] Devices: Smartphones, cameras, and applications using augmented reality (AR) technology

[1184] User: Operating the application, using the camera, using the chat function

[1185] Data acquisition and preprocessing

[1186] The server retrieves instruction manual data for a home appliance from the database. For example, it retrieves instruction manual data for an air conditioner. The retrieved data is often in PDF or image file format.

[1187] Text extraction and summarization

[1188] The server converts the acquired instruction manual data into text data using optical character recognition (OCR) technology. For example, it uses Tesseract OCR to extract text from PDFs and images. The extracted text data is then input into a natural language processing (NLP) model (e.g., BERT, GPT-3) to generate a summary. For example, it can extract only important setup steps and necessary information from long instructions and create a short summary.

[1189] Product Scanning and Display

[1190] The terminal provides an application that allows the user to select the model of a home appliance. When the user selects "air conditioner," the terminal requests the model information from the server, and the server sends related operating instructions and troubleshooting information to the terminal.

[1191] Inquiry response

[1192] The user uses the app's chatbot function to input a question, such as "How do I set a timer?" The server analyzes the question using a natural language processing (NLP) model, generates an appropriate answer, and sends it to the device.

[1193] Operation history recording and learning

[1194] The server records the user's operation history and setting changes, and uses a machine learning algorithm to learn the user's preferences and usage habits. For example, it records a history such as "Timer settings were changed on October 1, 2023," and based on this, suggests optimal settings for users who often use the device at night.

[1195] Recognition of emotional states

[1196] The emotion analysis engine analyzes the user's facial expressions and voice via a camera and microphone to assess their emotional state. For example, it may recognize that the user is confused. The server records this emotional data and generates appropriate feedback.

[1197] Providing advice and coordination

[1198] The server tailors operational advice and setting change suggestions based on the results of the emotion analysis engine. For example, if the user is confused, it generates detailed and easy-to-understand instructions and sends them to the device.

[1199] Examples of concrete examples and prompts

[1200] For example, when a user wants to change the remote control settings of an air conditioner, the specific operating procedure is as follows.

[1201] 1. The user opens the application and selects the air conditioner model.

[1202] 2. The device requests the selected model information from the server, and the server returns the operation instructions and setting data for the corresponding model.

[1203] 3. The device instructs the user to scan the air conditioner remote control with the camera and overlays operating instructions on the scanned image.

[1204] 4. When a user asks the chatbot, "I don't know how to set it up," the server analyzes the question, generates detailed instructions, and sends them to the device.

[1205] 5. The device displays the generated instructions on the chatbot screen, and the user follows them to complete the setup.

[1206] 6. The emotion engine detects confusion or stress from the user's facial expressions, and the server adjusts the advice based on this and sends a more understandable explanation to the device.

[1207] As an example of a prompt sentence, you could feed the generative AI model the following:

[1208] "How do I set the timer on the air conditioner remote control?"

[1209] "What should I do if a user is confused?"

[1210] In this way, the system not only provides users with intuitive and efficient operating instructions and troubleshooting information for home appliances, making it easier to understand instruction manuals, but also provides appropriate support according to the user's emotional state.

[1211] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1212] Step 1:

[1213] The server retrieves instruction manual data for home appliances from a database. The retrieved data may be in PDF or image format. The input data may be, for example, "air conditioner instruction manual data," which is then subjected to OCR processing on the server side. The server then converts this data into text data using optical character recognition (OCR) technology. The output is the text data for the instruction manual. For example, the text may include "How to operate the air conditioner" or "Troubleshooting information."

[1214] Step 2:

[1215] The server inputs the character data extracted by OCR into a natural language processing (NLP) model to generate a summary. The character data is used as input data and passed to an NLP model (e.g., BERT or GPT-3). Specifically, the text is summarized using an application with a summarization function. As output, a summary text is generated that extracts only the important steps from the detailed description. For example, "Description of all buttons on a remote control" is summarized as "Description of the main buttons."

[1216] Step 3:

[1217] The terminal launches an application to allow the user to select a model of a home appliance. The user operates the terminal's user interface to select a model, such as "air conditioner." The user's model selection information is used as input. Based on this, the terminal sends a request for model information to the system. The model selection information is sent to the server as output.

[1218] Step 4:

[1219] Based on the request from the terminal, the server returns the operation method and troubleshooting information for the relevant model. The input is the model information selected by the user, and based on that, the server retrieves the relevant information from the database. Specifically, this includes detailed data such as operation procedures and how to deal with errors. As output, the server sends the related operation method and troubleshooting information to the terminal.

[1220] Step 5:

[1221] The device instructs the user to scan the home appliance with a camera and visualizes the operation procedure using augmented reality (AR) technology. The input is the image data scanned by the user with the camera. Specifically, AR technology is used to overlay instructions such as "Press the POWER button" on the camera image. The output is the visualized operation procedure provided to the user.

[1222] Step 6:

[1223] Users can use the chatbot function within the app to input specific questions, such as, "I don't know how to set the timer on my air conditioner." This question is then sent to the chatbot.

[1224] Step 7:

[1225] The server analyzes questions received via the chatbot using a natural language processing (NLP) model, generates appropriate answers, and sends them to the device. The input is the user's question text. Based on this, the NLP model operates and generates text with specific setup procedures and operation instructions. As output, the generated answer text is sent to the device and displayed on the chatbot screen.

[1226] Step 8:

[1227] The server records the user's operation history and setting changes, and uses a machine learning algorithm to learn the user's preferences and usage habits. The input is the user's operation history data. Specifically, a history such as "Change timer settings on October 1, 2023" is saved. The output is recommended settings and operation methods based on the learning results.

[1228] Step 9:

[1229] The emotion analysis engine analyzes the user's facial expressions and voice via a camera and microphone to evaluate their emotional state. The inputs include facial expression data captured by the camera and voice data captured by the microphone. Specifically, it analyzes in real time whether the user is confused or not. The output is the emotion analysis result.

[1230] Step 10:

[1231] The server adjusts operational advice and setting change suggestions based on the results of the sentiment analysis engine. The input is the sentiment analysis data. Specifically, it prepares detailed and friendly explanations for confused users. The output is the adjusted advice and setting change suggestions sent to the device.

[1232] The above is the specific processing flow of the program of this system.

[1233] (Application example 2)

[1234] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1235] Conventional instruction manual systems for home appliances often make it difficult for users to intuitively understand how to operate them, especially when it comes to complex operations and troubleshooting. Furthermore, due to a lack of systems utilizing emotion engines and generative AI, users' emotional state and individual operational support are insufficient. This results in reduced user operational efficiency and longer troubleshooting times. Furthermore, there is a demand for more efficient picking operations in warehouses, but intuitive support using AR technology is lacking.

[1236] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1237] In this invention, the server includes means for acquiring instruction manual data and extracting text data by optical character recognition, means for summarizing the extracted text data by applying a natural language processing model, and means for selecting a device model through a user interface and providing related operating instructions and troubleshooting information, thereby enabling intuitive and rapid provision of operating instructions and troubleshooting information to the user.

[1238] The server also includes a means for scanning devices using a camera and visualizing operation procedures using augmented reality technology, a means for answering user questions using a support chatbot using a natural language processing model, a means for recording user operation history and setting changes and learning user preferences and usage habits using a machine learning algorithm, a means for providing individual operation advice and suggesting setting changes based on the learning results, a means for proposing work procedures using a generative AI model, a means for analyzing the user's emotional state using an emotion engine and adjusting the operation procedures, and a means for displaying item picking information and supporting the user using augmented reality technology. This makes it possible to provide optimal support according to the user's emotional state and behavioral patterns, thereby improving operation efficiency and work accuracy.

[1239] "Instruction Data" means digital instruction manual information including operating instructions and troubleshooting information for home appliances and other devices.

[1240] Optical character recognition is a technology that automatically reads letters and numbers from images or handwritten characters and converts them into digital text.

[1241] A "natural language processing model" is a machine learning model for understanding, analyzing, and generating human language, allowing it to summarize text data and answer questions.

[1242] A "user interface" is an interaction mechanism that provides a screen and operating methods for users to interact with a system.

[1243] "Augmented reality technology" is a technology that displays digital information overlaid on the real environment, and is used to overlay operating procedures and instructions on camera images.

[1244] A "support chatbot" is a program that uses a natural language processing model to automatically respond to questions from users.

[1245] A "machine learning algorithm" is an algorithm that learns patterns using large amounts of data and makes predictions and classifications for new data.

[1246] A "generative AI model" is an artificial intelligence model that generates new content based on existing data, and is particularly used to generate text, audio, images, etc.

[1247] An "emotion engine" is a system that analyzes a person's emotional state from facial expressions and voice, and responds adaptively based on the results.

[1248] "Item picking information" is information about the work of picking items from warehouses and logistics centers, and includes a picking list and the like.

[1249] This invention is a logistics center support system that utilizes generative AI models, augmented reality (AR) technology, and an emotion engine to enable users to intuitively and quickly perform item picking tasks. This system can be implemented as follows through the roles of the server, terminal, and user.

[1250] Overall system overview

[1251] The system acquires item picking information in digital format and extracts text data using optical character recognition (OCR) technology. The extracted data is summarized using a natural language processing (NLP) model and provided as information to assist the user in their work. AR technology also allows users to visually understand the picking process using smart glasses. A support chatbot answers user inquiries, and an emotion engine recognizes the user's emotions to provide optimal assistance.

[1252] Program processing overview

[1253] The server retrieves item picking information from a database. The retrieved data is converted into text data using OCR technology and summarized using an NLP model. The server then uses a generative AI model to suggest work procedures. The user obtains relevant picking information through a device equipped with a user interface. The device scans the item using smart glasses and overlays the picking list information using AR technology. The user can intuitively perform the work using the smart glasses, following the visual guide. When the user inputs a question using the smart glasses' chatbot function, the server uses an NLP model to generate an appropriate answer and sends it to the device. The server records the user's work history and setting changes and uses a machine learning algorithm to learn the user's preferences and usage habits. Based on this learning result, it generates individual work procedures and setting change suggestions and sends them to the device. The emotion engine analyzes the user's emotional state from their facial expressions and voice and adjusts the operating procedures. The server adjusts assistance advice and provides more appropriate explanations based on the emotional data obtained by the emotion engine.

[1254] Specific Examples

[1255] The following is a specific scenario.

[1256] Support for picking work in the warehouse

[1257] 1. The server retrieves item picking information from the database.

[1258] 2. The server converts the acquired picking information into text data using OCR technology.

[1259] 3. The server summarizes the converted text data using an NLP model.

[1260] 4. The server uses the generative AI model to generate the optimal picking procedure.

[1261] 5. The terminal overlays the picking list information to the user through the smart glasses.

[1262] 6. The user wears the smart glasses and follows visual guidance to pick the item.

[1263] 7. When a user uses the chatbot function to ask a question, the server uses an NLP model to generate an appropriate answer and sends it to the device.

[1264] 8. The emotion engine assesses the user's emotional state from their facial expressions and voice, and adjusts the assistance provided depending on the problem that has occurred.

[1265] Example prompts for generative AI models

[1266] "Picking List Information: Box 123, Section A1, Item 345

[1267] Please suggest the best warehouse operation procedure based on the information below:

[1268] In this way, the system not only provides users with intuitive and efficient support for picking tasks, improving work efficiency within the warehouse, but also provides appropriate support according to the user's emotional state.

[1269] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1270] Step 1:

[1271] The server retrieves item picking information from the database. It uses the picking list information stored in the database as input to obtain the necessary picking data. It obtains raw data for OCR processing as output. Specific operations include issuing a database query to obtain the list information.

[1272] Step 2:

[1273] The server converts the acquired picking information into text data using OCR technology. It uses the raw data obtained in step 1 as input and uses OCR software to obtain the output as digital text. Specifically, it calls an OCR module (e.g., Tesseract) to extract text information from images and documents.

[1274] Step 3:

[1275] The server summarizes the converted text data using a natural language processing (NLP) model. The text data is provided as input to the NLP model, which performs the summarization process and outputs a concise operating procedure. Specifically, an NLP library (e.g., SpaCy or OpenAI GPT-3) is used to generate the summary text.

[1276] Step 4:

[1277] The server uses a generative AI model to propose the optimal picking procedure. The generative AI model receives the summarized text data and the prompt as input. Based on the prompt, the server generates the optimal picking procedure and obtains the procedure as output. Specifically, the server invokes the generative AI (e.g., OpenAI GPT-3), inputs the prompt, and generates the recommended procedure.

[1278] Step 5:

[1279] The terminal overlays the picking list information to the user through the smart glasses. It uses the picking procedure sent from the server as input. It uses AR technology to overlay a visual guide on the smart glasses display and provides visual assistance as output. Specific operations include using AR software (e.g., Unity or ARKit) to display the information on the glasses display.

[1280] Step 6:

[1281] The user wears smart glasses and follows visual guidance to pick items. The input is the information displayed on the smart glasses. The output is to accurately pick the specified items. Specific actions involve looking at the display on the glasses, picking out the items as instructed, and placing them on a cart or other device for inspection.

[1282] Step 7:

[1283] When a user uses the chatbot function to ask a question, the server uses an NLP model to generate an appropriate answer and sends it to the device. The user's question text is used as input. The NLP model analyzes it and outputs the answer. Specifically, the natural language processing model analyzes the user's question, generates an appropriate answer text, and displays it on the chat screen.

[1284] Step 8:

[1285] The emotion engine evaluates the user's emotional state from their facial expressions and voice, and adjusts the assistance provided depending on the problem that has occurred. The input is the user's video and audio data. The emotion engine analyzes the data and outputs the user's emotional state. Specifically, it uses facial recognition technology (e.g., DeepFace or face_recognition) and voice analysis technology to analyze the user's emotions and adjust appropriate advice and procedures.

[1286] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1287] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1288] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1289] [Fourth embodiment]

[1290] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1291] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1292] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1293] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1294] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1295] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1296] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1297] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1298] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1299] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1300] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1301] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1302] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1303] This invention is a home appliance instruction manual system that utilizes generative AI and augmented reality (AR) technology, allowing users to intuitively and quickly obtain operation methods and troubleshooting information for home appliances. The invention can be implemented as follows through the roles of a server, a terminal, and a user.

[1304] Overall system overview

[1305] This system acquires digital versions of instruction manuals for home appliances and extracts text data using optical character recognition (OCR) technology. The extracted text data is summarized using a natural language processing (NLP) model and provided as information to assist users in their operations. It also utilizes AR technology to enable users to visually understand product operation procedures using their smartphones. Furthermore, a support chatbot responds to user inquiries, providing a means for users to instantly obtain the information they need.

[1306] Program processing overview

[1307] 1. The server retrieves the instruction manual data for the home appliance from the database. The retrieved data is converted into text data using OCR technology and summarized using an NLP model. This converts the long instructions into a short, easy-to-understand format.

[1308] 2. The device launches the application and prompts the user to select the model of the home appliance through the user interface. Based on the selected model, the device communicates with the server to obtain relevant operating instructions and troubleshooting information.

[1309] 3. The device allows users to scan home appliances using their smartphone camera, and uses AR technology to overlay operating instructions on the camera image, providing users with a visual guide.

[1310] 4. The user enters a question using the chatbot function within the application. The chatbot uses an NLP model to analyze the question and generate an appropriate answer.

[1311] 5. The server records the user's operation history and setting changes, and uses a machine learning algorithm to learn the user's preferences and usage habits. Based on this learning, it generates individualized operation advice and setting change suggestions and sends them to the device.

[1312] Specific Examples

[1313] The following is a specific scenario.

[1314] Changing the air conditioner remote control settings

[1315] 1. The user opens the application and selects the air conditioner model.

[1316] 2. The device requests the selected model information from the server, and the server returns the operation instructions and setting data for the corresponding model.

[1317] 3. The device instructs the user to scan the air conditioner remote control with the camera, and overlays operating instructions on the scanned image of the remote control, allowing the user to visually confirm the specific button operations.

[1318] 4. When a user asks the chatbot, "I don't know how to set it up," the server analyzes the question, generates detailed instructions, and sends them to the device.

[1319] 5. The device displays the generated instructions on the chatbot screen, and the user follows them to complete the setup.

[1320] In this way, the system intuitively and efficiently provides users with operation and troubleshooting information for home appliances, making it easier to understand instruction manuals and providing optimal operation advice tailored to the user's individual preferences and usage habits.

[1321] The processing flow will be explained below.

[1322] Step 1:

[1323] The server queries the database to retrieve the instruction manual (PDF format) for the home appliance, and stores the retrieved PDF file in local storage or temporarily in memory.

[1324] Step 2:

[1325] The server uses OCR (Optical Character Recognition) technology to extract text data from the PDF. Using an OCR library such as Tesseract, it analyzes all pages in the PDF and converts them into text data.

[1326] Step 3:

[1327] The server then runs the extracted text data through a natural language processing (NLP) model to summarize it, using models like BERT and GPT to reduce redundant descriptions and extract key information.

[1328] Step 4:

[1329] The device launches the application and displays the user interface, initially displaying a search bar and drop-down lists for selecting appliance categories and models.

[1330] Step 5:

[1331] The user selects the model of the home appliance they are using on the interface, and the selected model information is sent to the server, requesting related operating instructions and troubleshooting information.

[1332] Step 6:

[1333] Based on the received model information, the server retrieves relevant operating instructions and troubleshooting information and sends it to the device. The information is retrieved from a database and appropriately formatted before being sent.

[1334] Step 7:

[1335] The device responds to user requests and displays the received instructional and troubleshooting information, appropriately laid out in the app's interface.

[1336] Step 8:

[1337] The device will activate the camera, allowing the user to scan home appliances, capture camera footage in real time, and provide guidance to the user.

[1338] Step 9:

[1339] The device uses augmented reality (AR) technology to overlay operation instructions on the camera image, using libraries such as ARKit and ARCore to overlay instruction icons and text on the operation panel.

[1340] Step 10:

[1341] Users operate home appliances by following the instructions on the camera footage, and follow the on-screen instructions and guides to complete the required operations.

[1342] Step 11:

[1343] Users enter questions into the in-app chatbot, using natural language, and submit the question.

[1344] Step 12:

[1345] The server analyzes the questions received from users through a natural language processing model, interprets the input text to understand its meaning, and generates an appropriate answer.

[1346] Step 13:

[1347] The server sends the generated answer to the terminal, where it is formatted appropriately and displayed on the chatbot screen.

[1348] Step 14:

[1349] The device displays the received response on the chatbot's conversation screen, and the user can confirm the displayed information and perform any necessary operations or settings.

[1350] Step 15:

[1351] The device records the user's operation history and setting changes, and sends the operation details and timestamps to the server as log data.

[1352] Step 16:

[1353] The server compiles the received operation history and setting change data and uses machine learning algorithms to learn the user's preferences and usage habits.

[1354] Step 17:

[1355] The server generates individualized operation advice and setting change suggestions based on the learning results, and sends the suggestions to the device as a notification.

[1356] Step 18:

[1357] The device will notify the user of the received suggestions, either as a push notification or a pop-up message, suggesting new settings or operation methods to the user.

[1358] Example 1

[1359] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1360] Conventional paper manuals for home appliance operation and troubleshooting information are often difficult for users to understand and take a long time to understand. Furthermore, manuals are often lost, making it difficult to quickly obtain the necessary information. Furthermore, when operation methods or solutions are complex, users can easily become lost. There is a need for a system that can solve these problems and enable users to quickly and intuitively understand and operate home appliances.

[1361] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1362] In this invention, the server includes means for acquiring instruction manual data for the home appliance and extracting text data by optical character recognition, means for summarizing the extracted text data through a natural language processing model, means for providing the contents of the instruction manual in a format that is easy for the user to understand using a generative AI model, and means for prompting the user to ask specific questions using prompt sentences and generating detailed answers in response to the questions, thereby enabling the user to quickly and intuitively obtain complex operating instructions and troubleshooting information.

[1363] "Home appliances" refers to all household electronic devices that users use on a daily basis.

[1364] An "instruction manual" refers to a document that describes how to operate, use, maintain, and troubleshoot a home appliance.

[1365] "Digitalization" refers to the process of converting data in paper or physical form into electronic form.

[1366] "User interface" refers to the screen and input devices that users use to operate the device.

[1367] "Optical character recognition (OCR)" refers to the technology that identifies character information from image data and converts it into digital text.

[1368] "Text data" refers to data that is stored as character information and can be processed electronically.

[1369] A "natural language processing (NLP) model" refers to an artificial intelligence model for understanding, analyzing, and generating human language.

[1370] An "abstract" is a concise summary of the contents of a longer document.

[1371] "Augmented reality (AR) technology" refers to the technology of overlaying digital information onto images of the real world.

[1372] A "support chatbot" is a program that automatically responds to questions from users.

[1373] A "machine learning algorithm" refers to a method that allows a computer to learn on its own based on data and make predictions and classifications.

[1374] A "generative AI model" refers to an artificial intelligence model that generates new sentences and answers based on human language.

[1375] A "prompt sentence" refers to a question-type sentence that guides the user's input.

[1376] This invention is a home appliance instruction manual system that utilizes generative AI models and augmented reality (AR) technology, allowing users to intuitively and quickly obtain operating instructions and troubleshooting information for home appliances. This invention is realized through the roles of a server, a terminal, and a user.

[1377] Overall system overview

[1378] This system acquires the digital version of the appliance's instruction manual and extracts text data using optical character recognition (OCR) technology. The extracted text data is summarized using a natural language processing (NLP) model. Furthermore, a generative AI model is used to present the contents of the instruction manual in a format that is easy for users to understand. The user selects the appliance model through the user interface, and related operating instructions and troubleshooting information are provided. When the user scans the appliance using a smartphone, AR technology is used to overlay operating instructions on the camera image. Furthermore, a support chatbot answers questions from the user, quickly providing the necessary information. Through this process, the user's operation history and setting changes are recorded, and a machine learning algorithm learns the user's preferences and usage habits. Based on the learning results, personalized operating advice and setting change suggestions are provided.

[1379] Hardware and software used

[1380] Server: A common RDBMS (e.g., MySQL, PostgreSQL) is used as the database, Tesseract OCR is used as the OCR technology, and BERT or GPT-4 is used as the NLP model.

[1381] Device: A mobile device such as a smartphone or tablet that uses ARCore (for Android) or ARKit (for iOS) as AR technology.

[1382] Users: Use the user interface and chatbot functions within the application.

[1383] Specific examples

[1384] Changing the air conditioner remote control settings

[1385] 1. The user opens the application and selects the model of the air conditioner.

[1386] 2. The device requests the selected model information from the server, and the server returns the operation instructions and setting data for the corresponding model.

[1387] 3. The device instructs the user to scan the air conditioner remote control with the camera, and overlays operating instructions on the scanned image of the remote control, allowing the user to visually confirm the specific button operations.

[1388] 4. When a user asks the chatbot, "I don't know how to set it up," the server analyzes the question, generates detailed instructions, and sends them to the device.

[1389] 5. The device displays the generated instructions on the chatbot screen, and the user follows them to complete the setup.

[1390] Prompt Sentence Examples

[1391] Example question: "How do I switch my air conditioner to cooling mode?"

[1392] Example prompt: "My question is about how to operate my air conditioner. What are the specific steps to switch it to cooling mode?"

[1393] Based on these prompts, the generative AI model generates easy-to-understand answers for users, allowing them to intuitively and efficiently obtain instructions and troubleshooting information for their home appliances.

[1394] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1395] Step 1:

[1396] The server retrieves instruction manual data for the home appliance from the database.

[1397] Input: Product Model ID

[1398] Data processing: Search instruction manual data using database queries

[1399] Output: Instruction manual data

[1400] Specific operation: The server executes an SQL query based on the product model ID to retrieve the instruction manual data. For example, it uses a query like "SELECT FROM manuals WHERE product_id = 'AC1234';".

[1401] Step 2:

[1402] The instruction manual data acquired by the server is converted into text data using OCR technology.

[1403] Input: Instruction manual data

[1404] Data processing: Applying OCR technology to extract text data from image data

[1405] Output: Extracted text data

[1406] Specific operation: Convert image-format instruction manual data into text using Tesseract OCR. For example, read the image using Python's PIL library and extract the text using pytesseract.

[1407] Step 3:

[1408] The server then applies the extracted text data to an NLP model to summarize it.

[1409] Input: Text data

[1410] Data processing: Summarizing text using NLP models

[1411] Output: Summarized text data

[1412] How it works: Summarize long text data concisely using NLP models such as BERT and GPT-4, for example using the summarization pipeline in the transformers library.

[1413] Step 4:

[1414] The terminal allows the user to select the model of the home appliance through a user interface.

[1415] Input: User selected model information

[1416] Data processing: Generate API requests and send them to the server

[1417] Output: Relevant information received from the server

[1418] Specific operation: The user selects a product model on the initial screen of the application, and an API request is sent to the server based on that information. For example, an HTTP request such as "GET / manuals / AC1234".

[1419] Step 5:

[1420] The server transmits the acquired related information to the terminal.

[1421] Input: User selected model information

[1422] Data processing: Generate related information after search processing

[1423] Output: Related operating instructions and troubleshooting information

[1424] Specific operation: The server receives the request, retrieves the relevant operation instructions and troubleshooting information from the database, and sends them to the terminal.

[1425] Step 6:

[1426] The device uses its camera to scan home appliances and uses AR technology to overlay operating instructions.

[1427] Input: Acquired related information, camera footage

[1428] Data processing: AR technology overlays operation procedures onto camera images

[1429] Output: AR displayed operation procedure

[1430] Specific operation: When a user points the camera at a home appliance, ARCore or ARKit is used to overlay operating instructions on the real-world image.

[1431] Step 7:

[1432] A user enters a question using the chatbot function within the application.

[1433] Input: User's question text

[1434] Data processing: Parse the question with an NLP model

[1435] Output: Parsed question

[1436] Specific operation: A question is entered into the chatbot, and the question is analyzed using an NLP model.

[1437] Step 8:

[1438] The server generates an appropriate answer to the analyzed question and sends it to the terminal.

[1439] Input: Parsed question content

[1440] Data processing: Generative AI models generate answers

[1441] Output: The generated answer

[1442] Specific operation: The server inputs the received question into a generative AI model to generate an appropriate answer. For example, it uses GPT-4 to generate an answer and sends it to the device.

[1443] Step 9:

[1444] The device displays the generated answer on the chatbot screen.

[1445] Input: Generated Answer

[1446] Data processing: None

[1447] Output: The answer displayed on the chatbot screen

[1448] Specific operation: The generated answer text is displayed on the chatbot screen for the user to see.

[1449] Step 10:

[1450] The server records user operation history and setting changes and learns from them using machine learning algorithms.

[1451] Input: User operation history and setting change information

[1452] Data processing: Learning processing using machine learning algorithms

[1453] Output: Learning results

[1454] Specific operation: Collect operation history and setting change data, and cluster behavioral patterns using, for example, Scikit-learn.

[1455] Step 11:

[1456] The server generates individual operation advice and suggestions for setting changes based on the learning results and sends them to the device.

[1457] Input: Learning results

[1458] Data processing: generating personalized advice and suggestions for changing settings

[1459] Output: Generated advice and suggestions

[1460] Specific operation: Generates personalized advice based on user behavior data and sends it to the device.

[1461] (Application example 1)

[1462] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1463] Instruction manuals for home appliances are often provided in paper format, and understanding them takes time and effort. Furthermore, when users encounter difficulties setting up or operating the device, it is difficult to intuitively understand the specific operation method. Furthermore, with the recent spread of smart security systems, there is a demand for intuitive guidance for their installation and configuration. Conventional technologies lack systems that provide comprehensive, immediate, and appropriate support, so there is a need to solve these issues.

[1464] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1465] In this invention, the server includes means for acquiring instruction manual data and extracting text data using optical character recognition, means for summarizing the extracted text data using a natural language processing model, means for selecting a home appliance model through a user interface and providing related operating instructions and troubleshooting information, means for scanning the home appliance using a camera and visualizing operating procedures using augmented reality technology, means for answering questions from the user using a support chatbot using the natural language processing model, means for recording user operation history and setting changes and learning user preferences and usage habits using a machine learning algorithm, means for providing individual operating advice and suggesting setting changes based on the learning results, and means for visualizing installation and configuration methods for the smart security system using augmented reality technology, thereby enabling intuitive and efficient operation and configuration of home appliances and smart security systems.

[1466] An "instruction manual" is a document that provides users with information on how to operate, set up, and maintain home appliances and other devices.

[1467] "Optical character recognition" is a technology that extracts text data from images and documents.

[1468] A "natural language processing model" is an algorithm or system for understanding and generating human language.

[1469] "User interface" refers to the means or screen through which a user interacts with a system.

[1470] A "camera" is a device that receives light and captures image information.

[1471] "Augmented reality technology" is a technology that displays digital information superimposed on images of the real world.

[1472] A "support chatbot" is software that automatically answers users' questions via text or voice.

[1473] A "machine learning algorithm" is a computational method for learning patterns from data and making predictions and classifications.

[1474] "Preferences" refer to the preferences and tendencies of an individual.

[1475] A "smart security system" is a system that automatically monitors and manages safety using network-connected cameras and sensors.

[1476] An "operating procedure" refers to the steps or methods for using a piece of equipment or system.

[1477] "Troubleshooting Information" means information containing guidance or advice for resolving equipment malfunctions or errors.

[1478] "Operation history" is a record of when a user uses a system or device.

[1479] "Configuration change" refers to the change made to adjust the operation of a system or device.

[1480] This invention is an instruction manual system for a smart security system that utilizes generative AI and augmented reality (AR) technology, allowing users to intuitively and quickly obtain operating instructions and troubleshooting information for security devices. The invention can be implemented as follows through the roles of a server, a terminal, and a user.

[1481] Overall system overview

[1482] This system digitally acquires the instruction manual for the smart security system and extracts the text data using optical character recognition (OCR) technology. The extracted text data is summarized using a natural language processing (NLP) model and provided as information to assist the user in operation. It also utilizes AR technology to allow users to visually understand the product's operation procedures and installation methods using their smartphone. Furthermore, a support chatbot responds to user inquiries, providing a means for users to instantly obtain the information they need.

[1483] Program processing overview

[1484] 1. The server retrieves the instruction manual data for the smart security system from the database. The retrieved data is converted into text data using OCR technology (e.g., Google Cloud Vision API) and then summarized using an NLP model (e.g., OpenAI GPT-3). This converts the long instruction manual into a short, easy-to-understand format.

[1485] 2. The device (smartphone, smart glasses, head-mounted display) launches the application and prompts the user to select the security device model through the user interface. Based on the selected model, it communicates with the server to obtain relevant operating instructions and troubleshooting information.

[1486] 3. The device allows users to use the camera to scan for security devices, and uses AR technology (e.g., Apple ARKit, Google ARCore) to overlay operation instructions and installation locations on the camera image, providing users with a visual guide.

[1487] 4. The user enters a question using the chatbot function within the application. The chatbot uses an NLP model to analyze the question and generate an appropriate answer.

[1488] 5. The server records the user's operation history and setting changes, and uses machine learning algorithms to learn the user's preferences and usage habits. Based on this learning, it generates personalized operation advice and setting change suggestions and sends them to the device.

[1489] Specific Examples

[1490] Security camera Wi-Fi settings

[1491] 1. The user opens the application and selects the security camera model.

[1492] 2. The device requests the selected model information from the server, and the server returns the operation instructions and setting data for the corresponding model.

[1493] 3. The device prompts the user to scan the security camera and overlays Wi-Fi setup instructions on the scanned camera image, allowing the user to visually confirm the specific steps.

[1494] 4. When a user asks the chatbot, "I don't know how to set up Wi-Fi," the server analyzes the question, generates detailed instructions, and sends them to the device.

[1495] 5. The device displays the generated instructions on the chatbot screen, and the user follows them to complete the setup.

[1496] Prompt Sentence Examples

[1497] "Please tell me about the Wi-Fi settings for the security camera."

[1498] "I'd like to set up the initial settings for my security system. Please tell me the procedure."

[1499] "Please show me in AR where to place the motion sensor."

[1500] In this way, the system provides users with intuitive and efficient operation and troubleshooting information for their smart security system, making it easier to understand the instruction manual and providing optimal operation advice tailored to the user's individual preferences and usage habits.

[1501] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1502] Step 1:

[1503] The server retrieves the instruction manual data for the smart security system from the database. The input is a request for instruction manual data, and the output is the retrieved instruction manual data. Using OCR technology (e.g., Google Cloud Vision API), the image data is converted into text data. Specifically, the server accesses the database upon receiving the request, retrieves the image data of the specified instruction manual, and applies OCR technology to extract the text data.

[1504] Step 2:

[1505] The server summarizes the extracted text data by running it through a natural language processing (NLP) model (e.g., OpenAI GPT-3). The input is the OCR-processed text data, and the output is the summarized text data. Specifically, the server inputs the text data into the NLP model and receives the generated summary data.

[1506] Step 3:

[1507] The terminal (smartphone, smart glasses, head-mounted display) allows the user to select a security device model through a user interface. The input is the user's model selection operation, and the output is the selected model information. Specifically, the terminal presents multiple models through the display interface and collects the information selected by the user.

[1508] Step 4:

[1509] The terminal requests the selected model information from the server, and the server provides the related operation method and troubleshooting information. The input is the model information request, and the output is the operation method and troubleshooting information. In concrete terms, the terminal sends a request to the server, and the server searches for the corresponding model information and returns it to the terminal.

[1510] Step 5:

[1511] The device allows users to scan security devices using a camera. The input is the camera image, and the output is an overlay display of location information and operation procedures. The display is achieved using augmented reality technology (e.g., Apple ARKit, Google ARCore). Specifically, the device acquires real-time images from the camera and uses AR technology to overlay the necessary guide information on the image.

[1512] Step 6:

[1513] A user inputs a question using the chatbot function within the application. The chatbot uses an NLP model to analyze the input question and generate an appropriate answer. The input is the user's question, and the output is the generated answer. Specifically, the user inputs a question as text, the chatbot sends the question to the NLP model for analysis, and then displays the generated answer to the user.

[1514] Step 7:

[1515] The server records the user's operation history and setting changes, and uses a machine learning algorithm to learn the user's preferences and usage habits. The input is the operation history and setting change data, and the output is individual operation advice and setting change suggestions based on the learning results. Specifically, the server records the operation history in a database, analyzes the data using a machine learning algorithm, and generates individually optimized advice and setting change suggestions.

[1516] Step 8:

[1517] The device provides the user with personalized operation advice and suggestions for setting changes based on the learning results. The input is the suggestion data sent from the server, and the output is a presentation to the user. Specifically, the device receives the suggestion data and displays it on the user interface.

[1518] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1519] The present invention is a home appliance instruction manual system that utilizes generative AI, augmented reality (AR) technology, and an emotion engine, allowing users to intuitively and quickly obtain operation instructions and troubleshooting information for home appliances. The present invention can be implemented as follows through the roles of a server, a terminal, and a user.

[1520] Overall system overview

[1521] This system acquires digital versions of instruction manuals for home appliances, extracts text data using optical character recognition (OCR), summarizes it using natural language processing (NLP) models, and provides it as information to assist users in their operations. Augmented reality (AR) technology also allows users to visually understand product operation procedures using their smartphones. A support chatbot responds to user inquiries, and an emotion engine recognizes the user's emotions and provides optimal advice.

[1522] Program processing overview

[1523] 1. The server retrieves the instruction manual data for the home appliance from the database. The retrieved data is converted into text data using OCR technology and summarized using an NLP model, converting the long instructions into a short, easy-to-understand format.

[1524] 2. The device launches the application and prompts the user to select the model of the home appliance through the user interface. Based on the selected model, the device communicates with the server to obtain relevant operating instructions and troubleshooting information.

[1525] 3. The device allows users to scan home appliances using their smartphone camera, and uses AR technology to overlay operating instructions on the camera image, providing users with a visual guide.

[1526] 4. The user enters a question using the chatbot function within the application. The chatbot uses an NLP model to analyze the question and generate an appropriate answer.

[1527] 5. The server records the user's operation history and setting changes, and uses a machine learning algorithm to learn the user's preferences and usage habits. Based on this learning, it generates individualized operation advice and setting change suggestions and sends them to the device.

[1528] 6. Recognize user emotions using an emotion engine. Analyze facial expressions and voice recorded through the camera while the user is operating the device, and evaluate the user's emotional state in real time.

[1529] 7. The server adjusts its operational advice and suggestions for setting changes based on the emotional data obtained by the emotion engine. For example, if the user is confused, it will provide more detailed and gentler explanations.

[1530] Specific Examples

[1531] The following is a specific scenario.

[1532] Changing the air conditioner remote control settings

[1533] 1. The user opens the application and selects the air conditioner model.

[1534] 2. The device requests the selected model information from the server, and the server returns the operation instructions and setting data for the corresponding model.

[1535] 3. The device instructs the user to scan the air conditioner remote control with the camera, and overlays operating instructions on the scanned image of the remote control, allowing the user to visually confirm the specific button operations.

[1536] 4. When a user asks the chatbot, "I don't know how to set it up," the server analyzes the question, generates detailed instructions, and sends them to the device.

[1537] 5. The device displays the generated instructions on the chatbot screen, and the user follows them to complete the setup.

[1538] 6. If the emotion engine detects confusion or stress in the user's facial expression, the server will adjust the advice accordingly and send a more understandable explanation to the device.

[1539] In this way, the system not only provides users with intuitive and efficient instructions on how to operate home appliances and troubleshooting information, making it easier to understand instruction manuals, but also provides appropriate support according to the user's emotional state.

[1540] The processing flow will be explained below.

[1541] Step 1:

[1542] The server queries the database to retrieve the instruction manual (PDF format) for the home appliance, and stores the retrieved PDF file in local storage or temporarily in memory.

[1543] Step 2:

[1544] The server extracts text data from the PDF using OCR (Optical Character Recognition) technology. Using an OCR library such as Tesseract, it analyzes all pages in the PDF and converts them into text data.

[1545] Step 3:

[1546] The server then runs the extracted text data through a natural language processing (NLP) model to summarize it, using models like BERT and GPT to reduce redundant descriptions and extract key information.

[1547] Step 4:

[1548] The device launches the application and displays the user interface, initially displaying a search bar and drop-down lists for selecting appliance categories and models.

[1549] Step 5:

[1550] The user selects the model of the home appliance they are using on the interface, and the selected model information is sent to the server, requesting related operating instructions and troubleshooting information.

[1551] Step 6:

[1552] Based on the received model information, the server retrieves relevant operating instructions and troubleshooting information and sends it to the device. The information is retrieved from a database and appropriately formatted before being sent.

[1553] Step 7:

[1554] The device responds to user requests and displays the received instructional and troubleshooting information, appropriately laid out in the app's interface.

[1555] Step 8:

[1556] The device will activate the camera, allowing the user to scan home appliances, capture camera footage in real time, and provide guidance to the user.

[1557] Step 9:

[1558] The device uses augmented reality (AR) technology to overlay operation instructions on the camera image, using libraries such as ARKit and ARCore to overlay instruction icons and text on the operation panel.

[1559] Step 10:

[1560] Users operate home appliances by following the instructions on the camera footage, and follow the on-screen instructions and guides to complete the required operations.

[1561] Step 11:

[1562] Users enter questions into the in-app chatbot, using natural language, and submit the question.

[1563] Step 12:

[1564] The server analyzes the questions received from users through a natural language processing model, interprets the input text to understand its meaning, and generates an appropriate answer.

[1565] Step 13:

[1566] The server sends the generated answer to the terminal, where it is formatted appropriately and displayed on the chatbot screen.

[1567] Step 14:

[1568] The device displays the received response on the chatbot's conversation screen, and the user can confirm the displayed information and perform any necessary operations or settings.

[1569] Step 15:

[1570] The device records the user's operation history and setting changes, and sends the operation details and timestamps to the server as log data.

[1571] Step 16:

[1572] The server compiles the received operation history and setting change data and uses machine learning algorithms to learn the user's preferences and usage habits.

[1573] Step 17:

[1574] The server generates individualized operation advice and setting change suggestions based on the learning results, and sends the suggestions to the device as a notification.

[1575] Step 18:

[1576] The device will notify the user of the received suggestions, either as a push notification or a pop-up message, suggesting new settings or operation methods to the user.

[1577] Step 19:

[1578] The device captures facial expression data through the camera while the user is operating the device, and transmits it to the emotion engine, which analyzes facial expressions and voice to evaluate the user's emotional state in real time.

[1579] Step 20:

[1580] The emotion data recognized by the emotion engine is sent to the server, which then adjusts the advice and suggested settings changes accordingly.

[1581] Step 21:

[1582] If the user's emotional state is confused or stressed, the server generates a more detailed and gentle explanation and sends it to the device, improving the quality of support according to the user's emotional state.

[1583] Example 2

[1584] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1585] Modern home appliances are becoming increasingly multifunctional, making it difficult to accurately understand their operation methods and troubleshooting information. Traditional instruction manuals are provided in paper or digital format, but the volume of information makes it difficult for users to quickly find the information they need. Furthermore, the lack of flexible responses or personalized advice based on the user's emotional state leaves the user with a poor user experience. Furthermore, there is a need for systems that can learn users' preferences and usage habits and suggest optimal operations and settings.

[1586] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1587] In this invention, the server includes means for acquiring instruction manual data and extracting character data using optical character recognition, means for summarizing the extracted character data through a natural language processing model, means for recording a user's operation history and setting changes and learning the user's preferences and usage habits using a machine learning algorithm, means for recognizing the user's emotional state using an emotion analysis engine, and means for adjusting operation advice and setting change suggestions based on the results of the emotion analysis engine. This allows the user to quickly obtain the necessary information and operate the home appliance intuitively and efficiently. In addition, to improve the user experience, the user can receive optimal advice and setting suggestions based on the user's emotional state.

[1588] "Home appliances" refer to electronic devices and electrically powered devices that are generally used in the home, including, for example, air conditioners, refrigerators, washing machines, and the like.

[1589] An "instruction manual" is a document that describes how to use, install, and troubleshoot a home appliance device.

[1590] "Digitization" is the process of converting paper or analog information into digital form, making it possible to view, store, and process the information on computers and other digital devices.

[1591] Optical character recognition (OCR) is a technology that extracts character data from images and scanned documents, allowing handwritten or printed characters to be recognized as digital text.

[1592] A "natural language processing (NLP) model" is an artificial intelligence technique for understanding, analyzing, and processing natural language, and can perform tasks such as text summarization, translation, and sentiment analysis.

[1593] A "user interface" is a medium through which a user interacts with a system, and includes a graphical user interface (GUI) and a voice user interface (VUI).

[1594] An "image acquisition device" is a device for acquiring image data, such as a camera or scanner, and particularly includes cameras installed in smartphones.

[1595] Augmented reality (AR) technology is a technology that displays digital information overlaid on images of the real world, allowing users to visually recognize the real world and virtual information in an integrated manner.

[1596] A "support chatbot" is a program that automatically responds to questions and requests from users, and uses natural language processing technology to provide appropriate answers.

[1597] A "machine learning algorithm" is a computational method for analyzing data, learning patterns, and making predictions. It is particularly used to analyze user behavior patterns and suggest optimal operations and settings.

[1598] An "emotion analysis engine" is an artificial intelligence technology that analyzes a user's facial expressions and voice data to evaluate their emotional state, and can provide feedback according to the user's psychological state.

[1599] "Operation history" is a record of operations and setting changes performed by a user using the system, and this information is used to understand the user's preferences and behavioral patterns.

[1600] "Adjusting suggestions" means dynamically changing the appropriate advice and recommendations for setting changes to the user based on collected data and analysis results.

[1601] The present invention is a user manual system for home appliances that utilizes generative AI models, augmented reality (AR) technology, and an emotion engine, allowing users to intuitively and quickly obtain operating instructions and troubleshooting information for home appliances. The present invention is implemented through the roles of a server, a terminal, and a user. This improves user convenience and satisfaction.

[1602] Hardware and Software Use

[1603] Server: Database, Optical Character Recognition (OCR) technology, Natural Language Processing (NLP) models, Sentiment Analysis Engine, Machine Learning Algorithms

[1604] Devices: Smartphones, cameras, and applications using augmented reality (AR) technology

[1605] User: Operating the application, using the camera, using the chat function

[1606] Data acquisition and preprocessing

[1607] The server retrieves instruction manual data for a home appliance from the database. For example, it retrieves instruction manual data for an air conditioner. The retrieved data is often in PDF or image file format.

[1608] Text extraction and summarization

[1609] The server converts the acquired instruction manual data into text data using optical character recognition (OCR) technology. For example, it uses Tesseract OCR to extract text from PDFs and images. The extracted text data is then input into a natural language processing (NLP) model (e.g., BERT, GPT-3) to generate a summary. For example, it can extract only important setup steps and necessary information from long instructions and create a short summary.

[1610] Product Scanning and Display

[1611] The terminal provides an application that allows the user to select the model of a home appliance. When the user selects "air conditioner," the terminal requests the model information from the server, and the server sends related operating instructions and troubleshooting information to the terminal.

[1612] Inquiry response

[1613] The user uses the app's chatbot function to input a question, such as "How do I set a timer?" The server analyzes the question using a natural language processing (NLP) model, generates an appropriate answer, and sends it to the device.

[1614] Operation history recording and learning

[1615] The server records the user's operation history and setting changes, and uses a machine learning algorithm to learn the user's preferences and usage habits. For example, it records a history such as "Timer settings were changed on October 1, 2023," and based on this, suggests optimal settings for users who often use the device at night.

[1616] Recognition of emotional states

[1617] The emotion analysis engine analyzes the user's facial expressions and voice via a camera and microphone to assess their emotional state. For example, it may recognize that the user is confused. The server records this emotional data and generates appropriate feedback.

[1618] Providing advice and coordination

[1619] The server tailors operational advice and setting change suggestions based on the results of the emotion analysis engine. For example, if the user is confused, it generates detailed and easy-to-understand instructions and sends them to the device.

[1620] Examples of concrete examples and prompts

[1621] For example, when a user wants to change the remote control settings of an air conditioner, the specific operating procedure is as follows.

[1622] 1. The user opens the application and selects the air conditioner model.

[1623] 2. The device requests the selected model information from the server, and the server returns the operation instructions and setting data for the corresponding model.

[1624] 3. The device instructs the user to scan the air conditioner remote control with the camera and overlays operating instructions on the scanned image.

[1625] 4. When a user asks the chatbot, "I don't know how to set it up," the server analyzes the question, generates detailed instructions, and sends them to the device.

[1626] 5. The device displays the generated instructions on the chatbot screen, and the user follows them to complete the setup.

[1627] 6. The emotion engine detects confusion or stress from the user's facial expressions, and the server adjusts the advice based on this and sends a more understandable explanation to the device.

[1628] As an example of a prompt sentence, you could feed the generative AI model the following:

[1629] "How do I set the timer on the air conditioner remote control?"

[1630] "What should I do if a user is confused?"

[1631] In this way, the system not only provides users with intuitive and efficient operating instructions and troubleshooting information for home appliances, making it easier to understand instruction manuals, but also provides appropriate support according to the user's emotional state.

[1632] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1633] Step 1:

[1634] The server retrieves instruction manual data for home appliances from a database. The retrieved data may be in PDF or image format. The input data may be, for example, "air conditioner instruction manual data," which is then subjected to OCR processing on the server side. The server then converts this data into text data using optical character recognition (OCR) technology. The output is the text data for the instruction manual. For example, the text may include "How to operate the air conditioner" or "Troubleshooting information."

[1635] Step 2:

[1636] The server inputs the character data extracted by OCR into a natural language processing (NLP) model to generate a summary. The character data is used as input data and passed to an NLP model (e.g., BERT or GPT-3). Specifically, the text is summarized using an application with a summarization function. As output, a summary text is generated that extracts only the important steps from the detailed description. For example, "Description of all buttons on a remote control" is summarized as "Description of the main buttons."

[1637] Step 3:

[1638] The terminal launches an application to allow the user to select a model of a home appliance. The user operates the terminal's user interface to select a model, such as "air conditioner." The user's model selection information is used as input. Based on this, the terminal sends a request for model information to the system. The model selection information is sent to the server as output.

[1639] Step 4:

[1640] Based on the request from the terminal, the server returns the operation method and troubleshooting information for the relevant model. The input is the model information selected by the user, and based on that, the server retrieves the relevant information from the database. Specifically, this includes detailed data such as operation procedures and how to deal with errors. As output, the server sends the related operation method and troubleshooting information to the terminal.

[1641] Step 5:

[1642] The device instructs the user to scan the home appliance with a camera and visualizes the operation procedure using augmented reality (AR) technology. The input is the image data scanned by the user with the camera. Specifically, AR technology is used to overlay instructions such as "Press the POWER button" on the camera image. The output is the visualized operation procedure provided to the user.

[1643] Step 6:

[1644] Users can use the chatbot function within the app to input specific questions, such as, "I don't know how to set the timer on my air conditioner." This question is then sent to the chatbot.

[1645] Step 7:

[1646] The server analyzes questions received via the chatbot using a natural language processing (NLP) model, generates appropriate answers, and sends them to the device. The input is the user's question text. Based on this, the NLP model operates and generates text with specific setup procedures and operation instructions. As output, the generated answer text is sent to the device and displayed on the chatbot screen.

[1647] Step 8:

[1648] The server records the user's operation history and setting changes, and uses a machine learning algorithm to learn the user's preferences and usage habits. The input is the user's operation history data. Specifically, a history such as "Change timer settings on October 1, 2023" is saved. The output is recommended settings and operation methods based on the learning results.

[1649] Step 9:

[1650] The emotion analysis engine analyzes the user's facial expressions and voice via a camera and microphone to evaluate their emotional state. The inputs include facial expression data captured by the camera and voice data captured by the microphone. Specifically, it analyzes in real time whether the user is confused or not. The output is the emotion analysis result.

[1651] Step 10:

[1652] The server adjusts operational advice and setting change suggestions based on the results of the sentiment analysis engine. The input is the sentiment analysis data. Specifically, it prepares detailed and friendly explanations for confused users. The output is the adjusted advice and setting change suggestions sent to the device.

[1653] The above is the specific processing flow of the program of this system.

[1654] (Application example 2)

[1655] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1656] Conventional instruction manual systems for home appliances often make it difficult for users to intuitively understand how to operate them, especially when it comes to complex operations and troubleshooting. Furthermore, due to a lack of systems utilizing emotion engines and generative AI, users' emotional state and individual operational support are insufficient. This results in reduced user operational efficiency and longer troubleshooting times. Furthermore, there is a demand for more efficient picking operations in warehouses, but intuitive support using AR technology is lacking.

[1657] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1658] In this invention, the server includes means for acquiring instruction manual data and extracting text data by optical character recognition, means for summarizing the extracted text data by applying a natural language processing model, and means for selecting a device model through a user interface and providing related operating instructions and troubleshooting information, thereby enabling intuitive and rapid provision of operating instructions and troubleshooting information to the user.

[1659] The server also includes a means for scanning devices using a camera and visualizing operation procedures using augmented reality technology, a means for answering user questions using a support chatbot using a natural language processing model, a means for recording user operation history and setting changes and learning user preferences and usage habits using a machine learning algorithm, a means for providing individual operation advice and suggesting setting changes based on the learning results, a means for proposing work procedures using a generative AI model, a means for analyzing the user's emotional state using an emotion engine and adjusting the operation procedures, and a means for displaying item picking information and supporting the user using augmented reality technology. This makes it possible to provide optimal support according to the user's emotional state and behavioral patterns, thereby improving operation efficiency and work accuracy.

[1660] "Instruction Data" means digital instruction manual information including operating instructions and troubleshooting information for home appliances and other devices.

[1661] Optical character recognition is a technology that automatically reads letters and numbers from images or handwritten characters and converts them into digital text.

[1662] A "natural language processing model" is a machine learning model for understanding, analyzing, and generating human language, allowing it to summarize text data and answer questions.

[1663] A "user interface" is an interaction mechanism that provides a screen and operating methods for users to interact with a system.

[1664] "Augmented reality technology" is a technology that displays digital information overlaid on the real environment, and is used to overlay operating procedures and instructions on camera images.

[1665] A "support chatbot" is a program that uses a natural language processing model to automatically respond to questions from users.

[1666] A "machine learning algorithm" is an algorithm that learns patterns using large amounts of data and makes predictions and classifications for new data.

[1667] A "generative AI model" is an artificial intelligence model that generates new content based on existing data, and is particularly used to generate text, audio, images, etc.

[1668] An "emotion engine" is a system that analyzes a person's emotional state from facial expressions and voice, and responds adaptively based on the results.

[1669] "Item picking information" is information about the work of picking items from warehouses and logistics centers, and includes a picking list and the like.

[1670] This invention is a logistics center support system that utilizes generative AI models, augmented reality (AR) technology, and an emotion engine to enable users to intuitively and quickly perform item picking tasks. This system can be implemented as follows through the roles of the server, terminal, and user.

[1671] Overall system overview

[1672] The system acquires item picking information in digital format and extracts text data using optical character recognition (OCR) technology. The extracted data is summarized using a natural language processing (NLP) model and provided as information to assist the user in their work. AR technology also allows users to visually understand the picking process using smart glasses. A support chatbot answers user inquiries, and an emotion engine recognizes the user's emotions to provide optimal assistance.

[1673] Program processing overview

[1674] The server retrieves item picking information from a database. The retrieved data is converted into text data using OCR technology and summarized using an NLP model. The server then uses a generative AI model to suggest work procedures. The user obtains relevant picking information through a device equipped with a user interface. The device scans the item using smart glasses and overlays the picking list information using AR technology. The user can intuitively perform the work using the smart glasses, following the visual guide. When the user inputs a question using the smart glasses' chatbot function, the server uses an NLP model to generate an appropriate answer and sends it to the device. The server records the user's work history and setting changes and uses a machine learning algorithm to learn the user's preferences and usage habits. Based on this learning result, it generates individual work procedures and setting change suggestions and sends them to the device. The emotion engine analyzes the user's emotional state from their facial expressions and voice and adjusts the operating procedures. The server adjusts assistance advice and provides more appropriate explanations based on the emotional data obtained by the emotion engine.

[1675] Specific Examples

[1676] The following is a specific scenario.

[1677] Support for picking work in the warehouse

[1678] 1. The server retrieves item picking information from the database.

[1679] 2. The server converts the acquired picking information into text data using OCR technology.

[1680] 3. The server summarizes the converted text data using an NLP model.

[1681] 4. The server uses the generative AI model to generate the optimal picking procedure.

[1682] 5. The terminal overlays the picking list information to the user through the smart glasses.

[1683] 6. The user wears the smart glasses and follows visual guidance to pick the item.

[1684] 7. When a user uses the chatbot function to ask a question, the server uses an NLP model to generate an appropriate answer and sends it to the device.

[1685] 8. The emotion engine assesses the user's emotional state from their facial expressions and voice, and adjusts the assistance provided depending on the problem that has occurred.

[1686] Example prompts for generative AI models

[1687] "Picking List Information: Box 123, Section A1, Item 345

[1688] Please suggest the best warehouse operation procedure based on the information below:

[1689] In this way, the system not only provides users with intuitive and efficient support for picking tasks, improving work efficiency within the warehouse, but also provides appropriate support according to the user's emotional state.

[1690] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1691] Step 1:

[1692] The server retrieves item picking information from the database. It uses the picking list information stored in the database as input to obtain the necessary picking data. It obtains raw data for OCR processing as output. Specific operations include issuing a database query to obtain the list information.

[1693] Step 2:

[1694] The server converts the acquired picking information into text data using OCR technology. It uses the raw data obtained in step 1 as input and uses OCR software to obtain the output as digital text. Specifically, it calls an OCR module (e.g., Tesseract) to extract text information from images and documents.

[1695] Step 3:

[1696] The server summarizes the converted text data using a natural language processing (NLP) model. The text data is provided as input to the NLP model, which performs the summarization process and outputs a concise operating procedure. Specifically, an NLP library (e.g., SpaCy or OpenAI GPT-3) is used to generate the summary text.

[1697] Step 4:

[1698] The server uses a generative AI model to propose the optimal picking procedure. The generative AI model receives the summarized text data and the prompt as input. Based on the prompt, the server generates the optimal picking procedure and obtains the procedure as output. Specifically, the server invokes the generative AI (e.g., OpenAI GPT-3), inputs the prompt, and generates the recommended procedure.

[1699] Step 5:

[1700] The terminal overlays the picking list information to the user through the smart glasses. It uses the picking procedure sent from the server as input. It uses AR technology to overlay a visual guide on the smart glasses display and provides visual assistance as output. Specific operations include using AR software (e.g., Unity or ARKit) to display the information on the glasses display.

[1701] Step 6:

[1702] The user wears smart glasses and follows visual guidance to pick items. The input is the information displayed on the smart glasses. The output is to accurately pick the specified items. Specific actions involve looking at the display on the glasses, picking out the items as instructed, and placing them on a cart or other device for inspection.

[1703] Step 7:

[1704] When a user uses the chatbot function to ask a question, the server uses an NLP model to generate an appropriate answer and sends it to the device. The user's question text is used as input. The NLP model analyzes it and outputs the answer. Specifically, the natural language processing model analyzes the user's question, generates an appropriate answer text, and displays it on the chat screen.

[1705] Step 8:

[1706] The emotion engine evaluates the user's emotional state from their facial expressions and voice, and adjusts the assistance provided depending on the problem that has occurred. The input is the user's video and audio data. The emotion engine analyzes the data and outputs the user's emotional state. Specifically, it uses facial recognition technology (e.g., DeepFace or face_recognition) and voice analysis technology to analyze the user's emotions and adjust appropriate advice and procedures.

[1707] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1708] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1709] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1710] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1711] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1712] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1713] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1714] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiol...

Claims

1. A system for digitizing instruction manuals for home appliances and providing users with intuitive operation methods, a means for acquiring instruction data and extracting text data by optical character recognition; A means for summarizing the extracted text data through a natural language processing model; a means for selecting a model of appliance through a user interface and providing associated operating instructions and troubleshooting information; A means for scanning home appliances using a camera and visualizing operation procedures using augmented reality technology; A means of answering user questions through a support chatbot using a natural language processing model; A means of recording user operation history and setting changes and using machine learning algorithms to learn user preferences and usage habits; A means to provide individual operation advice and suggest setting changes based on the learning results, A system including:

2. The system according to claim 1 , further comprising means for recognizing the home appliance through a camera and overlaying instructions on the operation panel using augmented reality technology.

3. The system according to claim 1 , further comprising means for analyzing a user's behavioral patterns using a machine learning algorithm and proposing optimal settings and operation methods.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A