system

The system addresses complex assembly instructions by generating customizable assembly videos from product information, enhancing user understanding and reducing errors through intuitive visual guidance.

JP2026068419APending Publication Date: 2026-04-22SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-10
Publication Date
2026-04-22

AI Technical Summary

Technical Problem

Existing product assembly instructions are complex and difficult to understand, particularly for inexperienced users and foreign language speakers, often relying on text and images that can lead to incorrect assembly processes.

Method used

A system that analyzes product information and instruction manuals, extracts assembly procedures, and generates customized assembly videos in multiple languages using generative AI, allowing users to follow intuitive visual guides.

Benefits of technology

Improves the success rate of assembly by providing clear, language-independent visual instructions tailored to individual user needs, reducing errors and assembly time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026068419000001_ABST
    Figure 2026068419000001_ABST
Patent Text Reader

Abstract

Instructions for products requiring assembly are often complex and visually difficult to understand, leading to an increase in user failures during assembly. This system aims to resolve this issue. [Solution] The system receives product information and instruction manual data, analyzes them to extract assembly procedures, learns from existing video data, and automatically generates assembly videos based on the analyzed procedures. The system includes means for receiving instruction manual data and product information, means for analyzing the received instruction manuals and extracting assembly procedures, means for learning from existing video data and generating assembly videos based on the analyzed procedures, and means for distributing the generated videos to users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Instructions for products that require assembly are often complex and visually difficult to understand, and the cases where users fail in assembly are increasing. This problem is particularly prominent among inexperienced users and foreign language speakers, and there is a need for a method to lower the hurdle for assembly work. In existing instructions, there is a problem that they often rely on text and still images, making it difficult to intuitively understand and likely to cause incorrect assembly processes.

Means for Solving the Problems

[0005] This invention provides a system that receives product information and instruction manual data, analyzes them, and extracts assembly procedures. This system learns from existing video data and automatically generates assembly videos based on the analyzed procedures. As a result, users are guided through assembly via visually easy-to-understand videos, thereby improving the success rate of assembly. Furthermore, the generated videos are distributed in multiple languages, making assembly easier even in different language environments. In this way, the present invention aims to effectively solve the problems of the past.

[0006] "Instruction manual data" refers to digital information of documents that describe the procedures and precautions for products that require assembly.

[0007] "Product information" refers to basic information such as the name and model number used to identify a specific product.

[0008] "Analysis" is the process of carefully examining the contents of instruction manual data and extracting assembly procedures and important points.

[0009] "Existing video data" refers to a database of video clips and related footage of various assemblies collected in the past.

[0010] "Learning" is the process by which a generative AI model identifies patterns and trends from existing video data and accumulates that knowledge.

[0011] An "assembly video" is a visual guide video that is automatically generated based on the analyzed procedure.

[0012] "Distribution" refers to the act of providing the generated assembled video in a form accessible to users.

[0013] A "system" is a platform with a set of technical components for performing instruction manual analysis, video generation, and video distribution. [Brief explanation of the drawing]

[0014] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Modes for Carrying Out the Invention

[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0016] First, the terms used in the following description will be explained.

[0017] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0018] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0019] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0022] [First Embodiment]

[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0031] As shown in Figure 2, in the data processing device 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0035] The system for carrying out this invention uses a program that operates between the user, terminal, and server. The program processing and specific embodiments of this system are described below.

[0036] The user uses a terminal to input information about the product they wish to assemble and its accompanying instructions. The terminal sends this information to the server, which first receives the instruction manual data.

[0037] The server analyzes the text and images in the received instruction manual. Here, natural language processing techniques are used to extract instructions and procedures from the document, and image recognition techniques are used to identify the parts needed for assembly.

[0038] Based on the analysis, the server accesses an existing video database and loads relevant video clips into the learning model. The generative AI model is designed to automatically generate customized assembly videos based on these video clips, according to the user's assembly procedure.

[0039] The generated video is formatted, including multilingual options, and delivered to the user via their device. By watching this assembly video on their device, the user can assemble the product efficiently and accurately.

[0040] For example, if a user wants to assemble furniture, they can upload a specific product number and the associated instructions to their device to begin the process. The server analyzes the instructions to identify each part of the furniture and the assembly steps. Next, it learns from relevant assembly videos and generates an assembly video specifically for that furniture. This video clearly shows the furniture assembly process, allowing the user to follow the necessary steps without getting lost.

[0041] In this way, the present invention assists users in assembling products in a way that is intuitively easy to understand. By performing analysis on a server and generating automated videos using AI, a highly personalized assembly guide can be provided.

[0042] The following describes the processing flow.

[0043] Step 1:

[0044] The user uses a terminal to enter the identification information of the product they want to assemble (e.g., part number or model number) and upload the product's instruction manual. The terminal then sends this data to the server.

[0045] Step 2:

[0046] The server analyzes the instruction manual data received from the terminal. Using natural language processing, it reads assembly procedures and precautions from the manual and extracts them as text data.

[0047] Step 3:

[0048] The server analyzes images within the instruction manual and uses image recognition technology to identify necessary parts and assembly steps. This lays the foundation for generating assembly instructions that take into account not only textual information but also visual information.

[0049] Step 4:

[0050] The server retrieves existing video content from the database and feeds it into a learning model. The generative AI model then uses this to identify appropriate video patterns that meet the user's assembly needs.

[0051] Step 5:

[0052] The generative AI model combines analysis results with pre-trained videos to automatically generate assembly videos for specific products. The videos visually and intuitively demonstrate the assembly process.

[0053] Step 6:

[0054] The server converts the generated video to the optimal format and allows for the addition of subtitle options in multiple languages. This enables users to view the video in their chosen language.

[0055] Step 7:

[0056] The device receives assembly videos streamed from the server, allowing users to stream or download the videos. Users can then refer to the videos to assemble the product correctly.

[0057] (Example 1)

[0058] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0059] In traditional assembly processes, a challenge exists in that the procedures described in product manuals are often complex or insufficient, making it difficult for users to assemble products accurately. Furthermore, the time required to understand the procedures can hinder efficient assembly.

[0060] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0061] In this invention, the server includes a device for receiving instruction manual data and product information, a device for analyzing the received instruction manual and extracting assembly procedures, a device for reading relevant video information from a data store based on the analysis and automatically generating an assembly video, and a device for distributing the generated video to the user's device. This makes it possible for the user to assemble the product intuitively and efficiently.

[0062] "Instruction manual data" refers to documentary information provided to users for assembling a product, including procedures and parts information.

[0063] "Product information" refers to information necessary to identify the product to be assembled, and includes the model number and name.

[0064] "Device" refers to a system of hardware or software configured to perform a specific function.

[0065] "Analysis" refers to the calculations and algorithms used to process data and understand its structure and meaning.

[0066] "Assembly instructions" refer to a series of steps and instructions necessary to correctly assemble a product.

[0067] "Visual information" refers to visual data, including videos and images.

[0068] A "data store" refers to a place where information is stored, such as a database or storage system.

[0069] "Automated generation" refers to the process of creating products using programs or algorithms without human intervention.

[0070] "User equipment" refers to terminals or devices used by users to input or display information.

[0071] This system consists of three elements: user, terminal, and server, which together process information related to product assembly.

[0072] The user uses a terminal to input information about the product they wish to assemble and its accompanying instructions. The terminal can be a tablet or a personal computer. Common file formats include PDF and image files. The terminal formats this information appropriately and sends it to the server.

[0073] The servers are located, for example, on a cloud platform or in a proprietary data center. When processing the received instruction manual data, the servers utilize natural language processing and image recognition technologies. In this case, general API services are used as natural language processing software. For image recognition, various tools are used to identify parts.

[0074] Based on the received data, the server accesses an existing video database and collects the relevant video data. This video data is then reconstructed into a customized assembly video using a generative AI model, according to the user's assembly procedure. In this process, the video is automatically generated using a learned model.

[0075] The generated videos are converted into a multilingual format and delivered to the user via their device. This allows users to assemble products efficiently and accurately. For example, if a user wants to assemble furniture, they can upload the product model number and instructions to their device and start the process with a prompt that says, "Upload the instructions for the product that needs assembly and generate a customized assembly video."

[0076] In this way, by utilizing server-side data analysis and AI-driven automated generation technology, we can provide personalized assembly guides.

[0077] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0078] Step 1:

[0079] The user uses the terminal to input information about the product they want to assemble and its accompanying instruction manual. Specifically, they upload text data such as the product name and model number, as well as PDF or image data of the instruction manual to the terminal. The terminal receives the input data, verifies its format, and then prepares to send it to the server.

[0080] Step 2:

[0081] The terminal sends product information and instruction manual data received from the user to the server. Input data (text data, PDF, images) is sent to the server. The terminal encodes this data using the appropriate protocol and passes it to the server over the network.

[0082] Step 3:

[0083] The server analyzes the received instruction manual data. Specifically, it uses natural language processing technology to extract assembly instructions from the text data of the instruction manual and image recognition technology to identify necessary parts information from image data. Through these analyses, a list of necessary parts and an outline of the assembly procedure are generated for the product.

[0084] Step 4:

[0085] Based on the analysis results, the server accesses an existing video database and collects relevant video clips. The server uses the analyzed procedural information as input to search the video database and find the corresponding video clips. The found clips are then loaded for subsequent generative model processing.

[0086] Step 5:

[0087] The server uses a generative AI model to automatically generate customized assembled videos based on the loaded video clips and analysis results. Here, the generative AI model uses the analyzed procedures to edit the video clips and create videos tailored to the user's specifications. The generated videos may also include narration and text explanations to enhance user viewing experience.

[0088] Step 6:

[0089] The server converts the generated video into a format that can be delivered in the user's selected language and sends it to the device. The server also places multilingual subtitles in the video file and adjusts it to be easily understood by the user. The converted video data is sent to the device and delivered to the user.

[0090] Step 7:

[0091] The device receives a customized video streamed from the server and prepares to play it. Users can intuitively understand the assembly procedure by watching the assembly video on the device's screen. This allows users to assemble the product efficiently.

[0092] (Application Example 1)

[0093] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0094] Traditional assembly procedures are prone to human error and time-inefficient processes when assembling complex items. This hinders productivity improvements and increases the burden on on-site workers. In particular, there is a need for accurate and efficient procedures for assembling machinery and equipment used in factories and manufacturing sites.

[0095] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0096] In this invention, the server includes means for receiving instruction manual information and product data, means for analyzing the received instruction manual and extracting operating procedures, means for machine learning existing video data and generating operating videos based on the analyzed procedures, and means for displaying the generated operating videos on an augmented reality device. This allows field workers to obtain real-time visual assembly guides through the augmented reality device. This enables accurate assembly work and reduces assembly time.

[0097] "Instruction manual information" refers to document data containing instructions and procedures necessary for assembling or using an item.

[0098] "Item data" refers to information about the products or parts that are to be assembled or operated.

[0099] "Means of receiving data" refers to the methods and functions by which a server acquires data from external sources and uses it for analysis.

[0100] "Means of analysis and extraction of operating procedures" refers to the process of analyzing received data and concretizing and clarifying each step.

[0101] "Existing video data" refers to visual media information that has been collected or stored in the past.

[0102] "Machine learning" is a technique that allows computers to recognize patterns in data and build predictive models.

[0103] "Operation video" refers to video data that visually demonstrates the assembly and operation procedures.

[0104] "Means of generation" refers to the processes and technologies used to create new images.

[0105] An "augmented reality device" is an electronic device that uses technology to overlay digital information onto the real world's field of view.

[0106] "Means of display" refers to the method of presenting generated digital content to the user's field of vision.

[0107] In implementing this invention, the server receives instruction manual information and product data via a network. The received information is analyzed using natural language processing techniques such as NLTK and spaCy to extract operating procedures. Furthermore, visual information is analyzed using OpenCV to perform element recognition. This allows the server to clearly identify the necessary steps and parts.

[0108] Next, existing video data is used as part of machine learning, and a generative AI model using a deep learning framework such as PyTorch generates operation videos based on the operation procedures. These generated videos are then formatted using multilingual support technology and provided to the user's augmented reality device. This allows users to intuitively understand the work content and prevents errors.

[0109] As a concrete example, when an assembly worker manufactures a new home appliance, they input the instruction manual information into smart glasses. At this time, the server analyzes the received information and generates an assembly video specifically for that product. The generated video is displayed in real time on the worker's smart glasses, supporting assembly according to the correct procedure. An example of a prompt message used in this case would be, "Read the instruction manual for the home appliance to be assembled, identify the necessary parts, and generate a customized assembly video based on the contents of the manual."

[0110] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0111] Step 1:

[0112] The terminal receives product information and instruction manual data from the user as input. This aggregates specific item information that the user wishes to assemble. The terminal then transmits this information to the server.

[0113] Step 2:

[0114] The server receives product information and instruction manual data sent from the terminal as input. Using natural language processing technologies such as NLTK and spaCy, the instruction manual data is analyzed and operating procedures are extracted. The extracted procedure information is used for video generation in the next step.

[0115] Step 3:

[0116] The server uses image recognition technology based on OpenCV to analyze visual information from instruction manual data. This analysis allows for the identification of parts and the acquisition of layout information. The obtained part information is incorporated into the operating procedure and treated as information necessary for video generation.

[0117] Step 4:

[0118] The server utilizes a generative AI model based on PyTorch to generate video data from an existing video database according to the operating procedure. Taking the operating procedure and parts information as input, it outputs a detailed assembly video. This video is customized to be easily understood by the user.

[0119] Step 5:

[0120] The server converts the generated video feed into a multilingual format suitable for augmented reality devices. This allows the video to be displayed to different users in various languages. The formatted video data is then sent to the terminal or augmented reality device.

[0121] Step 6:

[0122] Users view streamed instructional videos using augmented reality devices (e.g., smart glasses). These devices overlay the video onto the user's field of view, providing real-time assembly support. This allows users to perform assembly tasks efficiently and accurately.

[0123] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0124] This invention is implemented using a program that operates between the user, a terminal, and a server. In particular, it provides more effective assembly assistance by recognizing the user's emotional state and adjusting the assembly guide accordingly.

[0125] First, the user uses a terminal to enter information about the product they want to assemble and uploads the product's instruction manual. The terminal sends this data to the server. The server analyzes the instruction manual data and extracts the assembly steps and important points.

[0126] The server further analyzes image information within the instruction manual to recognize parts. Based on this analysis, it retrieves appropriate clips from a pre-trained video database and uses a generation AI model to automatically generate assembly videos suitable for the specific product.

[0127] This system incorporates an emotion engine that uses the device's camera and sensors to recognize the user's emotional state in real time. The data obtained by the emotion engine is used to dynamically adjust the content and pace of the assembled video. For example, if the emotion engine detects that the user is confused, the video's pace will be slowed down and adjusted to provide more detailed explanations.

[0128] The assembly videos generated in this way have multilingual options and are delivered from the server to the terminal. Users perform the assembly work while watching the videos via streaming or download. User emotion data is recorded and used to improve the system's performance.

[0129] For example, if a user wants to assemble new furniture, they input the product information into the system, and a personalized video is provided to guide them through the assembly process. During this process, the system detects the user's emotions from their facial expressions and tone of voice, and if anxiety about assembly is detected, the explanation becomes more detailed, providing a guide tailored to the user.

[0130] In this way, the present invention realizes a flexible assembly guide that is tailored to the user's situation, providing an environment in which users can assemble products more comfortably.

[0131] The following describes the processing flow.

[0132] Step 1:

[0133] The user uses a terminal to input the identification information and instructions for the product they want to assemble into the system. The terminal then sends this data to the server.

[0134] Step 2:

[0135] The server analyzes the instruction manual data received from the terminal. It uses natural language processing technology to extract assembly instructions and precautions in text format, and then uses image recognition technology to identify parts.

[0136] Step 3:

[0137] Based on the analyzed procedure, the server selects relevant video clips from an existing video database and loads them into the generation AI model. This model learns the patterns necessary to generate appropriate assembled videos.

[0138] Step 4:

[0139] The generative AI model automatically generates product assembly videos using analysis results and trained data. The videos include scenarios that visually demonstrate each step of the assembly process.

[0140] Step 5:

[0141] The emotion engine analyzes the user's emotional state in real time through the device's camera and microphone. The engine analyzes the user's facial expressions and voice to identify their current emotional state.

[0142] Step 6:

[0143] The server receives feedback from the emotion engine and dynamically adjusts the content of the generated assembled video. For example, if the user is feeling anxious, the video speed is slowed down and additional explanations are added to improve support.

[0144] Step 7:

[0145] The device delivers a pre-configured assembly video to the user. The user can play the video and assemble the product according to the instructions. User sentiment data is recorded for future improvements.

[0146] By coordinating each step in this way, we can provide an assembly guide that resonates with the user's emotions, resulting in increased work efficiency and a better user experience.

[0147] (Example 2)

[0148] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0149] Conventional assembly support systems were unable to take into account the confusion and anxiety users experienced during the assembly process in real time, making it difficult to provide optimal assembly assistance. Furthermore, the generated assembly instructions were dependent on specific languages, resulting in a lack of multilingual support.

[0150] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0151] In this invention, the server includes means for acquiring instruction manual information and product data, means for analyzing the acquired instruction manual and extracting assembly procedures, and means for recognizing the user's emotional state and dynamically adjusting the content and speed of the assembly video. This makes it possible to provide personalized assembly support while taking into account the user's emotions. Furthermore, flexible support through multilingual capabilities allows for the provision of appropriate assembly guides to users who speak different languages.

[0152] "Instruction manual information" refers to document data containing a series of instructions and procedures necessary for the user to assemble the product.

[0153] "Item data" refers to information related to the goods or products to be assembled, and includes product name, model, specifications, etc.

[0154] "Part identification" is the process of analyzing instruction manuals and image data to identify the individual parts and components needed for assembly.

[0155] "Video information" refers to video data used to visually demonstrate assembly procedures and related work processes.

[0156] "User emotional state" refers to the type and intensity of emotions felt by the user performing the assembly, and is recognized through analysis of facial expressions and voice.

[0157] "Multilingual support" refers to a feature designed to allow users who speak different languages ​​to understand the same information, meaning that information can be provided in multiple languages.

[0158] This invention is implemented by a system that operates between a user, a terminal, and a server. The user inputs information about the product they wish to assemble via the terminal and uploads the corresponding instruction manual file. The terminal used here is a general-purpose computer or smart device equipped with a camera and sensors.

[0159] The terminal receives instructions from the user and sends that information and instruction manual data to the server. The server receives this data and uses optical character recognition (OCR) and image analysis technologies to extract assembly procedures, precautions, and parts information from the instruction manual.

[0160] For analyzing instruction manuals, we use Tesseract, a Python®-based OCR library, and for image analysis, we use OpenCV. The server utilizes the extracted data to select appropriate video clips from a pre-trained video database. Then, a generative AI model is used to automatically generate product-specific assembly videos. The AI ​​model, for example, is a Transformer-based model and creates videos based on prompts such as "Show the assembly procedure for part A of this product."

[0161] Furthermore, the emotion engine built into the device analyzes the user's facial expressions and voice tone in real time to recognize the emotional state the user is experiencing. This data is used on the server to dynamically adjust the content and speed of the assembled video. The video is provided in multiple languages ​​according to the user's needs and is delivered from the server to the device via streaming or download.

[0162] For example, if the emotion engine detects that a user is feeling anxious about assembling new furniture, the server will adjust the video instructions to be more detailed and slower. Through this process, the user receives an emotionally sensitive and personalized guide, enabling them to proceed with the assembly process more comfortably.

[0163] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0164] Step 1:

[0165] The user uses a terminal to input information and instructions for the product they want to assemble, and then uploads them. This sends the terminal the basic information necessary for the user to assemble the product. The input information includes the product name, model number, and a PDF file of the instruction manual. The terminal receives this input and prepares to send it to the server.

[0166] Step 2:

[0167] The device sends information received from the user and instruction manual data to the server. The HTTP protocol is used for transmission, ensuring the data is securely transferred to the server. This provides the server with data for analysis.

[0168] Step 3:

[0169] The server uses optical character recognition (OCR) technology to convert the received instruction manual data into text and extract the assembly instructions. Specifically, it uses the Tesseract library to scan the instruction manual's PDF or image data and obtain the information as text. This output is text data organized step by step.

[0170] Step 4:

[0171] The server uses an image analysis algorithm to analyze the image information in the instruction manual and identify parts. Using libraries such as OpenCV, it recognizes objects in the image and identifies parts based on their characteristics such as shape and color. As a result, data for each part is output.

[0172] Step 5:

[0173] The server uses the extracted procedure and parts data to select the appropriate video clip from a trained video database. The video database stores past assembly videos categorized by assembly procedure, and the server automatically retrieves the most suitable clip according to the required procedure.

[0174] Step 6:

[0175] The server uses a generation AI model to automatically generate assembly videos for a specific product based on selected video clips. The generation AI model is input with prompts, and edits and combines the video accordingly. For example, the prompt "Show the steps to assemble part A" is used, and the edited video is generated as output.

[0176] Step 7:

[0177] The device uses an onboard emotion engine to recognize the user's emotional state in real time. It employs an algorithm that analyzes the user's voice tone and facial expressions using the camera and microphone. This process yields user emotional data as output.

[0178] Step 8:

[0179] The server dynamically adjusts the content and pace of the generated video based on the emotional data received from the device. If the user expresses anxiety or confusion, adjustments are made, such as slowing down video playback and providing more detailed explanations. The adjusted video is then provided to the device as the final output.

[0180] Step 9:

[0181] Users view videos provided by the server via streaming or download on their devices. The videos are available in multiple languages, allowing users to proceed with assembly while receiving assembly guides in their own language.

[0182] (Application Example 2)

[0183] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0184] In modern manufacturing environments, workers are required to assemble products quickly and accurately when faced with complex work procedures. However, conventional manuals and fixed instructions make it difficult to flexibly adapt to the experience and emotional state of the operators. This can lead to a decrease in on-site efficiency. This invention aims to improve manufacturing efficiency by providing a system that takes the emotional state of the worker into consideration and instantly adapts work instructions.

[0185] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0186] In this invention, the server includes means for receiving explanatory data and item information; means for analyzing the received explanatory data and extracting work procedures; means for learning existing video data and generating action videos based on the analyzed procedures; means for detecting the user's emotional state and dynamically adjusting the content and speed of the action videos; and means for distributing the generated videos to the operator. This enables flexible work support that responds to the emotional state of the worker.

[0187] "Explanatory data" refers to documents or files that contain information necessary for assembling or operating an item.

[0188] "Item information" refers to data that describes the characteristics of a specific item, such as its name, model number, and specifications.

[0189] "Means of receiving" refers to functions or devices for taking in data from external sources.

[0190] "Analysis" is the process of breaking down received data into an easily understandable form and extracting the necessary information.

[0191] A "work procedure" is a set of instructions that outlines the steps required to complete a specific task.

[0192] "Existing video data" refers to video files and clips that were filmed or recorded in the past.

[0193] "Learning" is the process of finding patterns and rules from existing data and using them for new data processing.

[0194] "Means for generating motion videos" refers to devices or functions that create videos to provide visualized instructions and guidance based on analysis results and training data.

[0195] "Detecting the user's emotional state" refers to the process of identifying the user's emotions at that moment based on their facial expressions, voice, and other factors.

[0196] "Means of dynamic adjustment" refers to the ability to flexibly change processes and outputs in response to real-time changing situations and requirements.

[0197] "Means of distribution to operators" refers to technologies and methods for transmitting generated information to specific workers.

[0198] The system for realizing this application is configured as follows: First, the user uses a terminal to input descriptive data and item information about the items to be assembled. The terminal sends this data to a server. The server analyzes the received descriptive data and extracts the work procedure. Natural language processing (NLP) technology is used for the analysis. The server also learns from existing video data and generates action videos based on the extracted work procedure. Here, a generative AI model is used to create videos suitable for specific tasks.

[0199] The server further detects the user's emotional state in real time through the camera and microphone installed on the user's device. This detection uses an emotion recognition engine that analyzes the user's facial expressions and voice data. Based on the detected emotional state, the server dynamically adjusts the video and pace of the work instructions. Hardware used includes, for example, Microsoft® HoloLens®, and software such as Google® Cloud Vision API and Microsoft Azure® Face API are used.

[0200] The generated video footage is distributed to operators and supports multiple languages, making it easy to implement in international factories. For example, when introducing a new manufacturing line to a factory, workers wear HoloLens and perform tasks while watching guide videos generated by the system. If the user shows a confused expression, the system slows down the pace of the guide and adds detailed instructions such as, "Let me explain it again slowly."

[0201] Examples of prompts for a generative AI model are as follows:

[0202] Design an AI model that dynamically adjusts the guide video displayed on the HMD based on the user's emotions (e.g., frustration, anxiety). Consider how to make the user perform the task more smoothly and generate appropriate feedback.

[0203] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0204] Step 1:

[0205] The user uses a terminal to input description data and item information and sends it to the system. The input data consists of files and digital information, including specific assembly requirements and specifications for the item. The server then receives the item information.

[0206] Step 2:

[0207] The server analyzes the received explanatory data and extracts the work procedure. This analysis utilizes natural language processing techniques to identify the necessary steps from the input text-based explanatory data. As a result of the analysis, specific work procedures are extracted.

[0208] Step 3:

[0209] The server learns from existing video data and generates action videos based on the analyzed work procedures. Here, a generative AI model is used to create guide videos suitable for the procedures. The input is a conventional video clip, and the output is a new action video.

[0210] Step 4:

[0211] The server detects the user's emotional state in real time through the terminal's camera and microphone. An emotion recognition engine is used for detection, analyzing captured images and audio to determine the user's emotions. The input for this analysis is image and audio data, and the output is the user's emotional state.

[0212] Step 5:

[0213] The server dynamically adjusts the content and speed of the action video according to the detected emotional state. For example, if it detects that the user is confused, it slows down the pace of the guide video and adds detailed narration. This generates video instructions that are appropriate for the user.

[0214] Step 6:

[0215] The server delivers adjusted motion video to the terminal and supports multiple languages, making it usable in international environments. Ultimately, the video is presented to the user, who can use it safely and effectively. The input for delivery is the adjusted motion video, and the output is the video display on the user's terminal.

[0216] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0217] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0218] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0219] [Second Embodiment]

[0220] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0221] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0222] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0223] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0224] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0225] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0226] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0227] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0228] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0229] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0230] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0231] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0232] The system for carrying out this invention uses a program that operates between the user, terminal, and server. The program processing and specific embodiments of this system are described below.

[0233] The user uses a terminal to input information about the product they wish to assemble and its accompanying instructions. The terminal sends this information to the server, which first receives the instruction manual data.

[0234] The server analyzes the text and images in the received instruction manual. Here, natural language processing techniques are used to extract instructions and procedures from the document, and image recognition techniques are used to identify the parts needed for assembly.

[0235] Based on the analysis, the server accesses an existing video database and loads relevant video clips into the learning model. The generative AI model is designed to automatically generate customized assembly videos based on these video clips, according to the user's assembly procedure.

[0236] The generated video is formatted, including multilingual options, and delivered to the user via their device. By watching this assembly video on their device, the user can assemble the product efficiently and accurately.

[0237] For example, if a user wants to assemble furniture, they can upload a specific product number and the associated instructions to their device to begin the process. The server analyzes the instructions to identify each part of the furniture and the assembly steps. Next, it learns from relevant assembly videos and generates an assembly video specifically for that furniture. This video clearly shows the furniture assembly process, allowing the user to follow the necessary steps without getting lost.

[0238] In this way, the present invention assists users in assembling products in a way that is intuitively easy to understand. By performing analysis on a server and generating automated videos using AI, a highly personalized assembly guide can be provided.

[0239] The following describes the processing flow.

[0240] Step 1:

[0241] The user uses a terminal to enter the identification information of the product they want to assemble (e.g., part number or model number) and upload the product's instruction manual. The terminal then sends this data to the server.

[0242] Step 2:

[0243] The server analyzes the instruction manual data received from the terminal. Using natural language processing, it reads assembly procedures and precautions from the manual and extracts them as text data.

[0244] Step 3:

[0245] The server analyzes images within the instruction manual and uses image recognition technology to identify necessary parts and assembly steps. This lays the foundation for generating assembly instructions that take into account not only textual information but also visual information.

[0246] Step 4:

[0247] The server retrieves existing video content from the database and feeds it into a learning model. The generative AI model then uses this to identify appropriate video patterns that meet the user's assembly needs.

[0248] Step 5:

[0249] The generative AI model combines analysis results with pre-trained videos to automatically generate assembly videos for specific products. The videos visually and intuitively demonstrate the assembly process.

[0250] Step 6:

[0251] The server converts the generated video to the optimal format and allows for the addition of subtitle options in multiple languages. This enables users to view the video in their chosen language.

[0252] Step 7:

[0253] The device receives assembly videos streamed from the server, allowing users to stream or download the videos. Users can then refer to the videos to assemble the product correctly.

[0254] (Example 1)

[0255] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0256] In traditional assembly processes, a challenge exists in that the procedures described in product manuals are often complex or insufficient, making it difficult for users to assemble products accurately. Furthermore, the time required to understand the procedures can hinder efficient assembly.

[0257] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0258] In this invention, the server includes a device for receiving instruction manual data and product information, a device for analyzing the received instruction manual and extracting assembly procedures, a device for reading relevant video information from a data store based on the analysis and automatically generating an assembly video, and a device for distributing the generated video to the user's device. This makes it possible for the user to assemble the product intuitively and efficiently.

[0259] "Instruction manual data" refers to documentary information provided to users for assembling a product, including procedures and parts information.

[0260] "Product information" refers to information necessary to identify the product to be assembled, and includes the model number and name.

[0261] "Device" refers to a system of hardware or software configured to perform a specific function.

[0262] "Analysis" refers to the calculations and algorithms used to process data and understand its structure and meaning.

[0263] "Assembly instructions" refer to a series of steps and instructions necessary to correctly assemble a product.

[0264] "Visual information" refers to visual data, including videos and images.

[0265] A "data store" refers to a place where information is stored, such as a database or storage system.

[0266] "Automated generation" refers to the process of creating products using programs or algorithms without human intervention.

[0267] "User equipment" refers to terminals or devices used by users to input or display information.

[0268] This system consists of three elements: user, terminal, and server, which together process information related to product assembly.

[0269] The user uses a terminal to input information about the product they wish to assemble and its accompanying instructions. The terminal can be a tablet or a personal computer. Common file formats include PDF and image files. The terminal formats this information appropriately and sends it to the server.

[0270] The servers are located, for example, on a cloud platform or in a proprietary data center. When processing the received instruction manual data, the servers utilize natural language processing and image recognition technologies. In this case, general API services are used as natural language processing software. For image recognition, various tools are used to identify parts.

[0271] Based on the received data, the server accesses an existing video database and collects the relevant video data. This video data is then reconstructed into a customized assembly video using a generative AI model, according to the user's assembly procedure. In this process, the video is automatically generated using a learned model.

[0272] The generated videos are converted into a multilingual format and delivered to the user via their device. This allows users to assemble products efficiently and accurately. For example, if a user wants to assemble furniture, they can upload the product model number and instructions to their device and start the process with a prompt that says, "Upload the instructions for the product that needs assembly and generate a customized assembly video."

[0273] In this way, by utilizing server-side data analysis and AI-driven automated generation technology, we can provide personalized assembly guides.

[0274] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0275] Step 1:

[0276] The user uses the terminal to input information about the product they want to assemble and its accompanying instruction manual. Specifically, they upload text data such as the product name and model number, as well as PDF or image data of the instruction manual to the terminal. The terminal receives the input data, verifies its format, and then prepares to send it to the server.

[0277] Step 2:

[0278] The terminal sends product information and instruction manual data received from the user to the server. Input data (text data, PDF, images) is sent to the server. The terminal encodes this data using the appropriate protocol and passes it to the server over the network.

[0279] Step 3:

[0280] The server analyzes the received manual data. Specifically, it extracts the assembly procedures from the text data of the manual using natural language processing technology, and identifies the necessary part information from the image data using image recognition technology. Through these analyses, a parts list required for the product and an outline of the assembly procedures are generated.

[0281] Step 4:

[0282] Based on the analysis results, the server accesses the existing video database and collects relevant video clips. The server uses the procedure information analyzed for querying the video database as input to search for the corresponding video clips. The found clips are loaded for subsequent generation model processing.

[0283] Step 5:

[0284] The server uses the generation AI model to automatically generate a customized assembly video based on the loaded video clips and the analysis results. Here, the generation AI model uses the analyzed procedures to edit the video clips and produce a video that meets the user specifications. The generated video may also have a narration and text description added to consider the user's viewing convenience.

[0285] Step 6:

[0286] The server converts the generated video into a format that can be delivered in the language selected by the user and sends it to the terminal. The server places multilingual subtitles in the video file and adjusts it for easy understanding by the user. The converted video data is sent to the terminal and reaches the user.

[0287] Step 7:

[0288] The device receives a customized video streamed from the server and prepares to play it. Users can intuitively understand the assembly procedure by watching the assembly video on the device's screen. This allows users to assemble the product efficiently.

[0289] (Application Example 1)

[0290] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0291] Traditional assembly procedures are prone to human error and time-inefficient processes when assembling complex items. This hinders productivity improvements and increases the burden on on-site workers. In particular, there is a need for accurate and efficient procedures for assembling machinery and equipment used in factories and manufacturing sites.

[0292] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0293] In this invention, the server includes means for receiving instruction manual information and product data, means for analyzing the received instruction manual and extracting operating procedures, means for machine learning existing video data and generating operating videos based on the analyzed procedures, and means for displaying the generated operating videos on an augmented reality device. This allows field workers to obtain real-time visual assembly guides through the augmented reality device. This enables accurate assembly work and reduces assembly time.

[0294] "Instruction manual information" refers to document data containing instructions and procedures necessary for assembling or using an item.

[0295] "Item data" refers to information about the products or parts that are to be assembled or operated.

[0296] "Means of receiving data" refers to the methods and functions by which a server acquires data from external sources and uses it for analysis.

[0297] "Means of analysis and extraction of operating procedures" refers to the process of analyzing received data and concretizing and clarifying each step.

[0298] "Existing video data" refers to visual media information that has been collected or stored in the past.

[0299] "Machine learning" is a technique that allows computers to recognize patterns in data and build predictive models.

[0300] "Operation video" refers to video data that visually demonstrates the assembly and operation procedures.

[0301] "Means of generation" refers to the processes and technologies used to create new images.

[0302] An "augmented reality device" is an electronic device that uses technology to overlay digital information onto the real world's field of view.

[0303] "Means of display" refers to the method of presenting generated digital content to the user's field of vision.

[0304] In implementing this invention, the server receives instruction manual information and product data via a network. The received information is analyzed using natural language processing techniques such as NLTK and spaCy to extract operating procedures. Furthermore, visual information is analyzed using OpenCV to perform element recognition. This allows the server to clearly identify the necessary steps and parts.

[0305] Next, existing video data is utilized as part of machine learning, and an operation video based on the operation procedure is generated using a generative AI model with a deep learning framework such as PyTorch. The generated video is format-converted using multilingual technology and provided to the augmented reality device used by the user. As a result, the user can intuitively understand the work content and prevent incorrect operations.

[0306] As a specific example, when an assembly worker manufactures a new household appliance product, the instruction manual information is input into smart glasses. At this time, the server analyzes the received information and generates an assembly video dedicated to that product. The generated video is displayed in real time on the worker's smart glasses to support the assembly along the correct procedure. As an example of the prompt sentence at this time, an instruction such as "Read the instruction manual of the household appliance product to be assembled, identify the necessary parts, and generate a customized assembly video based on the content of the instruction manual." is used.

[0307] The flow of the specific process in Application Example 1 will be described using FIG. 12.

[0308] Step 1:

[0309] The terminal receives, as inputs, product information and instruction manual data from the user. As a result, the specific item information that the user wishes to assemble is aggregated in the terminal. The terminal transmits this information to the server.

[0310] Step 2:

[0311] The server receives, as inputs, the product information and instruction manual data transmitted from the terminal. Using NLTK or spaCy as natural language processing technologies, the instruction manual data is analyzed to extract the operation procedure. The extracted procedure information is used for video generation in the next step.

[0312] Step 3:

[0313] The server uses image recognition technology based on OpenCV to analyze visual information from instruction manual data. This analysis allows for the identification of parts and the acquisition of layout information. The obtained part information is incorporated into the operating procedure and treated as information necessary for video generation.

[0314] Step 4:

[0315] The server utilizes a generative AI model based on PyTorch to generate video data from an existing video database according to the operating procedure. Taking the operating procedure and parts information as input, it outputs a detailed assembly video. This video is customized to be easily understood by the user.

[0316] Step 5:

[0317] The server converts the generated video feed into a multilingual format suitable for augmented reality devices. This allows the video to be displayed to different users in various languages. The formatted video data is then sent to the terminal or augmented reality device.

[0318] Step 6:

[0319] Users view streamed instructional videos using augmented reality devices (e.g., smart glasses). These devices overlay the video onto the user's field of view, providing real-time assembly support. This allows users to perform assembly tasks efficiently and accurately.

[0320] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0321] This invention is implemented using a program that operates between the user, a terminal, and a server. In particular, it provides more effective assembly assistance by recognizing the user's emotional state and adjusting the assembly guide accordingly.

[0322] First, the user uses a terminal to enter information about the product they want to assemble and uploads the product's instruction manual. The terminal sends this data to the server. The server analyzes the instruction manual data and extracts the assembly steps and important points.

[0323] The server further analyzes image information within the instruction manual to recognize parts. Based on this analysis, it retrieves appropriate clips from a pre-trained video database and uses a generation AI model to automatically generate assembly videos suitable for the specific product.

[0324] This system incorporates an emotion engine that uses the device's camera and sensors to recognize the user's emotional state in real time. The data obtained by the emotion engine is used to dynamically adjust the content and pace of the assembled video. For example, if the emotion engine detects that the user is confused, the video's pace will be slowed down and adjusted to provide more detailed explanations.

[0325] The assembly videos generated in this way have multilingual options and are delivered from the server to the terminal. Users perform the assembly work while watching the videos via streaming or download. User emotion data is recorded and used to improve the system's performance.

[0326] For example, if a user wants to assemble new furniture, they input the product information into the system, and a personalized video is provided to guide them through the assembly process. During this process, the system detects the user's emotions from their facial expressions and tone of voice, and if anxiety about assembly is detected, the explanation becomes more detailed, providing a guide tailored to the user.

[0327] In this way, the present invention realizes a flexible assembly guide that is tailored to the user's situation, providing an environment in which users can assemble products more comfortably.

[0328] The following describes the processing flow.

[0329] Step 1:

[0330] The user uses a terminal to input the identification information and instructions for the product they want to assemble into the system. The terminal then sends this data to the server.

[0331] Step 2:

[0332] The server analyzes the instruction manual data received from the terminal. It uses natural language processing technology to extract assembly instructions and precautions in text format, and then uses image recognition technology to identify parts.

[0333] Step 3:

[0334] Based on the analyzed procedure, the server selects relevant video clips from an existing video database and loads them into the generation AI model. This model learns the patterns necessary to generate appropriate assembled videos.

[0335] Step 4:

[0336] The generative AI model automatically generates product assembly videos using analysis results and trained data. The videos include scenarios that visually demonstrate each step of the assembly process.

[0337] Step 5:

[0338] The emotion engine analyzes the user's emotional state in real time through the device's camera and microphone. The engine analyzes the user's facial expressions and voice to identify their current emotional state.

[0339] Step 6:

[0340] The server receives feedback from the emotion engine and dynamically adjusts the content of the generated assembled video. For example, if the user is feeling anxious, the video speed is slowed down and additional explanations are added to improve support.

[0341] Step 7:

[0342] The device delivers a pre-configured assembly video to the user. The user can play the video and assemble the product according to the instructions. User sentiment data is recorded for future improvements.

[0343] By coordinating each step in this way, we can provide an assembly guide that resonates with the user's emotions, resulting in increased work efficiency and a better user experience.

[0344] (Example 2)

[0345] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0346] Conventional assembly support systems were unable to take into account the confusion and anxiety users experienced during the assembly process in real time, making it difficult to provide optimal assembly assistance. Furthermore, the generated assembly instructions were dependent on specific languages, resulting in a lack of multilingual support.

[0347] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0348] In this invention, the server includes means for acquiring instruction manual information and product data, means for analyzing the acquired instruction manual and extracting assembly procedures, and means for recognizing the user's emotional state and dynamically adjusting the content and speed of the assembly video. This makes it possible to provide personalized assembly support while taking into account the user's emotions. Furthermore, flexible support through multilingual capabilities allows for the provision of appropriate assembly guides to users who speak different languages.

[0349] "Instruction manual information" refers to document data containing a series of instructions and procedures necessary for the user to assemble the product.

[0350] "Item data" refers to information related to the goods or products to be assembled, and includes product name, model, specifications, etc.

[0351] "Part identification" is the process of analyzing instruction manuals and image data to identify the individual parts and components needed for assembly.

[0352] "Video information" refers to video data used to visually demonstrate assembly procedures and related work processes.

[0353] "User emotional state" refers to the type and intensity of emotions felt by the user performing the assembly, and is recognized through analysis of facial expressions and voice.

[0354] "Multilingual support" refers to a feature designed to allow users who speak different languages ​​to understand the same information, meaning that information can be provided in multiple languages.

[0355] This invention is implemented by a system that operates between a user, a terminal, and a server. The user inputs information about the product they wish to assemble via the terminal and uploads the corresponding instruction manual file. The terminal used here is a general-purpose computer or smart device equipped with a camera and sensors.

[0356] The terminal receives instructions from the user and sends that information and instruction manual data to the server. The server receives this data and uses optical character recognition (OCR) and image analysis technologies to extract assembly procedures, precautions, and parts information from the instruction manual.

[0357] For analyzing instruction manuals, we use Tesseract, a Python-based OCR library, and for image analysis, we use OpenCV. The server utilizes the extracted data to select appropriate video clips from a pre-trained video database. Then, a generative AI model is used to automatically generate product-specific assembly videos. The AI ​​model, for example, is a Transformer-based model, and it creates videos based on prompts such as "Show the assembly procedure for part A of this product."

[0358] Furthermore, the emotion engine built into the device analyzes the user's facial expressions and voice tone in real time to recognize the emotional state the user is experiencing. This data is used on the server to dynamically adjust the content and speed of the assembled video. The video is provided in multiple languages ​​according to the user's needs and is delivered from the server to the device via streaming or download.

[0359] For example, if the emotion engine detects that a user is feeling anxious about assembling new furniture, the server will adjust the video instructions to be more detailed and slower. Through this process, the user receives an emotionally sensitive and personalized guide, enabling them to proceed with the assembly process more comfortably.

[0360] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0361] Step 1:

[0362] The user uses a terminal to input information and instructions for the product they want to assemble, and then uploads them. This sends the terminal the basic information necessary for the user to assemble the product. The input information includes the product name, model number, and a PDF file of the instruction manual. The terminal receives this input and prepares to send it to the server.

[0363] Step 2:

[0364] The device sends information received from the user and instruction manual data to the server. The HTTP protocol is used for transmission, ensuring the data is securely transferred to the server. This provides the server with data for analysis.

[0365] Step 3:

[0366] The server uses optical character recognition (OCR) technology to convert the received instruction manual data into text and extract the assembly instructions. Specifically, it uses the Tesseract library to scan the instruction manual's PDF or image data and obtain the information as text. This output is text data organized step by step.

[0367] Step 4:

[0368] The server uses an image analysis algorithm to analyze the image information in the instruction manual and identify parts. Using libraries such as OpenCV, it recognizes objects in the image and identifies parts based on their characteristics such as shape and color. As a result, data for each part is output.

[0369] Step 5:

[0370] The server uses the extracted procedure and parts data to select the appropriate video clip from a trained video database. The video database stores past assembly videos categorized by assembly procedure, and the server automatically retrieves the most suitable clip according to the required procedure.

[0371] Step 6:

[0372] The server uses a generation AI model to automatically generate assembly videos for a specific product based on selected video clips. The generation AI model is input with prompts, and edits and combines the video accordingly. For example, the prompt "Show the steps to assemble part A" is used, and the edited video is generated as output.

[0373] Step 7:

[0374] The device uses an onboard emotion engine to recognize the user's emotional state in real time. It employs an algorithm that analyzes the user's voice tone and facial expressions using the camera and microphone. This process yields user emotional data as output.

[0375] Step 8:

[0376] The server dynamically adjusts the content and pace of the generated video based on the emotional data received from the device. If the user expresses anxiety or confusion, adjustments are made, such as slowing down video playback and providing more detailed explanations. The adjusted video is then provided to the device as the final output.

[0377] Step 9:

[0378] Users view videos provided by the server via streaming or download on their devices. The videos are available in multiple languages, allowing users to proceed with assembly while receiving assembly guides in their own language.

[0379] (Application Example 2)

[0380] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0381] In modern manufacturing environments, workers are required to assemble products quickly and accurately when faced with complex work procedures. However, conventional manuals and fixed instructions make it difficult to flexibly adapt to the experience and emotional state of the operators. This can lead to a decrease in on-site efficiency. This invention aims to improve manufacturing efficiency by providing a system that takes the emotional state of the worker into consideration and instantly adapts work instructions.

[0382] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0383] In this invention, the server includes means for receiving explanatory data and item information; means for analyzing the received explanatory data and extracting work procedures; means for learning existing video data and generating action videos based on the analyzed procedures; means for detecting the user's emotional state and dynamically adjusting the content and speed of the action videos; and means for distributing the generated videos to the operator. This enables flexible work support that responds to the emotional state of the worker.

[0384] "Explanatory data" refers to documents or files that contain information necessary for assembling or operating an item.

[0385] "Item information" refers to data that describes the characteristics of a specific item, such as its name, model number, and specifications.

[0386] "Means of receiving" refers to functions or devices for taking in data from external sources.

[0387] "Analysis" is the process of breaking down received data into an easily understandable form and extracting the necessary information.

[0388] A "work procedure" is a set of instructions that outlines the steps required to complete a specific task.

[0389] "Existing video data" refers to video files and clips that were filmed or recorded in the past.

[0390] "Learning" is the process of finding patterns and rules from existing data and using them for new data processing.

[0391] "Means for generating motion videos" refers to devices or functions that create videos to provide visualized instructions and guidance based on analysis results and training data.

[0392] "Detecting the user's emotional state" refers to the process of identifying the user's emotions at that moment based on their facial expressions, voice, and other factors.

[0393] "Means of dynamic adjustment" refers to the ability to flexibly change processes and outputs in response to real-time changing situations and requirements.

[0394] "Means of distribution to operators" refers to technologies and methods for transmitting generated information to specific workers.

[0395] The system for realizing this application is configured as follows: First, the user uses a terminal to input descriptive data and item information about the items to be assembled. The terminal sends this data to a server. The server analyzes the received descriptive data and extracts the work procedure. Natural language processing (NLP) technology is used for the analysis. The server also learns from existing video data and generates action videos based on the extracted work procedure. Here, a generative AI model is used to create videos suitable for specific tasks.

[0396] The server further detects the user's emotional state in real time through the camera and microphone installed on the user's device. This detection uses an emotion recognition engine that analyzes the user's facial expressions and voice data. Based on the detected emotional state, the server dynamically adjusts the video and pace of the work instructions. Hardware used includes, for example, Microsoft HoloLens, and software such as Google Cloud Vision API and Microsoft Azure Face API are used.

[0397] The generated video footage is distributed to operators and supports multiple languages, making it easy to implement in international factories. For example, when introducing a new manufacturing line to a factory, workers wear HoloLens and perform tasks while watching guide videos generated by the system. If the user shows a confused expression, the system slows down the pace of the guide and adds detailed instructions such as, "Let me explain it again slowly."

[0398] Examples of prompts for a generative AI model are as follows:

[0399] Design an AI model that dynamically adjusts the guide video displayed on the HMD based on the user's emotions (e.g., frustration, anxiety). Consider how to make the user perform the task more smoothly and generate appropriate feedback.

[0400] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0401] Step 1:

[0402] The user uses a terminal to input description data and item information and sends it to the system. The input data consists of files and digital information, including specific assembly requirements and specifications for the item. The server then receives the item information.

[0403] Step 2:

[0404] The server analyzes the received explanatory data and extracts the work procedure. This analysis utilizes natural language processing techniques to identify the necessary steps from the input text-based explanatory data. As a result of the analysis, specific work procedures are extracted.

[0405] Step 3:

[0406] The server learns from existing video data and generates action videos based on the analyzed work procedures. Here, a generative AI model is used to create guide videos suitable for the procedures. The input is a conventional video clip, and the output is a new action video.

[0407] Step 4:

[0408] The server detects the user's emotional state in real time through the terminal's camera and microphone. An emotion recognition engine is used for detection, analyzing captured images and audio to determine the user's emotions. The input for this analysis is image and audio data, and the output is the user's emotional state.

[0409] Step 5:

[0410] The server dynamically adjusts the content and speed of the action video according to the detected emotional state. For example, if it detects that the user is confused, it slows down the pace of the guide video and adds detailed narration. This generates video instructions that are appropriate for the user.

[0411] Step 6:

[0412] The server delivers adjusted motion video to the terminal and supports multiple languages, making it usable in international environments. Ultimately, the video is presented to the user, who can use it safely and effectively. The input for delivery is the adjusted motion video, and the output is the video display on the user's terminal.

[0413] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0414] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0415] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0416] [Third Embodiment]

[0417] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0418] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0419] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0420] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0421] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0422] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0423] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0424] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0425] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0426] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0427] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0428] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0429] The system for carrying out this invention uses a program that operates between the user, terminal, and server. The program processing and specific embodiments of this system are described below.

[0430] The user uses a terminal to input information about the product they wish to assemble and its accompanying instructions. The terminal sends this information to the server, which first receives the instruction manual data.

[0431] The server analyzes the text and images in the received instruction manual. Here, natural language processing techniques are used to extract instructions and procedures from the document, and image recognition techniques are used to identify the parts needed for assembly.

[0432] Based on the analysis, the server accesses an existing video database and loads relevant video clips into the learning model. The generative AI model is designed to automatically generate customized assembly videos based on these video clips, according to the user's assembly procedure.

[0433] The generated video is formatted, including multilingual options, and delivered to the user via their device. By watching this assembly video on their device, the user can assemble the product efficiently and accurately.

[0434] For example, if a user wants to assemble furniture, they can upload a specific product number and the associated instructions to their device to begin the process. The server analyzes the instructions to identify each part of the furniture and the assembly steps. Next, it learns from relevant assembly videos and generates an assembly video specifically for that furniture. This video clearly shows the furniture assembly process, allowing the user to follow the necessary steps without getting lost.

[0435] In this way, the present invention assists users in assembling products in a way that is intuitively easy to understand. By performing analysis on a server and generating automated videos using AI, a highly personalized assembly guide can be provided.

[0436] The following describes the processing flow.

[0437] Step 1:

[0438] The user uses a terminal to enter the identification information of the product they want to assemble (e.g., part number or model number) and upload the product's instruction manual. The terminal then sends this data to the server.

[0439] Step 2:

[0440] The server analyzes the instruction manual data received from the terminal. Using natural language processing, it reads assembly procedures and precautions from the manual and extracts them as text data.

[0441] Step 3:

[0442] The server analyzes images within the instruction manual and uses image recognition technology to identify necessary parts and assembly steps. This lays the foundation for generating assembly instructions that take into account not only textual information but also visual information.

[0443] Step 4:

[0444] The server retrieves existing video content from the database and feeds it into a learning model. The generative AI model then uses this to identify appropriate video patterns that meet the user's assembly needs.

[0445] Step 5:

[0446] The generative AI model combines analysis results with pre-trained videos to automatically generate assembly videos for specific products. The videos visually and intuitively demonstrate the assembly process.

[0447] Step 6:

[0448] The server converts the generated video to the optimal format and allows for the addition of subtitle options in multiple languages. This enables users to view the video in their chosen language.

[0449] Step 7:

[0450] The device receives assembly videos streamed from the server, allowing users to stream or download the videos. Users can then refer to the videos to assemble the product correctly.

[0451] (Example 1)

[0452] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0453] In traditional assembly processes, a challenge exists in that the procedures described in product manuals are often complex or insufficient, making it difficult for users to assemble products accurately. Furthermore, the time required to understand the procedures can hinder efficient assembly.

[0454] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0455] In this invention, the server includes a device for receiving instruction manual data and product information, a device for analyzing the received instruction manual and extracting assembly procedures, a device for reading relevant video information from a data store based on the analysis and automatically generating an assembly video, and a device for distributing the generated video to the user's device. This makes it possible for the user to assemble the product intuitively and efficiently.

[0456] "Instruction manual data" refers to documentary information provided to users for assembling a product, including procedures and parts information.

[0457] "Product information" refers to information necessary to identify the product to be assembled, and includes the model number and name.

[0458] "Device" refers to a system of hardware or software configured to perform a specific function.

[0459] "Analysis" refers to the calculations and algorithms used to process data and understand its structure and meaning.

[0460] "Assembly instructions" refer to a series of steps and instructions necessary to correctly assemble a product.

[0461] "Visual information" refers to visual data, including videos and images.

[0462] A "data store" refers to a place where information is stored, such as a database or storage system.

[0463] "Automated generation" refers to the process of creating products using programs or algorithms without human intervention.

[0464] "User equipment" refers to terminals or devices used by users to input or display information.

[0465] This system consists of three elements: user, terminal, and server, which together process information related to product assembly.

[0466] The user uses a terminal to input information about the product they wish to assemble and its accompanying instructions. The terminal can be a tablet or a personal computer. Common file formats include PDF and image files. The terminal formats this information appropriately and sends it to the server.

[0467] The servers are located, for example, on a cloud platform or in a proprietary data center. When processing the received instruction manual data, the servers utilize natural language processing and image recognition technologies. In this case, general API services are used as natural language processing software. For image recognition, various tools are used to identify parts.

[0468] Based on the received data, the server accesses an existing video database and collects the relevant video data. This video data is then reconstructed into a customized assembly video using a generative AI model, according to the user's assembly procedure. In this process, the video is automatically generated using a learned model.

[0469] The generated videos are converted into a multilingual format and delivered to the user via their device. This allows users to assemble products efficiently and accurately. For example, if a user wants to assemble furniture, they can upload the product model number and instructions to their device and start the process with a prompt that says, "Upload the instructions for the product that needs assembly and generate a customized assembly video."

[0470] In this way, by utilizing server-side data analysis and AI-driven automated generation technology, we can provide personalized assembly guides.

[0471] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0472] Step 1:

[0473] The user uses the terminal to input information about the product they want to assemble and its accompanying instruction manual. Specifically, they upload text data such as the product name and model number, as well as PDF or image data of the instruction manual to the terminal. The terminal receives the input data, verifies its format, and then prepares to send it to the server.

[0474] Step 2:

[0475] The terminal sends product information and instruction manual data received from the user to the server. Input data (text data, PDF, images) is sent to the server. The terminal encodes this data using the appropriate protocol and passes it to the server over the network.

[0476] Step 3:

[0477] The server analyzes the received instruction manual data. Specifically, it uses natural language processing technology to extract assembly instructions from the text data of the instruction manual and image recognition technology to identify necessary parts information from image data. Through these analyses, a list of necessary parts and an outline of the assembly procedure are generated for the product.

[0478] Step 4:

[0479] Based on the analysis results, the server accesses an existing video database and collects relevant video clips. The server uses the analyzed procedural information as input to search the video database and find the corresponding video clips. The found clips are then loaded for subsequent generative model processing.

[0480] Step 5:

[0481] The server uses a generative AI model to automatically generate customized assembled videos based on the loaded video clips and analysis results. Here, the generative AI model uses the analyzed procedures to edit the video clips and create videos tailored to the user's specifications. The generated videos may also include narration and text explanations to enhance user viewing experience.

[0482] Step 6:

[0483] The server converts the generated video into a format that can be delivered in the user's selected language and sends it to the device. The server also places multilingual subtitles in the video file and adjusts it to be easily understood by the user. The converted video data is sent to the device and delivered to the user.

[0484] Step 7:

[0485] The device receives a customized video streamed from the server and prepares to play it. Users can intuitively understand the assembly procedure by watching the assembly video on the device's screen. This allows users to assemble the product efficiently.

[0486] (Application Example 1)

[0487] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0488] Traditional assembly procedures are prone to human error and time-inefficient processes when assembling complex items. This hinders productivity improvements and increases the burden on on-site workers. In particular, there is a need for accurate and efficient procedures for assembling machinery and equipment used in factories and manufacturing sites.

[0489] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0490] In this invention, the server includes means for receiving instruction manual information and product data, means for analyzing the received instruction manual and extracting operating procedures, means for machine learning existing video data and generating operating videos based on the analyzed procedures, and means for displaying the generated operating videos on an augmented reality device. This allows field workers to obtain real-time visual assembly guides through the augmented reality device. This enables accurate assembly work and reduces assembly time.

[0491] "Instruction manual information" refers to document data containing instructions and procedures necessary for assembling or using an item.

[0492] "Item data" refers to information about the products or parts that are to be assembled or operated.

[0493] "Means of receiving data" refers to the methods and functions by which a server acquires data from external sources and uses it for analysis.

[0494] "Means of analysis and extraction of operating procedures" refers to the process of analyzing received data and concretizing and clarifying each step.

[0495] "Existing video data" refers to visual media information that has been collected or stored in the past.

[0496] "Machine learning" is a technique that allows computers to recognize patterns in data and build predictive models.

[0497] "Operation video" refers to video data that visually demonstrates the assembly and operation procedures.

[0498] "Means of generation" refers to the processes and technologies used to create new images.

[0499] An "augmented reality device" is an electronic device that uses technology to overlay digital information onto the real world's field of view.

[0500] "Means of display" refers to the method of presenting generated digital content to the user's field of vision.

[0501] In implementing this invention, the server receives instruction manual information and product data via a network. The received information is analyzed using natural language processing techniques such as NLTK and spaCy to extract operating procedures. Furthermore, visual information is analyzed using OpenCV to perform element recognition. This allows the server to clearly identify the necessary steps and parts.

[0502] Next, existing video data is used as part of machine learning, and a generative AI model using a deep learning framework such as PyTorch generates operation videos based on the operation procedures. These generated videos are then formatted using multilingual support technology and provided to the user's augmented reality device. This allows users to intuitively understand the work content and prevents errors.

[0503] As a concrete example, when an assembly worker manufactures a new home appliance, they input the instruction manual information into smart glasses. At this time, the server analyzes the received information and generates an assembly video specifically for that product. The generated video is displayed in real time on the worker's smart glasses, supporting assembly according to the correct procedure. An example of a prompt message used in this case would be, "Read the instruction manual for the home appliance to be assembled, identify the necessary parts, and generate a customized assembly video based on the contents of the manual."

[0504] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0505] Step 1:

[0506] The terminal receives product information and instruction manual data from the user as input. This aggregates specific item information that the user wishes to assemble. The terminal then transmits this information to the server.

[0507] Step 2:

[0508] The server receives product information and instruction manual data sent from the terminal as input. Using natural language processing technologies such as NLTK and spaCy, the instruction manual data is analyzed and operating procedures are extracted. The extracted procedure information is used for video generation in the next step.

[0509] Step 3:

[0510] The server uses image recognition technology based on OpenCV to analyze visual information from instruction manual data. This analysis allows for the identification of parts and the acquisition of layout information. The obtained part information is incorporated into the operating procedure and treated as information necessary for video generation.

[0511] Step 4:

[0512] The server utilizes a generative AI model based on PyTorch to generate video data from an existing video database according to the operating procedure. Taking the operating procedure and parts information as input, it outputs a detailed assembly video. This video is customized to be easily understood by the user.

[0513] Step 5:

[0514] The server converts the generated video feed into a multilingual format suitable for augmented reality devices. This allows the video to be displayed to different users in various languages. The formatted video data is then sent to the terminal or augmented reality device.

[0515] Step 6:

[0516] Users view streamed instructional videos using augmented reality devices (e.g., smart glasses). These devices overlay the video onto the user's field of view, providing real-time assembly support. This allows users to perform assembly tasks efficiently and accurately.

[0517] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0518] This invention is implemented using a program that operates between the user, a terminal, and a server. In particular, it provides more effective assembly assistance by recognizing the user's emotional state and adjusting the assembly guide accordingly.

[0519] First, the user uses a terminal to enter information about the product they want to assemble and uploads the product's instruction manual. The terminal sends this data to the server. The server analyzes the instruction manual data and extracts the assembly steps and important points.

[0520] The server further analyzes image information within the instruction manual to recognize parts. Based on this analysis, it retrieves appropriate clips from a pre-trained video database and uses a generation AI model to automatically generate assembly videos suitable for the specific product.

[0521] This system incorporates an emotion engine that uses the device's camera and sensors to recognize the user's emotional state in real time. The data obtained by the emotion engine is used to dynamically adjust the content and pace of the assembled video. For example, if the emotion engine detects that the user is confused, the video's pace will be slowed down and adjusted to provide more detailed explanations.

[0522] The assembly videos generated in this way have multilingual options and are delivered from the server to the terminal. Users perform the assembly work while watching the videos via streaming or download. User emotion data is recorded and used to improve the system's performance.

[0523] For example, if a user wants to assemble new furniture, they input the product information into the system, and a personalized video is provided to guide them through the assembly process. During this process, the system detects the user's emotions from their facial expressions and tone of voice, and if anxiety about assembly is detected, the explanation becomes more detailed, providing a guide tailored to the user.

[0524] In this way, the present invention realizes a flexible assembly guide that is tailored to the user's situation, providing an environment in which users can assemble products more comfortably.

[0525] The following describes the processing flow.

[0526] Step 1:

[0527] The user uses a terminal to input the identification information and instructions for the product they want to assemble into the system. The terminal then sends this data to the server.

[0528] Step 2:

[0529] The server analyzes the instruction manual data received from the terminal. It uses natural language processing technology to extract assembly instructions and precautions in text format, and then uses image recognition technology to identify parts.

[0530] Step 3:

[0531] Based on the analyzed procedure, the server selects relevant video clips from an existing video database and loads them into the generation AI model. This model learns the patterns necessary to generate appropriate assembled videos.

[0532] Step 4:

[0533] The generative AI model automatically generates product assembly videos using analysis results and trained data. The videos include scenarios that visually demonstrate each step of the assembly process.

[0534] Step 5:

[0535] The emotion engine analyzes the user's emotional state in real time through the device's camera and microphone. The engine analyzes the user's facial expressions and voice to identify their current emotional state.

[0536] Step 6:

[0537] The server receives feedback from the emotion engine and dynamically adjusts the content of the generated assembled video. For example, if the user is feeling anxious, the video speed is slowed down and additional explanations are added to improve support.

[0538] Step 7:

[0539] The device delivers a pre-configured assembly video to the user. The user can play the video and assemble the product according to the instructions. User sentiment data is recorded for future improvements.

[0540] By coordinating each step in this way, we can provide an assembly guide that resonates with the user's emotions, resulting in increased work efficiency and a better user experience.

[0541] (Example 2)

[0542] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0543] Conventional assembly support systems were unable to take into account the confusion and anxiety users experienced during the assembly process in real time, making it difficult to provide optimal assembly assistance. Furthermore, the generated assembly instructions were dependent on specific languages, resulting in a lack of multilingual support.

[0544] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0545] In this invention, the server includes means for acquiring instruction manual information and product data, means for analyzing the acquired instruction manual and extracting assembly procedures, and means for recognizing the user's emotional state and dynamically adjusting the content and speed of the assembly video. This makes it possible to provide personalized assembly support while taking into account the user's emotions. Furthermore, flexible support through multilingual capabilities allows for the provision of appropriate assembly guides to users who speak different languages.

[0546] "Instruction manual information" refers to document data containing a series of instructions and procedures necessary for the user to assemble the product.

[0547] "Item data" refers to information related to the goods or products to be assembled, and includes product name, model, specifications, etc.

[0548] "Part identification" is the process of analyzing instruction manuals and image data to identify the individual parts and components needed for assembly.

[0549] "Video information" refers to video data used to visually demonstrate assembly procedures and related work processes.

[0550] "User emotional state" refers to the type and intensity of emotions felt by the user performing the assembly, and is recognized through analysis of facial expressions and voice.

[0551] "Multilingual support" refers to a feature designed to allow users who speak different languages ​​to understand the same information, meaning that information can be provided in multiple languages.

[0552] This invention is implemented by a system that operates between a user, a terminal, and a server. The user inputs information about the product they wish to assemble via the terminal and uploads the corresponding instruction manual file. The terminal used here is a general-purpose computer or smart device equipped with a camera and sensors.

[0553] The terminal receives instructions from the user and sends that information and instruction manual data to the server. The server receives this data and uses optical character recognition (OCR) and image analysis technologies to extract assembly procedures, precautions, and parts information from the instruction manual.

[0554] For analyzing instruction manuals, we use Tesseract, a Python-based OCR library, and for image analysis, we use OpenCV. The server utilizes the extracted data to select appropriate video clips from a pre-trained video database. Then, a generative AI model is used to automatically generate product-specific assembly videos. The AI ​​model, for example, is a Transformer-based model, and it creates videos based on prompts such as "Show the assembly procedure for part A of this product."

[0555] Furthermore, the emotion engine built into the device analyzes the user's facial expressions and voice tone in real time to recognize the emotional state the user is experiencing. This data is used on the server to dynamically adjust the content and speed of the assembled video. The video is provided in multiple languages ​​according to the user's needs and is delivered from the server to the device via streaming or download.

[0556] For example, if the emotion engine detects that a user is feeling anxious about assembling new furniture, the server will adjust the video instructions to be more detailed and slower. Through this process, the user receives an emotionally sensitive and personalized guide, enabling them to proceed with the assembly process more comfortably.

[0557] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0558] Step 1:

[0559] The user uses a terminal to input information and instructions for the product they want to assemble, and then uploads them. This sends the terminal the basic information necessary for the user to assemble the product. The input information includes the product name, model number, and a PDF file of the instruction manual. The terminal receives this input and prepares to send it to the server.

[0560] Step 2:

[0561] The device sends information received from the user and instruction manual data to the server. The HTTP protocol is used for transmission, ensuring the data is securely transferred to the server. This provides the server with data for analysis.

[0562] Step 3:

[0563] The server uses optical character recognition (OCR) technology to convert the received instruction manual data into text and extract the assembly instructions. Specifically, it uses the Tesseract library to scan the instruction manual's PDF or image data and obtain the information as text. This output is text data organized step by step.

[0564] Step 4:

[0565] The server uses an image analysis algorithm to analyze the image information in the instruction manual and identify parts. Using libraries such as OpenCV, it recognizes objects in the image and identifies parts based on their characteristics such as shape and color. As a result, data for each part is output.

[0566] Step 5:

[0567] The server uses the extracted procedure and parts data to select the appropriate video clip from a trained video database. The video database stores past assembly videos categorized by assembly procedure, and the server automatically retrieves the most suitable clip according to the required procedure.

[0568] Step 6:

[0569] The server uses a generation AI model to automatically generate assembly videos for a specific product based on selected video clips. The generation AI model is input with prompts, and edits and combines the video accordingly. For example, the prompt "Show the steps to assemble part A" is used, and the edited video is generated as output.

[0570] Step 7:

[0571] The device uses an onboard emotion engine to recognize the user's emotional state in real time. It employs an algorithm that analyzes the user's voice tone and facial expressions using the camera and microphone. This process yields user emotional data as output.

[0572] Step 8:

[0573] The server dynamically adjusts the content and pace of the generated video based on the emotional data received from the device. If the user expresses anxiety or confusion, adjustments are made, such as slowing down video playback and providing more detailed explanations. The adjusted video is then provided to the device as the final output.

[0574] Step 9:

[0575] Users view videos provided by the server via streaming or download on their devices. The videos are available in multiple languages, allowing users to proceed with assembly while receiving assembly guides in their own language.

[0576] (Application Example 2)

[0577] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0578] In modern manufacturing environments, workers are required to assemble products quickly and accurately when faced with complex work procedures. However, conventional manuals and fixed instructions make it difficult to flexibly adapt to the experience and emotional state of the operators. This can lead to a decrease in on-site efficiency. This invention aims to improve manufacturing efficiency by providing a system that takes the emotional state of the worker into consideration and instantly adapts work instructions.

[0579] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0580] In this invention, the server includes means for receiving explanatory data and item information; means for analyzing the received explanatory data and extracting work procedures; means for learning existing video data and generating action videos based on the analyzed procedures; means for detecting the user's emotional state and dynamically adjusting the content and speed of the action videos; and means for distributing the generated videos to the operator. This enables flexible work support that responds to the emotional state of the worker.

[0581] "Explanatory data" refers to documents or files that contain information necessary for assembling or operating an item.

[0582] "Item information" refers to data that describes the characteristics of a specific item, such as its name, model number, and specifications.

[0583] "Means of receiving" refers to functions or devices for taking in data from external sources.

[0584] "Analysis" is the process of breaking down received data into an easily understandable form and extracting the necessary information.

[0585] A "work procedure" is a set of instructions that outlines the steps required to complete a specific task.

[0586] "Existing video data" refers to video files and clips that were filmed or recorded in the past.

[0587] "Learning" is the process of finding patterns and rules from existing data and using them for new data processing.

[0588] "Means for generating motion videos" refers to devices or functions that create videos to provide visualized instructions and guidance based on analysis results and training data.

[0589] "Detecting the user's emotional state" refers to the process of identifying the user's emotions at that moment based on their facial expressions, voice, and other factors.

[0590] "Means of dynamic adjustment" refers to the ability to flexibly change processes and outputs in response to real-time changing situations and requirements.

[0591] "Means of distribution to operators" refers to technologies and methods for transmitting generated information to specific workers.

[0592] The system for realizing this application is configured as follows: First, the user uses a terminal to input descriptive data and item information about the items to be assembled. The terminal sends this data to a server. The server analyzes the received descriptive data and extracts the work procedure. Natural language processing (NLP) technology is used for the analysis. The server also learns from existing video data and generates action videos based on the extracted work procedure. Here, a generative AI model is used to create videos suitable for specific tasks.

[0593] The server further detects the user's emotional state in real time through the camera and microphone installed on the user's device. This detection uses an emotion recognition engine that analyzes the user's facial expressions and voice data. Based on the detected emotional state, the server dynamically adjusts the video and pace of the work instructions. Hardware used includes, for example, Microsoft HoloLens, and software such as Google Cloud Vision API and Microsoft Azure Face API are used.

[0594] The generated video footage is distributed to operators and supports multiple languages, making it easy to implement in international factories. For example, when introducing a new manufacturing line to a factory, workers wear HoloLens and perform tasks while watching guide videos generated by the system. If the user shows a confused expression, the system slows down the pace of the guide and adds detailed instructions such as, "Let me explain it again slowly."

[0595] Examples of prompts for a generative AI model are as follows:

[0596] Design an AI model that dynamically adjusts the guide video displayed on the HMD based on the user's emotions (e.g., frustration, anxiety). Consider how to make the user perform the task more smoothly and generate appropriate feedback.

[0597] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0598] Step 1:

[0599] The user uses a terminal to input description data and item information and sends it to the system. The input data consists of files and digital information, including specific assembly requirements and specifications for the item. The server then receives the item information.

[0600] Step 2:

[0601] The server analyzes the received explanatory data and extracts the work procedure. This analysis utilizes natural language processing techniques to identify the necessary steps from the input text-based explanatory data. As a result of the analysis, specific work procedures are extracted.

[0602] Step 3:

[0603] The server learns from existing video data and generates action videos based on the analyzed work procedures. Here, a generative AI model is used to create guide videos suitable for the procedures. The input is a conventional video clip, and the output is a new action video.

[0604] Step 4:

[0605] The server detects the user's emotional state in real time through the terminal's camera and microphone. An emotion recognition engine is used for detection, analyzing captured images and audio to determine the user's emotions. The input for this analysis is image and audio data, and the output is the user's emotional state.

[0606] Step 5:

[0607] The server dynamically adjusts the content and speed of the action video according to the detected emotional state. For example, if it detects that the user is confused, it slows down the pace of the guide video and adds detailed narration. This generates video instructions that are appropriate for the user.

[0608] Step 6:

[0609] The server delivers adjusted motion video to the terminal and supports multiple languages, making it usable in international environments. Ultimately, the video is presented to the user, who can use it safely and effectively. The input for delivery is the adjusted motion video, and the output is the video display on the user's terminal.

[0610] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0611] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0612] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0613] [Fourth Embodiment]

[0614] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0615] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0616] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0617] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0618] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0619] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0620] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0621] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0622] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0623] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0624] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0625] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0626] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0627] The system for carrying out this invention uses a program that operates between the user, terminal, and server. The program processing and specific embodiments of this system are described below.

[0628] The user uses a terminal to input information about the product they wish to assemble and its accompanying instructions. The terminal sends this information to the server, which first receives the instruction manual data.

[0629] The server analyzes the text and images in the received instruction manual. Here, natural language processing techniques are used to extract instructions and procedures from the document, and image recognition techniques are used to identify the parts needed for assembly.

[0630] Based on the analysis, the server accesses an existing video database and loads relevant video clips into the learning model. The generative AI model is designed to automatically generate customized assembly videos based on these video clips, according to the user's assembly procedure.

[0631] The generated video is formatted, including multilingual options, and delivered to the user via their device. By watching this assembly video on their device, the user can assemble the product efficiently and accurately.

[0632] For example, if a user wants to assemble furniture, they can upload a specific product number and the associated instructions to their device to begin the process. The server analyzes the instructions to identify each part of the furniture and the assembly steps. Next, it learns from relevant assembly videos and generates an assembly video specifically for that furniture. This video clearly shows the furniture assembly process, allowing the user to follow the necessary steps without getting lost.

[0633] In this way, the present invention assists users in assembling products in a way that is intuitively easy to understand. By performing analysis on a server and generating automated videos using AI, a highly personalized assembly guide can be provided.

[0634] The following describes the processing flow.

[0635] Step 1:

[0636] The user uses a terminal to enter the identification information of the product they want to assemble (e.g., part number or model number) and upload the product's instruction manual. The terminal then sends this data to the server.

[0637] Step 2:

[0638] The server analyzes the instruction manual data received from the terminal. Using natural language processing, it reads assembly procedures and precautions from the manual and extracts them as text data.

[0639] Step 3:

[0640] The server analyzes images within the instruction manual and uses image recognition technology to identify necessary parts and assembly steps. This lays the foundation for generating assembly instructions that take into account not only textual information but also visual information.

[0641] Step 4:

[0642] The server retrieves existing video content from the database and feeds it into a learning model. The generative AI model then uses this to identify appropriate video patterns that meet the user's assembly needs.

[0643] Step 5:

[0644] The generative AI model combines analysis results with pre-trained videos to automatically generate assembly videos for specific products. The videos visually and intuitively demonstrate the assembly process.

[0645] Step 6:

[0646] The server converts the generated video to the optimal format and allows for the addition of subtitle options in multiple languages. This enables users to view the video in their chosen language.

[0647] Step 7:

[0648] The device receives assembly videos streamed from the server, allowing users to stream or download the videos. Users can then refer to the videos to assemble the product correctly.

[0649] (Example 1)

[0650] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0651] In traditional assembly processes, a challenge exists in that the procedures described in product manuals are often complex or insufficient, making it difficult for users to assemble products accurately. Furthermore, the time required to understand the procedures can hinder efficient assembly.

[0652] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0653] In this invention, the server includes a device for receiving instruction manual data and product information, a device for analyzing the received instruction manual and extracting assembly procedures, a device for reading relevant video information from a data store based on the analysis and automatically generating an assembly video, and a device for distributing the generated video to the user's device. This makes it possible for the user to assemble the product intuitively and efficiently.

[0654] "Instruction manual data" refers to documentary information provided to users for assembling a product, including procedures and parts information.

[0655] "Product information" refers to information necessary to identify the product to be assembled, and includes the model number and name.

[0656] "Device" refers to a system of hardware or software configured to perform a specific function.

[0657] "Analysis" refers to the calculations and algorithms used to process data and understand its structure and meaning.

[0658] "Assembly instructions" refer to a series of steps and instructions necessary to correctly assemble a product.

[0659] "Visual information" refers to visual data, including videos and images.

[0660] A "data store" refers to a place where information is stored, such as a database or storage system.

[0661] "Automated generation" refers to the process of creating products using programs or algorithms without human intervention.

[0662] "User equipment" refers to terminals or devices used by users to input or display information.

[0663] This system consists of three elements: user, terminal, and server, which together process information related to product assembly.

[0664] The user uses a terminal to input information about the product they wish to assemble and its accompanying instructions. The terminal can be a tablet or a personal computer. Common file formats include PDF and image files. The terminal formats this information appropriately and sends it to the server.

[0665] The servers are located, for example, on a cloud platform or in a proprietary data center. When processing the received instruction manual data, the servers utilize natural language processing and image recognition technologies. In this case, general API services are used as natural language processing software. For image recognition, various tools are used to identify parts.

[0666] Based on the received data, the server accesses an existing video database and collects the relevant video data. This video data is then reconstructed into a customized assembly video using a generative AI model, according to the user's assembly procedure. In this process, the video is automatically generated using a learned model.

[0667] The generated videos are converted into a multilingual format and delivered to the user via their device. This allows users to assemble products efficiently and accurately. For example, if a user wants to assemble furniture, they can upload the product model number and instructions to their device and start the process with a prompt that says, "Upload the instructions for the product that needs assembly and generate a customized assembly video."

[0668] In this way, by utilizing server-side data analysis and AI-driven automated generation technology, we can provide personalized assembly guides.

[0669] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0670] Step 1:

[0671] The user uses the terminal to input information about the product they want to assemble and its accompanying instruction manual. Specifically, they upload text data such as the product name and model number, as well as PDF or image data of the instruction manual to the terminal. The terminal receives the input data, verifies its format, and then prepares to send it to the server.

[0672] Step 2:

[0673] The terminal sends product information and instruction manual data received from the user to the server. Input data (text data, PDF, images) is sent to the server. The terminal encodes this data using the appropriate protocol and passes it to the server over the network.

[0674] Step 3:

[0675] The server analyzes the received instruction manual data. Specifically, it uses natural language processing technology to extract assembly instructions from the text data of the instruction manual and image recognition technology to identify necessary parts information from image data. Through these analyses, a list of necessary parts and an outline of the assembly procedure are generated for the product.

[0676] Step 4:

[0677] Based on the analysis results, the server accesses an existing video database and collects relevant video clips. The server uses the analyzed procedural information as input to search the video database and find the corresponding video clips. The found clips are then loaded for subsequent generative model processing.

[0678] Step 5:

[0679] The server uses a generative AI model to automatically generate customized assembled videos based on the loaded video clips and analysis results. Here, the generative AI model uses the analyzed procedures to edit the video clips and create videos tailored to the user's specifications. The generated videos may also include narration and text explanations to enhance user viewing experience.

[0680] Step 6:

[0681] The server converts the generated video into a format that can be delivered in the user's selected language and sends it to the device. The server also places multilingual subtitles in the video file and adjusts it to be easily understood by the user. The converted video data is sent to the device and delivered to the user.

[0682] Step 7:

[0683] The device receives a customized video streamed from the server and prepares to play it. Users can intuitively understand the assembly procedure by watching the assembly video on the device's screen. This allows users to assemble the product efficiently.

[0684] (Application Example 1)

[0685] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0686] Traditional assembly procedures are prone to human error and time-inefficient processes when assembling complex items. This hinders productivity improvements and increases the burden on on-site workers. In particular, there is a need for accurate and efficient procedures for assembling machinery and equipment used in factories and manufacturing sites.

[0687] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0688] In this invention, the server includes means for receiving instruction manual information and product data, means for analyzing the received instruction manual and extracting operating procedures, means for machine learning existing video data and generating operating videos based on the analyzed procedures, and means for displaying the generated operating videos on an augmented reality device. This allows field workers to obtain real-time visual assembly guides through the augmented reality device. This enables accurate assembly work and reduces assembly time.

[0689] "Instruction manual information" refers to document data containing instructions and procedures necessary for assembling or using an item.

[0690] "Item data" refers to information about the products or parts that are to be assembled or operated.

[0691] "Means of receiving data" refers to the methods and functions by which a server acquires data from external sources and uses it for analysis.

[0692] "Means of analysis and extraction of operating procedures" refers to the process of analyzing received data and concretizing and clarifying each step.

[0693] "Existing video data" refers to visual media information that has been collected or stored in the past.

[0694] "Machine learning" is a technique that allows computers to recognize patterns in data and build predictive models.

[0695] "Operation video" refers to video data that visually demonstrates the assembly and operation procedures.

[0696] "Means of generation" refers to the processes and technologies used to create new images.

[0697] An "augmented reality device" is an electronic device that uses technology to overlay digital information onto the real world's field of view.

[0698] "Means of display" refers to the method of presenting generated digital content to the user's field of vision.

[0699] In implementing this invention, the server receives instruction manual information and product data via a network. The received information is analyzed using natural language processing techniques such as NLTK and spaCy to extract operating procedures. Furthermore, visual information is analyzed using OpenCV to perform element recognition. This allows the server to clearly identify the necessary steps and parts.

[0700] Next, existing video data is used as part of machine learning, and a generative AI model using a deep learning framework such as PyTorch generates operation videos based on the operation procedures. These generated videos are then formatted using multilingual support technology and provided to the user's augmented reality device. This allows users to intuitively understand the work content and prevents errors.

[0701] As a concrete example, when an assembly worker manufactures a new home appliance, they input the instruction manual information into smart glasses. At this time, the server analyzes the received information and generates an assembly video specifically for that product. The generated video is displayed in real time on the worker's smart glasses, supporting assembly according to the correct procedure. An example of a prompt message used in this case would be, "Read the instruction manual for the home appliance to be assembled, identify the necessary parts, and generate a customized assembly video based on the contents of the manual."

[0702] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0703] Step 1:

[0704] The terminal receives product information and instruction manual data from the user as input. This aggregates specific item information that the user wishes to assemble. The terminal then transmits this information to the server.

[0705] Step 2:

[0706] The server receives product information and instruction manual data sent from the terminal as input. Using natural language processing technologies such as NLTK and spaCy, the instruction manual data is analyzed and operating procedures are extracted. The extracted procedure information is used for video generation in the next step.

[0707] Step 3:

[0708] The server uses image recognition technology based on OpenCV to analyze visual information from instruction manual data. This analysis allows for the identification of parts and the acquisition of layout information. The obtained part information is incorporated into the operating procedure and treated as information necessary for video generation.

[0709] Step 4:

[0710] The server utilizes a generative AI model based on PyTorch to generate video data from an existing video database according to the operating procedure. Taking the operating procedure and parts information as input, it outputs a detailed assembly video. This video is customized to be easily understood by the user.

[0711] Step 5:

[0712] The server converts the generated video feed into a multilingual format suitable for augmented reality devices. This allows the video to be displayed to different users in various languages. The formatted video data is then sent to the terminal or augmented reality device.

[0713] Step 6:

[0714] Users view streamed instructional videos using augmented reality devices (e.g., smart glasses). These devices overlay the video onto the user's field of view, providing real-time assembly support. This allows users to perform assembly tasks efficiently and accurately.

[0715] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0716] This invention is implemented using a program that operates between the user, a terminal, and a server. In particular, it provides more effective assembly assistance by recognizing the user's emotional state and adjusting the assembly guide accordingly.

[0717] First, the user uses a terminal to enter information about the product they want to assemble and uploads the product's instruction manual. The terminal sends this data to the server. The server analyzes the instruction manual data and extracts the assembly steps and important points.

[0718] The server further analyzes image information within the instruction manual to recognize parts. Based on this analysis, it retrieves appropriate clips from a pre-trained video database and uses a generation AI model to automatically generate assembly videos suitable for the specific product.

[0719] This system incorporates an emotion engine that uses the device's camera and sensors to recognize the user's emotional state in real time. The data obtained by the emotion engine is used to dynamically adjust the content and pace of the assembled video. For example, if the emotion engine detects that the user is confused, the video's pace will be slowed down and adjusted to provide more detailed explanations.

[0720] The assembly videos generated in this way have multilingual options and are delivered from the server to the terminal. Users perform the assembly work while watching the videos via streaming or download. User emotion data is recorded and used to improve the system's performance.

[0721] For example, if a user wants to assemble new furniture, they input the product information into the system, and a personalized video is provided to guide them through the assembly process. During this process, the system detects the user's emotions from their facial expressions and tone of voice, and if anxiety about assembly is detected, the explanation becomes more detailed, providing a guide tailored to the user.

[0722] In this way, the present invention realizes a flexible assembly guide that is tailored to the user's situation, providing an environment in which users can assemble products more comfortably.

[0723] The following describes the processing flow.

[0724] Step 1:

[0725] The user uses a terminal to input the identification information and instructions for the product they want to assemble into the system. The terminal then sends this data to the server.

[0726] Step 2:

[0727] The server analyzes the instruction manual data received from the terminal. It uses natural language processing technology to extract assembly instructions and precautions in text format, and then uses image recognition technology to identify parts.

[0728] Step 3:

[0729] Based on the analyzed procedure, the server selects relevant video clips from an existing video database and loads them into the generation AI model. This model learns the patterns necessary to generate appropriate assembled videos.

[0730] Step 4:

[0731] The generative AI model automatically generates product assembly videos using analysis results and trained data. The videos include scenarios that visually demonstrate each step of the assembly process.

[0732] Step 5:

[0733] The emotion engine analyzes the user's emotional state in real time through the device's camera and microphone. The engine analyzes the user's facial expressions and voice to identify their current emotional state.

[0734] Step 6:

[0735] The server receives feedback from the emotion engine and dynamically adjusts the content of the generated assembled video. For example, if the user is feeling anxious, the video speed is slowed down and additional explanations are added to improve support.

[0736] Step 7:

[0737] The device delivers a pre-configured assembly video to the user. The user can play the video and assemble the product according to the instructions. User sentiment data is recorded for future improvements.

[0738] By coordinating each step in this way, we can provide an assembly guide that resonates with the user's emotions, resulting in increased work efficiency and a better user experience.

[0739] (Example 2)

[0740] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0741] Conventional assembly support systems were unable to take into account the confusion and anxiety users experienced during the assembly process in real time, making it difficult to provide optimal assembly assistance. Furthermore, the generated assembly instructions were dependent on specific languages, resulting in a lack of multilingual support.

[0742] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0743] In this invention, the server includes means for acquiring instruction manual information and product data, means for analyzing the acquired instruction manual and extracting assembly procedures, and means for recognizing the user's emotional state and dynamically adjusting the content and speed of the assembly video. This makes it possible to provide personalized assembly support while taking into account the user's emotions. Furthermore, flexible support through multilingual capabilities allows for the provision of appropriate assembly guides to users who speak different languages.

[0744] "Instruction manual information" refers to document data containing a series of instructions and procedures necessary for the user to assemble the product.

[0745] "Item data" refers to information related to the goods or products to be assembled, and includes product name, model, specifications, etc.

[0746] "Part identification" is the process of analyzing instruction manuals and image data to identify the individual parts and components needed for assembly.

[0747] "Video information" refers to video data used to visually demonstrate assembly procedures and related work processes.

[0748] "User emotional state" refers to the type and intensity of emotions felt by the user performing the assembly, and is recognized through analysis of facial expressions and voice.

[0749] "Multilingual support" refers to a feature designed to allow users who speak different languages ​​to understand the same information, meaning that information can be provided in multiple languages.

[0750] This invention is implemented by a system that operates between a user, a terminal, and a server. The user inputs information about the product they wish to assemble via the terminal and uploads the corresponding instruction manual file. The terminal used here is a general-purpose computer or smart device equipped with a camera and sensors.

[0751] The terminal receives instructions from the user and sends that information and instruction manual data to the server. The server receives this data and uses optical character recognition (OCR) and image analysis technologies to extract assembly procedures, precautions, and parts information from the instruction manual.

[0752] For analyzing instruction manuals, we use Tesseract, a Python-based OCR library, and for image analysis, we use OpenCV. The server utilizes the extracted data to select appropriate video clips from a pre-trained video database. Then, a generative AI model is used to automatically generate product-specific assembly videos. The AI ​​model, for example, is a Transformer-based model, and it creates videos based on prompts such as "Show the assembly procedure for part A of this product."

[0753] Furthermore, the emotion engine built into the device analyzes the user's facial expressions and voice tone in real time to recognize the emotional state the user is experiencing. This data is used on the server to dynamically adjust the content and speed of the assembled video. The video is provided in multiple languages ​​according to the user's needs and is delivered from the server to the device via streaming or download.

[0754] For example, if the emotion engine detects that a user is feeling anxious about assembling new furniture, the server will adjust the video instructions to be more detailed and slower. Through this process, the user receives an emotionally sensitive and personalized guide, enabling them to proceed with the assembly process more comfortably.

[0755] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0756] Step 1:

[0757] The user uses a terminal to input information and instructions for the product they want to assemble, and then uploads them. This sends the terminal the basic information necessary for the user to assemble the product. The input information includes the product name, model number, and a PDF file of the instruction manual. The terminal receives this input and prepares to send it to the server.

[0758] Step 2:

[0759] The device sends information received from the user and instruction manual data to the server. The HTTP protocol is used for transmission, ensuring the data is securely transferred to the server. This provides the server with data for analysis.

[0760] Step 3:

[0761] The server uses optical character recognition (OCR) technology to convert the received instruction manual data into text and extract the assembly instructions. Specifically, it uses the Tesseract library to scan the instruction manual's PDF or image data and obtain the information as text. This output is text data organized step by step.

[0762] Step 4:

[0763] The server uses an image analysis algorithm to analyze the image information in the instruction manual and identify parts. Using libraries such as OpenCV, it recognizes objects in the image and identifies parts based on their characteristics such as shape and color. As a result, data for each part is output.

[0764] Step 5:

[0765] The server uses the extracted procedure and parts data to select the appropriate video clip from a trained video database. The video database stores past assembly videos categorized by assembly procedure, and the server automatically retrieves the most suitable clip according to the required procedure.

[0766] Step 6:

[0767] The server uses a generation AI model to automatically generate assembly videos for a specific product based on selected video clips. The generation AI model is input with prompts, and edits and combines the video accordingly. For example, the prompt "Show the steps to assemble part A" is used, and the edited video is generated as output.

[0768] Step 7:

[0769] The device uses an onboard emotion engine to recognize the user's emotional state in real time. It employs an algorithm that analyzes the user's voice tone and facial expressions using the camera and microphone. This process yields user emotional data as output.

[0770] Step 8:

[0771] The server dynamically adjusts the content and pace of the generated video based on the emotional data received from the device. If the user expresses anxiety or confusion, adjustments are made, such as slowing down video playback and providing more detailed explanations. The adjusted video is then provided to the device as the final output.

[0772] Step 9:

[0773] Users view videos provided by the server via streaming or download on their devices. The videos are available in multiple languages, allowing users to proceed with assembly while receiving assembly guides in their own language.

[0774] (Application Example 2)

[0775] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0776] In modern manufacturing environments, workers are required to assemble products quickly and accurately when faced with complex work procedures. However, conventional manuals and fixed instructions make it difficult to flexibly adapt to the experience and emotional state of the operators. This can lead to a decrease in on-site efficiency. This invention aims to improve manufacturing efficiency by providing a system that takes the emotional state of the worker into consideration and instantly adapts work instructions.

[0777] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0778] In this invention, the server includes means for receiving explanatory data and item information; means for analyzing the received explanatory data and extracting work procedures; means for learning existing video data and generating action videos based on the analyzed procedures; means for detecting the user's emotional state and dynamically adjusting the content and speed of the action videos; and means for distributing the generated videos to the operator. This enables flexible work support that responds to the emotional state of the worker.

[0779] "Explanatory data" refers to documents or files that contain information necessary for assembling or operating an item.

[0780] "Item information" refers to data that describes the characteristics of a specific item, such as its name, model number, and specifications.

[0781] "Means of receiving" refers to functions or devices for taking in data from external sources.

[0782] "Analysis" is the process of breaking down received data into an easily understandable form and extracting the necessary information.

[0783] A "work procedure" is a set of instructions that outlines the steps required to complete a specific task.

[0784] "Existing video data" refers to video files and clips that were filmed or recorded in the past.

[0785] "Learning" is the process of finding patterns and rules from existing data and using them for new data processing.

[0786] "Means for generating motion videos" refers to devices or functions that create videos to provide visualized instructions and guidance based on analysis results and training data.

[0787] "Detecting the user's emotional state" refers to the process of identifying the user's emotions at that moment based on their facial expressions, voice, and other factors.

[0788] "Means of dynamic adjustment" refers to the ability to flexibly change processes and outputs in response to real-time changing situations and requirements.

[0789] "Means of distribution to operators" refers to technologies and methods for transmitting generated information to specific workers.

[0790] The system for realizing this application is configured as follows: First, the user uses a terminal to input descriptive data and item information about the items to be assembled. The terminal sends this data to a server. The server analyzes the received descriptive data and extracts the work procedure. Natural language processing (NLP) technology is used for the analysis. The server also learns from existing video data and generates action videos based on the extracted work procedure. Here, a generative AI model is used to create videos suitable for specific tasks.

[0791] The server further detects the user's emotional state in real time through the camera and microphone installed on the user's device. This detection uses an emotion recognition engine that analyzes the user's facial expressions and voice data. Based on the detected emotional state, the server dynamically adjusts the video and pace of the work instructions. Hardware used includes, for example, Microsoft HoloLens, and software such as Google Cloud Vision API and Microsoft Azure Face API are used.

[0792] The generated video footage is distributed to operators and supports multiple languages, making it easy to implement in international factories. For example, when introducing a new manufacturing line to a factory, workers wear HoloLens and perform tasks while watching guide videos generated by the system. If the user shows a confused expression, the system slows down the pace of the guide and adds detailed instructions such as, "Let me explain it again slowly."

[0793] Examples of prompts for a generative AI model are as follows:

[0794] Design an AI model that dynamically adjusts the guide video displayed on the HMD based on the user's emotions (e.g., frustration, anxiety). Consider how to make the user perform the task more smoothly and generate appropriate feedback.

[0795] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0796] Step 1:

[0797] The user uses a terminal to input description data and item information and sends it to the system. The input data consists of files and digital information, including specific assembly requirements and specifications for the item. The server then receives the item information.

[0798] Step 2:

[0799] The server analyzes the received explanatory data and extracts the work procedure. This analysis utilizes natural language processing techniques to identify the necessary steps from the input text-based explanatory data. As a result of the analysis, specific work procedures are extracted.

[0800] Step 3:

[0801] The server learns from existing video data and generates action videos based on the analyzed work procedures. Here, a generative AI model is used to create guide videos suitable for the procedures. The input is a conventional video clip, and the output is a new action video.

[0802] Step 4:

[0803] The server detects the user's emotional state in real time through the terminal's camera and microphone. An emotion recognition engine is used for detection, analyzing captured images and audio to determine the user's emotions. The input for this analysis is image and audio data, and the output is the user's emotional state.

[0804] Step 5:

[0805] The server dynamically adjusts the content and speed of the action video according to the detected emotional state. For example, if it detects that the user is confused, it slows down the pace of the guide video and adds detailed narration. This generates video instructions that are appropriate for the user.

[0806] Step 6:

[0807] The server delivers adjusted motion video to the terminal and supports multiple languages, making it usable in international environments. Ultimately, the video is presented to the user, who can use it safely and effectively. The input for delivery is the adjusted motion video, and the output is the video display on the user's terminal.

[0808] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0809] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0810] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0811] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0812] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0813] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0814] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0815] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0816] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0817] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0818] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0819] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0820] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0821] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0822] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0823] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0824] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0825] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0826] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0827] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0828] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0829] The following is further disclosed regarding the embodiments described above.

[0830] (Claim 1)

[0831] A means for receiving instruction manual data and product information,

[0832] A means for analyzing the received instruction manual and extracting the assembly procedure,

[0833] A means for generating assembled videos based on procedures learned from and analyzed from existing video data,

[0834] A means of distributing the generated video to the user,

[0835] A system that includes this.

[0836] (Claim 2)

[0837] The system according to claim 1, which analyzes image information from the received instruction manual and performs part recognition.

[0838] (Claim 3)

[0839] The system according to claim 1, which distributes the generated assembly video in multiple languages.

[0840] "Example 1"

[0841] (Claim 1)

[0842] A device that receives instruction manual data and product information,

[0843] A device that analyzes the received instruction manual and extracts the assembly procedure,

[0844] Based on the analysis, the device reads relevant video information from a data store and automatically generates an assembled video.

[0845] A device that distributes the generated video to the user's device,

[0846] A system that includes this.

[0847] (Claim 2)

[0848] The system according to claim 1, which analyzes image information from the received instruction manual and performs part recognition.

[0849] (Claim 3)

[0850] The system according to claim 1, which distributes the generated assembly video in multiple languages.

[0851] "Application Example 1"

[0852] (Claim 1)

[0853] Means for receiving instruction manual information and product data,

[0854] A means for analyzing the received instruction manual and extracting the operating procedures,

[0855] A means of generating operation videos based on machine learning applied to existing video data and the analyzed procedures,

[0856] A means of providing the generated video to the user,

[0857] A means for displaying generated operation images on an augmented reality device,

[0858] A system that includes this.

[0859] (Claim 2)

[0860] The system according to claim 1, which analyzes visual information from the received instruction manual and performs element recognition.

[0861] (Claim 3)

[0862] The system according to claim 1, which provides the generated operation video in a multilingual format.

[0863] "Example 2 of combining an emotion engine"

[0864] (Claim 1)

[0865] Means for acquiring instruction manual information and product data,

[0866] A means of analyzing the acquired instruction manual and extracting the assembly procedure,

[0867] A means for identifying parts based on analyzed data and generating an assembly video using existing video information,

[0868] A means of recognizing the user's emotional state and dynamically adjusting the content and speed of the assembly video,

[0869] A means of sending the generated video to the user,

[0870] A system that includes this.

[0871] (Claim 2)

[0872] The system according to claim 1, which analyzes image data from the acquired instruction manual and performs part identification.

[0873] (Claim 3)

[0874] The system according to claim 1, which transmits the generated assembly video in a multilingual format.

[0875] "Application example 2 when combining with an emotional engine"

[0876] (Claim 1)

[0877] Means for receiving descriptive data and item information,

[0878] A means for analyzing the received explanatory data and extracting the work procedure,

[0879] A means for generating motion video based on procedures learned from and analyzed from existing video data,

[0880] A means for detecting the user's emotional state and dynamically adjusting the content and speed of the video footage,

[0881] A means of distributing the generated video to the operator,

[0882] A system that includes this.

[0883] (Claim 2)

[0884] The system according to claim 1, which analyzes visual information from the received explanatory data and performs component recognition.

[0885] (Claim 3)

[0886] The system according to claim 1, which distributes the generated motion video in a multilingual format. [Explanation of Symbols]

[0887] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for receiving instruction manual data and product information, A means for analyzing the received instruction manual and extracting the assembly procedure, A means for generating assembled videos based on procedures learned from and analyzed from existing video data, A means of distributing the generated video to the user, A system that includes this.

2. The system according to claim 1, which analyzes image information from the received instruction manual and performs part recognition.

3. The system according to claim 1, which distributes the generated assembly video in multiple languages.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A