System
A system records and trains AI models with craftsmen's procedures, providing virtual reality training and matching, effectively addressing the challenges of skill transmission in traditional crafts.
Patent Information
- Application Number
- JP2024128458
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-02
- Publication Date
- 2026-02-16
AI Technical Summary
The transmission of traditional crafts and skilled techniques is hindered by the aging of artisans, lack of successors, difficulty in recording and demonstrating skills, and the inefficiency of current training systems lacking real-time feedback and simulated learning environments.
A system that records craftsmen's work procedures with video and photographs, trains a generative AI model using analyzed data, provides a simulated training environment, and matches craftsmen with potential successors, utilizing virtual reality for effective skill transfer.
Enables efficient recording and transfer of artisan skills to the next generation with real-time feedback and effective matching, promoting the continuation of traditional crafts.
Smart Images

Figure 2026025649000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] The transmission of skills is a major issue in traditional crafts and skilled techniques. Specifically, these include the aging of skilled artisans, a lack of successors, and the difficulty of recording and transmitting skills and knowledge. Furthermore, it is difficult for artisans to demonstrate their skills while explaining them in detail, as this requires concentration and consumes a great deal of time and effort. Furthermore, current training systems and matching platforms lack real-time feedback and simulated learning environments, making them inefficient. As a result, many traditional skills are not being passed on to the next generation and are in danger of being lost. [Means for solving the problem]
[0005] The present invention provides a system that includes a means for recording a craftsman's work procedures with video and photographs, a means for receiving and analyzing the recorded data, a means for training a generative AI model using the analyzed data, a means for providing a simulated training environment using the trained generative AI model, and a means for matching craftsmen with potential successors. This system enables the efficient recording of a craftsman's skills and knowledge and their transfer to the next generation. Furthermore, a training environment using a virtual space provides real-time feedback, enabling more effective learning of skills. Furthermore, the system effectively matches craftsmen with potential successors, promoting the transfer of skills.
[0006] "Artisan" refers to a worker or craftsman skilled in a particular skill or technique.
[0007] "Work procedure" refers to a series of steps or processes for performing a particular skill or technique.
[0008] "Video and photography" refers to moving and still image formats for recording visual information.
[0009] "Recording means" refers to devices or systems for capturing video or photographic footage of craftsmen's work procedures.
[0010] "Means for receiving and analyzing" refers to the algorithms and software used to receive and analyze recorded video and photographic data.
[0011] A "generative AI model" refers to an artificial intelligence model built using machine learning technology that has the ability to generate new information and procedures from specific data.
[0012] "Training means" refers to the process or algorithm used to provide data to an AI model and train the model based on that data.
[0013] A "pseudo training environment" refers to a system that virtually recreates a real-world work environment in which users can practice and simulate.
[0014] "Means for providing" refers to the interface or device that makes the simulated training environment accessible to users.
[0015] "Successor candidate" refers to a potential new skill acquirer who intends to inherit skills and techniques from a skilled craftsman.
[0016] "Matching means" refers to algorithms and platforms that connect artisans with potential successors.
[0017] "System" refers to a set of hardware and software that includes the elements described above and functions in an integrated manner. [Brief explanation of the drawings]
[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8]FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0020] First, the terms used in the following description will be explained.
[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0026] [First embodiment]
[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0039] The present invention relates to a system that records the detailed work procedures of craftsmen and allows them to be learned by AI to pass on their skills. This system provides various means for effectively passing on the skills of craftsmen to the next generation. The following describes the system configuration and operation for specifically implementing the present invention.
[0040] Basic system configuration
[0041] 1. Data collection device
[0042] Devices: Includes devices such as smartphones and video cameras that record the craftsman's work procedures with video and photos, allowing the craftsman's detailed movements and techniques to be recorded in detail.
[0043] 2. Data Management Server
[0044] Server: A server for receiving and analyzing recorded video and photo data. The server works with a database to organize and store the recorded data. It also has software for data analysis and preprocesses the uploaded data.
[0045] 3. Generative AI Models
[0046] Server: The preprocessed data is used to train a generative AI model, which then learns the specific work steps of the craftsman and generates new work steps. A machine learning algorithm is used to create a highly accurate model across multiple steps.
[0047] 4. Virtual Reality (VR) Training Environment
[0048] Device: A VR system that uses a trained generative AI model to provide a training environment in a virtual space. The user wears a VR headset and performs training based on the generated steps in the virtual space.
[0049] 5. Successor Matching Platform
[0050] Server: A platform for effectively matching artisans with potential successors. It implements a matching algorithm that matches the skill sets of artisans and the interests of successors, and proposes optimal pairs.
[0051] Program processing overview
[0052] 1. Data collection procedure
[0053] User: The craftsman uses a smartphone or video camera to take videos and photos of the work procedure. Once the recording is complete, the data can be viewed through the application on the device and uploaded to the server.
[0054] 2. Data preprocessing and storage
[0055] Server: Receives uploaded video and photo data, stores them in storage, records metadata in a database, and performs preprocessing for analysis.
[0056] 3. Training the AI model
[0057] Server: Trains the generative AI model using the preprocessed data. It uses machine learning algorithms to learn the craftsman's work procedures and generate a highly accurate model.
[0058] 4. Providing a training environment
[0059] Device: The trained generative AI model is integrated into a VR system, allowing users to practice in a virtual space. The user puts on a VR headset and practices by following the steps generated by the AI model.
[0060] 5. Feedback and Ratings
[0061] Server: Analyzes the user's actions in the VR environment in real time and provides feedback, allowing the user to improve their skills.
[0062] 6. Use of Matching Platforms
[0063] Users: By accessing the platform, craftsmen register their skills, and successors input their interests and goals. The system then suggests optimal matches, and users can start communicating based on those matches.
[0064] Specific examples
[0065] 1. Data collection: Potters use their smartphones to videotape the process of kneading the clay, shaping it, and firing it, and also take photographs of each detailed step.
[0066] 2. Data preprocessing: The server splits the received video data into frames and adds metadata about the work done to each frame.
[0067] 3. Training the AI model: The server trains a generative AI model that learns specific steps in pottery making, enabling it to generate new steps.
[0068] 4. Providing a training environment: Users use a VR headset to practice pottery making in a virtual space, following the steps they learned. AI provides real-time feedback.
[0069] 5. Matching Platform: Young potters are matched with experienced artisans through the platform and begin online sessions.
[0070] In this way, the system provides a comprehensive solution for realizing the effective inheritance of traditional skills.
[0071] The processing flow will be explained below.
[0072] Step 1: Data collection
[0073] User: Uses a smartphone or video camera to capture video and photos of the craftsman's work steps.
[0074] Device: Press the start recording button to capture video and photos of the craftsman's work. The captured data is temporarily saved.
[0075] User: After shooting, the user checks the video and photo data through the application and enters metadata (e.g., work process, date, time, etc.).
[0076] Terminal: Compresses video and photo data and uploads it to a server via the Internet.
[0077] Step 2: Data Preprocessing
[0078] Server: Stores the received video and photo data in storage. Creates a new entry in the database and records metadata related to the uploaded video (such as the craftsman's name, the work performed, and the date and time of the shoot).
[0079] Server: Splits the video data into frames, assigns a timestamp to each frame, extracts important frames for each task, and deletes unnecessary frames.
[0080] Server: Converts image data into tensor format and audio data into spectrograms. Prepares the data for analysis.
[0081] Step 3: Training the AI model
[0082] Server: Inputs the preprocessed dataset into the generative AI model and starts training the model. The training process is done using a supervised learning algorithm.
[0083] Server: Monitors the loss function and accuracy during training every epoch, adjusting hyperparameters as learning progresses and optimizing the model.
[0084] Server: Stores the trained model and makes it available for training and simulation.
[0085] Step 4: Providing a training environment
[0086] Server: Integrates the trained generative AI model into the VR system, generating a virtual space and configuring it so that users can simulate tasks in that space.
[0087] Device: Using a VR headset and controller, the user enters the virtual space and practices the generated steps.
[0088] User: Wear a VR headset and follow the steps guided by the AI in the virtual space to complete the task. Check whether you can proceed as planned.
[0089] Step 5: Real-time feedback and evaluation
[0090] Server: Analyzes the work users do in the virtual environment in real time and determines which areas need improvement.
[0091] Server: Based on the analysis results, the server provides visual and audio feedback to the user, such as advice like "Take more time to shape this part."
[0092] Users: Receive feedback and then rework the work based on that feedback, thereby improving their skills.
[0093] Step 6: Use a matching platform
[0094] Server: Runs a matching algorithm based on the skill sets and interests of the successor candidate and the craftsman, and proposes the optimal combination.
[0095] Device: Notifies users of matches and displays their profile and contact information.
[0096] User: Logs in to the platform and checks the proposed matching results. The craftsman and successor communicate directly and create a specific training plan.
[0097] Through these steps, the system provides a comprehensive solution for effectively passing on artisan skills to the next generation.
[0098] Example 1
[0099] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0100] Many traditional skills today are being lost due to the aging of artisans and a lack of successors. In particular, if there is no established method for effectively passing on artisanal techniques and know-how to the next generation, those skills are at high risk of disappearing. Traditional methods, such as simply recording artisans' work procedures with video or photographs, make it difficult to pass on skills to the next generation. Furthermore, a lack of feedback and a lack of training environments hinder the transfer of skills.
[0101] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0102] In this invention, the server includes means for recording the work procedures of craftsmen using a recording device, means for receiving and saving the recorded data, means for preprocessing the saved data, means for training a generative AI model using the preprocessed data, means for integrating the trained generative AI model into a virtual reality system, means for effectively matching craftsmen with potential successors, means for providing a training environment to users using the virtual reality system, and means for providing real-time feedback to users during training. This makes it possible to effectively pass on craftsman skills to the next generation and provide detailed, practical training using virtual reality.
[0103] A "craftsman" is a professional who has specific skills and techniques and uses them to create products or works.
[0104] A "work procedure" is a set of specific steps or processes for performing a particular task.
[0105] "Recording device" refers to a device for recording work procedures as video or photographs.
[0106] "Data" refers to video, photographs, and other information captured by recording devices.
[0107] "Receiving" refers to the process of transferring and receiving recorded data to a server or other system.
[0108] "Storage" is the process of storing received data in storage so that it can be used later.
[0109] "Preprocessing" refers to the process of performing a series of operations and filtering to transform raw data into a form that is easier to analyze.
[0110] A "generative AI model" is an artificial intelligence model generated based on training data to reproduce and analyze specific work procedures and actions.
[0111] "Training" refers to the process of teaching a generative AI model so that it learns specific task steps.
[0112] A "virtual reality system" is a system that allows users to experience and practice actual work procedures in a virtual space.
[0113] "Matching" refers to the process of connecting craftsmen with potential successors based on certain criteria.
[0114] A "training environment" is an environment in which a user can practice or perform exercises through a virtual reality system.
[0115] "Feedback" refers to the evaluation and improvement advice provided in real time to the user's actions and results during training.
[0116] This invention relates to a system that records the work procedures of craftsmen in detail and has AI learn from them to pass on skills. The mode for carrying out the invention is configured as follows.
[0117] 1. Data collection device
[0118] Devices: Smartphones and video cameras are used to record the craftsman's work procedures. This allows the craftsman's detailed movements and techniques to be recorded in detail using video and photographs. For example, a potter can use the video function on his smartphone to record the process of shaping clay, and the position and movement of his hands can be recorded in photographs.
[0119] 2. Uploading and saving data
[0120] User: After capturing the data, the user can check it through a dedicated smartphone application and upload it to the server. This upload process is performed by pressing a dedicated button within the app.
[0121] Server: Receives uploaded data and stores it in cloud storage. After storage, the data is automatically labeled and organized.
[0122] 3. Data Preprocessing
[0123] Server: Analyzes the stored video and photo data and splits it into frames. Each frame is then tagged with metadata that describes the task, including hand position and movement.
[0124] 4. Training the AI model
[0125] Server: Trains a generative AI model using the preprocessed data. It uses machine learning algorithms to learn specific work procedures performed by artisans and generate highly accurate models. For example, by learning pottery work procedures, an AI model can be created that can generate new work procedures.
[0126] 5. Creating a training environment
[0127] Device: Integrate the trained generative AI model into a virtual reality (VR) system so that when a user puts on a VR headset, they can follow the learned steps in the virtual space and practice.
[0128] 6. Feedback and Ratings
[0129] Server: Analyzes the work data performed by the user in the VR environment in real time and provides feedback. The AI model analyzes the user's movements and suggests accurate movements and areas for improvement.
[0130] 7. Successor matching
[0131] Server: Provides a platform for effectively matching craftsmen with potential successors. Proposes optimal matches based on the craftsman's skill set and the successor's interests.
[0132] Prompt Sentence Examples
[0133] "Create an AI model that learns specific steps from a video of a craftsman kneading clay. Analyze the movements of each step in detail and then provide practical training in a VR training environment."
[0134] In this way, this system realizes the effective inheritance of traditional skills and provides an innovative approach that utilizes AI and virtual reality technology, making it possible to efficiently pass on artisan skills to the next generation.
[0135] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0136] Step 1:
[0137] Data collection
[0138] Input: A craftsman films the work procedure using a smartphone or video camera.
[0139] Specific actions: The user, a craftsman, uses a smartphone or video camera to record video and photos of work procedures, such as shaping and firing pottery. Since it is necessary to capture detailed actions from various angles, multiple videos and images are taken.
[0140] Output: Video and photo data showing the work procedure.
[0141] Step 2:
[0142] Data upload
[0143] Input: Completed video and photo data.
[0144] How it works: The user checks the captured data through a dedicated smartphone application and uploads it to the server. By pressing a dedicated button on the app, the data is sent to the cloud.
[0145] Output: Video and photo data uploaded to the server.
[0146] Step 3:
[0147] Data reception and storage
[0148] Input: Uploaded video and photo data.
[0149] What it does: The server receives the uploads and stores them in cloud storage, automatically labeling and organizing the data at the same time.
[0150] Output: Structured data stored in cloud storage.
[0151] Step 4:
[0152] Data Preprocessing
[0153] Input: Video and photo data stored in cloud storage.
[0154] How it works: The server divides the data into frames and adds metadata (information about the hand's position and movement) to each frame. Image analysis algorithms are used to identify the hand's movement and position.
[0155] Output: A frame-by-frame dataset with metadata.
[0156] Step 5:
[0157] Training an AI model
[0158] Input: The preprocessed dataset.
[0159] How it works: The server uses machine learning algorithms to train a generative AI model, for example, for a specific task (such as shaping pottery) and then replicates the movements and positions of each step.
[0160] Output: A trained generative AI model.
[0161] Step 6:
[0162] Creating a training environment
[0163] Input: A trained generative AI model.
[0164] How it works: The device integrates the trained generative AI model into a virtual reality (VR) system, so that when the user puts on the VR headset, the work steps are accurately reproduced in the virtual space.
[0165] Output: Model integrated into a virtual reality system.
[0166] Step 7:
[0167] VR training
[0168] Input: A generative AI model integrated into a virtual reality system.
[0169] Specific actions: The user puts on a VR headset and performs exercises in the VR space by following the steps generated by the AI model. For example, the user recreates the action of molding clay in a virtual space and practices the steps.
[0170] Output: User's practice data.
[0171] Step 8:
[0172] Feedback and Ratings
[0173] Input: User's practice data.
[0174] Specific movements: The server analyzes the user's practice data in real time and provides feedback. The AI model analyzes the user's movements and provides a detailed evaluation of the exact movements and areas for improvement.
[0175] Output: Real-time feedback provided to the user.
[0176] Step 9:
[0177] Successor matching
[0178] Input: Craftsman skillset and potential successor information.
[0179] How it works: The server proposes optimal matches based on the craftsman's skill set and the potential successor's interests and goals. A matching algorithm is used to connect craftsmen and potential successors.
[0180] Output: Matching results between craftsmen and potential successors.
[0181] (Application example 1)
[0182] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0183] In today's world, where it is difficult to pass on the skills of craftsmen, there is a need for a method to reliably record work procedures and pass them on to the next generation. Furthermore, particularly in factory production environments, transferring the skills of skilled craftsmen to robots would lead to improved efficiency and quality, but no concrete method for doing so has been established. Furthermore, there is an issue of a gap between training in a virtual environment and actual robot operation.
[0184] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0185] In this invention, the server includes means for recording the work procedures of craftsmen, means for receiving and analyzing the recorded data, means for training a generative AI model using the analyzed data, means for providing a pseudo-training environment using the trained generative AI model, means for matching craftsmen with potential successors, and means for applying the trained generative AI model to a robot and having it execute the work procedures. This allows for the efficient transfer of craftsman skills and enables factory robots to perform advanced tasks.
[0186] A "craftsman" is someone who has specific skills or techniques and uses those skills to create products or crafts.
[0187] "Work procedure" refers to the series of steps and processes that a craftsman follows to create a product using his skilled techniques and skills.
[0188] "Means of recording" refers to devices or methods for recording a craftsman's work procedures as video or photographs using a smartphone, video camera, etc.
[0189] "Means of analysis" refers to the techniques and equipment used to analyze recorded video and photographs and extract important technical and skill elements from them.
[0190] A "generative AI model" is an AI that uses machine learning algorithms to learn from collected data and imitate or improve the work procedures of craftsmen.
[0191] A "simulated training environment" is a place or system that uses virtual reality (VR) and simulation to recreate the work performed by craftsmen, allowing them to train in an environment that is similar to the actual work environment.
[0192] "Matching methods" refer to systems and methods that match the skill sets of craftsmen with the interests and abilities of potential successors, and then select and connect the most suitable craftsmen and successors.
[0193] A "robot" is a mechanical device that can actually carry out work procedures learned using an AI model.
[0194] MODE FOR CARRYING OUT THE INVENTION
[0195] This invention is a system that records and analyzes the work procedures of craftsmen in detail, and then uses a generative AI model to have a robot carry out the work. The system configuration and operation are explained below.
[0196] Data collection methods
[0197] The device uses a smartphone or video camera to record video and photos of the craftsman's work procedures, allowing the craftsman's detailed movements and techniques to be recorded in detail.
[0198] Data analysis
[0199] The server receives and analyzes the recorded video and photo data, including splitting the footage and photos into frames and analyzing the movements of the craftsmen, using image processing techniques and machine learning algorithms.
[0200] Training generative AI models
[0201] The server trains a generative AI model based on the analyzed data, building a model that mimics or improves the craftsman's work procedures. The machine learning algorithms used include deep learning and reinforcement learning.
[0202] Providing a simulated training environment
[0203] The device provides a virtual reality (VR) environment using the trained generative AI model. In this environment, users can practice based on the learned procedures in a virtual space. Users wear a VR headset and train by following the instructions generated by the AI model.
[0204] Application to robots
[0205] The server integrates the trained generative AI model into the robot control system, allowing the robot to actually execute the craftsman's work steps that the AI model has learned. The robot can be an industrial robot arm or an autonomous robot.
[0206] Matching System
[0207] The server provides a platform for matching artisans with potential successors. This platform matches the skill sets of artisans with the interests and abilities of successors to create the optimal pairing. Users access the platform and register their own skills and interests to be matched.
[0208] Examples and prompts
[0209] For example, a skilled worker assembling a product uses a camera to record the steps of installing parts. The server analyzes the captured video, divides the actions into frames, and extracts important parts. The generative AI model uses this data to learn and formulates assembly steps. It then instructs the robot arm to execute the steps. The entire process is instructed to the AI model with the following prompt:
[0210] "Analyze this video frame by frame and generate step-by-step assembly instructions."
[0211] In this way, the present invention provides a series of means for effectively transferring the skills of craftsmen, making it possible to realize advanced work using factory robots.
[0212] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0213] Step 1:
[0214] The terminal uses a smartphone or video camera to record the craftsman's work procedures with video and photos. The input data are the videos and photos of the craftsman's movements and techniques. The output data are the video and photo files stored on the recording device.
[0215] Step 2:
[0216] The server receives the recorded video and photo data, divides it into frames, and performs analysis. The input data are the video and photo files obtained in step 1. The output data are the analyzed frame-by-frame motion data and its metadata. Specifically, image processing techniques are used to extract important motion points from each frame.
[0217] Step 3:
[0218] The server uses the analyzed data to train a generative AI model. The input data is the frame-by-frame motion data and metadata obtained in step 2. The output data is the trained generative AI model. Specifically, it uses machine learning algorithms (e.g., deep learning and reinforcement learning) to learn engineering features from the analyzed data.
[0219] Step 4:
[0220] The device provides a virtual reality (VR) environment using a trained generative AI model. The input data is the trained generative AI model. The output data is a virtual training environment accessible to users. Specifically, users wear a VR headset and practice in a virtual space by following the steps of a craftsman generated by the AI model.
[0221] Step 5:
[0222] The server integrates the trained generative AI model into the robot control system and has the robot execute the work procedure. The input data is the trained generative AI model. The output data is the specific operation instructions to be executed by the robot. Specifically, the robot arm is controlled to perform tasks such as assembly and processing.
[0223] Step 6:
[0224] The server provides a platform for matching craftsmen with potential successors. The input data is the skill set of the craftsman and the interests and abilities of the potential successor. The output data is the pairing results of the optimal craftsman and potential successor. Specifically, users log in to the platform and enter the necessary information, and the best match is suggested.
[0225] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0226] This invention relates to a system that records the detailed work procedures of craftsmen and allows them to be trained by AI to pass on their skills, and also combines this with an emotion engine that recognizes the user's emotions. This system provides various means for effectively passing on craftsmen's skills to the next generation. The system configuration and operation for specifically implementing this invention are described below.
[0227] Basic system configuration
[0228] 1. Data collection device
[0229] Devices: Includes devices such as smartphones and video cameras that record the craftsman's work procedures with video and photos, allowing the craftsman's detailed movements and techniques to be recorded in detail.
[0230] 2. Data Management Server
[0231] Server: A server for receiving and analyzing recorded video and photo data. The server works with a database to organize and store the recorded data. It also has software for data analysis and preprocesses the uploaded data.
[0232] 3. Generative AI Models
[0233] Server: The preprocessed data is used to train a generative AI model, which then learns the specific work steps of the craftsman and generates new work steps. A machine learning algorithm is used to create a highly accurate model across multiple steps.
[0234] 4. Virtual Reality (VR) Training Environment
[0235] Device: A VR system that uses a trained generative AI model to provide a training environment in a virtual space. The user wears a VR headset and performs training based on the generated steps in the virtual space.
[0236] 5. Successor Matching Platform
[0237] Server: A platform for effectively matching artisans with potential successors. It implements a matching algorithm based on the artisan's skill set and the successor's interests, and proposes optimal pairs.
[0238] 6. Emotion Engine
[0239] Server: Equipped with an emotion engine that analyzes the user's facial expressions and voice data to recognize their emotions. Analyzes the user's emotional data in real time and reflects it in training content and feedback.
[0240] Program processing overview
[0241] 1. Data collection procedure
[0242] User: The craftsman uses a smartphone or video camera to take videos and photos of the work procedure. Once the recording is complete, the data can be viewed through the application on the device and uploaded to the server.
[0243] 2. Data preprocessing and storage
[0244] Server: Receives uploaded video and photo data, stores them in storage, records metadata in a database, and performs preprocessing for analysis.
[0245] 3. Training the AI model
[0246] Server: Trains the generative AI model using the preprocessed data. It uses machine learning algorithms to learn the craftsman's work procedures and generate a highly accurate model.
[0247] 4. Providing a training environment
[0248] Device: The trained generative AI model is integrated into a VR system, allowing users to practice in a virtual space. The user puts on a VR headset and practices by following the steps generated by the AI model.
[0249] 5. Feedback and Ratings
[0250] Server: Analyzes the user's actions in the VR environment in real time and provides feedback, allowing the user to improve their skills.
[0251] 6. Emotional Recognition
[0252] Server: Captures the user's facial expressions and voice in real time, and the emotion engine analyzes the data. Based on the user's emotional state, the system adjusts the training progress.
[0253] Device: Receives feedback from the emotion engine. For example, if the user is feeling stressed, the system adjusts the training to reduce stress.
[0254] Users: By receiving appropriate feedback and adjustments based on their emotional state, they can train more efficiently and with greater focus.
[0255] Specific examples
[0256] 1. Data collection: Potters use their smartphones to videotape the process of kneading the clay, shaping it, and firing it, and also take photographs of each detailed step.
[0257] 2. Data preprocessing: The server splits the received video data into frames and adds metadata about the work done to each frame.
[0258] 3. Training the AI model: The server trains a generative AI model that learns specific steps in pottery making, enabling it to generate new steps.
[0259] 4. Providing a training environment: Users use a VR headset to practice pottery making in a virtual space, following the steps they learned. AI provides real-time feedback.
[0260] 5. Emotion recognition and adjustment: If the user is feeling stressed during training, the emotion engine will recognize this and the system will adjust the training content, for example by suggesting a short break to reduce the user's stress.
[0261] 6. Matching Platform: Young potters are matched with experienced craftsmen through the platform and begin online sessions.
[0262] In this way, the system effectively passes on the skills of craftsmen to the next generation and supports skill improvement while taking into consideration the user's emotional state.
[0263] The processing flow will be explained below.
[0264] Step 1: Data collection
[0265] User: Using a smartphone or video camera, the user takes videos and photographs of the craftsman's work steps, allowing the craftsman to record his work in detail.
[0266] Device: Press the start recording button to capture video and photos of the craftsman's work. The captured data is temporarily saved.
[0267] User: After shooting, check the video and photo data through the application and enter metadata (process, date, time, etc.).
[0268] Terminal: Compresses video and photo data and uploads it to a server via the Internet.
[0269] Step 2: Data Preprocessing
[0270] Server: Stores the received video and photo data in storage. Creates a new entry in the database and records metadata related to the uploaded video (such as the craftsman's name, the work performed, and the date and time of the shoot).
[0271] Server: Splits the video data into frames, assigns a timestamp to each frame, extracts important frames for analysis, and removes unnecessary frames.
[0272] Server: Converts the stored image data into tensor format and audio data into spectrograms, thereby adjusting the data format to be analyzable by machine learning algorithms.
[0273] Step 3: Training the AI model
[0274] Server: Inputs the preprocessed dataset into the generative AI model and begins training the model. Using a supervised learning algorithm, it learns the craftsman's work procedures in detail.
[0275] Server: Monitors the loss function and accuracy every epoch during the training process, adjusting hyperparameters as training progresses to optimize model performance.
[0276] Server: Stores the trained generative AI model for later use in training and simulation.
[0277] Step 4: Providing a training environment
[0278] Server: Integrates the trained generative AI model into the VR system, allowing users to simulate tasks in a virtual space.
[0279] Device: The user wears a VR headset and follows the steps generated by the AI model in a virtual space. The user uses the VR controller to recreate the tasks in the virtual environment.
[0280] User: Enter the VR environment and follow the steps guided by the AI model. Practice while checking whether the steps are progressing as expected.
[0281] Step 5: Real-time feedback and evaluation
[0282] Server: Analyzes the work users do in the virtual environment in real time and determines which areas need improvement.
[0283] Server: Based on the analysis results, the server provides visual and audio feedback to the user, such as advice like "Take more time to shape this part."
[0284] Users: Receive feedback and then rework the work based on that feedback, thereby improving their skills.
[0285] Step 6: Recognize emotions
[0286] Server: Captures the user's facial expressions and voice in real time, and the emotion engine analyzes the data to recognize the user's emotional state (stress level and concentration state).
[0287] Device: Receives feedback from the emotion engine. If the user is feeling stressed, the system can adjust the training, for example suggesting a short break.
[0288] Users: By receiving appropriate feedback and adjustments based on their emotional state, they can train more efficiently and with greater focus.
[0289] Step 7: Use a matching platform
[0290] Server: Runs a matching algorithm based on the skill sets and interests of the successor candidate and the craftsman, and proposes the optimal combination.
[0291] Device: Notifies users of matches and displays their profile and contact information.
[0292] User: Logs in to the platform and checks the proposed matching results. The craftsman and successor communicate directly and create a specific training plan.
[0293] Through these steps, the system effectively passes on the craftsman's skills to the next generation and supports skill improvement while taking into account the user's emotional state.
[0294] Example 2
[0295] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0296] In today's aging society, effectively passing on the skills of skilled craftsmen to the next generation is an important issue. However, traditional methods of skill transfer are inefficient, as they make it difficult to fully convey the finer details of the skills and the craftsman's know-how. Furthermore, stress and a decline in motivation for the successor during the skill transfer process can also be an issue.
[0297] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for recording the work procedures of the craftsman with video and photographs, a means for receiving the recorded data, preprocessing and analyzing the data, and a means for training a machine learning model using the analyzed data. This makes it possible to record and analyze the specific work procedures and know-how of the craftsman with high accuracy.
[0298] Furthermore, this invention includes a means for providing a virtual training environment using a trained machine learning model, a means for matching craftsmen with potential successors, and a means for recognizing the user's emotions in real time and adjusting the training content based on those emotions. This increases the efficiency of skill transfer, prevents stress and a decline in motivation for successors, and enables smooth skill improvement.
[0299] A "craftsman" is a professional who has specific skills and techniques and creates things based on those skills.
[0300] "Work procedures" refer to the specific steps and methods of operation that craftsmen use to perform work.
[0301] "Video" refers to video data captured using a camera or other device.
[0302] A "photograph" refers to still image data captured at a specific moment with a camera or other device.
[0303] "Means of recording" refers to means for recording video or still images using devices such as video cameras or smartphones.
[0304] "Means for receiving, pre-processing and analyzing data" refers to means by which the server receives uploaded video and photo data and performs the data processing necessary to analyze them.
[0305] A "machine learning model" refers to an algorithm that uses AI techniques to learn from a specific dataset and make predictions or generation decisions.
[0306] "Means for providing a virtual training environment" refers to means for providing an environment in which a user can recreate and practice the work procedures of a craftsman in a virtual space.
[0307] "Potential successors" refer to people who wish to inherit the skills of a particular craftsman.
[0308] "Matching means" refers to algorithms and platforms that efficiently connect artisans with potential successors.
[0309] "Means for recognizing emotions in real time" refers to means for analyzing the user's facial expressions and voice data to identify their current emotional state.
[0310] "Means for adjusting training content" refers to means for adaptively changing training content in a virtual training environment based on the user's emotional state.
[0311] The present invention is a system that effectively transfers skills by recording the work procedures of craftsmen in detail and building a generative AI model based on the records. This system has the ability to recognize the user's emotions in real time and dynamically adjust the training content. The following describes the system configuration and operation for specifically implementing the present invention.
[0312] Basic system configuration
[0313] 1. Data collection device
[0314] Devices: Recording devices such as smartphones and video cameras are used. These devices record the craftsman's work procedures as videos and photographs, allowing the craftsman's detailed movements and techniques to be recorded in detail.
[0315] 2. Data Management Server
[0316] Server: This is the server that receives and analyzes the recorded video and photo data. The server connects to a database to organize and store the recorded data. It also has software for data analysis and preprocesses the uploaded data. Specifically, it uses tools such as Python and OpenCV to segment the data and add metadata.
[0317] 3. Generative AI Models
[0318] Server: The generative AI model is trained using the preprocessed data. This allows it to learn the specific work procedures of craftsmen and generate new work procedures. A machine learning algorithm is used to create a highly accurate model across multiple steps. TensorFlow and PyTorch are suitable libraries to use.
[0319] 4. Virtual Reality (VR) Training Environment
[0320] Device: A VR system is used that uses a trained generative AI model to provide a training environment in a virtual space. The user wears a VR headset (e.g., Oculus Rift or HTC Vive) and performs training based on the generated steps in the virtual space. The VR environment is built using Unity or Unreal Engine.
[0321] 5. Successor Matching Platform
[0322] Server: Provides a platform for effectively matching artisans with potential successors. Implements a matching algorithm based on the artisan's skill set and the successor's interests, and proposes optimal pairs.
[0323] 6. Emotion Engine
[0324] Server: Equipped with an emotion engine that analyzes the user's facial expressions and voice data to recognize emotions. Analyzes the user's emotional data in real time and reflects it in training content and feedback. Emotion analysis is performed using Microsoft Azure Cognitive Services and the Affectiva SDK.
[0325] Specific examples
[0326] 1. Data collection: Potters use their smartphones to videotape the process of kneading the clay, shaping it, and firing it, and also take photographs of each detailed step.
[0327] 2. Data preprocessing: The server splits the received video data into frames and adds metadata about the work done to each frame.
[0328] 3. Training the AI model: The server trains a generative AI model that learns specific steps in pottery making, enabling it to generate new steps.
[0329] 4. Providing a training environment: Users use a VR headset (e.g., Oculus Rift) to practice pottery making in a virtual space, following the steps they learned. AI provides real-time feedback.
[0330] 5. Emotion recognition and adjustment: If the user is feeling stressed during training, the emotion engine will recognize this and the system will adjust the training content, for example by suggesting a short break to reduce the user's stress.
[0331] 6. Matching Platform: Young potters are matched with experienced craftsmen through the platform and begin online sessions.
[0332] Example prompt
[0333] Example prompt: "I would like to learn the steps of pottery making. I would like to provide videos of each step of the pottery making process, from mixing the clay to molding and firing, and train in a virtual space."
[0334] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0335] Step 1:
[0336] The user uses a smartphone or video camera to record video and photos of the craftsman's work procedures. Specifically, the user records the steps of the potter kneading the clay, shaping it, and firing it. The input data are video and photo files, which are then uploaded to a dedicated application on the device. The output data are verified video and photo files.
[0337] Step 2:
[0338] The device reviews the video and photo data received from the user through the application and edits it as needed, such as cutting out unnecessary parts and labeling important parts. The input data are video and photo files, and the output data are the edited video and photo files.
[0339] Step 3:
[0340] The device uploads the edited video and photo files to the server by sending them to a specified URL on the server using an HTTP request. The input data are the edited video and photo files, and the output data are the video and photo data uploaded to the server.
[0341] Step 4:
[0342] The server receives the uploaded video and photo data and stores them in storage. Specifically, it writes the received files to disk storage and records the metadata in a database. The input data is the uploaded video and photo data, and the output data is the stored video and photo data and their metadata.
[0343] Step 5:
[0344] The server analyzes and preprocesses the stored video and photo data. Specifically, it uses OpenCV to split the video data into frames and add metadata about the work done to each frame. It also performs noise removal and resolution adjustment. The input data is the stored video and photo data, and the output data is the preprocessed video and photo data.
[0345] Step 6:
[0346] The server uses the preprocessed data to train a generative AI model. Specifically, it uses TensorFlow and PyTorch to run machine learning algorithms to learn the craftsman's work procedures. The input data is the preprocessed video and photo data, and the output data is the trained generative AI model.
[0347] Step 7:
[0348] The device integrates the trained generative AI model into the VR system. Specifically, it uses Unity or Unreal Engine to recreate the steps generated by the model in a virtual space so that the user can experience them. The input data is the trained generative AI model, and the output data is the work steps recreated in the VR environment.
[0349] Step 8:
[0350] The user wears a VR headset and performs practical training in a virtual space by following the steps provided by the generative AI model. Specifically, the user practices pottery molding in a virtual space. The input data is the work steps reproduced in the VR environment, and the output data is the motion data of the work performed by the user.
[0351] Step 9:
[0352] The server analyzes the user's behavior data in real time and provides feedback. Specifically, it points out errors in behavior and suggests ways to correct them. It also performs real-time data processing using Apache Kafka and Apache Flink. The input data is the user's behavior data, and the output data is the feedback message.
[0353] Step 10:
[0354] The server captures the user's emotional data in real time and analyzes it using an emotion engine. Specifically, it uses Microsoft Azure Cognitive Services and the Affectiva SDK to recognize emotions from the user's facial expressions and voice. The input data is the user's facial expression and voice data, and the output data is the emotion analysis results.
[0355] Step 11:
[0356] The device dynamically adjusts the training content based on the results obtained from the emotion engine. Specifically, if the user feels stressed, the system will reduce the training content or suggest taking a break. The input data is the emotion analysis results, and the output data is the adjusted training content.
[0357] Step 12:
[0358] The server operates a platform that matches artisans with potential successors. Specifically, it records the skill sets of artisans and the interests of successors in a database and uses a recommendation algorithm to suggest optimal pairs. The input data are the skills and interests of artisans and successors, and the output data are the proposed matching pairs.
[0359] (Application example 2)
[0360] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0361] The transfer of craftsmanship skills requires specific skills and experience, which is time-consuming and costly. Furthermore, conventional methods have difficulty providing feedback that takes into account the craftsman's emotions and state. Furthermore, for a robot to accurately learn and execute these skills, advanced data collection and analysis techniques are required. Therefore, there is a need for a system that can efficiently and effectively transfer craftsmanship skills and enable robots to execute skills and provide feedback that reflects the user's emotions.
[0362] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0363] In this invention, the server includes means for recording the work procedures of craftsmen with video and photographs, means for receiving and analyzing the recorded data, means for training a generative AI model using the analyzed data, means for providing a simulated training environment using the trained generative AI model, means for matching craftsmen with potential successors, means for collecting work procedure data in real time from craftsmen wearing smart glasses, means for controlling a robot using the collected data to execute the learned work procedures, and means for recognizing the user's emotional state using emotion analysis means and providing feedback. This allows the craftsman's work procedures to be recorded in detail, enabling the robot to learn and execute the procedures, and providing feedback based on the user's emotional state in real time.
[0364] "Means for video and photographic recording of craftsmen's work procedures" is a general term for video cameras and camera-equipped devices used to record in detail the series of tasks and operations performed by craftsmen.
[0365] "Means for receiving and analyzing recorded data" refers to a combination of software and hardware for collecting said video and photographic data and analyzing it using appropriate algorithms.
[0366] "Means for training a generative AI model" is a general term for software and computer equipment that applies machine learning algorithms based on the analyzed data to train an AI model.
[0367] "Means for providing a pseudo-training environment using a trained generative AI model" refers to a system that uses a trained AI model to provide a training environment in a simulated format to users.
[0368] The "means for matching craftsmen with potential successors" refers to an algorithm and system that compares the skills of craftsmen with the interests and abilities of potential successors and proposes the optimal combination.
[0369] "Means for collecting work procedure data in real time from craftsmen wearing smart glasses" refers to technology and devices that use smart glasses to capture work procedures from the perspective of craftsmen in real time and collect them as data.
[0370] The "means for controlling the robot using collected data to execute the learned work procedure" is a system for causing the robot to execute the work procedure based on data collected from smart glasses and other sensor devices.
[0371] "Means for recognizing a user's emotional state using emotion analysis means and providing feedback" refers to software that analyzes a user's facial expressions and voice data to recognize their emotional state, and a system for providing feedback to the user based on that.
[0372] The system for implementing this invention records the work procedures of craftsmen in detail, trains a generative AI model based on that data, and applies it to a robot to achieve skill transfer and efficient work execution. It also includes an emotion analysis means for analyzing the user's emotional state in real time and providing appropriate feedback.
[0373] System Configuration
[0374] 1. Data collection device
[0375] Terminal: The craftsman wears smart glasses that capture the work procedure in real time. The smart glasses are equipped with a camera and a microphone to record both visual and audio data, which is then sent to a server via the Internet.
[0376] 2. Data Management Server
[0377] Server: Receives, stores, and analyzes recorded video and audio data. The server works with a database (e.g., MongoDB) to organize and store the data. It also preprocesses the data using machine learning algorithms such as TensorFlow for data analysis.
[0378] 3. Generative AI Models
[0379] Server: The generative AI model is trained using machine learning libraries such as TensorFlow based on the preprocessed data. This model replicates the specific steps of the craftsman's work and applies them to the robot.
[0380] 4. Robot Control
[0381] Terminal: The robot executes the collected work procedures based on the generative AI model. For example, it can use a general-purpose robot arm such as the UR5 or UR10 to perform precise movements.
[0382] 5. Emotion analysis method
[0383] Server: Receives the user's facial expressions and voice data from smart glasses and other devices, and performs emotion analysis using Microsoft Azure's Emotion API, etc. Based on the analysis results, provides feedback to the user in real time.
[0384] Specific examples
[0385] In a factory, a skilled craftsman performs welding work. During this process, the craftsman wears smart glasses that record his work steps in real time. The collected data is sent to a server where it is analyzed. A generative AI model learns the craftsman's work steps and generates new welding steps. This model is then used by a robotic arm to perform the work. Furthermore, an emotion engine analyzes the observer's emotional state, and if stress or doubt is recognized, the system provides appropriate feedback and corrections.
[0386] Prompt Sentence Examples
[0387] "Example of teaching AI new work procedures:
[0388] Detailed procedures for welding work that should be performed by factory robots
[0389] Important points to highlight in the video data
[0390] Expert instructions to be extracted from speech data
[0391] Specific methods for handling facial expression and voice data for emotion analysis
[0392] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0393] Step 1:
[0394] Terminal: The craftsman puts on the smart glasses and begins work. The smart glasses are equipped with a camera and microphone, which record video and audio data of the work in real time.
[0395] Input: Worker's workflow (visual and audio data)
[0396] Output: Recorded video and audio data
[0397] How it works: Smart glasses capture video and audio and collect data in real time.
[0398] Step 2:
[0399] Server: Stores the video and audio data received from the smart glasses in storage, and adds metadata (date, time, operation details, etc.) to the data.
[0400] Input: Video and audio data sent from smart glasses
[0401] Output: Data saved to storage (with metadata)
[0402] Specific operation: Stores data in a database (such as MongoDB) and organizes it by adding metadata.
[0403] Step 3:
[0404] Server: Analyzes the stored video and audio data and extracts the necessary information. Preprocesses the data using machine learning algorithms (e.g., TensorFlow).
[0405] Input: Stored video and audio data (with metadata)
[0406] Output: Preprocessed data (features extracted)
[0407] Specific operations: Split video data into frames and extract key points. Extract important reference phrases from audio data.
[0408] Step 4:
[0409] Server: Trains the generative AI model using the preprocessed data.
[0410] Input: Preprocessed data (features extracted)
[0411] Output: A trained generative AI model
[0412] How it works: Using TensorFlow, we train a model to learn the work procedures of craftsmen, using multi-layer perceptrons and convolutional neural networks (CNNs).
[0413] Step 5:
[0414] Terminal: The trained generative AI model is applied to the robot, which then executes the actual work steps.
[0415] Input: A trained generative AI model
[0416] Output: Robot performs work
[0417] Specific operation: Control a robot arm (e.g., UR5 or UR10) and perform tasks according to learned procedures.
[0418] Step 6:
[0419] Server: Receives user facial and voice data from smart glasses and other sensor devices and performs emotion analysis. Uses Microsoft Azure's Emotion API.
[0420] Input: User facial and voice data
[0421] Output: Emotional state analysis result
[0422] Specific operation: Analyzes user emotion data in real time using Microsoft Azure's Emotion API.
[0423] Step 7:
[0424] Server: Generates feedback based on the results of sentiment analysis and provides it to the user.
[0425] Input: Emotional state analysis result
[0426] Output: Feedback (adjustments to training, new instructions, etc.)
[0427] What it does: Depending on the user's emotional state, the system will adjust the intensity of their workout or suggest a break. Feedback is provided via voice and text.
[0428] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0429] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0430] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0431] [Second embodiment]
[0432] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0433] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0434] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0435] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0436] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0437] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0438] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0439] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0440] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0441] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0442] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0443] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0444] The present invention relates to a system that records the detailed work procedures of craftsmen and allows them to be learned by AI to pass on their skills. This system provides various means for effectively passing on the skills of craftsmen to the next generation. The following describes the system configuration and operation for specifically implementing the present invention.
[0445] Basic system configuration
[0446] 1. Data collection device
[0447] Devices: Includes devices such as smartphones and video cameras that record the craftsman's work procedures with video and photos, allowing the craftsman's detailed movements and techniques to be recorded in detail.
[0448] 2. Data Management Server
[0449] Server: A server for receiving and analyzing recorded video and photo data. The server works with a database to organize and store the recorded data. It also has software for data analysis and preprocesses the uploaded data.
[0450] 3. Generative AI Models
[0451] Server: The preprocessed data is used to train a generative AI model, which then learns the specific work steps of the craftsman and generates new work steps. A machine learning algorithm is used to create a highly accurate model across multiple steps.
[0452] 4. Virtual Reality (VR) Training Environment
[0453] Device: A VR system that uses a trained generative AI model to provide a training environment in a virtual space. The user wears a VR headset and performs training based on the generated steps in the virtual space.
[0454] 5. Successor Matching Platform
[0455] Server: A platform for effectively matching artisans with potential successors. It implements a matching algorithm that matches the skill sets of artisans and the interests of successors, and proposes optimal pairs.
[0456] Program processing overview
[0457] 1. Data collection procedure
[0458] User: The craftsman uses a smartphone or video camera to take videos and photos of the work procedure. Once the recording is complete, the data can be viewed through the application on the device and uploaded to the server.
[0459] 2. Data preprocessing and storage
[0460] Server: Receives uploaded video and photo data, stores them in storage, records metadata in a database, and performs preprocessing for analysis.
[0461] 3. Training the AI model
[0462] Server: Trains the generative AI model using the preprocessed data. It uses machine learning algorithms to learn the craftsman's work procedures and generate a highly accurate model.
[0463] 4. Providing a training environment
[0464] Device: The trained generative AI model is integrated into a VR system, allowing users to practice in a virtual space. The user puts on a VR headset and practices by following the steps generated by the AI model.
[0465] 5. Feedback and Ratings
[0466] Server: Analyzes the user's actions in the VR environment in real time and provides feedback, allowing the user to improve their skills.
[0467] 6. Use of Matching Platforms
[0468] Users: By accessing the platform, craftsmen register their skills, and successors input their interests and goals. The system then suggests optimal matches, and users can start communicating based on those matches.
[0469] Specific examples
[0470] 1. Data collection: Potters use their smartphones to videotape the process of kneading the clay, shaping it, and firing it, and also take photographs of each detailed step.
[0471] 2. Data preprocessing: The server splits the received video data into frames and adds metadata about the work done to each frame.
[0472] 3. Training the AI model: The server trains a generative AI model that learns specific steps in pottery making, enabling it to generate new steps.
[0473] 4. Providing a training environment: Users use a VR headset to practice pottery making in a virtual space, following the steps they learned. AI provides real-time feedback.
[0474] 5. Matching Platform: Young potters are matched with experienced artisans through the platform and begin online sessions.
[0475] In this way, the system provides a comprehensive solution for realizing the effective inheritance of traditional skills.
[0476] The processing flow will be explained below.
[0477] Step 1: Data collection
[0478] User: Uses a smartphone or video camera to capture video and photos of the craftsman's work steps.
[0479] Device: Press the start recording button to capture video and photos of the craftsman's work. The captured data is temporarily saved.
[0480] User: After shooting, the user checks the video and photo data through the application and enters metadata (e.g., work process, date, time, etc.).
[0481] Terminal: Compresses video and photo data and uploads it to a server via the Internet.
[0482] Step 2: Data Preprocessing
[0483] Server: Stores the received video and photo data in storage. Creates a new entry in the database and records metadata related to the uploaded video (such as the craftsman's name, the work performed, and the date and time of the shoot).
[0484] Server: Splits the video data into frames, assigns a timestamp to each frame, extracts important frames for each task, and deletes unnecessary frames.
[0485] Server: Converts image data into tensor format and audio data into spectrograms. Prepares the data for analysis.
[0486] Step 3: Training the AI model
[0487] Server: Inputs the preprocessed dataset into the generative AI model and starts training the model. The training process is done using a supervised learning algorithm.
[0488] Server: Monitors the loss function and accuracy during training every epoch, adjusting hyperparameters as learning progresses and optimizing the model.
[0489] Server: Stores the trained model and makes it available for training and simulation.
[0490] Step 4: Providing a training environment
[0491] Server: Integrates the trained generative AI model into the VR system, generating a virtual space and configuring it so that users can simulate tasks in that space.
[0492] Device: Using a VR headset and controller, the user enters the virtual space and practices the generated steps.
[0493] User: Wear a VR headset and follow the steps guided by the AI in the virtual space to complete the task. Check whether you can proceed as planned.
[0494] Step 5: Real-time feedback and evaluation
[0495] Server: Analyzes the work users do in the virtual environment in real time and determines which areas need improvement.
[0496] Server: Based on the analysis results, the server provides visual and audio feedback to the user, such as advice like "Take more time to shape this part."
[0497] Users: Receive feedback and then rework the work based on that feedback, thereby improving their skills.
[0498] Step 6: Use a matching platform
[0499] Server: Runs a matching algorithm based on the skill sets and interests of the successor candidate and the craftsman, and proposes the optimal combination.
[0500] Device: Notifies users of matches and displays their profile and contact information.
[0501] User: Logs in to the platform and checks the proposed matching results. The craftsman and successor communicate directly and create a specific training plan.
[0502] Through these steps, the system provides a comprehensive solution for effectively passing on artisan skills to the next generation.
[0503] Example 1
[0504] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0505] Many traditional skills today are being lost due to the aging of artisans and a lack of successors. In particular, if there is no established method for effectively passing on artisanal techniques and know-how to the next generation, those skills are at high risk of disappearing. Traditional methods, such as simply recording artisans' work procedures with video or photographs, make it difficult to pass on skills to the next generation. Furthermore, a lack of feedback and a lack of training environments hinder the transfer of skills.
[0506] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0507] In this invention, the server includes means for recording the work procedures of craftsmen using a recording device, means for receiving and saving the recorded data, means for preprocessing the saved data, means for training a generative AI model using the preprocessed data, means for integrating the trained generative AI model into a virtual reality system, means for effectively matching craftsmen with potential successors, means for providing a training environment to users using the virtual reality system, and means for providing real-time feedback to users during training. This makes it possible to effectively pass on craftsman skills to the next generation and provide detailed, practical training using virtual reality.
[0508] A "craftsman" is a professional who has specific skills and techniques and uses them to create products or works.
[0509] A "work procedure" is a set of specific steps or processes for performing a particular task.
[0510] "Recording device" refers to a device for recording work procedures as video or photographs.
[0511] "Data" refers to video, photographs, and other information captured by recording devices.
[0512] "Receiving" refers to the process of transferring and receiving recorded data to a server or other system.
[0513] "Storage" is the process of storing received data in storage so that it can be used later.
[0514] "Preprocessing" refers to the process of performing a series of operations and filtering to transform raw data into a form that is easier to analyze.
[0515] A "generative AI model" is an artificial intelligence model generated based on training data to reproduce and analyze specific work procedures and actions.
[0516] "Training" refers to the process of teaching a generative AI model so that it learns specific task steps.
[0517] A "virtual reality system" is a system that allows users to experience and practice actual work procedures in a virtual space.
[0518] "Matching" refers to the process of connecting craftsmen with potential successors based on certain criteria.
[0519] A "training environment" is an environment in which a user can practice or perform exercises through a virtual reality system.
[0520] "Feedback" refers to the evaluation and improvement advice provided in real time to the user's actions and results during training.
[0521] This invention relates to a system that records the work procedures of craftsmen in detail and has AI learn from them to pass on skills. The mode for carrying out the invention is configured as follows.
[0522] 1. Data collection device
[0523] Devices: Smartphones and video cameras are used to record the craftsman's work procedures. This allows the craftsman's detailed movements and techniques to be recorded in detail using video and photographs. For example, a potter can use the video function on his smartphone to record the process of shaping clay, and the position and movement of his hands can be recorded in photographs.
[0524] 2. Uploading and saving data
[0525] User: After capturing the data, the user can check it through a dedicated smartphone application and upload it to the server. This upload process is performed by pressing a dedicated button within the app.
[0526] Server: Receives uploaded data and stores it in cloud storage. After storage, the data is automatically labeled and organized.
[0527] 3. Data Preprocessing
[0528] Server: Analyzes the stored video and photo data and splits it into frames. Each frame is then tagged with metadata that describes the task, including hand position and movement.
[0529] 4. Training the AI model
[0530] Server: Trains a generative AI model using the preprocessed data. It uses machine learning algorithms to learn specific work procedures performed by artisans and generate highly accurate models. For example, by learning pottery work procedures, an AI model can be created that can generate new work procedures.
[0531] 5. Creating a training environment
[0532] Device: Integrate the trained generative AI model into a virtual reality (VR) system so that when a user puts on a VR headset, they can follow the learned steps in the virtual space and practice.
[0533] 6. Feedback and Ratings
[0534] Server: Analyzes the work data performed by the user in the VR environment in real time and provides feedback. The AI model analyzes the user's movements and suggests accurate movements and areas for improvement.
[0535] 7. Successor matching
[0536] Server: Provides a platform for effectively matching craftsmen with potential successors. Proposes optimal matches based on the craftsman's skill set and the successor's interests.
[0537] Prompt Sentence Examples
[0538] "Create an AI model that learns specific steps from a video of a craftsman kneading clay. Analyze the movements of each step in detail and then provide practical training in a VR training environment."
[0539] In this way, this system realizes the effective inheritance of traditional skills and provides an innovative approach that utilizes AI and virtual reality technology, making it possible to efficiently pass on artisan skills to the next generation.
[0540] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0541] Step 1:
[0542] Data collection
[0543] Input: A craftsman films the work procedure using a smartphone or video camera.
[0544] Specific actions: The user, a craftsman, uses a smartphone or video camera to record video and photos of work procedures, such as shaping and firing pottery. Since it is necessary to capture detailed actions from various angles, multiple videos and images are taken.
[0545] Output: Video and photo data showing the work procedure.
[0546] Step 2:
[0547] Data upload
[0548] Input: Completed video and photo data.
[0549] How it works: The user checks the captured data through a dedicated smartphone application and uploads it to the server. By pressing a dedicated button on the app, the data is sent to the cloud.
[0550] Output: Video and photo data uploaded to the server.
[0551] Step 3:
[0552] Data reception and storage
[0553] Input: Uploaded video and photo data.
[0554] What it does: The server receives the uploads and stores them in cloud storage, automatically labeling and organizing the data at the same time.
[0555] Output: Structured data stored in cloud storage.
[0556] Step 4:
[0557] Data Preprocessing
[0558] Input: Video and photo data stored in cloud storage.
[0559] How it works: The server divides the data into frames and adds metadata (information about the hand's position and movement) to each frame. Image analysis algorithms are used to identify the hand's movement and position.
[0560] Output: A frame-by-frame dataset with metadata.
[0561] Step 5:
[0562] Training an AI model
[0563] Input: The preprocessed dataset.
[0564] How it works: The server uses machine learning algorithms to train a generative AI model, for example, for a specific task (such as shaping pottery) and then replicates the movements and positions of each step.
[0565] Output: A trained generative AI model.
[0566] Step 6:
[0567] Creating a training environment
[0568] Input: A trained generative AI model.
[0569] How it works: The device integrates the trained generative AI model into a virtual reality (VR) system, so that when the user puts on the VR headset, the work steps are accurately reproduced in the virtual space.
[0570] Output: Model integrated into a virtual reality system.
[0571] Step 7:
[0572] VR training
[0573] Input: A generative AI model integrated into a virtual reality system.
[0574] Specific actions: The user puts on a VR headset and performs exercises in the VR space by following the steps generated by the AI model. For example, the user recreates the action of molding clay in a virtual space and practices the steps.
[0575] Output: User's practice data.
[0576] Step 8:
[0577] Feedback and Ratings
[0578] Input: User's practice data.
[0579] Specific movements: The server analyzes the user's practice data in real time and provides feedback. The AI model analyzes the user's movements and provides a detailed evaluation of the exact movements and areas for improvement.
[0580] Output: Real-time feedback provided to the user.
[0581] Step 9:
[0582] Successor matching
[0583] Input: Craftsman skillset and potential successor information.
[0584] How it works: The server proposes optimal matches based on the craftsman's skill set and the potential successor's interests and goals. A matching algorithm is used to connect craftsmen and potential successors.
[0585] Output: Matching results between craftsmen and potential successors.
[0586] (Application example 1)
[0587] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0588] In today's world, where it is difficult to pass on the skills of craftsmen, there is a need for a method to reliably record work procedures and pass them on to the next generation. Furthermore, particularly in factory production environments, transferring the skills of skilled craftsmen to robots would lead to improved efficiency and quality, but no concrete method for doing so has been established. Furthermore, there is an issue of a gap between training in a virtual environment and actual robot operation.
[0589] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0590] In this invention, the server includes means for recording the work procedures of craftsmen, means for receiving and analyzing the recorded data, means for training a generative AI model using the analyzed data, means for providing a pseudo-training environment using the trained generative AI model, means for matching craftsmen with potential successors, and means for applying the trained generative AI model to a robot and having it execute the work procedures. This allows for the efficient transfer of craftsman skills and enables factory robots to perform advanced tasks.
[0591] A "craftsman" is someone who has specific skills or techniques and uses those skills to create products or crafts.
[0592] "Work procedure" refers to the series of steps and processes that a craftsman follows to create a product using his skilled techniques and skills.
[0593] "Means of recording" refers to devices or methods for recording a craftsman's work procedures as video or photographs using a smartphone, video camera, etc.
[0594] "Means of analysis" refers to the techniques and equipment used to analyze recorded video and photographs and extract important technical and skill elements from them.
[0595] A "generative AI model" is an AI that uses machine learning algorithms to learn from collected data and imitate or improve the work procedures of craftsmen.
[0596] A "simulated training environment" is a place or system that uses virtual reality (VR) and simulation to recreate the work performed by craftsmen, allowing them to train in an environment that is similar to the actual work environment.
[0597] "Matching methods" refer to systems and methods that match the skill sets of craftsmen with the interests and abilities of potential successors, and then select and connect the most suitable craftsmen and successors.
[0598] A "robot" is a mechanical device that can actually carry out work procedures learned using an AI model.
[0599] MODE FOR CARRYING OUT THE INVENTION
[0600] This invention is a system that records and analyzes the work procedures of craftsmen in detail, and then uses a generative AI model to have a robot carry out the work. The system configuration and operation are explained below.
[0601] Data collection methods
[0602] The device uses a smartphone or video camera to record video and photos of the craftsman's work procedures, allowing the craftsman's detailed movements and techniques to be recorded in detail.
[0603] Data analysis
[0604] The server receives and analyzes the recorded video and photo data, including splitting the footage and photos into frames and analyzing the movements of the craftsmen, using image processing techniques and machine learning algorithms.
[0605] Training generative AI models
[0606] The server trains a generative AI model based on the analyzed data, building a model that mimics or improves the craftsman's work procedures. The machine learning algorithms used include deep learning and reinforcement learning.
[0607] Providing a simulated training environment
[0608] The device provides a virtual reality (VR) environment using the trained generative AI model. In this environment, users can practice based on the learned procedures in a virtual space. Users wear a VR headset and train by following the instructions generated by the AI model.
[0609] Application to robots
[0610] The server integrates the trained generative AI model into the robot control system, allowing the robot to actually execute the craftsman's work steps that the AI model has learned. The robot can be an industrial robot arm or an autonomous robot.
[0611] Matching System
[0612] The server provides a platform for matching artisans with potential successors. This platform matches the skill sets of artisans with the interests and abilities of successors to create the optimal pairing. Users access the platform and register their own skills and interests to be matched.
[0613] Examples and prompts
[0614] For example, a skilled worker assembling a product uses a camera to record the steps of installing parts. The server analyzes the captured video, divides the actions into frames, and extracts important parts. The generative AI model uses this data to learn and formulates assembly steps. It then instructs the robot arm to execute the steps. The entire process is instructed to the AI model with the following prompt:
[0615] "Analyze this video frame by frame and generate step-by-step assembly instructions."
[0616] In this way, the present invention provides a series of means for effectively transferring the skills of craftsmen, making it possible to realize advanced work using factory robots.
[0617] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0618] Step 1:
[0619] The terminal uses a smartphone or video camera to record the craftsman's work procedures with video and photos. The input data are the videos and photos of the craftsman's movements and techniques. The output data are the video and photo files stored on the recording device.
[0620] Step 2:
[0621] The server receives the recorded video and photo data, divides it into frames, and performs analysis. The input data are the video and photo files obtained in step 1. The output data are the analyzed frame-by-frame motion data and its metadata. Specifically, image processing techniques are used to extract important motion points from each frame.
[0622] Step 3:
[0623] The server uses the analyzed data to train a generative AI model. The input data is the frame-by-frame motion data and metadata obtained in step 2. The output data is the trained generative AI model. Specifically, it uses machine learning algorithms (e.g., deep learning and reinforcement learning) to learn engineering features from the analyzed data.
[0624] Step 4:
[0625] The device provides a virtual reality (VR) environment using a trained generative AI model. The input data is the trained generative AI model. The output data is a virtual training environment accessible to users. Specifically, users wear a VR headset and practice in a virtual space by following the steps of a craftsman generated by the AI model.
[0626] Step 5:
[0627] The server integrates the trained generative AI model into the robot control system and has the robot execute the work procedure. The input data is the trained generative AI model. The output data is the specific operation instructions to be executed by the robot. Specifically, the robot arm is controlled to perform tasks such as assembly and processing.
[0628] Step 6:
[0629] The server provides a platform for matching craftsmen with potential successors. The input data is the skill set of the craftsman and the interests and abilities of the potential successor. The output data is the pairing results of the optimal craftsman and potential successor. Specifically, users log in to the platform and enter the necessary information, and the best match is suggested.
[0630] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0631] This invention relates to a system that records the detailed work procedures of craftsmen and allows them to be trained by AI to pass on their skills, and also combines this with an emotion engine that recognizes the user's emotions. This system provides various means for effectively passing on craftsmen's skills to the next generation. The system configuration and operation for specifically implementing this invention are described below.
[0632] Basic system configuration
[0633] 1. Data collection device
[0634] Devices: Includes devices such as smartphones and video cameras that record the craftsman's work procedures with video and photos, allowing the craftsman's detailed movements and techniques to be recorded in detail.
[0635] 2. Data Management Server
[0636] Server: A server for receiving and analyzing recorded video and photo data. The server works with a database to organize and store the recorded data. It also has software for data analysis and preprocesses the uploaded data.
[0637] 3. Generative AI Models
[0638] Server: The preprocessed data is used to train a generative AI model, which then learns the specific work steps of the craftsman and generates new work steps. A machine learning algorithm is used to create a highly accurate model across multiple steps.
[0639] 4. Virtual Reality (VR) Training Environment
[0640] Device: A VR system that uses a trained generative AI model to provide a training environment in a virtual space. The user wears a VR headset and performs training based on the generated steps in the virtual space.
[0641] 5. Successor Matching Platform
[0642] Server: A platform for effectively matching artisans with potential successors. It implements a matching algorithm based on the artisan's skill set and the successor's interests, and proposes optimal pairs.
[0643] 6. Emotion Engine
[0644] Server: Equipped with an emotion engine that analyzes the user's facial expressions and voice data to recognize their emotions. Analyzes the user's emotional data in real time and reflects it in training content and feedback.
[0645] Program processing overview
[0646] 1. Data collection procedure
[0647] User: The craftsman uses a smartphone or video camera to take videos and photos of the work procedure. Once the recording is complete, the data can be viewed through the application on the device and uploaded to the server.
[0648] 2. Data preprocessing and storage
[0649] Server: Receives uploaded video and photo data, stores them in storage, records metadata in a database, and performs preprocessing for analysis.
[0650] 3. Training the AI model
[0651] Server: Trains the generative AI model using the preprocessed data. It uses machine learning algorithms to learn the craftsman's work procedures and generate a highly accurate model.
[0652] 4. Providing a training environment
[0653] Device: The trained generative AI model is integrated into a VR system, allowing users to practice in a virtual space. The user puts on a VR headset and practices by following the steps generated by the AI model.
[0654] 5. Feedback and Ratings
[0655] Server: Analyzes the user's actions in the VR environment in real time and provides feedback, allowing the user to improve their skills.
[0656] 6. Emotional Recognition
[0657] Server: Captures the user's facial expressions and voice in real time, and the emotion engine analyzes the data. Based on the user's emotional state, the system adjusts the training progress.
[0658] Device: Receives feedback from the emotion engine. For example, if the user is feeling stressed, the system adjusts the training to reduce stress.
[0659] Users: By receiving appropriate feedback and adjustments based on their emotional state, they can train more efficiently and with greater focus.
[0660] Specific examples
[0661] 1. Data collection: Potters use their smartphones to videotape the process of kneading the clay, shaping it, and firing it, and also take photographs of each detailed step.
[0662] 2. Data preprocessing: The server splits the received video data into frames and adds metadata about the work done to each frame.
[0663] 3. Training the AI model: The server trains a generative AI model that learns specific steps in pottery making, enabling it to generate new steps.
[0664] 4. Providing a training environment: Users use a VR headset to practice pottery making in a virtual space, following the steps they learned. AI provides real-time feedback.
[0665] 5. Emotion recognition and adjustment: If the user is feeling stressed during training, the emotion engine will recognize this and the system will adjust the training content, for example by suggesting a short break to reduce the user's stress.
[0666] 6. Matching Platform: Young potters are matched with experienced craftsmen through the platform and begin online sessions.
[0667] In this way, the system effectively passes on the skills of craftsmen to the next generation and supports skill improvement while taking into consideration the user's emotional state.
[0668] The processing flow will be explained below.
[0669] Step 1: Data collection
[0670] User: Using a smartphone or video camera, the user takes videos and photographs of the craftsman's work steps, allowing the craftsman to record his work in detail.
[0671] Device: Press the start recording button to capture video and photos of the craftsman's work. The captured data is temporarily saved.
[0672] User: After shooting, check the video and photo data through the application and enter metadata (process, date, time, etc.).
[0673] Terminal: Compresses video and photo data and uploads it to a server via the Internet.
[0674] Step 2: Data Preprocessing
[0675] Server: Stores the received video and photo data in storage. Creates a new entry in the database and records metadata related to the uploaded video (such as the craftsman's name, the work performed, and the date and time of the shoot).
[0676] Server: Splits the video data into frames, assigns a timestamp to each frame, extracts important frames for analysis, and removes unnecessary frames.
[0677] Server: Converts the stored image data into tensor format and audio data into spectrograms, thereby adjusting the data format to be analyzable by machine learning algorithms.
[0678] Step 3: Training the AI model
[0679] Server: Inputs the preprocessed dataset into the generative AI model and begins training the model. Using a supervised learning algorithm, it learns the craftsman's work procedures in detail.
[0680] Server: Monitors the loss function and accuracy every epoch during the training process, adjusting hyperparameters as training progresses to optimize model performance.
[0681] Server: Stores the trained generative AI model for later use in training and simulation.
[0682] Step 4: Providing a training environment
[0683] Server: Integrates the trained generative AI model into the VR system, allowing users to simulate tasks in a virtual space.
[0684] Device: The user wears a VR headset and follows the steps generated by the AI model in a virtual space. The user uses the VR controller to recreate the tasks in the virtual environment.
[0685] User: Enter the VR environment and follow the steps guided by the AI model. Practice while checking whether the steps are progressing as expected.
[0686] Step 5: Real-time feedback and evaluation
[0687] Server: Analyzes the work users do in the virtual environment in real time and determines which areas need improvement.
[0688] Server: Based on the analysis results, the server provides visual and audio feedback to the user, such as advice like "Take more time to shape this part."
[0689] Users: Receive feedback and then rework the work based on that feedback, thereby improving their skills.
[0690] Step 6: Recognize emotions
[0691] Server: Captures the user's facial expressions and voice in real time, and the emotion engine analyzes the data to recognize the user's emotional state (stress level and concentration state).
[0692] Device: Receives feedback from the emotion engine. If the user is feeling stressed, the system can adjust the training, for example suggesting a short break.
[0693] Users: By receiving appropriate feedback and adjustments based on their emotional state, they can train more efficiently and with greater focus.
[0694] Step 7: Use a matching platform
[0695] Server: Runs a matching algorithm based on the skill sets and interests of the successor candidate and the craftsman, and proposes the optimal combination.
[0696] Device: Notifies users of matches and displays their profile and contact information.
[0697] User: Logs in to the platform and checks the proposed matching results. The craftsman and successor communicate directly and create a specific training plan.
[0698] Through these steps, the system effectively passes on the craftsman's skills to the next generation and supports skill improvement while taking into account the user's emotional state.
[0699] Example 2
[0700] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0701] In today's aging society, effectively passing on the skills of skilled craftsmen to the next generation is an important issue. However, traditional methods of skill transfer are inefficient, as they make it difficult to fully convey the finer details of the skills and the craftsman's know-how. Furthermore, stress and a decline in motivation for the successor during the skill transfer process can also be an issue.
[0702] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for recording the work procedures of the craftsman with video and photographs, a means for receiving the recorded data, preprocessing and analyzing the data, and a means for training a machine learning model using the analyzed data. This makes it possible to record and analyze the specific work procedures and know-how of the craftsman with high accuracy.
[0703] Furthermore, this invention includes a means for providing a virtual training environment using a trained machine learning model, a means for matching craftsmen with potential successors, and a means for recognizing the user's emotions in real time and adjusting the training content based on those emotions. This increases the efficiency of skill transfer, prevents stress and a decline in motivation for successors, and enables smooth skill improvement.
[0704] A "craftsman" is a professional who has specific skills and techniques and creates things based on those skills.
[0705] "Work procedures" refer to the specific steps and methods of operation that craftsmen use to perform work.
[0706] "Video" refers to video data captured using a camera or other device.
[0707] A "photograph" refers to still image data captured at a specific moment with a camera or other device.
[0708] "Means of recording" refers to means for recording video or still images using devices such as video cameras or smartphones.
[0709] "Means for receiving, pre-processing and analyzing data" refers to means by which the server receives uploaded video and photo data and performs the data processing necessary to analyze them.
[0710] A "machine learning model" refers to an algorithm that uses AI techniques to learn from a specific dataset and make predictions or generation decisions.
[0711] "Means for providing a virtual training environment" refers to means for providing an environment in which a user can recreate and practice the work procedures of a craftsman in a virtual space.
[0712] "Potential successors" refer to people who wish to inherit the skills of a particular craftsman.
[0713] "Matching means" refers to algorithms and platforms that efficiently connect artisans with potential successors.
[0714] "Means for recognizing emotions in real time" refers to means for analyzing the user's facial expressions and voice data to identify their current emotional state.
[0715] "Means for adjusting training content" refers to means for adaptively changing training content in a virtual training environment based on the user's emotional state.
[0716] The present invention is a system that effectively transfers skills by recording the work procedures of craftsmen in detail and building a generative AI model based on the records. This system has the ability to recognize the user's emotions in real time and dynamically adjust the training content. The following describes the system configuration and operation for specifically implementing the present invention.
[0717] Basic system configuration
[0718] 1. Data collection device
[0719] Devices: Recording devices such as smartphones and video cameras are used. These devices record the craftsman's work procedures as videos and photographs, allowing the craftsman's detailed movements and techniques to be recorded in detail.
[0720] 2. Data Management Server
[0721] Server: This is the server that receives and analyzes the recorded video and photo data. The server connects to a database to organize and store the recorded data. It also has software for data analysis and preprocesses the uploaded data. Specifically, it uses tools such as Python and OpenCV to segment the data and add metadata.
[0722] 3. Generative AI Models
[0723] Server: The generative AI model is trained using the preprocessed data. This allows it to learn the specific work procedures of craftsmen and generate new work procedures. A machine learning algorithm is used to create a highly accurate model across multiple steps. TensorFlow and PyTorch are suitable libraries to use.
[0724] 4. Virtual Reality (VR) Training Environment
[0725] Device: A VR system is used that uses a trained generative AI model to provide a training environment in a virtual space. The user wears a VR headset (e.g., Oculus Rift or HTC Vive) and performs training based on the generated steps in the virtual space. The VR environment is built using Unity or Unreal Engine.
[0726] 5. Successor Matching Platform
[0727] Server: Provides a platform for effectively matching artisans with potential successors. Implements a matching algorithm based on the artisan's skill set and the successor's interests, and proposes optimal pairs.
[0728] 6. Emotion Engine
[0729] Server: Equipped with an emotion engine that analyzes the user's facial expressions and voice data to recognize emotions. Analyzes the user's emotional data in real time and reflects it in training content and feedback. Emotion analysis is performed using Microsoft Azure Cognitive Services and the Affectiva SDK.
[0730] Specific examples
[0731] 1. Data collection: Potters use their smartphones to videotape the process of kneading the clay, shaping it, and firing it, and also take photographs of each detailed step.
[0732] 2. Data preprocessing: The server splits the received video data into frames and adds metadata about the work done to each frame.
[0733] 3. Training the AI model: The server trains a generative AI model that learns specific steps in pottery making, enabling it to generate new steps.
[0734] 4. Providing a training environment: Users use a VR headset (e.g., Oculus Rift) to practice pottery making in a virtual space, following the steps they learned. AI provides real-time feedback.
[0735] 5. Emotion recognition and adjustment: If the user is feeling stressed during training, the emotion engine will recognize this and the system will adjust the training content, for example by suggesting a short break to reduce the user's stress.
[0736] 6. Matching Platform: Young potters are matched with experienced craftsmen through the platform and begin online sessions.
[0737] Example prompt
[0738] Example prompt: "I would like to learn the steps of pottery making. I would like to provide videos of each step of the pottery making process, from mixing the clay to molding and firing, and train in a virtual space."
[0739] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0740] Step 1:
[0741] The user uses a smartphone or video camera to record video and photos of the craftsman's work procedures. Specifically, the user records the steps of the potter kneading the clay, shaping it, and firing it. The input data are video and photo files, which are then uploaded to a dedicated application on the device. The output data are verified video and photo files.
[0742] Step 2:
[0743] The device reviews the video and photo data received from the user through the application and edits it as needed, such as cutting out unnecessary parts and labeling important parts. The input data are video and photo files, and the output data are the edited video and photo files.
[0744] Step 3:
[0745] The device uploads the edited video and photo files to the server by sending them to a specified URL on the server using an HTTP request. The input data are the edited video and photo files, and the output data are the video and photo data uploaded to the server.
[0746] Step 4:
[0747] The server receives the uploaded video and photo data and stores them in storage. Specifically, it writes the received files to disk storage and records the metadata in a database. The input data is the uploaded video and photo data, and the output data is the stored video and photo data and their metadata.
[0748] Step 5:
[0749] The server analyzes and preprocesses the stored video and photo data. Specifically, it uses OpenCV to split the video data into frames and add metadata about the work done to each frame. It also performs noise removal and resolution adjustment. The input data is the stored video and photo data, and the output data is the preprocessed video and photo data.
[0750] Step 6:
[0751] The server uses the preprocessed data to train a generative AI model. Specifically, it uses TensorFlow and PyTorch to run machine learning algorithms to learn the craftsman's work procedures. The input data is the preprocessed video and photo data, and the output data is the trained generative AI model.
[0752] Step 7:
[0753] The device integrates the trained generative AI model into the VR system. Specifically, it uses Unity or Unreal Engine to recreate the steps generated by the model in a virtual space so that the user can experience them. The input data is the trained generative AI model, and the output data is the work steps recreated in the VR environment.
[0754] Step 8:
[0755] The user wears a VR headset and performs practical training in a virtual space by following the steps provided by the generative AI model. Specifically, the user practices pottery molding in a virtual space. The input data is the work steps reproduced in the VR environment, and the output data is the motion data of the work performed by the user.
[0756] Step 9:
[0757] The server analyzes the user's behavior data in real time and provides feedback. Specifically, it points out errors in behavior and suggests ways to correct them. It also performs real-time data processing using Apache Kafka and Apache Flink. The input data is the user's behavior data, and the output data is the feedback message.
[0758] Step 10:
[0759] The server captures the user's emotional data in real time and analyzes it using an emotion engine. Specifically, it uses Microsoft Azure Cognitive Services and the Affectiva SDK to recognize emotions from the user's facial expressions and voice. The input data is the user's facial expression and voice data, and the output data is the emotion analysis results.
[0760] Step 11:
[0761] The device dynamically adjusts the training content based on the results obtained from the emotion engine. Specifically, if the user feels stressed, the system will reduce the training content or suggest taking a break. The input data is the emotion analysis results, and the output data is the adjusted training content.
[0762] Step 12:
[0763] The server operates a platform that matches artisans with potential successors. Specifically, it records the skill sets of artisans and the interests of successors in a database and uses a recommendation algorithm to suggest optimal pairs. The input data are the skills and interests of artisans and successors, and the output data are the proposed matching pairs.
[0764] (Application example 2)
[0765] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0766] The transfer of craftsmanship skills requires specific skills and experience, which is time-consuming and costly. Furthermore, conventional methods have difficulty providing feedback that takes into account the craftsman's emotions and state. Furthermore, for a robot to accurately learn and execute these skills, advanced data collection and analysis techniques are required. Therefore, there is a need for a system that can efficiently and effectively transfer craftsmanship skills and enable robots to execute skills and provide feedback that reflects the user's emotions.
[0767] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0768] In this invention, the server includes means for recording the work procedures of craftsmen with video and photographs, means for receiving and analyzing the recorded data, means for training a generative AI model using the analyzed data, means for providing a simulated training environment using the trained generative AI model, means for matching craftsmen with potential successors, means for collecting work procedure data in real time from craftsmen wearing smart glasses, means for controlling a robot using the collected data to execute the learned work procedures, and means for recognizing the user's emotional state using emotion analysis means and providing feedback. This allows the craftsman's work procedures to be recorded in detail, enabling the robot to learn and execute the procedures, and providing feedback based on the user's emotional state in real time.
[0769] "Means for video and photographic recording of craftsmen's work procedures" is a general term for video cameras and camera-equipped devices used to record in detail the series of tasks and operations performed by craftsmen.
[0770] "Means for receiving and analyzing recorded data" refers to a combination of software and hardware for collecting said video and photographic data and analyzing it using appropriate algorithms.
[0771] "Means for training a generative AI model" is a general term for software and computer equipment that applies machine learning algorithms based on the analyzed data to train an AI model.
[0772] "Means for providing a pseudo-training environment using a trained generative AI model" refers to a system that uses a trained AI model to provide a training environment in a simulated format to users.
[0773] The "means for matching craftsmen with potential successors" refers to an algorithm and system that compares the skills of craftsmen with the interests and abilities of potential successors and proposes the optimal combination.
[0774] "Means for collecting work procedure data in real time from craftsmen wearing smart glasses" refers to technology and devices that use smart glasses to capture work procedures from the perspective of craftsmen in real time and collect them as data.
[0775] The "means for controlling the robot using collected data to execute the learned work procedure" is a system for causing the robot to execute the work procedure based on data collected from smart glasses and other sensor devices.
[0776] "Means for recognizing a user's emotional state using emotion analysis means and providing feedback" refers to software that analyzes a user's facial expressions and voice data to recognize their emotional state, and a system for providing feedback to the user based on that.
[0777] The system for implementing this invention records the work procedures of craftsmen in detail, trains a generative AI model based on that data, and applies it to a robot to achieve skill transfer and efficient work execution. It also includes an emotion analysis means for analyzing the user's emotional state in real time and providing appropriate feedback.
[0778] System Configuration
[0779] 1. Data collection device
[0780] Terminal: The craftsman wears smart glasses that capture the work procedure in real time. The smart glasses are equipped with a camera and a microphone to record both visual and audio data, which is then sent to a server via the Internet.
[0781] 2. Data Management Server
[0782] Server: Receives, stores, and analyzes recorded video and audio data. The server works with a database (e.g., MongoDB) to organize and store the data. It also preprocesses the data using machine learning algorithms such as TensorFlow for data analysis.
[0783] 3. Generative AI Models
[0784] Server: The generative AI model is trained using machine learning libraries such as TensorFlow based on the preprocessed data. This model replicates the specific steps of the craftsman's work and applies them to the robot.
[0785] 4. Robot Control
[0786] Terminal: The robot executes the collected work procedures based on the generative AI model. For example, it can use a general-purpose robot arm such as the UR5 or UR10 to perform precise movements.
[0787] 5. Emotion analysis method
[0788] Server: Receives the user's facial expressions and voice data from smart glasses and other devices, and performs emotion analysis using Microsoft Azure's Emotion API, etc. Based on the analysis results, provides feedback to the user in real time.
[0789] Specific examples
[0790] In a factory, a skilled craftsman performs welding work. During this process, the craftsman wears smart glasses that record his work steps in real time. The collected data is sent to a server where it is analyzed. A generative AI model learns the craftsman's work steps and generates new welding steps. This model is then used by a robotic arm to perform the work. Furthermore, an emotion engine analyzes the observer's emotional state, and if stress or doubt is recognized, the system provides appropriate feedback and corrections.
[0791] Prompt Sentence Examples
[0792] "Example of teaching AI new work procedures:
[0793] Detailed procedures for welding work that should be performed by factory robots
[0794] Important points to highlight in the video data
[0795] Expert instructions to be extracted from speech data
[0796] Specific methods for handling facial expression and voice data for emotion analysis
[0797] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0798] Step 1:
[0799] Terminal: The craftsman puts on the smart glasses and begins work. The smart glasses are equipped with a camera and microphone, which record video and audio data of the work in real time.
[0800] Input: Worker's workflow (visual and audio data)
[0801] Output: Recorded video and audio data
[0802] How it works: Smart glasses capture video and audio and collect data in real time.
[0803] Step 2:
[0804] Server: Stores the video and audio data received from the smart glasses in storage, and adds metadata (date, time, operation details, etc.) to the data.
[0805] Input: Video and audio data sent from smart glasses
[0806] Output: Data saved to storage (with metadata)
[0807] Specific operation: Stores data in a database (such as MongoDB) and organizes it by adding metadata.
[0808] Step 3:
[0809] Server: Analyzes the stored video and audio data and extracts the necessary information. Preprocesses the data using machine learning algorithms (e.g., TensorFlow).
[0810] Input: Stored video and audio data (with metadata)
[0811] Output: Preprocessed data (features extracted)
[0812] Specific operations: Split video data into frames and extract key points. Extract important reference phrases from audio data.
[0813] Step 4:
[0814] Server: Trains the generative AI model using the preprocessed data.
[0815] Input: Preprocessed data (features extracted)
[0816] Output: A trained generative AI model
[0817] How it works: Using TensorFlow, we train a model to learn the work procedures of craftsmen, using multi-layer perceptrons and convolutional neural networks (CNNs).
[0818] Step 5:
[0819] Terminal: The trained generative AI model is applied to the robot, which then executes the actual work steps.
[0820] Input: A trained generative AI model
[0821] Output: Robot performs work
[0822] Specific operation: Control a robot arm (e.g., UR5 or UR10) and perform tasks according to learned procedures.
[0823] Step 6:
[0824] Server: Receives user facial and voice data from smart glasses and other sensor devices and performs emotion analysis. Uses Microsoft Azure's Emotion API.
[0825] Input: User facial and voice data
[0826] Output: Emotional state analysis result
[0827] Specific operation: Analyzes user emotion data in real time using Microsoft Azure's Emotion API.
[0828] Step 7:
[0829] Server: Generates feedback based on the results of sentiment analysis and provides it to the user.
[0830] Input: Emotional state analysis result
[0831] Output: Feedback (adjustments to training, new instructions, etc.)
[0832] What it does: Depending on the user's emotional state, the system will adjust the intensity of their workout or suggest a break. Feedback is provided via voice and text.
[0833] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0834] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0835] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0836] [Third embodiment]
[0837] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0838] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0839] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0840] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0841] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0842] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0843] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0844] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0845] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0846] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0847] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0848] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0849] The present invention relates to a system that records the detailed work procedures of craftsmen and allows them to be learned by AI to pass on their skills. This system provides various means for effectively passing on the skills of craftsmen to the next generation. The following describes the system configuration and operation for specifically implementing the present invention.
[0850] Basic system configuration
[0851] 1. Data collection device
[0852] Devices: Includes devices such as smartphones and video cameras that record the craftsman's work procedures with video and photos, allowing the craftsman's detailed movements and techniques to be recorded in detail.
[0853] 2. Data Management Server
[0854] Server: A server for receiving and analyzing recorded video and photo data. The server works with a database to organize and store the recorded data. It also has software for data analysis and preprocesses the uploaded data.
[0855] 3. Generative AI Models
[0856] Server: The preprocessed data is used to train a generative AI model, which then learns the specific work steps of the craftsman and generates new work steps. A machine learning algorithm is used to create a highly accurate model across multiple steps.
[0857] 4. Virtual Reality (VR) Training Environment
[0858] Device: A VR system that uses a trained generative AI model to provide a training environment in a virtual space. The user wears a VR headset and performs training based on the generated steps in the virtual space.
[0859] 5. Successor Matching Platform
[0860] Server: A platform for effectively matching artisans with potential successors. It implements a matching algorithm that matches the skill sets of artisans and the interests of successors, and proposes optimal pairs.
[0861] Program processing overview
[0862] 1. Data collection procedure
[0863] User: The craftsman uses a smartphone or video camera to take videos and photos of the work procedure. Once the recording is complete, the data can be viewed through the application on the device and uploaded to the server.
[0864] 2. Data preprocessing and storage
[0865] Server: Receives uploaded video and photo data, stores them in storage, records metadata in a database, and performs preprocessing for analysis.
[0866] 3. Training the AI model
[0867] Server: Trains the generative AI model using the preprocessed data. It uses machine learning algorithms to learn the craftsman's work procedures and generate a highly accurate model.
[0868] 4. Providing a training environment
[0869] Device: The trained generative AI model is integrated into a VR system, allowing users to practice in a virtual space. The user puts on a VR headset and practices by following the steps generated by the AI model.
[0870] 5. Feedback and Ratings
[0871] Server: Analyzes the user's actions in the VR environment in real time and provides feedback, allowing the user to improve their skills.
[0872] 6. Use of Matching Platforms
[0873] Users: By accessing the platform, craftsmen register their skills, and successors input their interests and goals. The system then suggests optimal matches, and users can start communicating based on those matches.
[0874] Specific examples
[0875] 1. Data collection: Potters use their smartphones to videotape the process of kneading the clay, shaping it, and firing it, and also take photographs of each detailed step.
[0876] 2. Data preprocessing: The server splits the received video data into frames and adds metadata about the work done to each frame.
[0877] 3. Training the AI model: The server trains a generative AI model that learns specific steps in pottery making, enabling it to generate new steps.
[0878] 4. Providing a training environment: Users use a VR headset to practice pottery making in a virtual space, following the steps they learned. AI provides real-time feedback.
[0879] 5. Matching Platform: Young potters are matched with experienced artisans through the platform and begin online sessions.
[0880] In this way, the system provides a comprehensive solution for realizing the effective inheritance of traditional skills.
[0881] The processing flow will be explained below.
[0882] Step 1: Data collection
[0883] User: Uses a smartphone or video camera to capture video and photos of the craftsman's work steps.
[0884] Device: Press the start recording button to capture video and photos of the craftsman's work. The captured data is temporarily saved.
[0885] User: After shooting, the user checks the video and photo data through the application and enters metadata (e.g., work process, date, time, etc.).
[0886] Terminal: Compresses video and photo data and uploads it to a server via the Internet.
[0887] Step 2: Data Preprocessing
[0888] Server: Stores the received video and photo data in storage. Creates a new entry in the database and records metadata related to the uploaded video (such as the craftsman's name, the work performed, and the date and time of the shoot).
[0889] Server: Splits the video data into frames, assigns a timestamp to each frame, extracts important frames for each task, and deletes unnecessary frames.
[0890] Server: Converts image data into tensor format and audio data into spectrograms. Prepares the data for analysis.
[0891] Step 3: Training the AI model
[0892] Server: Inputs the preprocessed dataset into the generative AI model and starts training the model. The training process is done using a supervised learning algorithm.
[0893] Server: Monitors the loss function and accuracy during training every epoch, adjusting hyperparameters as learning progresses and optimizing the model.
[0894] Server: Stores the trained model and makes it available for training and simulation.
[0895] Step 4: Providing a training environment
[0896] Server: Integrates the trained generative AI model into the VR system, generating a virtual space and configuring it so that users can simulate tasks in that space.
[0897] Device: Using a VR headset and controller, the user enters the virtual space and practices the generated steps.
[0898] User: Wear a VR headset and follow the steps guided by the AI in the virtual space to complete the task. Check whether you can proceed as planned.
[0899] Step 5: Real-time feedback and evaluation
[0900] Server: Analyzes the work users do in the virtual environment in real time and determines which areas need improvement.
[0901] Server: Based on the analysis results, the server provides visual and audio feedback to the user, such as advice like "Take more time to shape this part."
[0902] Users: Receive feedback and then rework the work based on that feedback, thereby improving their skills.
[0903] Step 6: Use a matching platform
[0904] Server: Runs a matching algorithm based on the skill sets and interests of the successor candidate and the craftsman, and proposes the optimal combination.
[0905] Device: Notifies users of matches and displays their profile and contact information.
[0906] User: Logs in to the platform and checks the proposed matching results. The craftsman and successor communicate directly and create a specific training plan.
[0907] Through these steps, the system provides a comprehensive solution for effectively passing on artisan skills to the next generation.
[0908] Example 1
[0909] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0910] Many traditional skills today are being lost due to the aging of artisans and a lack of successors. In particular, if there is no established method for effectively passing on artisanal techniques and know-how to the next generation, those skills are at high risk of disappearing. Traditional methods, such as simply recording artisans' work procedures with video or photographs, make it difficult to pass on skills to the next generation. Furthermore, a lack of feedback and a lack of training environments hinder the transfer of skills.
[0911] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0912] In this invention, the server includes means for recording the work procedures of craftsmen using a recording device, means for receiving and saving the recorded data, means for preprocessing the saved data, means for training a generative AI model using the preprocessed data, means for integrating the trained generative AI model into a virtual reality system, means for effectively matching craftsmen with potential successors, means for providing a training environment to users using the virtual reality system, and means for providing real-time feedback to users during training. This makes it possible to effectively pass on craftsman skills to the next generation and provide detailed, practical training using virtual reality.
[0913] A "craftsman" is a professional who has specific skills and techniques and uses them to create products or works.
[0914] A "work procedure" is a set of specific steps or processes for performing a particular task.
[0915] "Recording device" refers to a device for recording work procedures as video or photographs.
[0916] "Data" refers to video, photographs, and other information captured by recording devices.
[0917] "Receiving" refers to the process of transferring and receiving recorded data to a server or other system.
[0918] "Storage" is the process of storing received data in storage so that it can be used later.
[0919] "Preprocessing" refers to the process of performing a series of operations and filtering to transform raw data into a form that is easier to analyze.
[0920] A "generative AI model" is an artificial intelligence model generated based on training data to reproduce and analyze specific work procedures and actions.
[0921] "Training" refers to the process of teaching a generative AI model so that it learns specific task steps.
[0922] A "virtual reality system" is a system that allows users to experience and practice actual work procedures in a virtual space.
[0923] "Matching" refers to the process of connecting craftsmen with potential successors based on certain criteria.
[0924] A "training environment" is an environment in which a user can practice or perform exercises through a virtual reality system.
[0925] "Feedback" refers to the evaluation and improvement advice provided in real time to the user's actions and results during training.
[0926] This invention relates to a system that records the work procedures of craftsmen in detail and has AI learn from them to pass on skills. The mode for carrying out the invention is configured as follows.
[0927] 1. Data collection device
[0928] Devices: Smartphones and video cameras are used to record the craftsman's work procedures. This allows the craftsman's detailed movements and techniques to be recorded in detail using video and photographs. For example, a potter can use the video function on his smartphone to record the process of shaping clay, and the position and movement of his hands can be recorded in photographs.
[0929] 2. Uploading and saving data
[0930] User: After capturing the data, the user can check it through a dedicated smartphone application and upload it to the server. This upload process is performed by pressing a dedicated button within the app.
[0931] Server: Receives uploaded data and stores it in cloud storage. After storage, the data is automatically labeled and organized.
[0932] 3. Data Preprocessing
[0933] Server: Analyzes the stored video and photo data and splits it into frames. Each frame is then tagged with metadata that describes the task, including hand position and movement.
[0934] 4. Training the AI model
[0935] Server: Trains a generative AI model using the preprocessed data. It uses machine learning algorithms to learn specific work procedures performed by artisans and generate highly accurate models. For example, by learning pottery work procedures, an AI model can be created that can generate new work procedures.
[0936] 5. Creating a training environment
[0937] Device: Integrate the trained generative AI model into a virtual reality (VR) system so that when a user puts on a VR headset, they can follow the learned steps in the virtual space and practice.
[0938] 6. Feedback and Ratings
[0939] Server: Analyzes the work data performed by the user in the VR environment in real time and provides feedback. The AI model analyzes the user's movements and suggests accurate movements and areas for improvement.
[0940] 7. Successor matching
[0941] Server: Provides a platform for effectively matching craftsmen with potential successors. Proposes optimal matches based on the craftsman's skill set and the successor's interests.
[0942] Prompt Sentence Examples
[0943] "Create an AI model that learns specific steps from a video of a craftsman kneading clay. Analyze the movements of each step in detail and then provide practical training in a VR training environment."
[0944] In this way, this system realizes the effective inheritance of traditional skills and provides an innovative approach that utilizes AI and virtual reality technology, making it possible to efficiently pass on artisan skills to the next generation.
[0945] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0946] Step 1:
[0947] Data collection
[0948] Input: A craftsman films the work procedure using a smartphone or video camera.
[0949] Specific actions: The user, a craftsman, uses a smartphone or video camera to record video and photos of work procedures, such as shaping and firing pottery. Since it is necessary to capture detailed actions from various angles, multiple videos and images are taken.
[0950] Output: Video and photo data showing the work procedure.
[0951] Step 2:
[0952] Data upload
[0953] Input: Completed video and photo data.
[0954] How it works: The user checks the captured data through a dedicated smartphone application and uploads it to the server. By pressing a dedicated button on the app, the data is sent to the cloud.
[0955] Output: Video and photo data uploaded to the server.
[0956] Step 3:
[0957] Data reception and storage
[0958] Input: Uploaded video and photo data.
[0959] What it does: The server receives the uploads and stores them in cloud storage, automatically labeling and organizing the data at the same time.
[0960] Output: Structured data stored in cloud storage.
[0961] Step 4:
[0962] Data Preprocessing
[0963] Input: Video and photo data stored in cloud storage.
[0964] How it works: The server divides the data into frames and adds metadata (information about the hand's position and movement) to each frame. Image analysis algorithms are used to identify the hand's movement and position.
[0965] Output: A frame-by-frame dataset with metadata.
[0966] Step 5:
[0967] Training an AI model
[0968] Input: The preprocessed dataset.
[0969] How it works: The server uses machine learning algorithms to train a generative AI model, for example, for a specific task (such as shaping pottery) and then replicates the movements and positions of each step.
[0970] Output: A trained generative AI model.
[0971] Step 6:
[0972] Creating a training environment
[0973] Input: A trained generative AI model.
[0974] How it works: The device integrates the trained generative AI model into a virtual reality (VR) system, so that when the user puts on the VR headset, the work steps are accurately reproduced in the virtual space.
[0975] Output: Model integrated into a virtual reality system.
[0976] Step 7:
[0977] VR training
[0978] Input: A generative AI model integrated into a virtual reality system.
[0979] Specific actions: The user puts on a VR headset and performs exercises in the VR space by following the steps generated by the AI model. For example, the user recreates the action of molding clay in a virtual space and practices the steps.
[0980] Output: User's practice data.
[0981] Step 8:
[0982] Feedback and Ratings
[0983] Input: User's practice data.
[0984] Specific movements: The server analyzes the user's practice data in real time and provides feedback. The AI model analyzes the user's movements and provides a detailed evaluation of the exact movements and areas for improvement.
[0985] Output: Real-time feedback provided to the user.
[0986] Step 9:
[0987] Successor matching
[0988] Input: Craftsman skillset and potential successor information.
[0989] How it works: The server proposes optimal matches based on the craftsman's skill set and the potential successor's interests and goals. A matching algorithm is used to connect craftsmen and potential successors.
[0990] Output: Matching results between craftsmen and potential successors.
[0991] (Application example 1)
[0992] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0993] In today's world, where it is difficult to pass on the skills of craftsmen, there is a need for a method to reliably record work procedures and pass them on to the next generation. Furthermore, particularly in factory production environments, transferring the skills of skilled craftsmen to robots would lead to improved efficiency and quality, but no concrete method for doing so has been established. Furthermore, there is an issue of a gap between training in a virtual environment and actual robot operation.
[0994] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0995] In this invention, the server includes means for recording the work procedures of craftsmen, means for receiving and analyzing the recorded data, means for training a generative AI model using the analyzed data, means for providing a pseudo-training environment using the trained generative AI model, means for matching craftsmen with potential successors, and means for applying the trained generative AI model to a robot and having it execute the work procedures. This allows for the efficient transfer of craftsman skills and enables factory robots to perform advanced tasks.
[0996] A "craftsman" is someone who has specific skills or techniques and uses those skills to create products or crafts.
[0997] "Work procedure" refers to the series of steps and processes that a craftsman follows to create a product using his skilled techniques and skills.
[0998] "Means of recording" refers to devices or methods for recording a craftsman's work procedures as video or photographs using a smartphone, video camera, etc.
[0999] "Means of analysis" refers to the techniques and equipment used to analyze recorded video and photographs and extract important technical and skill elements from them.
[1000] A "generative AI model" is an AI that uses machine learning algorithms to learn from collected data and imitate or improve the work procedures of craftsmen.
[1001] A "simulated training environment" is a place or system that uses virtual reality (VR) and simulation to recreate the work performed by craftsmen, allowing them to train in an environment that is similar to the actual work environment.
[1002] "Matching methods" refer to systems and methods that match the skill sets of craftsmen with the interests and abilities of potential successors, and then select and connect the most suitable craftsmen and successors.
[1003] A "robot" is a mechanical device that can actually carry out work procedures learned using an AI model.
[1004] MODE FOR CARRYING OUT THE INVENTION
[1005] This invention is a system that records and analyzes the work procedures of craftsmen in detail, and then uses a generative AI model to have a robot carry out the work. The system configuration and operation are explained below.
[1006] Data collection methods
[1007] The device uses a smartphone or video camera to record video and photos of the craftsman's work procedures, allowing the craftsman's detailed movements and techniques to be recorded in detail.
[1008] Data analysis
[1009] The server receives and analyzes the recorded video and photo data, including splitting the footage and photos into frames and analyzing the movements of the craftsmen, using image processing techniques and machine learning algorithms.
[1010] Training generative AI models
[1011] The server trains a generative AI model based on the analyzed data, building a model that mimics or improves the craftsman's work procedures. The machine learning algorithms used include deep learning and reinforcement learning.
[1012] Providing a simulated training environment
[1013] The device provides a virtual reality (VR) environment using the trained generative AI model. In this environment, users can practice based on the learned procedures in a virtual space. Users wear a VR headset and train by following the instructions generated by the AI model.
[1014] Application to robots
[1015] The server integrates the trained generative AI model into the robot control system, allowing the robot to actually execute the craftsman's work steps that the AI model has learned. The robot can be an industrial robot arm or an autonomous robot.
[1016] Matching System
[1017] The server provides a platform for matching artisans with potential successors. This platform matches the skill sets of artisans with the interests and abilities of successors to create the optimal pairing. Users access the platform and register their own skills and interests to be matched.
[1018] Examples and prompts
[1019] For example, a skilled worker assembling a product uses a camera to record the steps of installing parts. The server analyzes the captured video, divides the actions into frames, and extracts important parts. The generative AI model uses this data to learn and formulates assembly steps. It then instructs the robot arm to execute the steps. The entire process is instructed to the AI model with the following prompt:
[1020] "Analyze this video frame by frame and generate step-by-step assembly instructions."
[1021] In this way, the present invention provides a series of means for effectively transferring the skills of craftsmen, making it possible to realize advanced work using factory robots.
[1022] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1023] Step 1:
[1024] The terminal uses a smartphone or video camera to record the craftsman's work procedures with video and photos. The input data are the videos and photos of the craftsman's movements and techniques. The output data are the video and photo files stored on the recording device.
[1025] Step 2:
[1026] The server receives the recorded video and photo data, divides it into frames, and performs analysis. The input data are the video and photo files obtained in step 1. The output data are the analyzed frame-by-frame motion data and its metadata. Specifically, image processing techniques are used to extract important motion points from each frame.
[1027] Step 3:
[1028] The server uses the analyzed data to train a generative AI model. The input data is the frame-by-frame motion data and metadata obtained in step 2. The output data is the trained generative AI model. Specifically, it uses machine learning algorithms (e.g., deep learning and reinforcement learning) to learn engineering features from the analyzed data.
[1029] Step 4:
[1030] The device provides a virtual reality (VR) environment using a trained generative AI model. The input data is the trained generative AI model. The output data is a virtual training environment accessible to users. Specifically, users wear a VR headset and practice in a virtual space by following the steps of a craftsman generated by the AI model.
[1031] Step 5:
[1032] The server integrates the trained generative AI model into the robot control system and has the robot execute the work procedure. The input data is the trained generative AI model. The output data is the specific operation instructions to be executed by the robot. Specifically, the robot arm is controlled to perform tasks such as assembly and processing.
[1033] Step 6:
[1034] The server provides a platform for matching craftsmen with potential successors. The input data is the skill set of the craftsman and the interests and abilities of the potential successor. The output data is the pairing results of the optimal craftsman and potential successor. Specifically, users log in to the platform and enter the necessary information, and the best match is suggested.
[1035] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1036] This invention relates to a system that records the detailed work procedures of craftsmen and allows them to be trained by AI to pass on their skills, and also combines this with an emotion engine that recognizes the user's emotions. This system provides various means for effectively passing on craftsmen's skills to the next generation. The system configuration and operation for specifically implementing this invention are described below.
[1037] Basic system configuration
[1038] 1. Data collection device
[1039] Devices: Includes devices such as smartphones and video cameras that record the craftsman's work procedures with video and photos, allowing the craftsman's detailed movements and techniques to be recorded in detail.
[1040] 2. Data Management Server
[1041] Server: A server for receiving and analyzing recorded video and photo data. The server works with a database to organize and store the recorded data. It also has software for data analysis and preprocesses the uploaded data.
[1042] 3. Generative AI Models
[1043] Server: The preprocessed data is used to train a generative AI model, which then learns the specific work steps of the craftsman and generates new work steps. A machine learning algorithm is used to create a highly accurate model across multiple steps.
[1044] 4. Virtual Reality (VR) Training Environment
[1045] Device: A VR system that uses a trained generative AI model to provide a training environment in a virtual space. The user wears a VR headset and performs training based on the generated steps in the virtual space.
[1046] 5. Successor Matching Platform
[1047] Server: A platform for effectively matching artisans with potential successors. It implements a matching algorithm based on the artisan's skill set and the successor's interests, and proposes optimal pairs.
[1048] 6. Emotion Engine
[1049] Server: Equipped with an emotion engine that analyzes the user's facial expressions and voice data to recognize their emotions. Analyzes the user's emotional data in real time and reflects it in training content and feedback.
[1050] Program processing overview
[1051] 1. Data collection procedure
[1052] User: The craftsman uses a smartphone or video camera to take videos and photos of the work procedure. Once the recording is complete, the data can be viewed through the application on the device and uploaded to the server.
[1053] 2. Data preprocessing and storage
[1054] Server: Receives uploaded video and photo data, stores them in storage, records metadata in a database, and performs preprocessing for analysis.
[1055] 3. Training the AI model
[1056] Server: Trains the generative AI model using the preprocessed data. It uses machine learning algorithms to learn the craftsman's work procedures and generate a highly accurate model.
[1057] 4. Providing a training environment
[1058] Device: The trained generative AI model is integrated into a VR system, allowing users to practice in a virtual space. The user puts on a VR headset and practices by following the steps generated by the AI model.
[1059] 5. Feedback and Ratings
[1060] Server: Analyzes the user's actions in the VR environment in real time and provides feedback, allowing the user to improve their skills.
[1061] 6. Emotional Recognition
[1062] Server: Captures the user's facial expressions and voice in real time, and the emotion engine analyzes the data. Based on the user's emotional state, the system adjusts the training progress.
[1063] Device: Receives feedback from the emotion engine. For example, if the user is feeling stressed, the system adjusts the training to reduce stress.
[1064] Users: By receiving appropriate feedback and adjustments based on their emotional state, they can train more efficiently and with greater focus.
[1065] Specific examples
[1066] 1. Data collection: Potters use their smartphones to videotape the process of kneading the clay, shaping it, and firing it, and also take photographs of each detailed step.
[1067] 2. Data preprocessing: The server splits the received video data into frames and adds metadata about the work done to each frame.
[1068] 3. Training the AI model: The server trains a generative AI model that learns specific steps in pottery making, enabling it to generate new steps.
[1069] 4. Providing a training environment: Users use a VR headset to practice pottery making in a virtual space, following the steps they learned. AI provides real-time feedback.
[1070] 5. Emotion recognition and adjustment: If the user is feeling stressed during training, the emotion engine will recognize this and the system will adjust the training content, for example by suggesting a short break to reduce the user's stress.
[1071] 6. Matching Platform: Young potters are matched with experienced craftsmen through the platform and begin online sessions.
[1072] In this way, the system effectively passes on the skills of craftsmen to the next generation and supports skill improvement while taking into consideration the user's emotional state.
[1073] The processing flow will be explained below.
[1074] Step 1: Data collection
[1075] User: Using a smartphone or video camera, the user takes videos and photographs of the craftsman's work steps, allowing the craftsman to record his work in detail.
[1076] Device: Press the start recording button to capture video and photos of the craftsman's work. The captured data is temporarily saved.
[1077] User: After shooting, check the video and photo data through the application and enter metadata (process, date, time, etc.).
[1078] Terminal: Compresses video and photo data and uploads it to a server via the Internet.
[1079] Step 2: Data Preprocessing
[1080] Server: Stores the received video and photo data in storage. Creates a new entry in the database and records metadata related to the uploaded video (such as the craftsman's name, the work performed, and the date and time of the shoot).
[1081] Server: Splits the video data into frames, assigns a timestamp to each frame, extracts important frames for analysis, and removes unnecessary frames.
[1082] Server: Converts the stored image data into tensor format and audio data into spectrograms, thereby adjusting the data format to be analyzable by machine learning algorithms.
[1083] Step 3: Training the AI model
[1084] Server: Inputs the preprocessed dataset into the generative AI model and begins training the model. Using a supervised learning algorithm, it learns the craftsman's work procedures in detail.
[1085] Server: Monitors the loss function and accuracy every epoch during the training process, adjusting hyperparameters as training progresses to optimize model performance.
[1086] Server: Stores the trained generative AI model for later use in training and simulation.
[1087] Step 4: Providing a training environment
[1088] Server: Integrates the trained generative AI model into the VR system, allowing users to simulate tasks in a virtual space.
[1089] Device: The user wears a VR headset and follows the steps generated by the AI model in a virtual space. The user uses the VR controller to recreate the tasks in the virtual environment.
[1090] User: Enter the VR environment and follow the steps guided by the AI model. Practice while checking whether the steps are progressing as expected.
[1091] Step 5: Real-time feedback and evaluation
[1092] Server: Analyzes the work users do in the virtual environment in real time and determines which areas need improvement.
[1093] Server: Based on the analysis results, the server provides visual and audio feedback to the user, such as advice like "Take more time to shape this part."
[1094] Users: Receive feedback and then rework the work based on that feedback, thereby improving their skills.
[1095] Step 6: Recognize emotions
[1096] Server: Captures the user's facial expressions and voice in real time, and the emotion engine analyzes the data to recognize the user's emotional state (stress level and concentration state).
[1097] Device: Receives feedback from the emotion engine. If the user is feeling stressed, the system can adjust the training, for example suggesting a short break.
[1098] Users: By receiving appropriate feedback and adjustments based on their emotional state, they can train more efficiently and with greater focus.
[1099] Step 7: Use a matching platform
[1100] Server: Runs a matching algorithm based on the skill sets and interests of the successor candidate and the craftsman, and proposes the optimal combination.
[1101] Device: Notifies users of matches and displays their profile and contact information.
[1102] User: Logs in to the platform and checks the proposed matching results. The craftsman and successor communicate directly and create a specific training plan.
[1103] Through these steps, the system effectively passes on the craftsman's skills to the next generation and supports skill improvement while taking into account the user's emotional state.
[1104] Example 2
[1105] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1106] In today's aging society, effectively passing on the skills of skilled craftsmen to the next generation is an important issue. However, traditional methods of skill transfer are inefficient, as they make it difficult to fully convey the finer details of the skills and the craftsman's know-how. Furthermore, stress and a decline in motivation for the successor during the skill transfer process can also be an issue.
[1107] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for recording the work procedures of the craftsman with video and photographs, a means for receiving the recorded data, preprocessing and analyzing the data, and a means for training a machine learning model using the analyzed data. This makes it possible to record and analyze the specific work procedures and know-how of the craftsman with high accuracy.
[1108] Furthermore, this invention includes a means for providing a virtual training environment using a trained machine learning model, a means for matching craftsmen with potential successors, and a means for recognizing the user's emotions in real time and adjusting the training content based on those emotions. This increases the efficiency of skill transfer, prevents stress and a decline in motivation for successors, and enables smooth skill improvement.
[1109] A "craftsman" is a professional who has specific skills and techniques and creates things based on those skills.
[1110] "Work procedures" refer to the specific steps and methods of operation that craftsmen use to perform work.
[1111] "Video" refers to video data captured using a camera or other device.
[1112] A "photograph" refers to still image data captured at a specific moment with a camera or other device.
[1113] "Means of recording" refers to means for recording video or still images using devices such as video cameras or smartphones.
[1114] "Means for receiving, pre-processing and analyzing data" refers to means by which the server receives uploaded video and photo data and performs the data processing necessary to analyze them.
[1115] A "machine learning model" refers to an algorithm that uses AI techniques to learn from a specific dataset and make predictions or generation decisions.
[1116] "Means for providing a virtual training environment" refers to means for providing an environment in which a user can recreate and practice the work procedures of a craftsman in a virtual space.
[1117] "Potential successors" refer to people who wish to inherit the skills of a particular craftsman.
[1118] "Matching means" refers to algorithms and platforms that efficiently connect artisans with potential successors.
[1119] "Means for recognizing emotions in real time" refers to means for analyzing the user's facial expressions and voice data to identify their current emotional state.
[1120] "Means for adjusting training content" refers to means for adaptively changing training content in a virtual training environment based on the user's emotional state.
[1121] The present invention is a system that effectively transfers skills by recording the work procedures of craftsmen in detail and building a generative AI model based on the records. This system has the ability to recognize the user's emotions in real time and dynamically adjust the training content. The following describes the system configuration and operation for specifically implementing the present invention.
[1122] Basic system configuration
[1123] 1. Data collection device
[1124] Devices: Recording devices such as smartphones and video cameras are used. These devices record the craftsman's work procedures as videos and photographs, allowing the craftsman's detailed movements and techniques to be recorded in detail.
[1125] 2. Data Management Server
[1126] Server: This is the server that receives and analyzes the recorded video and photo data. The server connects to a database to organize and store the recorded data. It also has software for data analysis and preprocesses the uploaded data. Specifically, it uses tools such as Python and OpenCV to segment the data and add metadata.
[1127] 3. Generative AI Models
[1128] Server: The generative AI model is trained using the preprocessed data. This allows it to learn the specific work procedures of craftsmen and generate new work procedures. A machine learning algorithm is used to create a highly accurate model across multiple steps. TensorFlow and PyTorch are suitable libraries to use.
[1129] 4. Virtual Reality (VR) Training Environment
[1130] Device: A VR system is used that uses a trained generative AI model to provide a training environment in a virtual space. The user wears a VR headset (e.g., Oculus Rift or HTC Vive) and performs training based on the generated steps in the virtual space. The VR environment is built using Unity or Unreal Engine.
[1131] 5. Successor Matching Platform
[1132] Server: Provides a platform for effectively matching artisans with potential successors. Implements a matching algorithm based on the artisan's skill set and the successor's interests, and proposes optimal pairs.
[1133] 6. Emotion Engine
[1134] Server: Equipped with an emotion engine that analyzes the user's facial expressions and voice data to recognize emotions. Analyzes the user's emotional data in real time and reflects it in training content and feedback. Emotion analysis is performed using Microsoft Azure Cognitive Services and the Affectiva SDK.
[1135] Specific examples
[1136] 1. Data collection: Potters use their smartphones to videotape the process of kneading the clay, shaping it, and firing it, and also take photographs of each detailed step.
[1137] 2. Data preprocessing: The server splits the received video data into frames and adds metadata about the work done to each frame.
[1138] 3. Training the AI model: The server trains a generative AI model that learns specific steps in pottery making, enabling it to generate new steps.
[1139] 4. Providing a training environment: Users use a VR headset (e.g., Oculus Rift) to practice pottery making in a virtual space, following the steps they learned. AI provides real-time feedback.
[1140] 5. Emotion recognition and adjustment: If the user is feeling stressed during training, the emotion engine will recognize this and the system will adjust the training content, for example by suggesting a short break to reduce the user's stress.
[1141] 6. Matching Platform: Young potters are matched with experienced craftsmen through the platform and begin online sessions.
[1142] Example prompt
[1143] Example prompt: "I would like to learn the steps of pottery making. I would like to provide videos of each step of the pottery making process, from mixing the clay to molding and firing, and train in a virtual space."
[1144] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1145] Step 1:
[1146] The user uses a smartphone or video camera to record video and photos of the craftsman's work procedures. Specifically, the user records the steps of the potter kneading the clay, shaping it, and firing it. The input data are video and photo files, which are then uploaded to a dedicated application on the device. The output data are verified video and photo files.
[1147] Step 2:
[1148] The device reviews the video and photo data received from the user through the application and edits it as needed, such as cutting out unnecessary parts and labeling important parts. The input data are video and photo files, and the output data are the edited video and photo files.
[1149] Step 3:
[1150] The device uploads the edited video and photo files to the server by sending them to a specified URL on the server using an HTTP request. The input data are the edited video and photo files, and the output data are the video and photo data uploaded to the server.
[1151] Step 4:
[1152] The server receives the uploaded video and photo data and stores them in storage. Specifically, it writes the received files to disk storage and records the metadata in a database. The input data is the uploaded video and photo data, and the output data is the stored video and photo data and their metadata.
[1153] Step 5:
[1154] The server analyzes and preprocesses the stored video and photo data. Specifically, it uses OpenCV to split the video data into frames and add metadata about the work done to each frame. It also performs noise removal and resolution adjustment. The input data is the stored video and photo data, and the output data is the preprocessed video and photo data.
[1155] Step 6:
[1156] The server uses the preprocessed data to train a generative AI model. Specifically, it uses TensorFlow and PyTorch to run machine learning algorithms to learn the craftsman's work procedures. The input data is the preprocessed video and photo data, and the output data is the trained generative AI model.
[1157] Step 7:
[1158] The device integrates the trained generative AI model into the VR system. Specifically, it uses Unity or Unreal Engine to recreate the steps generated by the model in a virtual space so that the user can experience them. The input data is the trained generative AI model, and the output data is the work steps recreated in the VR environment.
[1159] Step 8:
[1160] The user wears a VR headset and performs practical training in a virtual space by following the steps provided by the generative AI model. Specifically, the user practices pottery molding in a virtual space. The input data is the work steps reproduced in the VR environment, and the output data is the motion data of the work performed by the user.
[1161] Step 9:
[1162] The server analyzes the user's behavior data in real time and provides feedback. Specifically, it points out errors in behavior and suggests ways to correct them. It also performs real-time data processing using Apache Kafka and Apache Flink. The input data is the user's behavior data, and the output data is the feedback message.
[1163] Step 10:
[1164] The server captures the user's emotional data in real time and analyzes it using an emotion engine. Specifically, it uses Microsoft Azure Cognitive Services and the Affectiva SDK to recognize emotions from the user's facial expressions and voice. The input data is the user's facial expression and voice data, and the output data is the emotion analysis results.
[1165] Step 11:
[1166] The device dynamically adjusts the training content based on the results obtained from the emotion engine. Specifically, if the user feels stressed, the system will reduce the training content or suggest taking a break. The input data is the emotion analysis results, and the output data is the adjusted training content.
[1167] Step 12:
[1168] The server operates a platform that matches artisans with potential successors. Specifically, it records the skill sets of artisans and the interests of successors in a database and uses a recommendation algorithm to suggest optimal pairs. The input data are the skills and interests of artisans and successors, and the output data are the proposed matching pairs.
[1169] (Application example 2)
[1170] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1171] The transfer of craftsmanship skills requires specific skills and experience, which is time-consuming and costly. Furthermore, conventional methods have difficulty providing feedback that takes into account the craftsman's emotions and state. Furthermore, for a robot to accurately learn and execute these skills, advanced data collection and analysis techniques are required. Therefore, there is a need for a system that can efficiently and effectively transfer craftsmanship skills and enable robots to execute skills and provide feedback that reflects the user's emotions.
[1172] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1173] In this invention, the server includes means for recording the work procedures of craftsmen with video and photographs, means for receiving and analyzing the recorded data, means for training a generative AI model using the analyzed data, means for providing a simulated training environment using the trained generative AI model, means for matching craftsmen with potential successors, means for collecting work procedure data in real time from craftsmen wearing smart glasses, means for controlling a robot using the collected data to execute the learned work procedures, and means for recognizing the user's emotional state using emotion analysis means and providing feedback. This allows the craftsman's work procedures to be recorded in detail, enabling the robot to learn and execute the procedures, and providing feedback based on the user's emotional state in real time.
[1174] "Means for video and photographic recording of craftsmen's work procedures" is a general term for video cameras and camera-equipped devices used to record in detail the series of tasks and operations performed by craftsmen.
[1175] "Means for receiving and analyzing recorded data" refers to a combination of software and hardware for collecting said video and photographic data and analyzing it using appropriate algorithms.
[1176] "Means for training a generative AI model" is a general term for software and computer equipment that applies machine learning algorithms based on the analyzed data to train an AI model.
[1177] "Means for providing a pseudo-training environment using a trained generative AI model" refers to a system that uses a trained AI model to provide a training environment in a simulated format to users.
[1178] The "means for matching craftsmen with potential successors" refers to an algorithm and system that compares the skills of craftsmen with the interests and abilities of potential successors and proposes the optimal combination.
[1179] "Means for collecting work procedure data in real time from craftsmen wearing smart glasses" refers to technology and devices that use smart glasses to capture work procedures from the perspective of craftsmen in real time and collect them as data.
[1180] The "means for controlling the robot using collected data to execute the learned work procedure" is a system for causing the robot to execute the work procedure based on data collected from smart glasses and other sensor devices.
[1181] "Means for recognizing a user's emotional state using emotion analysis means and providing feedback" refers to software that analyzes a user's facial expressions and voice data to recognize their emotional state, and a system for providing feedback to the user based on that.
[1182] The system for implementing this invention records the work procedures of craftsmen in detail, trains a generative AI model based on that data, and applies it to a robot to achieve skill transfer and efficient work execution. It also includes an emotion analysis means for analyzing the user's emotional state in real time and providing appropriate feedback.
[1183] System Configuration
[1184] 1. Data collection device
[1185] Terminal: The craftsman wears smart glasses that capture the work procedure in real time. The smart glasses are equipped with a camera and a microphone to record both visual and audio data, which is then sent to a server via the Internet.
[1186] 2. Data Management Server
[1187] Server: Receives, stores, and analyzes recorded video and audio data. The server works with a database (e.g., MongoDB) to organize and store the data. It also preprocesses the data using machine learning algorithms such as TensorFlow for data analysis.
[1188] 3. Generative AI Models
[1189] Server: The generative AI model is trained using machine learning libraries such as TensorFlow based on the preprocessed data. This model replicates the specific steps of the craftsman's work and applies them to the robot.
[1190] 4. Robot Control
[1191] Terminal: The robot executes the collected work procedures based on the generative AI model. For example, it can use a general-purpose robot arm such as the UR5 or UR10 to perform precise movements.
[1192] 5. Emotion analysis method
[1193] Server: Receives the user's facial expressions and voice data from smart glasses and other devices, and performs emotion analysis using Microsoft Azure's Emotion API, etc. Based on the analysis results, provides feedback to the user in real time.
[1194] Specific examples
[1195] In a factory, a skilled craftsman performs welding work. During this process, the craftsman wears smart glasses that record his work steps in real time. The collected data is sent to a server where it is analyzed. A generative AI model learns the craftsman's work steps and generates new welding steps. This model is then used by a robotic arm to perform the work. Furthermore, an emotion engine analyzes the observer's emotional state, and if stress or doubt is recognized, the system provides appropriate feedback and corrections.
[1196] Prompt Sentence Examples
[1197] "Example of teaching AI new work procedures:
[1198] Detailed procedures for welding work that should be performed by factory robots
[1199] Important points to highlight in the video data
[1200] Expert instructions to be extracted from speech data
[1201] Specific methods for handling facial expression and voice data for emotion analysis
[1202] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1203] Step 1:
[1204] Terminal: The craftsman puts on the smart glasses and begins work. The smart glasses are equipped with a camera and microphone, which record video and audio data of the work in real time.
[1205] Input: Worker's workflow (visual and audio data)
[1206] Output: Recorded video and audio data
[1207] How it works: Smart glasses capture video and audio and collect data in real time.
[1208] Step 2:
[1209] Server: Stores the video and audio data received from the smart glasses in storage, and adds metadata (date, time, operation details, etc.) to the data.
[1210] Input: Video and audio data sent from smart glasses
[1211] Output: Data saved to storage (with metadata)
[1212] Specific operation: Stores data in a database (such as MongoDB) and organizes it by adding metadata.
[1213] Step 3:
[1214] Server: Analyzes the stored video and audio data and extracts the necessary information. Preprocesses the data using machine learning algorithms (e.g., TensorFlow).
[1215] Input: Stored video and audio data (with metadata)
[1216] Output: Preprocessed data (features extracted)
[1217] Specific operations: Split video data into frames and extract key points. Extract important reference phrases from audio data.
[1218] Step 4:
[1219] Server: Trains the generative AI model using the preprocessed data.
[1220] Input: Preprocessed data (features extracted)
[1221] Output: A trained generative AI model
[1222] How it works: Using TensorFlow, we train a model to learn the work procedures of craftsmen, using multi-layer perceptrons and convolutional neural networks (CNNs).
[1223] Step 5:
[1224] Terminal: The trained generative AI model is applied to the robot, which then executes the actual work steps.
[1225] Input: A trained generative AI model
[1226] Output: Robot performs work
[1227] Specific operation: Control a robot arm (e.g., UR5 or UR10) and perform tasks according to learned procedures.
[1228] Step 6:
[1229] Server: Receives user facial and voice data from smart glasses and other sensor devices and performs emotion analysis. Uses Microsoft Azure's Emotion API.
[1230] Input: User facial and voice data
[1231] Output: Emotional state analysis result
[1232] Specific operation: Analyzes user emotion data in real time using Microsoft Azure's Emotion API.
[1233] Step 7:
[1234] Server: Generates feedback based on the results of sentiment analysis and provides it to the user.
[1235] Input: Emotional state analysis result
[1236] Output: Feedback (adjustments to training, new instructions, etc.)
[1237] What it does: Depending on the user's emotional state, the system will adjust the intensity of their workout or suggest a break. Feedback is provided via voice and text.
[1238] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1239] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1240] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1241] [Fourth embodiment]
[1242] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1243] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1244] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1245] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1246] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1247] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1248] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1249] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1250] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1251] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1252] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1253] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1254] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1255] The present invention relates to a system that records the detailed work procedures of craftsmen and allows them to be learned by AI to pass on their skills. This system provides various means for effectively passing on the skills of craftsmen to the next generation. The following describes the system configuration and operation for specifically implementing the present invention.
[1256] Basic system configuration
[1257] 1. Data collection device
[1258] Devices: Includes devices such as smartphones and video cameras that record the craftsman's work procedures with video and photos, allowing the craftsman's detailed movements and techniques to be recorded in detail.
[1259] 2. Data Management Server
[1260] Server: A server for receiving and analyzing recorded video and photo data. The server works with a database to organize and store the recorded data. It also has software for data analysis and preprocesses the uploaded data.
[1261] 3. Generative AI Models
[1262] Server: The preprocessed data is used to train a generative AI model, which then learns the specific work steps of the craftsman and generates new work steps. A machine learning algorithm is used to create a highly accurate model across multiple steps.
[1263] 4. Virtual Reality (VR) Training Environment
[1264] Device: A VR system that uses a trained generative AI model to provide a training environment in a virtual space. The user wears a VR headset and performs training based on the generated steps in the virtual space.
[1265] 5. Successor Matching Platform
[1266] Server: A platform for effectively matching artisans with potential successors. It implements a matching algorithm that matches the skill sets of artisans and the interests of successors, and proposes optimal pairs.
[1267] Program processing overview
[1268] 1. Data collection procedure
[1269] User: The craftsman uses a smartphone or video camera to take videos and photos of the work procedure. Once the recording is complete, the data can be viewed through the application on the device and uploaded to the server.
[1270] 2. Data preprocessing and storage
[1271] Server: Receives uploaded video and photo data, stores them in storage, records metadata in a database, and performs preprocessing for analysis.
[1272] 3. Training the AI model
[1273] Server: Trains the generative AI model using the preprocessed data. It uses machine learning algorithms to learn the craftsman's work procedures and generate a highly accurate model.
[1274] 4. Providing a training environment
[1275] Device: The trained generative AI model is integrated into a VR system, allowing users to practice in a virtual space. The user puts on a VR headset and practices by following the steps generated by the AI model.
[1276] 5. Feedback and Ratings
[1277] Server: Analyzes the user's actions in the VR environment in real time and provides feedback, allowing the user to improve their skills.
[1278] 6. Use of Matching Platforms
[1279] Users: By accessing the platform, craftsmen register their skills, and successors input their interests and goals. The system then suggests optimal matches, and users can start communicating based on those matches.
[1280] Specific examples
[1281] 1. Data collection: Potters use their smartphones to videotape the process of kneading the clay, shaping it, and firing it, and also take photographs of each detailed step.
[1282] 2. Data preprocessing: The server splits the received video data into frames and adds metadata about the work done to each frame.
[1283] 3. Training the AI model: The server trains a generative AI model that learns specific steps in pottery making, enabling it to generate new steps.
[1284] 4. Providing a training environment: Users use a VR headset to practice pottery making in a virtual space, following the steps they learned. AI provides real-time feedback.
[1285] 5. Matching Platform: Young potters are matched with experienced artisans through the platform and begin online sessions.
[1286] In this way, the system provides a comprehensive solution for realizing the effective inheritance of traditional skills.
[1287] The processing flow will be explained below.
[1288] Step 1: Data collection
[1289] User: Uses a smartphone or video camera to capture video and photos of the craftsman's work steps.
[1290] Device: Press the start recording button to capture video and photos of the craftsman's work. The captured data is temporarily saved.
[1291] User: After shooting, the user checks the video and photo data through the application and enters metadata (e.g., work process, date, time, etc.).
[1292] Terminal: Compresses video and photo data and uploads it to a server via the Internet.
[1293] Step 2: Data Preprocessing
[1294] Server: Stores the received video and photo data in storage. Creates a new entry in the database and records metadata related to the uploaded video (such as the craftsman's name, the work performed, and the date and time of the shoot).
[1295] Server: Splits the video data into frames, assigns a timestamp to each frame, extracts important frames for each task, and deletes unnecessary frames.
[1296] Server: Converts image data into tensor format and audio data into spectrograms. Prepares the data for analysis.
[1297] Step 3: Training the AI model
[1298] Server: Inputs the preprocessed dataset into the generative AI model and starts training the model. The training process is done using a supervised learning algorithm.
[1299] Server: Monitors the loss function and accuracy during training every epoch, adjusting hyperparameters as learning progresses and optimizing the model.
[1300] Server: Stores the trained model and makes it available for training and simulation.
[1301] Step 4: Providing a training environment
[1302] Server: Integrates the trained generative AI model into the VR system, generating a virtual space and configuring it so that users can simulate tasks in that space.
[1303] Device: Using a VR headset and controller, the user enters the virtual space and practices the generated steps.
[1304] User: Wear a VR headset and follow the steps guided by the AI in the virtual space to complete the task. Check whether you can proceed as planned.
[1305] Step 5: Real-time feedback and evaluation
[1306] Server: Analyzes the work users do in the virtual environment in real time and determines which areas need improvement.
[1307] Server: Based on the analysis results, the server provides visual and audio feedback to the user, such as advice like "Take more time to shape this part."
[1308] Users: Receive feedback and then rework the work based on that feedback, thereby improving their skills.
[1309] Step 6: Use a matching platform
[1310] Server: Runs a matching algorithm based on the skill sets and interests of the successor candidate and the craftsman, and proposes the optimal combination.
[1311] Device: Notifies users of matches and displays their profile and contact information.
[1312] User: Logs in to the platform and checks the proposed matching results. The craftsman and successor communicate directly and create a specific training plan.
[1313] Through these steps, the system provides a comprehensive solution for effectively passing on artisan skills to the next generation.
[1314] Example 1
[1315] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1316] Many traditional skills today are being lost due to the aging of artisans and a lack of successors. In particular, if there is no established method for effectively passing on artisanal techniques and know-how to the next generation, those skills are at high risk of disappearing. Traditional methods, such as simply recording artisans' work procedures with video or photographs, make it difficult to pass on skills to the next generation. Furthermore, a lack of feedback and a lack of training environments hinder the transfer of skills.
[1317] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1318] In this invention, the server includes means for recording the work procedures of craftsmen using a recording device, means for receiving and saving the recorded data, means for preprocessing the saved data, means for training a generative AI model using the preprocessed data, means for integrating the trained generative AI model into a virtual reality system, means for effectively matching craftsmen with potential successors, means for providing a training environment to users using the virtual reality system, and means for providing real-time feedback to users during training. This makes it possible to effectively pass on craftsman skills to the next generation and provide detailed, practical training using virtual reality.
[1319] A "craftsman" is a professional who has specific skills and techniques and uses them to create products or works.
[1320] A "work procedure" is a set of specific steps or processes for performing a particular task.
[1321] "Recording device" refers to a device for recording work procedures as video or photographs.
[1322] "Data" refers to video, photographs, and other information captured by recording devices.
[1323] "Receiving" refers to the process of transferring and receiving recorded data to a server or other system.
[1324] "Storage" is the process of storing received data in storage so that it can be used later.
[1325] "Preprocessing" refers to the process of performing a series of operations and filtering to transform raw data into a form that is easier to analyze.
[1326] A "generative AI model" is an artificial intelligence model generated based on training data to reproduce and analyze specific work procedures and actions.
[1327] "Training" refers to the process of teaching a generative AI model so that it learns specific task steps.
[1328] A "virtual reality system" is a system that allows users to experience and practice actual work procedures in a virtual space.
[1329] "Matching" refers to the process of connecting craftsmen with potential successors based on certain criteria.
[1330] A "training environment" is an environment in which a user can practice or perform exercises through a virtual reality system.
[1331] "Feedback" refers to the evaluation and improvement advice provided in real time to the user's actions and results during training.
[1332] This invention relates to a system that records the work procedures of craftsmen in detail and has AI learn from them to pass on skills. The mode for carrying out the invention is configured as follows.
[1333] 1. Data collection device
[1334] Devices: Smartphones and video cameras are used to record the craftsman's work procedures. This allows the craftsman's detailed movements and techniques to be recorded in detail using video and photographs. For example, a potter can use the video function on his smartphone to record the process of shaping clay, and the position and movement of his hands can be recorded in photographs.
[1335] 2. Uploading and saving data
[1336] User: After capturing the data, the user can check it through a dedicated smartphone application and upload it to the server. This upload process is performed by pressing a dedicated button within the app.
[1337] Server: Receives uploaded data and stores it in cloud storage. After storage, the data is automatically labeled and organized.
[1338] 3. Data Preprocessing
[1339] Server: Analyzes the stored video and photo data and splits it into frames. Each frame is then tagged with metadata that describes the task, including hand position and movement.
[1340] 4. Training the AI model
[1341] Server: Trains a generative AI model using the preprocessed data. It uses machine learning algorithms to learn specific work procedures performed by artisans and generate highly accurate models. For example, by learning pottery work procedures, an AI model can be created that can generate new work procedures.
[1342] 5. Creating a training environment
[1343] Device: Integrate the trained generative AI model into a virtual reality (VR) system so that when a user puts on a VR headset, they can follow the learned steps in the virtual space and practice.
[1344] 6. Feedback and Ratings
[1345] Server: Analyzes the work data performed by the user in the VR environment in real time and provides feedback. The AI model analyzes the user's movements and suggests accurate movements and areas for improvement.
[1346] 7. Successor matching
[1347] Server: Provides a platform for effectively matching craftsmen with potential successors. Proposes optimal matches based on the craftsman's skill set and the successor's interests.
[1348] Prompt Sentence Examples
[1349] "Create an AI model that learns specific steps from a video of a craftsman kneading clay. Analyze the movements of each step in detail and then provide practical training in a VR training environment."
[1350] In this way, this system realizes the effective inheritance of traditional skills and provides an innovative approach that utilizes AI and virtual reality technology, making it possible to efficiently pass on artisan skills to the next generation.
[1351] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1352] Step 1:
[1353] Data collection
[1354] Input: A craftsman films the work procedure using a smartphone or video camera.
[1355] Specific actions: The user, a craftsman, uses a smartphone or video camera to record video and photos of work procedures, such as shaping and firing pottery. Since it is necessary to capture detailed actions from various angles, multiple videos and images are taken.
[1356] Output: Video and photo data showing the work procedure.
[1357] Step 2:
[1358] Data upload
[1359] Input: Completed video and photo data.
[1360] How it works: The user checks the captured data through a dedicated smartphone application and uploads it to the server. By pressing a dedicated button on the app, the data is sent to the cloud.
[1361] Output: Video and photo data uploaded to the server.
[1362] Step 3:
[1363] Data reception and storage
[1364] Input: Uploaded video and photo data.
[1365] What it does: The server receives the uploads and stores them in cloud storage, automatically labeling and organizing the data at the same time.
[1366] Output: Structured data stored in cloud storage.
[1367] Step 4:
[1368] Data Preprocessing
[1369] Input: Video and photo data stored in cloud storage.
[1370] How it works: The server divides the data into frames and adds metadata (information about the hand's position and movement) to each frame. Image analysis algorithms are used to identify the hand's movement and position.
[1371] Output: A frame-by-frame dataset with metadata.
[1372] Step 5:
[1373] Training an AI model
[1374] Input: The preprocessed dataset.
[1375] How it works: The server uses machine learning algorithms to train a generative AI model, for example, for a specific task (such as shaping pottery) and then replicates the movements and positions of each step.
[1376] Output: A trained generative AI model.
[1377] Step 6:
[1378] Creating a training environment
[1379] Input: A trained generative AI model.
[1380] How it works: The device integrates the trained generative AI model into a virtual reality (VR) system, so that when the user puts on the VR headset, the work steps are accurately reproduced in the virtual space.
[1381] Output: Model integrated into a virtual reality system.
[1382] Step 7:
[1383] VR training
[1384] Input: A generative AI model integrated into a virtual reality system.
[1385] Specific actions: The user puts on a VR headset and performs exercises in the VR space by following the steps generated by the AI model. For example, the user recreates the action of molding clay in a virtual space and practices the steps.
[1386] Output: User's practice data.
[1387] Step 8:
[1388] Feedback and Ratings
[1389] Input: User's practice data.
[1390] Specific movements: The server analyzes the user's practice data in real time and provides feedback. The AI model analyzes the user's movements and provides a detailed evaluation of the exact movements and areas for improvement.
[1391] Output: Real-time feedback provided to the user.
[1392] Step 9:
[1393] Successor matching
[1394] Input: Craftsman skillset and potential successor information.
[1395] How it works: The server proposes optimal matches based on the craftsman's skill set and the potential successor's interests and goals. A matching algorithm is used to connect craftsmen and potential successors.
[1396] Output: Matching results between craftsmen and potential successors.
[1397] (Application example 1)
[1398] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1399] In today's world, where it is difficult to pass on the skills of craftsmen, there is a need for a method to reliably record work procedures and pass them on to the next generation. Furthermore, particularly in factory production environments, transferring the skills of skilled craftsmen to robots would lead to improved efficiency and quality, but no concrete method for doing so has been established. Furthermore, there is an issue of a gap between training in a virtual environment and actual robot operation.
[1400] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1401] In this invention, the server includes means for recording the work procedures of craftsmen, means for receiving and analyzing the recorded data, means for training a generative AI model using the analyzed data, means for providing a pseudo-training environment using the trained generative AI model, means for matching craftsmen with potential successors, and means for applying the trained generative AI model to a robot and having it execute the work procedures. This allows for the efficient transfer of craftsman skills and enables factory robots to perform advanced tasks.
[1402] A "craftsman" is someone who has specific skills or techniques and uses those skills to create products or crafts.
[1403] "Work procedure" refers to the series of steps and processes that a craftsman follows to create a product using his skilled techniques and skills.
[1404] "Means of recording" refers to devices or methods for recording a craftsman's work procedures as video or photographs using a smartphone, video camera, etc.
[1405] "Means of analysis" refers to the techniques and equipment used to analyze recorded video and photographs and extract important technical and skill elements from them.
[1406] A "generative AI model" is an AI that uses machine learning algorithms to learn from collected data and imitate or improve the work procedures of craftsmen.
[1407] A "simulated training environment" is a place or system that uses virtual reality (VR) and simulation to recreate the work performed by craftsmen, allowing them to train in an environment that is similar to the actual work environment.
[1408] "Matching methods" refer to systems and methods that match the skill sets of craftsmen with the interests and abilities of potential successors, and then select and connect the most suitable craftsmen and successors.
[1409] A "robot" is a mechanical device that can actually carry out work procedures learned using an AI model.
[1410] MODE FOR CARRYING OUT THE INVENTION
[1411] This invention is a system that records and analyzes the work procedures of craftsmen in detail, and then uses a generative AI model to have a robot carry out the work. The system configuration and operation are explained below.
[1412] Data collection methods
[1413] The device uses a smartphone or video camera to record video and photos of the craftsman's work procedures, allowing the craftsman's detailed movements and techniques to be recorded in detail.
[1414] Data analysis
[1415] The server receives and analyzes the recorded video and photo data, including splitting the footage and photos into frames and analyzing the movements of the craftsmen, using image processing techniques and machine learning algorithms.
[1416] Training generative AI models
[1417] The server trains a generative AI model based on the analyzed data, building a model that mimics or improves the craftsman's work procedures. The machine learning algorithms used include deep learning and reinforcement learning.
[1418] Providing a simulated training environment
[1419] The device provides a virtual reality (VR) environment using the trained generative AI model. In this environment, users can practice based on the learned procedures in a virtual space. Users wear a VR headset and train by following the instructions generated by the AI model.
[1420] Application to robots
[1421] The server integrates the trained generative AI model into the robot control system, allowing the robot to actually execute the craftsman's work steps that the AI model has learned. The robot can be an industrial robot arm or an autonomous robot.
[1422] Matching System
[1423] The server provides a platform for matching artisans with potential successors. This platform matches the skill sets of artisans with the interests and abilities of successors to create the optimal pairing. Users access the platform and register their own skills and interests to be matched.
[1424] Examples and prompts
[1425] For example, a skilled worker assembling a product uses a camera to record the steps of installing parts. The server analyzes the captured video, divides the actions into frames, and extracts important parts. The generative AI model uses this data to learn and formulates assembly steps. It then instructs the robot arm to execute the steps. The entire process is instructed to the AI model with the following prompt:
[1426] "Analyze this video frame by frame and generate step-by-step assembly instructions."
[1427] In this way, the present invention provides a series of means for effectively transferring the skills of craftsmen, making it possible to realize advanced work using factory robots.
[1428] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1429] Step 1:
[1430] The terminal uses a smartphone or video camera to record the craftsman's work procedures with video and photos. The input data are the videos and photos of the craftsman's movements and techniques. The output data are the video and photo files stored on the recording device.
[1431] Step 2:
[1432] The server receives the recorded video and photo data, divides it into frames, and performs analysis. The input data are the video and photo files obtained in step 1. The output data are the analyzed frame-by-frame motion data and its metadata. Specifically, image processing techniques are used to extract important motion points from each frame.
[1433] Step 3:
[1434] The server uses the analyzed data to train a generative AI model. The input data is the frame-by-frame motion data and metadata obtained in step 2. The output data is the trained generative AI model. Specifically, it uses machine learning algorithms (e.g., deep learning and reinforcement learning) to learn engineering features from the analyzed data.
[1435] Step 4:
[1436] The device provides a virtual reality (VR) environment using a trained generative AI model. The input data is the trained generative AI model. The output data is a virtual training environment accessible to users. Specifically, users wear a VR headset and practice in a virtual space by following the steps of a craftsman generated by the AI model.
[1437] Step 5:
[1438] The server integrates the trained generative AI model into the robot control system and has the robot execute the work procedure. The input data is the trained generative AI model. The output data is the specific operation instructions to be executed by the robot. Specifically, the robot arm is controlled to perform tasks such as assembly and processing.
[1439] Step 6:
[1440] The server provides a platform for matching craftsmen with potential successors. The input data is the skill set of the craftsman and the interests and abilities of the potential successor. The output data is the pairing results of the optimal craftsman and potential successor. Specifically, users log in to the platform and enter the necessary information, and the best match is suggested.
[1441] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1442] This invention relates to a system that records the detailed work procedures of craftsmen and allows them to be trained by AI to pass on their skills, and also combines this with an emotion engine that recognizes the user's emotions. This system provides various means for effectively passing on craftsmen's skills to the next generation. The system configuration and operation for specifically implementing this invention are described below.
[1443] Basic system configuration
[1444] 1. Data collection device
[1445] Devices: Includes devices such as smartphones and video cameras that record the craftsman's work procedures with video and photos, allowing the craftsman's detailed movements and techniques to be recorded in detail.
[1446] 2. Data Management Server
[1447] Server: A server for receiving and analyzing recorded video and photo data. The server works with a database to organize and store the recorded data. It also has software for data analysis and preprocesses the uploaded data.
[1448] 3. Generative AI Models
[1449] Server: The preprocessed data is used to train a generative AI model, which then learns the specific work steps of the craftsman and generates new work steps. A machine learning algorithm is used to create a highly accurate model across multiple steps.
[1450] 4. Virtual Reality (VR) Training Environment
[1451] Device: A VR system that uses a trained generative AI model to provide a training environment in a virtual space. The user wears a VR headset and performs training based on the generated steps in the virtual space.
[1452] 5. Successor Matching Platform
[1453] Server: A platform for effectively matching artisans with potential successors. It implements a matching algorithm based on the artisan's skill set and the successor's interests, and proposes optimal pairs.
[1454] 6. Emotion Engine
[1455] Server: Equipped with an emotion engine that analyzes the user's facial expressions and voice data to recognize their emotions. Analyzes the user's emotional data in real time and reflects it in training content and feedback.
[1456] Program processing overview
[1457] 1. Data collection procedure
[1458] User: The craftsman uses a smartphone or video camera to take videos and photos of the work procedure. Once the recording is complete, the data can be viewed through the application on the device and uploaded to the server.
[1459] 2. Data preprocessing and storage
[1460] Server: Receives uploaded video and photo data, stores them in storage, records metadata in a database, and performs preprocessing for analysis.
[1461] 3. Training the AI model
[1462] Server: Trains the generative AI model using the preprocessed data. It uses machine learning algorithms to learn the craftsman's work procedures and generate a highly accurate model.
[1463] 4. Providing a training environment
[1464] Device: The trained generative AI model is integrated into a VR system, allowing users to practice in a virtual space. The user puts on a VR headset and practices by following the steps generated by the AI model.
[1465] 5. Feedback and Ratings
[1466] Server: Analyzes the user's actions in the VR environment in real time and provides feedback, allowing the user to improve their skills.
[1467] 6. Emotional Recognition
[1468] Server: Captures the user's facial expressions and voice in real time, and the emotion engine analyzes the data. Based on the user's emotional state, the system adjusts the training progress.
[1469] Device: Receives feedback from the emotion engine. For example, if the user is feeling stressed, the system adjusts the training to reduce stress.
[1470] Users: By receiving appropriate feedback and adjustments based on their emotional state, they can train more efficiently and with greater focus.
[1471] Specific examples
[1472] 1. Data collection: Potters use their smartphones to videotape the process of kneading the clay, shaping it, and firing it, and also take photographs of each detailed step.
[1473] 2. Data preprocessing: The server splits the received video data into frames and adds metadata about the work done to each frame.
[1474] 3. Training the AI model: The server trains a generative AI model that learns specific steps in pottery making, enabling it to generate new steps.
[1475] 4. Providing a training environment: Users use a VR headset to practice pottery making in a virtual space, following the steps they learned. AI provides real-time feedback.
[1476] 5. Emotion recognition and adjustment: If the user is feeling stressed during training, the emotion engine will recognize this and the system will adjust the training content, for example by suggesting a short break to reduce the user's stress.
[1477] 6. Matching Platform: Young potters are matched with experienced craftsmen through the platform and begin online sessions.
[1478] In this way, the system effectively passes on the skills of craftsmen to the next generation and supports skill improvement while taking into consideration the user's emotional state.
[1479] The processing flow will be explained below.
[1480] Step 1: Data collection
[1481] User: Using a smartphone or video camera, the user takes videos and photographs of the craftsman's work steps, allowing the craftsman to record his work in detail.
[1482] Device: Press the start recording button to capture video and photos of the craftsman's work. The captured data is temporarily saved.
[1483] User: After shooting, check the video and photo data through the application and enter metadata (process, date, time, etc.).
[1484] Terminal: Compresses video and photo data and uploads it to a server via the Internet.
[1485] Step 2: Data Preprocessing
[1486] Server: Stores the received video and photo data in storage. Creates a new entry in the database and records metadata related to the uploaded video (such as the craftsman's name, the work performed, and the date and time of the shoot).
[1487] Server: Splits the video data into frames, assigns a timestamp to each frame, extracts important frames for analysis, and removes unnecessary frames.
[1488] Server: Converts the stored image data into tensor format and audio data into spectrograms, thereby adjusting the data format to be analyzable by machine learning algorithms.
[1489] Step 3: Training the AI model
[1490] Server: Inputs the preprocessed dataset into the generative AI model and begins training the model. Using a supervised learning algorithm, it learns the craftsman's work procedures in detail.
[1491] Server: Monitors the loss function and accuracy every epoch during the training process, adjusting hyperparameters as training progresses to optimize model performance.
[1492] Server: Stores the trained generative AI model for later use in training and simulation.
[1493] Step 4: Providing a training environment
[1494] Server: Integrates the trained generative AI model into the VR system, allowing users to simulate tasks in a virtual space.
[1495] Device: The user wears a VR headset and follows the steps generated by the AI model in a virtual space. The user uses the VR controller to recreate the tasks in the virtual environment.
[1496] User: Enter the VR environment and follow the steps guided by the AI model. Practice while checking whether the steps are progressing as expected.
[1497] Step 5: Real-time feedback and evaluation
[1498] Server: Analyzes the work users do in the virtual environment in real time and determines which areas need improvement.
[1499] Server: Based on the analysis results, the server provides visual and audio feedback to the user, such as advice like "Take more time to shape this part."
[1500] Users: Receive feedback and then rework the work based on that feedback, thereby improving their skills.
[1501] Step 6: Recognize emotions
[1502] Server: Captures the user's facial expressions and voice in real time, and the emotion engine analyzes the data to recognize the user's emotional state (stress level and concentration state).
[1503] Device: Receives feedback from the emotion engine. If the user is feeling stressed, the system can adjust the training, for example suggesting a short break.
[1504] Users: By receiving appropriate feedback and adjustments based on their emotional state, they can train more efficiently and with greater focus.
[1505] Step 7: Use a matching platform
[1506] Server: Runs a matching algorithm based on the skill sets and interests of the successor candidate and the craftsman, and proposes the optimal combination.
[1507] Device: Notifies users of matches and displays their profile and contact information.
[1508] User: Logs in to the platform and checks the proposed matching results. The craftsman and successor communicate directly and create a specific training plan.
[1509] Through these steps, the system effectively passes on the craftsman's skills to the next generation and supports skill improvement while taking into account the user's emotional state.
[1510] Example 2
[1511] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1512] In today's aging society, effectively passing on the skills of skilled craftsmen to the next generation is an important issue. However, traditional methods of skill transfer are inefficient, as they make it difficult to fully convey the finer details of the skills and the craftsman's know-how. Furthermore, stress and a decline in motivation for the successor during the skill transfer process can also be an issue.
[1513] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for recording the work procedures of the craftsman with video and photographs, a means for receiving the recorded data, preprocessing and analyzing the data, and a means for training a machine learning model using the analyzed data. This makes it possible to record and analyze the specific work procedures and know-how of the craftsman with high accuracy.
[1514] Furthermore, this invention includes a means for providing a virtual training environment using a trained machine learning model, a means for matching craftsmen with potential successors, and a means for recognizing the user's emotions in real time and adjusting the training content based on those emotions. This increases the efficiency of skill transfer, prevents stress and a decline in motivation for successors, and enables smooth skill improvement.
[1515] A "craftsman" is a professional who has specific skills and techniques and creates things based on those skills.
[1516] "Work procedures" refer to the specific steps and methods of operation that craftsmen use to perform work.
[1517] "Video" refers to video data captured using a camera or other device.
[1518] A "photograph" refers to still image data captured at a specific moment with a camera or other device.
[1519] "Means of recording" refers to means for recording video or still images using devices such as video cameras or smartphones.
[1520] "Means for receiving, pre-processing and analyzing data" refers to means by which the server receives uploaded video and photo data and performs the data processing necessary to analyze them.
[1521] A "machine learning model" refers to an algorithm that uses AI techniques to learn from a specific dataset and make predictions or generation decisions.
[1522] "Means for providing a virtual training environment" refers to means for providing an environment in which a user can recreate and practice the work procedures of a craftsman in a virtual space.
[1523] "Potential successors" refer to people who wish to inherit the skills of a particular craftsman.
[1524] "Matching means" refers to algorithms and platforms that efficiently connect artisans with potential successors.
[1525] "Means for recognizing emotions in real time" refers to means for analyzing the user's facial expressions and voice data to identify their current emotional state.
[1526] "Means for adjusting training content" refers to means for adaptively changing training content in a virtual training environment based on the user's emotional state.
[1527] The present invention is a system that effectively transfers skills by recording the work procedures of craftsmen in detail and building a generative AI model based on the records. This system has the ability to recognize the user's emotions in real time and dynamically adjust the training content. The following describes the system configuration and operation for specifically implementing the present invention.
[1528] Basic system configuration
[1529] 1. Data collection device
[1530] Devices: Recording devices such as smartphones and video cameras are used. These devices record the craftsman's work procedures as videos and photographs, allowing the craftsman's detailed movements and techniques to be recorded in detail.
[1531] 2. Data Management Server
[1532] Server: This is the server that receives and analyzes the recorded video and photo data. The server connects to a database to organize and store the recorded data. It also has software for data analysis and preprocesses the uploaded data. Specifically, it uses tools such as Python and OpenCV to segment the data and add metadata.
[1533] 3. Generative AI Models
[1534] Server: The generative AI model is trained using the preprocessed data. This allows it to learn the specific work procedures of craftsmen and generate new work procedures. A machine learning algorithm is used to create a highly accurate model across multiple steps. TensorFlow and PyTorch are suitable libraries to use.
[1535] 4. Virtual Reality (VR) Training Environment
[1536] Device: A VR system is used that uses a trained generative AI model to provide a training environment in a virtual space. The user wears a VR headset (e.g., Oculus Rift or HTC Vive) and performs training based on the generated steps in the virtual space. The VR environment is built using Unity or Unreal Engine.
[1537] 5. Successor Matching Platform
[1538] Server: Provides a platform for effectively matching artisans with potential successors. Implements a matching algorithm based on the artisan's skill set and the successor's interests, and proposes optimal pairs.
[1539] 6. Emotion Engine
[1540] Server: Equipped with an emotion engine that analyzes the user's facial expressions and voice data to recognize emotions. Analyzes the user's emotional data in real time and reflects it in training content and feedback. Emotion analysis is performed using Microsoft Azure Cognitive Services and the Affectiva SDK.
[1541] Specific examples
[1542] 1. Data collection: Potters use their smartphones to videotape the process of kneading the clay, shaping it, and firing it, and also take photographs of each detailed step.
[1543] 2. Data preprocessing: The server splits the received video data into frames and adds metadata about the work done to each frame.
[1544] 3. Training the AI model: The server trains a generative AI model that learns specific steps in pottery making, enabling it to generate new steps.
[1545] 4. Providing a training environment: Users use a VR headset (e.g., Oculus Rift) to practice pottery making in a virtual space, following the steps they learned. AI provides real-time feedback.
[1546] 5. Emotion recognition and adjustment: If the user is feeling stressed during training, the emotion engine will recognize this and the system will adjust the training content, for example by suggesting a short break to reduce the user's stress.
[1547] 6. Matching Platform: Young potters are matched with experienced craftsmen through the platform and begin online sessions.
[1548] Example prompt
[1549] Example prompt: "I would like to learn the steps of pottery making. I would like to provide videos of each step of the pottery making process, from mixing the clay to molding and firing, and train in a virtual space."
[1550] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1551] Step 1:
[1552] The user uses a smartphone or video camera to record video and photos of the craftsman's work procedures. Specifically, the user records the steps of the potter kneading the clay, shaping it, and firing it. The input data are video and photo files, which are then uploaded to a dedicated application on the device. The output data are verified video and photo files.
[1553] Step 2:
[1554] The device reviews the video and photo data received from the user through the application and edits it as needed, such as cutting out unnecessary parts and labeling important parts. The input data are video and photo files, and the output data are the edited video and photo files.
[1555] Step 3:
[1556] The device uploads the edited video and photo files to the server by sending them to a specified URL on the server using an HTTP request. The input data are the edited video and photo files, and the output data are the video and photo data uploaded to the server.
[1557] Step 4:
[1558] The server receives the uploaded video and photo data and stores them in storage. Specifically, it writes the received files to disk storage and records the metadata in a database. The input data is the uploaded video and photo data, and the output data is the stored video and photo data and their metadata.
[1559] Step 5:
[1560] The server analyzes and preprocesses the stored video and photo data. Specifically, it uses OpenCV to split the video data into frames and add metadata about the work done to each frame. It also performs noise removal and resolution adjustment. The input data is the stored video and photo data, and the output data is the preprocessed video and photo data.
[1561] Step 6:
[1562] The server uses the preprocessed data to train a generative AI model. Specifically, it uses TensorFlow and PyTorch to run machine learning algorithms to learn the craftsman's work procedures. The input data is the preprocessed video and photo data, and the output data is the trained generative AI model.
[1563] Step 7:
[1564] The device integrates the trained generative AI model into the VR system. Specifically, it uses Unity or Unreal Engine to recreate the steps generated by the model in a virtual space so that the user can experience them. The input data is the trained generative AI model, and the output data is the work steps recreated in the VR environment.
[1565] Step 8:
[1566] The user wears a VR headset and performs practical training in a virtual space by following the steps provided by the generative AI model. Specifically, the user practices pottery molding in a virtual space. The input data is the work steps reproduced in the VR environment, and the output data is the motion data of the work performed by the user.
[1567] Step 9:
[1568] The server analyzes the user's behavior data in real time and provides feedback. Specifically, it points out errors in behavior and suggests ways to correct them. It also performs real-time data processing using Apache Kafka and Apache Flink. The input data is the user's behavior data, and the output data is the feedback message.
[1569] Step 10:
[1570] The server captures the user's emotional data in real time and analyzes it using an emotion engine. Specifically, it uses Microsoft Azure Cognitive Services and the Affectiva SDK to recognize emotions from the user's facial expressions and voice. The input data is the user's facial expression and voice data, and the output data is the emotion analysis results.
[1571] Step 11:
[1572] The device dynamically adjusts the training content based on the results obtained from the emotion engine. Specifically, if the user feels stressed, the system will reduce the training content or suggest taking a break. The input data is the emotion analysis results, and the output data is the adjusted training content.
[1573] Step 12:
[1574] The server operates a platform that matches artisans with potential successors. Specifically, it records the skill sets of artisans and the interests of successors in a database and uses a recommendation algorithm to suggest optimal pairs. The input data are the skills and interests of artisans and successors, and the output data are the proposed matching pairs.
[1575] (Application example 2)
[1576] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1577] The transfer of craftsmanship skills requires specific skills and experience, which is time-consuming and costly. Furthermore, conventional methods have difficulty providing feedback that takes into account the craftsman's emotions and state. Furthermore, for a robot to accurately learn and execute these skills, advanced data collection and analysis techniques are required. Therefore, there is a need for a system that can efficiently and effectively transfer craftsmanship skills and enable robots to execute skills and provide feedback that reflects the user's emotions.
[1578] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1579] In this invention, the server includes means for recording the work procedures of craftsmen with video and photographs, means for receiving and analyzing the recorded data, means for training a generative AI model using the analyzed data, means for providing a simulated training environment using the trained generative AI model, means for matching craftsmen with potential successors, means for collecting work procedure data in real time from craftsmen wearing smart glasses, means for controlling a robot using the collected data to execute the learned work procedures, and means for recognizing the user's emotional state using emotion analysis means and providing feedback. This allows the craftsman's work procedures to be recorded in detail, enabling the robot to learn and execute the procedures, and providing feedback based on the user's emotional state in real time.
[1580] "Means for video and photographic recording of craftsmen's work procedures" is a general term for video cameras and camera-equipped devices used to record in detail the series of tasks and operations performed by craftsmen.
[1581] "Means for receiving and analyzing recorded data" refers to a combination of software and hardware for collecting said video and photographic data and analyzing it using appropriate algorithms.
[1582] "Means for training a generative AI model" is a general term for software and computer equipment that applies machine learning algorithms based on the analyzed data to train an AI model.
[1583] "Means for providing a pseudo-training environment using a trained generative AI model" refers to a system that uses a trained AI model to provide a training environment in a simulated format to users.
[1584] The "means for matching craftsmen with potential successors" refers to an algorithm and system that compares the skills of craftsmen with the interests and abilities of potential successors and proposes the optimal combination.
[1585] "Means for collecting work procedure data in real time from craftsmen wearing smart glasses" refers to technology and devices that use smart glasses to capture work procedures from the perspective of craftsmen in real time and collect them as data.
[1586] The "means for controlling the robot using collected data to execute the learned work procedure" is a system for causing the robot to execute the work procedure based on data collected from smart glasses and other sensor devices.
[1587] "Means for recognizing a user's emotional state using emotion analysis means and providing feedback" refers to software that analyzes a user's facial expressions and voice data to recognize their emotional state, and a system for providing feedback to the user based on that.
[1588] The system for implementing this invention records the work procedures of craftsmen in detail, trains a generative AI model based on that data, and applies it to a robot to achieve skill transfer and efficient work execution. It also includes an emotion analysis means for analyzing the user's emotional state in real time and providing appropriate feedback.
[1589] System Configuration
[1590] 1. Data collection device
[1591] Terminal: The craftsman wears smart glasses that capture the work procedure in real time. The smart glasses are equipped with a camera and a microphone to record both visual and audio data, which is then sent to a server via the Internet.
[1592] 2. Data Management Server
[1593] Server: Receives, stores, and analyzes recorded video and audio data. The server works with a database (e.g., MongoDB) to organize and store the data. It also preprocesses the data using machine learning algorithms such as TensorFlow for data analysis.
[1594] 3. Generative AI Models
[1595] Server: The generative AI model is trained using machine learning libraries such as TensorFlow based on the preprocessed data. This model replicates the specific steps of the craftsman's work and applies them to the robot.
[1596] 4. Robot Control
[1597] Terminal: The robot executes the collected work procedures based on the generative AI model. For example, it can use a general-purpose robot arm such as the UR5 or UR10 to perform precise movements.
[1598] 5. Emotion analysis method
[1599] Server: Receives the user's facial expressions and voice data from smart glasses and other devices, and performs emotion analysis using Microsoft Azure's Emotion API, etc. Based on the analysis results, provides feedback to the user in real time.
[1600] Specific examples
[1601] In a factory, a skilled craftsman performs welding work. During this process, the craftsman wears smart glasses that record his work steps in real time. The collected data is sent to a server where it is analyzed. A generative AI model learns the craftsman's work steps and generates new welding steps. This model is then used by a robotic arm to perform the work. Furthermore, an emotion engine analyzes the observer's emotional state, and if stress or doubt is recognized, the system provides appropriate feedback and corrections.
[1602] Prompt Sentence Examples
[1603] "Example of teaching AI new work procedures:
[1604] Detailed procedures for welding work that should be performed by factory robots
[1605] Important points to highlight in the video data
[1606] Expert instructions to be extracted from speech data
[1607] Specific methods for handling facial expression and voice data for emotion analysis
[1608] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1609] Step 1:
[1610] Terminal: The craftsman puts on the smart glasses and begins work. The smart glasses are equipped with a camera and microphone, which record video and audio data of the work in real time.
[1611] Input: Worker's workflow (visual and audio data)
[1612] Output: Recorded video and audio data
[1613] How it works: Smart glasses capture video and audio and collect data in real time.
[1614] Step 2:
[1615] Server: Stores the video and audio data received from the smart glasses in storage, and adds metadata (date, time, operation details, etc.) to the data.
[1616] Input: Video and audio data sent from smart glasses
[1617] Output: Data saved to storage (with metadata)
[1618] Specific operation: Stores data in a database (such as MongoDB) and organizes it by adding metadata.
[1619] Step 3:
[1620] Server: Analyzes the stored video and audio data and extracts the necessary information. Preprocesses the data using machine learning algorithms (e.g., TensorFlow).
[1621] Input: Stored video and audio data (with metadata)
[1622] Output: Preprocessed data (features extracted)
[1623] Specific operations: Split video data into frames and extract key points. Extract important reference phrases from audio data.
[1624] Step 4:
[1625] Server: Trains the generative AI model using the preprocessed data.
[1626] Input: Preprocessed data (features extracted)
[1627] Output: A trained generative AI model
[1628] How it works: Using TensorFlow, we train a model to learn the work procedures of craftsmen, using multi-layer perceptrons and convolutional neural networks (CNNs).
[1629] Step 5:
[1630] Terminal: The trained generative AI model is applied to the robot, which then executes the actual work steps.
[1631] Input: A trained generative AI model
[1632] Output: Robot performs work
[1633] Specific operation: Control a robot arm (e.g., UR5 or UR10) and perform tasks according to learned procedures.
[1634] Step 6:
[1635] Server: Receives user facial and voice data from smart glasses and other sensor devices and performs emotion analysis. Uses Microsoft Azure's Emotion API.
[1636] Input: User facial and voice data
[1637] Output: Emotional state analysis result
[1638] Specific operation: Analyzes user emotion data in real time using Microsoft Azure's Emotion API.
[1639] Step 7:
[1640] Server: Generates feedback based on the results of sentiment analysis and provides it to the user.
[1641] Input: Emotional state analysis result
[1642] Output: Feedback (adjustments to training, new instructions, etc.)
[1643] What it does: Depending on the user's emotional state, the system will adjust the intensity of their workout or suggest a break. Feedback is provided via voice and text.
[1644] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1645] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1646] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1647] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1648] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1649] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1650] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1651] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1652] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1653] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1654] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1655] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1656] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1657] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1658] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1659] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1660] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1661] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1662] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1663] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1664] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1665] The following is further disclosed regarding the above embodiment.
[1666] (Claim 1)
[1667] a means of recording the work procedures of craftsmen through video and photographs;
[1668] means for receiving and analyzing the recorded data;
[1669] means for training a generative AI model using the analyzed data;
[1670] A means for providing a pseudo-training environment using the trained generative AI model;
[1671] A system that includes a means of matching craftsmen with potential successors.
[1672] (Claim 2)
[1673] 2. The system according to claim 1, further comprising means for reproducing a work procedure of a craftsman in a virtual space, and for a user to perform training in the virtual space.
[1674] (Claim 3)
[1675] 10. The system of claim 1, further comprising means for providing real-time feedback to the user while training in the virtual space.
[1676] "Example 1"
[1677] (Claim 1)
[1678] a means for recording the work procedures of the craftsman with a recording device;
[1679] means for receiving and storing said recorded data;
[1680] means for preprocessing the stored data;
[1681] means for training a generative AI model using the preprocessed data;
[1682] means for integrating the trained generative AI model into a virtual reality system;
[1683] A means of effectively matching craftsmen with potential successors,
[1684] means for providing a training environment to a user using the virtual reality system;
[1685] means for providing real-time feedback to the user during said training;
[1686] A system including:
[1687] (Claim 2)
[1688] 2. The system according to claim 1, further comprising means for recording the work procedures of the craftsman in detail with a recording device, and for the user to practice according to the learned procedures in the virtual space.
[1689] (Claim 3)
[1690] 10. The system according to claim 1, further comprising means for analyzing the user's practice data in real time and providing feedback during training in the virtual space.
[1691] "Application Example 1"
[1692] (Claim 1)
[1693] a means for recording the work procedures of the craftsmen;
[1694] means for receiving and analyzing the recorded data;
[1695] means for training a generative AI model using the analyzed data;
[1696] A means for providing a pseudo-training environment using the trained generative AI model;
[1697] A means of matching craftsmen with potential successors,
[1698] A system including a means for applying the trained generative AI model to a robot and causing it to execute a work procedure.
[1699] (Claim 2)
[1700] A means for reproducing the work procedures of a craftsman in a virtual space and for users to train in the virtual space;
[1701] 2. The system according to claim 1, further comprising means for reflecting training results in the virtual space in robot operation.
[1702] (Claim 3)
[1703] means for providing real-time feedback to a user during training in the virtual space;
[1704] The system of claim 1 further comprising means for using said feedback information to improve the behavior of the robot.
[1705] "Example 2: Combining Emotion Engines"
[1706] (Claim 1)
[1707] a means of recording the work procedures of craftsmen through video and photographs;
[1708] means for receiving, pre-processing and analyzing said recorded data;
[1709] means for training a machine learning model using the analyzed data;
[1710] means for providing a virtual training environment using the trained machine learning model;
[1711] A means of matching craftsmen with potential successors,
[1712] The system includes a means for recognizing a user's emotions in real time and adjusting training content based on those emotions.
[1713] (Claim 2)
[1714] 2. The system according to claim 1, further comprising means for reproducing a work procedure of a craftsman in a virtual space, and for a user to perform training in the virtual space.
[1715] (Claim 3)
[1716] means for providing real-time feedback to a user during training in the virtual space;
[1717] 10. The system of claim 1, further comprising means for analyzing user motion data and providing feedback.
[1718] "Application example 2 when combining emotion engines"
[1719] (Claim 1)
[1720] a means of recording the work procedures of craftsmen through video and photographs;
[1721] means for receiving and analyzing the recorded data;
[1722] means for training a generative AI model using the analyzed data;
[1723] A means for providing a pseudo-training environment using the trained generative AI model;
[1724] A means of matching craftsmen with potential successors,
[1725] A means for collecting work procedure data in real time from craftsmen wearing smart glasses;
[1726] a means for controlling a robot using the collected data to execute the learned work procedure;
[1727] A system including means for recognizing a user's emotional state using emotion analysis means and providing feedback.
[1728] (Claim 2)
[1729] 2. The system according to claim 1, further comprising means for reproducing a work procedure of a craftsman in a virtual space, and for a user to perform training in the virtual space.
[1730] (Claim 3)
[1731] 10. The system of claim 1, further comprising means for providing real-time feedback to the user while training in the virtual space. [Explanation of symbols]
[1732] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. a means of recording the work procedures of craftsmen through video and photographs; means for receiving and analyzing the recorded data; means for training a generative AI model using the analyzed data; A means for providing a pseudo-training environment using the trained generative AI model; A system that includes a means of matching craftsmen with potential successors.
2. 2. The system according to claim 1, further comprising means for reproducing a work procedure of a craftsman in a virtual space, and for a user to perform training in the virtual space.
3. The system of claim 1 , further comprising means for providing real-time feedback to the user during training in the virtual space.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A