system
The system addresses the inflexibility of conventional machine control by creating a customizable motion model with user feedback, allowing machines to efficiently adapt to diverse conditions and improve over time.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-16
- Publication Date
- 2026-04-28
AI Technical Summary
Conventional machine control systems lack flexibility, requiring extensive labor and expertise for adjustments to operate in different environments and tasks, and do not effectively incorporate user feedback for improvements.
A system that constructs a general-purpose motion model in a physical simulation environment, adjusts it to specific machine characteristics, and allows user feedback for further optimization, enabling flexible and efficient machine operation.
Enables machines to operate adaptively in various environments without specialized programming, enhancing efficiency and accuracy through continuous feedback integration.
Smart Images

Figure 2026070883000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In conventional machine control, operations are often controlled by individually set rule-based programs, which has the problem of lacking flexibility. Therefore, in order to cope with different operating environments and tasks, it is necessary to modify and adjust the program each time, which requires a great deal of labor and expertise. The object of this invention is to solve such problems and improve the flexibility in the operation control of machines and reduce the design cost.
Means for Solving the Problems
[0005] This invention proposes constructing a general-purpose motion model using motion data generated in a physical simulation environment, and then adjusting that model to suit the characteristics of a specific machine. This system allows the adjusted model to be applied to actual machine operation, enabling the machine to operate flexibly and efficiently. Furthermore, users can monitor the machine's operation and provide feedback, enabling further improvements and optimizations.
[0006] A "physical simulation environment" is a platform for reproducing the behavior of machines and objects in a virtual space that mimics actual physical laws.
[0007] "Motion data" refers to a collection of information about specific actions of machines or robots, presented as numerical data or records.
[0008] A "general-purpose behavioral model" is an algorithm or program that learns and applies a wide range of behavioral patterns in a way that is independent of specific machines.
[0009] "Adjusting to specific characteristics" refers to the process of optimizing a pre-created model according to the physical attributes and specific functional needs of individual machines.
[0010] A "tuned model" is a model that has been optimized according to specific characteristics, resulting in efficient machine operation.
[0011] "Feedback" refers to opinions and data provided based on observations and information from users or systems, with the aim of improving processes and systems. [Brief explanation of the drawing]
[0012] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3]It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
MODE FOR CARRYING OUT THE INVENTION
[0013] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0014] First, the terms used in the following description will be explained.
[0015] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0016] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0017] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs, various parameters, and the like. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0018] In the following embodiments, the numbered communication I / F (Interface) is an interface that includes a communication processor, an antenna, and the like. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0019] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0020] [First Embodiment]
[0021] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0022] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0023] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0024] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0025] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0027] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0028] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0029] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0030] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0031] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0032] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0033] To implement this invention, first, the server constructs a physical simulation environment and generates diverse motion data. This simulation environment mimics the operation of actual machinery and enables recording of operation in various scenarios. Based on this simulation, the server executes a machine learning algorithm to construct a general-purpose motion model.
[0034] Next, the server adjusts this general-purpose model to suit the characteristics of a specific robot. This process involves fine-tuning the model according to the robot's size, power, and functional requirements. This results in a motion model that operates efficiently in a specific operating environment.
[0035] Subsequently, the robot, acting as the terminal, receives this adjusted model and begins operation. The robot can perform tasks such as moving objects on a manufacturing line. In this embodiment, the robot utilizes sensors to perceive environmental information and flexibly adapts its actions based on that information.
[0036] Users monitor the robot's movements and provide feedback as needed. This feedback can be used to fine-tune the model via the server, helping to further improve the robot's accuracy and efficiency.
[0037] As a concrete example, consider an assembly line in the manufacturing industry. The server simulates various assembly operations and builds a model based on them. The robot, acting as the terminal, uses this model to handle parts of different sizes accurately and quickly. The user can monitor the assembly process and provide feedback on accuracy and speed, further improving the process.
[0038] This configuration enables advanced robot motion control without requiring specialized programming knowledge, facilitating the development of machines that can function adaptively in various environments.
[0039] The following describes the processing flow.
[0040] Step 1:
[0041] The server sets up a physical simulation environment and generates motion data. In this process, a virtual robot model is used to design various motion scenarios and run simulations. For example, it simulates a robot arm grasping and lifting an object.
[0042] Step 2:
[0043] The server uses the generated behavioral data to execute machine learning algorithms and build a general-purpose behavioral model. Using deep learning techniques, it learns behavioral characteristics from the generated data and creates a model that can be applied to operation in various environments.
[0044] Step 3:
[0045] The server fine-tunes a general-purpose operating model to suit the specific characteristics of the machine. This process takes into account the specific operating requirements and physical characteristics of the actual machine, optimizing the model for the robot's operating environment.
[0046] Step 4:
[0047] The terminal robot receives a finely tuned motion model from the server and performs actions based on it. This model is used to efficiently perform tasks such as handling parts or assembly on a manufacturing line.
[0048] Step 5:
[0049] Users monitor the robot's movements in real time and provide feedback as needed. This feedback is used for fine-tuning to further improve the robot's accuracy and efficiency. Through this process, the overall system performance is enhanced.
[0050] (Example 1)
[0051] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0052] In modern automated systems, sophisticated motion models are necessary to flexibly adapt to different environments and conditions. However, current technology can only provide static models that depend on specific environments and conditions, making efficient and highly accurate operation difficult. Furthermore, the lack of sufficient use of feedback and readjustment functions to improve accuracy hinders improvements in versatility and efficiency.
[0053] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0054] In this invention, the server includes means for generating operational data and recording various operational patterns in a physical simulation environment, means for adjusting operational parameters of a general-purpose operational model to suit the characteristics of a specific automated device, and means for an operator to monitor the operation of the automated device, provide feedback on operational accuracy and efficiency, and use that feedback to further fine-tune the model. This makes it possible to adapt to a specific operating environment and continue to operate efficiently.
[0055] A "physical simulation environment" refers to a system or platform that generates virtual operating data by simulating real-world conditions and then analyzes that data.
[0056] "Motion data" refers to a dataset containing information about the various movements and paths performed by automated equipment, which is used for motion analysis and model construction.
[0057] A "motion model" is an algorithm or program built based on collected motion data, providing indicators or a framework for automated equipment to perform specific tasks efficiently and accurately.
[0058] An "automated device" refers to a machine or robot that can autonomously perform tasks based on a pre-set operating model.
[0059] "Operational parameters" are settings or variables that are customized to match the operating conditions and functions of automated equipment, and by being reflected in the operation model, they enable optimal operation.
[0060] An "operator" refers to a person whose role is to monitor automated equipment and its operation, and to provide feedback and make corrections as needed.
[0061] "Feedback" refers to the evaluations and suggestions provided by monitors to improve the operation of automated equipment, which then leads to the readjustment and improvement of the model.
[0062] A description of the embodiment for carrying out the invention will be provided.
[0063] First, the server constructs a physical simulation environment. This environment is used to simulate various scenarios in which real-world automated equipment operates. Specifically, it utilizes simulation software such as "Gazebo" and "Unity." Within this simulation environment, the server generates operational data and records various operational patterns. This allows for verification of operation under various conditions, independent of the environment.
[0064] Next, the server executes machine learning algorithms such as "TENSORFLOW®" and "PyTorch" based on the generated operational data to build a general-purpose operational model. This model forms the basis for achieving efficient and highly accurate operation of automated equipment in various environments.
[0065] Furthermore, the server adjusts this general operating model to suit the characteristics of the specific automated equipment. At this stage, various factors such as the size, range of motion, and power source of the equipment are taken into consideration. This adjustment optimizes the model to fit the specific operating environment.
[0066] After adjustment, the automated device, acting as the terminal, receives this motion model and starts the configured task. The automated device perceives the environment using sensors and performs the task while making decisions based on the motion model. For example, this could be a task such as accurately assembling multiple parts on a manufacturing line.
[0067] Finally, the user monitors the entire execution process and provides feedback on the accuracy and efficiency of the device's operation. This feedback data is sent to the server and used to further refine and retrain the operating model, thereby improving the device's performance.
[0068] A concrete example is a process in manufacturing assembly lines where the movements required at each stage are simulated, and the optimal assembly movements are modeled based on that data. By using this system, the efficiency and precision of the assembly line can be improved.
[0069] An example of an input prompt sentence for a generated AI model might be, "Design a behavioral model that optimizes the placement and handling of parts on a manufacturing line."
[0070] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0071] Step 1:
[0072] The server constructs a physical simulation environment. It receives simulation software configuration data as input and sets up a virtual environment. Here, it simulates the operation of a virtual automated device and records operation patterns in various environmental scenarios. As output, it generates operation data corresponding to each operation pattern. This operation data forms the basis of the system's operation model. Specifically, it reproduces the movement of arms and the movement of objects.
[0073] Step 2:
[0074] The server uses the generated behavioral data to execute a machine learning algorithm and build a general-purpose behavioral model. It receives the behavioral data generated in step 1 as input and performs data analysis. Specifically, it extracts major behavioral patterns and organizes the data to learn the most efficient behavior. This process generates a general-purpose behavioral model that can handle a variety of behaviors as output.
[0075] Step 3:
[0076] The server adapts a general-purpose motion model to the specific characteristics of the automation equipment. It uses parameter data, taking into account the equipment's physical characteristics and required performance, as input. Based on this data, the server fine-tunes the model's parameters to suit the specific operating environment. This results in a model that operates efficiently even under specific tasks and conditions. Specific operations include adjusting the motion profile according to the equipment's size and power.
[0077] Step 4:
[0078] The automated device, acting as the terminal, receives a pre-tuned motion model and initiates the actual task. It receives a pre-tuned motion model as input. Based on this model, the device utilizes sensor information to adapt its actions to the current environment. The output is the efficient execution of the specified task. Specifically, this involves autonomously handling parts and performing assembly processes on a manufacturing line.
[0079] Step 5:
[0080] Users monitor the operation of automated equipment and provide feedback on its accuracy and efficiency. Inputs include historical equipment operation data and visual monitoring information. Users analyze this data to identify areas for improvement. The feedback is sent to a server and used for model refinement and retraining. The output is a more accurate operating model. Specific actions include reviewing equipment operation logs and video feeds and reporting any necessary corrections.
[0081] (Application Example 1)
[0082] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0083] In today's industrial environment, there is a demand for diverse machinery to operate efficiently under varying conditions and for its operation to be easily adjusted. However, conventional methods have faced challenges such as requiring advanced expertise to fine-tune machine operation and lacking flexibility. In particular, the lack of mechanisms to quickly incorporate user feedback has made it difficult to improve operational efficiency.
[0084] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0085] In this invention, the server includes means for generating motion data in a physical simulation environment, means for constructing a general-purpose motion model using the generated motion data, and means for transmitting motion feedback from individual information terminals and using it for model adjustment. This enables the rapid incorporation of user-provided feedback into the motion model, resulting in efficient and flexible control of machine motion.
[0086] A "physical simulation environment" is a system that digitally mimics real-world physical phenomena and virtually reproduces the operation of machines under various conditions.
[0087] "Motion data" refers to a collection of information about the operation of a machine obtained through a physical simulation environment, and is used to construct motion models.
[0088] A "general-purpose motion model" is a model that represents motion patterns that can be commonly applied to various machines, and serves as a foundation for adapting to specific conditions or machines.
[0089] "Adjusting to the characteristics of a specific machine" means optimizing a general operating model to the performance and specifications of an individual machine and adapting it to the actual operating environment.
[0090] "Applying to actual machine operation" means implementing a tuned motion model into a machine that is actually in operation and controlling its operation in real time.
[0091] A "personal information terminal" is a digital device that a user carries with them to monitor and provide feedback on the operation of a machine.
[0092] "Sending operational feedback" is the process of a user communicating their opinions and suggestions for improvement regarding the machine's operation to a server via digital communication.
[0093] "Using it for model adjustment" means updating the motion model based on the feedback received to achieve more efficient and accurate machine operation.
[0094] The server builds a physical simulation environment and generates operational data over the network. This environment uses software to virtually reproduce real-world physical phenomena and simulate the operation of machines under different conditions. Specific software such as "Simulink" and "ANSYS," which are suitable for physical simulation and data analysis, are sometimes used.
[0095] Users monitor the machine's operation through applications installed on their individual information terminals and send feedback to the server. This feedback is used to improve the robot's accuracy and efficiency. The server receives this feedback, adapts the generated motion model using machine learning algorithms, and retrains the motion model. Representative machine learning frameworks used here include "TensorFlow" and "PyTorch."
[0096] The robot, acting as the terminal, receives an optimized motion model from the server and operates in the actual work environment. The robot uses built-in sensors to perceive environmental information in real time and performs adaptive actions based on that information. Sensor technologies such as "LiDAR" and "IMU (Inertial Measurement Unit)" are utilized in this process.
[0097] A concrete example is the operation of robots on a factory assembly line. Users monitor the assembly process using their smartphones and provide feedback to improve the speed of parts handling. The server then adjusts the operation model based on this information, enabling efficient assembly.
[0098] An example of a prompt for the generated AI model might be: "Please construct an optimal operating scenario for the robot to efficiently transport parts, and suggest ways to fine-tune the model based on user feedback."
[0099] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0100] Step 1:
[0101] The server constructs a physical simulation environment and generates machine motion data under various conditions. The inputs are virtual models and physical laws, and the output is detailed motion data obtained from the simulation. This data is generated using a physics engine to comprehensively cover the machine's motion patterns.
[0102] Step 2:
[0103] The server uses the generated motion data to build a general-purpose motion model. The input is the motion data obtained in step 1, and the output is a model of motion patterns applicable to various machines. Machine learning algorithms are used to analyze the data and perform pattern recognition and model optimization.
[0104] Step 3:
[0105] Users monitor actual machine operation and provide feedback via individual information terminals. Input is the machine operation information observed by the user, and output is the feedback data received by the server. Users use smart terminals to clearly communicate areas where improvement or adjustment of operation is needed.
[0106] Step 4:
[0107] The server adjusts its behavioral model based on the feedback it receives. The input is user feedback, and the output is the finely tuned behavioral model. Through data analysis and a retraining process, the model is updated to better suit the user's requirements.
[0108] Step 5:
[0109] The terminal robot applies a pre-tuned motion model to the actual work environment and starts operating efficiently. The input is the latest motion model provided by the server, and the output is the robot's specific action result. The robot collects environmental information through sensors and performs the optimal action based on the model in real time.
[0110] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0111] To implement this invention, the system consists of three components: a server, a terminal (robot), and a user. The server generates motion data in various scenarios using a physical simulation environment and constructs a general-purpose motion model by executing machine learning algorithms based on that data. Next, the server fine-tunes this model to suit the characteristics of a specific robot, providing a tuned model applicable to real-world motion.
[0112] The robot, acting as the terminal, uses this adjusted model to flexibly perform real-world tasks. For example, it can perform tasks such as moving objects in a manufacturing environment. The robot is equipped with an emotion engine that can recognize the user's emotions in real time. This emotional information is taken into consideration by the robot during its operations, and is used to improve work efficiency and execution accuracy.
[0113] Users can monitor the robot's movements and provide real-time feedback. The emotion engine senses the user's emotional state, such as joy or dissatisfaction, and sends this information back to the server, which then uses the feedback to further optimize the movement model. Furthermore, the robot's movement process is improved based on the user's emotional information, enhancing the quality of its work.
[0114] As a concrete example, consider a robotic assistant in a nursing care facility. This system allows the care robot to sense the emotions of the residents and perform actions that provide a sense of security. By providing services that respond to the residents' emotions, it aims to reduce their stress and anxiety and create a better user experience.
[0115] Thus, systems that incorporate emotion engines not only improve operational efficiency but also enable the provision of optimal services based on user interaction, and are expected to be used in a variety of application fields.
[0116] The following describes the processing flow.
[0117] Step 1:
[0118] The server sets up a physical simulation environment and generates motion data by simulating various motion scenarios. These scenarios include actions such as the robot grasping, moving, and assembling objects.
[0119] Step 2:
[0120] The server uses the generated behavioral data to execute machine learning algorithms and build a general-purpose behavioral model. This model can learn various behavioral patterns and has broad applicability.
[0121] Step 3:
[0122] The server fine-tunes a general-purpose motion model according to the specific robot and application. This results in an optimal model that matches the robot's actual physical characteristics and operating environment.
[0123] Step 4:
[0124] The robot, acting as the terminal, is equipped with an emotion engine that recognizes the emotional information conveyed by the user in real time. This allows the robot to select actions and optimize its response according to the user's emotions.
[0125] Step 5:
[0126] Users observe the robot's movements and see how their emotions are being communicated to the robot. Users evaluate their satisfaction with the robot's work and services, and this evaluation is fed back into the system as emotional feedback.
[0127] Step 6:
[0128] The server uses user feedback and emotion data to further refine the motion model and emotion engine. This continuously improves the robot's motion accuracy and user interaction capabilities.
[0129] Through the processing steps described above, the system provides behavior that aligns with the user's emotions, achieving both machine flexibility and an improved user experience.
[0130] (Example 2)
[0131] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0132] Efficiently and flexibly controlling the operation of automated machinery in the real world is technically challenging. In particular, constructing motion models that can handle diverse scenarios and adapting those models to the characteristics of specific devices has not been adequately achieved with conventional technologies. Solving this challenge is necessary to improve the efficiency and accuracy of automated work and enable effective interaction with users.
[0133] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0134] In this invention, the server includes means for generating information about actions in a virtual environment, means for creating a structure for mimicking general actions using the generated information about actions, and means for adapting the structure for mimicking general actions to the characteristics of a specific device. This enables efficient and flexible control of the operation in the implementation device, and also allows for further improvements utilizing real-time feedback to the user.
[0135] A "virtual environment" is a simulation system used to reproduce physical phenomena and situations in the real world on a computer.
[0136] "Information about behavior" refers to specific action data such as movements, actions, and reactions, and is a set of data that forms the basis for motion analysis and model construction of automated devices.
[0137] "A structure for mimicking general behavior" refers to algorithms and frameworks that constitute behavioral models applicable to a variety of behavioral scenarios.
[0138] "Adapting to the characteristics of a specific device" refers to the process of adjusting a general operating model to optimize it for the physical and technical constraints and characteristics of a particular device.
[0139] "Applying to the operation of a real device" means using a modified motion model on an actual machine to control its operation.
[0140] An "external observer" refers to a user or operator whose role is to observe the operation of a system or machine and to provide feedback on its performance and functionality.
[0141] "Opinions" refer to evaluations and feedback from external observers regarding the operation and results of a device, and are information that plays a part in optimizing the system.
[0142] "Recognition technology" refers to methods and techniques for improving the performance of models by incorporating information about behavior generated using machine learning algorithms and the like.
[0143] This invention is implemented by a system primarily composed of three main elements: a server, a terminal, and a user. The server generates information about actions in a virtual scenario by using a physical simulation environment. Specifically, it uses a high-performance computer and simulation software to reproduce complex physical phenomena and behavioral patterns. Using this generated information, it utilizes machine learning frameworks such as TensorFlow and PyTorch to construct a structure, or behavioral model, for simulating general behavior.
[0144] Next, the server adapts this behavioral model to the characteristics of a specific terminal device. The terminal is specifically an automated device such as a robot, which uses the adapted model received from the server to flexibly and efficiently perform real-world tasks. The robot recognizes the user's emotions in real time through a built-in emotion recognition module and optimizes its actions based on that information. The emotion recognition technology used here employs common image processing and speech recognition technologies.
[0145] The user monitors the robot's movements and provides feedback on the results. This feedback, including emotions and actions, is returned to the server and used to further optimize the motion model. This entire process enables flexible control of the device's movements and provides high-quality service to the user.
[0146] A concrete example is a robotic assistant in a nursing home. This robot improves the psychological comfort of users by identifying their emotions and providing gentle actions that alleviate anxiety. Another example of a prompt for a generative AI model is the instruction, "Please tell me the optimal action strategy for a robot to provide comfort to users in a nursing home." Through such prompts, the AI model makes specific suggestions and contributes to the overall system.
[0147] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0148] Step 1:
[0149] The server builds a physical simulation environment and generates behavioral data for virtual scenarios. Scenario parameters (e.g., gravity, friction, initial conditions) are set as input. The simulation engine uses these parameters to simulate movement and generates motion trajectories and interaction data as output. This data forms the basis for subsequent motion model construction. Specific examples include the reproduction of vehicle and robotic arm movements.
[0150] Step 2:
[0151] The server uses the generated behavioral data to build a general-purpose behavioral model using a machine learning algorithm. The behavioral data obtained in step 1 is used as input. A machine learning framework (e.g., TensorFlow) is used to train a neural network, obtaining a general-purpose behavioral model as output. This model provides a foundation capable of handling a wide variety of scenarios. Specifically, data is input into a feedforward neural network, and appropriate weights are adjusted.
[0152] Step 3:
[0153] The server fine-tunes the constructed general-purpose motion model to match the characteristics of a specific terminal device. The input is the physical and environmental characteristics of the terminal device. The server adjusts the model taking these characteristics into account, generating an optimized motion model for the specific device as output. This process allows the robot to adapt to specific work environments and tasks. As part of the adjustment, the device's sensor and actuator characteristics are incorporated into the model.
[0154] Step 4:
[0155] The terminal (robot) receives a pre-configured motion model from the server and performs the actual task. Inputs include the motion model from the server and real-time data from the real environment (e.g., camera images, distance sensor information). The robot uses this data to perform the task and generates task results and log information as output. Specifically, the robot performs actions such as actually transporting an object.
[0156] Step 5:
[0157] The user monitors the robot's movements and provides feedback. The user observes the robot's movements and task performance as input. Based on this, the user provides feedback on the quality of the movements. This feedback information is sent to the server as output. Specifically, the user evaluates the robot's smoothness and speed.
[0158] Step 6:
[0159] The server receives feedback from the user and further optimizes the behavioral model. The user feedback information is used as input. Based on this information, the server retrains its learning algorithm and adjusts the adaptive model. The output is a more improved behavioral model. Specifically, this includes readjusting weights and testing new behavioral patterns.
[0160] (Application Example 2)
[0161] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0162] In recent years, the operation of automated devices in the real world has required not only efficiency but also flexible responses that respond to the emotions of users. However, conventional systems have difficulty recognizing user emotions in real time and optimizing their operation based on that. In this situation, the challenge is to build systems that enable automated devices to provide more personalized services to users.
[0163] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0164] In this invention, the server includes means for generating behavioral data in a physical environment, means for constructing a general-purpose behavioral model using the generated behavioral data, and means for detecting and analyzing the customer's emotional state in real time. This enables the automated device to perform actions on-site that take customer emotions into consideration, allowing for the provision of more appropriate services.
[0165] The "physical environment" refers to the real-world environment with its own physical characteristics and constraints, and is the area for verifying the operation of automated devices through simulations and behavioral data.
[0166] "Behavioral data" refers to a collection of information that records the specific actions performed by automated equipment in a particular environment, along with the various parameters associated with those actions.
[0167] A "general-purpose behavioral model" refers to a predictive model that abstracts and generalizes the operation of automated devices applicable to various environments and conditions.
[0168] An "automated device" is a general term for a machine or device that can perform programmed actions autonomously, and is particularly designed to perform a specific task.
[0169] "Emotional state" refers to a person's internal psychological state and includes information such as joy and anxiety, which can be identified from cues such as facial expressions and voice.
[0170] "Methods for real-time analysis" refer to methods and technologies that process data immediately upon acquisition and obtain results, enabling rapid feedback.
[0171] In an embodiment for carrying out this invention, the system is mainly configured as follows.
[0172] The server generates behavioral data for automated devices in a physical environment and constructs a general-purpose behavioral model based on the generated data. This model is executed using machine learning libraries such as TensorFlow and PyTorch. A physical simulation environment is utilized to generate the behavioral data, resulting in a model that can handle a variety of scenarios.
[0173] The refined model will be applied to specific automated devices and is intended for use in physical stores and other similar settings. In this scenario, a device equipped with a camera and microphone will be used to detect emotional states, and real-time sentiment analysis will be performed using image processing libraries such as OpenCV. The sentiment data will be processed using a generative AI model, which will generate prompt messages as needed. Specifically, these prompts may include phrases like, "When the customer's facial expression is relaxed, recommend new products in a friendly tone."
[0174] The device uses these tuned models to perform its actual operations. If the customer shows favor, it selects the next product or service to suggest and makes recommendations in natural language. For example, if the customer smiles, it might say, "Here are some incense sticks that would go well with the aromatherapy candle you just saw."
[0175] Users are required to monitor the overall operation of the system and provide feedback. This feedback is then sent back to the server and used to improve the model. This two-way process allows the automated system to operate with increasing precision.
[0176] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0177] Step 1:
[0178] The server sets up scenarios in a physical environment and generates behavioral data. It receives parameters of the simulation environment as input and simulates the operation of automated equipment under various conditions based on these parameters. As output, it provides a set of generated behavioral data, which later serves as foundational data for training behavioral models.
[0179] Step 2:
[0180] The server constructs a general-purpose behavioral model using the generated behavioral data. It receives the behavioral data generated in step 1 as input and trains the model using a machine learning algorithm (e.g., TensorFlow or PyTorch). The output is a general-purpose behavioral model applicable under various conditions. This model predicts the basic operating patterns of automated equipment.
[0181] Step 3:
[0182] The server adapts a general-purpose behavioral model to optimize it for a specific automated device. It receives device characteristic information and the aforementioned general-purpose model as input. As data processing, it applies an adjustment algorithm and provides an optimized model for that device as output. This model enables operation tailored to the characteristics of the actual operating environment.
[0183] Step 4:
[0184] The terminal loads a pre-configured model and is used in real-world settings such as physical stores. It receives the pre-configured model and real-time emotional data as input. Data processing analyzes customer facial expressions and voices via camera and microphone to perform emotion recognition. The output consists of appropriate action commands tailored to the situation, specifically generating prompt messages to provide service to the customer.
[0185] Step 5:
[0186] Users monitor the device's operation and provide feedback. Inputs include device operation results and customer response data. Based on this information, users evaluate the system and create feedback that identifies areas for improvement. This feedback is sent to the server and used for subsequent model adjustments.
[0187] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0188] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0189] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0190] [Second Embodiment]
[0191] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0192] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0193] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0194] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0195] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0196] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0197] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0198] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0199] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0200] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0201] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0202] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0203] To implement this invention, first, the server constructs a physical simulation environment and generates diverse motion data. This simulation environment mimics the operation of actual machinery and enables recording of operation in various scenarios. Based on this simulation, the server executes a machine learning algorithm to construct a general-purpose motion model.
[0204] Next, the server adjusts this general-purpose model to suit the characteristics of a specific robot. This process involves fine-tuning the model according to the robot's size, power, and functional requirements. This results in a motion model that operates efficiently in a specific operating environment.
[0205] Subsequently, the robot, acting as the terminal, receives this adjusted model and begins operation. The robot can perform tasks such as moving objects on a manufacturing line. In this embodiment, the robot utilizes sensors to perceive environmental information and flexibly adapts its actions based on that information.
[0206] Users monitor the robot's movements and provide feedback as needed. This feedback can be used to fine-tune the model via the server, helping to further improve the robot's accuracy and efficiency.
[0207] As a concrete example, consider an assembly line in the manufacturing industry. The server simulates various assembly operations and builds a model based on them. The robot, acting as the terminal, uses this model to handle parts of different sizes accurately and quickly. The user can monitor the assembly process and provide feedback on accuracy and speed, further improving the process.
[0208] This configuration enables advanced robot motion control without requiring specialized programming knowledge, facilitating the development of machines that can function adaptively in various environments.
[0209] The following describes the processing flow.
[0210] Step 1:
[0211] The server sets up a physical simulation environment and generates motion data. In this process, a virtual robot model is used to design various motion scenarios and run simulations. For example, it simulates a robot arm grasping and lifting an object.
[0212] Step 2:
[0213] The server uses the generated behavioral data to execute machine learning algorithms and build a general-purpose behavioral model. Using deep learning techniques, it learns behavioral characteristics from the generated data and creates a model that can be applied to operation in various environments.
[0214] Step 3:
[0215] The server fine-tunes a general-purpose operating model to suit the specific characteristics of the machine. This process takes into account the specific operating requirements and physical characteristics of the actual machine, optimizing the model for the robot's operating environment.
[0216] Step 4:
[0217] The terminal robot receives a finely tuned motion model from the server and performs actions based on it. This model is used to efficiently perform tasks such as handling parts or assembly on a manufacturing line.
[0218] Step 5:
[0219] Users monitor the robot's movements in real time and provide feedback as needed. This feedback is used for fine-tuning to further improve the robot's accuracy and efficiency. Through this process, the overall system performance is enhanced.
[0220] (Example 1)
[0221] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0222] In modern automated systems, sophisticated motion models are necessary to flexibly adapt to different environments and conditions. However, current technology can only provide static models that depend on specific environments and conditions, making efficient and highly accurate operation difficult. Furthermore, the lack of sufficient use of feedback and readjustment functions to improve accuracy hinders improvements in versatility and efficiency.
[0223] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0224] In this invention, the server includes means for generating operational data and recording various operational patterns in a physical simulation environment, means for adjusting operational parameters of a general-purpose operational model to suit the characteristics of a specific automated device, and means for an operator to monitor the operation of the automated device, provide feedback on operational accuracy and efficiency, and use that feedback to further fine-tune the model. This makes it possible to adapt to a specific operating environment and continue to operate efficiently.
[0225] A "physical simulation environment" refers to a system or platform that generates virtual operating data by simulating real-world conditions and then analyzes that data.
[0226] "Motion data" refers to a dataset containing information about the various movements and paths performed by automated equipment, which is used for motion analysis and model construction.
[0227] A "motion model" is an algorithm or program built based on collected motion data, providing indicators or a framework for automated equipment to perform specific tasks efficiently and accurately.
[0228] An "automated device" refers to a machine or robot that can autonomously perform tasks based on a pre-set operating model.
[0229] "Operational parameters" are settings or variables that are customized to match the operating conditions and functions of automated equipment, and by being reflected in the operation model, they enable optimal operation.
[0230] An "operator" refers to a person whose role is to monitor automated equipment and its operation, and to provide feedback and make corrections as needed.
[0231] "Feedback" refers to the evaluations and suggestions provided by monitors to improve the operation of automated equipment, which then leads to the readjustment and improvement of the model.
[0232] A description of the embodiment for carrying out the invention will be provided.
[0233] First, the server constructs a physical simulation environment. This environment is used to simulate various scenarios in which real-world automated equipment operates. Specifically, it utilizes simulation software such as "Gazebo" and "Unity." Within this simulation environment, the server generates operational data and records various operational patterns. This allows for verification of operation under various conditions, independent of the environment.
[0234] Next, the server executes machine learning algorithms such as "TensorFlow" and "PyTorch" based on the generated operational data to build a general-purpose operational model. This model forms the basis for achieving efficient and highly accurate operation of automated devices in various environments.
[0235] Furthermore, the server adjusts this general operating model to suit the characteristics of the specific automated equipment. At this stage, various factors such as the size, range of motion, and power source of the equipment are taken into consideration. This adjustment optimizes the model to fit the specific operating environment.
[0236] After adjustment, the automated device, acting as the terminal, receives this motion model and starts the configured task. The automated device perceives the environment using sensors and performs the task while making decisions based on the motion model. For example, this could be a task such as accurately assembling multiple parts on a manufacturing line.
[0237] Finally, the user monitors the entire execution process and provides feedback on the accuracy and efficiency of the device's operation. This feedback data is sent to the server and used to further refine and retrain the operating model, thereby improving the device's performance.
[0238] A concrete example is a process in manufacturing assembly lines where the movements required at each stage are simulated, and the optimal assembly movements are modeled based on that data. By using this system, the efficiency and precision of the assembly line can be improved.
[0239] An example of an input prompt sentence for a generated AI model might be, "Design a behavioral model that optimizes the placement and handling of parts on a manufacturing line."
[0240] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0241] Step 1:
[0242] The server constructs a physical simulation environment. It receives simulation software configuration data as input and sets up a virtual environment. Here, it simulates the operation of a virtual automated device and records operation patterns in various environmental scenarios. As output, it generates operation data corresponding to each operation pattern. This operation data forms the basis of the system's operation model. Specifically, it reproduces the movement of arms and the movement of objects.
[0243] Step 2:
[0244] The server uses the generated behavioral data to execute a machine learning algorithm and build a general-purpose behavioral model. It receives the behavioral data generated in step 1 as input and performs data analysis. Specifically, it extracts major behavioral patterns and organizes the data to learn the most efficient behavior. This process generates a general-purpose behavioral model that can handle a variety of behaviors as output.
[0245] Step 3:
[0246] The server adapts a general-purpose motion model to the specific characteristics of the automation equipment. It uses parameter data, taking into account the equipment's physical characteristics and required performance, as input. Based on this data, the server fine-tunes the model's parameters to suit the specific operating environment. This results in a model that operates efficiently even under specific tasks and conditions. Specific operations include adjusting the motion profile according to the equipment's size and power.
[0247] Step 4:
[0248] The automated device, acting as the terminal, receives a pre-tuned motion model and initiates the actual task. It receives a pre-tuned motion model as input. Based on this model, the device utilizes sensor information to adapt its actions to the current environment. The output is the efficient execution of the specified task. Specifically, this involves autonomously handling parts and performing assembly processes on a manufacturing line.
[0249] Step 5:
[0250] Users monitor the operation of automated equipment and provide feedback on its accuracy and efficiency. Inputs include historical equipment operation data and visual monitoring information. Users analyze this data to identify areas for improvement. The feedback is sent to a server and used for model refinement and retraining. The output is a more accurate operating model. Specific actions include reviewing equipment operation logs and video feeds and reporting any necessary corrections.
[0251] (Application Example 1)
[0252] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0253] In today's industrial environment, there is a demand for diverse machinery to operate efficiently under varying conditions and for its operation to be easily adjusted. However, conventional methods have faced challenges such as requiring advanced expertise to fine-tune machine operation and lacking flexibility. In particular, the lack of mechanisms to quickly incorporate user feedback has made it difficult to improve operational efficiency.
[0254] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0255] In this invention, the server includes means for generating motion data in a physical simulation environment, means for constructing a general-purpose motion model using the generated motion data, and means for transmitting motion feedback from individual information terminals and using it for model adjustment. This enables the rapid incorporation of user-provided feedback into the motion model, resulting in efficient and flexible control of machine motion.
[0256] A "physical simulation environment" is a system that digitally mimics real-world physical phenomena and virtually reproduces the operation of machines under various conditions.
[0257] "Motion data" refers to a collection of information about the operation of a machine obtained through a physical simulation environment, and is used to construct motion models.
[0258] A "general-purpose motion model" is a model that represents motion patterns that can be commonly applied to various machines, and serves as a foundation for adapting to specific conditions or machines.
[0259] "Adjusting to the characteristics of a specific machine" means optimizing a general operating model to the performance and specifications of an individual machine and adapting it to the actual operating environment.
[0260] "Applying to actual machine operation" means implementing a tuned motion model into a machine that is actually in operation and controlling its operation in real time.
[0261] A "personal information terminal" is a digital device that a user carries with them to monitor and provide feedback on the operation of a machine.
[0262] "Sending operational feedback" is the process of a user communicating their opinions and suggestions for improvement regarding the machine's operation to a server via digital communication.
[0263] "Using it for model adjustment" means updating the motion model based on the feedback received to achieve more efficient and accurate machine operation.
[0264] The server builds a physical simulation environment and generates operational data over the network. This environment uses software to virtually reproduce real-world physical phenomena and simulate the operation of machines under different conditions. Specific software such as "Simulink" and "ANSYS," which are suitable for physical simulation and data analysis, are sometimes used.
[0265] Users monitor the machine's operation through applications installed on their individual information terminals and send feedback to the server. This feedback is used to improve the robot's accuracy and efficiency. The server receives this feedback, adapts the generated motion model using machine learning algorithms, and retrains the motion model. Representative machine learning frameworks used here include "TensorFlow" and "PyTorch."
[0266] The robot, acting as the terminal, receives an optimized motion model from the server and operates in the actual work environment. The robot uses built-in sensors to perceive environmental information in real time and performs adaptive actions based on that information. Sensor technologies such as "LiDAR" and "IMU (Inertial Measurement Unit)" are utilized in this process.
[0267] A concrete example is the operation of robots on a factory assembly line. Users monitor the assembly process using their smartphones and provide feedback to improve the speed of parts handling. The server then adjusts the operation model based on this information, enabling efficient assembly.
[0268] An example of a prompt for the generated AI model might be: "Please construct an optimal operating scenario for the robot to efficiently transport parts, and suggest ways to fine-tune the model based on user feedback."
[0269] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0270] Step 1:
[0271] The server constructs a physical simulation environment and generates machine motion data under various conditions. The inputs are virtual models and physical laws, and the output is detailed motion data obtained from the simulation. This data is generated using a physics engine to comprehensively cover the machine's motion patterns.
[0272] Step 2:
[0273] The server uses the generated motion data to build a general-purpose motion model. The input is the motion data obtained in step 1, and the output is a model of motion patterns applicable to various machines. Machine learning algorithms are used to analyze the data and perform pattern recognition and model optimization.
[0274] Step 3:
[0275] Users monitor actual machine operation and provide feedback via individual information terminals. Input is the machine operation information observed by the user, and output is the feedback data received by the server. Users use smart terminals to clearly communicate areas where improvement or adjustment of operation is needed.
[0276] Step 4:
[0277] The server adjusts its behavioral model based on the feedback it receives. The input is user feedback, and the output is the finely tuned behavioral model. Through data analysis and a retraining process, the model is updated to better suit the user's requirements.
[0278] Step 5:
[0279] The terminal robot applies a pre-tuned motion model to the actual work environment and starts operating efficiently. The input is the latest motion model provided by the server, and the output is the robot's specific action result. The robot collects environmental information through sensors and performs the optimal action based on the model in real time.
[0280] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0281] To implement this invention, the system consists of three components: a server, a terminal (robot), and a user. The server generates motion data in various scenarios using a physical simulation environment and constructs a general-purpose motion model by executing machine learning algorithms based on that data. Next, the server fine-tunes this model to suit the characteristics of a specific robot, providing a tuned model applicable to real-world motion.
[0282] The robot, which is a terminal, uses this adjusted model to flexibly perform actual tasks. For example, it executes tasks such as moving objects at a manufacturing site. At this time, the robot is equipped with an emotion engine and can recognize the user's emotions in real time. This emotion information is considered when the robot operates and is used to improve work efficiency and execution accuracy.
[0283] The user can monitor the operation of the robot and provide real-time feedback. The emotion engine senses the user's emotional states such as joy and dissatisfaction, and returns the information to the server. Based on this feedback, the server further optimizes the operation model. Also, based on the user's emotion information, the operation process of the robot is improved to enhance the quality of the work.
[0284] As a specific example, a robot assistant in a nursing facility is cited. This system enables the care robot to sense the emotions of the user and perform operations that give a sense of security. By providing services according to emotions to reduce the stress and anxiety of the user, a better user experience is realized.
[0285] Thus, the system combined with the emotion engine not only improves the efficiency of operations but also enables the provision of optimal services based on interactions with the user, and is expected to be used in various application fields.
[0286] The following describes the processing flow.
[0287] Step 1:
[0288] The server sets up a physical simulation environment and generates operation data by simulating various operation scenarios. The scenarios include operations such as the robot grasping, moving, and assembling objects.
[0289] Step 2:
[0290] The server uses the generated behavioral data to execute machine learning algorithms and build a general-purpose behavioral model. This model can learn various behavioral patterns and has broad applicability.
[0291] Step 3:
[0292] The server fine-tunes a general-purpose motion model according to the specific robot and application. This results in an optimal model that matches the robot's actual physical characteristics and operating environment.
[0293] Step 4:
[0294] The robot, acting as the terminal, is equipped with an emotion engine that recognizes the emotional information conveyed by the user in real time. This allows the robot to select actions and optimize its response according to the user's emotions.
[0295] Step 5:
[0296] Users observe the robot's movements and see how their emotions are being communicated to the robot. Users evaluate their satisfaction with the robot's work and services, and this evaluation is fed back into the system as emotional feedback.
[0297] Step 6:
[0298] The server uses user feedback and emotion data to further refine the motion model and emotion engine. This continuously improves the robot's motion accuracy and user interaction capabilities.
[0299] Through the processing steps described above, the system provides behavior that aligns with the user's emotions, achieving both machine flexibility and an improved user experience.
[0300] (Example 2)
[0301] Next, Example 2 will be described. In the following description, the data processing device 12 is referred to as a "server", and the smart glasses 214 are referred to as a "terminal".
[0302] Efficiently and flexibly controlling the operation of automated mechanical devices in the real world is technically difficult. In particular, constructing an operation model capable of handling various scenarios and adapting that model to the characteristics of a specific device has not been fully realized by conventional techniques. By solving this problem, it is necessary to improve the efficiency and accuracy of automated work and enable effective interaction with users.
[0303] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0304] In this invention, the server includes means for generating information related to actions in a virtual environment, means for creating a structure for mimicking general actions using the generated information related to actions, and means for adapting the structure for mimicking general actions to the characteristics of a specific device. Thereby, efficient and flexible control of the operation in the implementation device becomes possible, and further improvement utilizing real-time feedback to the user can also be realized.
[0305] The "virtual environment" is a simulation system for reproducing physical phenomena and situations in the real world on a computer.
[0306] The "information related to actions" refers to specific action data such as operations, movements, and reactions, and is a group of data that forms the basis for operation analysis and model construction of automated devices.
[0307] The "structure for mimicking general actions" refers to algorithms and frameworks that constitute an operation model applicable in common to various operation scenarios.
[0308] "Adapting to the characteristics of a specific device" refers to the process of adjusting a general operating model to optimize it for the physical and technical constraints and characteristics of a particular device.
[0309] "Applying to the operation of a real device" means using a modified motion model on an actual machine to control its operation.
[0310] An "external observer" refers to a user or operator whose role is to observe the operation of a system or machine and to provide feedback on its performance and functionality.
[0311] "Opinions" refer to evaluations and feedback from external observers regarding the operation and results of a device, and are information that plays a part in optimizing the system.
[0312] "Recognition technology" refers to methods and techniques for improving the performance of models by incorporating information about behavior generated using machine learning algorithms and the like.
[0313] This invention is implemented by a system primarily composed of three main elements: a server, a terminal, and a user. The server generates information about actions in a virtual scenario by using a physical simulation environment. Specifically, it uses a high-performance computer and simulation software to reproduce complex physical phenomena and behavioral patterns. Using this generated information, it utilizes machine learning frameworks such as TensorFlow and PyTorch to construct a structure, or behavioral model, for simulating general behavior.
[0314] Next, the server adapts this behavioral model to the characteristics of a specific terminal device. The terminal is specifically an automated device such as a robot, which uses the adapted model received from the server to flexibly and efficiently perform real-world tasks. The robot recognizes the user's emotions in real time through a built-in emotion recognition module and optimizes its actions based on that information. The emotion recognition technology used here employs common image processing and speech recognition technologies.
[0315] The user monitors the robot's movements and provides feedback on the results. This feedback, including emotions and actions, is returned to the server and used to further optimize the motion model. This entire process enables flexible control of the device's movements and provides high-quality service to the user.
[0316] A concrete example is a robotic assistant in a nursing home. This robot improves the psychological comfort of users by identifying their emotions and providing gentle actions that alleviate anxiety. Another example of a prompt for a generative AI model is the instruction, "Please tell me the optimal action strategy for a robot to provide comfort to users in a nursing home." Through such prompts, the AI model makes specific suggestions and contributes to the overall system.
[0317] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0318] Step 1:
[0319] The server builds a physical simulation environment and generates behavioral data for virtual scenarios. Scenario parameters (e.g., gravity, friction, initial conditions) are set as input. The simulation engine uses these parameters to simulate movement and generates motion trajectories and interaction data as output. This data forms the basis for subsequent motion model construction. Specific examples include the reproduction of vehicle and robotic arm movements.
[0320] Step 2:
[0321] The server uses the generated behavioral data to build a general-purpose behavioral model using a machine learning algorithm. The behavioral data obtained in step 1 is used as input. A machine learning framework (e.g., TensorFlow) is used to train a neural network, obtaining a general-purpose behavioral model as output. This model provides a foundation capable of handling a wide variety of scenarios. Specifically, data is input into a feedforward neural network, and appropriate weights are adjusted.
[0322] Step 3:
[0323] The server fine-tunes the constructed general-purpose motion model to match the characteristics of a specific terminal device. The input is the physical and environmental characteristics of the terminal device. The server adjusts the model taking these characteristics into account, generating an optimized motion model for the specific device as output. This process allows the robot to adapt to specific work environments and tasks. As part of the adjustment, the device's sensor and actuator characteristics are incorporated into the model.
[0324] Step 4:
[0325] The terminal (robot) receives a pre-configured motion model from the server and performs the actual task. Inputs include the motion model from the server and real-time data from the real environment (e.g., camera images, distance sensor information). The robot uses this data to perform the task and generates task results and log information as output. Specifically, the robot performs actions such as actually transporting an object.
[0326] Step 5:
[0327] The user monitors the robot's movements and provides feedback. The user observes the robot's movements and task performance as input. Based on this, the user provides feedback on the quality of the movements. This feedback information is sent to the server as output. Specifically, the user evaluates the robot's smoothness and speed.
[0328] Step 6:
[0329] The server receives feedback from the user and further optimizes the behavioral model. The user feedback information is used as input. Based on this information, the server retrains its learning algorithm and adjusts the adaptive model. The output is a more improved behavioral model. Specifically, this includes readjusting weights and testing new behavioral patterns.
[0330] (Application Example 2)
[0331] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0332] In recent years, the operation of automated devices in the real world has required not only efficiency but also flexible responses that respond to the emotions of users. However, conventional systems have difficulty recognizing user emotions in real time and optimizing their operation based on that. In this situation, the challenge is to build systems that enable automated devices to provide more personalized services to users.
[0333] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0334] In this invention, the server includes means for generating behavioral data in a physical environment, means for constructing a general-purpose behavioral model using the generated behavioral data, and means for detecting and analyzing the customer's emotional state in real time. This enables the automated device to perform actions on-site that take customer emotions into consideration, allowing for the provision of more appropriate services.
[0335] The "physical environment" refers to the real-world environment with its own physical characteristics and constraints, and is the area for verifying the operation of automated devices through simulations and behavioral data.
[0336] "Behavioral data" refers to a collection of information that records the specific actions performed by automated equipment in a particular environment, along with the various parameters associated with those actions.
[0337] A "general-purpose behavioral model" refers to a predictive model that abstracts and generalizes the operation of automated devices applicable to various environments and conditions.
[0338] An "automated device" is a general term for a machine or device that can perform programmed actions autonomously, and is particularly designed to perform a specific task.
[0339] "Emotional state" refers to a person's internal psychological state and includes information such as joy and anxiety, which can be identified from cues such as facial expressions and voice.
[0340] "Methods for real-time analysis" refer to methods and technologies that process data immediately upon acquisition and obtain results, enabling rapid feedback.
[0341] In an embodiment for carrying out this invention, the system is mainly configured as follows.
[0342] The server generates behavioral data for automated devices in a physical environment and constructs a general-purpose behavioral model based on the generated data. This model is executed using machine learning libraries such as TensorFlow and PyTorch. A physical simulation environment is utilized to generate the behavioral data, resulting in a model that can handle a variety of scenarios.
[0343] The refined model will be applied to specific automated devices and is intended for use in physical stores and other similar settings. In this scenario, a device equipped with a camera and microphone will be used to detect emotional states, and real-time sentiment analysis will be performed using image processing libraries such as OpenCV. The sentiment data will be processed using a generative AI model, which will generate prompt messages as needed. Specifically, these prompts may include phrases like, "When the customer's facial expression is relaxed, recommend new products in a friendly tone."
[0344] The device uses these tuned models to perform its actual operations. If the customer shows favor, it selects the next product or service to suggest and makes recommendations in natural language. For example, if the customer smiles, it might say, "Here are some incense sticks that would go well with the aromatherapy candle you just saw."
[0345] Users are required to monitor the overall operation of the system and provide feedback. This feedback is then sent back to the server and used to improve the model. This two-way process allows the automated system to operate with increasing precision.
[0346] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0347] Step 1:
[0348] The server sets up scenarios in a physical environment and generates behavioral data. It receives parameters of the simulation environment as input and simulates the operation of automated equipment under various conditions based on these parameters. As output, it provides a set of generated behavioral data, which later serves as foundational data for training behavioral models.
[0349] Step 2:
[0350] The server constructs a general-purpose behavioral model using the generated behavioral data. It receives the behavioral data generated in step 1 as input and trains the model using a machine learning algorithm (e.g., TensorFlow or PyTorch). The output is a general-purpose behavioral model applicable under various conditions. This model predicts the basic operating patterns of automated equipment.
[0351] Step 3:
[0352] The server adapts a general-purpose behavioral model to optimize it for a specific automated device. It receives device characteristic information and the aforementioned general-purpose model as input. As data processing, it applies an adjustment algorithm and provides an optimized model for that device as output. This model enables operation tailored to the characteristics of the actual operating environment.
[0353] Step 4:
[0354] The terminal loads a pre-configured model and is used in real-world settings such as physical stores. It receives the pre-configured model and real-time emotional data as input. Data processing analyzes customer facial expressions and voices via camera and microphone to perform emotion recognition. The output consists of appropriate action commands tailored to the situation, specifically generating prompt messages to provide service to the customer.
[0355] Step 5:
[0356] Users monitor the device's operation and provide feedback. Inputs include device operation results and customer response data. Based on this information, users evaluate the system and create feedback that identifies areas for improvement. This feedback is sent to the server and used for subsequent model adjustments.
[0357] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0358] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0359] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0360] [Third Embodiment]
[0361] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0362] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0363] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0364] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0365] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0366] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0367] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0368] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0369] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0370] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0371] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0372] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0373] To implement this invention, first, the server constructs a physical simulation environment and generates diverse motion data. This simulation environment mimics the operation of actual machinery and enables recording of operation in various scenarios. Based on this simulation, the server executes a machine learning algorithm to construct a general-purpose motion model.
[0374] Next, the server adjusts this general-purpose model to suit the characteristics of a specific robot. This process involves fine-tuning the model according to the robot's size, power, and functional requirements. This results in a motion model that operates efficiently in a specific operating environment.
[0375] Subsequently, the robot, acting as the terminal, receives this adjusted model and begins operation. The robot can perform tasks such as moving objects on a manufacturing line. In this embodiment, the robot utilizes sensors to perceive environmental information and flexibly adapts its actions based on that information.
[0376] Users monitor the robot's movements and provide feedback as needed. This feedback can be used to fine-tune the model via the server, helping to further improve the robot's accuracy and efficiency.
[0377] As a concrete example, consider an assembly line in the manufacturing industry. The server simulates various assembly operations and builds a model based on them. The robot, acting as the terminal, uses this model to handle parts of different sizes accurately and quickly. The user can monitor the assembly process and provide feedback on accuracy and speed, further improving the process.
[0378] This configuration enables advanced robot motion control without requiring specialized programming knowledge, facilitating the development of machines that can function adaptively in various environments.
[0379] The following describes the processing flow.
[0380] Step 1:
[0381] The server sets up a physical simulation environment and generates motion data. In this process, a virtual robot model is used to design various motion scenarios and run simulations. For example, it simulates a robot arm grasping and lifting an object.
[0382] Step 2:
[0383] The server uses the generated behavioral data to execute machine learning algorithms and build a general-purpose behavioral model. Using deep learning techniques, it learns behavioral characteristics from the generated data and creates a model that can be applied to operation in various environments.
[0384] Step 3:
[0385] The server fine-tunes a general-purpose operating model to suit the specific characteristics of the machine. This process takes into account the specific operating requirements and physical characteristics of the actual machine, optimizing the model for the robot's operating environment.
[0386] Step 4:
[0387] The terminal robot receives a finely tuned motion model from the server and performs actions based on it. This model is used to efficiently perform tasks such as handling parts or assembly on a manufacturing line.
[0388] Step 5:
[0389] Users monitor the robot's movements in real time and provide feedback as needed. This feedback is used for fine-tuning to further improve the robot's accuracy and efficiency. Through this process, the overall system performance is enhanced.
[0390] (Example 1)
[0391] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0392] In modern automated systems, sophisticated motion models are necessary to flexibly adapt to different environments and conditions. However, current technology can only provide static models that depend on specific environments and conditions, making efficient and highly accurate operation difficult. Furthermore, the lack of sufficient use of feedback and readjustment functions to improve accuracy hinders improvements in versatility and efficiency.
[0393] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0394] In this invention, the server includes means for generating operational data and recording various operational patterns in a physical simulation environment, means for adjusting operational parameters of a general-purpose operational model to suit the characteristics of a specific automated device, and means for an operator to monitor the operation of the automated device, provide feedback on operational accuracy and efficiency, and use that feedback to further fine-tune the model. This makes it possible to adapt to a specific operating environment and continue to operate efficiently.
[0395] A "physical simulation environment" refers to a system or platform that generates virtual operating data by simulating real-world conditions and then analyzes that data.
[0396] "Motion data" refers to a dataset containing information about the various movements and paths performed by automated equipment, which is used for motion analysis and model construction.
[0397] A "motion model" is an algorithm or program built based on collected motion data, providing indicators or a framework for automated equipment to perform specific tasks efficiently and accurately.
[0398] An "automated device" refers to a machine or robot that can autonomously perform tasks based on a pre-set operating model.
[0399] "Operational parameters" are settings or variables that are customized to match the operating conditions and functions of automated equipment, and by being reflected in the operation model, they enable optimal operation.
[0400] An "operator" refers to a person whose role is to monitor automated equipment and its operation, and to provide feedback and make corrections as needed.
[0401] "Feedback" refers to the evaluations and suggestions provided by monitors to improve the operation of automated equipment, which then leads to the readjustment and improvement of the model.
[0402] A description of the embodiment for carrying out the invention will be provided.
[0403] First, the server constructs a physical simulation environment. This environment is used to simulate various scenarios in which real-world automated equipment operates. Specifically, it utilizes simulation software such as "Gazebo" and "Unity." Within this simulation environment, the server generates operational data and records various operational patterns. This allows for verification of operation under various conditions, independent of the environment.
[0404] Next, the server executes machine learning algorithms such as "TensorFlow" and "PyTorch" based on the generated operational data to build a general-purpose operational model. This model forms the basis for achieving efficient and highly accurate operation of automated devices in various environments.
[0405] Furthermore, the server adjusts this general operating model to suit the characteristics of the specific automated equipment. At this stage, various factors such as the size, range of motion, and power source of the equipment are taken into consideration. This adjustment optimizes the model to fit the specific operating environment.
[0406] After adjustment, the automated device, acting as the terminal, receives this motion model and starts the configured task. The automated device perceives the environment using sensors and performs the task while making decisions based on the motion model. For example, this could be a task such as accurately assembling multiple parts on a manufacturing line.
[0407] Finally, the user monitors the entire execution process and provides feedback on the accuracy and efficiency of the device's operation. This feedback data is sent to the server and used to further refine and retrain the operating model, thereby improving the device's performance.
[0408] A concrete example is a process in manufacturing assembly lines where the movements required at each stage are simulated, and the optimal assembly movements are modeled based on that data. By using this system, the efficiency and precision of the assembly line can be improved.
[0409] An example of an input prompt sentence for a generated AI model might be, "Design a behavioral model that optimizes the placement and handling of parts on a manufacturing line."
[0410] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0411] Step 1:
[0412] The server constructs a physical simulation environment. It receives simulation software configuration data as input and sets up a virtual environment. Here, it simulates the operation of a virtual automated device and records operation patterns in various environmental scenarios. As output, it generates operation data corresponding to each operation pattern. This operation data forms the basis of the system's operation model. Specifically, it reproduces the movement of arms and the movement of objects.
[0413] Step 2:
[0414] The server uses the generated behavioral data to execute a machine learning algorithm and build a general-purpose behavioral model. It receives the behavioral data generated in step 1 as input and performs data analysis. Specifically, it extracts major behavioral patterns and organizes the data to learn the most efficient behavior. This process generates a general-purpose behavioral model that can handle a variety of behaviors as output.
[0415] Step 3:
[0416] The server adapts a general-purpose motion model to the specific characteristics of the automation equipment. It uses parameter data, taking into account the equipment's physical characteristics and required performance, as input. Based on this data, the server fine-tunes the model's parameters to suit the specific operating environment. This results in a model that operates efficiently even under specific tasks and conditions. Specific operations include adjusting the motion profile according to the equipment's size and power.
[0417] Step 4:
[0418] The automated device, acting as the terminal, receives a pre-tuned motion model and initiates the actual task. It receives a pre-tuned motion model as input. Based on this model, the device utilizes sensor information to adapt its actions to the current environment. The output is the efficient execution of the specified task. Specifically, this involves autonomously handling parts and performing assembly processes on a manufacturing line.
[0419] Step 5:
[0420] Users monitor the operation of automated equipment and provide feedback on its accuracy and efficiency. Inputs include historical equipment operation data and visual monitoring information. Users analyze this data to identify areas for improvement. The feedback is sent to a server and used for model refinement and retraining. The output is a more accurate operating model. Specific actions include reviewing equipment operation logs and video feeds and reporting any necessary corrections.
[0421] (Application Example 1)
[0422] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0423] In today's industrial environment, there is a demand for diverse machinery to operate efficiently under varying conditions and for its operation to be easily adjusted. However, conventional methods have faced challenges such as requiring advanced expertise to fine-tune machine operation and lacking flexibility. In particular, the lack of mechanisms to quickly incorporate user feedback has made it difficult to improve operational efficiency.
[0424] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0425] In this invention, the server includes means for generating motion data in a physical simulation environment, means for constructing a general-purpose motion model using the generated motion data, and means for transmitting motion feedback from individual information terminals and using it for model adjustment. This enables the rapid incorporation of user-provided feedback into the motion model, resulting in efficient and flexible control of machine motion.
[0426] A "physical simulation environment" is a system that digitally mimics real-world physical phenomena and virtually reproduces the operation of machines under various conditions.
[0427] "Motion data" refers to a collection of information about the operation of a machine obtained through a physical simulation environment, and is used to construct motion models.
[0428] A "general-purpose motion model" is a model that represents motion patterns that can be commonly applied to various machines, and serves as a foundation for adapting to specific conditions or machines.
[0429] "Adjusting to the characteristics of a specific machine" means optimizing a general operating model to the performance and specifications of an individual machine and adapting it to the actual operating environment.
[0430] "Applying to actual machine operation" means implementing a tuned motion model into a machine that is actually in operation and controlling its operation in real time.
[0431] A "personal information terminal" is a digital device that a user carries with them to monitor and provide feedback on the operation of a machine.
[0432] "Sending operational feedback" is the process of a user communicating their opinions and suggestions for improvement regarding the machine's operation to a server via digital communication.
[0433] "Using it for model adjustment" means updating the motion model based on the feedback received to achieve more efficient and accurate machine operation.
[0434] The server builds a physical simulation environment and generates operational data over the network. This environment uses software to virtually reproduce real-world physical phenomena and simulate the operation of machines under different conditions. Specific software such as "Simulink" and "ANSYS," which are suitable for physical simulation and data analysis, are sometimes used.
[0435] Users monitor the machine's operation through applications installed on their individual information terminals and send feedback to the server. This feedback is used to improve the robot's accuracy and efficiency. The server receives this feedback, adapts the generated motion model using machine learning algorithms, and retrains the motion model. Representative machine learning frameworks used here include "TensorFlow" and "PyTorch."
[0436] The robot, acting as the terminal, receives an optimized motion model from the server and operates in the actual work environment. The robot uses built-in sensors to perceive environmental information in real time and performs adaptive actions based on that information. Sensor technologies such as "LiDAR" and "IMU (Inertial Measurement Unit)" are utilized in this process.
[0437] A concrete example is the operation of robots on a factory assembly line. Users monitor the assembly process using their smartphones and provide feedback to improve the speed of parts handling. The server then adjusts the operation model based on this information, enabling efficient assembly.
[0438] An example of a prompt for the generated AI model might be: "Please construct an optimal operating scenario for the robot to efficiently transport parts, and suggest ways to fine-tune the model based on user feedback."
[0439] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0440] Step 1:
[0441] The server constructs a physical simulation environment and generates machine motion data under various conditions. The inputs are virtual models and physical laws, and the output is detailed motion data obtained from the simulation. This data is generated using a physics engine to comprehensively cover the machine's motion patterns.
[0442] Step 2:
[0443] The server uses the generated motion data to build a general-purpose motion model. The input is the motion data obtained in step 1, and the output is a model of motion patterns applicable to various machines. Machine learning algorithms are used to analyze the data and perform pattern recognition and model optimization.
[0444] Step 3:
[0445] Users monitor actual machine operation and provide feedback via individual information terminals. Input is the machine operation information observed by the user, and output is the feedback data received by the server. Users use smart terminals to clearly communicate areas where improvement or adjustment of operation is needed.
[0446] Step 4:
[0447] The server adjusts its behavioral model based on the feedback it receives. The input is user feedback, and the output is the finely tuned behavioral model. Through data analysis and a retraining process, the model is updated to better suit the user's requirements.
[0448] Step 5:
[0449] The terminal robot applies a pre-tuned motion model to the actual work environment and starts operating efficiently. The input is the latest motion model provided by the server, and the output is the robot's specific action result. The robot collects environmental information through sensors and performs the optimal action based on the model in real time.
[0450] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0451] To implement this invention, the system consists of three components: a server, a terminal (robot), and a user. The server generates motion data in various scenarios using a physical simulation environment and constructs a general-purpose motion model by executing machine learning algorithms based on that data. Next, the server fine-tunes this model to suit the characteristics of a specific robot, providing a tuned model applicable to real-world motion.
[0452] The robot, acting as the terminal, uses this adjusted model to flexibly perform real-world tasks. For example, it can perform tasks such as moving objects in a manufacturing environment. The robot is equipped with an emotion engine that can recognize the user's emotions in real time. This emotional information is taken into consideration by the robot during its operations, and is used to improve work efficiency and execution accuracy.
[0453] Users can monitor the robot's movements and provide real-time feedback. The emotion engine senses the user's emotional state, such as joy or dissatisfaction, and sends this information back to the server, which then uses the feedback to further optimize the movement model. Furthermore, the robot's movement process is improved based on the user's emotional information, enhancing the quality of its work.
[0454] As a concrete example, consider a robotic assistant in a nursing care facility. This system allows the care robot to sense the emotions of the residents and perform actions that provide a sense of security. By providing services that respond to the residents' emotions, it aims to reduce their stress and anxiety and create a better user experience.
[0455] Thus, systems that incorporate emotion engines not only improve operational efficiency but also enable the provision of optimal services based on user interaction, and are expected to be used in a variety of application fields.
[0456] The following describes the processing flow.
[0457] Step 1:
[0458] The server sets up a physical simulation environment and generates motion data by simulating various motion scenarios. These scenarios include actions such as the robot grasping, moving, and assembling objects.
[0459] Step 2:
[0460] The server uses the generated behavioral data to execute machine learning algorithms and build a general-purpose behavioral model. This model can learn various behavioral patterns and has broad applicability.
[0461] Step 3:
[0462] The server fine-tunes a general-purpose motion model according to the specific robot and application. This results in an optimal model that matches the robot's actual physical characteristics and operating environment.
[0463] Step 4:
[0464] The robot, acting as the terminal, is equipped with an emotion engine that recognizes the emotional information conveyed by the user in real time. This allows the robot to select actions and optimize its response according to the user's emotions.
[0465] Step 5:
[0466] Users observe the robot's movements and see how their emotions are being communicated to the robot. Users evaluate their satisfaction with the robot's work and services, and this evaluation is fed back into the system as emotional feedback.
[0467] Step 6:
[0468] The server uses user feedback and emotion data to further refine the motion model and emotion engine. This continuously improves the robot's motion accuracy and user interaction capabilities.
[0469] Through the processing steps described above, the system provides behavior that aligns with the user's emotions, achieving both machine flexibility and an improved user experience.
[0470] (Example 2)
[0471] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0472] Efficiently and flexibly controlling the operation of automated machinery in the real world is technically challenging. In particular, constructing motion models that can handle diverse scenarios and adapting those models to the characteristics of specific devices has not been adequately achieved with conventional technologies. Solving this challenge is necessary to improve the efficiency and accuracy of automated work and enable effective interaction with users.
[0473] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0474] In this invention, the server includes means for generating information about actions in a virtual environment, means for creating a structure for mimicking general actions using the generated information about actions, and means for adapting the structure for mimicking general actions to the characteristics of a specific device. This enables efficient and flexible control of the operation in the implementation device, and also allows for further improvements utilizing real-time feedback to the user.
[0475] A "virtual environment" is a simulation system used to reproduce physical phenomena and situations in the real world on a computer.
[0476] "Information about behavior" refers to specific action data such as movements, actions, and reactions, and is a set of data that forms the basis for motion analysis and model construction of automated devices.
[0477] "A structure for mimicking general behavior" refers to algorithms and frameworks that constitute behavioral models applicable to a variety of behavioral scenarios.
[0478] "Adapting to the characteristics of a specific device" refers to the process of adjusting a general operating model to optimize it for the physical and technical constraints and characteristics of a particular device.
[0479] "Applying to the operation of a real device" means using a modified motion model on an actual machine to control its operation.
[0480] An "external observer" refers to a user or operator whose role is to observe the operation of a system or machine and to provide feedback on its performance and functionality.
[0481] "Opinions" refer to evaluations and feedback from external observers regarding the operation and results of a device, and are information that plays a part in optimizing the system.
[0482] "Recognition technology" refers to methods and techniques for improving the performance of models by incorporating information about behavior generated using machine learning algorithms and the like.
[0483] This invention is implemented by a system primarily composed of three main elements: a server, a terminal, and a user. The server generates information about actions in a virtual scenario by using a physical simulation environment. Specifically, it uses a high-performance computer and simulation software to reproduce complex physical phenomena and behavioral patterns. Using this generated information, it utilizes machine learning frameworks such as TensorFlow and PyTorch to construct a structure, or behavioral model, for simulating general behavior.
[0484] Next, the server adapts this behavioral model to the characteristics of a specific terminal device. The terminal is specifically an automated device such as a robot, which uses the adapted model received from the server to flexibly and efficiently perform real-world tasks. The robot recognizes the user's emotions in real time through a built-in emotion recognition module and optimizes its actions based on that information. The emotion recognition technology used here employs common image processing and speech recognition technologies.
[0485] The user monitors the robot's movements and provides feedback on the results. This feedback, including emotions and actions, is returned to the server and used to further optimize the motion model. This entire process enables flexible control of the device's movements and provides high-quality service to the user.
[0486] A concrete example is a robotic assistant in a nursing home. This robot improves the psychological comfort of users by identifying their emotions and providing gentle actions that alleviate anxiety. Another example of a prompt for a generative AI model is the instruction, "Please tell me the optimal action strategy for a robot to provide comfort to users in a nursing home." Through such prompts, the AI model makes specific suggestions and contributes to the overall system.
[0487] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0488] Step 1:
[0489] The server builds a physical simulation environment and generates behavioral data for virtual scenarios. Scenario parameters (e.g., gravity, friction, initial conditions) are set as input. The simulation engine uses these parameters to simulate movement and generates motion trajectories and interaction data as output. This data forms the basis for subsequent motion model construction. Specific examples include the reproduction of vehicle and robotic arm movements.
[0490] Step 2:
[0491] The server uses the generated behavioral data to build a general-purpose behavioral model using a machine learning algorithm. The behavioral data obtained in step 1 is used as input. A machine learning framework (e.g., TensorFlow) is used to train a neural network, obtaining a general-purpose behavioral model as output. This model provides a foundation capable of handling a wide variety of scenarios. Specifically, data is input into a feedforward neural network, and appropriate weights are adjusted.
[0492] Step 3:
[0493] The server fine-tunes the constructed general-purpose motion model to match the characteristics of a specific terminal device. The input is the physical and environmental characteristics of the terminal device. The server adjusts the model taking these characteristics into account, generating an optimized motion model for the specific device as output. This process allows the robot to adapt to specific work environments and tasks. As part of the adjustment, the device's sensor and actuator characteristics are incorporated into the model.
[0494] Step 4:
[0495] The terminal (robot) receives a pre-configured motion model from the server and performs the actual task. Inputs include the motion model from the server and real-time data from the real environment (e.g., camera images, distance sensor information). The robot uses this data to perform the task and generates task results and log information as output. Specifically, the robot performs actions such as actually transporting an object.
[0496] Step 5:
[0497] The user monitors the robot's movements and provides feedback. The user observes the robot's movements and task performance as input. Based on this, the user provides feedback on the quality of the movements. This feedback information is sent to the server as output. Specifically, the user evaluates the robot's smoothness and speed.
[0498] Step 6:
[0499] The server receives feedback from the user and further optimizes the behavioral model. The user feedback information is used as input. Based on this information, the server retrains its learning algorithm and adjusts the adaptive model. The output is a more improved behavioral model. Specifically, this includes readjusting weights and testing new behavioral patterns.
[0500] (Application Example 2)
[0501] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0502] In recent years, the operation of automated devices in the real world has required not only efficiency but also flexible responses that respond to the emotions of users. However, conventional systems have difficulty recognizing user emotions in real time and optimizing their operation based on that. In this situation, the challenge is to build systems that enable automated devices to provide more personalized services to users.
[0503] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0504] In this invention, the server includes means for generating behavioral data in a physical environment, means for constructing a general-purpose behavioral model using the generated behavioral data, and means for detecting and analyzing the customer's emotional state in real time. This enables the automated device to perform actions on-site that take customer emotions into consideration, allowing for the provision of more appropriate services.
[0505] The "physical environment" refers to the real-world environment with its own physical characteristics and constraints, and is the area for verifying the operation of automated devices through simulations and behavioral data.
[0506] "Behavioral data" refers to a collection of information that records the specific actions performed by automated equipment in a particular environment, along with the various parameters associated with those actions.
[0507] A "general-purpose behavioral model" refers to a predictive model that abstracts and generalizes the operation of automated devices applicable to various environments and conditions.
[0508] An "automated device" is a general term for a machine or device that can perform programmed actions autonomously, and is particularly designed to perform a specific task.
[0509] "Emotional state" refers to a person's internal psychological state and includes information such as joy and anxiety, which can be identified from cues such as facial expressions and voice.
[0510] "Methods for real-time analysis" refer to methods and technologies that process data immediately upon acquisition and obtain results, enabling rapid feedback.
[0511] In an embodiment for carrying out this invention, the system is mainly configured as follows.
[0512] The server generates behavioral data for automated devices in a physical environment and constructs a general-purpose behavioral model based on the generated data. This model is executed using machine learning libraries such as TensorFlow and PyTorch. A physical simulation environment is utilized to generate the behavioral data, resulting in a model that can handle a variety of scenarios.
[0513] The refined model will be applied to specific automated devices and is intended for use in physical stores and other similar settings. In this scenario, a device equipped with a camera and microphone will be used to detect emotional states, and real-time sentiment analysis will be performed using image processing libraries such as OpenCV. The sentiment data will be processed using a generative AI model, which will generate prompt messages as needed. Specifically, these prompts may include phrases like, "When the customer's facial expression is relaxed, recommend new products in a friendly tone."
[0514] The device uses these tuned models to perform its actual operations. If the customer shows favor, it selects the next product or service to suggest and makes recommendations in natural language. For example, if the customer smiles, it might say, "Here are some incense sticks that would go well with the aromatherapy candle you just saw."
[0515] Users are required to monitor the overall operation of the system and provide feedback. This feedback is then sent back to the server and used to improve the model. This two-way process allows the automated system to operate with increasing precision.
[0516] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0517] Step 1:
[0518] The server sets up scenarios in a physical environment and generates behavioral data. It receives parameters of the simulation environment as input and simulates the operation of automated equipment under various conditions based on these parameters. As output, it provides a set of generated behavioral data, which later serves as foundational data for training behavioral models.
[0519] Step 2:
[0520] The server constructs a general-purpose behavioral model using the generated behavioral data. It receives the behavioral data generated in step 1 as input and trains the model using a machine learning algorithm (e.g., TensorFlow or PyTorch). The output is a general-purpose behavioral model applicable under various conditions. This model predicts the basic operating patterns of automated equipment.
[0521] Step 3:
[0522] The server adapts a general-purpose behavioral model to optimize it for a specific automated device. It receives device characteristic information and the aforementioned general-purpose model as input. As data processing, it applies an adjustment algorithm and provides an optimized model for that device as output. This model enables operation tailored to the characteristics of the actual operating environment.
[0523] Step 4:
[0524] The terminal loads a pre-configured model and is used in real-world settings such as physical stores. It receives the pre-configured model and real-time emotional data as input. Data processing analyzes customer facial expressions and voices via camera and microphone to perform emotion recognition. The output consists of appropriate action commands tailored to the situation, specifically generating prompt messages to provide service to the customer.
[0525] Step 5:
[0526] Users monitor the device's operation and provide feedback. Inputs include device operation results and customer response data. Based on this information, users evaluate the system and create feedback that identifies areas for improvement. This feedback is sent to the server and used for subsequent model adjustments.
[0527] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0528] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0529] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0530] [Fourth Embodiment]
[0531] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0532] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0533] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0534] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0535] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0536] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0537] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0538] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0539] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0540] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0541] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0542] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0543] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0544] To implement this invention, first, the server constructs a physical simulation environment and generates diverse motion data. This simulation environment mimics the operation of actual machinery and enables recording of operation in various scenarios. Based on this simulation, the server executes a machine learning algorithm to construct a general-purpose motion model.
[0545] Next, the server adjusts this general-purpose model to suit the characteristics of a specific robot. This process involves fine-tuning the model according to the robot's size, power, and functional requirements. This results in a behavioral model that operates efficiently in a specific operating environment.
[0546] Subsequently, the robot, acting as the terminal, receives this adjusted model and begins operation. The robot can perform tasks such as moving objects on a manufacturing line. In this embodiment, the robot utilizes sensors to perceive environmental information and flexibly adapts its actions based on that information.
[0547] Users monitor the robot's movements and provide feedback as needed. This feedback can be used to fine-tune the model via the server, helping to further improve the robot's accuracy and efficiency.
[0548] As a concrete example, consider an assembly line in the manufacturing industry. The server simulates various assembly operations and builds a model based on them. The robot, acting as the terminal, uses this model to handle parts of different sizes accurately and quickly. The user can monitor the assembly process and provide feedback on accuracy and speed, further improving the process.
[0549] This configuration enables advanced robot motion control without requiring specialized programming knowledge, facilitating the development of machines that can function adaptively in various environments.
[0550] The following describes the processing flow.
[0551] Step 1:
[0552] The server sets up a physical simulation environment and generates motion data. In this process, a virtual robot model is used to design various motion scenarios and run simulations. For example, it simulates a robot arm grasping and lifting an object.
[0553] Step 2:
[0554] The server uses the generated behavioral data to execute machine learning algorithms and build a general-purpose behavioral model. Using deep learning techniques, it learns behavioral characteristics from the generated data and creates a model that can be applied to operation in various environments.
[0555] Step 3:
[0556] The server fine-tunes a general-purpose operating model to suit the specific characteristics of the machine. This process takes into account the specific operating requirements and physical characteristics of the actual machine, optimizing the model for the robot's operating environment.
[0557] Step 4:
[0558] The terminal robot receives a finely tuned motion model from the server and performs actions based on it. This model is used to efficiently perform tasks such as handling parts or assembly on a manufacturing line.
[0559] Step 5:
[0560] Users monitor the robot's movements in real time and provide feedback as needed. This feedback is used for fine-tuning to further improve the robot's accuracy and efficiency. Through this process, the overall system performance is enhanced.
[0561] (Example 1)
[0562] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0563] In modern automated systems, sophisticated motion models are necessary to flexibly adapt to different environments and conditions. However, current technology can only provide static models that depend on specific environments and conditions, making efficient and highly accurate operation difficult. Furthermore, the lack of sufficient use of feedback and readjustment functions to improve accuracy hinders improvements in versatility and efficiency.
[0564] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0565] In this invention, the server includes means for generating operational data and recording various operational patterns in a physical simulation environment, means for adjusting operational parameters of a general-purpose operational model to suit the characteristics of a specific automated device, and means for an operator to monitor the operation of the automated device, provide feedback on operational accuracy and efficiency, and use that feedback to further fine-tune the model. This makes it possible to adapt to a specific operating environment and continue to operate efficiently.
[0566] A "physical simulation environment" refers to a system or platform that generates virtual operating data by simulating real-world conditions and then analyzes that data.
[0567] "Motion data" refers to a dataset containing information about the various movements and paths performed by automated equipment, which is used for motion analysis and model construction.
[0568] A "motion model" is an algorithm or program built based on collected motion data, providing indicators or a framework for automated equipment to perform specific tasks efficiently and accurately.
[0569] An "automated device" refers to a machine or robot that can autonomously perform tasks based on a pre-set operating model.
[0570] "Operational parameters" are settings or variables that are customized to match the operating conditions and functions of automated equipment, and by being reflected in the operation model, they enable optimal operation.
[0571] An "operator" refers to a person whose role is to monitor automated equipment and its operation, and to provide feedback and make corrections as needed.
[0572] "Feedback" refers to the evaluations and suggestions provided by monitors to improve the operation of automated equipment, which then leads to the readjustment and improvement of the model.
[0573] A description of the embodiment for carrying out the invention will be provided.
[0574] First, the server constructs a physical simulation environment. This environment is used to simulate various scenarios in which real-world automated equipment operates. Specifically, it utilizes simulation software such as "Gazebo" and "Unity." Within this simulation environment, the server generates operational data and records various operational patterns. This allows for verification of operation under various conditions, independent of the environment.
[0575] Next, the server executes machine learning algorithms such as "TensorFlow" and "PyTorch" based on the generated operational data to build a general-purpose operational model. This model forms the basis for achieving efficient and highly accurate operation of automated devices in various environments.
[0576] Furthermore, the server adjusts this general operating model to suit the characteristics of the specific automated equipment. At this stage, various factors such as the size, range of motion, and power source of the equipment are taken into consideration. This adjustment optimizes the model to fit the specific operating environment.
[0577] After adjustment, the automated device, acting as the terminal, receives this motion model and starts the configured task. The automated device perceives the environment using sensors and performs the task while making decisions based on the motion model. For example, this could be a task such as accurately assembling multiple parts on a manufacturing line.
[0578] Finally, the user monitors the entire execution process and provides feedback on the accuracy and efficiency of the device's operation. This feedback data is sent to the server and used to further refine and retrain the operating model, thereby improving the device's performance.
[0579] A concrete example is a process in manufacturing assembly lines where the movements required at each stage are simulated, and the optimal assembly movements are modeled based on that data. By using this system, the efficiency and precision of the assembly line can be improved.
[0580] An example of an input prompt sentence for a generated AI model might be, "Design a behavioral model that optimizes the placement and handling of parts on a manufacturing line."
[0581] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0582] Step 1:
[0583] The server constructs a physical simulation environment. It receives simulation software configuration data as input and sets up a virtual environment. Here, it simulates the operation of a virtual automated device and records operation patterns in various environmental scenarios. As output, it generates operation data corresponding to each operation pattern. This operation data forms the basis of the system's operation model. Specifically, it reproduces the movement of arms and the movement of objects.
[0584] Step 2:
[0585] The server uses the generated behavioral data to execute a machine learning algorithm and build a general-purpose behavioral model. It receives the behavioral data generated in step 1 as input and performs data analysis. Specifically, it extracts major behavioral patterns and organizes the data to learn the most efficient behavior. This process generates a general-purpose behavioral model that can handle a variety of behaviors as output.
[0586] Step 3:
[0587] The server adapts a general-purpose motion model to the specific characteristics of the automation equipment. It uses parameter data, taking into account the equipment's physical characteristics and required performance, as input. Based on this data, the server fine-tunes the model's parameters to suit the specific operating environment. This results in a model that operates efficiently even under specific tasks and conditions. Specific operations include adjusting the motion profile according to the equipment's size and power.
[0588] Step 4:
[0589] The automated device, acting as the terminal, receives a pre-tuned motion model and initiates the actual task. It receives a pre-tuned motion model as input. Based on this model, the device utilizes sensor information to adapt its actions to the current environment. The output is the efficient execution of the specified task. Specifically, this involves autonomously handling parts and performing assembly processes on a manufacturing line.
[0590] Step 5:
[0591] Users monitor the operation of automated equipment and provide feedback on its accuracy and efficiency. Inputs include historical equipment operation data and visual monitoring information. Users analyze this data to identify areas for improvement. The feedback is sent to a server and used for model refinement and retraining. The output is a more accurate operating model. Specific actions include reviewing equipment operation logs and video feeds and reporting any necessary corrections.
[0592] (Application Example 1)
[0593] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0594] In today's industrial environment, there is a demand for diverse machinery to operate efficiently under varying conditions and for its operation to be easily adjusted. However, conventional methods have faced challenges such as requiring advanced expertise to fine-tune machine operation and lacking flexibility. In particular, the lack of mechanisms to quickly incorporate user feedback has made it difficult to improve operational efficiency.
[0595] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0596] In this invention, the server includes means for generating motion data in a physical simulation environment, means for constructing a general-purpose motion model using the generated motion data, and means for transmitting motion feedback from individual information terminals and using it for model adjustment. This enables the rapid incorporation of user-provided feedback into the motion model, resulting in efficient and flexible control of machine motion.
[0597] A "physical simulation environment" is a system that digitally mimics real-world physical phenomena and virtually reproduces the operation of machines under various conditions.
[0598] "Motion data" refers to a collection of information about the operation of a machine obtained through a physical simulation environment, and is used to construct motion models.
[0599] A "general-purpose motion model" is a model that represents motion patterns that can be commonly applied to various machines, and serves as a foundation for adapting to specific conditions or machines.
[0600] "Adjusting to the characteristics of a specific machine" means optimizing a general operating model to the performance and specifications of an individual machine and adapting it to the actual operating environment.
[0601] "Applying to actual machine operation" means implementing a tuned motion model into a machine that is actually in operation and controlling its operation in real time.
[0602] A "personal information terminal" is a digital device that a user carries with them to monitor and provide feedback on the operation of a machine.
[0603] "Sending operational feedback" is the process of a user communicating their opinions and suggestions for improvement regarding the machine's operation to a server via digital communication.
[0604] "Using it for model adjustment" means updating the motion model based on the feedback received to achieve more efficient and accurate machine operation.
[0605] The server builds a physical simulation environment and generates operational data over the network. This environment uses software to virtually reproduce real-world physical phenomena and simulate the operation of machines under different conditions. Specific software such as "Simulink" and "ANSYS," which are suitable for physical simulation and data analysis, are sometimes used.
[0606] Users monitor the machine's operation through applications installed on their individual information terminals and send feedback to the server. This feedback is used to improve the robot's accuracy and efficiency. The server receives this feedback, adapts the generated motion model using machine learning algorithms, and retrains the motion model. Representative machine learning frameworks used here include "TensorFlow" and "PyTorch."
[0607] The robot, acting as the terminal, receives an optimized motion model from the server and operates in the actual work environment. The robot uses built-in sensors to perceive environmental information in real time and performs adaptive actions based on that information. Sensor technologies such as "LiDAR" and "IMU (Inertial Measurement Unit)" are utilized in this process.
[0608] A concrete example is the operation of robots on a factory assembly line. Users monitor the assembly process using their smartphones and provide feedback to improve the speed of parts handling. The server then adjusts the operation model based on this information, enabling efficient assembly.
[0609] An example of a prompt for the generated AI model might be: "Please construct an optimal operating scenario for the robot to efficiently transport parts, and suggest ways to fine-tune the model based on user feedback."
[0610] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0611] Step 1:
[0612] The server constructs a physical simulation environment and generates machine motion data under various conditions. The inputs are virtual models and physical laws, and the output is detailed motion data obtained from the simulation. This data is generated using a physics engine to comprehensively cover the machine's motion patterns.
[0613] Step 2:
[0614] The server uses the generated motion data to build a general-purpose motion model. The input is the motion data obtained in step 1, and the output is a model of motion patterns applicable to various machines. Machine learning algorithms are used to analyze the data and perform pattern recognition and model optimization.
[0615] Step 3:
[0616] Users monitor actual machine operation and provide feedback via individual information terminals. Input is the machine operation information observed by the user, and output is the feedback data received by the server. Users use smart terminals to clearly communicate areas where improvement or adjustment of operation is needed.
[0617] Step 4:
[0618] The server adjusts its behavioral model based on the feedback it receives. The input is user feedback, and the output is the finely tuned behavioral model. Through data analysis and a retraining process, the model is updated to better suit the user's requirements.
[0619] Step 5:
[0620] The terminal robot applies a pre-tuned motion model to the actual work environment and starts operating efficiently. The input is the latest motion model provided by the server, and the output is the robot's specific action result. The robot collects environmental information through sensors and performs the optimal action based on the model in real time.
[0621] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0622] To implement this invention, the system consists of three components: a server, a terminal (robot), and a user. The server generates motion data in various scenarios using a physical simulation environment and constructs a general-purpose motion model by executing machine learning algorithms based on that data. Next, the server fine-tunes this model to suit the characteristics of a specific robot, providing a tuned model applicable to real-world motion.
[0623] The robot, acting as the terminal, uses this adjusted model to flexibly perform real-world tasks. For example, it can perform tasks such as moving objects in a manufacturing environment. The robot is equipped with an emotion engine that can recognize the user's emotions in real time. This emotional information is taken into consideration by the robot during its operations, and is used to improve work efficiency and execution accuracy.
[0624] Users can monitor the robot's movements and provide real-time feedback. The emotion engine senses the user's emotional state, such as joy or dissatisfaction, and sends this information back to the server, which then uses the feedback to further optimize the movement model. Furthermore, the robot's movement process is improved based on the user's emotional information, enhancing the quality of its work.
[0625] As a concrete example, consider a robotic assistant in a nursing care facility. This system allows the care robot to sense the emotions of the residents and perform actions that provide a sense of security. By providing services that respond to the residents' emotions, it aims to reduce their stress and anxiety and create a better user experience.
[0626] Thus, systems that incorporate emotion engines not only improve operational efficiency but also enable the provision of optimal services based on user interaction, and are expected to be used in a variety of application fields.
[0627] The following describes the processing flow.
[0628] Step 1:
[0629] The server sets up a physical simulation environment and generates motion data by simulating various motion scenarios. These scenarios include actions such as the robot grasping, moving, and assembling objects.
[0630] Step 2:
[0631] The server uses the generated behavioral data to execute machine learning algorithms and build a general-purpose behavioral model. This model can learn various behavioral patterns and has broad applicability.
[0632] Step 3:
[0633] The server fine-tunes a general-purpose motion model according to the specific robot and application. This results in an optimal model that matches the robot's actual physical characteristics and operating environment.
[0634] Step 4:
[0635] The robot, acting as the terminal, is equipped with an emotion engine that recognizes the emotional information conveyed by the user in real time. This allows the robot to select actions and optimize its response according to the user's emotions.
[0636] Step 5:
[0637] Users observe the robot's movements and see how their emotions are being communicated to the robot. Users evaluate their satisfaction with the robot's work and services, and this evaluation is fed back into the system as emotional feedback.
[0638] Step 6:
[0639] The server uses user feedback and emotion data to further refine the motion model and emotion engine. This continuously improves the robot's motion accuracy and user interaction capabilities.
[0640] Through the processing steps described above, the system provides behavior that aligns with the user's emotions, achieving both machine flexibility and an improved user experience.
[0641] (Example 2)
[0642] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0643] Efficiently and flexibly controlling the operation of automated machinery in the real world is technically challenging. In particular, constructing motion models that can handle diverse scenarios and adapting those models to the characteristics of specific devices has not been adequately achieved with conventional technologies. Solving this challenge is necessary to improve the efficiency and accuracy of automated work and enable effective interaction with users.
[0644] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0645] In this invention, the server includes means for generating information about actions in a virtual environment, means for creating a structure for mimicking general actions using the generated information about actions, and means for adapting the structure for mimicking general actions to the characteristics of a specific device. This enables efficient and flexible control of the operation in the implementation device, and also allows for further improvements utilizing real-time feedback to the user.
[0646] A "virtual environment" is a simulation system used to reproduce physical phenomena and situations in the real world on a computer.
[0647] "Information about behavior" refers to specific action data such as movements, actions, and reactions, and is a set of data that forms the basis for motion analysis and model construction of automated devices.
[0648] "A structure for mimicking general behavior" refers to algorithms and frameworks that constitute behavioral models applicable to a variety of behavioral scenarios.
[0649] "Adapting to the characteristics of a specific device" refers to the process of adjusting a general operating model to optimize it for the physical and technical constraints and characteristics of a particular device.
[0650] "Applying to the operation of a real device" means using a modified motion model on an actual machine to control its operation.
[0651] An "external observer" refers to a user or operator whose role is to observe the operation of a system or machine and to provide feedback on its performance and functionality.
[0652] "Opinions" refer to evaluations and feedback from external observers regarding the operation and results of a device, and are information that plays a part in optimizing the system.
[0653] "Recognition technology" refers to methods and techniques for improving the performance of models by incorporating information about behavior generated using machine learning algorithms and the like.
[0654] This invention is implemented by a system primarily composed of three main elements: a server, a terminal, and a user. The server generates information about actions in a virtual scenario by using a physical simulation environment. Specifically, it uses a high-performance computer and simulation software to reproduce complex physical phenomena and behavioral patterns. Using this generated information, it utilizes machine learning frameworks such as TensorFlow and PyTorch to construct a structure, or behavioral model, for simulating general behavior.
[0655] Next, the server adapts this behavioral model to the characteristics of a specific terminal device. The terminal is specifically an automated device such as a robot, which uses the adapted model received from the server to flexibly and efficiently perform real-world tasks. The robot recognizes the user's emotions in real time through a built-in emotion recognition module and optimizes its actions based on that information. The emotion recognition technology used here employs common image processing and speech recognition technologies.
[0656] The user monitors the robot's movements and provides feedback on the results. This feedback, including emotions and actions, is returned to the server and used to further optimize the motion model. This entire process enables flexible control of the device's movements and provides high-quality service to the user.
[0657] A concrete example is a robotic assistant in a nursing home. This robot improves the psychological comfort of users by identifying their emotions and providing gentle actions that alleviate anxiety. Another example of a prompt for a generative AI model is the instruction, "Please tell me the optimal action strategy for a robot to provide comfort to users in a nursing home." Through such prompts, the AI model makes specific suggestions and contributes to the overall system.
[0658] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0659] Step 1:
[0660] The server builds a physical simulation environment and generates behavioral data for virtual scenarios. Scenario parameters (e.g., gravity, friction, initial conditions) are set as input. The simulation engine uses these parameters to simulate movement and generates motion trajectories and interaction data as output. This data forms the basis for subsequent motion model construction. Specific examples include the reproduction of vehicle and robotic arm movements.
[0661] Step 2:
[0662] The server uses the generated behavioral data to build a general-purpose behavioral model using a machine learning algorithm. The behavioral data obtained in step 1 is used as input. A machine learning framework (e.g., TensorFlow) is used to train a neural network, obtaining a general-purpose behavioral model as output. This model provides a foundation capable of handling a wide variety of scenarios. Specifically, data is input into a feedforward neural network, and appropriate weights are adjusted.
[0663] Step 3:
[0664] The server fine-tunes the constructed general-purpose motion model to match the characteristics of a specific terminal device. The input is the physical and environmental characteristics of the terminal device. The server adjusts the model taking these characteristics into account, generating an optimized motion model for the specific device as output. This process allows the robot to adapt to specific work environments and tasks. As part of the adjustment, the device's sensor and actuator characteristics are incorporated into the model.
[0665] Step 4:
[0666] The terminal (robot) receives a pre-configured motion model from the server and performs the actual task. Inputs include the motion model from the server and real-time data from the real environment (e.g., camera images, distance sensor information). The robot uses this data to perform the task and generates task results and log information as output. Specifically, the robot performs actions such as actually transporting an object.
[0667] Step 5:
[0668] The user monitors the robot's movements and provides feedback. The user observes the robot's movements and task performance as input. Based on this, the user provides feedback on the quality of the movements. This feedback information is sent to the server as output. Specifically, the user evaluates the robot's smoothness and speed.
[0669] Step 6:
[0670] The server receives feedback from the user and further optimizes the behavioral model. The user feedback information is used as input. Based on this information, the server retrains its learning algorithm and adjusts the adaptive model. The output is a more improved behavioral model. Specifically, this includes readjusting weights and testing new behavioral patterns.
[0671] (Application Example 2)
[0672] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0673] In recent years, the operation of automated devices in the real world has required not only efficiency but also flexible responses that respond to the emotions of users. However, conventional systems have difficulty recognizing user emotions in real time and optimizing their operation based on that. In this situation, the challenge is to build systems that enable automated devices to provide more personalized services to users.
[0674] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0675] In this invention, the server includes means for generating behavioral data in a physical environment, means for constructing a general-purpose behavioral model using the generated behavioral data, and means for detecting and analyzing the customer's emotional state in real time. This enables the automated device to perform actions on-site that take customer emotions into consideration, allowing for the provision of more appropriate services.
[0676] The "physical environment" refers to the real-world environment with its own physical characteristics and constraints, and is the area for verifying the operation of automated devices through simulations and behavioral data.
[0677] "Behavioral data" refers to a collection of information that records the specific actions performed by automated equipment in a particular environment, along with the various parameters associated with those actions.
[0678] A "general-purpose behavioral model" refers to a predictive model that abstracts and generalizes the operation of automated devices applicable to various environments and conditions.
[0679] An "automated device" is a general term for a machine or device that can perform programmed actions autonomously, and is particularly designed to perform a specific task.
[0680] "Emotional state" refers to a person's internal psychological state and includes information such as joy and anxiety, which can be identified from cues such as facial expressions and voice.
[0681] "Methods for real-time analysis" refer to methods and technologies that process data immediately upon acquisition and obtain results, enabling rapid feedback.
[0682] In an embodiment for carrying out this invention, the system is mainly configured as follows.
[0683] The server generates behavioral data for automated devices in a physical environment and constructs a general-purpose behavioral model based on the generated data. This model is executed using machine learning libraries such as TensorFlow and PyTorch. A physical simulation environment is utilized to generate the behavioral data, resulting in a model that can handle a variety of scenarios.
[0684] The refined model will be applied to specific automated devices and is intended for use in physical stores and other similar settings. In this scenario, a device equipped with a camera and microphone will be used to detect emotional states, and real-time sentiment analysis will be performed using image processing libraries such as OpenCV. The sentiment data will be processed using a generative AI model, which will generate prompt messages as needed. Specifically, these prompts may include phrases like, "When the customer's facial expression is relaxed, recommend new products in a friendly tone."
[0685] The device uses these tuned models to perform its actual operations. If the customer shows favor, it selects the next product or service to suggest and makes recommendations in natural language. For example, if the customer smiles, it might say, "Here are some incense sticks that would go well with the aromatherapy candle you just saw."
[0686] Users are required to monitor the overall operation of the system and provide feedback. This feedback is then sent back to the server and used to improve the model. This two-way process allows the automated system to operate with increasing precision.
[0687] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0688] Step 1:
[0689] The server sets up scenarios in a physical environment and generates behavioral data. It receives parameters of the simulation environment as input and simulates the operation of automated equipment under various conditions based on these parameters. As output, it provides a set of generated behavioral data, which later serves as foundational data for training behavioral models.
[0690] Step 2:
[0691] The server constructs a general-purpose behavioral model using the generated behavioral data. It receives the behavioral data generated in step 1 as input and trains the model using a machine learning algorithm (e.g., TensorFlow or PyTorch). The output is a general-purpose behavioral model applicable under various conditions. This model predicts the basic operating patterns of automated equipment.
[0692] Step 3:
[0693] The server adapts a general-purpose behavioral model to optimize it for a specific automated device. It receives device characteristic information and the aforementioned general-purpose model as input. As data processing, it applies an adjustment algorithm and provides an optimized model for that device as output. This model enables operation tailored to the characteristics of the actual operating environment.
[0694] Step 4:
[0695] The terminal loads a pre-configured model and is used in real-world settings such as physical stores. It receives the pre-configured model and real-time emotional data as input. Data processing analyzes customer facial expressions and voices via camera and microphone to perform emotion recognition. The output consists of appropriate action commands tailored to the situation, specifically generating prompt messages to provide service to the customer.
[0696] Step 5:
[0697] Users monitor the device's operation and provide feedback. Inputs include device operation results and customer response data. Based on this information, users evaluate the system and create feedback that identifies areas for improvement. This feedback is sent to the server and used for subsequent model adjustments.
[0698] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0699] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0700] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0701] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0702] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0703] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0704] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0705] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0706] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0707] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0708] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0709] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0710] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0711] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0712] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0713] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0714] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0715] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0716] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0717] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0718] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0719] The following is further disclosed regarding the embodiments described above.
[0720] (Claim 1)
[0721] A means for generating operational data in a physical simulation environment,
[0722] A means for constructing a general-purpose motion model using the generated motion data,
[0723] A means of adjusting a general-purpose operating model to suit the characteristics of a specific machine,
[0724] Means for applying the adjusted model to actual machine operation,
[0725] A system that includes means for the user to monitor the results of machine operation and provide feedback.
[0726] (Claim 2)
[0727] The system according to claim 1, which enables flexible motion control of a machine using the adjusted model.
[0728] (Claim 3)
[0729] The system according to claim 1, which runs a machine learning algorithm based on generated motion data and trains a model.
[0730] "Example 1"
[0731] (Claim 1)
[0732] A means for generating motion data in a physical simulation environment and recording various motion patterns,
[0733] A means of executing a machine learning algorithm to construct a general-purpose behavior model using the generated behavioral data,
[0734] A means for adjusting the operating parameters of a general operating model to suit the characteristics of a specific automated device,
[0735] A means for applying the adjusted model to the actual automated equipment's operation and starting it up,
[0736] A system that includes means for an operator to monitor the operation of automated equipment, provide feedback on operational accuracy and efficiency, and use that feedback to further refine the model.
[0737] (Claim 2)
[0738] The system according to claim 1, which enables flexible operation control of an automated device using the adjusted model.
[0739] (Claim 3)
[0740] The system according to claim 1, which implements a machine learning algorithm based on generated motion data, trains a model, and further retrains it based on feedback to adapt it to a specific work environment.
[0741] "Application Example 1"
[0742] (Claim 1)
[0743] A means for generating operational data in a physical simulation environment,
[0744] A means for constructing a general-purpose motion model using the generated motion data,
[0745] A means of adjusting a general-purpose operating model to suit the characteristics of a specific machine,
[0746] Means for applying the adjusted model to actual machine operation,
[0747] A means for the user to monitor the machine's operation results and provide feedback,
[0748] A system that includes means for transmitting operational feedback from individual information terminals and using it for model adjustment.
[0749] (Claim 2)
[0750] The system according to claim 1, which uses the adjusted model to achieve flexible motion control of a machine and reflects improvements based on remote instructions from a user.
[0751] (Claim 3)
[0752] The system according to claim 1, which runs a machine learning algorithm based on generated motion data, trains a motion model, and optimizes it in response to user input.
[0753] "Example 2 of combining an emotion engine"
[0754] (Claim 1)
[0755] A means for generating information about actions in a virtual environment,
[0756] A means for creating a structure to mimic general behavior using generated information about behavior,
[0757] Means for adapting a structure for mimicking general behavior to the characteristics of a specific device,
[0758] Means for applying the adapted structure to the operation of a real device,
[0759] A system that includes means for an external observer to monitor the operating results of a device and provide feedback.
[0760] (Claim 2)
[0761] The system according to claim 1, which enables flexible operation control of the device using the adapted structure described above.
[0762] (Claim 3)
[0763] The system according to claim 1, which uses recognition technology to enhance a structure based on information about generated behavior.
[0764] "Application example 2 when combining with an emotional engine"
[0765] (Claim 1)
[0766] A means of generating behavioral data in a physical environment,
[0767] A means of constructing a general-purpose behavioral model using generated behavioral data,
[0768] A means for adjusting a general-purpose behavioral model to suit the characteristics of a specific automated device,
[0769] Means for applying the adjusted model to the actual operation of automated equipment,
[0770] A means to detect and analyze the emotional state of customers in real time,
[0771] A system that includes means for monitoring the operational results of automated equipment and providing feedback based on customer sentiment.
[0772] (Claim 2)
[0773] The system according to claim 1, which uses the adjusted model to achieve flexible behavioral control of an automated device and optimize customer service.
[0774] (Claim 3)
[0775] The system according to claim 1, which runs a learning algorithm based on generated behavioral data to train a model. [Explanation of Symbols]
[0776] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for generating operational data in a physical simulation environment, A means for constructing a general-purpose motion model using the generated motion data, A means of adjusting a general-purpose operating model to suit the characteristics of a specific machine, Means for applying the adjusted model to actual machine operation, A system that includes means for the user to monitor the results of machine operation and provide feedback.
2. The system according to claim 1, which enables flexible motion control of a machine using the adjusted model.
3. The system according to claim 1, which executes a machine learning algorithm based on generated motion data and trains a model.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A