Information processing system
By preprocessing real-time sensor data and using generative artificial intelligence models to predict future behavior, the problems of data noise and outliers in existing technologies have been solved, enabling efficient risk assessment and timely alarms for industrial equipment behavior, thereby improving safety and management efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2025-10-16
- Publication Date
- 2026-04-21
AI Technical Summary
Existing technologies suffer from significant noise and numerous outliers when dealing with large-scale real-time sensor data, making it impossible to achieve high-precision prediction of complex temporal behaviors. This results in the inability to identify abnormal or dangerous behaviors in a timely manner, affecting the safety of operators and equipment as well as production efficiency.
Generative artificial intelligence models are used to complete missing values, remove outliers, and normalize real-time collected time-series data. Combined with preset prompts, future behavior is predicted, and risk assessment is performed based on the prediction results. Personalized warning messages are automatically generated and notified to users in a multimodal manner.
It enables accurate prediction of device behavior and timely intelligent alarms in dynamic environments, improving safety and management efficiency in the industrial sector.
Smart Images

Figure CN121901562A_ABST
Abstract
Description
Technical Field
[0001] The technology disclosed herein relates to an information processing system. Background Technology
[0002] Japanese Patent Application Publication No. 2022-180282 discloses a method for controlling a role-based chatbot executed by at least one processor. The method includes the following steps: receiving a user's speech; adding the user's speech to a prompt word, the prompt word containing instruction statements associated with an explanation of the chatbot's role; encoding the prompt word; and inputting the encoded prompt word into a language model to generate a chatbot response to the user's speech. Summary of the Invention
[0003] This invention proposes an information processing system, comprising: a device for collecting real-time data from a robot; a device for preprocessing the collected data; a device for predicting the robot's future behavior using the preprocessed data; a device for performing risk assessment based on the predicted behavior; a device for generating an alarm based on the risk assessment results; and a device for notifying the user of the generated alarm. Through the synergistic effect of the above devices, real-time prediction and risk analysis of robot behavior can be achieved, and early warning information can be promptly conveyed to the user, helping the user take appropriate measures to effectively prevent risk events from occurring, thereby improving the overall safety and reliability of the system.
[0004] A "robot" is an intelligent mechanical device that can automatically perform tasks and interact with the external environment through sensors.
[0005] "Real-time data" refers to data that reflects the robot's current state and environmental conditions, and is collected and transmitted within a short period of time.
[0006] "Data collection device" refers to a hardware or software system used to acquire various sensor information, working status and environmental data from a robot.
[0007] "Preprocessing" refers to the process of cleaning, denoising, and standardizing the raw collected data in order to facilitate further analysis and utilization.
[0008] "A device for predicting future behavior" refers to a system that can determine and output the robot's possible future actions or states based on historical and current data, through data analysis or artificial intelligence algorithms.
[0009] "Time series data" refers to a data sequence arranged in chronological order, used to show the changes of an object over a continuous period of time.
[0010] A "risk assessment device" is a system that calculates and analyzes the potential dangers of a robot based on its predicted future behavior.
[0011] "Risk score" is a quantitative value that reflects the level of risk and is used to evaluate the safety of robot behavior.
[0012] An "alarm" refers to a reminder or warning message that is automatically generated and provided to relevant personnel or systems when the system detects a potential risk.
[0013] "A device for notifying users" refers to a device or system used to transmit alarm information to the robot's manager or operator via a terminal, interface, or other communication method. Attached Figure Description
[0014] Figure 1 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the first embodiment.
[0015] Figure 2 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and smart device according to the first embodiment.
[0016] Figure 3 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the second embodiment.
[0017] Figure 4 This is a conceptual diagram illustrating an example of the main functions of the data processing device and smart glasses according to the second embodiment.
[0018] Figure 5 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the third embodiment.
[0019] Figure 6 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and head-mounted terminal according to the third embodiment.
[0020] Figure 7 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the fourth embodiment.
[0021] Figure 8 This is a conceptual diagram illustrating an example of the main functions of the data processing device and robot according to the fourth embodiment.
[0022] Figure 9 This represents an emotion map that maps multiple emotions.
[0023] Figure 10 This represents an emotion map that maps multiple emotions.
[0024] Figure 11This is a sequence diagram illustrating the processing flow of the data processing system of Embodiment 1.
[0025] Figure 12 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 1.
[0026] Figure 13 This is a sequence diagram illustrating the processing flow of the data processing system in Embodiment 2.
[0027] Figure 14 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 2. Detailed Implementation
[0028] Hereinafter, an example of an implementation of the system according to the present disclosure will be described with reference to the accompanying drawings.
[0029] First, let me explain the terminology used in the following instructions.
[0030] In the following embodiments, the processor (hereinafter referred to as "processor") with reference numerals may be a single computing device or a combination of multiple computing devices. Furthermore, the processor may be a single computing device or a combination of multiple computing devices. Examples of computing devices include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.
[0031] In the following embodiments, RAM (Random Access Memory), as indicated in the figures, is a memory that temporarily stores information and is used as working memory by the processor.
[0032] In the following embodiments, the memory, as indicated by the reference numerals, is one or more non-volatile storage devices that store various programs and parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), disks (e.g., hard disks), or magnetic tapes.
[0033] In the following embodiments, the communication I / F (Interface) with reference numerals is an interface that includes a communication processor and an antenna, etc. The communication I / F is responsible for communication between multiple computers. As an example of a communication specification applicable to the communication I / F, wireless communication specifications such as 5G (5th Generation Mobile Communication System), Wi-Fi (wireless fidelity) (registered trademark), or Bluetooth (registered trademark) can be listed.
[0034] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it can be only A, only B, or a combination of A and B. Furthermore, in this specification, when "and / or" is used to connect and express more than three items, the same interpretation as "A and / or B" applies.
[0035] First Implementation Method Figure 1 An example of the configuration of the data processing system 10 according to the first embodiment is shown.
[0036] like Figure 1 As shown, the data processing system 10 includes a data processing device 12 and an intelligent device 14. A server can be cited as an example of the data processing device 12.
[0037] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0038] The smart device 14 includes a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. In addition, the receiving device 38, output device 40, camera 42, and communication I / F 44 are also connected to the bus 52.
[0039] The receiving device 38 includes a touchscreen 38A and a microphone 38B, and receives user input. The touchscreen 38A receives user input via touch by detecting contact with an indicator (e.g., a pen or finger). The microphone 38B receives user input via sound by detecting the user's voice. The control unit 46A in the processor 46 sends data representing the user input received by the touchscreen 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data representing the user input.
[0040] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting data in a form perceptible to the user 20 (e.g., sound and / or text). The display 40A displays visual information such as text and images according to instructions from the processor 46. The speaker 40B outputs sound according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0041] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for sending and receiving various information between processor 46 and processor 28 via network 54.
[0042] Figure 2 The diagram shows an example of the main functions of the data processing device 12 and the smart device 14.
[0043] like Figure 2 As shown, in the data processing apparatus 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the memory 32. The specific processing program 56 is an example of a "program" as understood in this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0044] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).
[0045] In the smart device 14, the processor 46 performs the acceptance output processing. The memory 50 stores the acceptance output program 60. The acceptance output program 60 is used in conjunction with the data processing system 10 and the specific processing program 56. The processor 46 reads the acceptance output program 60 from the memory 50 and executes the read acceptance output program 60 on the RAM 48. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48. Furthermore, the smart device 14 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48.
[0046] Alternatively, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. Furthermore, the data processing device 12 may be a server device or a user-held terminal device (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of the processing of the data processing system 10 of the first embodiment will be described.
[0047] Example 1 The flow of a specific process in Example 1 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. Furthermore, the data processing device 12 is referred to as the "server," and the smart device 14 is referred to as the "terminal."
[0048] With the development of smart factories and automated equipment, traditional data acquisition and risk management systems have many shortcomings in responding to dynamic environmental changes, real-time risk prediction, and efficient alarm response. Specifically, existing technologies suffer from significant data noise and numerous outliers when dealing with large-scale real-time sensor data. Furthermore, their ability to identify and accurately predict complex temporal behaviors is limited, making it impossible to provide timely warnings of abnormal or dangerous behaviors, thus hindering the safety of operators and equipment and production efficiency. Therefore, there is an urgent need for an information processing system capable of automatically and efficiently preprocessing various temporal information, performing high-order behavior prediction based on generative artificial intelligence models, and quantitatively assessing risks and providing adaptive alarms.
[0049] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 1 is achieved by the following means.
[0050] In this invention, the server includes a device for real-time acquisition of time-series data from the physical information generation device; a device for preprocessing the time-series data, such as filling in missing values, removing outliers, and normalizing; a device for inputting the processed data and predefined prompts into a generative artificial intelligence model to predict future behavior; a device for calculating a quantitative risk score based on the prediction results and standard safety benchmarks and generating personalized warning messages; and a device for notifying the warning message to the information prompting terminal and visually or audibly alerting the recipient. This enables accurate prediction of the future behavior of equipment in dynamic environments, efficient quantitative assessment of risks, and timely intelligent alarms, thereby effectively improving safety and management efficiency in the industrial field.
[0051] "Physical information generation device" refers to a hardware device or combined system that can collect, generate and output environmental or equipment-related data such as location, temperature, humidity and pressure.
[0052] "Time sequence information" refers to a data sequence arranged in chronological order, which usually reflects the changes of the monitored object over time.
[0053] "Preprocessing" refers to operations performed on raw data, such as filling in missing values, removing outliers, normalizing, and smoothing, in order to improve the accuracy and reliability of subsequent analysis.
[0054] "Generative AI models" refer to AI analysis tools that can autonomously generate predictions about future behaviors, states, or trends based on input prompts and structured data.
[0055] "Prompt statements" refer to the informational guidance text input into generative artificial intelligence models, used to clarify modeling goals and data analysis directions.
[0056] "Future behavior prediction" refers to the speculation and calculation of the actions or states that a physical information generating device may take at a later time.
[0057] "Standard benchmarks" refer to reference values or criteria for judging risk levels, established based on industry norms, equipment safety standards, etc.
[0058] "Risk assessment" refers to the process of quantitatively or qualitatively analyzing the probability and impact of possible abnormal or dangerous events in the future by combining prediction results with standard benchmarks.
[0059] A "warning message" is a message generated based on risk assessment results, used to alert recipients to potential dangers and guide them in taking appropriate measures.
[0060] "Information alert device" refers to a terminal device that can present warning messages to users in a visual or audible manner, such as a monitor, tablet, mobile phone or alarm.
[0061] "Receiver" refers to a user or administrator who can receive warning messages through an information notification device and respond accordingly.
[0062] This invention relates to a system for predicting and risk-assessing the future behavior of a physical information generation device using a generative artificial intelligence model. Typical embodiments of this system are described in detail below.
[0063] This system comprises servers, terminal devices, and users. The server is the core of the entire system, responsible for data collection, analysis, and alert generation. The server can be a general-purpose computing device with high-performance computing capabilities, such as a server equipped with a multi-core processor and GPU, and the operating system can be Linux. On the software side, Python can be used as the basic language for data processing and AI model invocation; MySQL or InfluxDB can be used as the database; the software components for sensor data acquisition can be a Socket communication module or an MQTT protocol library; data preprocessing can be performed using toolkits such as Pandas and NumPy; generative AI models can be locally deployed LSTM or Transformer models, or cloud-based generative AI (such as GPT-4 services) can be invoked through open APIs.
[0064] The server communicates with physical information generating devices (such as industrial robots and sensors in automated equipment) via a network to collect various time-series information, including location, temperature, humidity, and motion status. The server automatically performs preprocessing operations on the collected raw data, such as missing value completion, outlier removal, and normalization, to ensure data quality. Then, the server inputs the preprocessed data, along with predefined prompts, into a generative artificial intelligence model to predict future behavior. The model output includes information such as the device's motion trend over several subsequent time points, possible abnormal events, and their probabilities.
[0065] The server performs a quantitative risk assessment based on future behavior predictions generated by a generative artificial intelligence model, as well as industry safety standards and equipment operation specifications. For predictions exceeding a predetermined risk threshold, the server automatically generates a warning message containing the risk type, a detailed description, and recommended measures. The server then promptly delivers the warning message to the terminal device via WebSocket or a push notification API.
[0066] Terminal devices (such as industrial tablets, smartphones, etc.) receive warning messages pushed by the server and present the warning information to the user intuitively through graphical interfaces, pop-ups, sounds, etc., to ensure that the user can learn about potential risks and corresponding operation suggestions as soon as possible.
[0067] Based on the information displayed on the terminal, users can take necessary measures such as emergency shutdown, equipment inspection, and personnel evacuation in a timely manner, and can also report the response results to the server through the terminal to achieve closed-loop information management.
[0068] A specific example of this embodiment is as follows: The server collects time-series data such as position, angle, and speed from the factory robot every second. After preprocessing, the data is input into the LSTM model. Examples of prompts set by the server include: "Based on the following robot motion data, analyze whether there is any abnormal behavior in the next 5 seconds." or "Based on this time-series sensor data, predict the risk level of the robot's next move and explain the reason." The server compares the model's predicted output with the factory safety benchmark, calculates the risk score, and pushes the warning content to the operator's tablet terminal, which displays: "Warning: The robot may approach a crowd at high speed in the future. Please take immediate action." The user then responds accordingly.
[0069] Through the above implementation methods, the system can efficiently and intelligently monitor the dynamics of physical devices, improving the safety and management efficiency of equipment operation.
[0070] use Figure 11 The processing procedure is explained.
[0071] Step 1: The server collects real-time time-series data from physical information generating devices (such as robots or sensors) via a network interface. Inputs include various types of raw sensor data (such as position, temperature, velocity, acceleration, etc.), and actions include sending data requests to the data source and parsing returned JSON or binary data packets. Outputs are raw time-series data arranged with timestamps and saved to a database.
[0072] Step 2: The server preprocesses the stored raw time-series data. The input is the latest batch of raw data from the database. Specific actions include using Pandas for missing value completion, outlier removal, moving average smoothing, and normalization. Appropriate algorithms are selected based on the input data type; for example, temperature data is standardized to 0-1, and outliers are imputed using interpolation. The output is a high-quality, uniformly formatted time-series dataset, stored in memory or a database cache.
[0073] Step 3: The server inputs preprocessed time-series data and predefined prompts into the generative AI model. The input consists of formatted time-series data and text prompts (e.g., "Please predict whether the robot will exhibit dangerous behavior in the next 5 seconds"). Specific actions include constructing a model request body, packaging the data into a structure conforming to the model interface format, and calling the local or cloud-based generative AI model via API. The data is processed into embedded vectors, and the model generates behavioral predictions for several future time points based on the input. The output includes predictions such as the behavior trend and anomaly probability for the corresponding future time period.
[0074] Step 4: The server assesses the risk of future actions based on the output of the artificial intelligence model. The input is the model's predicted output (such as the risk probability at each time point). Specific actions include comparing the predicted behavior with safety standards and calculating a quantitative risk score; for example, if the robot is less than a safety threshold away from a person and its speed is high, it is considered high-risk. Data processing involves extracting key features, performing logical judgments, and assigning scores. The output is the risk assessment result, including the risk level, probability of occurrence, and a detailed description.
[0075] Step 5: The server generates a warning message based on the risk assessment results. The input is the detailed output of the risk assessment. Specific actions include assembling a text message, including the risk type, suggested actions, and the time of occurrence. The server saves the message as a push notification. The output is a structured warning data packet for display on the terminal.
[0076] Step 6: The server pushes warning messages to end devices via push protocols such as WebSocket, REST API, or message queues. The input is the warning data to be sent. Specific actions include calling the push service API to send the data to a specific end device ID or user group. The output is the warning message arriving at the target end device's storage queue.
[0077] Step 7: Upon receiving a warning message, the terminal immediately alerts the user via pop-up windows, color highlighting, beeping, or voice announcement. The input is the warning message data packet pushed by the server. Specifically, after parsing the message content, it displays a warning notification on the screen, plays a warning sound, or vibrates. The output is a visually or audibly displayed safety warning interface for the user.
[0078] Step 8: After seeing the terminal prompt, the user takes action based on the warning. The input is the warning content displayed on the terminal. Specific actions may include pressing the device's emergency stop button, inspecting the site, evacuating personnel, etc., and the user can receive a "handled" confirmation message on the terminal. The output is the on-site risk intervention action and user feedback record, which is sent to the server for subsequent management and recording.
[0079] Application Example 1 The process flow corresponding to the specific processing in Use Case 1 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. Furthermore, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".
[0080] With the increasing level of industrial automation, the number of robots and the complexity of their operations in factories are constantly rising. Real-time monitoring of the behavior of multiple robots, timely identification of abnormal actions, rapid risk assessment, and timely, safe, and intuitive delivery of targeted warnings and response suggestions to operators have become urgent challenges. Existing technologies lack an efficient overall system for data preprocessing, intelligent prediction, flexible push notifications, and user interaction, leading to delayed risk response, inaccurate information delivery, and increased user workload, thus impacting production safety and efficiency. Furthermore, traditional systems cannot dynamically adjust notification methods based on user emotional states, hindering improved user experience and accurate assurance in emergency scenarios.
[0081] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 1 is achieved by the following means.
[0082] In this invention, the server includes various devices and functional modules for real-time acquisition of robot time-series data, intelligent data preprocessing, predicting future actions through generative artificial intelligence models, conducting risk assessments based on prediction results and generating warnings and recommended measures, pushing multimodal notifications to user terminals via the network, monitoring and analyzing user status and emotionally adapting notification content and methods, and collecting user response feedback to optimize system performance. This enables intelligent dynamic perception and prediction of robot behavior in the factory area, automated risk identification and emergency early warning, and provides customized human-machine interaction experiences based on user on-site conditions, significantly improving production safety, reliability, and management efficiency.
[0083] "Data acquisition device" refers to a hardware or software system that can acquire information such as robot position, speed, status, and environment in real time.
[0084] "Time series information" refers to a set of data arranged in chronological order that reflects the changes in the state of things over time.
[0085] "Data preprocessing algorithms" refer to data processing methods used to remove noise, correct anomalies or missing data, standardize raw data, and generate suitable structures for subsequent analysis.
[0086] "Generative AI models" refer to AI analysis models that can generate or predict possible future scenarios and action sequences by learning from a large amount of historical data.
[0087] "Future actions" refers to the behavior and state changes that robots and other objects may undergo in the future, based on current and historical data.
[0088] "Risk assessment" refers to the process of determining the hazard of the prediction results with reference to safety benchmarks, quantifying the scores, and analyzing feasible countermeasures.
[0089] "Risk score" refers to a numerical indicator that reflects the likelihood of a system or object facing an unsafe or abnormal state.
[0090] "Notification information" refers to a message automatically generated by the system based on the assessment results, which includes alarm descriptions and response suggestions, and is sent to users via the network.
[0091] "Network communication path" refers to the communication infrastructure or protocol that enables the transmission of data and information between devices such as servers and terminals.
[0092] "User terminal" refers to a smart mobile or desktop device used to receive and display notification information and to allow users to interact with it.
[0093] "Visual notification" refers to delivering information to users through visual means such as screens and indicator lights.
[0094] "Audio notification" refers to reminding users of information through auditory means such as sound, voice, or alarm sounds.
[0095] "Haptic notification" refers to issuing alerts to users through vibration, force feedback, or other means.
[0096] The "user interface" refers to the software interface through which users interact with the system, view notifications, and perform operational controls.
[0097] A "user state recognition engine" refers to an algorithm or module that identifies a user's emotions and state by analyzing the user's input information such as voice, facial expressions, and operations.
[0098] "Optimizing notification content and methods" refers to adjusting the content and delivery format of alerts and suggestions issued by the system based on dynamic factors such as user emotions and the on-site environment.
[0099] "User response history" refers to the historical records of a user's actions and feedback after receiving a system notification.
[0100] "Feedback information" refers to information such as user evaluations, suggestions, or responses to system prompts.
[0101] This invention relates to a real-time monitoring and risk management system for industrial robots based on a generative artificial intelligence model. The system consists of a server, a terminal, and a user. Through data collection, intelligent prediction, risk assessment, and multimodal notification, it significantly improves the safety and operational efficiency of factory robots.
[0102] First, the server is equipped with a data acquisition module that connects to multiple industrial robots via industrial Ethernet or wireless communication protocols, enabling it to collect various time-series information in real time, including position, speed, attitude, and environmental sensing data. The server uses a high-speed processor and a large-capacity storage device, and runs the Linux operating system.
[0103] The server utilizes AI platforms such as TensorFlow and PyTorch to preprocess the collected raw data, including noise removal, normalization, and time series completion. Then, the server inputs the cleaned time-series information into a generative AI model (such as LSTM or Transformer) trained on historical data, enabling high-precision predictions of the robot's possible actions and scenarios within the next few minutes. The server also incorporates a risk assessment algorithm that compares various generated predictions with production safety benchmarks, calculates risk scores, and determines whether there are any significant current or future safety hazards.
[0104] When a high-risk situation arises, the server automatically generates a notification message containing an alert description and response suggestions. Using mainstream push services such as Firebase Cloud Messaging or WebSocket, the server can push the notification message in real-time to user devices (such as smartphones, tablets, and wearable devices) in multimodal formats including text, images, sound, and vibration. The terminal application (Android, iOS, or Web platform) will automatically display a pop-up window, play an alert sound effect, and simultaneously display targeted operation instructions.
[0105] In addition, the terminal integrates a user state recognition engine, which can capture user facial expressions through the camera or analyze speech using the microphone. Combined with emotion recognition software such as OpenVINO and Baidu EasyDL, it infers the user's current emotion (such as anxiety, tension, calmness, etc.) and sends the recognition results to the server. The server dynamically optimizes notification content based on the user's current state. For example, when the user is under pressure, it uses more soothing warning wording and prompts to help the user better understand and respond to risk warnings.
[0106] Users can quickly view risk information and recommended actions through the terminal interface, such as "pause the robot with one click," "switch to manual control," or "immediately evacuate the danger zone," and can confirm and provide feedback on received warnings. All user actions and feedback information are collected by the server for system iteration and optimization, AI model retraining, and long-term safety management improvements.
[0107] Specific example: For instance, if a robot suddenly approaches a densely populated area at an abnormal speed in a workshop, the server detects the abnormal movement through sensor data and, using a generative artificial intelligence model, infers that the robot is highly likely to enter a restricted area, with a risk score of 95. At this point, the server automatically generates a high-risk warning and pushes it to the smartphone of the person in charge on-site. Simultaneously, the terminal displays the alarm information to the user currently working through a red highlight, vibration, and voice announcement. If the user's facial expression is identified as "anxious," the server further reinforces the notification with a strong reminder, prompting the user to "leave the current area immediately; the robot has been suspended."
[0108] Generative artificial intelligence models can be trained and used for inference using the following type of prompts: "Based on the factory robot's motion and sensor data over the past 24 hours, predict whether it will exhibit any abnormal movements or safety risks in the next 5 minutes, and provide a description of the main risks and suggested response measures." "Please analyze the robot's historical sensor data and determine under what circumstances a safety alert needs to be sent immediately." "If the robot is about to enter a restricted area, please provide detailed warning information and emergency suggestions to ensure that users can respond in a timely and effective manner." "Based on the context, infer the user's possible emotional state when receiving the alert, and generate the most appropriate alert content for different emotions." This invention can be applied to scenarios such as manufacturing, automated warehousing, and smart logistics. It can be implemented using standard server hardware, mainstream artificial intelligence software platforms, and various terminal devices, which is beneficial for intelligent safety management and risk prevention in various complex industrial sites.
[0109] use Figure 12 The processing procedure is explained.
[0110] Step 1: The server collects raw data such as position, speed, status, and environment from multiple robots and external sensors in real time as input. The server periodically calls the data acquisition interface via industrial Ethernet, and the received data is saved to the database with timestamps and device numbers, outputting the raw dataset.
[0111] Step 2: The server preprocesses the original dataset as input. It uses tools such as TensorFlow for noise removal, missing value imputation, outlier correction, and normalization. It also synchronizes and corrects timestamps from different sensors, outputting a structured and time-series complete dataset.
[0112] Step 3: The server takes the preprocessed time-series data as input and feeds it into a generative artificial intelligence model (such as a time-series prediction model based on LSTM or Transformer). The server then initiates the model inference process, simulating and calculating the robot's behavior over the next 5 minutes, and outputting multiple predicted action paths and their probabilities.
[0113] Step 4: The server takes multiple predictions from the AI model as input and combines them with factory safety benchmarks and historical risk rules to perform a risk assessment. The server uses a built-in algorithm to calculate a risk score for each predicted scenario, enabling quantitative analysis of potential future hazards and outputting a prediction results table with risk levels.
[0114] Step 5: The server evaluates the risk assessment output, taking a prediction result table with risk levels as input. The server determines if any risk level exceeds a set threshold; if so, it automatically generates an alarm message (including risk description, specific data, and recommended measures). The output is a digital message containing all alarm elements.
[0115] Step 6: The server takes the alert information as input and pushes it to the target terminal using a network API (such as Firebase Cloud Messaging). Upon receiving the alert, the terminal automatically displays a pop-up window showing the alert content, accompanied by multimodal alerts such as sound and vibration. The output is a real-time alert notification on the terminal interface.
[0116] Step 7: The terminal takes the user's response to the alarm as input, such as clicking to pause the robot, switching to manual mode, or filling in feedback. The terminal's operation results are synchronized back to the server, which then adds the operation logs and feedback information to its historical database. The output consists of user response logs and system feedback data.
[0117] Step 8: The terminal uses a camera and microphone to collect user image and sound data as input, and a built-in user state recognition engine performs emotion analysis. Based on the emotion recognition results, the terminal adjusts its alert methods; for example, it increases the alert level when it detects that the user is emotionally stressed, outputting an optimized alarm message.
[0118] Step 9: The server takes user operation logs, feedback data, and terminal emotion recognition information as input, analyzes the collected user responses and behaviors, and continuously optimizes risk assessment strategies and generative artificial intelligence model parameters. The output is new system parameters and an optimized AI model.
[0119] Alternatively, an emotion engine for inferring user emotions can be combined. That is, the specific processing unit 290 can also use the emotion-specific model 59 to infer user emotions and perform specific processing using user emotions.
[0120] Example 2 The flow of a specific process in Example 2 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. The data processing device 12 will be referred to as the "server," and the smart device 14 as the "terminal."
[0121] In existing technologies, robot systems typically provide one-way alerts based solely on device data when predicting and issuing risks. They lack effective perception and feedback regarding the user's emotional state and cannot adjust warning methods and content in a timely manner according to the user's actual psychological reactions, resulting in insufficient user experience and operational safety. Furthermore, traditional alerts lack flexibility, easily causing user anxiety, misjudgment, or ignoring risk warnings, thus hindering the achievement of safe and efficient management under intelligent human-machine collaboration.
[0122] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 2 is achieved by the following means.
[0123] In this invention, the server includes means for acquiring information from a device, means for temporarily storing the acquired information in a storage medium, means for processing the information stored in the storage medium, means for inputting the processed information into a generative artificial intelligence model for predicting future behavior, means for assessing the risk based on the prediction results and benchmark values, means for generating warning information when the assessment results exceed a threshold, means for outputting the warning information to a display device, means for detecting the user's state operating the display device, and means for adjusting the content or display method of the warning information according to the user's state. This enables the robot system to intelligently predict future behavior and risks based on a generative artificial intelligence model, and flexibly adjust the warning content in conjunction with the user's real-time emotional state, effectively improving the safety and user experience of human-machine collaboration.
[0124] "Device" refers to an electronic hardware or software system used to achieve a specific function, including but not limited to robots, sensor modules, servers, terminal equipment, etc.
[0125] "Information" refers to the data content collected, generated, or processed by a device, including raw or processed digital data such as sensor data, location information, and status parameters.
[0126] "Storage medium" refers to a data carrier that can record and preserve information, including internal memory, external storage devices, or cloud storage platforms.
[0127] "Data processing" refers to the process of preprocessing, cleaning, format conversion, noise reduction, completion, and statistical analysis of collected information.
[0128] "Generative artificial intelligence models" refer to artificial intelligence systems that can automatically generate predictions, judgments, or content based on input information, such as models built based on neural networks or deep learning algorithms.
[0129] "Future behavior prediction" refers to using historical and real-time information to predict the actions or state changes that a device may perform in the future through generative artificial intelligence models.
[0130] "Risk assessment" refers to the process of comparing potential risks with predetermined safety benchmarks based on prediction results, quantifying and determining their risk level.
[0131] "Warning messages" refer to messages generated based on risk assessment results and used to prompt or remind users, including text, images, or voice messages.
[0132] "Display device" refers to an output device used to display warning messages or other content to users, including displays, lights, speakers, etc.
[0133] "User status" refers to the state information such as emotion, attention or psychological reaction obtained by the terminal after analyzing the user's behavioral characteristics such as voice, facial expression, and posture.
[0134] "Adjustment" refers to the process of optimizing, personalizing, or changing the display of warning messages based on user status.
[0135] The embodiments of the present invention are described below.
[0136] This invention relates to a system that combines generative artificial intelligence models and user emotional state recognition for device risk prediction and personalized display of warning information. Through collaboration between the server, terminal, and user, this system achieves high security and a positive user experience in an intelligent human-computer interaction environment.
[0137] The server employs general-purpose processing units (such as x86 architecture servers), network communication modules (such as Ethernet cards and Wi-Fi modules), and storage media (such as SSDs or cloud storage) as its hardware foundation. The server uses a data acquisition module to obtain real-time data from the device, which may include location, speed, acceleration, temperature, device status, etc., in JSON, CSV, or custom binary formats. After the data is transmitted to the server via the network, it is temporarily stored using databases such as MySQL and MongoDB.
[0138] When the server performs data processing, it uses software tools such as Python combined with NumPy and Pandas for data cleaning (including noise reduction, anomaly detection, and format conversion), and then feeds the processed time-series data as input to the generative artificial intelligence model. The server-side AI model can be developed based on open-source frameworks such as TensorFlow, PyTorch, and Keras, with specific models such as LSTM and Transformer structures, enabling high-precision predictions of device behavior within future time windows.
[0139] The server uses a dedicated algorithm module (implemented via Python scripts) to assess the risk level based on the prediction results. For example, it sets thresholds to determine high risk based on factors such as the distance between the device and obstacles, and speed limits. If the assessment score exceeds the safety settings, the server generates a corresponding warning message.
[0140] Based on the evaluation results, the server invokes text generation tools (such as integrating MiniGPT, BERT, or based on the enterprise's own generative large language model API) to automatically generate diverse warning content. Warning messages can be pushed to the terminal in real time via protocols such as WebSocket, HTTP, and MQTT.
[0141] The terminal is typically a smart device equipped with a camera, microphone, display, and processor, such as a mobile workstation, tablet, or dedicated control panel. The terminal captures the user's facial expressions through the camera and their voice through the microphone. It then uses a local or cloud-based facial expression analysis engine (such as OpenCV+Dlib or TensorFlow Lite) to process the expressions and voice, and employs sentiment analysis algorithms (such as deep learning-based sentiment classifiers or external AI services) to identify the user's current emotional state (e.g., anxiety, surprise, calmness) in real time.
[0142] The terminal reports the analyzed user status to the server. The server, based on this status, uses a generative artificial intelligence model to generate secondary prompts, adjusting the expression and display of the warning content (such as using milder wording, adding explanations, and adjusting visual colors), and finally sends the adaptive warning information to the terminal for display, ensuring that the user can understand and take appropriate action.
[0143] For example, in a factory environment, when the server predicts that a robot will approach a dangerous area in three seconds and the terminal detects that the user is in an anxious state, the server will automatically generate a warning such as "The robot is automatically detecting and ensuring safety. Please rest assured and there is no need to be nervous," and display it on the terminal in a clear and easy-to-understand manner.
[0144] Examples of prompts supported by this system include: "Based on the following user emotional state and future action predictions, please generate a heartwarming safety warning message: User is anxious, robot will approach danger zone in three seconds." "How can I create a notification that reassures users when the robot's data is abnormal?" "When a high-risk action occurs and the user shows fear, how can we adjust the warning to include both encouragement and reassurance?" Through the above system structure and process, the present invention can efficiently realize intelligent prediction of future risks of the device and generate personalized warning prompts based on the user's real-time emotional state, thereby improving the overall safety and comfort of human-machine collaboration.
[0145] use Figure 13 The processing procedure is explained.
[0146] Step 1: The server acquires sensor data in real time from devices (such as robots), including parameters such as position, speed, acceleration, and temperature. The input is raw data collected from the sensors (e.g., in JSON format). After receiving this raw data over the network, the server writes it to a database or stores it temporarily on a local disk. The output is the raw data record stored in the database. For example, the server automatically collects and stores the robot's current position and ambient temperature every second.
[0147] Step 2: The server preprocesses the stored raw data. The input is the raw data stored in the previous step. Using Python's pandas and NumPy data processing libraries, the server first performs noise reduction on the data, such as filtering out abnormal high-temperature values or abrupt acceleration changes. Then, it checks for and completes missing data, for example, by using the average of the preceding and following data. The output is the cleaned and completed structured data. A specific operation is: "After removing outliers, the missing acceleration data segment has been completed." Step 3: The server organizes the preprocessed data into time-series features and inputs them into a generative artificial intelligence model. The input is formatted time-series data (e.g., a data window of the past 30 seconds). The server calls an LSTM / Transformer model implemented in PyTorch or TensorFlow to predict the device's behavior over the next few seconds to tens of seconds, deriving the predicted behavioral trajectory and state sequence. The output is the predicted future motion path and state (e.g., "predicting approaching an obstacle in 5 seconds"). For example, the model predicts the robot will approach the workstation area within 10 seconds.
[0148] Step 4: The server reads the predicted data output by the generative artificial intelligence model and performs a risk assessment. The input is the model's predicted behavior. The server compares these results with safety rules (such as safe distance from obstacles, speed limits, etc.), calculates a risk score using a custom Python script, and determines whether an alarm condition has been met. The output is an assessment result with a risk score and related explanations. A specific action might be, "Robot path detected as being below the threshold distance to obstacles; risk score: 90." Step 5: The server generates warning messages based on the hazard assessment results. The input is the risk score and related description from the previous step. The server uses template messages or invokes a generative AI model to generate warning messages including alert text and suggested actions. The output is structured warning content (e.g., the text message "WARNING: Robot is approaching, please maintain a safe distance"). The server pushes the warning to the terminal device via a WebSocket interface.
[0149] Step 6: The terminal collects the user's facial expressions and voice data through a camera and microphone. The input is real-time video and sound. The terminal runs a local facial expression analysis and voice emotion recognition program to analyze the image and audio data and extract the user's emotional state information. The output is a structured user emotional state (such as "anxious", "calm", "surprised"). For example, if the terminal detects that the user frowns and raises their voice, it determines that they are "anxious".
[0150] Step 7: The server receives user emotional states uploaded from the terminal and, combined with previously generated warning information, intelligently adjusts the warning content or manner using generative AI models or specific templates. The input consists of the user's emotional state and initial warning content. Based on the user's current emotion, the server adjusts the tone, explanation, or visual presentation of the warning, such as adding reassuring expressions or using friendly colors. The output is the final personalized warning message. For example, "Based on the user's anxiety level, generate a comforting warning: 'The robot has initiated a self-check; please rest assured, there is no need to be anxious.'" Step 8: Users receive, read, or listen to personalized warning messages through the terminal. Input is the warning message displayed or read aloud by the terminal. Users can respond based on the warning content, such as clicking the "Understood" button or replying via voice. The terminal collects user feedback and uploads the relevant information to the server. Output is the new user feedback data. For example, if a user clicks the "Understood" button, the terminal records the click event and notifies the server.
[0151] Application Example 2 The process flow corresponding to the specific processing in Use Case 2 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. In addition, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".
[0152] Existing electromechanical equipment and user interaction systems have shortcomings in handling changes in user emotions and responding to real-time risks. In particular, when users are in a state of anxiety, confusion, or other negative emotions, the system struggles to adjust its interaction strategies in a timely and effective manner, leading to decreased user satisfaction and safety levels. Furthermore, how to combine multimodal data to reasonably predict the future actions of electromechanical equipment and provide personalized prompts based on dynamic context and user emotions remains a significant technical challenge.
[0153] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 2 is achieved by the following means.
[0154] In this invention, the server includes a device for acquiring electromechanical equipment status information and environmental information; a device for preprocessing the acquired dynamic information; a device for predicting the future actions of the electromechanical equipment based on the preprocessed data using a machine learning model; a device for performing safety analysis and risk assessment on the predicted actions; a device for automatically generating and sending prompt information; a device for estimating the emotion of user facial expressions or voice; a device for dynamically adjusting the content and method of notifications according to the emotional state; and a device for automatically generating prompt statements based on a generative artificial intelligence model. This enables the system to flexibly predict and control the behavior of electromechanical equipment based on the real-time environment and the user's current emotional state, generating interactive prompts that conform to the scenario and the user's psychology, thereby improving user experience and overall system security.
[0155] "Information processing device" refers to electronic equipment capable of collecting, processing, storing and transmitting data, including but not limited to servers, computers and gateway devices.
[0156] "Mechanical and electrical equipment" refers to mechanical, electrical, or electronic devices with sensing, actuation, and automatic control functions, such as robots and automated guided vehicles.
[0157] "Dynamic information" refers to data that reflects the changes in electromechanical equipment and its operating environment at different times, including sensor data, status information, environmental parameters, etc.
[0158] Signal processing refers to the technical process of filtering, denoising, and extracting features from raw acquired data to improve data quality and application efficiency.
[0159] "Data completion" refers to the process of repairing and improving missing, abnormal, or erroneous parts of collected information to make the data structure complete and have analytical value.
[0160] "Machine learning model" refers to a mathematical model and algorithm system that is trained on a large number of data samples to acquire reasoning ability and make predictions or classification judgments on input data.
[0161] "Risk assessment" refers to the technical means of analyzing, quantifying, and classifying the unsafe factors that may arise from electromechanical equipment in order to determine the potential hazards.
[0162] "Prompt messages" refer to notifications, voice messages, or images that provide users with information about device status, operational suggestions, or risk alerts.
[0163] "User identification device" refers to sensing hardware that can acquire user identity, behavior, facial expressions, or voice characteristics, such as cameras, microphones, and wearable devices.
[0164] "Emotion estimation processing" refers to the process of judging a user's emotional state, such as calm, doubt, or anxiety, by analyzing multimodal information such as facial expressions and voice.
[0165] "Dynamic adjustment of notification content and notification method" refers to the system's ability to change the display method, tone, or notification frequency of prompts based on the user's current emotional state and scenario needs.
[0166] "Generative AI models" refer to AI algorithm models that can automatically generate natural language text, speech, or image content based on the input context to enhance the interactive experience.
[0167] This invention relates to an interactive system based on an information processing device, a user terminal, and a user. This system combines real-time data processing, machine learning, and emotion recognition to provide intelligent services to users. The system primarily utilizes a server (information processing device), terminal devices (such as smart glasses and wearable devices), and electromechanical equipment with data acquisition and interaction functions (such as robots) to collaboratively provide context-adaptive and user-friendly risk warnings and operational guidance.
[0168] The server is equipped with a high-performance computer running Linux or Windows operating systems, and features machine learning frameworks such as TensorFlow and PyTorch, as well as data and image processing libraries such as OpenCV and scikit-learn. It can also interact with relational databases (such as MySQL) or NoSQL databases. The terminal includes a smart device with a camera and microphone, facilitating real-time acquisition of users' facial expressions, voice, and other multimodal emotional features, and transmitting the information to the server via a wireless communication module (such as Wi-Fi or Bluetooth).
[0169] During operation, the server acquires status and environmental dynamic information from sensors, positioning modules, and cameras of the electromechanical equipment, and performs preprocessing techniques such as filtering, noise reduction, anomaly correction, and structuring. After feature extraction and time-series analysis, the data is input into a machine learning-based model (such as a generative artificial intelligence model implemented using TensorFlow) to predict the future actions of the electromechanical equipment. The server further incorporates custom safety standards to conduct risk assessments and quantitative analysis of these future actions.
[0170] The server also utilizes real-time user facial expressions and voice data collected from terminal devices to determine emotional states using methods such as the Microsoft Azure Emotion API or a self-developed sentiment analysis model. Based on the determination results, the server can automatically adjust the expression, display frequency, and tone of notification content. The system employs a generative artificial intelligence model to automatically generate prompts, making the prompts more relevant to the user's current psychological state and actual needs. The system is highly flexible, supporting multi-channel push notifications of various information such as text and voice, enhancing the overall interactive experience and service security.
[0171] For example, in a real-world application scenario, suppose a user appears confused upon seeing a self-service ordering robot in a coffee shop. The terminal device detects that the user's brow is furrowed and their speech rate has slowed down, and uploads this emotional data to the server. After analysis, the server determines that the user is confused and generates a prompt statement using a generative artificial intelligence model, such as: "Hello, do you need help? Please don't worry, I will guide you step by step to complete the operation." This prompt can be displayed as text on smart glasses or read aloud by the robot in a gentle tone.
[0172] Example of a prompt statement for generating an AI model: "If the customer looks confused, please explain the ordering process to them in a respectful and gentle tone." "When a customer appears anxious, offer words of encouragement and briefly inquire about their needs." "If a user remains silent for more than 30 seconds, please generate a message to initiate interaction." Through the above technical solutions, the system can realize real-time perception of the environment and user status, dynamically predict risks and generate personalized prompts, greatly improving the service quality and security capabilities for users.
[0173] use Figure 14 The processing procedure is explained.
[0174] Step 1: The server acquires real-time dynamic information from electromechanical equipment. Inputs include equipment status data, environmental data, camera images, and audio signals. The server performs initial data collection from sensors and cameras, storing it in a database. Output is a raw dynamic dataset containing all raw sensor information.
[0175] Step 2: The server preprocesses the raw dynamic data. The input is the original dataset. The server uses OpenCV for image denoising and normalization, and uses NumPy or Pandas for outlier cleaning and missing data completion. The output is high-quality, standardized and corrected data.
[0176] Step 3: The server performs feature extraction and format conversion on the preprocessed data. The input is high-quality dynamic information. The server uses OpenCV for face detection and expression recognition, and scikit-learn for spectral analysis of the audio data, while encoding all features into a unified input vector structure. The output is a set of data vectors that can be used as input to machine learning models.
[0177] Step 4: The server uses a machine learning model to predict future actions of electromechanical equipment. The input is a formatted feature vector. The server imports the input vector into a generative AI model trained in TensorFlow or PyTorch, which then predicts possible subsequent actions. The output is the predicted future action, such as the behavior category or the expected trigger time.
[0178] Step 5: The server performs a risk assessment on the prediction results. The input is the predicted future action output by the machine learning model. The server analyzes the action risk based on built-in safety standards, quantifies each risk, and derives a risk score. The output is a risk assessment report, including the risk level and recommended measures.
[0179] Step 6: The server automatically generates prompts based on the risk assessment results. The input is the risk assessment report. The server extracts key information through rule-based algorithms or directly calls a generative artificial intelligence model to generate prompts suitable for the current context. The output is structured prompt text or audio.
[0180] Step 7: The terminal collects and uploads user emotional information in real time. Input consists of the user's facial expressions and voice signals. The terminal uses a camera to capture images and a microphone to collect sound, then transmits the data to the server in real time via a wireless network. Output is raw emotion-related data.
[0181] Step 8: The server performs sentiment estimation on the received user sentiment information. The input consists of raw data from the user's facial expressions and speech. The server calls a sentiment analysis API or a self-developed sentiment analysis model to identify the user's current emotion (such as confusion, anxiety, calmness, etc.) and generates a sentiment state label. The output is the determination of the user's current sentiment state.
[0182] Step 9: The server dynamically adjusts the content and manner of notification messages based on risk assessment results and the user's emotional state. Inputs include the notification text and the user's emotional state. The server uses a generative AI model to optimize the notification statements, making the tone more gentle or detailed, and adjusting the push frequency and delivery channels. The output is the final customized notification message.
[0183] Step 10: Users receive optimized prompts via terminals or electromechanical devices. The input is the adjusted notification content. Users see on their terminal screen or hear via voice friendly prompts generated based on their emotions and context, facilitating quick understanding and response. The output is the actual guidance and service experience received by the user.
[0184] The specific processing unit 290 sends the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires sound representing user input regarding the result of the specific processing. The control unit 46A sends the sound data representing user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0185] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0186] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart device 14 or external devices, and the smart device 14 acquires or collects information required for processing from the data processing device 12 or external devices.
[0187] For example, the collection unit is implemented by the control unit 46A of the smart device 14 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart device 14 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the output device 40 of the smart device 14 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0188] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart device 14.
[0189] Second Implementation Method Figure 3 An example of the configuration of the data processing system 210 according to the second embodiment is shown.
[0190] like Figure 3 As shown, the data processing system 210 includes a data processing device 12 and smart glasses 214. A server can be cited as an example of the data processing device 12.
[0191] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0192] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, and communication I / F 44 are also connected to the bus 52.
[0193] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0194] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).
[0195] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0196] Figure 4 This illustrates an example of the main functions of the data processing device 12 and the smart glasses 214. For example... Figure 4 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0197] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0198] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).
[0199] In the smart glasses 214, the processor 46 performs reception and output processing. The memory 50 stores the reception and output program 60. The processor 46 reads the reception and output program 60 from the memory 50 and executes the read reception and output program 60 on the RAM 48. The reception and output processing is implemented by the processor 46 operating as a control unit 46A according to the reception and output program 60 executed on the RAM 48. Furthermore, the smart glasses 214 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290.
[0200] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart glasses 214. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0201] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0202] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0203] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0204] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0205] The specific processing unit 290 sends the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A outputs the result of the specific processing to the speaker 240. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0206] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0207] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or external devices, and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or external devices.
[0208] For example, the collection unit is implemented by the control unit 46A of the smart glasses 214 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart glasses 214 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the smart glasses 214 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0209] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart glasses 214.
[0210] Third Implementation Method Figure 5 An example of the configuration of the data processing system 310 according to the third embodiment is shown.
[0211] like Figure 5 As shown, the data processing system 310 includes a data processing device 12 and a head-mounted terminal 314. A server can be cited as an example of the data processing device 12.
[0212] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0213] The head-mounted terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, display 343, and communication I / F 44 are also connected to the bus 52.
[0214] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0215] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).
[0216] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0217] Figure 6 This illustrates an example of the main functions of the data processing device 12 and the head-mounted terminal 314. For example... Figure 6 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0218] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0219] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.
[0220] In the head-mounted terminal 314, the processor 46 performs the acceptance / output processing. The memory 50 stores the acceptance / output program 60. The processor 46 reads the acceptance / output program 60 from the memory 50 and executes the read acceptance / output program 60 on the RAM 48. The acceptance / output processing is implemented by the processor 46 operating as a control unit 46A according to the acceptance / output program 60 executed on the RAM 48.
[0221] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the head-mounted terminal 314. In the following description, the data processing device 12 will be referred to as the "server" and the head-mounted terminal 314 will be referred to as the "terminal".
[0222] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0223] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0224] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0225] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0226] The specific processing unit 290 sends the result of the specific processing to the head-mounted terminal 314. In the head-mounted terminal 314, the control unit 46A outputs the result of the specific processing to the speaker 240 and the display 343. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0227] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 includes prompt words containing instructions, as well as inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0228] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the head-mounted terminal 314, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the head-mounted terminal 314. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the head-mounted terminal 314 or external devices, and the head-mounted terminal 314 acquires or collects information required for processing from the data processing device 12 or external devices.
[0229] For example, the collection unit is implemented by the control unit 46A of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the head-mounted terminal 314 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12 to analyze the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12 to generate a menu using a generation AI. For example, the serving unit is implemented by the speaker 240 and display 343 of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12 to provide the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0230] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the head-mounted terminal 314.
[0231] Fourth Implementation Method Figure 7 An example of the configuration of the data processing system 410 according to the fourth embodiment is shown.
[0232] like Figure 7 As shown, the data processing system 410 includes a data processing device 12 and a robot 414. A server can be cited as an example of the data processing device 12.
[0233] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0234] Robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, controlled object 443, and communication I / F 44 are also connected to the bus 52.
[0235] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.
[0236] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, to photograph the area around robot 414 (e.g., the field of view defined by a perspective equivalent to the field of vision of an average healthy person).
[0237] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.
[0238] The controlled object 443 includes a display device, LEDs (light-emitting diodes) for the eyes, and motors for driving the arms, hands, and feet. The posture or movement of the robot 414 is controlled by controlling the motors in the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. In addition, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0239] Figure 8 This illustrates an example of the main functions of the data processing device 12 and the robot 414. For example... Figure 8 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.
[0240] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0241] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.
[0242] In robot 414, the processor 46 performs the acceptance and output processing. The memory 50 stores the acceptance and output program 60. The processor 46 reads the acceptance and output program 60 from the memory 50 and executes the read acceptance and output program 60 on RAM 48. The acceptance and output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance and output program 60 executed on RAM 48.
[0243] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the robot 414. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 will be referred to as the "terminal".
[0244] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.
[0245] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.
[0246] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.
[0247] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.
[0248] The specific processing unit 290 sends the result of the specific processing to the robot 414. In the robot 414, the control unit 46A outputs the result of the specific processing to the speaker 240 and the controlled object 443. The microphone 238 acquires sound input representing the result of the specific processing. The control unit 46A sends the sound data representing the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.
[0249] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI including the generation AI.
[0250] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or external devices, and the robot 414 acquires or collects information required for processing from the data processing device 12 or external devices.
[0251] For example, the collection unit is implemented by the control unit 46A of the robot 414 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the robot 414 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the robot 414 and the control object 443 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.
[0252] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the robot 414.
[0253] Furthermore, the emotion-specific model 59, acting as an emotion engine, can determine a user's emotion based on a specific mapping. Specifically, the emotion-specific model 59 can determine a user's emotion based on an emotion graph that serves as a specific mapping (see [reference]). Figure 9 The emotion-specific model 59 can also determine the robot's emotion, and the specific processing unit 290 performs specific processing based on the robot's emotions.
[0254] Figure 9 This is a diagram representing an emotion map 400 that maps multiple emotions. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotion is. On the outer side of the concentric circles, emotions representing states or behaviors arising from mood are arranged. Emotions are concepts that include feelings and mental states. Emotions generated by reactions occurring in the brain are arranged roughly to the left of the concentric circles. Emotions derived from situational judgments are arranged roughly to the right of the concentric circles. Emotions generated by reactions occurring in the brain and derived from situational judgments are arranged roughly above and below the concentric circles. Furthermore, "pleasant" emotions are arranged above the concentric circles, and "unpleasant" emotions are arranged below them. Thus, in the emotion map 400, multiple emotions are mapped based on the structure that generates emotions, and emotions that are likely to occur simultaneously are mapped close to each other.
[0255] These emotions are distributed at the three o'clock position of the emotion map 400, typically fluctuating between peace and anxiety. In the right half of the emotion map 400, situational awareness dominates over internal sensation, thus resulting in an impression of calm.
[0256] The inner side of the emotion map 400 represents the inner state, while the outer side represents behavior. Therefore, the further outward you are from the emotion map 400, the more visible the emotion becomes (manifested in behavior).
[0257] Here, human emotions are based on various balances such as posture and blood sugar levels. When these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotions in robots, cars, motorcycles, etc., can also be created in the following way: based on various balances such as posture and remaining battery power, when these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotion maps can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a Brain Physiological Signal Analysis System for Voice Emotion Recognition and Emotion, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). In the left half of the emotion map, emotions belonging to the sensory-dominated region, called "response," are arranged. Furthermore, in the right half of the emotion map, emotions belonging to the situational cognition-dominated region, called "situation," are arranged.
[0258] In the emotion map, two types of emotions that promote learning are defined. One is a negative emotion on the situational side, in the middle or peripheral region of "repentance" or "reflection." This occurs when the robot experiences negative emotions such as "I don't want to experience this feeling again" or "I don't want to be blamed again." The other is a positive emotion on the response side, near the "desire" region. This occurs when there are positive feelings such as "wanting more" or "wanting to know more."
[0259] The emotion-specific model 59 inputs user input into a pre-trained neural network to obtain emotion values representing each emotion shown in the emotion map 400, thereby determining the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network... Figure 10 As shown in the sentiment graph 900, it was trained in a way that sentiments that are configured close to each other have similar values. Figure 10 The text shows examples of emotions such as "peace of mind", "stability", and "reassurance" that have similar emotion values.
[0260] The above description focuses on the functions of the data processing device 12, but the system of this disclosure is not necessarily installed on a server. The system of this disclosure can also be installed as a general information processing system. This disclosure can also be installed, for example, as a software program running on a personal computer, an application running on a smartphone, etc. The method of this disclosure can also be provided to users in the form of SaaS (Software as a Service).
[0261] In the above embodiments, an example of a specific process being performed by a single computer 22 is given. However, the technology disclosed herein is not limited to this, and the specific process can also be distributed among multiple computers, including computer 22. For example, the data generation model 58 can be located on an external device of the data processing apparatus 12, where data is generated based on the input data.
[0262] In the above embodiments, examples of storing a specific processing program 56 in the memory 32 have been described, but the technology disclosed herein is not limited thereto. For example, the specific processing program 56 may also be stored in a portable computer-readable non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed into the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0263] Alternatively, a specific processing program 56 may be pre-stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 according to the requirements of the data processing device 12.
[0264] In addition, it is not necessary to store all the specific processing program 56 in the storage device such as the server connected to the data processing device 12 via the network 54 or in the memory 32; a portion of the specific processing program 56 may be stored in advance.
[0265] As hardware resources for performing specific processes, various processors, as shown below, can be used. For example, a CPU can be listed as a processor, which functions as a general-purpose processor that performs specific processes by executing software, i.e., a program. Furthermore, processors can be listed as special-purpose circuits such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application-Specific Integrated Circuits), which are processors with circuitry specifically designed to perform specific processes. Each processor has built-in or connected memory, and each processor executes specific processes using that memory.
[0266] The hardware resources for performing a specific process can consist of one of these various processors, or a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resources for performing a specific process can be a single processor.
[0267] As an example of a single processor, there are two approaches: First, a processor is composed of a combination of one or more CPUs and software, which functions as a hardware resource to perform a specific process; second, as represented by a SoC (System-on-a-chip), a processor is used to implement the functionality of the entire system, which includes multiple hardware resources for performing a specific process, using a single IC (Integrated Circuit) chip. In this way, the specific process is implemented by using one or more of the aforementioned processors as hardware resources.
[0268] Furthermore, the hardware architecture of these various processors, more specifically, can utilize circuits that combine semiconductor elements and other circuit components. Moreover, the specific process described above is just one example. Therefore, without departing from the main point, unnecessary steps can certainly be deleted, new steps added, or the processing order changed.
[0269] The descriptions and illustrations above are detailed explanations of a portion of the technology disclosed herein, and are merely one example of the technology disclosed herein. For example, the above descriptions of the structure, function, effect, and results are just one example of the structure, function, effect, and results of a portion of the technology disclosed herein. Therefore, without departing from the spirit of the technology disclosed herein, unnecessary parts may be deleted, new elements added, or replacements may be made to the descriptions and illustrations above. Furthermore, to avoid confusion and facilitate understanding of a portion of the technology disclosed herein, explanations of common technical knowledge that do not require special explanation under the premise of being able to implement the technology disclosed herein have been omitted from the descriptions and illustrations above.
[0270] All documents, patent applications and technical specifications set forth in this specification are incorporated herein by reference to the same extent that each document, patent application and technical specification is specifically and individually described therein and referenced by reference.
[0271] In addition, the following notes are provided in response to the above explanation.
[0272] Example 1 (Note 1) An information processing system includes: a device for acquiring time-series information in real time from a physical information generation device; a device for preprocessing the acquired time-series information, such as filling in missing values, removing outliers, and normalizing; a device for inputting the preprocessed time-series information as input to a generative artificial intelligence model with predetermined prompt statements to predict future behavior; a device for comparing the predicted future behavior with a standard benchmark and performing a risk assessment; a device for generating customized warning messages based on the risk assessment results; and a device for notifying the generated warning messages to a notification information prompting device and presenting them to a receiver via visual or auditory means.
[0273] (Note 2) The information processing system according to Appendix 1 further includes: a device for using time-series information as input data for a generative artificial intelligence model and predicting future behavior in a multi-step time-series manner.
[0274] (Note 3) The information processing system according to Appendix 1 further includes: means for quantitatively scoring the predicted future behavior based on multiple factors, and generating a warning message when the risk score exceeds a predetermined threshold.
[0275] Application Example 1 (Note 1) An information processing system includes: a device for acquiring real-time information from a data acquisition device; a device for removing noise and correcting data from the acquired time-series information using a data preprocessing algorithm; a device for inputting the preprocessed time-series information into a generative artificial intelligence model to probabilistically predict future actions; a device for comparing the predicted multiple action patterns with a safety benchmark and performing risk assessment to obtain a risk score; a device for generating a notification message containing warning content and recommended responses when the risk score exceeds a threshold; a device for sending the generated notification message to a user terminal via a network communication path and notifying the user visually, audibly, or tactilely; a device for assisting the user in adjusting the actions of a working device or responding on-site based on the notification message through a user interface; a device for inferring the user's emotional state based on input data using a user state recognition engine and optimizing the notification content and method based on the inference results; and a device for collecting the user's response history and feedback information for system performance optimization and learning processing.
[0276] (Note 2) According to the information processing system described in Note 1, the preprocessed time series information is input into the generative artificial intelligence model, which outputs the future predicted action pattern and generates the corresponding risk scenario accordingly.
[0277] (Note 3) According to the information processing system described in Note 1, during the risk assessment process, the action prediction results of the generative artificial intelligence model are compared with the safety benchmark value of the working environment to calculate the risk score. When the risk score exceeds the threshold, warnings and user assistance information are automatically generated.
[0278] Example 2 (Note 1) An information processing system includes: means for acquiring information from a device; means for temporarily storing the acquired information in a storage medium; means for processing the information stored in the storage medium; means for inputting the processed information into a generative artificial intelligence model for predicting future behavior; means for assessing the risk based on the prediction result and a benchmark value; means for generating a warning message when the assessment result exceeds a threshold; means for outputting the warning message to a display device; means for detecting the user's state while operating the display device; and means for adjusting the content or display mode of the warning message according to the user's state.
[0279] (Note 2) According to the information processing system described in Note 1, the generative artificial intelligence model makes predictions based on time-series information.
[0280] (Note 3) According to the information processing system described in Appendix 1, the system identifies the user's emotional state by parsing the user's voice or image information, and makes corresponding adjustments based on the identified state.
[0281] Application Example 2 (Note 1) An information processing system includes: a device for acquiring status information of electromechanical equipment and environmental information from an information processing unit; a preprocessing device for performing signal processing and data completion on the acquired dynamic information; a device for using the preprocessed dynamic information as input and predicting future actions of the electromechanical equipment using a machine learning model; a device for performing safety analysis and quantitatively assessing potential risks based on the predicted actions according to predetermined standards; a device for automatically generating prompt information based on the assessment results; a device for performing emotion estimation processing based on facial expression information or voice information acquired from a user recognition device; a device for dynamically adjusting the notification content and notification method based on the estimated emotional state and the generated prompt information and sending the notification information to a user terminal; and a device for automatically generating prompt information based on scene and human emotion information using a generative artificial intelligence model.
[0282] (Note 2) According to the information processing system described in Appendix 1, the machine learning model predicts the future actions of electromechanical equipment based on dynamic information time-series data at multiple time points.
[0283] (Note 3) According to the information processing system described in Appendix 1, when the predicted action exceeds a preset standard value, the security analysis device calculates a risk score and reflects the score in the notification information.
Claims
1. An information processing system, characterized in that, include: A device for collecting real-time data from a robot; A device for preprocessing collected data; A device for predicting the future behavior of a robot using preprocessed data; Device for conducting risk assessment based on predicted behavior; Device used to generate alarms based on risk assessment results; and A device used to notify users of generated alarms.
2. The information processing system according to claim 1, characterized in that, The device for predicting the robot's future behavior uses time-series data for prediction.
3. The information processing system according to claim 1, characterized in that, The risk assessment device calculates a risk score when the predicted behavior exceeds the safety standard.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A