Process monitoring system and process monitoring method

JP2026144971APending Publication Date: 2026-09-09HITACHI LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025246358
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-28
Filing Date
2025-12-12
Publication Date
2026-09-09

Smart Images

  • Figure 2026144971000001_ABST
    Figure 2026144971000001_ABST
Patent Text Reader

Abstract

It provides an innovative method for implementing process monitoring. [Solution] The process monitoring system includes one or more sensors and a processor that communicates with one or more sensors, wherein the processor is configured to perform anomaly detection by selecting a first system checkpoint for performing process monitoring on a robot system, measuring system variables using one or more sensors at the first system checkpoint, capturing one or more images of the robot system using a vision sensor at the first system checkpoint, generating at least one predicted image of the robot system using a predictive model that uses the system variables as input to the predictive model, and comparing one or more images of the robot system with at least one predicted image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to methods and systems for performing process monitoring and anomaly detection. Background Art

[0002] In automation, it is important to monitor the state of a process and detect any anomalies, which is important for purposes such as evaluating system performance (e.g., throughput, success rate, etc.), improving system design based on errors, and preventing prolonged downtime caused by process errors.

[0003] For example, in robotic pick-and-place applications, one may want to know whether objects are being safely handled and transported. When manipulating packages during a process, unexpected events such as package dropping or package damage may occur. These errors may further affect downstream operations. Therefore, to detect any anomalies, it is necessary to monitor operations including not only the state (e.g., functionality, appearance, etc.) of system components but also their motions.

[0004] Some process states can be directly measured and checked using sensors (e.g., weight, temperature, etc.), while others cannot be directly measured, or may be difficult / costly to measure even with special sensors, such as the state of a package or the grasping stability when a robot handles materials. Leaving some states of a process / system unmonitored because they cannot be directly measured may leave process anomalies undetected, which can cause significant downtime or damage to the system. Summary of the Invention Problem to be Solved by the Invention

[0005] Conventional technologies disclose methods for detecting robot anomalies by generating and comparing simulation images. Generating simulation images requires a complete operation / motion plan, which can be difficult and burdensome to provide in practice. Furthermore, anomaly detection is limited to the robot itself and does not consider the environment in which the robot operates.

[0006] Prior art has disclosed methods for monitoring and detecting system / robot anomalies. Images of the robot are captured while it is operating normally and compared to previously captured images of the robot. This method requires capturing a reference image for comparison, lacks flexibility, and has a limited detection range. [Means for solving the problem]

[0007] Aspects of the present disclosure include an innovative method for performing process monitoring. The method includes the steps of: a processor selecting a first system checkpoint for performing process monitoring on a system; the processor measuring system variables using one or more sensors at the first system checkpoint; the processor capturing one or more images of the system using vision sensors at the first system checkpoint; the processor generating at least one predicted image of the system using a predictive model that takes the system variables as input to the predictive model; and the processor performing anomaly detection by comparing one or more images of the system with at least one predicted image.

[0008] In some exemplary implementations, the system may include a robotic arm configured to grasp and move objects.

[0009] In some exemplary implementations, system variables may include one or more of the following: the joint angles of the robot arm, the weight of the object, the dimensions of the object, the stock-keeping unit (SKU) of the object, or the barcode of the object.

[0010] In some exemplary implementations, the method may further include a step in which the processor automatically issues instructions to the robot arm to perform a recovery process in response to a detected anomaly.

[0011] In some exemplary implementations, the processor may be configured to select a first system checkpoint based on the posture of the robot arm or the position of an object.

[0012] In some exemplary implementations, one or more images captured using a vision sensor may include one or more images of a robotic arm and / or an object at a first system checkpoint.

[0013] In some exemplary implementations, the method further includes the steps of a processor deriving a relative position between a robot arm and a vision sensor using one or more images of the system, and a processor deriving joint angles of the robot arm using one or more images of the system, and generating at least one predictive image, wherein the relative position between the robot arm and the vision sensor and the joint angles of the robot arm are used as inputs to a predictive model in order to generate at least one predictive image.

[0014] In some exemplary implementations, the predictive model is an artificial intelligence (AI) model trained using historical system data, including (i) past joint angles of the robot arm and (ii) past motion images of the robot arm.

[0015] In some exemplary implementations, past motion images of the robot arm may include two-dimensional motion images of the robot arm.

[0016] In some exemplary implementations, one or more images of the robot system may include images of the robot arm and images of the object.

[0017] In some exemplary implementations, at least one predictive image of the robot system may include a robot arm predictive image and an object predictive image.

[0018] In some exemplary implementations, the processor may be configured to compare one or more images of a robotic system with at least one predicted image by comparing an image of the robotic arm with a predicted image of the robotic arm and an image of an object with a predicted image of the object.

[0019] In some exemplary implementations, the processor may be configured to select a first system checkpoint based on the reception of sensor signals indicating process transitions.

[0020] Aspects of the present disclosure include an innovative system for performing process monitoring. The system includes one or more sensors and a processor that communicates with one or more sensors, the processor being configured to perform anomaly detection by selecting a first system checkpoint for performing process monitoring on a robot system, measuring system variables using one or more sensors at the first system checkpoint, capturing one or more images of the robot system using vision sensors at the first system checkpoint, generating at least one predicted image of the robot system using a predictive model that takes the system variables as input to the predictive model, and comparing one or more images of the robot system with at least one predicted image.

[0021] In some exemplary implementations, a robotic system may include a robotic arm configured to grasp and move objects.

[0022] In some exemplary implementations, the system variables may include one or more of a joint angle of the robot arm, a weight of an object, a dimension of an object, a stock keeping unit (SKU) of an object, or a barcode of an object.

[0023] In some exemplary implementations, the processor may be further configured to automatically issue an instruction for executing a recovery process on the robot arm in response to a detected abnormality.

[0024] In some exemplary implementations, the processor may be configured to select a first system checkpoint based on a posture of the robot arm or a position of an object.

[0025] In some exemplary implementations, the one or more images captured using a vision sensor may include one or more images of the robot arm and / or the object at the first system checkpoint.

[0026] In some exemplary implementations, the processor may be further configured to derive a relative position between the robot arm and the vision sensor using one or more images of the robot system, and derive the joint angle of the robot arm using the one or more images of the robot system, wherein generating at least one predicted image may include using the relative position between the robot arm and the vision sensor and the joint angle of the robot arm as inputs to a prediction model to generate the at least one predicted image.

[0027] In some exemplary implementations, the prediction model may be an artificial intelligence (AI) model trained using historical system data including (i) historical joint angles of the robot arm and (ii) historical operation images of the robot arm.

[0028] In some exemplary implementations, the historical operation images of the robot arm may include two-dimensional operation images of the robot arm.

[0029] In some exemplary implementations, the one or more images of the robot system may include an image of a robot arm and an image of an object.

[0030] In some exemplary implementations, the at least one predicted image of the robot system may include a robot arm predicted image and an object predicted image.

[0031] In some exemplary implementations, the processor may be configured to compare the one or more images of the robot system with the at least one predicted image by comparing the image of the robot arm with the robot arm predicted image and comparing the image of the object with the object predicted image.

[0032] In some exemplary implementations, the processor may be configured to select the first system checkpoint based on receiving a sensor signal indicating a process transition.

[0033] Aspects of the present disclosure include an innovative non-transitory computer-readable medium storing instructions for performing process monitoring. The instructions may include: selecting a first system checkpoint for performing process monitoring on a system; measuring system variables using one or more sensors at the first system checkpoint; capturing one or more images of the system using a vision sensor at the first system checkpoint; generating at least one predicted image of the system using a prediction model that uses the system variables as an input to the prediction model; and performing anomaly detection by comparing the one or more images of the system with the at least one predicted image.

[0034] In some exemplary implementations, the system may include a robot arm configured to grip and move an object.

[0035] In some exemplary implementations, system variables may include one or more of the following: the joint angles of the robot arm, the weight of the object, the dimensions of the object, the stock-keeping unit (SKU) of the object, or the barcode of the object.

[0036] In some exemplary implementations, the instructions may further include the processor automatically issuing instructions to the robot arm to perform a recovery process in response to detected anomalies.

[0037] In some exemplary implementations, the instruction may further include selecting a first system checkpoint based on the posture of the robot arm or the position of an object.

[0038] In some exemplary implementations, one or more images captured using a vision sensor may include one or more images of a robotic arm and / or an object at a first system checkpoint.

[0039] In some exemplary implementations, the instruction further includes the processor deriving the relative position between the robot arm and the vision sensor using one or more images of the system, and the processor deriving the joint angles of the robot arm using one or more images of the system, and generating at least one predictive image, using the relative position between the robot arm and the vision sensor and the joint angles of the robot arm as inputs to a predictive model to generate at least one predictive image.

[0040] In some exemplary implementations, the predictive model is an artificial intelligence (AI) model trained using historical system data, including (i) past joint angles of the robot arm and (ii) past motion images of the robot arm.

[0041] In some exemplary implementations, past motion images of the robot arm may include two-dimensional motion images of the robot arm.

[0042] In some exemplary implementations, one or more images of the robot system may include images of the robot arm and images of the object.

[0043] In some exemplary implementations, at least one predictive image of the robot system may include a robot arm predictive image and an object predictive image.

[0044] In some exemplary implementations, the instruction may further include comparing one or more images of the robotic system to at least one predicted image by comparing an image of the robotic arm to a predicted image of the robotic arm and an image of the object to a predicted image of the object.

[0045] In some exemplary implementations, the instruction may further include selecting a first system checkpoint based on the reception of a sensor signal indicating a process transition. [Brief explanation of the drawing]

[0046] [Figure 1] An exemplary system configuration diagram 100 for performing process evaluation and monitoring is shown, based on an exemplary implementation. [Figure 2] This section shows exemplary checkpoints selected based on the robot's posture, based on an exemplary implementation. [Figure 3] This shows exemplary checkpoints selected based on sensor signals. [Figure 4] This shows illustrative images captured at various checkpoints. [Figure 5] This document presents an exemplary predictive image generation process 500 that utilizes a mathematical model, using an exemplary implementation. [Figure 6] This shows an exemplary 3D robot motion visualization 600 with an exemplary implementation. [Figure 7] This document presents an exemplary process 700 for training a machine learning model that generates predictive images, using an exemplary implementation. [Figure 8] This document presents an exemplary process flow 800 for training a machine learning model that generates predictive images / 2D motion images, using an exemplary implementation. [Figure 9] An exemplary process flow 900 for generating a 2D motion image using the trained machine learning model 860 shown in Figure 8 is presented, based on an exemplary implementation. [Figure 10] Figure 1000 shows an exemplary implementation that illustrates a method for comparing predicted images and captured images. [Figure 11] Figure 1100 shows an example of an implementation configuration illustrating how to perform image comparison. [Figure 12] Diagram 1200 shows an alternative exemplary system configuration for performing process evaluation and monitoring using an exemplary implementation. [Figure 13] This illustrates an exemplary computing environment with exemplary computer devices suitable for use in several exemplary implementations. [Figure 14] An alternative exemplary process 1400 is shown for training a machine learning model to generate predictive images, using an exemplary implementation. [Figure 15] This document presents an alternative process flow 1500 for training a machine learning model that generates predictive images / 2D motion images, using an exemplary implementation. [Modes for carrying out the invention]

[0047] Next, a general architecture for implementing various features of this disclosure will be described with reference to the drawings. The drawings and related descriptions are provided to illustrate exemplary implementations of this disclosure and are not intended to limit the scope of this disclosure. Throughout the drawings, reference numbers are reused to indicate correspondences between the referenced elements.

[0048] The following detailed description provides details of the drawings and exemplary implementations of this application. Reference numbers and descriptions of elements that overlap between drawings are omitted for clarity. Terms used throughout this description are provided as examples and are not intended to be limiting. For example, the term “automatic” may include fully automatic or semi-automatic, depending on the implementation desired by any person skilled in the art to implement the implementation of this application, with user or administrator control only in certain embodiments of this implementation. Selection may be made by the user via a user interface or other input means, or implemented via a desired algorithm. The exemplary implementations described herein may be used individually or in combination, and the functions of the exemplary implementations may be implemented by any means in a desired implementation.

[0049] Figure 1 shows an exemplary system configuration diagram 100 for performing process evaluation and monitoring in an exemplary implementation. As shown in Figure 1, the system configuration diagram 100 may include a first component 110, a second component 120, a third component 130, and so on. In the first component 110, checkpoints to be monitored are selected in the automated process. The selected checkpoints correspond to one or more stages in the process that need to be checked in order to determine whether the process is in a normal state or to perform performance evaluation, etc.

[0050] In some exemplary implementations, one or more robots may be used for handling / manipulating objects. Checkpoints can be selected based on one or more factors, such as the robot's posture or the position of the object being manipulated. For example, if the focus of monitoring is on the state of the object being handled / manipulated by the robot, one checkpoint could be when the object is first grasped by the robot, and another could be, for example, when the robot releases the object. Figure 2 shows exemplary checkpoints selected based on the robot's posture in an exemplary implementation. The moment the robot grasps the object is set as the first checkpoint 210. The waypoint for object transport is set as the second checkpoint 220. The moment the object is released is set as the third checkpoint 230. If the focus of monitoring is on the robot, multiple robot postures can be selected as checkpoints.

[0051] In some exemplary implementations, checkpoints can be set using sensor signals. Specifically, a particular sensor signal indicating a critical transition point in the process can be received and used for checkpoint setting. For example, if a sensor signal indicates that the process is entering the next phase (e.g., when a barcode scanner detects an object, a proximity sensor detects an object, or a laser sensor reads a specific value), it is important to monitor whether the transition is proceeding as planned. The sensors used to generate the sensor signals may include, but are not limited to, one or more, vision sensors (video cameras, cameras, etc.), radio frequency identification (RFID) scanners, scanners, and quick response (QR) scanners. Figure 3 shows exemplary checkpoints selected based on sensor signals. As shown in Figure 3, a data scan / read performed by a laser sensor may be the first checkpoint 310, and a barcode scan may be the second checkpoint 320. In alternative exemplary implementations, checkpoints may be selected using variables that depend on sensor measurements. In some exemplary implementations, the selection of checkpoints may be obtained from, or performed by, other management systems / software, such as warehouse management systems (WMS) or warehouse operations management systems (WES), although this is not limited to these systems.

[0052] In the second component 120, at each checkpoint, specified system variables are measured (measured values ​​122), and specified images of the system components (images / image-based information 126) are captured. Images may be captured using vision sensors such as video cameras, cameras, and mobile devices, but are not limited to these. Furthermore, predictive images 124 are generated based on the measured values ​​122. The measured values ​​122 include, but are not limited to, the joint angles of a robot / robot arm, the weight of an object grasped by the robot, the dimensions of an object, an object barcode, an object's stockkeeping unit (SKU), or radio frequency identification (RFID) on the object's packaging, which may be acquired via one or more sensors (scanners, mobile devices, scales, video cameras, cameras, etc.).

[0053] Image / image-based information 126 contains information necessary to determine the state and / or operation of the system components. Such information may include, but is not limited to, images of the robot / robot arm grasping the object, images of sensor light, and images of the pallet containing the grasped object. Figure 4 shows exemplary images captured at various checkpoints. As shown in Figure 4, images 412 and 414 may be taken at the first checkpoint 410. Image 412 is an overall image of the robot 402 and the object 404 being grasped by the robot 402. Image 414 contains only an image of the object 404. At the second checkpoint 420 (transport phase), images 422 and 424 are captured. Image 422 includes the robot 402 and the object 404 and is taken at the transport waypoint. Image 424 contains only an image of the object 404 taken at the transport waypoint. At the third checkpoint 430 (release phase), image 432 is captured. Image 432 includes an image of object 404 and a destination region 406 where object 404 is located.

[0054] In some exemplary implementations, the region of the system from which images are captured can be determined dynamically. For example, images of the area / region where the most significant motion is occurring may be captured. Significant motion could be the velocity of an object in 3D space along the direction of observation, or the motion of an object visualized in 2D from a given camera view.

[0055] The predicted image 124 is an image derived based on the measured values ​​122, predicting what the system (robot, environment, object, etc.) will look like under normal circumstances. This assumes that no failures or anomalies occur. The goal is to predict what the normal state will look like if the process proceeds as planned. In some exemplary implementations, the prediction is performed using a mathematical model. Figure 5 shows an exemplary predictive image generation process 500 using a mathematical model in an exemplary implementation. This is suitable when all states to be checked are predictable based on a linear mathematical relationship between system components and measured values. The advantage is that it is usually computationally inexpensive. For example, given measured robot joint angles (measured values ​​502), robot model information 504, and recognized object pose 506, it is possible to calculate what the system will look like when an object is grasped by the robot (predicted system 508).

[0056] In some exemplary implementations, a digital model of the system is generated and used to derive predictive images 124 by visualizing the system in 2D / 3D. The predictive images 124 can be derived from the digital model after updating the model with measured system variables. For example, 3D and 2D visualizations of a robotic arm are available if a digital model of that robotic arm exists. 3D visualization of the robotic arm is available if measured joint angles are available, and 2D images of the robotic arm are available once captured by a camera. The geometric shape of the pallet can also be predicted once packages / objects are placed on the pallet according to a planned order. A digital twin, which is an exact copy of the physical system, is another example of such a digital model. The advantage of using a digital model is that predictive images can be obtained directly.

[0057] In some exemplary implementations, 2D and 3D motion can be monitored, such as the movement of a (mobile or stationary) robot, the movement of an object being transported, or a combination of a robot and an object. 3D motion can be calculated using a mathematical model of the system. Figure 6 shows an exemplary 3D robot motion visualization 600 from an exemplary implementation. As shown in Figure 6, the 3D motion of a robot arm can be visualized in 2D using optical flow. Specifically, 3D motion can be predicted and represented using 2D images. The advantage of motion monitoring is that, compared to discrete checkpoints, any process anomaly can be detected because the motion is continuous.

[0058] In some exemplary implementations, the predicted image 124 can be generated using a trained machine learning model. This is recommended when there is no clear mathematical model to make predictions, or when it is faster to make predictions using a trained model than using a mathematical model. Using a trained machine learning model allows for greater flexibility in responding to changes in normal operations. For example, if the robot's reach extends to various positions on a pallet to pick up an object, a trained machine learning model can be used to generate predictions that take into account variations in the robot's posture.

[0059] Figure 7 shows an exemplary process 700 for training a machine learning model to generate predictive images, according to an exemplary implementation. During the offline training phase, a trainable machine learning model 706 is trained using (i) robot joint angles observed on the robot 704 and (ii) a 2D view of the robot 704 in camera 702. Once the trainable machine learning model 706 is trained, a trained machine learning model 708 is generated. In some exemplary implementations, the trainable machine learning model 706 is further trained using the relative position between the robot and the camera. Figure 14 shows an alternative exemplary process 1400 for training a machine learning model to generate predictive images, according to an exemplary implementation. In some exemplary implementations, the relative position 1410 between the robot and the camera may be determined by 3D motion capture using camera 702 on the robot 704.

[0060] A trainable machine learning model 706 / trained machine learning model 708 may include, but is not limited to, one or more artificial intelligence (AI) / machine learning (ML) models, such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), deep RNNs (DRNNs), Q-learning networks (QNs), deep Q-learning networks (DQNs), linear regression, decision trees, and K-nearest neighbors. RNNs may include long short-term memory (LSTMs), large language models (LLMs), etc.

[0061] During the online prediction phase, the current robot joint angles 710 and the relative position 712 between the robot and the camera can be input into a trained machine learning model 708 to generate a 2D view 714 of the robot. The 2D view 714 (predicted image 124) of the robot is then compared to the actual view (image / image-based information 126) to identify anomalies. Because 2D / 3D digital models and visualizations are unnecessary, the required computational resources can be significantly reduced.

[0062] Figure 8 shows an exemplary process flow 800 for training a machine learning model that generates predictive images / 2D motion images, according to an exemplary implementation. During the offline training phase, the trainable machine learning model 850 is trained using (i) robot joint angles 820-1, 820-2, ..., 820-n, and (ii) 2D motion images 830 of the robot / system. In some exemplary implementations, the trainable machine learning model 850 is further trained using the relative position between the robot and the camera. Figure 15 shows an alternative process flow 1500 for training a machine learning model that generates predictive images / 2D motion images, according to an exemplary implementation. In some exemplary implementations, the relative position between the robot and the camera 1510 may be determined by 3D motion capture using a camera pointed at the robot / system.

[0063] Robot joint angles 820-1 to 820-n and 2D motion images 830 of the robot / system are captured / measured at various timestamps 1-n. In some exemplary implementations, the 2D motion images 830 may be optical flow images. The 2D motion images 830 of the robot / system may be derived using a 3D motion sequence of the robot / system captured at timestamps 1-n and / or a 2D view of the robot / system taken at timestamps 1-n.

[0064] Once the trainable machine learning model 850 is trained, a trained machine learning model 860 is generated, which generates a 2D motion image of the robot / system as a predictive image. The trainable machine learning model 850 / trained machine learning model 860 may include, but is not limited to, one or more artificial intelligence (AI) / machine learning (ML) models, such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), deep RNNs (DRNNs), Q-learning networks (QNs), deep Q-learning networks (DQNs), linear regression, decision trees, and K-nearest neighbors. RNNs may include long short-term memory (LSTMs), large-scale language models (LLMs), etc.

[0065] Figure 9 shows an exemplary process flow 900 for generating a 2D motion image using the trained machine learning model 860 from Figure 8, in an exemplary implementation. During the online prediction phase, a sequence of robot joint angles 910 and the relative position 920 between the robot and the camera are input to the trained machine learning model 860 to generate a predicted 2D motion image 930 (predicted image 124), which can then be compared to the actual view (image / image-based information 126) to identify anomalies. Because 2D / 3D digital models and visualizations are not required, the computational resources required can be significantly reduced.

[0066] Referring back to Figure 1, the third component 130 performs image evaluation by comparing the image / image-based information 126 with the predicted image 124 to identify anomalies. Figure 10 shows an exemplary figure 1000 illustrating a method for comparing the predicted image and the captured image in an exemplary implementation. As shown in Figure 10, a predicted image 1002 can be generated from a digital model and compared with the captured image 1004 for anomaly detection. By comparing the predicted image 1002 with the captured image 1004, anomalies and irregularities can be detected in the process.

[0067] While not limited to these, anomalies such as machine / robot behavior, objects, and object appearance are detected when the similarity between the predicted image 124 and the image / image-based information 126 is lower than a threshold (e.g., a preset percentage threshold, an adjustable numerical threshold) or according to a customized metric.

[0068] Figure 11 shows an exemplary Figure 1100 illustrating how image comparison is performed in an exemplary implementation. As shown in Figure 11, based on the received measurements, two predicted images 1102 and 1104 are generated at checkpoint 1 and can be compared with actual images 1106 and 1108 captured at checkpoint 1 for anomaly detection. Images 1102 and 1106 may include the entire robot / system, while images 1104 and 1108 may be directed at the grasped object. By comparing the predicted images 1102 and 1104 with the captured actual images 1106 and 1108, anomalies and irregularities in the process can be detected.

[0069] In some exemplary implementations, the comparison is performed by directly comparing image pixels. For example, if the image contains depth information, the depth of each pixel can be compared. If the image has multiple channels and contains color information, each pixel in each channel of the image can be compared.

[0070] In some exemplary implementations, the comparison is performed by first transforming the images. For example, the images may first be converted into vectors, and the similarity of the vectors is used to represent the similarity of the images. In alternative exemplary implementations, the images may also be converted into words that describe them, and then the image descriptions can be compared.

[0071] In some exemplary implementations, comparisons are made using information extracted from images. For example, the distribution of pixel values ​​in two images can be compared, or the behavior of a robot / object visualized in a 2D image (such as the optical flow of an image) can be compared.

[0072] Figure 12 shows an alternative exemplary system configuration diagram 1200 for performing process evaluation and monitoring in an exemplary implementation. As illustrated in Figure 12, the system configuration diagram 1200 may include a first component 110, a second component 120, a third component 130, a fourth component 1210, and so on. The first component 110, the second component 120, and the third component 130 are treated as identical to those in Figure 1. The evaluation is performed by the third component 130, and once completed, a determination is then made to determine whether an anomaly exists / is detected. If an anomaly exists / is detected, a recovery process is performed, which corresponds to the fourth component 1210.

[0073] In some exemplary implementations, the recovery process may be performed / initiated automatically without operator / human input. For example, a robot / robot arm may be controlled to set aside a damaged object / package, place it in a designated location, re-grasp the object, or pick up a dropped object / package. In alternative exemplary implementations, the recovery process is performed with human assistance. For example, an operator may determine the best recovery process based on information received from the system (e.g., warnings). In some exemplary implementations, the system may alert the operator to an anomaly and provide recommended recovery actions for the operator to review and perform.

[0074] The aforementioned exemplary implementations may have various advantages and merits, such as an unconventional method for detecting system anomalies by generating predictive images and comparing them to actual operation / images. This method provides an efficient solution for monitoring process conditions that cannot be directly measured. Furthermore, it allows for timely detection of anomalies to avoid prolonged downtime, system damage, and quality problems. Moreover, recovery actions can be implemented immediately (manually or automatically) without causing further delays in the process.

[0075] Figure 13 shows an exemplary computing environment having exemplary computer devices suitable for use in several exemplary implementation forms. The computer device 1305 within the computing environment 1300 may include one or more processing units, cores, or processors 1310, memory 1315 (e.g., RAM, ROM, and / or similar), internal storage 1320 (e.g., magnetic, optical, solid-state storage, and / or organic storage), and / or I / O interfaces 1325, any of which may be connected to a communication mechanism or bus 1330 for communicating information, or incorporated into the computer device 1305. The I / O interface 1325 may also be configured, depending on the desired implementation form, to receive images from a camera or provide images to a projector or display.

[0076] The computer device 1305 can be communicatively coupled to an input / user interface 1335 and an output device / interface 1340. Either or both of the input / user interface 1335 and the output device / interface 1340 can be wired or wireless interfaces and can be detachable. The input / user interface 1335 may include any physical or virtual device, component, sensor, or interface that can be used to provide input (e.g., buttons, touchscreen interfaces, keyboards, pointing / cursor controls, microphones, cameras, Braille, motion sensors, accelerometers, optical readers, and / or similar). The output device / interface 1340 may include displays, televisions, monitors, printers, speakers, Braille, etc. In some exemplary implementations, the input / user interface 1335 and the output device / interface 1340 can be built into or physically connected to the computer device 1305. In other exemplary implementations, other computer devices may function as, or provide, an input / user interface 1335 and an output device / interface 1340 for computer device 1305.

[0077] Examples of computer devices 1305 may include, but are not limited to, highly mobile devices (e.g., smartphones, devices in vehicles and other machines, devices carried by humans or animals), mobile devices (e.g., tablets, notebooks, laptops, personal computers, portable televisions, radios, etc.), and devices not designed for mobility (e.g., desktop computers, other computers, information kiosks, televisions, radios, etc., with one or more processors incorporated and / or combined).

[0078] Computer device 1305 can be communicatively connected (for example, via IO interface 1325) to external storage 1345 and network 1350 to communicate with any number of network components, devices, and systems, including one or more computer devices of the same or different configurations. Computer device 1305 or any connected computer device may function as, provide, or be referred to as, a server, client, thin server, general-purpose machine, dedicated machine, or by any other name.

[0079] The IO interface 1325 may include, but is not limited to, wired and / or wireless interfaces using any communication or IO protocol or standard (e.g., Ethernet, 802.11x, Universal System Bus, WiMAX®, modem, cellular network protocol, etc.) for communicating information to and from at least all connected components, devices, and networks within the computing environment 1300. The network 1350 may be any network or combination of networks (e.g., the Internet, local area network, wide area network, telephone network, cellular network, satellite network, etc.).

[0080] Computer device 1305 may use / communicate using computer-usable media or computer-readable media, including temporary and non-temporary media. Temporary media include transmission media (e.g., metal cables, optical fibers), signals, carrier waves, etc. Non-temporary media include magnetic media (e.g., disks and tapes), optical recording media (e.g., CD-ROMs, digital video discs, Blu-ray discs), solid-state media (e.g., RAM, ROMs, flash memory, solid-state storage), and other non-volatile storage devices or memories.

[0081] Computer device 1305 can be used to implement techniques, methods, applications, processes, or computer executable instructions in several exemplary computing environments. Computer executable instructions can be obtained from temporary media, stored in non-temporary media, and retrieved from non-temporary media. Executable instructions can be generated from one or more arbitrary programming languages, scripting languages, and machine code (e.g., C, C++, C#, Java, Visual Basic, Python, Perl, JavaScript, etc.).

[0082] The processor 1310 can run under any operating system (OS) (not shown) in a native or virtual environment. One or more applications can be deployed, which include a logical unit 1360, an application programming interface (API) unit 1365, an input unit 1370 and an output unit 1375, and an inter-unit communication mechanism 1395 for communication between different units, for communication with the OS, and for communication with other applications (not shown). The units and elements described are subject to change by design, function, configuration, or implementation, and are not limited to the description provided. The processor 1310 can take the form of a hardware processor such as a central processing unit (CPU), or a combination of hardware and software units.

[0083] In some exemplary implementations, when information or execution instructions are received by the API unit 1365, they may be communicated to one or more other units (e.g., a logical unit 1360, an input unit 1370, and an output unit 1375). In some cases, the logical unit 1360 may be configured to control the flow of information between units and to direct the services provided by the API unit 1365, the input unit 1370, and the output unit 1375 in some exemplary implementations described above. For example, the flow of one or more processes or implementations may be controlled by the logical unit 1360 alone or in conjunction with the API unit 1365. The input unit 1370 may be configured to take the inputs necessary for the computation described in the exemplary implementation, and the output unit 1375 may be configured to provide outputs based on the computation described in the exemplary implementation.

[0084] The processor 1310 can be configured to select a first system checkpoint for performing process monitoring on the robot system, as shown in Figures 1 and 7-8. The processor 1310 can also be configured to measure system variables using one or more sensors at the first system checkpoint, as shown in Figures 1 and 7-8. The processor 1310 can also be configured to capture one or more images of the robot system using vision sensors at the first system checkpoint, as shown in Figures 1 and 7-8. The processor 1310 can also be configured to generate at least one predicted image of the robot system using a predictive model that takes system variables as input to the predictive model, as shown in Figures 1 and 7-8. The processor 1310 can also be configured to perform anomaly detection by comparing one or more images of the robot system with at least one predicted image, as shown in Figures 1 and 7-8.

[0085] The processor 1310 can be configured to automatically issue commands to the robot arm to perform a recovery process if an anomaly is detected, as shown in Figure 12. The processor 1310 can also be configured to select a first system checkpoint based on the posture of the robot arm or the position of an object, as shown in Figures 1 and 7-8.

[0086] The processor 1310 can be configured to derive the relative position between the robot arm and the vision sensor using one or more images of the robot system, as shown in Figures 6 to 8. The processor 1310 can be configured to derive the joint angles of the robot arm using one or more images of the robot system, as shown in Figures 6 to 8. The processor 1310 can be configured to select a first system checkpoint based on the reception of sensor signals indicating process transitions, as shown in Figure 3.

[0087] Some parts of the detailed explanation are presented with respect to algorithms and symbolic representations of computer operations. These descriptions and symbolic representations of algorithms are a means used by those skilled in the field of data processing technology to communicate the essence of their innovations to others skilled in the field. An algorithm is a set of defined steps that bring about a desired final state or result. In exemplary implementations, the steps performed require the physical manipulation of specific quantities to achieve a specific outcome.

[0088] Unless otherwise specified, as is evident from the description, descriptions using terms such as “process,” “computer process,” “calculate,” “determine,” and “display” throughout the description may include actions and processes of a computer system or other information processing device that manipulate and convert data represented as physical (electronic) quantities in the registers and memory of a computer system into other data similarly represented as physical quantities in the memory or registers of a computer system or other information storage device, transmission device, or display device.

[0089] Exemplary implementations may also relate to apparatus for performing the operations described herein. This apparatus may include one or more general-purpose computers, which may be specifically configured for a desired purpose or selectively invoked or reconfigured by one or more computer programs. Such computer programs may be stored on computer-readable media, such as computer-readable storage media or computer-readable signal media. Computer-readable storage media may include tangible media, such as, but not limited to, optical disks, magnetic disks, read-only memory, random-access memory, solid-state devices, and drives, or any other type of tangible media or non-temporary media suitable for storing electronic information. Computer-readable signal media may include media such as carrier waves. The algorithms and displays presented herein are not inherently related to any particular computer or other apparatus. Computer programs may include pure software implementations containing instructions for performing the operations of the desired implementation.

[0090] Various general-purpose systems may be used with the programs and modules illustrated herein, or it may be convenient to construct more specialized devices to carry out the steps of the desired method. Furthermore, the exemplary implementations are not described with reference to a specific programming language. It will be understood that various programming languages ​​may be used to carry out the teachings of the exemplary implementations described herein. Instructions in a programming language may be executed by one or more processing devices, such as a central processing unit (CPU), processor, or controller.

[0091] As is known in the art, the operations described above can be performed by hardware, software, or any combination of software and hardware. Various aspects of the exemplary implementations may be implemented using circuit and logic devices (hardware), and in other aspects, they may be implemented using instructions stored in a machine-readable medium (software), which, when executed by a processor, cause the processor to perform the method for carrying out the implementations of this application. Furthermore, some exemplary implementations of this application may be implemented in hardware only, and other exemplary implementations may be implemented in software only. Furthermore, the various functions described may be implemented in a single unit or distributed among multiple components in any various way. When implemented in software, the method may be executed by a processor such as a general-purpose computer based on instructions stored in a computer-readable medium. If necessary, the instructions may be stored in the medium in a compressed and / or encrypted form.

[0092] Furthermore, other implementations of this application will be apparent to those skilled in the art from the considerations herein and the practice of the teachings herein. Various aspects and / or components of the exemplary implementations described may be used individually or in any combination. This specification and the exemplary implementations are intended to be considered as examples only, and the true scope and spirit of this application are set forth by the following claims. [Explanation of Symbols]

[0093] 100 System Configuration Diagram 110 Components (1) 120 Components (2) 122 measurement values 124 Predicted Images 126 Images / Image-based information 130 Components (3) 210 First checkpoint 220 Second checkpoint 230 Third Checkpoint 310 First checkpoint 320 Second checkpoint 402 Robots 404 Object 406 Destination Area 410 First checkpoint 412 images 414 images 420 Second checkpoint 422 images 424 pixels 430 Third Checkpoint 432 images 500 Predictive Image Generation Processes 502 measurement values 504 Robot Model Information 506 Object pose 508 predicted systems 600 3D robot motion visualizations 700 processes 702 Camera 704 Robot 706 trainable models, trainable machine learning models 708 Trained Models, Trained Machine Learning Models 710 Robot joint angles 712 Relative position between robot and camera 714 Robot 2D View 800 Process Flows 820-1 Joint angle at timestamp 1 820-2 Joint angle at timestamp 2 820-n Joint angle at timestamp n 830 2D motion images 850 trainable models, trainable machine learning models 860 trained models, trained machine learning models 900 Process Flows 910 Robot joint angle sequence 920 Relative position between the robot and the camera 930 Predictive 2D motion images 1000 Figures 1002 Predicted Images 1004 Captured Images Figure 1100 1102 Predicted image 1104 Predicted image 1106 Actual image 1108 Actual image 1200 System Configuration Diagram 1210 Components (4) 1300 Computing Environments 1305 Computer Devices 1310 Processor 1315 memory 1320 internal storage 1325 I / O Interfaces 1330 Bus 1335 Input / User Interface 1340 Output Devices / Interfaces 1345 External Storage 1350 Network 1360 Logical Units 1365 API units 1370 Input Unit 1375 Output Unit 1395 Inter-unit communication mechanism 1400 processes 1410 Relative position between the robot and the camera 1500 Alternative Process Flows 1510 Relative position between the robot and the camera

Claims

1. A process monitoring system, One or more sensors, The processor includes a processor that communicates with one or more of the sensors, and the processor Select a first system checkpoint for performing process monitoring on the robot system. At the first system checkpoint, the system variables are measured using one or more of the sensors. At the first system checkpoint, one or more images of the robot system are captured using a vision sensor. Using the prediction model that uses the aforementioned system variables as input to the prediction model, at least one predicted image of the robot system is generated. A process monitoring system configured to perform anomaly detection by comparing one or more images of the robot system with at least one predicted image.

2. The robot system includes a robotic arm configured to grasp and move an object, The system variables include one or more of the joint angles of the robot arm, the weight of the object, the dimensions of the object, the stockkeeping unit (SKU) of the object, or the barcode of the object. The process monitoring system according to claim 1.

3. The aforementioned processor The process monitoring system according to claim 2, further configured to automatically issue a command to the robot arm to perform a recovery process when an abnormality is detected.

4. The process monitoring system according to claim 2, wherein the processor is configured to select the first system checkpoint based on the posture of the robot arm or the position of the object.

5. The process monitoring system according to claim 2, wherein the one or more images captured using the vision sensor include one or more images of the robot arm and / or the object at the first system checkpoint.

6. The aforementioned processor Using one or more images of the robot system, the relative position between the robot arm and the vision sensor is derived. The system is further configured to derive the joint angles of the robot arm using one or more images of the robot system, The process monitoring system according to claim 2, wherein generating the at least one predictive image includes using the relative position between the robot arm and the vision sensor and the joint angle of the robot arm as inputs to the predictive model in order to generate the at least one predictive image.

7. The process monitoring system according to claim 6, wherein the predictive model is an artificial intelligence (AI) model trained using historical system data including (i) past joint angles of the robot arm and (ii) past motion images of the robot arm.

8. The process monitoring system according to claim 7, wherein the past motion images of the robot arm include two-dimensional motion images of the robot arm.

9. The one or more images of the robot system include an image of the robot arm and an image of the object, The at least one predicted image of the robot system includes a robot arm predicted image and an object predicted image, Comparing one or more images of the robot system with at least one predicted image includes comparing the image of the robot arm with a predicted image of the robot arm, and comparing the image of the object with a predicted image of the object. The process monitoring system according to claim 2.

10. The process monitoring system according to claim 1, wherein the processor is configured to select the first system checkpoint based on the reception of a sensor signal indicating a process transition.

11. A process monitoring method, The processor selects a first system checkpoint for performing process monitoring on the system, The processor performs the steps of measuring system variables using one or more sensors at the first system checkpoint, The processor performs the steps of capturing one or more images of the system using a vision sensor at the first system checkpoint, The processor generates at least one predicted image of the system using the prediction model which uses the system variables as input to the prediction model. A process monitoring method comprising the step of performing anomaly detection by comparing one or more images of the system with at least one predicted image using the processor.

12. The system includes a robotic arm configured to grasp and move an object. The system variables include one or more of the joint angles of the robot arm, the weight of the object, the dimensions of the object, the stockkeeping unit (SKU) of the object, or the barcode of the object. The process monitoring method according to claim 11.

13. The further step includes the processor automatically issuing a command to the robot arm to perform a recovery process in response to the detected anomaly, The process monitoring method according to claim 12.

14. The process monitoring method according to claim 12, wherein the processor is configured to select the first system checkpoint based on the posture of the robot arm or the position of the object.

15. The process monitoring method according to claim 12, wherein the one or more images captured using the vision sensor include one or more images of the robot arm and / or the object at the first system checkpoint.

16. The processor performs the steps of using one or more images of the system to derive the relative position between the robot arm and the vision sensor, The processor further includes the step of deriving the joint angles of the robot arm using one or more images of the system, The step of generating the at least one predicted image includes using the relative position between the robot arm and the vision sensor and the joint angle of the robot arm as inputs to the prediction model in order to generate the at least one predicted image. The process monitoring method according to claim 12.

17. The process monitoring method according to claim 16, wherein the predictive model is an artificial intelligence (AI) model trained using past system data including (i) past joint angles of the robot arm and (ii) past motion images of the robot arm.

18. The process monitoring method according to claim 17, wherein the past motion images of the robot arm include two-dimensional motion images of the robot arm.

19. The one or more images of the system include the image of the robot arm and the image of the object, The at least one predicted image of the system includes a robot arm predicted image and an object predicted image, The processor is configured to compare one or more images of the system with at least one predicted image by comparing the image of the robot arm with the predicted image of the robot arm and the image of the object with the predicted image of the object. The process monitoring method according to claim 12.

20. The process monitoring method according to claim 11, wherein the processor is configured to select the first system checkpoint based on the reception of a sensor signal indicating a process transition.