Method for determining depth information data of an object for a robot

The method uses two independent imaging sensors and processing components for separate depth estimation and object recognition, addressing the challenge of differentiating obstacles in robotic systems, enhancing safety and accuracy while meeting industrial safety standards.

WO2026057159A1PCT designated stage Publication Date: 2026-03-19ABB (SCHWEIZ) AG
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-12
Publication Date
2026-03-19

AI Technical Summary

Technical Problem

Existing robotic systems face challenges in distinguishing between different types of obstacles, leading to unnecessary workflow interruptions, as they rely on expensive laser scanners that detect objects without differentiation, and existing imaging sensor approaches either require complex sensor fusion or single detection channels, making it difficult to achieve the required level of functional safety.

Method used

A method utilizing two independent imaging sensors and processing components to capture and process image data separately, employing machine learning algorithms for depth estimation and object recognition, ensuring redundancy and compliance with safety standards by independent depth information determination.

Benefits of technology

This approach enhances safety and accuracy by providing redundant depth information processing, allowing the robot to reliably distinguish between objects like persons and walls, ensuring timely and safe responses, and adhering to industrial safety standards with minimal additional cost and complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024075445_19032026_PF_FP_ABST
    Figure EP2024075445_19032026_PF_FP_ABST
Patent Text Reader

Abstract

The disclosure relates to a method (100) for determining depth information data of an object for a robot (1), the method comprising: - obtaining image data of an object in the environment of the robot (1) obtained by at least two imaging sensors (10, 10'), wherein the image data is obtained independently; - providing the image data for separate processing to each one of two or more processing components (42, 42'); - determining depth information data based on the processed image data, wherein the depth information data is indicative of the depth of the object in the environment of the robot (1).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] P240527W001 - 1 - 11.09.2024

[0002] ABB Schweiz AG A19286WO

[0003] METHOD FOR DETERMINING DEPTH INFORMATION DATA OF AN OBJECT FOR A ROBOT

[0004] TECHNICAL FIELD

[0005] The present invention relates to a method for determining depth information data of an object for a robot, one or more computer program products, a data processing system, and a robot.

[0006] BACKGROUND

[0007] Robots, such as autonomous mobile robots and certain types of collaborative industrial robots, are dependent on seamless monitoring of their surroundings, as they often operate in an environment with persons. To ensure the safety of personnel and comply with stringent industrial safety standards, these robots must reliably recognize the person in their vicinity and respond appropriately by slowing down or stopping as needed. Typically, this person recognition is realized with safety-tested laser scanners that detect obstacles within a defined field and provide two-channel inputs to the robot controller to trigger the required safety measures.

[0008] However, these laser scanners have significant drawbacks. They are expensive and only capable of detecting objects in general, without distinguishing between different types of obstacles, such as persons, other robots, or fixed structures like walls. This often leads to unnecessary interruptions in workflow, as the robot slows down or stops for every detected object.

[0009] To address these challenges, modern robots increasingly utilize imaging sensors for tasks like navigation and object identification. These sensors could also be leveraged for personnel detection.

[0010] However, existing approaches either rely on complex and costly sensor fusion techniques or use a single detection channel, making it difficult to achieve the required level of functional safety. P240527W001 - 2 - 11.09.2024

[0011] SUMMARY

[0012] The above problem or need is at least partially solved or alleviated by the subject matters of the independent claims of the present disclosure, wherein further examples are incorporated in the dependent claims.

[0013] According to an aspect of the present disclosure, there is provided a method for determining depth information data of an object for a robot, the method comprising:

[0014] - obtaining image data of an object in the environment of the robot obtained by at least two imaging sensors, wherein the image data is obtained independently;

[0015] - providing the image data for separate processing to each one of two or more processing components;

[0016] - determining depth information data based on the processed image data, wherein the depth information data is indicative of the depth of the object in the environment of the robot.

[0017] Image data may refer to digital information captured by an imaging sensor, such as a camera, particularly a monocular camera. Two or more imaging sensors capture image data of the robot's surroundings. These two or more imaging sensors may operate in a predefined field of view. This image data may comprise an array of pixel values that represent the visual characteristics of a scene or object, including brightness, colour, and other relevant image attributes in the environment of the robot. The image data may be obtained by at least two independent imaging sensors in the environment of the robot. Each imaging sensor collects image data independently, which means that the imaging sensors operate separately to capture visual information from the robot's environment, e.g. an object or a person. The independent operation of the imaging sensors may allow them to capture different perspectives or viewpoints of the environment, for example. The image data may comprise a single image, parts of an image or even image series or sequences, for example.

[0018] The image data collected by the independent imaging sensors may be sent to two or more processing components for separate processing. For example, this may mean that each processing component may use its own set of image data for further processing. These processing components may e.g. analyze and interpret the image P240527W001 - 3 - 11.09.2024 data to extract relevant information, such as the depth of objects in the robot’s environment. The processing of this data may be performed separately by each processing component. This may mean that each processing component may process the image data autonomously. For the processing, each processing component may use several image data processing methods. Examples for such image data processing methods may be edge or feature recognition, methods based on machine learning, such as neural networks for depth estimation, or similar. In particular, the two or more processing components may comprise algorithms for monocular depth estimation. These algorithms may be based on machine learning and / or trainings data and estimate the depth solely from the information of a single image, which is particularly important for monocular cameras. The algorithms may also take into account additional information such as the camera position, the movement vectors of the robot, or the lighting conditions, or similar in the environment of the robot, for example.

[0019] Depth information data of an object may be determined from each of the processed image data. By using two separate imaging sensors and two independent processing components, two independently working processing channels may be generated. Each channel may process the image data separately to determine the depth of objects in the robot's environment, for example. This separate capturing and processing may create redundancy, which may be important for safe operating of the robot, for example. Redundancy here may mean that a system consists of several independent components that fulfil the same task so that the overall system functions correctly even if one component fails. In particular, the invention may be based on a multi-channel architecture in which two or more independent imaging sensors and two or more separate processing components are used to determine depth information data. This may mean that the image data may be captured and processed in two or more parallel channels. Each channel may operate independently of the other and perform the depth determination of the object based on the image data it receives from one of the imaging sensors. Thus, the independent processing of the image data may generate an additional layer of safety that prevents a single point of failure from jeopardising the entire system, for example. Many industrial safety standards require redundant systems in order to obtain certification. The independent determination of depth P240527W001 - 4 - 11.09.2024 information data may comply with functional safety standards required in industrial applications such as ISO 13849-1 category 3, or I EC 61508 hardware fault tolerance of 1 , for example. In most industrial applications, such imaging sensors requiring personnel detection are already present. Therefore, the method may facilitate the fulfilment of safety requirements with minimal additional cost and complexity, as the sensors used are already present in most industrial applications.

[0020] Further, by processing the image data independently by two separate processing components, e.g. potential inaccuracies or distortions that could occur in a single processing path may be compensated or prevented. Analysing or processing the image data in parallel may allow to obtain more precise and reliable depth information, as different algorithms or models can be used that complement each other, for example. Thus, independent processing of the image data may also increase the reliability and accuracy of the process of determining the depth information data.

[0021] The depth information data of the object may be provided as an output to the robot. In particular, the depth information data may be provided to a control system of the robot. By determining the depth information data independently, the control system of the robot may be able to verify the consistency of the results or the depth information data and thus increase its accuracy. For example, if both processing components deliver the same result or the same depth information data, it is more likely that the depth information data determined is accurate. This may enable the robot to accurately determine the position of the object in its environment, for example. For example, the robot may use the depth information data to make decisions such as stopping the robot, slowing down or avoiding obstacles, or similar.

[0022] The method of the first aspect may in particular be an at least partially or fully computer implemented method. This means that at least one, multiple or all of the steps of the method may be carried out by a data processing system, which may comprise one or more data processing apparatuses, which may be in the form of computers or computing units, which may comprise one or more processors and data storages or memories. Different steps may be carried out by the same or by different data processing apparatuses of the data processing system. For example, the data P240527W001 - 5 - 11.09.2024 processing system may comprise processors or computing units (e.g. an onboard computer or GPUs for image processing or similar), machine learning accelerators and general computing hardware. It may also include software frameworks specifically designed for machine learning, image processing and real-time data processing, for example. The two or more processing components may be executed independently on a single processing apparatus. This may mean, two or more redundant processing components may be executed e.g. by one CPU. Therefore, the method may not depend on the number of processing apparatuses or computers, as long as the independency between the processing components is ensured.

[0023] In an example, the method may further comprise recognizing the object as one of two or more object types based on the depth information data, wherein the recognizing as an object type comprises at least one type for person recognition. For example, monocular depth estimation techniques may be used to determine differences in the image data e.g. caused by movements of the robot or by small differences in the images (when the sensors are slightly offset) to determine the depth of the objects. Machine learning models that specialise in image processing may use the depth information data and raw image data to recognise and classify objects. These machine learning models may be trained with large datasets containing annotated images on which different object types are labelled, for example. A specialized model or a special component within the model may be trained to recognize persons as an object type, for example. This may be utilized by combining depth information data and characteristic features of people (e.g. body shapes, movement patterns, face recognition or similar). Various algorithms could be used, such as convolutional neural networks (CNNs), which may be particularly well suited to image processing tasks. These networks may learn to recognise specific patterns in the image and depth data that indicate the presence of persons, for example.

[0024] For example, distinguishing between a wall and a person may be based on a combination of depth information data, image data and machine learning models. First, imaging sensors may capture the robot's surroundings or environment and provide this data to depth estimation algorithms that create depth information data. A wall may show a uniform depth distribution over a larger area, while a person generates different P240527W001 - 6 - 11.09.2024 depth values in the depth information data based on their body shape. In addition, machine learning models may analyze the image data for specific features such as edges, contours, textures, and colours. Walls often have straight lines and uniform textures, while the outlines and textures of people are more complex and varied. These differences may be recognised by e.g. algorithms that are trained to identify these specific patterns. Another aspect considered may be motion analysis. While walls are static, people usually move in the scene. For example, this movement may be recognized by analysing image sequences, to identify the presence of a person. The distinction between walls and people is only exemplary and not limiting. All possible objects may be used for differentiation.

[0025] In an example, a control command relating to a safety-related action may be provided to the robot based on the depth information data and / or the recognized object type. A control command may refer to instructions transmitted to the robot to perform a specific action based on the depth information and / or the recognised object type. For example, based on the depth information data and possibly additional image features (such as shape, texture, movement), a machine learning model identifies the objects in the robot's field of view. The processing system may classify these objects into different categories, e.g. ‘person’, ‘wall’, ‘obstacle’. Based on the depth information data obtained and / or the classification of the objects, the system may decide on safetyrelevant actions. This logic may be implemented in a software component that may act as a safety controller, for example. The safety controller may generate control commands that are forwarded to the robot's control unit. These control commands may then be translated into physical actions, e.g. a safety-related action, that ensure the safety of the robot and its environment. By using two separate imaging sensors and two independent processing components, two separate processing channels may be provided, wherein each channel may process the image data separately to determine the depth of objects in the robot's environment. This dual processing creates redundancy, which may be crucial for safety. Both processing components may work in parallel to ensure that the depth information data may be processed in real time, for example. This parallel processing may minimize delays and ensures that the robot may react quickly to changes in its environment. In some implementations, the processed depth information data of the two or more processing components could be combined P240527W001 - 7 - 11.09.2024 to produce a final, refined depth estimate. For example, an average of the two or more depth information data may be calculated. The decision on which safety measure to take (e.g. stop, slow down) may be based on the processed depth information data. By using two or more independent processing components or processing channels, the processing system may be able to continue to provide correct depth information data in the event of a failure or error on one channel, for example. This may be important in safety-critical environments, for example. For example, if a difference in depth information data of the two or more processing components is detected, the system may log this error and take action to identify and correct the erroneous channel if necessary.

[0026] In an example, the method may further comprise performing a safety- related action when a person is detected. The safety-related action may be at least one of the following: reducing the speed of the robot, changing a direction, switching to a safe operating mode, generating an acoustic signal, stopping the robot, or enabling any other suitable safety-related action. For example, the processing system may recognize a person in the robot's field of view using the dual determined depth information data and / or object recognition described before. As soon as the processing system recognises a person in the robot's travel path, a control command may be sent e.g. to a drive unit of the robot to reduce the speed of motion or similar. This may be achieved by reducing the motor power or by adjusting the control parameters, for example. Based on the detection of a person, the robot could switch to a special safe operating mode designed to minimise the risk of collisions or accidents. In this mode, the robot could severely restrict its movements, for example by travelling more slowly, limiting certain functions (e.g. lifting heavy loads), or by activating additional safety protocols. Safe mode could also include increased monitoring of the environment to continuously assess the situation. This means that the robot may switch to different modes based on object recognition. If for example a person is very close to the robot or standing directly in its path, the processing system may cause the robot to stop immediately. This stop may be achieved by quickly switching off the motors or by activating braking mechanisms, for example. Thus, the implementation of safety- related actions when a person is detected may ensure a safe working environment. These measures may enable the robot to react flexibly to the presence of people in its P240527W001 - 8 - 11.09.2024 environment and avoid potential hazards by automatically adapting its behaviour or similar, for example.

[0027] In an example, the method further comprises identifying similarities and / or differences in the received depth information data of the two or more processing components and; providing a control command to the robot based on the identified similarities and / or differences. As the depth information data is determined independently of each other by two different processing components, the results may be compared with each other to detect inconsistencies, for example. This may be referred to as cross-checking. An example may be that the two or more processing components provide different depth information data for the same object. For example, another algorithm may compare the depth information data of the two processing components point by point or on an aggregated basis (e.g. by comparing depth profiles or maps). This algorithm may be designed to recognize differences e.g. in the depth information data that could indicate possible errors or deviations. For example, differences could occur if a channel provides faulty data due to a sensor failure. The algorithm may also identify similarities in the depth information data that may indicate that the two or more processing components are providing consistent and correct information. Based on the similarities and / or differences, the processing system may calculate a confidence level for the depth information data obtained, for example. If the depth information data of both channels is very similar, a high confidence level may be achieved. If there are significant differences, the processing system may recognize a possible error in one of the two or more processing components or channels. In such cases, the processing system may decide to use the depth information data from the more reliable processing component or issue an error message, for example. If the depth information data is consistent in both channels, the processing system may generate a control command for the robot based on the reliable data. This could be, for example, the continuation of the robot's current movement or an adjustment of its speed, or a safety-related action, or similar, for example. The processing system could keep a log that may document all differences and similarities in the depth information data in order to be able to carry out error analyses afterwards and ensure the safety of the system. In cases where an error is detected in one channel, the erroneous data may be corrected or it could be switched to a safer alternative, such as a conservative estimate of depth, for example. P240527W001 - 9 - 11.09.2024

[0028] The use of two independent processing components or channels may provide redundancy that allows to make safe decisions even in the event of a partial failure.

[0029] In an example, the depth of the object may be determined based on a combination of the depth information data and / or based on the recognized object obtained from the two or more processing components. The depth information data provided by the two or more processing components may be combined to provide a more accurate and robust estimate of the object's depth. This may be achieved e.g. by using various methods such as averaging, weighted combination or similar. The depth could also be calculated based on the recognized object type. For example, the system could use different depth estimation models for different object types (e.g. people, walls, obstacles). For example, if a person may be recognized, a specific assumption about the person's size and shape to improve depth estimation may be provided. The depth information data may be combined using algorithms that analyse the differences and similarities in the depth information data to create an optimised depth estimation, for example. Further, the combined depth information data may be adjusted based on additional information, such as the type of object detected or the environment in which the robot is located. By combining depth information data from the two or more processing components, a more accurate depth estimate may be achieved. This may be important in complex environments where individual sensors or algorithms may not be reliable enough, for example. Additionally, by combining the depth information data with the recognized object types, it may be better assessed which reaction may be appropriate, e.g. stopping when a person is detected in the immediate vicinity, or similar.

[0030] In an example, each of the two or more processing components may be based on different machine learning models and / or training data. Each of the at least two processing components may use a separate machine learning model to analyse the image data and depth information data. These models may be designed differently to emphasise different aspects of these data. For example, one model may be a Convolutional Neural Network (CNN) specialized in object and depth information recognition in standard environments. It may be trained on a large dataset covering typical industrial scenarios. Another model may be a Recurrent Neural Network (RNN) P240527W001 - IQ - 11.09.2024 or a hybrid model specializing in dynamic scenarios, such as tracking moving objects. This model may be trained with data focused on moving objects and changing environmental conditions. The mentioned machine learning models are not limiting, any other model suitable for providing determining depth information may be conceivable. Further, the machine learning models may be trained on different datasets, e.g. to increase their specialization and adaptability. For example, one data set may contain mainly images and depth information from controlled, static environments, such as a standardized production area. Another data set may comprise images and depth information from dynamic, real-world environments, including variable lighting conditions, movements, and unpredictable obstacles. This diversity may ensure that the machine learning models may perform effectively in a variety of situations. The predictions from the different machine learning models may be combined to provide a more accurate and robust estimation of depth information and object recognition, for example. By utilizing different machine learning models and training data, resilience to various types of disturbances or changes in the environment may be achieved, for example.

[0031] In an example, the method may further comprise determining a distance between the object and the robot based on the generated depth information data. At least two processing components, each of which independently processes the image data captured by the robot’s imaging sensors may be used to generate depth information data about the objects in the environment. Each processing component may apply its own algorithm or machine learning models to analyze the sensor data and produce depth information data. The depth information data from the different processing components may be combined to determine a precise distance between the object and the robot, for example. This may be done through several methods such as averaging, where the average of the depth estimates is calculated and provided by each processing component. Another method may be outlier detection mechanisms to identify and discard any anomalous depth estimates that significantly differ from others, thereby improving the accuracy of the final distance measurement, for example. Determining a distance between the object and the robot may allow the robot to accurately determine the distance to objects in its environment by leveraging depth information data from multiple processing components. By combining these depth P240527W001 - 11 - 11.09.2024 information data, a more reliable and precise measurement may be achieved, which may be important for safe and effective robot operation.

[0032] In an example, a depth map of the object in the environment of the robot may be generated for each of the two imaging sensors. Each imaging sensor may independently process the captured images using depth estimation algorithms. These algorithms may analyze the image data to calculate the distance of each pixel in the image from the sensor, resulting in a depth map. If monocular cameras may be used, machine learning models, such as convolutional neural networks (CNNs), may be trained to infer depth from single images based on learned patterns and features. The output of this process may be a depth map. A depth map is essentially a two- dimensional grid where each point (pixel) represents the distance from the sensor to the corresponding point in the environment. Each sensor produces its own depth map, providing two independent views of the environment. A depth map may provide a pixel- by-pixel representation of the distance of objects from the sensor, offering a highly detailed view of the environment. This may allow the robot to understand not just the presence of obstacles, but their precise shape, size, and spatial relationship to other objects and the robot itself. For example, by analyzing the depth map, the shape and size of objects may be determined more accurately. For example, a person typically has a distinct vertical profile with varying depth (head, torso, limbs), whereas a wall or other static object has a more uniform depth profile. This makes it easier e.g. to differentiate between a person and other objects.

[0033] In an example, the image data by at least two imaging sensors may correspond to a field of view in an environment of the robot. The fields of view of the two imaging sensors may be strategically designed to either overlap or cover adjacent areas within the robot’s environment. This may ensure that the robot receives comprehensive image data, even though each imaging sensor operates independently. The imaging sensors may work in synchronization, meaning that they capture images of the same environment at the same time, allowing for consistent and comparable data across the different sensors. Each imaging sensor may independently capture visual cues like object size, texture gradient, and occlusion, which may be used to estimate depth. By comparing these cues from both imaging sensors, the accuracy of depth perception P240527W001 - 12 - 11.09.2024 may be enhanced. By capturing the environment from two slightly different perspectives, these image data may be combined to infer depth information data more effectively. Additionally, the combination of image data from two or more imaging sensors may provide more information about the relative positioning of objects, leading to more accurate depth information data than what a single monocular camera could achieve, for example.

[0034] In an example, additional sensor data comprising kinematic data of the robot may be provided for separate processing to each one of two or more processing component, wherein the kinematic data is indicative for one or more the position, velocity, and acceleration of the robot. This means, additional sensors providing kinematic data (e.g., position, velocity, acceleration) may be utilized to enhance the accuracy of depth information data and object recognition, for example. For instance, knowing the robot's velocity and position may help in adjusting depth information data or object detection algorithms to account for the robot's movement. For example, when the robot is moving, the kinematic data may help in compensating for motion, allowing for more accurate depth estimation. Additionally, kinematic data may allow to predict the future position of the robot, which may be used to adjust the depth estimation in real-time, enhancing accuracy. Further, the kinematic data may help in distinguishing between the robot’s own movement and the movement of objects in the environment, for example. This distinction may be important for tracking objects accurately, especially when both the robot and objects are in motion. Kinematic data may further improve the robot's ability to predict and avoid collisions by understanding how its own movement will affect its interaction with nearby objects. Thus, the robot may plan more efficient and safer paths by integrating its current position, velocity, and acceleration into the processing of the depth information data. The additional sensors providing the kinematic data are already available in most industrial applications, which means that the kinematic data may be implemented in the processing with minimal additional effort. These additional sensors may therefore be used to support the determining of depth information data and object recognition to further ensure functional safety.

[0035] In an example, the additional sensor data may comprise time stamps, indicative for the time at which the image data was obtained by at least two imaging sensors. This may P240527W001 - 13 - 11.09.2024 mean that additional sensors provide time stamps to improve the accuracy of depth information data and object recognition. For example, time stamps may be metadata that record the precise time when each image is captured by the imaging sensor. The imaging sensors or the processing components may generate timestamps when each image is captured. Therefore, time stamps for the image data may be useful for comparing image series or between two imaging sensors, for example. These time stamps may be provided by a high-precision clock synchronized across the processing system, for example. Time stamps may be embedded in the image data as metadata, which travels along with the image data through the processing system. This metadata may then be used in various processing stages, such as depth estimation, object recognition, and navigation, for example. Time stamps may allow to synchronize image data with other sensor data (e.g., kinematic data) and control commands, for example. This synchronization may be important for cohesive and accurate decision-making. Further, time stamps may provide a record of when each image was captured, which may be valuable for event logging and post-event analysis, for example. This may be important for debugging, diagnostics, or understanding the sequence of events leading to a particular action or decision, or similar. The use of kinematic data and time stamps may provide an additional level of accuracy and safety that may be important in industrial applications.

[0036] According to a second aspect of this disclosure, there are provided one or more computer program products comprising instructions which, when executed by one or more data processing apparatuses, cause the one or more data processing apparatuses to carry out the method of the first aspect of this disclosure.

[0037] The computer program products may be a computer program or computer programs as such, meaning a computer program consisting of or comprising program code to be executed by the data processing apparatus, in particular computer.

[0038] Alternatively, the one or more computer program products may be products such as data storages, in particular computer-readable data storage mediums, on which the computer programs may be temporarily or permanently stored. P240527W001 - 14 - 11.09.2024

[0039] According to a third aspect of this disclosure, there is provided a data processing system configured to carry out the method according to the first aspect of this disclosure.

[0040] According to a fourth aspect, there is provided a robot configured to carry out the method according to the first aspect of this disclosure. The term robot in the patent application is understood broadly and can include a variety of autonomous or semi- autonomous machines used in different environments. In the context of the invention described, the robot could take various forms such as Industrial Robots, Collaborative Robots (Cobots), Autonomous Mobile Robots (AMRs), Service Robots, Logistics and Delivery Robots, Agricultural Robots, Construction Robots, Robots for use in public spaces, for example.

[0041] It is noted that the above aspects, examples, and features may be combined with each other irrespective of the aspect involved.

[0042] The above and other aspects of the present disclosure will become apparent from and elucidated with reference to the examples described hereinafter.

[0043] BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Exemplary embodiments will be further described with reference to Figures, wherein:

[0045] Figure 1 shows a method for determining depth information data of an object for a robot;

[0046] Figure 2 shows a data processing system; and Figure 3 shows a robot with a processing system.

[0047] The Figures are schematic only and not true to scale. In principle, identical or like parts, elements and / or steps are provided with identical or like reference numerals in the Figures. P240527W001 - 15 - 11.09.2024

[0048] DETAILED DESCRIPTION OF THE INVENTION

[0049] Figure 1 schematically shows a method 100 for determining depth information data of an object for a robot 1 . In a first step, 102, 102’ image data of an object is obtained. The image data may be obtained by at least two independent imaging sensors 10, 10’ in the environment of the robot. The independent operation of the imaging sensors 10, 10’ may allow them to capture different perspectives or viewpoints of the environment, for example. In a next step 103, 103’ the image data is provided for separate processing to each one of two or more processing components 42, 42’. In a further step 104, 104’ depth information data based on the processed image data is obtained, wherein the depth information data is indicative of the depth of the object in the environment of the robot 1. By using two separate imaging sensors and two independent processing components, two independently working processing channels may be generated. Each channel may process the image data separately to determine the depth of objects in the robot's environment, for example. This redundancy may be important for safety, ensuring that if one processing component 42, 42’ fails, the other may still provide accurate depth information, thereby preventing potential hazards or other risks, for example.

[0050] Figure 2 schematically shows a data processing system 50, which may comprise one or more data processing apparatuses 30, e.g., on board computers, two of which are shown for the purpose of example. The data processing system 50, in particular one or both of the data processing apparatuses 30, in particular their processors 32, may be used to carry out the method 100 for obtaining depth information data as schematically illustrated in Fig. 1. Each one of the exemplary two data processing apparatuses 30 comprises at least one processing unit or processor 32, e.g., a CPU, and at least one computer program product 34, e.g., in the form of a computer-readable storage medium. Computer programs 40 are stored on the computer program products 34. In this example, two computer elements or processing components 42, 42’ of the computer programs 40 are provided within the data processing system 50, which may form parts of the computer programs 40, e.g., different program code or algorithms for different functions or steps of the method 100. Specifically, a computer program 40 of one of the data processing apparatuses 30 may be comprising one processing component 42, which may be in the form of software codes or instructions for the P240527W001 - 16 - 11.09.2024 processors 32, such as a depth estimation algorithm. The other computer program 40 of the other data processing apparatus 30 in the data processing system 50 may comprise the second processing component 42’, which may be in a form of another depth estimation algorithm. The data processing system 50 may be a distributed computing environment, where, for example, the processing component 42 is executed on a local user data processing apparatus 30, such as any stationary or mobile computer, and the processing component 42’ may be executed by a remote or server type of data processing apparatus 30, having increased processing capabilities for processing of the computer elements 44. Alternatively, the processing components 42, 42’ may be all provided as part of the same computer program 40 or executed by the same data processing apparatus 30, which may be part of the robot, for example.

[0051] Figure 3 shows a robot 1 comprising two imaging sensors 10, 10’, a processing system 50 as shown in Fig. 2, and a robot control system 60. Further, the different steps of method 100 and optional method steps 105 and 106 are shown. In this example, the two imaging sensors 10, 10’ obtain image data in a first step 102, 102’, e.g. of an object in an environment of the robot, indicated by the two arrows pointing at the imaging sensors 10, 10’. This image data is in a next step 103, 103’ provided for separate processing to each one of two or more processing components 42, 42’ which are part of the processing system 50. In a next step 104, 104’ based on the processed image data, depth information data is determined, which is indicative of the depth of the object in the environment of the robot 1. By having two independent processing components 42, 42’ to determine depth information data, it may be ensured that a failure in one channel does not compromise the overall safety of the robot 1. If one processing component 42, 42’ fails or provides incorrect data, the other may still function correctly, allowing the robot 1 to continue operating safely. This redundancy may be important in environments where safety is important, such as in industrial settings with human-robot interaction. This redundancy not only may prevent errors, but also may serve to comply with safety standards such as functional safety standards required in industrial applications such as ISO 13849-1 category 3, or IEC 61508, for example. The use of two independent processing component 42, 42’ may allow for cross-verification of depth information data. The processing system 50 may compare the depth information data from both processing component 42, 42’ to identify discrepancies or confirm the P240527W001 - 17 - 11.09.2024 accuracy of the depth information data. This dual verification process may lead to more accurate and reliable depth perception, reducing the likelihood of errors that could lead to accidents or inefficiencies. This separate determined depth information data may e.g. be used for recognizing the object as one of two or more object types. For example, monocular depth estimation techniques may be used to determine differences in the image data e.g. caused by movements of the robot 1 or by small differences in the images (when the sensors are slightly offset) to determine the depth of the objects. Machine learning models that specialise in image processing may use the depth information data and raw image data to recognise and classify objects. For example, distinguishing between a wall and a person may be based on a combination of depth information data, image data and machine learning models. Thus, each processing component 42, 42’ may provide a category of the object mapped to the depth information data, e.g. a person at position A and a wall at position B. Optionally, the depth information data and object categorization obtained from each processing component 42, 42’ may be compared to detect faults affecting one of the processing components 42, 42’ or other hardware components, for example. Optionally, kinematic data of the robot 1 may be provided for separate processing to each one of two or more processing component 42, 42’, to generate depth information data and object categorization information, respectively. By incorporating kinematic data, each processing component 42, 42’ may perform processing the depth information data with greater accuracy, for example. Further, the image data captured by the two imaging sensors 10, 10’ may comprise time stamps, indicative for the time of the imaging sensor 10, 10’ measurement. In an optional next step 105, the depth information data of the object may be provided as an output to the robot 1. The depth information data may be provided to a control system 60 of the robot 1. For example, in an optional next step 106, the robot 1 may use the depth information data to make decisions and take actions such as stopping the robot 1, slowing down or avoiding obstacles, or similar. Depending on the recognised object, different operating modes may be executed by the robot. For example, by recognizing a person, the robot 1 may switch to a safe operating mode designed to minimise the risk of collisions or accidents. In this mode, the robot could severely restrict its movements, for example by travelling more slowly, limiting certain functions (e.g. lifting heavy loads), or by activating additional safety P240527W001 - 18 - 11.09.2024 protocols. If the object is something other than a person, such as a wall, it may be possible to switch to a different mode that performs less safety- re leva nt actions.

[0052] While the invention has been illustrated and described in detail in the drawings and foregoing description, such illustration and description are to be considered illustrative or exemplary and not restrictive. The invention is not limited to the disclosed embodiments. Other variations to the disclosed embodiments can be understood and effected by those skilled in the art and practicing the claimed invention, from a study of the drawings, the disclosure, and the claims.

[0053] As used herein, the word “comprising” does not exclude other elements or steps, and the indefinite article “a” or “an” does not exclude a plurality. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage. Further, as used herein, the phrase “at least one” or similar, e.g., “one or more of”, in reference to a list of one or more entities should be understood to mean at least one entity selected from any one or more of the entities in the list of entities, but not necessarily including at least one of each and every entity specifically listed within the list of entities and not excluding any combinations of entities in the list of entities. This definition also allows that such entities may optionally be present other than the entities specifically identified within the list of entities to which the phrase “at least one” or similar refers, whether related or unrelated to those entities specifically identified. Thus, as a non-limiting example, “at least one of A and B” (or, equivalently, “at least one of A or B” or, equivalently “at least one of A and / or B” or, equivalently “one or more of A and B”, “one or more of A or B”, or “one or more of A and / or B”) may refer, in one example, to at least one, optionally including more than one, A, with no B present (and optionally including entities other than B); in another example, to at least one, optionally including more than one, B, with no A present (and optionally including entities other than A); in yet another example, to at least one, optionally including more than one, A, and at least one, optionally including more than one, B (and optionally including other entities). In other words, the phrases “at least one,” “one or more,” and “and / or” are open-ended expressions that are both conjunctive and disjunctive in operation. For example, each of the expressions “at least one of A, B, and C,” “at least one of A, B, or C,” “one or more of A, B, and C,” P240527W001 - 19 - 11.09.2024

[0054] “one or more of A, B, or C,” and “A, B, and / or C” may mean A alone, B alone, C alone, A and B together, A and C together, B and C together, A, B, and C together, and optionally any of the above in combination with at least one other entity.

[0055] As used herein, the phrase “being indicative of” may for example mean “reflecting” and / or “comprising”. Accordingly, an entity, element and / or step referred to herein as “being indicative of [...]” can be synonymously or interchangeably used herein with one, two or all of said entity, element and / or step “comprising [...]” and said entity, element and / or step “reflecting [...]”. Further, as used herein, phrases such as “based on”, “related” or “relating”, “associated” and similar are not to be seen exclusively in terms of the entities, elements and / or steps to which they are referring, unless otherwise stated. Instead, these phrases are to be understood inclusively, unless otherwise stated, in that, for example, an entity, element or step referring by any of these phrases or similar, e.g., being “based on”, an or another entity, element or step, does not exclude that the respective entity, element or step may be further or also “based on” any other entity, element or step than the one to which it refers.

[0056] The designation of methods and steps as first, second, etc. as provided herein is merely intended to make the methods and their steps referenceable and distinguishable from one another. By no means does the designation of methods and steps constitute a limitation of the scope of this disclosure. For example, when this disclosure describes a third step of a method, a first or second step of the method do not need to be present yet alone be performed before the third step unless they are explicitly referred to as being required per se or before the third step. Moreover, the presentation of methods or steps in a certain order is merely intended to facilitate one example of this disclosure and by no means constitutes a limitation of the scope of this disclosure. Generally, unless no explicitly required order is being mentioned, the methods and steps may be carried out in any feasible order. Specifically, the terms first, second, third or (a), (b), (c) and the like in the description and in the claims are used for distinguishing between similar elements and not necessarily for describing a sequential or chronological order. It is to be understood that the terms so used are interchangeable under appropriate circumstances and that the embodiments of the P240527W001 - 20 - 11.09.2024 invention described herein are capable of operation in other sequences than described or illustrated herein.

[0057] In the context of the present invention any numerical value indicated is typically associated with an interval of accuracy that the person skilled in the art will understand to still ensure the technical effect of the feature in question. As used herein, the deviation from the indicated numerical value is in the range of ± 10%, and preferably of ± 5%. The aforementioned deviation from the indicated numerical interval of ± 10%, and preferably of ± 5% is also indicated by the terms “about” and “approximately” used herein with respect to a numerical value.

[0058] Any reference signs in the claims should not be construed as limiting the scope.

[0059] P240527W001 - 21 - 11.09.2024

[0060] LIST OF REFERENCE SIGNS

[0061] 1 robot

[0062] 10, 10’ imaging sensors

[0063] 30 processing apparatus

[0064] 34, 40 computer program products

[0065] 42, 42’ processing components

[0066] 50 processing system

[0067] 60 control system

[0068] 100 method

[0069] S102 - 106 steps of the method

Claims

P240527W001 - 22 - 11.09.2024Claims:

1. A method (100) for determining depth information data of an object for a robot (1), the method comprising:- obtaining image data of an object in the environment of the robot (1) obtained by at least two imaging sensors (10, 10’), wherein the image data is obtained independently;- providing the image data for separate processing to each one of two or more processing components (42, 42’);- determining depth information data based on the processed image data, wherein the depth information data is indicative of the depth of the object in the environment of the robot (1).

2. The method (100) according to claim 1 , wherein the method further comprises:- recognizing the object as one of two or more object types based on the depth information data, wherein the recognizing as an object type comprises at least one type for person recognition.

3. The method (100) of claim 1 or 2, wherein a control command relating to a safety-related action is provided to the robot (1) based on the depth information data and / or the recognized object type.

4. The method (100) according to claim 3, wherein the method further comprises:- performing a safety-related action when a person is detected.

5. The method (100) of any one of the previous claims, wherein the method further comprises:- identifying similarities and / or differences in the received depth information data of the two or more processing components (42, 42’) and;- providing a control command to the robot (1) based on the identified similarities and / or differences.

6. The method (100) of any one of the previous claims, wherein the depth of the object is determined based on a combination of the depth information dataP240527W001 - 23 - 11.09.2024 and / or based on the recognized object obtained from the two or more processing components (42, 42’).

7. The method (100) of any one of the previous claims, wherein each of the two or more processing components (42, 42’) are based on different machine learning models and / or training data.

8. The method (100) of any one of the previous claims, wherein the method further comprises:- determining a distance between the object and the robot (1) based on the generated depth information data.

9. The method (100) of any one of the preceding claims, wherein a depth map of the object in the environment of the robot (1) is generated for each of the two imaging sensors (10, 10’).

10. The method (100) of any one of the previous claims, wherein the image data by at least two imaging sensors (10, 10’) corresponds to a field of view in an environment of the robot (1).

11. The method (100) according to any one of the preceding claims, wherein additional sensor data comprises kinematic data of the robot (1) is provided for separate processing to each one of two or more processing components (42, 42’), wherein the kinematic data is indicative for one or more of the position, velocity, and acceleration of the robot (1).

12. The method (100) according to any one of the preceding claims, wherein the additional sensor data comprises time stamps, indicative for the time at which the image data was obtained by at least two imaging sensors (10, 10’).

13. Two or more computer program products (34, 40) comprising instructions which, when executed by one or more data processing apparatuses (30), causeP240527W001 - 24 - 11.09.2024 the one or more data processing apparatuses (30) to carry out the method (100) of any one of the previous claims.

14. A data processing system (50) configured to carry out the method (100) of any one of claims 1 to 13.

15. A robot (1) configured to carry out the method (100) of any one of claims 1 to

Citation Information

Patent Citations

  • Image volume for object pose estimation

    US20210158561A1

  • Machine learning techniques for predicting depth information in image data

    US20220292699A1