Identification and response to unknown context in industrial environments
By acquiring environmental information and using trained neural networks to determine action options, combined with auxiliary entity dialogue to optimize operations, the problem of autonomous identification and operation of robots in unknown environments has been solved, improving operational efficiency and safety.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-24
- Publication Date
- 2026-03-27
AI Technical Summary
Existing robots struggle to autonomously identify objects and perform appropriate operations when faced with unknown environments and changing requirements, necessitating human intervention, which leads to complex and costly operations.
By acquiring information about the robot's environment, training neural networks is used to determine action options. When these options are not fully defined, the robot can engage in dialogue with auxiliary entities to obtain more information and optimize its actions to adapt to unknown situations.
It enables robots to operate autonomously in unknown environments with safety and efficiency, reducing human intervention and configuration costs.
Smart Images

Figure CN121743955A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the identification and response to unknown situations in an industrial environment. Background Technology
[0002] Robots are widely used in modern industrial facilities to perform a variety of tasks. These applications range from performing routine tasks to tasks deemed potentially hazardous to human health. Robots can operate manually or execute predefined sequences of actions, and their degree of autonomy can vary.
[0003] A common problem is that robots can typically only recognize objects they already know, and can usually only perform predefined operations on those objects.
[0004] Adapting existing robots to new environments with changing requirements (such as altered operational steps), unknown tasks, and / or the need to manipulate new objects is typically complex. Robots often cannot make decisions autonomously, thus requiring coordination with human operators and / or operating only within defined scopes.
[0005] In other cases, human workers may be required to complete the main part of the task to be performed, and the robot may only be able to assist in solving related tasks (such as removing screws based on the robot and / or classifying objects based on known properties) within a predefined range of operating steps.
[0006] For a robot, it may be necessary to autonomously identify objects (e.g., type, model, etc.) before determining the appropriate action options based on that identification. For example, this might involve the robot determining whether the object is a mobile phone or a calculator in order to disassemble and recycle it. Based on this, the robot can infer the location of each screw on the object and which screws need to be loosened during disassembly. Therefore, based on object recognition, different action options can be executed to achieve optimal disassembly of the object.
[0007] For example, this can be a limiting factor when robots are used for battery recycling. Batteries typically come in different sizes or models and are manufactured by different companies, resulting in a wide variety of battery types. Ideally, the robot must be able to handle these different models to enable a robot-based battery recycling process. Because different battery manufacturers use different assembly methods, several possible and unknown steps arise that must be known to the robot's battery recycling process. Even if the recycling processes for different types of batteries may overlap or share some commonalities, it is still necessary to equip the robot with a variety of motion options to ensure that the corresponding battery recycling can be achieved.
[0008] Programming robots (e.g., through artificial intelligence) to adapt to changing environments is of particular importance for the safe disassembly or recycling of batteries, including previously unknown types. A fundamental problem frequently encountered in practical applications, especially in the ideal scenario where no human operator intervention is required, is how to address the uncertainties inherent in the robot's operational procedures.
[0009] While automation is achievable in very simple applications, robots still encounter obstacles when handling more complex problems. Therefore, human operators are still needed to perform certain production steps or manipulate objects before the robot can execute further operations. In some cases, human operators may need to perform object manipulation-related tasks themselves. To ensure the robot can handle similar problems recurring, appropriate configuration (and / or training) is typically required. This configuration then needs comprehensive testing in test environments and potentially new environments. Subsequently, the robot may need to be shut down and updated. However, this process is often labor-intensive, time-consuming, and involves high robot configuration costs.
[0010] If a disassembly plan is provided for the object to be disassembled (such as a battery), it can be used as a reference during the disassembly process. However, not all objects have such disassembly plans, so these plans can only be used to a limited extent.
[0011] To provide more dismantling options, policies are pushing for battery manufacturers to provide dismantling solutions (i.e., so-called battery passports). While this may make dismantling of batteries for future sale possible, it does not have retroactive effect on batteries already sold, thus failing to improve the current situation where batteries in circulation generally lack dismantling solutions. Furthermore, equipment containing batteries is not included in the mandatory provision of dismantling solutions, even though this equipment will also need to be dismantled and recycled in the future.
[0012] Therefore, current methods for robot manipulation cannot be guaranteed to work in all (desired) application scenarios.
[0013] Therefore, it is necessary to further improve the control of robots, especially in unknown environments. Summary of the Invention
[0014] In this context, the object of the present invention is to enable safe interaction between a robot and an object when there is uncertainty about the options for the action to be performed.
[0015] According to a first aspect, a computer-implemented method is proposed for determining at least one action option for a movable robot component in an environment-specific manner. The computer-implemented method includes: acquiring information related to the environment of the movable robot component, and converting the acquired information into a description of the environment. Furthermore, the method includes: providing the description to a first trained neural network; having the first trained neural network determine at least one action option based on the provided description; and determining whether the provided at least one action option is sufficiently defined for execution by the movable robot component. The computer-implemented method further includes: if it is determined that the provided at least one action option is not sufficiently defined, initiating a dialogue with an auxiliary entity other than the robot component to obtain more scene information, and providing the at least one action option to a controller of the movable robot component.
[0016] The robotic component can be, for example, a robotic arm capable of movement along at least one degree of freedom (such as rotation or linear motion). Alternatively, the movable robotic component can also be a tool component of the robot, capable of performing rotational and / or linear motion. In this case, the movable component can be equipped with a tool that can loosen or tighten screws on the object to be operated by rotation of the movable component. In other cases, the tool can be provided such that it can perform a sawing function when the movable robotic component moves back and forth, for example, when the tool is provided as a saw (component).
[0017] At least one motion option of a movable robot component can be understood as a motion of the movable robot component. This motion can be, for example, rotational motion and / or linear motion.
[0018] The information acquired can be obtained through one or more sensors. Alternatively, the acquired information can also be loaded from a server (e.g., from a database).
[0019] In some cases, descriptions can be provided as text descriptions. A text description can be understood as a description of the acquired information in text form. For example, if the acquired information includes one or more images, it can be provided as a text description, describing the content of the images in text. For example, it could state that the image shows a mobile phone, calculator, or other device.
[0020] Determining whether at least one action option is sufficiently defined can include determining whether the first trained neural network possesses sufficient information to determine at least one action option based on this information, thereby enabling successful or targeted execution of the desired operation on the object. Determining whether at least one action option is sufficiently defined can include determining parameters that indicate whether the definition of at least one action option is sufficient. If the determined parameters exceed a predetermined threshold, at least one action option is considered sufficiently defined.
[0021] This is determined when at least one action option is not sufficiently defined, for example, when it is impossible to clearly determine which object a movable robot part should operate based on the information obtained. Therefore, when sufficient information cannot be provided to the first training neural network about the object to be operated on (e.g., type, model, size, etc.) and / or the specific parts of the object (e.g., which screw (e.g., size, type, etc.)), it can be considered that at least one action option is not sufficiently defined.
[0022] In some cases, a dialogue can be initiated by posing at least one question to an auxiliary entity. This question can be output to the auxiliary entity via a display device (e.g., a display screen, monitor, etc.) and / or a speaker. Alternatively, the dialogue can be initiated via a data connection (e.g., via the Internet and / or an intranet). Preferably, the dialogue can be initiated without a calibration dataset. In some cases, an answer to at least one question can be obtained. The obtained answer can serve as the basis for determining at least one action option (and / or at least one further action option). For example, the answer can be received by the auxiliary entity connected to a neural network (e.g., a first neural network, a second neural network, and / or a third neural network). The associated neural network can, for example, include speech recognition (e.g., a large language model).
[0023] This enhances the ability to handle unknown situations from the perspective of movable robotic components. This allows human operators to engage in dialogue with an auxiliary entity and request further information when situations arise that are unknown to the movable robotic components (e.g., at least one action option is deemed insufficiently defined). This enables the smooth execution and completion of required object operations even in the presence of unknown situations.
[0024] According to one embodiment, the auxiliary entity can be a human operator of a movable robotic component; a human operator remotely connected to a movable robotic component; a human operator remotely connected to an auxiliary robot in the environment of the movable robotic component; and / or can be a second training neural network whose training database is larger than that of the first training neural network.
[0025] A human operator can be physically present, meaning that the human operator can be near the movable robotic parts.
[0026] Remote connections between movable robot components can be achieved, for example, via wireless and / or wired means. Remote connections can be established via the Internet and / or an intranet.
[0027] The assistive robot can be positioned near the movable robotic component. The assistive robot can be equipped with at least one sensor that can be used to acquire more information related to the environment of the movable robotic component.
[0028] Compared to the first neural network, the second training neural network can be trained on more training data, thus enabling it to utilize a larger training database. This helps to provide an improved data foundation through the second neural network, which acts as an auxiliary entity, thereby enabling the provision of at least one action option in a more optimized manner.
[0029] This enables the robot to gather information, with the aid of an auxiliary entity, that was previously insufficient to fully define at least one action option, without interrupting the object manipulation process performed by the movable robotic components and / or having it performed by a human operator. In this way, the determination of at least one action option can be improved, and ultimately, the manipulation process performed by the movable robotic components can be improved.
[0030] According to another implementation, at least one motion option of the movable robot component can be associated with a dismantling plan for electrical equipment and / or batteries.
[0031] In some cases, the object can be electrical equipment and / or a battery.
[0032] A dismantling plan serves as a guide, outlining the sequential steps that can be performed to dismantle electrical equipment and / or batteries, breaking them down into their individual components. This facilitates the recycling of the electrical equipment and / or batteries.
[0033] This method enables efficient and targeted disassembly of electrical equipment and / or batteries.
[0034] According to another embodiment, the described provision is capable of including a dismantling plan that takes into account the provided electrical equipment and / or batteries.
[0035] In some cases, the provided disassembly plan can be converted into a description. In this way, the disassembly process to be performed can be described in text form. This resulting description of the disassembly plan can be combined with, and / or appended to, or spliced into, the description of the acquired information.
[0036] In this way, the determination of at least one action option can be made as close as possible to the analogy of the provided disassembly plan. Overall, this enables an improved disassembly process through moving robotic components.
[0037] According to another implementation, the computer-implemented method can be executed for each disassembly step associated with the disassembly plan.
[0038] In some cases, at least one action option can correspond to a certain disassembly step according to the disassembly plan.
[0039] In this way, at least one optimized motion option can be determined for each disassembly step for the movable robotic component. This optimizes the disassembly process.
[0040] According to another implementation, the first training neural network and / or the second training neural network can include a large language model (LLM), preferably a generative pre-trained transformer (GPT).
[0041] This helps to optimize the initiation of dialogue with auxiliary entities.
[0042] The first training neural network and / or the second training neural network can consist of only an LLM (preferably a GPT), or include other components in addition to an LLM (e.g., another third neural network).
[0043] This enables efficient processing of text information.
[0044] According to another embodiment, the acquisition of information can include acquiring the object type of the object to be operated by the movable robot component and / or acquiring the manufacturer of the object and / or the location of the object in the environment.
[0045] In this description, object type can be understood as the type of the current object that the movable robot part needs to manipulate. Here, type can be understood as, for example, object category, such as the object being a mobile phone, calculator, battery, etc.
[0046] In some cases, "getting" can be understood as obtaining the size of the object to be operated on.
[0047] An object's position in the environment can be understood as its orientation relative to a reference system (such as a reference axis). Information acquisition can include determining by what spatial angles the object has rotated (and / or translated) relative to the reference system (or reference axis).
[0048] By obtaining the object type, it is possible to determine the object that the movable robot part will manipulate. Furthermore, it is possible to infer the object's dimensions (e.g., metric and / or imperial dimensions) and which parts of the object might contain which components (e.g., screws, adhesive points, etc.) from the obtained object type. Determining the object's position can include determining the object's orientation relative to, for example, a tool mounted on the movable robot part.
[0049] According to another implementation, if information acquisition includes acquiring the object type of an object, the computer-based method can further include determining a first uncertainty, which represents the uncertainty in classifying a detected object into a certain object type, and when the determined first uncertainty exceeds a first threshold, a movable robotic component executes at least one determined action option. When the determined first uncertainty is below the first threshold but above a second threshold, a human operator can be requested to confirm whether the determined at least one action option should be executed, and upon receiving confirmation, at least one action option can be executed based on the confirmation. Additionally, if the specific first uncertainty is less than the second threshold, at least one predefined action option can be provided to the human operator, the operator's confirmation regarding the execution of the provided predefined action option can be obtained, and the predefined action option can be executed based on the obtained confirmation.
[0050] In this context, the first uncertainty can be understood as a (numerical) parameter representing the probability of an error when assigning an object to a certain object type; that is, the probability that the assigned object should actually be assigned to a different object type than the one actually assigned. This numerical parameter can vary from 0% to 100%, where a probability of 0% indicates the assignment is considered unreliable or uncertain, while a probability of 100% indicates the assignment is considered highly reliable.
[0051] In some cases, object type acquisition can involve classifying object types based on the acquired information. In this example, classification can be understood as a method by which a computer system automatically assigns objects, data, or contexts to predefined categories or classes based on algorithms and machine learning models. This method can be based on the analysis of features or attributes of the elements to be classified, and the recognition of patterns in the input data. Systems typically derive rules or decision criteria from a large number of already classified examples, enabling them to assign new, unknown instances to appropriate categories. The goal is to achieve the most accurate and reliable assignment possible to automate complex decision-making and recognition tasks, and to generalize to new contexts.
[0052] Object types can be determined based on the classification method described in this article.
[0053] In this way, context-specific execution of at least one determined action option can be performed.
[0054] According to another implementation, the determination of at least one action option can also include determining a second uncertainty associated with the action option.
[0055] Here, the second uncertainty can be understood as a (numerical) parameter indicating whether a particular action option (at least one) can be considered the optimal action option. This numerical parameter can vary from 0% to 100%, where 0% probability indicates that the allocation is considered non-optimal, and 100% probability indicates that the allocation is considered optimal.
[0056] The second uncertainty may arise from the fact that the first trained neural network, when repeatedly providing a description, offers at least one different action option. In some cases, the first trained neural network cannot be directly queried but can only be interacted with via an application programming interface (API). However, if the first neural network can be directly queried, it can provide an output representing the discrete probability distribution of the activation strengths of all tokens (i.e., letters) within the alphabet it is based on. In some cases, this can be used to determine the uncertainty of the first neural network. If the activation strengths are uniformly distributed across different tokens, this can be seen as an indication that the first neural network is uncertain about the allocation of its outputs to their associated inputs.
[0057] This can help improve the descriptiveness of at least one action option that has been identified.
[0058] According to another implementation, the description of the provision and the determination of at least one action option can be achieved from N times, where N≥1, wherein the provision of at least one action option can be realized by repeating the description of the provision and the determination of at least one action option N times.
[0059] The N repetitions provide the ability to provide a description to the first trained neural network N times, thereby providing at least one action option N times accordingly.
[0060] In this way, the determination of at least one action option can be statistically evaluated efficiently, thereby improving the reliability of determining the at least one action option that is considered optimal.
[0061] According to a further embodiment, providing at least one action option may also include: determining the frequency distribution of N determined action options, determining the action option with the highest frequency among the N determined action options, and providing at least one determined action option.
[0062] The frequency distribution can be provided in the form of a histogram, wherein the frequency of a certain at least one action option can be plotted relative to each of the at least one action options.
[0063] Providing at least one action option can include providing at least one action option having the highest frequency according to a determined frequency distribution.
[0064] In this way, the reliability of ultimately providing the at least one action option can be improved.
[0065] According to the second aspect, a computer program product is proposed, which includes instructions that, when executed by a computer, cause the program to perform the methods of any aspect / embodiment described herein.
[0066] A computer program product, such as a computer program medium, can be provided as a storage medium (e.g., memory card, USB flash drive, CD-ROM, DVD) or as a file downloadable from a server on a network. This can be achieved by transmitting a corresponding file containing the computer program product or computer program medium over a wireless communication network.
[0067] According to a third aspect, a computer-implemented apparatus is proposed for determining action options for at least one movable robot component in an environment-specific manner. The computer-implemented apparatus includes: an acquisition unit for acquiring information related to the environment of the movable robot component; a conversion unit for converting the acquired information into a description of the environment; a first providing unit for providing the text description to a first trained neural network; and a first determining unit for determining at least one action option based on the provided description using the first trained neural network. Furthermore, the computer-implemented apparatus includes: a second determining unit for determining whether the provided at least one action option is sufficiently defined for execution by the movable robot component; an initiation unit for initiating a dialogue with an auxiliary entity other than the robot component to obtain more scene information when it is determined that the provided at least one action option is not sufficiently defined; and a second providing unit for providing at least one action option to a controller of the movable robot component.
[0068] Each unit, such as an acquisition unit, a conversion unit, a first determination unit, a second determination unit, a first provisioning unit, a startup unit, and / or a second provisioning unit, can be implemented in hardware and / or software. In hardware implementation, each unit can be implemented as a device or part of a device, such as as a computer, microprocessor, or vehicle control computer. In software implementation, each unit can be implemented as a computer program product, function, routine, program code, or executable object.
[0069] The environment of a movable robotic component can be understood as the factory floor in which it is located. Alternatively, the environment can also include objects that need to be manipulated by the movable robotic component.
[0070] According to one embodiment, the computer-implemented apparatus may further include a first execution unit for performing the computer-implemented method under any aspect / implementation described herein, and / or a second execution unit for performing a computer program product.
[0071] The first execution unit and / or the second execution unit can be, for example, a computer, a processor, a field-programmable gate array (FPGA), or a combination thereof.
[0072] According to the fourth aspect, a system is proposed for determining at least one motion option for a movable robotic component in a specific environment. The system includes the computer-implemented apparatus described herein and the computer program product described herein.
[0073] Although the embodiments described herein are shown separately, they can be combined in any way.
[0074] The embodiments and features described for the proposed apparatus are equally applicable to the proposed method, and vice versa.
[0075] Other possible implementations of the invention include any combination of features or embodiments not explicitly mentioned, but described above or below with reference to the embodiments. In this process, those skilled in the art can also add individual aspects as improvements or supplements to the basic form of the invention. Attached Figure Description
[0076] Other advantageous embodiments and aspects of the invention constitute the dependent claims of the invention and the content of the embodiments described below. The invention will now be further described with reference to the accompanying drawings and preferred embodiments.
[0077] Figure 1 A schematic flowchart is shown for determining at least one motion option for a movable robot component in a specific environment;
[0078] Figure 2 An exemplary application of the GPT-based language model is shown;
[0079] Figure 3 The computer-implemented method is shown;
[0080] Figure 4 A computer-implemented device is shown;
[0081] Figure 5 The system is shown.
[0082] In the figure, components that are the same or have the same function are labeled with the same reference numerals unless otherwise specified. Detailed Implementation
[0083] Figure 1 A schematic flowchart 100 is shown for determining at least one motion option for a movable robot component in a specific environment.
[0084] This schematic process begins by acquiring information related to the environment of the movable robotic component. This information can be acquired through at least one sensor 110. Sensor 110 can be, for example, a camera (true-color camera and / or infrared camera), a microphone, a vibration sensor, a distance sensor (based on light and / or ultrasound), a temperature sensor, etc. Based on at least one sensor 110, the acquired information can therefore include image information, audio information, etc. The acquired information can be a camera image and / or an audio sequence. The acquired information can include information about the object to be manipulated (such as object type and / or object location information), which exists in the environment.
[0085] At least one sensor 110 is capable of one-way or two-way communication 111 with the robot 120. For the communication 111 between the at least one sensor 110 and the robot 120, the information acquired by the sensor 110 can be converted into a description of the corresponding environment (e.g., a text description), such as through an image-to-text generator (e.g., LLM or GPT as described herein). In this way, environmental information acquired by at least one sensor 110 can be described. In this method, for example, an image captured by a camera can be converted into text that a language model can understand using a known neural network model, thereby providing an accurate description of the captured image. It may be necessary to train the base model accordingly for the task (e.g., by passing meta-information to the model to be trained: "There is a battery of model XYZ manufactured by manufacturer ABC at the location of the bounding box ((x, y), (x+delta_x, y+delta_y)) in the image. The surface of the battery is fixed with screws at locations (p1, p2, p3, p4), requiring the use of tool CDE"). Ideally, the camera used to capture relevant images should be able to acquire all relevant surface features. The resulting text or description can then serve as the basis for prompts given to the GPT (as described in this article).
[0086] The information here can include a large number of camera images and / or a large number of audio sequences.
[0087] Information acquired by at least one sensor 110 can be shared with robot 120 via communication. Robot 120 can include at least one movable robotic component and / or communicate with it to control and move that component.
[0088] Robot 120 can convert the description acquired by sensor 110 into an explicit environment description 121 and pass it to artificial intelligence 130. The explicit environment description 121 can include descriptions of environmental features necessary for subsequent manipulation of objects by robot 120 (at least of its movable robotic parts). For example, this can include a description of the form "the housing has screws xy at position (a, b)". This means that the description can include, on the one hand, information about what type of object the manipulated object is (e.g., the housing), and on the other hand, information about which components the object contains that are manipulated by movable robotic parts (e.g., screws "xy"). Furthermore, the description can also include information about the location of these components manipulated by movable robotic parts. Additionally, the explicit environment description can be supplemented with information from disassembly instructions (e.g., for relevant and / or similar models). The resulting input may need to be formatted according to a standardized input format (e.g., "csv" format (comma-separated value format)).
[0089] In some cases, the textual context information provided by communication 111, in collaboration with context description 121, can be understood as a textual description as described herein.
[0090] Artificial intelligence 130 can be provided, for example, as a first neural network. Artificial intelligence 130 can be configured to determine at least one action option 131 based on a description.
[0091] At least one action option 131 can describe the motion of at least one movable part of the robot 120 based on the provided description, which can achieve optimal operation on the object in the current context.
[0092] In some cases, as described herein, the description can be provided to the artificial intelligence 130 N times. In this case, at least one action option 131 can be provided N times.
[0093] The determination of at least one action option 131 can also include determining the uncertainty associated with that action option 131. The uncertainty can be provided as described herein.
[0094] In some cases, artificial intelligence 130 can refer to previous contexts as prompts and / or training data 132.
[0095] A cue word can be understood as a description or input of the generation, processing, and / or optimization of cue words, which serve as the basis for interaction with an artificial intelligence system or language model. This process can include the formulation, structuring, and adjustment of text elements to generate accurate, targeted, and context-sensitive outputs or responses from the artificial intelligence (such as AI 130). Here, the cue word can include input information to AI 130, resulting in the output of at least one action option 131 considered optimal.
[0096] This means that at least one action option 131 can be determined for the current application scenario based on the correspondence between descriptions that have been used in the past and at least one action option identified (i.e., training data).
[0097] The corresponding training data or information associated with the prompt words can be stored in the database 140.
[0098] At least one action option 131 (optionally along with a corresponding uncertainty) can be transmitted as message 122 to robot 120 (and movable robot parts).
[0099] Based on at least one action option 131 transmitted via message 122, robot 120 (or its movable component) can determine whether the provided at least one action option 131 is sufficiently defined for execution by robot 120 or its movable component.
[0100] If it is determined that at least one action option 131 is not adequately defined, the robot 120 (or a movable robotic component) can initiate a dialogue, for example, by sending question 123 to the auxiliary entity 150. Question 123 can be designed specifically to obtain information that adequately defines at least one action option 131 so that the robot 120 (or a movable robotic component) can perform the at least one action option 131.
[0101] The auxiliary entity 150 can be provided as described herein. For example, it can be provided as a human operator located near robot 120 (e.g., within the same building). Alternatively, it can be a person capable of remotely connecting to and interacting with the robot. The human operator can access the input signals of robot 120 (camera, real-time audio sequences) and, if necessary, additional information (e.g., attributes, protocol files, etc.). In some cases, the auxiliary entity 150 can also be a human operator who can remotely connect to another auxiliary robot located near the first robot (i.e., robot 120), which can be configured to perform additional tasks. In some cases, the auxiliary entity 150 can be artificial intelligence (e.g., a second trained neural network), capable of accessing a larger database (e.g., a larger training dataset), or possessing more background knowledge about the objects and materials in the task at hand (e.g., an online GPT service).
[0102] Based on the input of the auxiliary entity 150, it is possible to provide assistance 124 for the decision-making of selecting action options.
[0103] In some cases, the determination of at least one action option 131 may be related to the determination of a (first) uncertainty, which indicates uncertainty in classifying the acquired object into a certain object type. The determined uncertainty can be associated with the uncertainty in performing the classification of the object.
[0104] For example, this uncertainty can be determined by the classifier used. In some cases, the classifier can be configured to compute and provide uncertainty in addition to the classification result (e.g., uncertainty of the model output trained on a negative log-likelihood loss function, or uncertainty derived from the distribution of activation strengths across all possible output classes of the classifier network). The resulting uncertainty can then determine how subsequent processing should be handled.
[0105] Certain uncertainty (expressed as a percentage) can be added to certain security (expressed as a percentage) to reach 100% (or equal to 1). This specifically means that the corresponding security can be derived by subtracting certain uncertainty from 100%, and vice versa.
[0106] If the first determined safety exceeds a first threshold (e.g., 95%), and the object type is not preferentially classified as "unknown," then at least one action option 131 can be performed by a movable robotic component, i.e., the determined at least one action option is operated. In this case, the classification result on which it is based can be considered reliable. In some cases, a database of disassembly schemes for known objects can be utilized. Such disassembly schemes can be provided by the manufacturer of the relevant object. Additionally or alternatively, disassembly plans can be learned through previous disassembly processes or from other sources (e.g., disassembly videos or text disassembly guides from the object manufacturer, online videos and / or disassembly guides published on relevant portals, and / or other suitable sources).
[0107] If the first determined safety level is below a first threshold but above a second threshold (e.g., 75%), a human operator can be requested to confirm whether the at least one action option should be executed, and the at least one action option can be executed based on the confirmation received. In such cases, the classification result can be considered relatively reliable. In some situations, the object classification result may be assigned to one or more object categories. In this case, robot 120 can request confirmation from a human operator and / or auxiliary entity 150 whether the object should be assigned to a specific object category. This confirmation can be made verbally by the human operator and / or auxiliary entity 150 (e.g., "You have correctly identified the situation"). Based on this, robot 120 (or its movable parts) can execute the relevant at least one action option accordingly. In some cases, the generated data (e.g., confirmation of a given situation by a human operator and / or auxiliary entity 150) can be saved and used for retraining when necessary, thereby enabling improved responses in future (similar) situations.
[0108] Alternatively, if the determined first security level is less than the second threshold, at least one predefined action option can be provided to the human operator, their confirmation that the provided predefined action option should be executed can be obtained, and the predefined action option can be executed based on the obtained confirmation. In some cases, multiple possible predefined action options can also be presented to the human operator (such as auxiliary entity 150) for selection. In such cases, the classification result can be considered (very) uncertain, i.e., the probability that an object can be correctly classified into any object category is low for all categories, or the classification result is "unknown". In this case, the obtained information can also be saved and utilized in (similar) future situations.
[0109] As mentioned earlier, auxiliary entity 150 may also verify and confirm the validity of previously unknown action options. This may involve storing the new context 151 in a database (such as database 140).
[0110] Figure 2 An exemplary application 200 of a GPT-based language model is shown. This application 200 is able to address the uncertainty arising from repeated (e.g., N times) feeding of text descriptions to a first neural network.
[0111] In the current situation, input 220 "CONTEXT" can be passed to language model 210 (e.g., GPT language model). Input 220 can be passed to language model 210 N times. Based on this, language model 210 can generate outputs for each input, thus ultimately generating N outputs. A distribution function can be formed based on the outputs obtained in this way.
[0112] Each of at least one action option can be composed of multiple inference steps because the language model 210 can only provide, for example, a distribution of the probability of the next token (e.g., the next letter), which can be obtained, for example, by applying the Softmax function to a log-probability representation. By passing the input "CONTEXT" to the model input N times, m different action options can be obtained, where m ≤ n (due to the standardized output format). This may result in the generation of different letters 240 (here, "R") when "CONTEXT" is repeatedly input to the language model 210 and combined with the obtained response 230 (here, "ANSWER"). Based on this, a frequency distribution 250 can be generated from the resulting letters, where the (relative) frequencies of the different outputs generated by the language model 210 are plotted for their respective outputs (e.g., Q, R, S, T, etc.). This ultimately helps in the determination of at least one action option for each different type.
[0113] Based on this, the potential random uncertainty of action options can be derived from the relative frequency of action options obtained through the reasoning process, for example, by statistically analyzing the frequency (in percentage) of the most frequently occurring action options and comparing it with other possible m-1 options.
[0114] If the model’s output is known (and / or the output that the model is expected to produce is known), then it is also possible to directly observe the distribution of activation intensities generated for all labels.
[0115] In some cases, the responses from language model 210 may become very long (e.g., containing more than 5, 10, 15, 20, 25, 30, 50, 100, or 150 letters), which could cause the number of responses that language model 210 can generate to grow exponentially. By providing a standardized output format, this can be compensated for at least some extent.
[0116] This is helpful for comparing different predicted action options. For example, at least one action option, "Remove the third screw on the left side of the housing," can be semantically considered equivalent to at least one action option, "There are four screws on the housing. Looking from left to right, the third screw should be removed first."
[0117] By using a standardized format such as "disassembly, screw, position (x, y)", syntactic uniformity can also be achieved with at least one action option. This significantly reduces the number of possible outputs. Similarly, the method can also be configured to remove existing filler words from at least one action option (provided in text form), thereby further reducing the space of the possible m action options.
[0118] The relative frequency of each action option can be correlated with safety to determine whether the action option can truly be considered the correct action option. As mentioned earlier, robot 120 (such as...) Figure 1 The aforementioned entity can execute action options, or, in the presence of uncertainty, request more information from auxiliary entities and / or human operators.
[0119] In some cases, according to Figure 2 The results determined by the described method can also be stored in a database as a decomposition plan, which can be associated with a new and previously unknown object type.
[0120] The method discussed herein can be repeated until robot 120 receives notification as an action option and / or by a human operator and / or auxiliary entity that the desired action to be performed by the object has been successfully completed.
[0121] During this iteration, all determined action options and their potential decisions are recorded in a log to ensure that every step of the robot's operation can be traced. This log can be used as additional information, in the form of input (i.e., prompts), or as new training data for retraining.
[0122] In some cases, such as when it can be determined after the fact whether other action options would have been better in a specific context (e.g., in the sense of feedback functionality), retraining can further improve and enhance the artificial intelligence upon which it relies (e.g., ...). Figure 1 The effectiveness and robustness of the aforementioned KI 130.
[0123] In some cases, the decomposition process can be particularly complex. In such cases, it may be beneficial to divide the decomposition process into sub-processes. In some situations, it is possible to break down the object's decomposition process into individual parts.
[0124] In this complex situation, the robot can first detect the relevant environment using at least one sensor (such as a camera) and transmit the acquired information to artificial intelligence (as described in this document). The AI can be configured to perform classification to identify the current task (e.g., a disassembly process). For this purpose, trained neural networks such as YOLO, CNN, etc., can be considered. The disassembly process to be performed can include the identification of object type, manufacturer, and / or object location. Here, similar objects can be grouped into the same object category, and the classifier can determine which category a specific object falls into based on the acquired information related to the environment of the movable robot parts. A specific object can be predefined to belong to a predefined category (e.g., "battery" or "casing" categories). Furthermore, an unknown category can be set for "unknown" objects (e.g., when a processed object only results in a low activation value, i.e., an activation value less than the average activation value associated with objects assigned to other categories besides "unknown"). Unknown objects can be classified (e.g., when a processed object only produces a low activation value (i.e., less than the average activation value of objects already classified as non-"unknown" categories), and then assigned to the "unknown" category. Additionally or alternatively, a "Other Objects" category can also be set. In some cases, the neural network used for classification may have been trained for the "Unknown" and / or "Other Objects" categories.
[0125] Figure 3 A computer-implemented method 300 is shown for determining at least one motion option for a movable robotic component in a specific environment.
[0126] In step 310, information related to the environment of the movable robot component is obtained.
[0127] In step 320, the acquired information is converted into a description of the environment.
[0128] In step 330, this description is provided to the first trained neural network.
[0129] In step 340, the first trained neural network determines at least one action option based on the provided description.
[0130] In step 350, it is determined whether at least one action option provided is sufficiently defined so that it can be executed by a movable robot component.
[0131] In step 360, a dialogue with an auxiliary entity that is different from the robot component is initiated to obtain more scene information if it is determined that at least one action option provided is not adequately defined.
[0132] In step 370, at least one motion option is provided to the controller of the movable robot component.
[0133] Figure 4 An exemplary computer-implemented apparatus 400 is shown for determining at least one action option for a movable robotic component in a specific environment. The computer-implemented apparatus 400 includes an acquisition unit 410, a conversion unit 420, a first providing unit 430, a first determining unit 440, a second determining unit 450, a starting unit 460, and a second providing unit 470.
[0134] The acquisition unit 410 is configured to acquire information related to the environment of the movable robot component.
[0135] The conversion unit 420 is configured to convert the acquired information into an environment description.
[0136] The first providing unit 430 is configured to provide the description to the first trained neural network.
[0137] The first determining unit 440 is configured to determine at least one action option based on the provided description using a first trained neural network.
[0138] The second determining unit 450 is configured to determine whether at least one provided action option is sufficiently defined so that it can be executed by a movable robot component.
[0139] The initiation unit 460 is configured to initiate a dialogue with an auxiliary entity that is different from the robot component to obtain more scene information, provided that it is determined that at least one action option provided is not adequately defined.
[0140] The second providing unit 470 is configured to provide at least one motion option to the controller of the movable robot component.
[0141] Figure 5 A system 500 is shown for determining at least one motion option for a movable robot component in a specific environment.
[0142] System 500 includes a computer-implemented device 510 as described herein. In some cases, the computer-implemented device 510 may be provided in the same manner as the computer-implemented device 400.
[0143] System 500 includes computer program product 520 as described herein.
[0144] Although the invention has been described by way of examples, various modifications are possible.
[0145] Reference number list
[0146] 100 Flowchart
[0147] 110 sensor
[0148] 111Communication
[0149] 120 robot
[0150] 121 Environment Description
[0151] 122 News
[0152] Question 123
[0153] 124 Assistance
[0154] 130 Artificial Intelligence
[0155] 131 Action Options
[0156] 132 prompt words / training data
[0157] 140 database
[0158] 150 auxiliary entities
[0159] 151 Storage of New Contexts
[0160] 200 Applications of GPT-based Language Models
[0161] 210 Language Model
[0162] 220 input
[0163] 230 answers
[0164] 240 output
[0165] 250 frequency distribution
[0166] 300 computer implementation methods
[0167] 310 steps
[0168] 320 steps
[0169] 330 steps
[0170] 340 steps
[0171] 350 steps
[0172] 360 steps
[0173] 370 steps
[0174] 400 computer-implemented devices
[0175] 410 Acquisition Unit
[0176] 420 conversion unit
[0177] 430 First Supply Unit
[0178] 440 First Determined Unit
[0179] 450 Second Determined Unit
[0180] 460 Startup Unit
[0181] 470 Second Providing Unit
[0182] 500 system
[0183] 510 computer-implemented device
[0184] 520 computer program products.
Claims
1. A computer-implemented method (300) for determining at least one motion option for a movable robotic component in a specific environment, comprising: (310) Obtain (310) information related to the environment of the movable robot component; The acquired information is converted (320) into a description of the environment; The description is provided (330) to the first trained neural network; The first trained neural network determines (340) the at least one action option based on the provided description; Determine whether the at least one action option provided by (350) is sufficiently defined to be executed by the movable robot component; If it is determined that the at least one action option provided is not adequately defined, then initiate (360) a dialogue with an auxiliary entity different from the robot component to obtain more scene information; and The at least one action option is provided (370) to the controller of the movable robot component.
2. The computer-implemented method according to claim 1, wherein, The auxiliary entity is: The human operator of the movable robotic component; A human operator capable of remotely connecting to the movable robotic components; A human operator capable of remotely connecting to the movable robotic components in the environment of the assistive robot; and / or The second training neural network is trained using a training database that is larger than the first training neural network.
3. The computer-implemented method according to claim 1 or 2, wherein, The at least one motion option of the movable robot component is associated with a dismantling plan for the electrical equipment and / or battery.
4. The computer-implemented method according to claim 3, wherein, The description includes consideration of dismantling plans for the electrical equipment and / or the battery.
5. The computer-implemented method according to any one of claims 3 or 4, wherein, The computer-implemented method is executed for each step of the disassembly process associated with the disassembly plan.
6. The computer-implemented method according to any one of claims 1 to 5, wherein, The first training neural network and / or the second training neural network include a large language model, i.e., LLM, and preferably include a generative pre-trained transformer, i.e., GPT.
7. The computer-implemented method according to any one of claims 1 to 6, wherein, The acquisition of the information also includes: Obtain the object type of the object to be manipulated by the movable robot component; Obtain the manufacturer of the object; and / or Obtain the pose of the object in the environment.
8. The computer-implemented method according to claim 7, wherein, When acquiring the information includes acquiring the object type of the object, the computer-implemented method further includes determining a first uncertainty, the first uncertainty representing the uncertainty for assigning the acquired object to an object type; and The uncertainty identified in the assessment, among which ,: When the determined uncertainty is higher than the first threshold: The movable robotic component performs the determined at least one action option; When the determined uncertainty is below a first threshold and above a second threshold: A request is made for a human operator to confirm that at least one of the identified action options should be executed. Based on the confirmation, execute at least one action option; or When the determined uncertainty is less than the second threshold: Provide the human operator with at least one predefined action option. Obtain confirmation from the human operator that the provided predefined action options should be executed. Based on the obtained confirmation, the predefined action option is executed.
9. The computer-implemented method according to any one of claims 1 to 8, wherein, The determination of the at least one action option also includes determining the uncertainties associated with the determination of the at least one action option.
10. The computer-implemented method according to any one of claims 1 to 9, wherein, The description and the determination of the at least one action option are repeated N times, where N≥1; as well as The at least one action option is provided based on the description and the determination of the at least one action option through the N repetitions.
11. The computer-implemented method according to claim 10, wherein, The provision of the at least one action option also includes: Determine the frequency distribution of N determined action options; Determine the action option that is selected with the highest frequency among the N determined action options; Provide the determined at least one action option.
12. A computer program product comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 11.
13. A computer-implemented apparatus (400) for determining, in a specific environment, at least one action option for a movable robotic component, the apparatus comprising: Acquisition unit (410), the acquisition unit is used to acquire information related to the environment of the movable robot component; A conversion unit (420) is used to convert the acquired information into a description of the environment; A first providing unit (430) is configured to provide the description to a first trained neural network; A first determining unit (440) is configured to determine the at least one action option based on the provided description using a first trained neural network. A second determining unit (450) is configured to determine whether at least one provided action option is sufficiently defined to be executed by the movable robot component; A startup unit (460) is configured to initiate a dialogue with an auxiliary entity different from the robot component to obtain more scene information if it is determined that the at least one action option provided is not sufficiently defined. as well as A second providing unit (470) is used to provide the at least one motion option to the controller of the movable robot component.
14. The computer-implemented apparatus of claim 13, further comprising: A first execution unit, configured to execute a computer-implemented method according to any one of claims 1 to 11; and / or The second execution unit is configured to execute the computer program product according to claim 12.
15. A system (500) for determining at least one motion option for a movable robotic component in a specific environment, comprising: The computer-implemented apparatus (510) according to claim 13 or 14. as well as The computer program product (520) according to claim 12.