Training of artificial intelligence models

By guiding the robot through a user interface to perform dynamic path and trajectory planning, and combining artificial intelligence technology, the problem of personalized arrangement for fastener removal of aircraft parts was solved, achieving precise fastener operation and efficient operation of the automated system.

CN120897831AInactive Publication Date: 2025-11-04WILDER SYST INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480008952.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-01-24
Filing Date
2024-01-25
Publication Date
2025-11-04
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies make it difficult to achieve personalized placement and efficient removal of fasteners on aircraft parts, leading to the risk of improper drill bit selection or damage to parts. Furthermore, automated systems struggle to generate personalized paths and trajectories for different aircraft.

Method used

By using a user interface to guide the robot in dynamic path and trajectory planning, combined with artificial intelligence technology, multidimensional representations are generated and the attitude data of the target object is updated. Machine learning models are used for target detection and classification, and training data is generated to fine-tune the model, thereby achieving precise operation of aircraft parts.

Benefits of technology

It enables precise removal of fasteners from aircraft parts, reduces the risk of incorrect drill bit selection and part damage, and improves the operational efficiency and accuracy of the automated system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120897831A_ABST
    Figure CN120897831A_ABST
Patent Text Reader

Abstract

Aspects of the present disclosure relate to artificial intelligence-based modeling of a target object, such as an aircraft part. In one example, a system initially trains a machine learning (ML) model based on a composite image generated based on a multi-dimensional representation of a target object. The same system or a different system then further trains the ML model based on actual images generated by a camera in which the robot is positioned relative to the target object. The ML model may be used to process images generated by a camera in which a robot is positioned relative to a target object based on a multi-dimensional representation of the target object. The output of the ML model may indicate location data, target type, and / or visual inspection attributes for the detected target. This output may then be used to update the multi-dimensional representation, which is then used to perform robotic operations on the target object.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications This application claims priority under section 119(e) of Chapter 35 of the United States Code to U.S. Provisional Application No. 63 / 481,576, filed January 25, 2023; U.S. Provisional Application No. 63 / 481,563, filed January 25, 2023; U.S. Non-Provisional Application No. 18 / 447,244, filed August 9, 2023; U.S. Non-Provisional Application No. 18 / 447,230, filed August 9, 2023; and U.S. Non-Provisional Application No. 18 / 421,141, filed January 24, 2024, the contents of which are incorporated herein by reference in their entirety. Technical Field

[0002] This application relates to the field of autonomous robots, and more specifically to an artificial intelligence model for training said autonomous robot using an artificial intelligence model. Background Technology

[0003] Industrial robots represent an expanding field for a wide range of industries looking to improve both their internal and customer-facing processes. Industrial robots can be manufactured and programmed to perform a variety of tasks for different applications. This customizability has prompted many companies to extend the integration of robots from manufacturing to other processes to improve worker safety and efficiency. Attached Figure Description

[0004] Various embodiments according to this disclosure will be described with reference to the accompanying drawings, in which: Figure 1 This is an illustration of a user interface (UI) for guiding a robot according to one or more embodiments.

[0005] Figure 2 This is an illustration of a UI for guiding a robot according to one or more embodiments.

[0006] Figure 3 This is an illustration of a UI for guiding a robot according to one or more embodiments.

[0007] Figure 4 This is an illustration of a UI for guiding a robot according to one or more embodiments.

[0008] Figure 5 This is an illustration of a UI for guiding a robot according to one or more embodiments.

[0009] Figure 6 This is an illustration of a UI for guiding a robot according to one or more embodiments.

[0010] Figure 7This is an illustration of a UI for guiding a robot according to one or more embodiments.

[0011] Figure 8 This is an illustration of an environment for guiding robot operation using a UI, according to one or more embodiments.

[0012] Figure 9 This is a process flow for using a UI to guide a robot, according to one or more embodiments.

[0013] Figure 10 This is a process flow for guiding a robot according to one or more embodiments.

[0014] Figure 11 This is a process flow for guiding a robot according to one or more embodiments.

[0015] Figure 12 This is a process flow for guiding a robot according to one or more embodiments.

[0016] Figure 13 It is an illustration of an environment for modeling a target object using artificial intelligence, according to one or more embodiments.

[0017] Figure 14 This is an illustration of using a machine learning model as part of using artificial intelligence to model a target object, according to one or more embodiments.

[0018] Figure 15 This is an illustration of the estimated pose of a target based on one or more embodiments as part of using artificial intelligence to model the target object.

[0019] Figure 16 It is an illustration of reclassifying targets and updating paths for robot operation according to one or more embodiments.

[0020] Figure 17 It is a process flow according to one or more embodiments for using artificial intelligence to model a target object.

[0021] Figure 18 This is an illustration of an environment for artificial intelligence training according to one or more embodiments.

[0022] Figure 19 This is an illustration of a training method using synthetic image training data and detected target training data according to one or more embodiments.

[0023] Figure 20 It is a process flow for training artificial intelligence according to one or more embodiments.

[0024] Figure 21 It is a block diagram of an example of a computing system according to one or more embodiments. Detailed Implementation

[0025] In the following description, various embodiments will be described. Specific configurations and details are set forth for illustrative purposes in order to provide a thorough understanding of the embodiments. However, those skilled in the art will also understand that these embodiments can be practiced without specific details. Furthermore, well-known features may be omitted or simplified so as not to obscure the described embodiments.

[0026] Aircraft fasteners are mechanical devices used to assemble two or more aircraft parts into components that form an aircraft. When considering the type of fastener and the number of fasteners used for each component, engineers take into account various factors such as shear stress and aircraft loads. Fasteners can differ from one another based on diameter, material, and shape. Fasteners are inserted into holes arranged through the aircraft parts. A single wing may include thousands of holes, and an entire aircraft may include hundreds of thousands. As an aircraft ages or becomes damaged, damaged or worn parts need to be removed and replaced with new ones. In some cases, drilling tools are used to remove these aircraft fasteners from worn or damaged aircraft parts.

[0027] Drilling tools are selected based on fastener type, and the drill trajectory can be based on the orientation of the fastener and hole relative to the aircraft part. If an incorrect drill bit is selected or the drill bit is not properly inserted into the fastener, the fastener may not be removed and the aircraft part may be damaged. One issue related to fastener removal is that fasteners are not evenly distributed across different aircraft, even those of the same model. A first aircraft part can be mounted on an aircraft using a first set of fasteners. The same type of aircraft part can be mounted on another aircraft of the same model, but using a second set of fasteners with an orientation different from the first set. Since some aircraft parts include multiple types of fasteners, the individualized arrangement of fasteners for each aircraft part becomes amplified. Therefore, when inserting a drill bit into a fastener, the individual orientation and type of the fastener must be considered to remove it correctly. This problem is further complicated by the fact that the orientation of one set of fasteners in one aircraft part cannot be used as a template for another set of fasteners in a similar aircraft part. This could be a problem for automated systems designed for aircraft fastener removal, as the system needs to generate paths and trajectories to reach each fastener on an individual basis in order to properly remove the aircraft fasteners.

[0028] The embodiments described herein address the aforementioned problems via a system that guides a robot through dynamic path planning and trajectory planning. A user interface (UI) can be used to guide the robot to receive point cloud data representing a target object. The point cloud data can be segmented to associate point cloud data points with different objects within the point cloud data. The system can extract features from the segmented objects to identify the target object, such as an aircraft panel. The system can further identify real-world targets within the target object, such as holes or fastener heads. The user can further use the UI to guide the system to estimate the target's 3D position in space and calculate normals extending away from the target surface. The system can use the 3D position and normals to guide an end effector (such as a drill) into the real-world target. The user can further use the UI for registration to obtain a transformation matrix for transforming the target object's working coordinate system to the robot's reference coordinate system. The user can further use the UI to guide the system to generate a set of virtual points and use these virtual points to calculate a path over the target object's surface. Based on the calculated path, the user can use the UI to guide the system to calculate a trajectory traversing the calculated path. The trajectory can include a set of movements for each joint of the robot to traverse the path while avoiding collisions and taking into account the robot's limitations.

[0029] By using a UI, a multidimensional representation of a target object (e.g., its three-dimensional model) becomes possible. This multidimensional representation can indicate the pose (e.g., position and orientation) of each target object in the object's working coordinate system (e.g., a locally defined coordinate system relative to a reference point included in the object).

[0030] Embodiments of this disclosure further relate to updating multidimensional representations to enable more precise robotic manipulation of target objects (e.g., fastener removal operations applied to aircraft parts). This update may rely on artificial intelligence (AI) techniques. For example, the multidimensional representation can be used for path and trajectory planning to control the robot. The robot may include a camera (e.g., a two-dimensional camera) set as the tool center point (TCP). Given a trajectory, the robot can be controlled to position the camera at a predefined distance relative to the target object (e.g., a fastener on an aircraft part). During positioning, the camera may generate an image showing at least a portion of the target object, where that portion includes the target. The input to a machine learning (ML) model (such as a computer vision neural network) may be generated based on the image. The ML model may generate an output based on the input. This output may indicate target data detected from the image. The output may also indicate the target's classification (e.g., the target has a target type) and / or the target's visual inspection attributes (e.g., attributes of the fastener to be determined by visual inspection of the fastener, such as the level of corrosion of the fastener). If the position data is two-dimensional (e.g., along the XY plane), it can be augmented to three-dimensional based on a predefined distance. If the target is already three-dimensional (e.g., based on a depth sensor that may or may not be part of a camera), such enhancements may not be necessary. Furthermore, if no depth data is available, the camera orientation can be determined based on TCP controls, thereby normalizing the camera to the target. The camera orientation can be set to the target orientation. If depth data is available, the target orientation can be derived from the depth data. In both cases, the target's position and orientation in space are determined and correspond to the target's pose in a coordinate system (e.g., the robot's reference coordinate system, such as a locally defined coordinate system relative to a reference point included in the robot). This pose can be transformed into the actual pose defined in the working coordinate system based on inverse and forward kinematic transformations. The actual pose can then be stored in an updated multidimensional representation of the target object. This multidimensional representation can also be enhanced to include the target's classification and visual inspection attributes. Subsequently, a set of robot operations (e.g., to remove and replace corroded fasteners or fasteners with a specific classification) can be planned based on the updated multidimensional representation of the target object. Path generation and trajectory planning can be used for this purpose.

[0031] Furthermore, embodiments of this disclosure can fine-tune the ML model by generating training data from actual detections of targets across one or more target objects. For illustration, consider an example showing an image of a target. This image can be processed as described above to determine the target's pose data, classification, and / or visual inspection attributes. A human-machine loopback machine learning approach can be implemented, whereby inputs that correctly indicate corrected pose data, corrected classification, and / or corrected visual inspection attributes can be received. The corrected pose data, corrected classification, and / or corrected visual inspection attributes, along with the image, can be used to generate additional training data. This additional training data can be included in a training dataset, and this dataset can be used to further train the machine learning model.

[0032] The embodiments of this disclosure are described in conjunction with aircraft part and fastener removal operations. However, the embodiments are not limited thereto, but are similarly applicable to any other type of target object and / or any operation relying on a multidimensional model. Target objects may include automotive parts, marine parts, rocket parts, and / or other types of parts with simple or complex geometries. Target objects may also include entire automobiles, ships, and / or rockets. In addition to or as an alternative to fastener removal, operations may include fastener addition, fastener hole inspection, fastener hole sizing or rework, surface treatments (e.g., for etching), drilling, stamping, etching, deburring, welding, spot welding, gluing, stitching, sealing, painting, paint stripping, cleaning, etc.

[0033] Figure 1This is a user interface (UI) 100 for an application used to guide a robot, according to one or more embodiments. UI 100 is the interaction point between the user and the computing elements of the system. The computing elements can be a collection of hardware elements (such as user computing devices and / or bare-metal servers) and software elements (such as applications) for path generation, trajectory generation, and robot guidance. UI 100 can be a collection of interface elements, including input controls (e.g., buttons, text fields, checkboxes, dropdown lists), navigation components (e.g., sliders, search fields, markers, icons), information components (progress bars, notifications, message boxes), and containers (e.g., accordions). As shown, UI 100 is a multi-pane UI, where multiple pages can be contained within a single window 102. Each page can be accessed via a corresponding page tab. UI 100 may include an action pane 104, which includes a collection of input controls, navigation components, and information components. UI 100 may further include a visualization pane 106 for providing visualizations, such as point cloud segmentation, path generation, and trajectory generation. Action pane 104 may be adjacent to visualization pane 106, and both may be displayed on a single window 102. The UI may include one or more additional panes, such as the display selection pane described below.

[0034] Action pane 104 may include navigation pane 108, which includes tabs for navigating to the corresponding pages of UI 100. As shown, UI 100 may include a feature identifier (feature ID) tab 110 for navigating to the feature identifier page, a registration tab 112 for navigating to the registration page, a motion planning tab 114 for navigating to the motion planning page, and a perception tab 116 for navigating to the perception page.

[0035] exist Figure 1 In the UI 100, the user has selected the Feature Identification tab 110, and the Feature Identification page is thus displayed. The user can use the Feature Identification page 118 to select input controls for segmenting the point cloud, associating the point cloud data with objects included in the target object, and calculating the surface of the target object and / or the normals from the surface of the target object. The user can store a 3D representation of the target object (e.g., a point cloud or computer-aided design (CAD) file) in the computing system for later display as a visualization 120. The user can click the Load Point Cloud button 122 and access the computing system's file system. In response to receiving a selection of the Load Point Cloud button 122, the UI 100 can display the file system (e.g., files and folders), and the user can select the 3D representation to retrieve for the application.

[0036] The application can retrieve a 3D representation, which can be displayed as visualization 120 on visualization pane 106. As shown, point cloud data represented as visualization 120 is displayed on visualization pane 106. In this case, visualization 120 is a real-world panel with an array of holes. In practical applications, visualization 120 can be any real-world part or component (aircraft part, vehicle part, structural part). In other cases, visualization 120 can be any object on which paths and trajectories are to be generated. Visualization 120 can be a 3D representation that can be manipulated by a user through UI 100. The user can manipulate the visualization to change its position relative to visualization 120. For example, the user can use a peripheral device such as a keyboard to input commands, or use a mouse to click a cursor on visualization plane 106 to rotate, flip, zoom in, or zoom out of visualization 120. In response to input signals from the peripheral device, UI 100 can change the 3D position and orientation of visualization 120.

[0037] Users can further utilize UI 100 to participate in the segmentation process for associating point cloud data points with objects in visualization 120. UI 100 may include a first set of fields 124 for inputting segmentation parameters. Users can manually select fields and type target values ​​for the parameters. Segmentation parameters may include a tolerance, which may include the maximum distance between point cloud data points that will be considered part of a cluster. Segmentation parameters may include a minimum size, which may be the minimum number of data points required to form a cluster. Segmentation parameters may further include a maximum cluster size, which may be the maximum number of data points required to form a cluster. In some cases, certain objects may have an estimated number of data points. Therefore, users can use segmentation parameters to select a minimum cluster size lower than the estimated cluster size and a maximum cluster size higher than the estimated cluster size. In this sense, users can use UI 100 to filter out objects with an estimated number of data points lower than the minimum cluster size and objects with an estimated number of data points higher than the maximum cluster size. Users can then click the Generate Cluster button 126 to initiate the segmentation phase. In response to a single click of the Generate Cluster button 126, the application can receive each user-defined parameter value from the user-defined parameter values ​​and segment the 3D representation accordingly. The segmentation phase may include executing a distance measurement function (e.g., an Euclidean distance function) from a library and determining the distance between data points. The application can select data points that are within a threshold distance from each other to be considered clusters. The segmentation phase may further include identifying clusters based on a minimum cluster size and a maximum cluster size. As used herein, the identified clusters can be considered objects. For example, such as... Figure 1 As shown, the object may include a hole or side of the visible 120.

[0038] UI 100 may further allow the user to perform an identification phase after the segmentation phase, in which the application can determine whether the identified clusters correspond to real-world targets (e.g., holes, fasteners). UI 100 may include input control buttons for performing an object recognition model. The model may be a rule-based model or a machine learning model. The model may be further configured to extract features from the clusters as input and predict a category for each cluster. As described above, the embodiments herein are illustrated with respect to real-world targets such as holes or fasteners. Therefore, the model may be a circle fitting model configured to fit circles to the identified clusters. Feature identification page 118 of UI 100 may include a second set of fields 128 for inputting model parameters. Fields may include, for example, a maximum diameter field and a tolerance field. The user can input target values ​​for each field in the second set of fields. For example, the user can select a value for the maximum diameter size of the clusters that will be considered as holes or fasteners. Feature identification page 118 may further include a Fit Model button 130 for initiating model fitting. In response to receiving an input signal from the Fit Model button 130, the application can initiate an object recognition model (e.g., a circle fitting algorithm). The object recognition model can receive clusters and user-based parameters as input from the second set of fields and fit a predefined model (e.g., a circle) to the clusters. Clusters that fit the model and are within the user-based parameters of the second set of fields 128 are considered real-world objects. The application can modify the visualization 120 corresponding to the clusters of real-world objects in the visualization pane 106 (e.g., change the color, change the texture, add 3D structures, or otherwise visually label them). Therefore, a user viewing visualization 120 can visually identify the objects (e.g., real-world objects).

[0039] Feature identification page 118 can further allow the user to initiate the identification of the relative position of each target on visualization 120. Feature identification page 118 may include a calculate feature position button 132 for initiating the application to execute an algorithm to determine the relative position of each target. It should be understood that the names and functions of the input control buttons are not limiting. Each element in UI 100 can perform the general functions described, regardless of its name. For example, the load point cloud button 122 allows the user to select the point cloud to retrieve for the application. If the application's task is to locate targets in different 3D representations, such as computer-aided design (CAD) , UI 100 may include a load CAD button. Once the application has calculated the relative positions of the targets, UI 100 can allow the user to calculate the corresponding normals for each cluster corresponding to real-world targets. The calculation and visualization of the normals will be combined Figure 2The action pane 104 may further include a log window 134 for displaying a log of actions performed by the application.

[0040] Figure 2 This is an illustration of a UI 200 for guiding a robot according to one or more embodiments. For illustrative purposes, only the action pane and visualization pane of the UI 200 are shown. The feature identification page 202 of the UI 200 may include a third set of fields 204 for inputting parameters used to calculate corresponding normals. A normal is a line orthogonal to a cluster of surfaces corresponding to a real-world target. Normals can be used to guide a robot end effector (such as a drill) into or away from a target. The third set of fields 204 may include, for example, a search radius field, x, y, and z viewpoint fields, and a tool rotation field. The user can input target values ​​into each field of the third set of fields 204 via the UI 100. The feature identification page 202 may include a calculate feature orientation button 206. The user can click the calculate feature orientation button 206, and the application can initiate a normal calculation phase in response. The feature identification page 202 may further include a pose estimation phase to define a set of waypoints (e.g., virtual points) on the surface of the target object. Waypoints can correspond to each cluster of real-world objects (e.g., holes, fasteners, or locations for drilling), and the number of waypoints can be greater than or less than the number of clusters corresponding to real-world objects. The application can assume no coplanar waypoints and calculate normals for each cluster corresponding to real-world objects. The orientation of each normal relative to the surface of the target object can be based on the object identified in the point cloud data. Thus, a normal can be perfectly orthogonal to the surface of the target object, while adjacent normals can be slightly angled and not orthogonal to the surface. As described above, holes can be drilled and fasteners inserted manually. Therefore, the direction in which the hole or fastener enters the target object can be different from the orientation of the target object. Therefore, a single click on the calculate feature orientation button 206 can initiate a normal calculation phase relative to each waypoint for correct normal calculation.

[0041] The application can further specify the viewpoint, or specify whether the normal is calculated in a direction toward or away from the target. This can be calculated based on user-defined parameters input into the third set of fields 204. The application can further specify the orientation of the cluster corresponding to the real-world target. The motion of the robot's end effector can have more degrees of freedom than are sufficient to locate the target in three-dimensional space. Therefore, in order to calculate the movement of the robotic arm and end effector in three-dimensional space, the orientation of the target may have to be taken into account to correctly approach and move away from the real-world target.

[0042] UI 200 can further provide visualization of the calculated normals on visualization pane 210. Feature identification page 202 includes a Show Feature Attitude button 208 (e.g., an input control button). A user can click the Show Feature Attitude button 208, and in response, the application can display the calculated normals on visualization pane 210. The visualization pane 210 of UI 200 displays corresponding waypoints 212 located at targets in a cluster array corresponding to real-world targets. Waypoints are virtual points, where the number of waypoints can be less than, greater than, or the same as the number of targets. Waypoints 212 include normals 214 with directional arrows indicating the direction of normals 214 extending away from waypoints 212. As shown, each cluster corresponding to a real-world target includes normals extending away from waypoints 216 at the target. This visualization pane feature allows the user to verify that the calculated normals extend correctly toward or away from each target before launching the actual robot.

[0043] In addition to providing a visualization of the normal (which can be considered as a z-direction vector), the UI can also provide x-direction and y-direction vectors. Figure 3 This is an illustration of a UI 300 for guiding a robot according to one or more embodiments. For illustrative purposes, only the action pane and visualization pane of the UI 300 are shown. The visualization pane 302 of the UI 300 may display a visualization 304 including an array of targets. Clusters corresponding to real-world targets 306 (e.g., holes, fasteners, locations for drilling) may include x-direction vector visualizations 308 and y-direction vector visualizations 310. The x-direction vector visualization 308 may be perpendicular to the y-direction vector visualization 310. Each x-direction vector visualization in the x-direction vector visualization 308 may be perpendicular to the y-direction vector visualization 310 and may be perpendicular to the normal. These vector visualizations allow the user to verify that the orientation of real-world objects relative to the robot's end effector is correct before starting the robot.

[0044] Figure 4 This is an illustration of a UI 400 for guiding a robot according to one or more embodiments. Figure 4The registration page 402 of UI 400 is shown. Users can access registration page 402 by moving the cursor to the navigation pane and clicking the registration tab. Registration page 402 can initially be positioned below another page (e.g., the feature identification page). Users can click the registration tab and display registration page 402 above other pages in UI 400. Users can use registration page 402 to enable the application to perform a registration phase. During the registration phase, the application can position and orient a real-world object relative to the robot. Specifically, the registration phase can calculate a transfer function to transform the working coordinate system (e.g., the target object's coordinate system) into a reference coordinate system (the robot's base coordinate system). The registration page may include a "Generate Coordinate System from Selected Pose" button 404. "Generate Coordinate System from Selected Pose" button 404 can be used to initiate the application to calculate a reference coordinate system for the real-world object. In the real world, the target object can be oriented in three-dimensional space. For example, if the real-world object is an aircraft panel, the panel can be positioned on the fuselage, wing, or other location on the aircraft. Therefore, UI 400 can provide users with UI elements to calculate a reference coordinate system for the real-world objects that will be used to calculate the transfer function.

[0045] Applications can use various methods to generate a working coordinate system. For example, applications can use methods such as the common-point method or the reference marking method. For instance, a user can use UI 400 to select multiple clusters representing real-world targets (e.g., three targets). UI 400 can further display visual markers for the targets on its visualization pane 406. For example, each selected cluster corresponding to a real-world target can be displayed with a unique color different from each of the other targets. The user can move the robot's end effector to each target (e.g., a hole, fastener / drilling location) to orient the end effector's tool center point (TCP) to each target. The application can then calculate the distance measurements to the target objects to use for calculating a transfer function. The application can further use the transfer function to transform the working coordinate system of the target objects into the robot's reference coordinate system.

[0046] The registration page 402 may further include a registration workpiece button 408 for launching the application to calculate the transfer function between the robot's reference coordinate system and the real-world object. The UI 400 may further display the reference coordinate systems of the robot 410 and the real-world object 412. The display presented on the visualization pane 406 can be user-defined through interaction with the display selection pane 414. The display selection pane 414 may include a window 416 that can include a set of elements that the user can select. If there are too many elements to display together, the window 416 may include scrolling up and down to view the functionality of each element. As shown, the user has selected the base-linked coordinate system displayed on the visualization pane 406. This UI functionality can provide the user with a guarantee that the transfer function is associated with the correct reference object. The display selection pane 414 may further include a raw camera image window 418 for presenting video. For example, if the robot includes an image capture device, the raw camera image window 418 can display a live stream of images captured by that device.

[0047] Figure 5 This is an illustration of a UI 500 for guiding a robot according to one or more embodiments. Figure 5 The UI 500's exercise plan page 502 is displayed. Users can access the exercise plan page 502 by moving the cursor to the navigation pane and clicking the exercise plan tab. The exercise plan page 502 can initially be positioned below another page (e.g., the feature identification page and the registration page). Users can click the exercise plan tab and make the UI 500 display the exercise plan page 502 above other pages.

[0048] The motion plan page 502 may include a "Check Reachability" button 504 for launching the application to determine the reachability of each target. The user can click the "Check Reachability" button 504, and in response, the application can determine the reachability of each target. The "Check Reachability" button 504 allows the application to coordinate the robot's dimensions relative to the size of a real-world object and retrieve data related to potential collisions between the robot and the real-world target object. For example, the "Check Reachability" button 504 allows the application to retrieve data related to an expected collision between the robot and the real-world target object. The application can then use this data to further determine whether the robot is likely to collide with another object.

[0049] The motion planning page 502 may further include a calculate path button 506 for calculating the path over the target. The path can be a series of directions that traverse the target on a real-world object. A visualization pane 508 can display a visualization 510 of the real-world target object. The visualization 510 may include an initial target 512, which represents the first target to be traversed over the path. The visualization 510 may further include direction markers 514, which show the direction from one target to another over the path. In this sense, the user can visually inspect the path before starting the robot. If the user wishes to modify the path, the user can do so via UI 500. The user can engage UI 500, for example, via a peripheral device, and modify one or more targets (e.g., move the visualization of the target in 3D space, delete waypoints). As shown, a waypoint has been removed, and the visualization pane 508 can display the absence of the target 516. In response to the modification of the waypoint via UI 500, the application can recalculate the path. Figure 5 As shown, the waypoint has been removed, and there is no direction marker 514 pointing to the absence of the waypoint 516.

[0050] The motion planning page 502 may further include a "Plan Trajectory" button 518 for launching the application to generate a trajectory. A path may include a series of movements over the surface of a real-world object, without regard to the individual movements of robot parts. A trajectory may include a series of movements that include the movements of robot parts (e.g., the movements of joints in a robotic arm). The user can click the "Plan Trajectory" button 518, and in response, the application can calculate a trajectory for the robot. For example, the application may employ inverse kinematics functions to calculate the movements for each part of the robot (e.g., the upper arm, elbow, lower arm, wrist, and end effector).

[0051] The motion plan page 502 may further include a DB Export button for exporting paths and trajectories to a database. Paths and trajectories can identify the type of robot, a real-world object. For example, if the real-world object is the wing of an aircraft, the path and trajectory can further identify the individual aircraft including the wing (e.g., using identifiers such as tail numbers). In this sense, paths and trajectories can be stored in a database for future use. Thus, if the aircraft needs re-maintenance, personalized paths and trajectories for the aircraft can be retrieved. The motion plan page 502 may further include a Database (DB) Import button 522. A user can click the Import Paths from DB button 522 and launch the application to retrieve previously stored paths and trajectories from the database for the robot to execute. The motion plan page 502 may further include a Create Job from Path button 524 for launching the application to create robot jobs, where a robot job can be a file with one or more parts configured to be executed by the robot. A user can click the Create Job from Path button, and in response, the application can configure the file to be executed by the robot based on the paths and trajectories. The exercise plan page 502 may further include an "Upload Job to Robot" button 526 for uploading robot jobs to the robot's controller. The user can click the "Upload Job to Robot" button 526, and in response, the application can upload the robot job to the robot controller. The exercise plan page 502 may include a "Run Job" button 528 for starting the robot to perform robot jobs. The user can click the "Run Job" button 528, and in response, the application can make the robot perform robot jobs.

[0052] The motion plan page 502 may further include various input control buttons 530 for real-time control of the robot. These input control buttons may include a move to main pose button, a move to zero pose button, a pause motion button, a resume motion button, and a stop motion task button. The user can click any of these buttons, and in response, the application can initiate responsive motion of the robot. For example, if the user clicks the pause motion button, the robot can responsively pause its motion.

[0053] Figure 6This is an illustration of a UI 600 for guiding a robot according to one or more embodiments. As shown, a visualization 602 is displayed on a visualization pane 604. The visualization 602 displays a path 606 for traversing a target on a real-world object. The UI 600 may further include an inference image window 608 that can display visual markers in real time each time the robot detects a target. As shown, a first visual marker 610 is displayed next to a second visual marker 612. In practice, the robot may traverse a real-world object using a calculated trajectory. The visual markers appear when the robot's end effector passes by and detects a target. Therefore, when the end effector passes by the first target, the first visual marker 610 may appear on the inference image window 608 at a first time t0; and when the end effector passes by the second target, the second visual marker 612 may appear at a second time t1. Thus, the user can see in real time that the robot is detecting targets.

[0054] Figure 7 This is an illustration of a UI 700 for guiding a robot according to one or more embodiments. The perception page 702 of the UI 700 may include a start camera button 704 and a restart camera button 706. Additionally, the user can input camera parameters, such as camera surface offset 708. The start camera button 704 can be used to start the camera. The restart camera button can be used to restart the camera. As described above, the display selection pane may further include a raw camera image window for presenting video. For example, if the robot includes an image capture device, the raw camera image window may display a live stream of images captured by that device. The live image may correspond to a visualization displayed on the UI 700. The perception page 702 may further include a start feature detection service button 710. As described above, the application can detect features (e.g., holes or fasteners) on the surface of a target object. The start feature detection service button 710 can initiate a service for detecting features on the surface of the target object.

[0055] The perception page 702 may further include a Run Artificial Intelligence (AI) Detection button 712 for executing AI feature identification algorithms. The application can receive image data of the target object. For example, the application can receive captured images from a camera activated by the Start Camera button 704. In other embodiments, the application can receive data from other sensors, such as thermal sensors or spectral sensors. The application can receive the data and use it as input to a neural network, such as a convolutional neural network (CNN). The neural network can receive the data and predict the category for each target. The perception page may further include a Detect Reference Marker button 714. As described above, the application can use various methods to generate a working coordinate system. The application can use, for example, the concurrent point method or the reference marker method. The Detect Reference Marker button 714 can initialize the application to instruct the robot to detect the reference marker.

[0056] Figure 8 This is an illustration 800 of an example of an environment for guiding robot operations using a UI, according to one or more embodiments. Environment 800 may include a robot 802, a server 804, and a user computing device 806 for accessing the UI 808. The robot 802 may operatively communicate with the server 804, which in turn may operatively communicate with the user computing device 806. The server 804 may be located in the same location as the robot 802, or it may be located remotely from the robot 802. A user may engage with the UI 808 to transmit messages to and from the server, which in turn may transmit and receive messages from the robot 802. The robot 802, server 804, and user computing device 806 may be configured to allow the robot to perform autonomous operations on an aircraft 810 and / or its parts (e.g., wings and fuselage) or another real-world object. Autonomous operations may include, for example, removing or installing fasteners for the aircraft 810, cleaning the aircraft 810, and painting the aircraft 810. It should be understood that, as shown in the figure, the user computing device 806 can communicate with the robot 802 via the server 804. In other embodiments, the user computing device 806 can communicate directly with the robot 804, and further, the user computing device 806 performs the following functions of the server 804.

[0057] Robot 802 can navigate to the operating area and, once there, perform a set of operations to register aircraft 810 (or a portion thereof) for subsequent operation on aircraft 810 (and / or aircraft parts). Some of these operations may be computationally expensive (e.g., feature recognition, registration, path and trajectory generation), while others may be less computationally expensive and more latency-sensitive (e.g., drilling fastener holes). Computationally expensive operations can be offloaded to server 804, while the remaining operations can be performed locally by robot 802. Once the different operations are completed, robot 802 can autonomously return to the parking lot or be summoned to another operating area.

[0058] In one example, robot 802 may include a movable base, a power system, a powertrain system, a navigation system, a sensor system, a robotic arm, an end effector, input and output (I / O) interfaces, and a computer system. An end effector that can support a specific autonomous operation (e.g., drilling) may be a line-replaceable unit with a standard interface, allowing the end effector to be replaced with another unit supporting a different autonomous operation (e.g., sealing). End effector replacement may be performed by robot 802 itself or via a manual process where an operator can perform the replacement. The I / O interface may include a communication interface to communicate with server 804 and user computing device 806 based on input selection at UI 808 for selecting the autonomous operation to be performed by robot 802. The computer system may include one or more processors and one or more memories storing instructions that, when executed by the one or more processors, configure robot 802 to perform different operations. These instructions may correspond to program code for navigation, control of the power system, control of the powertrain system, collection and processing of sensor data, control of the robotic arm, control of the end effector, and / or communication.

[0059] Robot 802 may include a Light Detection and Ranging (LiDAR) sensor to emit pulsed lasers toward aircraft 810. The LiDAR sensor may be further configured to collect reflected signals from aircraft 810. The LiDAR sensor can determine the reflection angle of the reflected signal and the elapsed time (time of flight) between the transmission and reception of the reflected signal to determine the position of a reflecting point on the surface of aircraft 810 relative to the LiDAR sensor. Robot 802 can continuously emit laser pulses and collect reflected signals. Robot 802 can transmit sensor data to server 804, which can further generate a point cloud of aircraft 810. The point cloud can be displayed on UI 808 via a visualization pane. In other cases, server 804 can retrieve stored point clouds of aircraft 810 from a database.

[0060] Server 804 may be a hardware computer system including one or more I / O interfaces for communicating with robot 802 and user computing device 806, one or more processors, and one or more memories storing instructions that, when executed by the one or more processors, configure server 804 to perform various operations. The instructions may correspond to program code for communication and for processes executed locally on server 804 for robot 802 given data sent by robot 802.

[0061] User computing device 806 may be a hardware computer system including one or more I / O interfaces for communicating with server 804 and robot 802. The user computing device may include a display for displaying UI 808. The user computing device may further include one or more input interfaces (e.g., mouse, keyboard, touchscreen) for receiving user input from UI 808.

[0062] In response to one or more inputs at UI 808, multiple operations may need to be performed, and these operations may be interdependent. For example, to identify a fastener and / or drill a fastener hole in aircraft 810, robot 802 may detect a target (e.g., a hole), register aircraft 810 so that it is located in robot 802's local coordinate system, control the robotic arm to move to a position according to a specific trajectory, and control the end effector to drill the hole. Some of these operations may be computationally expensive and performed less frequently (e.g., generating a Simultaneous Localization and Mapping (SLAM) map, registration), while others may be computationally less expensive but latency-sensitive and performed more frequently (e.g., controlling the robotic arm and end effector). Therefore, server 804 may perform the procedures for the computationally expensive / less frequent operations, while robot 802 may perform the procedures for the remaining operations locally.

[0063] In some cases, a user can engage UI 808 and select one or more inputs for controlling robot 802. User computing device 806 can receive the inputs and transmit messages including the inputs to server 804. The server may include an application for performing one or more operations to control robot 802. Server 804 can receive messages from robot 802 and transmit messages to the robot. Messages may include operational instructions generated based on inputs to UI 808. For local operations, robot 802 can execute corresponding processes locally and can notify server 804 of the results of these local operations (e.g., drilling fastener holes at specific locations on aircraft 810). Server 804 can transmit messages to user computing device 806, where the messages may include control instructions for updating the visuals displayed on UI 808.

[0064] Figure 9 This is a process flow 900 for guiding a robot using a UI, according to one or more embodiments. At 902, the method may include a computing device (such as a user computing device) displaying a first page on a first pane, wherein the first page provides first control input for aligning the working coordinate system of a target object with a reference from the robot. The UI may be an interface between a user and an application executing on a computing device such as a server. The first page may be a registration page, and registration may include calculating a transfer function to transform the coordinate system of the target object into the coordinate system of the robot. The UI may display a three-dimensional representation of the target object on a second pane (such as a visualization pane). The UI may further include visual markers indicating that the target object is registered such that the coordinate system of the target object can be transformed into the coordinate system of the robot.

[0065] At 904, the method may include receiving a first user selection of a first control input via a UI based on the detection of a first user selection, to align the working coordinate system with the reference coordinate system. The computing device may further receive via the UI an indication that the user has selected the control input based on the detection. In some cases, the user may have manually entered parameter values ​​for aligning the target object. In these cases, the computing device may further receive the parameter values ​​based on the detection.

[0066] At 906, the method may include a UI displaying a second page on a first pane, wherein the second page can provide a second control input for generating a path for the robot to traverse the surface of the motion planning page, and the path may include a set of generated virtual points and directions to reach each subsequent virtual point to traverse the surface of the target object. The second control input may be, for example, a calculate path button that a user can click to generate a path over the surface of the target object.

[0067] At 908, the method may include the computing device receiving, via a UI, a second indication of a second user selection to a second control input for generating a path to the computing device. The UI may detect the user selection to the second control input. For example, the UI may detect that the user has selected a path calculation button. Based on the detection, the computing device may receive an indication of the selection.

[0068] At 910, the method may include a UI displaying a path on a second pane, wherein the path includes corresponding directional markers between pairs of virtual points along the path and virtual points describing the direction of the path. The second pane may be a visual pane and is adjacent to the first pane.

[0069] Figure 10 This is a process flow 1000 for guiding a robot according to one or more embodiments. At 1002, the method may include a computing device (such as a server or user computing device) clustering points in a point cloud, wherein the point cloud corresponds to a target object. The computing device may receive a point cloud of the target object (such as an aircraft, vehicle, or other structure). In some cases, the computing device may retrieve the point cloud of the target object from a storage repository based on the specification of the target object. The computing device may further proceed to a clustering or segmentation stage and generate one or more clusters based on a first input. In addition to receiving instructions to generate clusters, the computing device may also receive user-defined clustering or segmentation parameters. The user may input the clustering or segmentation parameters into a UI, and the UI may transmit the parameters to the computing device. The computing device may translate the instructions and parameters for generating clusters into instructions for clustering the data points of the point cloud.

[0070] At 1004, the method may include determining whether each cluster or fragment represents a real-world target based on features that identify a real-world target. Features may be associated with characteristics of real-world objects, such as a circular pattern for a hole or fastener. The computing device may extract features from each cluster to determine whether the cluster includes a desired set of features. For example, the computing device may determine whether the cluster has a circular shape. The computing device may further identify clusters that include the desired features as targets.

[0071] At point 1006, the method may include a computing device calculating normals for each cluster corresponding to real-world targets. The computing device may utilize a series of algorithms to generate waypoints. Waypoints are virtual points located at the targets, which can be used to generate paths and trajectories on the surface of the target objects. Waypoints can also be used to guide the robot toward or away from points on the surface of the target objects. The number of waypoints can be configured by the user based on various parameters. For example, the user may select multiple waypoints based on the shape of the target object's surface, the number of targets, or the size of the target objects. The computing device may further calculate normals extending to or from each waypoint.

[0072] At 1008, the method may include a computing device registering the working coordinate system of the target object with the robot's reference coordinate system. This may include calculating a transfer function for associating the target object's reference coordinate system with the coordinate system of the robot's base. In other words, registration may include determining the robot's position relative to the target object. In some cases, a user may mark multiple targets on a real-world object and on a computer representation of the object. The user may further guide the robot to each target to determine the positional relationship between each marked target object and the robot's tool center point (TCP). The robot may transmit the positional relationship data back to the computing device. Based on the positional relationship data, the computing device may calculate a transfer function to model the relationship between the robot's position and the position on the target object in the robot's coordinate system.

[0073] At 1010, the method may include a computing device generating a path over the surface of a target object, wherein the path comprises a set of virtual points and a corresponding normal at each virtual point, and wherein each virtual point corresponds to a real-world target. The computing device may generate the path on the surface of the target object based on waypoints. The path may be a collection of three-dimensional movements traversing the surface of the target object. The path may be robot-agnostic and does not necessarily take into account robot limitations. The path may cover the entire surface of the target object or a portion of the surface of the target object. While the waypoint generation step creates waypoints to be used as the basis for the path, the waypoints are not sorted to provide an indication of how the path should proceed. One technique for sorting waypoints is to minimize desired parameters such as execution time or path length. The computing device may minimize the path length by using the waypoint graph as input to an optimization technique, such as an optimization algorithm for the Traveling Salesman Problem (TSP). The computing device may use the output of the optimization technique to determine the path across waypoints. In other embodiments, the computing system may employ other optimization techniques, such as graph-based techniques or other suitable techniques.

[0074] At 1012, the method may include a computing device generating a trajectory over the surface of a target object based on registration, a path, and each corresponding normal, wherein the trajectory includes a set of robot joint parameters for traversing the surface of the target object, and wherein the trajectory traverses the surface of a real-world target. The computing device may translate the path into a trajectory for moving the robot in a real-world scenario. The trajectory may take into account, for example, the robot's available motions, the desired pose of the end effector relative to the surface of the target object, and collision avoidance with the target object. It should be understood that the entire robot may need to move to avoid collisions as the end effector moves over the surface of the target object. The computing device may consider anticipated changes in the relative position of the target object and the robot when generating the trajectory. In some cases, the computing device may need to repeatedly generate the trajectory to reach the position of the target object or avoid collisions.

[0075] Computing devices can use inverse kinematics to calculate trajectories. Inverse kinematics is the process of taking the desired pose of a robot arm and determining the position of each robot component (e.g., base, upper arm, elbow, forearm, wrist, and end effector). Through inverse kinematics, the computing device can determine the configuration of the components to achieve the robot's desired pose. Various methods for inverse kinematics can be applied, such as closed-form solution methods or optimization methods. Once the computing device determines the desired position of each of the aforementioned components, it can determine the motion for each component to reach that desired position. A robot can be actuated to move about each of its axes to achieve multiple degrees of freedom (DoF), and multiple configurations can achieve the desired pose; this phenomenon is called kinematic redundancy. Therefore, the computing device can apply optimization techniques to calculate the optimal motion for each component to reach the desired position.

[0076] The computing device can further identify one or more collision-free continuous trajectories by employing various techniques to avoid collisions with obstacles by optimizing certain criteria. The computing device can simulate candidate trajectories and determine, based on registration, whether any robotic part will collide with another object. If a candidate trajectory will lead to a collision, the computing device can continue moving to the next candidate trajectory until it identifies a collision-free trajectory. The computing device can store collision data based on past collisions or past calculations. The computing device can further eliminate candidate trajectories based on the stored collision data. For example, the computing device can compare candidate trajectories with collision data to determine whether a collision is likely. If the probability of a collision is less than a threshold, the candidate trajectory can be eliminated. However, if the probability of a collision is greater than the threshold, the candidate trajectory is not eliminated.

[0077] At 1014, the method may include a computing device classifying real-world target types using a machine learning model based on scanned data of the surface of a target object. The computing device may receive image data of the target object. For example, the computing device may transmit a trajectory along with instructions to capture image data of the target to a robot. The computing device may receive the image data and use it as input to a neural network, such as a convolutional neural network (CNN). The neural network may receive an image and category for each target. For example, the neural network may predict the category of each target, such as fastener type, fastener diameter, or fastener material.

[0078] At 1016, the method may include a computing device generating a robot job file, wherein the robot job file includes a trajectory and autonomous operation instructions. The computing device may convert the trajectory, instructions for executing an end effector, and other relevant information into a file formatted for execution by the robot. The robot job may identify target categories and include instructions for each category. For example, a robot job may instruct the robot to perform certain operations for some fastener types and other operations for other fastener types. At 1018, the method may include a computing device transmitting the robot job file to a robot controller.

[0079] Figure 11 This is process flow 1100 for a robot to perform autonomous operations. At 1102, the method may include the robot receiving a trajectory for traversing a surface of a target object. The robot may receive the trajectory from a server, such as a server that calculates the trajectory.

[0080] At 1104, the method may include a robot using a trajectory to traverse the surface of a target object. In addition to traversing the surface, the robot may collect data, such as image data from the surface of the target object. For example, the robot may include an image capture device to capture images of the surface of the target object. The image data may include images of the target that will be used for target classification. The image data may also be transmitted back to a server for real-time display on a UI.

[0081] At 1106, the method may include a robot receiving a robotic task to be performed on a target object. The robotic task may include instructions for manipulating at least one target on the surface of the target object (e.g., drilling). The robotic task may be configured based on a trajectory, robot limitations, the intended task, and the classification of the target. In some embodiments, the robotic task further includes a recalculated trajectory for traversing the surface of the target object. At 1108, the method may include a robot performing the robotic task on the target object.

[0082] Figure 12This is a process flow 1200 for purifying a target object according to one or more embodiments. Figure 11 The illustration provided is for a process other than inspecting holes, drilling, or removing fasteners, using the aforementioned system and steps. At 1202, the method may include a computing device registering a target object to the robot's coordinate system. The computing device may be a server selected and guided by a user at the UI. For example, through the UI, the user can instruct the server to load a point cloud of the target object and identify targets on the surface of the target object. The user can further use the UI to instruct the server to generate waypoints on the three-dimensional representation of the target object. Waypoints may be movable virtual markers that can be used to generate normals to and from the waypoints. Waypoints may further be used as points for generating paths. The user can further use the UI to register the target object to the robot's coordinate system, particularly the robot's base.

[0083] At 1204, the method may include a computing device generating a path traversing the surface of the target. A user can use a UI to select input controls to send instructions to the computing device to generate the path. To generate the path, the computing device can utilize a series of algorithms to generate waypoints. Waypoints are virtual points on the surface of the target object that can be used to generate a path over the surface of the target object. The computing device can further employ optimization algorithms (e.g., a TSP solver) to identify the optimal path over the surface of the target object. As described above, the user can use the UI to move one or more waypoints, and thus, a path can be generated above, above, or below the surface of the target object based on the user's positioning of the waypoints. The path can be robot-agnostic so that it can be used by multiple robots. The generated path can be further stored in a database for later retrieval. For example, the path can be reused for subsequent robot operations on the target object.

[0084] At 1206, the method may include a computing device generating a first trajectory traversing the surface of a target object based on a path and registration. A user may use a UI to select input controls to send instructions to the computing device to generate the trajectory. The trajectory may be a set of motions along a path for each robot part, taking into account the robot's limitations and avoiding collisions. The computing device may generate the trajectory based on implementing inverse kinematics techniques. The computing device may use inverse kinematics to generate a trajectory that guides the robot in a real-world manner, such that the end effector has a desired posture along the trajectory. As described above, the computing device may use waypoints to generate corresponding normals at each target. The normals can be used to correctly guide the robot to approach and move away from the target in the correct orientation. The computing device may further transmit the trajectory to the robot for traversal. For example, the computing device may transmit a trajectory with instructions to scan the surface of the target object using sensors such as image capture devices or sniffers. The scan data may be based on sensors used to collect the data. For example, the scan data may be images collected by an image capture device. In other cases, scan data may be scan data measuring one or more properties of the surface of the target object or the material on or within the surface of the target object. The scan data may be further presented to the user via a UI. For example, the UI may include a visualization pane for displaying scan data. In some cases, scanning may be performed by a robot, and the UI can present the scan data to the user in real time.

[0085] At 1208, the method may include a computing device identifying contaminants on the surface of a target object. A user can use a UI to select input to send instructions to the computing device to identify the contaminants. The computing device may format the data for use with neural networks (e.g., CNNs, Gated Recurrent Unit (GRU) networks, Long Short-Term Memory (LSTM) networks), and use the network to identify the contaminants. The UI can display the identified contaminants to the user. In other cases, instead of receiving identifications or otherwise, the user can view the scan data in a visualization pane and manually input the contaminant identifications into the UI.

[0086] At 1210, the method may include a computing device generating a second path to the isolation zone based on an identifier. A user can use a UI to select input to send instructions for generating the second path. In some cases, the robot end effector is positioned above the contaminant after scanning the surface of the target object, and in other cases, the robot end effector is away from the contaminant after scanning the surface of the target object. Therefore, the computing device can generate a second path from the end effector's position to the contaminant. The second path can be generated using waypoints in the same manner as the first path. However, unlike the first path, the second path is not interested in traversing the surface of the target object. Instead, the second path relates to the optimal path for reaching the contaminant. Similar to the first path, the second path is robot-agnostic and can be stored in a database for reuse.

[0087] At 1212, the method may include a computing device generating a second trajectory based on a second path. A user can use a UI to select input to send instructions for generating the second trajectory. The computing device can use inverse kinematics techniques to determine the appropriate motion for a robot part to follow to reach the contaminated area, taking into account the robot's limitations while avoiding collisions. The robotic arm can follow the trajectory and move the tool to the contaminated area. Although the trajectory is not robot-aware, the second trajectory can also be stored in a database for reuse.

[0088] At 1214, the method may include a computing device generating a local pattern to clean a contaminated area. A user may use a UI to select input to send instructions for generating the local pattern. The local pattern may include any sequence of movements for applying or distributing a cleaning solution on the contaminated area. The pattern may be achieved through the movement of an end effector, the movement of a robotic arm, or a combination of both. The local pattern may be based on, for example, a contaminant, a solution for removing the contaminant, the surface of a target object, or a combination thereof. The local pattern may be, for example, a raster pattern, concentric circles, or other patterns.

[0089] In some cases, the computing device can further verify that the cleanup operation resulted in the cleanup of the target object. The computing device can move sensors (e.g., sniffer sensors) over the contaminated area with a robotic arm and determine if they can detect the contaminant. In other cases, the computing device can move sensors over the contaminated area with a robotic arm and transmit the collected data to a neural network, such as one used to identify contaminants. The neural network can be trained to process the sensor data and determine whether contaminants are present on the target object. If either method concludes that the contaminant remains, the computing device can restart the cleanup.

[0090] Figure 13This is an illustration of an environment for modeling a target object using artificial intelligence, according to one or more embodiments. In one example, the target object is an aircraft part, although the embodiments are similarly applicable to other types of target objects. The aircraft part includes multiple targets, such as fasteners, although the embodiments are similarly applicable to other types of targets. Each target may have a target type, such as a fastener type.

[0091] A multidimensional representation of an aircraft part can exist, such as a three-dimensional model of the aircraft part. This multidimensional representation can be generated and stored in a data repository according to the techniques described above. Typically, the multidimensional representation includes multidimensional modeling of each target included in the aircraft part. For example, each target corresponds to a target modeled in the multidimensional representation. The attitude data (e.g., position and orientation) of each modeled target can be stored in the multidimensional representation. This attitude data can be defined in a working coordinate system (e.g., in the first coordinate system of the aircraft part). Other types of data, such as dimensional data, geometric data, etc., can also be stored.

[0092] Artificial intelligence (AI) techniques can be used to generate updated multidimensional representations based on multidimensional representations, where the updated multidimensional representations may include, where applicable, more accurate attitude data corresponding to each target of an aircraft part. To this end, AI techniques involve using machine learning models in the processing of images, each image showing at least one target of an aircraft part and generated based on a camera positioning relative to the aircraft part. Positioning can be based on the multidimensional representation. The output of the machine learning model can indicate at least position data for each target. For each target, the output also indicates target type classification and / or visual inspection attributes. The position data can be used to generate updated attitude data, and this updated attitude data can be included, where applicable, in the updated multidimensional representation along with the classification type and / or visual inspection attributes.

[0093] exist Figure 13 The diagram illustrates artificial intelligence technology using multiple stages. Each stage will be described below.

[0094] In the model-based instruction phase 1301, server 1310 may receive a multidimensional representation 1312 of the aircraft part. Server 1310 may be the same server implementing the techniques described above for generating the multidimensional representation 1312. Receiving the multidimensional representation 1312 may include retrieving the multidimensional representation 1312 from a local or remote data repository of server 1310. Also in this model-based instruction phase 1301, server 1310 may generate a robot job based on the multidimensional representation 1312 in a manner similar to the techniques described above. Specifically, the robot job indicates a path traversing the surface of the aircraft part to access each target among the targets of the aircraft part and / or a trajectory for controlling the operation of robot 1320 such that each target is accessed. In this case, robot 1320 may be equipped with a camera (e.g., a camera attached thereto as an end effector), and the camera may be configured to TCP. The robot job may be sent to robot 1320 as robot file 1314.

[0095] Depending on the computing power of robot 1320, robot 1320, rather than server 1310, may generate and / or receive multidimensional representations 1312. Furthermore, robot 1320 may generate robot tasks and / or trajectories.

[0096] In image processing stage 1302, robot 1320 moves the camera based on job file 1314 to image an aircraft part (e.g., the wing of aircraft 1330). For example, the aircraft part is registered to robot 1320 such that it can be positioned in robot 1320's reference coordinate system (e.g., in robot 1320's second coordinate system). Multidimensional representation 1312 can indicate the target pose of the aircraft part, defined in the working coordinate system. Based on forward kinematics, this pose can be redefined in the robot's reference coordinate system. Using the camera as the TCP and the redefined pose, the camera can be positioned at a predefined distance from the target (e.g., 1 cm or some other distance that could be a function of the camera's focal length) and at a relative orientation to the target (e.g., the focal length normal to the target's surface) based on inverse kinematics.

[0097] As explained above, inverse kinematics involves calculating the variable joint parameters required to position the end of a kinematic chain (e.g., a camera in a kinematic chain comprising the robot 1320 arm and its components) at a given position and orientation relative to the start of the chain (e.g., the origin of the local coordinate system). Forward kinematics involves calculating the camera position using equations of motion based on specified values ​​for the joint parameters. Inverse kinematics takes the Cartesian camera position and orientation (e.g., as defined in the local coordinate system and corresponding to TCP) as input and calculates the joint angles, while forward kinematics (for the arm) takes the joint angles as input and calculates the Cartesian position and orientation of the camera. Through inverse and forward kinematics, the robot 1320 can determine the configuration of its base, arm components, and camera to achieve the desired pose of the arm and camera relative to a target. Various methods can be applied to the robot 1320 for inverse and forward kinematics, such as closed-form solution methods or optimization methods. The arm and camera can be actuated to move about each of their axes to achieve multiple DoFs, and multiple configurations can achieve the desired pose; this phenomenon is called kinematic redundancy. Therefore, robot 1320 can apply optimization techniques to calculate the optimal motion for each component to reach the desired posture.

[0098] Once the camera is in the desired orientation relative to the target of the aircraft component (e.g., such that the target is centered in the camera's field of view and its surface is normalized relative to the camera), the camera can generate one or more images showing the target. For example, robot 1320 can instruct the camera to generate images when it detects that the camera has reached the desired orientation based on arm and camera actuation. Thus, images 1322 of different targets can be generated and sent to server 1310.

[0099] Based on each image showing the target, server 1310 can generate input to ML model 1316. Input can be, for example, the image itself or a processed version of the image (e.g., through cropping, blur reduction, brightness enhancement, or any other image processing techniques applied to the image). Server 1310 can execute program code for ML model 1316 and can provide input for execution (although ML model 1316 may be hosted on a different computer system, such as a cloud service, and may be accessible by server 1310 via application programming interface (API) calls). Output of ML model 1316 is then generated. In one instance, the camera is a two-dimensional camera. In this instance, the image showing the target is two-dimensional. The output indicates two-dimensional position data of the detected target in the image, where the detected target corresponds to the target. In one instance, the camera is a three-dimensional camera (e.g., a depth camera). In this instance, the image showing the target is three-dimensional. The output indicates three-dimensional position data and orientation data of the detected target in the image, where the detected target corresponds to the target. In one instance, the camera is a two-dimensional camera, and a depth sensor is also used to generate depth data. In this instance, the image showing the target is two-dimensional. Depth data can indicate the depth of each part of a target (e.g., at the pixel level of a depth sensor). Inputs can also include such depth data. In this example, the output also indicates three-dimensional position and orientation data of detected targets in the image, where the detected targets correspond to the target. The ML model 1316 can also be trained to output classification and / or visual inspection attributes of targets based on the inputs. Position data (e.g., two-dimensional or three-dimensional, as applicable) and orientation data (if applicable) can be defined in a camera coordinate system (e.g., a third coordinate system of the camera, such as UV coordinates and rotation). Classification can indicate the type of target (e.g., the type of fastener). Visual inspection attributes can indicate attributes of the target that would otherwise be detected using a visual process that would otherwise physically inspect the target (e.g., to determine fastener corrosion, corrosion of the area surrounding the fastener, the presence of the fastener, the absence of the fastener, fastener damage, damage to the surrounding area, or any other detectable attribute during physical visual inspection).

[0100] In the model update phase 1320, server 1310 can execute program code for attitude estimator 1318 (although attitude estimator 1318 may be hosted on a different computer system such as a cloud service and may be accessible by server 1310 via application programming interface (API) calls). At least a portion of the output model 1316 is used as input to attitude estimator 1318. For example, for each target, this input may include position data (e.g., two-dimensional or three-dimensional, as applicable) and orientation data (if applicable). If the input is two-dimensional position data for the target, attitude estimator 1318 can estimate the target's three-dimensional attitude data. This estimation may be based on a predefined distance and orientation of the camera when the target is imaged. Based on inverse kinematics, the three-dimensional attitude data can be expressed in the coordinate system of robot 1320 (e.g., a reference coordinate system), and based on forward kinematics, it can be further expressed in the coordinate system of the aircraft parts (e.g., a working coordinate system). If input is 3D position data for the target, the attitude estimator 1318 can update it to be expressed in the coordinate system of the robot 1320 and / or the aircraft path. If orientation data is not available, the target's orientation can be estimated based on the camera's orientation at the time of imaging, and this orientation can be expressed in the coordinate system of the robot 1320 and / or the aircraft path. If orientation data is also available, the attitude estimator 1318 can update it to be expressed in the coordinate system of the robot 1320 and / or the aircraft path. In any case, the attitude estimator 1318 can output attitude data (e.g., 3D position data and orientation data) for each imaged and detected target, where the attitude data can be expressed in the coordinate system of the robot 1320 and / or the aircraft part.

[0101] Server 1310 generates updated multidimensional representations 1319 of the aircraft parts and stores them in data repository 1340. For each detected target, the updated multidimensional representation 1319 may include the target's attitude data (e.g., in the coordinate system of the aircraft part) and its classification and / or visual inspection attributes. In data repository 1340, the updated multidimensional representation 1319 may be associated with identifiers of the aircraft part (e.g., product number) and / or the aircraft (e.g., tail number).

[0102] In the operation instruction phase 1304, server 1310 can use the updated multidimensional representation 1319 to generate a robotic job for a set of robotic operations to be performed on an aircraft part (although another computer system may do so alternatively or additionally). For example, for a specific target type (e.g., fastener type) and / or a specific visual inspection attribute (e.g., damaged fastener), server 1310 can determine a set of targets for the aircraft part and can generate paths and / or trajectories based on such paths that traverse the surface of the aircraft part to access each of these targets. The robotic job can instruct the path (or trajectory) and operation (e.g., fastener removal) to be performed at each of these targets. The robotic job can be sent to robot 1320 as job file 1315 (although job file 1315 can be sent to different robots). The camera can be replaced with an end effector configured to perform this type of robotic operation (e.g., a fastener drill end effector), so that robot 1320 can use the end effector as a TCP to calculate the trajectory (if it is not yet available in job file 1315) and control the movement of the arm and end effector to perform robotic operations.

[0103] Figure 14 This is an illustration of using ML model 1410 as part of using artificial intelligence to model a target object, according to one or more embodiments. ML model 1410 is... Figure 13 An example of ML model 1316 is shown. The training of ML model 1410 is further illustrated in the following figures.

[0104] In one instance, machine learning model 1410 includes an artificial neural network (ANN) trained to at least detect targets shown in an image and output location data of the detected targets in the image. The location data may correspond to bounding boxes surrounding the detected targets and may be defined as camera UV coordinates. The ANN (or possibly, another model, such as a classifier included in machine learning model 1410) may also be trained to classify the detected targets as having a target type. Additionally or alternatively, the ANN (or possibly, another model, such as another ANN included in machine learning model 1410) may also be trained to indicate visual inspection attributes of the detected targets.

[0105] Therefore, input to the machine learning model 1410 is received. This input may include multiple images 1402. Each image 1402 may be generated by a camera, with the camera positioned at a predefined distance and orientation relative to a target included in the aircraft part, and thus image 1402 may show the target. For example, image 1402 showing the target is generated when the camera is positioned such that the target is at a predefined distance from the camera, normalized to the camera, and at the center of the camera's field of view.

[0106] Images 1402 can be processed sequentially or in batches by being input into ML model 1410. Sequential input may include inputting an image 1402 showing a target into ML model 1410 in real time relative to when image 1402 is generated. Sequential input may also include preventing the camera from being repositioned to image another target until the output of ML model 1410 is generated and that output corresponds to the successful processing of image 1402. If processing is unsuccessful, the camera may be instructed to generate a new image showing the target while still in the same position, or the camera may be repositioned to improve the imaging of the target (e.g., by changing the distance and / or by changing the camera orientation so that the target is centered in its field of view). If processing is successful, the camera may be repositioned relative to the next target and instructed to generate an image 1402 showing that target. Batch input may include positioning the camera relative to each target, generating a corresponding image 1402, and then inputting these images 1402 as a batch load into ML model 1410.

[0107] After successfully processing the image 1402 showing the target, the output of the ML model 1410 may include position data 1404 of the target (e.g., the target's bounding box around the target). As explained above, the position data 1404 may be two-dimensional or three-dimensional depending on whether the image 1402 is two-dimensional or three-dimensional. Further, if the image is three-dimensional and / or if depth data is included in the input, the output may include rotation data of the target. Additionally or alternatively, the output may include target classification 1406 and / or visual inspection attributes of the target 1408.

[0108] Figure 15 This is an illustration of the estimated pose of a target according to one or more embodiments as part of using artificial intelligence to model the target object. In this illustration, a machine learning model (e.g., Figure 14 The output of the machine learning model 1410 includes two-dimensional position data 1502. In addition to the position data 1502, distance data 1504 and camera orientation data 1506 can be input into the pose estimator 1510. The output of the pose estimator 1408 can include pose data 1508.

[0109] In one example, position data 1502 may indicate the two-dimensional position of a bounding box around a target in the camera's UV coordinate system, where the bounding box corresponds to the detection of the target in an image. The image can be generated when the camera is positioned relative to the target at a predefined distance and orientation (e.g., such that the camera is normalized to the target). Distance data 1504 may indicate a predefined distance between the camera (e.g., its focal point) and the target. Camera orientation data 1506 may indicate the camera's orientation in the camera's coordinate system and / or in the coordinate system of a robot that controls the camera via TCP.

[0110] The camera's pose is known in the robot's coordinate system. Based on this pose and inverse kinematics, the pose estimator 1510 can convert the target's position data into two-dimensional coordinates (e.g., X and Y coordinates) in the robot's coordinate system. Based on the camera's pose and inverse kinematics, the pose estimator 1510 can also convert a predefined distance into a third-dimensional coordinate (e.g., Z coordinate) in the robot's coordinate system. It can be assumed that the camera's orientation is the same as or substantially similar to the target's orientation. Therefore, the pose estimator 1510 can set the camera orientation data 1506 as the target's orientation data in the robot's coordinate system. Thus, the pose estimator 1510 can estimate the target's pose in the robot's coordinate system. Based on forward kinematics, the pose estimator 1510 can convert this estimated pose into the coordinate system of the aircraft part. Therefore, the resulting pose data 1508 can include the target's three-dimensional position and orientation data in the robot's coordinate system and / or the aircraft part's coordinate system.

[0111] The above illustrates a use case for two-dimensional position data output by an ML model. If the ML model outputs three-dimensional position data, this data will be expressed in the camera's coordinate system. Therefore, by using inverse kinematics and / or forward kinematics transformations, the attitude estimator 1510 can estimate the target's three-dimensional position in the robot's coordinate system and / or the aircraft part's coordinate system. Similarly, if the ML model outputs orientation data, this data will be expressed in the camera's coordinate system. Therefore, by using inverse kinematics and / or forward kinematics transformations, the attitude estimator 1510 can estimate the target's orientation in the robot's coordinate system and / or the aircraft part's coordinate system.

[0112] Once the target's attitude data 1508 is estimated, this attitude data 1508 can be included in the updated multidimensional representation of the aircraft part. Alternatively, the updated multidimensional representation of the aircraft can be initialized to be the same as the multidimensional representation of the modeled aircraft part. In this case, once the attitude data 1508 is updated, it can be compared with the modeled attitude data of the corresponding modeled target. This comparison can indicate a threshold offset (e.g., position offset and / or rotation offset). If the offset is greater than the threshold offset, the attitude data 1508 can replace the modeled attitude data in the updated multidimensional representation. Otherwise (e.g., the offset is less than the threshold offset), the modeled attitude data is retained in the updated multidimensional representation. The threshold offset can depend on the accuracy of the ML model and / or acceptable accuracy error.

[0113] Figure 16 This is an illustration of reclassifying targets and updating paths for robot manipulation according to one or more embodiments. In one example, an aircraft part model 1602A is defined. For example, the aircraft part model 1602A includes a multi-dimensional (e.g., three-dimensional) representation of an aircraft part generated according to the techniques described above herein. Path modeling can be applied to the modeled targets (e.g., waypoints) by defining paths 1604 for traversing the targets being modeled using optimization algorithms (e.g., a TSP solver).

[0114] The aircraft parts can be imaged, and the images can be processed as described above to generate updated aircraft part models 1602B (e.g., updated 3D representations of the aircraft parts). The updated aircraft part model 1602B can be further processed for path replanning. Both types of processing are... Figure 16 The example shown is ML processing and path replanning 1610.

[0115] The updated aircraft part model 1602B can indicate the target type and / or visual inspection attributes for each target. Filters can be set on one or both of the target type and / or visual inspection attributes. For example, filters can be set for specific types of fasteners and fastener damage. The targets modeled in the updated aircraft part model 1602B can be filtered accordingly to generate a set of modeled targets (e.g., a set of fasteners with a specific fastener type and that are damaged). Path replanning 1610 can be applied to this set, thereby applying an optimization algorithm to the set to define paths for traversing only these modeled targets.

[0116] exist Figure 16The illustration uses two filters. The first filter corresponds to the first target type A 1606A. The corresponding modeled target is shown in a solid black shape. Path A 1608A is generated for these modeled targets. Similarly, the second filter corresponds to the second target type B 1606B. The corresponding modeled target is shown in a dashed shape. Path B 1608B is generated for these modeled targets. Therefore, assuming the first target type A 1606A is a first fastener type, the robot can be instructed to perform a first robotic operation on the corresponding target using the first fastener drill end effector based on path A1608A. Similarly, assuming the first target type B 1606B is a second fastener type, the robot can be instructed to perform a second robotic operation on the corresponding target using the second fastener drill end effector based on path B 1608B.

[0117] Figure 17 This is a process flow 1700 according to one or more embodiments for modeling a target object using artificial intelligence. Process flow 1700 can be implemented as a method by a system. The system can be a server, such as... Figure 13 Server 1310. At 1702, the method may include receiving a multidimensional representation of an aircraft part including a target, the multidimensional representation including first attitude data of the target in a first coordinate system of the aircraft part. For example, the multidimensional representation is a three-dimensional representation generated using the techniques described above.

[0118] At 1704, the method may include enabling the robot to position a camera relative to a target based on first pose data and the robot's second coordinate system. For example, a task file is generated based on a multidimensional representation and sent to the robot. Based on the task file, the robot can determine a path for traversing the surface of the aircraft part, where the path includes a series of pose data corresponding to each target on the aircraft part. The camera is configured as a TCP. Based on inverse and forward kinematics applied to the sequence, the robot determines a trajectory for the camera relative to each target, such that the camera can be positioned relative to the target at a predefined distance and oriented in a normalized manner.

[0119] At 1706, the method may include receiving an image generated by the camera while the camera is positioned relative to a target, the image showing at least the target. For example, at each location, the robot may instruct the camera to generate the next image, and may send that image to the system. In sequential image processing, the robot does not reposition the camera until it receives an indication from the system of successful processing of the current image. In batch image processing, the robot may continue repositioning the camera without any indication from the system of successful image processing.

[0120] At 1708, the method may include generating input to the machine learning model based on images. For example, image processing may be applied to the images (e.g., deblurring, cropping, etc.), and the resulting image data may be included in the input. If the depth data is generated based on a depth sensor, the depth data may also be included in the input.

[0121] At 1710, the method may include determining second attitude data of the target in a first coordinate system of the aircraft part based on a first output of a machine learning model in response to the input. For example, the first output may indicate position data in a camera's UV coordinate system. If the position data is two-dimensional, the attitude estimator may estimate a third dimension of the position data based on a predefined distance from the camera. The three-dimensional position data can be transformed into the first coordinate system using inverse and forward kinematic transformations. If the position data is already three-dimensional, the inverse and forward kinematic transformations are used to transform it into the first coordinate system. If the first output indicates orientation data of the part, the inverse and forward kinematic transformations are used to transform the orientation data into the first coordinate system. Otherwise, orientation data may be estimated from the camera's orientation and / or from depth data for subsequent transformation into the first coordinate system. Further, the first output may indicate the target's classification and / or visual inspection attributes.

[0122] At 1712, the method may include generating a second output based on second attitude data, which can be used to control robot manipulation of a target. For example, first attitude data and second attitude data are compared to determine an offset. If the offset is greater than a threshold offset, the second attitude data may be labeled as measured attitude data suitable for replacing the first attitude data. Otherwise, the first attitude data does not need to be updated. Similarly, the classification (or target type) indicated by the first output may be compared with the modeled classification (or modeled target type) included in the multidimensional representation. If they differ, the classification may be labeled as suitable for replacing the modeled classification. Likewise, the visual inspection attribute indicated by the first output may be compared with the modeled visual inspection attribute included in the multidimensional representation. If they differ, the visual inspection attribute may be labeled as suitable for replacing the modeled visual inspection attribute. An updated multidimensional representation of an aircraft part may be generated by updating the updated multidimensional representation of the aircraft part, whereby the modeled attitude data, the modeled classification, and / or the modeled visual inspection attributes are replaced by the measured attitude data, the determined classification, and / or the determined visual inspection attributes. Then, the updated multidimensional representation can be used for path replanning and subsequent trajectory planning to control robot operations.

[0123] At 1714, the method may include generating an updated multidimensional representation of the aircraft part based on the second output. For example, the modeled attitude data, modeled classification, and / or modeled visual inspection attributes may be replaced, as applicable, with measured attitude data, determined classification, and / or determined visual inspection attributes.

[0124] At 1716, the method may include storing updated multidimensional representations of aircraft parts in a data repository. For example, the updated multidimensional representations of aircraft parts may be stored in association with identifiers of aircraft parts and / or identifiers of aircraft that include aircraft parts, as applicable.

[0125] At 1718, the method may include enabling a robot to perform a series of robotic operations on an aircraft part based on an updated multidimensional representation of the part. The robot may, but is not necessarily, the same as the robot used to locate the camera. Path and trajectory planning may be generated based on the updated multidimensional representation to control the localization of the robot's end effector relative to at least some of the targets. The end effector may be configured to perform robotic operations on such targets.

[0126] Figure 18 This is an illustration of an environment for artificial intelligence training according to one or more embodiments. Artificial intelligence training can utilize synthetic image-based training and real image-based training. Figure 18 The diagram illustrates AI training using multiple stages. Each stage will be described below.

[0127] In the first training phase 1801, server 1810 can receive and train ML model 1812. Server 1810 can communicate with... Figure 13 The server 1310 may be the same as or different from the target. The ML model 1812 may be pre-trained for at least object detection, wherein the training is general and not specific to aircraft parts and / or targets included in aircraft parts. For example, the ML model 1812 may include a Single-Shot Multi-Box Detection (SSD) network, such as MobileNet SSD, trained to perform object detection. A first training phase 1801 further trains the ML model 1812 to at least detect targets included in aircraft parts, wherein the training uses training data 1802 based on synthetic images. Target detection may include two-dimensional position data of detected targets in two-dimensional images and / or three-dimensional position data of detected targets in three-dimensional images based on depth data, and possibly pose data of detected targets in two-dimensional or three-dimensional images. The ML model 1812 may also be trained to classify targets into target types and / or indicate visual inspection attributes of targets. Training may remove the output layer of the ML model 1812 and replace it with a new output layer for target detection and classification and / or visual inspection attribute determination (if applicable).

[0128] The training data 1802 based on synthetic images may include synthetic images showing modeled targets of different types and across different aircraft parts and / or aircraft. In one instance, a multidimensional representation of an aircraft part is used to generate a synthetic image showing a modeled target corresponding to that aircraft part. This multidimensional representation may indicate the attitude data, target type, and visual inspection attributes of the modeled target. A two-dimensional image can be generated by generating a two-dimensional snapshot of the modeled target from the multidimensional representation. Two-dimensional position data from the attitude data is associated with the two-dimensional image. Similarly, the target type and visual inspection attributes are associated with the two-dimensional image. The two-dimensional image can be set as a synthetic image showing the modeled target. This synthetic image and the associated position data, target type, and visual inspection attributes may be included in the training data 1802 based on the synthetic images. The position data, target type, and visual inspection attributes serve as ground reality associated with the synthetic image (e.g., as training labels). Furthermore, image transformations (e.g., blurring, distortion, mirroring, etc.) may be applied to the two-dimensional image to generate another synthetic image. Here, if the image transformation alters the position data (e.g., based on some transformation or rotation function), the same function can be applied to the position data to subsequently generate updated position data. Target type and visual inspection attributes do not change with the image transformation. The synthetic image and associated position data (updated where applicable), target type, and visual inspection attributes can be included in the synthetic image-based training data 1802, wherein the associated position data (updated where applicable), target type, and visual inspection attributes are also set to ground reality for the synthetic image. Even using a two-dimensional synthetic image, target orientation data can be included in the training labels, allowing the ML model 1812 to be trained to estimate target orientation based on the two-dimensional image. In addition to generating two-dimensional images, or alternatively, a three-dimensional image can be generated by applying a three-dimensional snapshot to a multidimensional representation of the aircraft parts and applying an image transformation to the three-dimensional snapshot, and this three-dimensional image can be included in the synthetic image-based training data 1802. For a 3D synthetic image, associated 3D data (updated where applicable if image transformations are applied and affect such location data), target type, and visual inspection attributes can be included in the training data 1802 based on the synthetic image as ground reality for that 3D synthesis.

[0129] Furthermore, in the first training phase 1801, a loss function can be used to train the ML model 1812. For example, the loss function can be defined using the offset between the estimated location data and the ground-based location data (two-dimensional or three-dimensional). Where applicable, the loss function can also be defined using the offset between the estimated orientation data and the ground-based orientation data, the classification mismatch between the determined target type and the ground-based target type, and / or the classification mismatch between the determined visual inspection attribute and the ground-based visual inspection attribute. Training can be iterative to minimize the loss function by updating the parameters of the ML model 1812 (e.g., the weights between connections between nodes across different layers) using a backpropagation algorithm.

[0130] Next, in image processing stage 1802, server 1810 receives data from robot 1820 (similar to...). Figure 13 The robot 1320 receives images 1822, which image targets that may be included in aircraft parts within the aircraft 1830. These images can be generated according to the techniques described above and can be used to generate inputs to the ML model 1812. Based on the inputs corresponding to the images showing the targets, the ML model 1812 can output position data 1804 (e.g., two-dimensional or three-dimensional, as applicable) and, if applicable, orientation data, target type, and / or visual inspection attributes.

[0131] In the human-machine loopback phase 1803, human-machine loopback training can be triggered. For example, the output of the ML model 1812 (or at least the position data 1804) can be subject to human review and correction (if applicable). Each output can be reviewed, or a sample of outputs can be reviewed. Additionally or alternatively, a selection process can be implemented to select the set of outputs for review. For outputs related to aircraft parts, this selection process can rely on a multidimensional representation of the aircraft part. In particular, the output can be specific to a target of the aircraft part and can include determined position data of the target, and optionally, determined orientation data, determined target type, and determined visual inspection characteristics of the target. For the target, the multidimensional representation can include modeled position data, modeled orientation data, modeled target type, and modeled visual inspection attributes of the target. The determined position data can be compared with the modeled position data to determine an offset, and this offset can be compared with a threshold offset. If it is greater than the threshold offset, the output is marked for human review. Similarly, the determined orientation data can be compared with the modeled orientation data to determine an offset, and this offset can be compared with a threshold offset. If the offset exceeds a threshold, the output is marked for human review. Alternatively, the determined target type can be compared to the modeled target type to determine if a match exists. Conversely, if a mismatch exists, the output is marked for human review. Alternatively, the determined visual inspection attributes can be compared to the modeled visual inspection attributes to determine if a match exists. Conversely, if a mismatch exists, the output is marked for human review.

[0132] Human review of the output may involve correcting the output. Server 1810 may receive human review from user computing device 1840 operated by a human user. Specifically, the output and the corresponding image processed to generate the output are presented via a UI on user computing device 1840. User input is received by user computing device and instructs for correction data. Correction data may be any or a combination of corrected location data, corrected orientation data, corrected target type, and / or corrected visual inspection attributes. The image is then added to training data 1814. The corrected location data, corrected orientation data, corrected target type, and / or corrected visual inspection attributes are also added to training data 1814 as ground truth of the image (e.g., as training labels associated with the image).

[0133] Next, in the second training phase 1804, the ML model 1812 is further trained based on the training data 1814. Here, unlike the first training phase 1801, it is not necessary to remove or replace the output layer. Instead, the same loss function can be reused to iteratively refine the parameters of the ML model 1812, minimizing the loss function.

[0134] Figure 19 This is an illustration of a training method using synthetic image training data 1920 and detected target training data 1930 according to one or more embodiments. The synthetic image training data 1920 may correspond to... Figure 18 The training data 1802 is based on synthetic images, while the training data 1930 for detected targets can correspond to... Figure 18 The training data is 1814.

[0135] exist Figure 19 In the illustration, the training method involves a first training 1901 and a second training 1902. In the first training 1901, synthetic image training data 1920 is used to reconfigure a generally trained ML model 1910 (e.g., a convolutional neural network) so that the trained ML model 1910 can process images showing targets depicting aircraft parts. The synthetic image training data 1920 may include synthetic images and ground realities associated with the synthetic images. In the second training 1902, detected target training data 1930 is used to refine the parameters of the ML model 1910 previously trained in the first training 1901 so that the ML model 1910 can more accurately process images showing targets depicting aircraft parts. The detected target training data 1930 may include actual images showing targets depicting aircraft parts and corrections to the output of the ML model 1910, where these corrections may be based on human-machine loopback training and may be used as ground realities associated with actual images. Once the first training 1901 is complete, actual images can be processed, and then corrections can be received to initiate the second training 1902. The second training 1902 can be continuous (e.g., repeated at time intervals within a planned time frame), where the ML model 1910 can be used for image processing.

[0136] Figure 20 This is a process flow 2000 for training artificial intelligence according to one or more embodiments. Process flow 2000 can be implemented as a method by a system. The system can be a server, such as... Figure 18 Server 1810. At 2002, the method may include generating synthetic images and ground fact data based on a multidimensional representation of aircraft parts. As described above, a two-dimensional or three-dimensional snapshot may be applied to the multidimensional representation to generate an image (two-dimensional or three-dimensional) showing the modeled target. This image may be set as a synthetic image, and position data, orientation data, target type, and visual inspection attributes indicated in the multidimensional representation for the modeled target may be associated with the synthetic image as its ground fact data. Image transformations may be applied to this image, and to the position data and / or orientation data (if applicable), to generate another synthetic image and another set of ground fact data.

[0137] In 2004, the method could include inducing an initial training of the machine learning model based on first training data. For example, the first training data is input into the machine learning model. A loss function is defined, and the parameters of the machine learning model are updated to minimize the loss function.

[0138] In 2006, the method could include a multidimensional representation based on aircraft parts to enable the robot to locate a camera relative to a first target included in the aircraft parts. This operation of the method could be similar to... Figure 17 Operation 1704.

[0139] At 2008, the method could include receiving a first image generated by the camera while the camera is positioned relative to a first target, the first image showing at least the first target. This operation of the method could be similar to... Figure 17 Operation 1706.

[0140] In 2010, the method could include generating a first input to a machine learning model based on a first image. This operation of the method could be similar to... Figure 17 Operation 1708.

[0141] At 2012, the method may include determining a first output of a machine learning model based on a first input, the first output indicating at least one of a first location data of a first detected target in a first image, a first classification of the first detected target, or a first visual inspection attribute of the first detected target corresponding to a first target. This operation of the method may be similar to... Figure 17 Operation 1710.

[0142] At 2014, the method may include determining at least one of the following: location data of the first correction of the detection of the first correction of the first target from the first image, classification of the first correction of the detection of the first correction, or visual inspection attributes of the first correction of the detection of the first correction. For example, a human-machine loopback training method may be invoked, thereby receiving correction data from a user computing device. A server may send a first image and a first output to the computing device, which then presents such data on a UI. User input may be received in response to an indication of correction data. The first image and the first output may be selected for review based on sampling or based on a selection procedure dependent on a multidimensional representation of aircraft parts.

[0143] In 2016, the method could include generating second training data based on first-corrected location data. For example, a first image is included in the second training data. Corrected data is also included in the second training data as ground-based data associated with the first image.

[0144] At 2018, the method may include a second training of the machine learning model based on second training data. This operation is similar to operation 2004, where the second training data is used instead of the first training data.

[0145] Figure 21 This is a block diagram of an example of a computing system 2100 capable of implementing some aspects of this disclosure. Components of the computing system 2100 may be part of a computing device, a server, a robot, and / or distributed between any two of the computing device, server, and robot. The computing system 2100 includes a processor 2104 coupled to memory 2104 via bus 2112. The processor 2102 may include one or more processing devices. Examples of the processor 2102 include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), microprocessors, or any combination thereof. The processor 2102 can execute instructions 2106 stored in memory 2104 to perform operations. In some instances, instructions 2106 may include processor-specific instructions generated by a compiler or interpreter from code written in any suitable computer programming language, such as C, C++, C#, Python, or Java.

[0146] Memory 2104 may include one or more memory devices. Memory 2104 may be non-volatile and includes any type of memory device that retains stored information when power is off. Examples of memory 2104 may include electrically erasable and programmable read-only memory (EEPROM), flash memory, or any other type of non-volatile memory. At least some of the memories in memory 2104 include non-transitory computer-readable media from which processor 2102 can read instructions 2106. Computer-readable media may include electronic storage devices, optical storage devices, magnetic storage devices, or other storage devices capable of providing computer-readable instructions or other program code to processor 2102. Examples of computer-readable media include magnetic disks, memory chips, ROM, random access memory (RAM), ASICs, configured processors, optical storage devices, or any other media from which a computer processor can read instructions 2106.

[0147] The computing system 2100 may also include other input and output (I / O) components. Input components 2108 may include a mouse, keyboard, trackball, touchpad, touchscreen display, or any combination thereof. Output components 2110 may include a visual display, audio display, haptic display, or any combination thereof. Examples of visual displays may include liquid crystal displays (LCDs), light-emitting diode (LED) displays, and touchscreen displays. Examples of audio displays may include speakers. Examples of haptic displays may include piezoelectric devices or eccentric rotating mass (ERM) devices.

[0148] The above description of certain examples (including illustrative examples) has been presented for illustrative and descriptive purposes only and is not intended to be exhaustive or to limit this disclosure to the precise forms disclosed. Modifications, adaptations, and uses thereof will be apparent to those skilled in the art without departing from the scope of this disclosure. For example, any example described herein can be combined with any other example.

[0149] Although specific embodiments have been described, various modifications, alterations, alternative constructions, and equivalents are also covered within the scope of this disclosure. The embodiments are not limited to operation within certain specific data processing environments, but can be freely operated within multiple data processing environments. Furthermore, although the embodiments have been described using a specific set of transactions and steps, it will be apparent to those skilled in the art that the scope of this disclosure is not limited to the described set of transactions and steps. Various features and aspects of the above embodiments can be used individually or in combination.

[0150] Furthermore, while embodiments have been described using specific combinations of hardware and software, it should be recognized that other combinations of hardware and software are also within the scope of this disclosure. Embodiments may be implemented solely in hardware, solely in software, or using a combination thereof. The various processes described herein can be implemented on the same or different processors in any combination. Thus, where a component or module is described as being configured to perform certain operations, such configuration can be implemented, for example, by designing electronic circuitry to operate, by programming programmable electronic circuitry (such as a microprocessor) to operate, or any combination thereof. Processes may communicate using various technologies, including but not limited to conventional technologies for inter-process communication, and different pairs of processes may use different technologies, or the same pair of processes may use different technologies at different times.

[0151] Therefore, the specification and drawings are to be considered illustrative rather than restrictive. However, it will be apparent that additions, subtractions, deletions, and other modifications and changes may be made therein without departing from the broader spirit and scope set forth in the claims. Thus, although specific disclosed embodiments have been described, they are not intended to be restrictive. Various modifications and equivalents are within the scope of the appended claims.

[0152] Unless otherwise stated herein or explicitly contradicted by the context, the terms “a / an” and “the”, and similar pronouns, used in the context of describing the disclosed embodiments (especially in the context of the following claims) should be interpreted to cover both the singular and plural. Unless otherwise stated, the terms “comprising,” “having,” “including,” and “containing” should be interpreted as open-ended terms (i.e., meaning “including but not limited to”). The term “connected” should be interpreted as partially or wholly contained in, attached to, or joined together, even if something else is involved. Unless otherwise indicated, the description of value ranges herein is intended only as a shorthand method of individually referring to each individual value falling within the range, and each individual value is incorporated into the specification as if individually described herein. Unless otherwise stated herein or explicitly contradicted by the context, all methods described herein can be performed in any suitable order. Unless otherwise required, the use of any and all instances or exemplary language (e.g., “such as”) provided herein is intended only to better illustrate the embodiments and does not constitute a limitation on the scope of this disclosure. Nothing in this specification should be construed as indicating any unclaimed element as necessary for practicing this disclosure.

[0153] Unless otherwise specifically stated, disjunctive languages ​​such as the phrase “at least one of X, Y, or Z” are intended to be understood in context as generally used to represent that items, terms, etc., can be X, Y, or Z or any combination thereof (e.g., X, Y, and / or Z). Therefore, such disjunctive languages ​​are generally not intended and should not imply that some embodiments require that at least one of X, at least one of Y, or at least one of Z each be present.

[0154] This document describes preferred embodiments of the present disclosure, including the best modes known for carrying out the present disclosure. Variations of those preferred embodiments may become apparent to those skilled in the art after reading the foregoing description. Those skilled in the art should be able to suitably employ such variations and may practice the present disclosure in ways other than those specifically described herein. Therefore, the present disclosure includes all modifications and equivalents to the subject matter set forth in the appended claims, where permitted by applicable law. Furthermore, unless otherwise stated herein, the present disclosure covers any combination of the foregoing elements with all their possible variations.

[0155] All references cited in this document (including publications, patent applications and patents) are hereby incorporated by reference to the extent that each reference is individually and specifically indicated by reference and is presented in the whole document.

[0156] In the foregoing description, various aspects of this disclosure have been described with reference to specific embodiments thereof; however, those skilled in the art will recognize that this disclosure is not limited thereto. The various features and aspects of the foregoing disclosure may be used individually or in combination. Furthermore, embodiments may be utilized in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of this specification. Therefore, the specification and drawings are to be considered illustrative rather than restrictive.

Claims

1. A computer-implemented method comprising: The robot is positioned relative to a first target camera included in the aircraft parts by using a multidimensional representation of the aircraft parts. A first image generated by the camera is received when the camera is positioned relative to the first target, the first image showing at least the first target; The first image is used to generate the first input to the machine learning model; The first output of the machine learning model is determined based on the first input, and the first output indicates the first location data of a first detected target in the first image, the first detected target corresponding to the first target. The first corrected position data of the detection of the first target is determined from the first image; First training data is generated based on the first corrected position data; as well as The first training of the machine learning model is initiated based on the first training data.

2. The computer-implemented method according to claim 1, further comprising: Determine the offset between the first position data and the first corrected position data; It is determined that the offset exceeds the threshold offset; as well as The first training data includes at least one of the first image and the offset or the first correction location data, wherein the first training of the machine learning model is based on a loss function that minimizes the offset.

3. The computer-implemented method according to any one of claims 1 or 2, further comprising: The robot positions the camera relative to a second target included in the aircraft component based on the multidimensional representation of the aircraft component. A second image generated by the camera is received while the camera is positioned relative to the second target, the second image showing at least the second target; Prior to the first training, a second input to the machine learning model is generated based on the second image; Prior to the first training, a second output of the machine learning model is determined based on the second input, the second output indicating the second location data of a second detected target in the second image, the second detected target corresponding to the second target; The offset between the second position data of the second target determined from the second image and the second corrected position data of the detected second correction; Determine that the offset is less than the threshold offset; as well as The second image, the offset, and the second corrected position data are excluded from the first training data.

4. The computer-implemented method according to any one of claims 1 to 3, further comprising: The robot positions the camera relative to a second target included in the aircraft component based on the multidimensional representation of the aircraft component. A second image generated by the camera is received while the camera is positioned relative to the second target, the second image showing at least the second target; The second image is used to generate a second input to the machine learning model; The second output of the machine learning model is determined based on the second input, and the second output indicates the third location data of the second detected target in the second image, the second detected target corresponding to the second target; The attitude data of the second target is generated based on the third position data; as well as An updated multidimensional representation of the aircraft component is generated by updating the multidimensional representation based at least on the attitude data.

5. The computer-implemented method according to claim 4, further comprising: The robot performs the robotic operation on the second target by using an end effector configured for robotic operation, based on the updated multidimensional representation of the aircraft parts.

6. The computer-implemented method according to any one of claims 1 to 5, further comprising: Synthetic images are generated based on the multidimensional representation of the aircraft parts, each synthetic image showing at least one modeled target corresponding to at least one target of the aircraft parts; The second training data is generated based on the synthesized image; as well as Prior to the first training, the machine learning model is subjected to a second training based on the second training data.

7. The computer-implemented method of claim 6, wherein the synthesized image is generated by applying an image transformation to the modeled target.

8. The computer-implemented method according to claim 6, further comprising: The third position data of the target being modeled is determined based on the multidimensional representation of the aircraft parts. as well as A synthetic image showing the modeled target and the third location data is included in the second training data.

9. The computer-implemented method according to any one of claims 1 to 8, further comprising: The robot positions the camera relative to a second target included in the aircraft component based on the multidimensional representation of the aircraft component. A second image generated by the camera is received while the camera is positioned relative to the second target, the second image showing at least the second target; The second image is used to generate a second input to the machine learning model; A second output of the machine learning model is determined based on the second input, the second output indicating a first classification of a second detected target in the second image, the second detected target corresponding to the second target; Determine the classification of the correction for the second target; as well as The second image and the corrected classification are included in the first training data, wherein the first training includes the classification using the correction.

10. The computer-implemented method according to any one of claims 1 to 9, further comprising: The robot positions the camera relative to a second target included in the aircraft component based on the multidimensional representation of the aircraft component. A second image generated by the camera is received while the camera is positioned relative to the second target, the second image showing at least the second target; The second image is used to generate a second input to the machine learning model; A second output of the machine learning model is determined based on the second input, the second output indicating a first visual inspection attribute of a second detected target in the second image, the second detected target corresponding to the second target; Determine the visual inspection attributes of the correction for the second target; as well as The second image and the corrected visual inspection attributes are included in the first training data, wherein the first training includes using the corrected visual inspection attributes.

11. The computer-implemented method according to any one of claims 1 to 10, wherein the first target comprises a fastener or a fastener hole, and wherein the machine learning model is trained to classify the target as a fastener type, determine the target location, and determine the target visual inspection attributes.

12. The computer-implemented method according to any one of claims 1 to 11, wherein the first location data indicates a bounding box surrounding the first detected target.

13. A system comprising: One or more processors; and One or more memories, the one or more memories storing instructions that, when executed by the one or more processors, configure the system to: The robot is positioned relative to a first target camera included in the aircraft parts by using a multidimensional representation of the aircraft parts. A first image generated by the camera is received when the camera is positioned relative to the first target, the first image showing at least the first target; The first image is used to generate the first input to the machine learning model; The first output of the machine learning model is determined based on the first input, and the first output indicates the first location data of a first detected target in the first image, the first detected target corresponding to the first target. The first corrected position data of the detection of the first target is determined from the first image; First training data is generated based on the first corrected position data; as well as The first training of the machine learning model is initiated based on the first training data.

14. The system of claim 13, wherein the first target comprises a fastener or a fastener hole, and wherein the machine learning model is trained to classify the target as a fastener type, determine the target location, and determine the target visual inspection attributes.

15. The system of any one of claims 13 or 14, wherein the one or more memories store further instructions that, when executed by the one or more processors, configure the system to: The robot positions the camera relative to a second target included in the aircraft component based on the multidimensional representation of the aircraft component. A second image generated by the camera is received while the camera is positioned relative to the second target, the second image showing at least the second target; The second image is used to generate a second input to the machine learning model; A second output of the machine learning model is determined based on the second input, the second output indicating a first classification of a second detected target in the second image, the second detected target corresponding to the second target; Determine the classification of the correction for the second target; as well as The second image and the corrected classification are included in the first training data, wherein the first training includes the classification using the correction.

16. The system according to any one of claims 13 to 15, wherein the one or more memories store further instructions, which, when executed by the one or more processors, configure the system to: The robot positions the camera relative to a second target included in the aircraft component based on the multidimensional representation of the aircraft component. A second image generated by the camera is received while the camera is positioned relative to the second target, the second image showing at least the second target; The second image is used to generate a second input to the machine learning model; A second output of the machine learning model is determined based on the second input, the second output indicating a first visual inspection attribute of a second detected target in the second image, the second detected target corresponding to the second target; Determine the visual inspection attributes of the correction for the second target; as well as The second image and the corrected visual inspection attributes are included in the first training data, wherein the first training includes using the corrected visual inspection attributes.

17. The system according to any one of claims 13 to 16, wherein the one or more memories store further instructions, which, when executed by the one or more processors, configure the system to: Determine the offset between the first position data and the first corrected position data; It is determined that the offset exceeds the threshold offset; as well as The first training data includes at least one of the first image and the offset or the first correction location data, wherein the first training of the machine learning model is based on a loss function that minimizes the offset.

18. One or more non-transitory computer-readable storage media, said non-transitory computer-readable storage media storing instructions, said instructions, when executed on a system, cause the system to perform operations including: The robot is positioned relative to a first target camera included in the aircraft parts by using a multidimensional representation of the aircraft parts. A first image generated by the camera is received when the camera is positioned relative to the first target, the first image showing at least the first target; The first image is used to generate the first input to the machine learning model; The first output of the machine learning model is determined based on the first input, and the first output indicates the first location data of a first detected target in the first image, the first detected target corresponding to the first target. The first corrected position data of the detection of the first target is determined from the first image; First training data is generated based on the first corrected position data; as well as The first training of the machine learning model is initiated based on the first training data.

19. One or more non-transitory computer-readable storage media according to claim 18, wherein the operation further comprises: The robot positions the camera relative to a second target included in the aircraft component based on the multidimensional representation of the aircraft component. A second image generated by the camera is received while the camera is positioned relative to the second target, the second image showing at least the second target; Prior to the first training, a second input to the machine learning model is generated based on the second image; Prior to the first training, a second output of the machine learning model is determined based on the second input, the second output indicating the second location data of a second detected target in the second image, the second detected target corresponding to the second target; The offset between the second position data of the second target determined from the second image and the second corrected position data of the detected second correction; Determine that the offset is less than the threshold offset; as well as The second image, the offset, and the second corrected position data are excluded from the first training data.

20. One or more non-transitory computer-readable media according to any one of claims 18 or 19, wherein the operation further comprises: Synthetic images are generated based on the multidimensional representation of the aircraft parts, each synthetic image showing at least one modeled target corresponding to at least one target of the aircraft parts; The second training data is generated based on the synthesized image; as well as Prior to the first training, the machine learning model is subjected to a second training based on the second training data.