Cross-domain metric learning system and method

By using cross-domain deep metric learning algorithms and convolutional neural networks, the problem of identifying invalid steps in augmented reality systems was solved, enabling accurate identification and correction of user operations and improving the system's workflow execution efficiency.

CN112329510BActive Publication Date: 2025-12-12ROBERT BOSCH GMBH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202010772017.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-08-05
Filing Date
2020-08-04
Publication Date
2025-12-12
Estimated Expiration
2040-08-04

AI Technical Summary

Technical Problem

Existing machine learning algorithms struggle to effectively identify and correct invalid user actions in augmented reality systems, especially in complex work procedures such as vehicle repair step identification and detection.

Method used

A cross-domain deep metric learning algorithm is adopted. A two-dimensional RGB image is processed by a convolutional neural network (CNN) to generate image features in the semantic space. A triple loss algorithm is used to train the CNN to reduce the distance between positive vectors and anchor vectors and increase the distance between anchor vectors and negative vectors. By combining Siamese neural network and deep metric learning, invalid steps of user operations can be detected.

Benefits of technology

This improves the accuracy of augmented reality systems in recognizing and correcting user actions, ensuring that users follow the correct repair sequence and reducing the occurrence of invalid steps.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112329510B_ABST
    Figure CN112329510B_ABST
Patent Text Reader

Abstract

Cross-domain metric learning systems and methods are disclosed. An augmented reality (AR) system and method can include a controller operable to process one or more convolutional neural networks (CNNs) and a visualization device operable to acquire one or more 2-D RGB images. In response to providing an anchor image to a first convolutional neural network (CNN), the controller can generate an anchor vector in a semantic space. The anchor image can be one of the 2-D RGB images. In response to providing a negative image and a positive image to a second CNN, the controller can generate a positive vector and a negative vector in the semantic space. The negative image and the positive image can be provided as 3-D CAD images. The controller can apply a cross-domain deep metric learning algorithm operable to extract image features in the semantic space using the anchor vector, the positive vector, and the negative vector.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The following generally relates to machine learning augmented reality (AR) systems and methods. BACKGROUND

[0002] Machine learning (ML) algorithms are generally operable to use patterns and inferences to carry out a given task. Machine learning algorithms are generally based on mathematical models that rely on "training data" to make predictions or decisions without being explicitly programmed to carry out the requested task. SUMMARY

[0003] In one embodiment, an augmented reality system (AR system) and method are disclosed that can include a controller operable to process one or more convolutional neural networks (CNNs) and a visualization device operable to acquire one or more two-dimensional RGB images. The controller can also generate an anchor vector in a semantic space in response to providing an anchor image to a first convolutional neural network (CNN). The anchor image can be one of the two-dimensional RGB images acquired by the visualization device. The controller can also generate a positive vector and a negative vector in the semantic space in response to providing a negative image and a positive image to a second CNN. The negative image can be a first three-dimensional computer-aided design (CAD) image and the positive image can be a second three-dimensional CAD image. Both the first CAD image and the second CAD image can be provided to the AR system from a database.

[0004] It is contemplated that the first CNN and the second CNN can include one or more second convolutional layers, one or more second max-pooling layers, a second flattening layer, a second dropout layer, and a second fully connected layer. It is also contemplated that the controller can apply a cross-domain deep metric learning algorithm that is operable to extract image features in the semantic space using the anchor image, the positive image, and the negative image. It is further contemplated that the controller can be operable to extract one or more image features from different modalities using the anchor vector, the positive vector, and the negative vector.

[0005] The cross-domain deep metric learning algorithm can be implemented as a triple loss algorithm that is operable to reduce a first distance between the anchor vector and the positive vector in the semantic space and increase a second distance between the anchor vector and the negative vector in the semantic space. Moreover, the convolutional layers included within the first CNN and the second CNN can be implemented using one or more activation functions that can include linear rectifier units.

[0006] It is contemplated that the first CNN and the second CNN can employ a skip-connection architecture. The second CNN can also be designed as a Siamese network. The controller can perform step recognition by analyzing the image features extracted in the semantic space. Finally, the controller can also be operable to determine whether an invalid repair sequence has occurred based on the analysis of the image features in the semantic space.

[0007] It is contemplated that the controller can be further operable to determine a pose of an image object within the one or more RGB images. The controller can also apply a post-processing image algorithm to the one or more RGB images. BRIEF DESCRIPTION OF DRAWINGS

[0008] Figure 1 is an exemplary augmented reality (AR) system according to one embodiment;

[0009] Figure 2A and Figure 2B is an exemplary illustration of the learning process of a triple loss function applied by a deep learning network;

[0010] Figure 3 is an exemplary illustration of a convolutional neural network operating during a training phase;

[0011] Figure 4 is an exemplary illustration of a convolutional neural network operating during a testing or runtime phase;

[0012] Figure 5 is an illustration of an RGB image processed during an invalid step detection process; and

[0013] Figure 6 is an illustration of a bounding box identifying a part from an RGB image. DETAILED DESCRIPTION

[0014] Detailed embodiments are disclosed herein; however, it is understood that the disclosed embodiments are merely examples and can be carried out in various forms and alternatives. The drawings are not necessarily to scale; some features can be exaggerated or minimized for clarity. The specific structural and functional details disclosed herein are not to be interpreted as limiting, but merely as a representative basis for teaching one skilled in the art to variously employ the embodiments.

[0015] It is contemplated that a vision-based AR system can be desirable that is operable to identify different states or steps of a work procedure (e.g., the correct steps required to repair a vehicle). The vision-based AR system can be operable to use a 3-dimensional (3D) model of the entire procedure (e.g., a 3D model of the vehicle and a subset of the vehicle parts) to identify the different steps. It is contemplated that the AR system can employ a machine learning algorithm that receives input data from different domains and encodes the input data into a high-dimensional feature vector. The machine learning algorithm can also be operable to transform the encoded features into a semantic space for distance measurement.

[0016] Figure 1 An AR system 100 is illustrated in accordance with one example embodiment. The client system 110 can include a visualization device 112 that can be operable to capture one or more two-dimensional RGB images 113 (i.e., two-dimensional color images) when worn by a user. For example, the visualization device 112 can include a pair of AR glasses having an integrated digital camera or video camera (e.g., a DSLR camera or a mirrorless digital camera) that is operable to capture the RGB images 113.

[0017] The client system 110 can further include a controller 114 and a memory 116. The controller 114 can be one or more computing devices, such as a quad-core processor for processing commands, such as a computer processor, microprocessor, or any other device, device series, or other mechanism capable of carrying out the operations discussed herein. The memory 116 can be operable to store instructions and commands. The instructions can take the form of software, firmware, computer code, or some combination thereof. The memory 116 can take any form of one or more data storage devices, such as volatile memory, non-volatile memory, electronic memory, magnetic memory, optical memory, or any other form of data storage device. The memory 116 can be an internal piece of the client system 110 (e.g., DDR memory), or the memory can include a removable memory component (e.g., a microSD card memory).

[0018] The RGB images 113 can be captured by the visualization device 112 for further processing by the controller 114. For example, the controller 114 can be operable to determine the position and orientation (i.e., pose) of the objects captured by the RGB images 113 using a vision-based odometry algorithm, a simultaneous localization and mapping (SLAM) algorithm, and / or a model-based tracking algorithm.

[0019] The client system 110 can connect to and communicate with the server system 120. For example, the client system 110 can include a client transceiver 118 operable to transmit and receive data using a wired network (e.g., a LAN) or using a wireless network communication (e.g., WiFi, cellular, Bluetooth, or Zigbee). The server system 120 can include a server transceiver 122 also operable to transmit and receive data using a wired network (e.g., a LAN) or using a wireless network communication (e.g., WiFi, cellular, Bluetooth, or Zigbee). It is contemplated that the client transceiver 118 can be operable to transmit data, such as the RGB image 113 and the pose data, to the server transceiver 122. However, it is further contemplated that the client system 110 and the server system 120 can be included in one unit, thus not requiring the client transceiver 118 and the server transceiver 122.

[0020] It is contemplated that the server system 120 can also include a memory 126, such as the memory 116, for storing the pose data and the RGB image 113 received from the client system 110. The server system 120 can also include a controller 124, such as the controller 114, operable to perform post-processing algorithms on the RGB image 113. The server system 120 can also connect to and communicate with a database 128 for storing 3-dimensional computer aided design (3D-CAD) models (i.e., 3D images) of target work procedures. The server system 120 can communicate with the database 128 using the server transceiver 122. It is also contemplated that the database 128 can be included within the server system 120, thus not requiring the server transceiver 122. The controller 124 can be operable to apply computer graphics rendering algorithms to generate one or more normal map images of the 3D-CAD models based on the view defined by the pose data received from the client system 110, which can be compared to the RGB image 113. The controller 124 can also be operable to apply post-processing image algorithms to the RGB image 113 received from the client system 110.

[0021] The controller 124 can also be operable to apply one or more deep neural network (DNN) algorithms. It is contemplated that the controller 124 can use the normal map images and the RGB image 113 to apply a DNN state prediction algorithm to predict a state (e.g., a current step) of the work procedure. The server system 120 can send instructions back to the client system 110, which can be visually displayed by the visualization device 112.

[0022] It is contemplated that the DNN can be a convolutional neural network (CNN) that includes convolutional layers and pooling layers for handling image recognition applications. The CNN can be a two-branch CNN operable to obtain a bilinear feature vector for fine-grained classification. The CNN can also be operable to learn appropriate features to compare a fine-grained class of a given image. It is also contemplated that the machine learning algorithm can be operable to capture small changes in important regions of interest (ROIs) and avoid noisy background images, lighting, and viewpoint changes. It is further contemplated that the controller 124 can be operable to employ deep metric learning (e.g., cross-domain deep metric learning) to further improve the performance of the CNN by not constraining the recognition process to only procedures provided during a training phase.

[0023] For example, the controller 124 can be operable to employ a CNN to determine a current step being performed by the user and provide instructions for the next step that the user will need to perform (i.e., “step recognition”). The controller 124 can also be operable to employ a CNN to detect whether the user’s actions deviate from a pre-stored sequence of steps (i.e., “invalid step detection”). Upon detecting that the user has performed an invalid step, the controller 124 can send instructions to the client system 110 to inform the user how to correct any incorrect actions that have been taken. It is contemplated that the controller 124 can be operable to perform “step recognition” and “invalid step detection” using the 3D-CAD model stored in the database 128 that corresponds to the target work procedure being evaluated.

[0024] It is contemplated that the CNN can be designed as a cross-domain deep metric learning algorithm that is trained to determine a similarity measure between the 2-D RGB images 113 acquired by the visualization device 112 and the 3D-CAD data stored in the database 128. By training the CNN, the controller 124 can be operable to compare the similarity or distance between the original data (i.e., the original 2-D images and 3-D CAD models) and the data transformed to semantic space using a triple loss algorithm, which is represented by Equation (1) below:

[0025]

[0026] wherein, is the computed triple loss function, is a branch of the CNN feature encoding of the RGB images 113, and is a branch of the triple loss function that extracts features from normal map images stored in the database 128 as 3D CAD models. It is contemplated that the triple loss function can operate on a triple feature vector that includes an anchor image, a positive image, and a negative image.

[0027] Figure 2A and Figure 2B are exemplary illustrations of how a triple loss algorithm can be used to train a CNN. Again, a triple image can be generated that includes an anchor image 210, a positive image 220, and a negative image 230. Figure 2A The distance from the positive image 220 to the anchor image 210 is illustrated as being greater than the distance from the negative image 230 to the anchor image 210. As Figure 2B indicated, a CNN can apply a triple loss function (Equation 1) to reduce the vector distance between the anchor image 210 and the positive image 220, and to increase the distance of the anchor image 210 and the negative image 230.

[0028] Figure 3 An anchor image 310, a positive image 320, and a negative image 330 that can be used during a triple loss algorithm are illustrated. As illustrated, one of the RGB images 113 acquired by the client system 110 can be assigned as the anchor image 310. The CNN can then receive the positive image 320 and the negative image 330 from a pre-stored normal map image that is generated from a 3D-CAD model stored in the database 128. The CNN algorithm can be trained to extract features from the anchor image 310 and the normal map image (i.e., the positive image 320 and the negative image 330) in the semantic space.

[0029] It is also contemplated that the controller 124 can be operable to implement a two-branch CNN that extracts high-dimensional features from the RGB images 113 acquired by the client system 110 and the normal map images stored as 3D-CAD models in the database 128. It is contemplated that each branch of the CNN can include one or more convolutional layers with kernels of various sizes (e.g., 3x3 or 5x5). The number and size of the kernels can be adjusted depending on the given application.

[0030] It is also contemplated that the CNN can include one or more convolutional layers that employ linear rectifier units (ReLU) or tanh as activation functions, one or more normalization layers that are operable to improve the performance and stability of the CNN, and max-pooling layers that can reduce the dimensionality. The CNN can include one or more stacked modules based on the complexity of the target data distribution. The CNN can also employ a skip-connection architecture in which the output of one layer of the CNN can not be directly input to the next sequential layer of the CNN. Instead, the skip-connection architecture can allow the output of one layer to be connected to the input of a non-sequential layer.

[0031] Table 1 below illustrates an exemplary architecture that can be used by the CNN for one branch that can include a design with 8,468,784 parameters.

[0032] Layer In size Out size Kernel size Activation Convolution 1 256x256x3 256x256x16 3x3x16 Relu Max pooling 256x256x16 128x128x16 - - Convolution 2 128x128x16 128x128x32 3x3x32 Relu Convolution 3 128x128x32 128x128x32 3x3x32 Relu Max pooling 128x128x32 64x64x32 - - Convolution 4 64x64x32 64x64x48 3x3x48 Relu Convolution 5 64x64x48 64x64x48 3x3x48 Relu Max pooling 64x64x48 32x32x48 - - Convolution 6 32x32x48 32x32x64 3x3x64 Relu Max pooling 32x32x64 16x16x64 - - Flatten 16x16x64 16384 - - Dropout - - - - Dense 16384 128 - -

[0033] Table 1.

[0034] As illustrated, the CNN can include one or more convolutional layers, one or more max-pooling layers, a flatten layer, a dropout layer, and a dense layer (i.e., a fully connected or linear layer). The size of the data input into each layer of the CNN (i.e.,“Size-in”) and the corresponding size of the data output by a given layer of the CNN (i.e.,“Size-out”). It is contemplated that the kernel size of each convolutional layer can vary, but it is also contemplated that the kernel size can be the same depending on the given application. The CNN can also employ a ReLU activation function for each convolutional layer. However, it is contemplated that the activation function can vary based on the given application, and the CNN can employ other known activation functions (e.g., a tanh activation function).

[0035] Figure 4 A two-branch CNN 400 that can be employed during the training phase is illustrated. It is contemplated that the CNN 400 can be designed according to the illustrated architecture. It is also contemplated that the number of layers, the size of each layer, and the activation function can vary depending on the given application. Figure 4

[0036] The CNN 400 can include an RGB network 410 branch that can receive an anchor image 420 as input data (e.g., the RGB image 113). The RGB network 410 can apply a function f RGB which represents a branch of the CNN used to feature encode one of the RGB images 113. The CNN 400 can also include a branch with a normalization network 430 that can receive a positive image 440 and a negative image 450 (e.g., one of the 3D-CAD models stored in the database 128) as input data.

[0037] The output data of the RGB network 410 can be provided to RGB encoded features 460. Also, the output data of the normalization network 430 can be provided to positive encoded features 470 and negative encoded features 480. A triple loss algorithm 490 (discussed above with respect to Equation 1) can then be employed to decrease the vector distance between the anchor image 420 and the positive image 440, and increase the vector distance between the anchor image 420 and the negative image 450.

[0038] It is contemplated that the normalization network 430 can be implemented as a“twin neural network” capable of operating in tandem to compute different vector distances of the provided positive image 440 and negative image 450. It is also contemplated that the RGB network 410 and the normalization network 430 can include different layers (i.e., convolutional layers, max-pooling layers, flatten layers, dropout layers, dense layers, and implementation layers) that employ different training data depending on the given application.​

[0039] Figure 5 A CNN 500 that can be employed during the testing phase or during runtime operations is illustrated. The CNN 500 can be designed again using a“twin neural network,” although it is contemplated that other network designs can be used depending on the given application. It is contemplated that the RGB network 510 and the normalization network 530 can include different variations of convolutional layers, max-pooling layers, flattening layers, dropout layers, and dense layers depending on the given application.

[0040] The RGB network 510 (i.e., f RGB ) can be provided the RGB image 520 as input. Again, the RGB image 520 can be the RGB image 113 received from the client system 110. The CNN 500 can also include a normalization network 530 (i.e., f normalization ) that receives one or more normalized images 540 extracted during the training phase. The RGB network 510 can provide output vector data to RGB-encoded features 550, and the normalization network 530 can provide output vector data to one or more normalized-encoded features 560, 570.

[0041] The CNN 500 can then use the output vector data provided by the RGB-encoded features 550 and the one or more normalized-encoded features 560, 570 to compute a set of distance vectors 580 from the feature vector of the normalized image 540 and the feature vector encoded within the RGB image 520. The CNN 500 can then select the distance vector with the smallest distance in the semantic space.

[0042] Again, it is also contemplated that the CNN can be operable to detect an“invalid” state, which indicates that the given step or sequence being performed by the user can not match any step in the procedure. The CNN can determine that the invalid state can be a result of an incorrect sequence of operations performed by the user. It is contemplated that by analyzing the RGB images 113 acquired by the client system 110, the CNN can be able to analyze the critical parts of the procedure to identify the invalid state.

[0043] Figure 6An RGB image 600 that can be acquired by the client system 110 and sent to the server system 120 is illustrated. The controller 124 can process the RGB image 600 to detect an unexpected part due to an invalid operation by the user. For example, the controller can apply an image cropping algorithm to use the RGB image 600 to generate a first bounding box 610, a second bounding box 620, a third bounding box 630, and a fourth bounding box 640 for the presence or absence of a part. Then, the CNN can determine whether an invalid state exists based on the presence or absence of the part in the RGB image 600. Alternatively, the CNN can determine whether the user performed an incorrect step based on the presence or absence of the part in the RGB image 600.

[0044] It is contemplated that a two-branch CNN network can be used to process the presence or absence of a part in the RGB image 600 using the presence or absence of the part in the RGB image 600. Figure 4 The two-branch CNN network is trained using the triple loss algorithm discussed. It is contemplated that the controller 124 can be operable to apply an image cropping algorithm that uses one or more 3D-CAD models stored in the database 128 to generate one or more 3D bounding boxes (e.g., the first bounding box 610) for things that can be preprogrammed as essential and critical parts. It is contemplated that when training and testing the CNN, the 3D bounding boxes can be projected to the image space based on the current pose. It is also contemplated that the CNN can use the two-dimensional projection of the 3D bounding boxes to define a region of interest in the RGB image 600 that can include certain parts. The parts can be generated by the CNN by cropping the RGB image 600 based on the predefined region of interest and rendering the cropped image.

[0045] It is also contemplated that the CNN can use one of the 3D bounding boxes (e.g., the first bounding box 610) as an anchor image. It is also contemplated that the anchor image and the negative and positive images provided by the database 128 are used to train the CNN 400 to employ an invalid step recognition process. During the training phase, the anchor image and the positive image can contain the same part, while the negative image can contain a different part than the anchor image. It is contemplated that the CNN 400 can be trained again using the anchor image, the positive image, and the negative image using the triple loss algorithm for determining when a part is detected or not present within a given region of interest.

[0046] Once trained, the CNN 500 can be employed during a testing phase or runtime operation to provide invalid or incorrect step detection. The testing or runtime operation can crop captured images (e.g., RGB images 113 received by the client system 110) into different parts. The CNN can then be operable to compare the parts in the images to corresponding areas in two or more normal map images. It is contemplated that one normal map image can include the part, while other normal images can not include the part. If the features extracted from the image are closer to the normal map image with the part, the CNN can determine that the part is detected. If the features extracted from the image are closer to the normal map image without the part, the CNN can determine that the part is not detected (i.e., none).

[0047] Table 2 below illustrates outputs that the CNN 500 can generate during the invalid step identification process.

[0048] Image Part 1 Part 2 Part 3 Part 4 Part 5 Predicted 1 Detected Detected No No No State 1 2 No No Detected Detected Detected State 2 3 No No No Detected Detected State 3 4 No No Detected No Detected Invalid state

[0049] Table 2.

[0050] As illustrated, the CNN 500 can employ the invalid step identification process on more than one image (e.g., image 1, image 2, image 3, image 4). For example, Table 2 illustrates that for “image 1,” the CNN 500 can generate the following outputs for “part 1” through “part 5”: [detected, detected, none, none, none]. Based on these outputs, the invalid step identification process can be operable to map the outputs to standard repair steps. As shown, for image 1, image 2, and image 3, the invalid step identification process determines that no invalid steps have occurred. In other words, for image 1 through image 3, the CNN 500 has determined that the user has followed the repair steps in the correct sequence.

[0051] However, for image “4,” the invalid step identification process has determined that an “invalid state” has occurred. The invalid state output can have been generated because, incorrectly, the user has performed the repair sequence incorrectly. For example, the fourth step in the repair sequence can require “part 1” to be detected. Since the CNN 500 did not detect “part 1” as being present, the invalid sequence was detected. Alternatively, the invalid state can also be generated if the user incorrectly attempts to re-add a part from the correct repair sequence. For example, for image 5, it can have been detected that “part 1” is not present because it is currently behind or blocked by an additional part (e.g., part 5). Since “part 1” is not visible, the invalid step identification process can still output “invalid state” because it is not until later in the repair sequence that “part 5” needs to be reassembled.

[0052] The processes, methods, or algorithms disclosed herein can be embodied in, and fully or partially automated by, processing devices, controllers, or computers that can include any existing programmable electronic control unit or dedicated electronic control units. Similarly, the processes, methods, or algorithms can be embodied in, and fully or partially automated by, a data storage medium containing instructions executable by a processing device, controller, or computer, a computer, or a dedicated electronic control unit. The data storage medium can be a write-able non-transitory data storage medium including a volatile or non-volatile memory device. The instructions can include any set of instructions that, when executed by a processing device, controller, or computer, are capable of causing the processing device, controller, or computer to perform a set of operations. The instructions can be stored in any memory device, volatile or non-volatile inside or outside a processing device, controller, or computer, including a RAM, a ROM, a hard disk, a floppy disk, a magnetic tape, an optical disc, a flash memory device, or a compact disc. Similarly, the processes, methods, or algorithms can be embodied in, and fully or partially automated by, a hardware logic circuit, such as a discrete component circuit, an ASIC, a FPGA, a state machine, a controller, or a computer, or a combination of the above.

[0053] While the example embodiments have been described above, it is not intended that these embodiments describe all possible forms of the application. Rather, the words used in the specification are words of description rather than limitation, and it is understood that various changes can be made without departing from the spirit and scope of the application. Additionally the features of various implementing embodiments can be combined to form further embodiments of the application.

Claims

1. A convolutional neural network (CNN) method for identifying an invalid state in a work procedure, comprising: generating an anchor vector in a semantic space in response to providing an anchor image to a first CNN, wherein the anchor image is a two-dimensional RGB image of a work procedure, the work procedure comprising one or more parts that are preprogrammed as key parts, wherein the first CNN comprises one or more first convolutional layers, one or more first max pooling layers, a first flattening layer, a first dropout layer, and a first fully connected layer; generating a positive vector and a negative vector in the semantic space in response to providing a negative image and a positive image to a second CNN, wherein the negative image is a first three-dimensional CAD image and the positive image is a second three-dimensional CAD image, wherein the second CNN comprises one or more second convolutional layers, one or more second max pooling layers, a second flattening layer, a second dropout layer, and a second fully connected layer; applying a cross-domain deep metric learning algorithm operable to extract image features in the semantic space using the anchor vector, the positive vector, and the negative vector; applying an image cropping algorithm to generate one or more bounding boxes for a presence or an absence of the one or more key parts of the work procedure using the anchor image; determining whether an invalid repair sequence has occurred based on an analysis of the presence or the absence of the one or more key parts of the work procedure; and outputting an invalid state if (1) a key part that is required to be detected in the repair sequence is not detected as present; or (2) a key part that is added back from a correct repair sequence behind or blocked by another key part.

2. The method of claim 1, wherein, The cross-domain deep metric learning algorithm is a triple loss algorithm operable to reduce a first distance between the anchor vector and the positive vector in the semantic space and increase a second distance between the anchor vector and the negative vector in the semantic space.

3. The method of claim 1, wherein, The one or more first convolutional layers and the one or more second convolutional layers are operable to apply one or more activation functions.

4. The method of claim 3, wherein, The one or more activation functions are implemented using a linear rectifier unit.

5. The method of claim 3, wherein, The first CNN and the second CNN further comprise one or more normalization layers.

6. The method of claim 1, wherein, The second CNN is designed using a siamese neural network.

7. The method of claim 1, wherein, The first CNN and the second CNN employ a skip connection architecture.

8. An augmented reality system for identifying an invalid state in a work procedure, comprising: a visualization device operable to acquire one or more RGB images of a work procedure, the work procedure comprising one or more parts that are preprogrammed as key parts; and a controller operable to, generate an anchor vector in a semantic space in response to providing an anchor image to a first CNN, wherein the anchor image is a two-dimensional RGB image, wherein the first CNN comprises one or more first convolutional layers, one or more first max pooling layers, a first flattening layer, a first dropout layer, and a first fully connected layer; generating, in response to providing the negative image and the positive image to a second CNN, the positive vector and the negative vector in a semantic space, wherein the negative image is a first three-dimensional CAD image and the positive image is a second three-dimensional CAD image, wherein the second CNN comprises one or more second convolutional layers, one or more second max-pooling layers, a second flattening layer, a second dropout layer, and a second fully connected layer; applying a cross-domain deep metric learning algorithm operable to extract image features in the semantic space using the anchor vector, the positive vector, and the negative vector; applying an image cropping algorithm to generate one or more bounding boxes for the presence or absence of one or more key parts of the work procedure using the anchor image; determining whether an invalid repair sequence has occurred based on the analysis of the presence or absence of the one or more key parts of the work procedure; and outputting an invalid state if (1) one key part required to be detected in the repair sequence is not detected as present; or (2) one key part that was removed from a correct repair sequence is re-added behind or blocked by another key part.

9. The augmented reality system of claim 8, wherein, the controller is further operable to determine a pose of an image object within the one or more RGB images.

10. The augmented reality system of claim 8, wherein, the controller is further operable to reduce a first distance between the anchor vector and the positive vector in the semantic space and increase a second distance between the anchor vector and the negative vector in the semantic space.

11. The augmented reality system of claim 8, wherein, the controller is further operable to apply a post-processing image algorithm to the one or more RGB images.

12. The augmented reality system of claim 8, wherein, the controller is further operable to display instructions to a visualization device based on a current step of the work procedure.

13. The augmented reality system of claim 8, wherein, the second CNN is designed using a siamese neural network.

14. An augmented reality method for identifying an invalid state in a work procedure, comprising: generating, in response to providing the anchor image to a first CNN, an anchor vector in a semantic space, wherein the anchor image is a two-dimensional RGB image of a work procedure, the work procedure comprising one or more parts that are preprogrammed as key parts, wherein the first CNN comprises one or more first convolutional layers, one or more first max-pooling layers, a first flattening layer, a first dropout layer, and a first fully connected layer; generating, in response to providing the negative image and the positive image to a second CNN, the positive vector and the negative vector in a semantic space, wherein the negative image is a first three-dimensional CAD image and the positive image is a second three-dimensional CAD image, wherein the second CNN comprises one or more second convolutional layers, one or more second max-pooling layers, a second flattening layer, a second dropout layer, and a second fully connected layer; extracting one or more image features from different modalities using the anchor vector, the positive vector, and the negative vector; applying an image cropping algorithm to generate one or more bounding boxes for the presence or absence of one or more key parts of the work procedure using the anchor image; determining whether an invalid repair sequence has occurred based on the analysis of the presence or absence of the one or more key parts of the work procedure; and outputting an invalid state if (1) one key part required to be detected in the repair sequence is not detected as present; or (2) one key part that was removed from a correct repair sequence is re-added behind or blocked by another key part. If (1) one of the critical parts required to be detected in the repair sequence is not detected as present; or (2) a critical part that was removed from the correct repair sequence is re-added behind or blocked by another critical part, then an invalid state is output.

15. The method of claim 14, further comprising: A triple loss algorithm is applied, which is operable to reduce a first distance between an anchor vector and a positive vector in a semantic space, and to increase a second distance between the anchor vector and a negative vector in the semantic space.

Citation Information

Patent Citations

  • System and method for three-dimensional augmented reality guidance for use of medical equipment

    US20180225993A1