A 3D inspection method for augmented reality-assisted assembly
By adopting a three-dimensional inspection method in the augmented reality-assisted assembly system, using the cross-domain image registration network model for texture registration and spatial pose registration, the problem of insufficient automation in the existing system is solved, and a more efficient assembly process and more accurate inspection results are achieved.
Patent Information
- Application Number
- CN202411559121.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-04
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-11-04
AI Technical Summary
The existing augmented reality assisted assembly system has insufficient automation, which leads to frequent voice or gesture interactions during the assembly process, distracting attention and reducing assembly efficiency.
A three-dimensional inspection method is adopted to obtain real camera images and three-dimensional digital models, and use the cross-domain image registration network model for texture registration and spatial pose registration to calculate the assembly score, thereby judging the completion status of the assembly and automatically jumping to the next step to guide.
It improves the automation level of augmented reality assisted assembly systems, reduces the interactive needs of operators, improves assembly efficiency, and enhances the accuracy of inspection results.
Smart Images

Figure CN119399177B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and particularly to a three-dimensional inspection method for augmented reality-assisted assembly. Background Art
[0002] In augmented reality-assisted assembly, virtual assembly instructions are placed in a real assembly environment using augmented reality technology to assist an operator in completing the assembly task step by step.
[0003] Currently, in an augmented reality-assisted assembly system, a guidance program is controlled by an operator's voice or gesture interaction.
[0004] However, this will distract the operator's attention and reduce the assembly efficiency. Therefore, an inspection method is needed to determine the completion status of an assembled body, that is, to determine whether the assembled body is correctly assembled after each assembly operation. Summary of the Invention
[0005] An embodiment of this application provides a three-dimensional inspection method for augmented reality-assisted assembly, which can solve the technical problem of poor automation in existing augmented reality-assisted assembly.
[0006] To achieve the above object, the embodiments of this application adopt the following technical solutions:
[0007] In a first aspect, an embodiment of the present application provides a three-dimensional inspection method for augmented reality-assisted assembly. The three-dimensional inspection method for augmented reality-assisted assembly includes: obtaining a real camera image and a three-dimensional digital model of an assembly to be inspected in the same state; obtaining a virtual rendering image by virtual camera rendering according to the three-dimensional digital model; inputting the real camera image and the virtual rendering image into a cross-domain image registration network model to train the cross-domain image registration network model; wherein, the cross-domain image registration network model includes a first matching layer, a second matching layer, a third matching layer, and a fourth matching layer; the first matching layer is configured to input a first-size feature map of the real camera image and a first-size feature map of the virtual rendering image to obtain a first deconvolution result; the second matching layer is configured to input a second-size feature map of the real camera image, a second-size feature map of the virtual rendering image, and the first deconvolution result to obtain a second deconvolution result; the third matching layer is configured to input a third-size feature map of the real camera image, a third-size feature map of the virtual rendering image, and the second deconvolution result to obtain a third deconvolution result; the fourth matching layer is configured to input a fourth-size feature map of the real camera image, a fourth-size feature map of the virtual rendering image, and the third deconvolution result to obtain a fourth deconvolution result; the fourth deconvolution result is used to obtain an estimated optical flow result; the estimated optical flow result is used for spatial pose registration; the size of the first-size feature map is smaller than the size of the second-size feature map; the size of the second-size feature map is smaller than the size of the third-size feature map; the size of the third-size feature map is smaller than the size of the fourth-size feature map; inputting an image to be inspected of the assembly to be inspected into the cross-domain image registration network model to match a first virtual image; obtaining a texture registration result according to the image to be inspected and the first virtual image; and obtaining a texture similarity score according to the texture registration result; the texture registration result includes the estimated optical flow result; performing spatial pose registration on the assembly to be inspected and the three-dimensional digital model according to the texture registration result to obtain a spatial pose similarity score; calculating an assembly score according to the texture similarity score and the spatial pose similarity score.
[0008] Based on the above description of the three-dimensional inspection method for augmented reality-assisted assembly provided by the embodiment of the present application, it can be seen that the three-dimensional inspection method for augmented reality-assisted assembly includes training a cross-domain image registration network model, inputting an image to be inspected into the cross-domain image registration network model to obtain a texture registration result, and obtaining a spatial pose registration result of the image to be inspected. Furthermore, an assembly score is obtained through calculation according to the texture similarity score and the spatial pose similarity score. Through the assembly score, it is possible to obtain the result of whether the operator has assembled correctly, so that the augmented reality-assisted assembly system can automatically jump to the next guiding program when the operator assembles correctly without any instruction from the operator, improving the automation degree of augmented reality-assisted assembly.
[0009] Moreover, the assembly situation is inspected by integrating the texture registration result and the spatial pose registration result. They are not simply spliced together. The spatial pose registration is carried out on the basis of the texture registration result. In this way, the accuracy of the inspection result is improved.
[0010] Furthermore, by taking into account two-dimensional features and three-dimensional features, the convenience of actual operation is improved.
[0011] In addition, it not only has high accuracy, does not rely on 3D technology, has a small pose estimation error, takes a short time, and is applicable to various industrial assembly scenarios.
[0012] In a feasible implementation manner of the first aspect, the first matching layer includes a coarse correlation layer and a coarse decoder; the coarse correlation layer is built according to the matching layer; the coarse decoder is built according to the mapping decoder; the three-dimensional inspection method for augmented reality-assisted assembly further includes: sequentially inputting the first-size feature map of the real camera image and the first-size feature map of the virtual rendering image into the coarse correlation layer and the coarse decoder to obtain a first optical flow result; deconvolving the first optical flow result to obtain a first deconvolution result.
[0013] In a feasible implementation manner of the first aspect, the second matching layer includes a deformation layer, a fine correlation layer, and a fine decoder; the deformation layer is used to deform the feature map by the bilinear interpolation method; the fine correlation layer is built according to the cost volume layer; the fine decoder is built according to the optical flow estimator; the three-dimensional inspection method for augmented reality-assisted assembly further includes: inputting the second-size feature map of the virtual rendering image and the first deconvolution result into the deformation layer to obtain a first deformation result; inputting the second-size feature map of the real camera image and the first deformation result into the fine correlation layer to obtain a first cost volume estimation result; inputting the first cost volume estimation result and the first deconvolution result into the fine decoder to obtain a first fine decoding result; obtaining a second optical flow result according to the first fine decoding result and the first deconvolution result; deconvolving the second optical flow result to obtain a second deconvolution result.
[0014] In a feasible implementation manner of the first aspect, the assembly to be inspected includes a first component and a second component; the three-dimensional inspection method for augmented reality-assisted assembly further includes: performing image segmentation on the real camera image and the virtual rendering image to obtain an image segmentation mask of the first component and an image segmentation mask of the second component; inputting the image segmentation mask of the first component and the image segmentation mask of the second component into the cross-domain image registration network model to obtain an estimated optical flow result.
[0015] In a feasible implementation of the first aspect, the three-dimensional inspection method for augmented reality-assisted assembly further includes: obtaining a first dense registration point pair according to the estimated optical flow result; the first dense registration point pair is a two-dimensional to two-dimensional dense registration point pair corresponding to the image to be inspected and the first virtual image; obtaining a second dense registration point pair according to the preset rendering parameters and the first dense registration point pair; the second dense registration point pair is a two-dimensional to three-dimensional dense registration point pair corresponding to the image to be inspected and the three-dimensional digital model; the second dense registration point pair includes multiple pairs; randomly select some of the second dense registration point pairs and iteratively calculate the six-dimensional pose; based on the six-dimensional pose, obtain the displacement similarity and the rotation similarity to obtain the spatial pose similarity score.
[0016] In a feasible implementation of the first aspect, the calculation formula for displacement similarity includes:
[0017]
[0018] where N represents the number of components; represents the estimated translation value of component i on the τ (X / Y / Z) axis; represents the relative translation value of components i and j in the three-dimensional digital model on the τ (X / Y / Z) axis; ΔT i represents the average displacement error of each component i, that is, the displacement similarity, which is determined by the average value of the displacement errors on the X, Y, and Z axes.
[0019] In a feasible implementation of the first aspect, the calculation formula for rotation similarity includes:
[0020]
[0021] where N represents the number of components, represents the estimated rotation value of the component on the axis, represents the estimated rotation value of component i on the τ (X / Y / Z) axis, represents the relative rotation value of components i and j in the three-dimensional digital model on the τ (X / Y / Z) axis, ΔR i represents the average rotation error of component i, that is, the rotation similarity, which is determined by the average value of the rotation errors on the X, Y, and Z axes.
[0022] In a feasible implementation of the first aspect, the calculation formula for texture similarity includes:
[0023]
[0024] where, for each component i, randomly select I matching point pairs and calculate the similarity within their L neighborhoods, N and respectively represent the pixel points within the neighborhood range of the matching point pairs in the original real camera image and the deformed real camera image.
[0025] In a second aspect, an embodiment of the present application provides a three-dimensional inspection system for augmented reality-assisted assembly. The three-dimensional inspection system for augmented reality-assisted assembly includes: at least one processor; a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method provided in the first aspect.
[0026] The three-dimensional inspection system for augmented reality-assisted assembly trains a cross-domain image registration network model by executing the method provided in the first aspect, inputs the image to be inspected into the cross-domain image registration network model to obtain a texture registration result, and obtains a spatial pose registration result of the image to be inspected. Furthermore, according to the texture similarity score and the spatial pose similarity score, an assembly score is obtained through calculation. Whether the operator assembles correctly can be obtained through the assembly score, so that the augmented reality-assisted assembly system can automatically jump to the next guiding program when the operator assembles correctly without any instruction from the operator, improving the automation degree of augmented reality-assisted assembly.
[0027] In a third aspect, an embodiment of the present application provides a computer-readable medium, on which computer program instructions are stored, and the computer program instructions can be executed by a processor to implement the method provided in the first aspect.
[0028] The computer program instructions in the computer-readable medium train a cross-domain image registration network model by implementing the method provided in the first aspect, input the image to be inspected into the cross-domain image registration network model to obtain a texture registration result, and obtain a spatial pose registration result of the image to be inspected. Furthermore, according to the texture similarity score and the spatial pose similarity score, an assembly score is obtained through calculation. Whether the operator assembles correctly can be obtained through the assembly score, so that the augmented reality-assisted assembly system can automatically jump to the next guiding program when the operator assembles correctly without any instruction from the operator, improving the automation degree of augmented reality-assisted assembly. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 It is a schematic structural diagram of a three-dimensional inspection system for augmented reality-assisted assembly provided by an embodiment of the present application;
[0030] Figure 2 It is a schematic flowchart of a three-dimensional inspection method for augmented reality-assisted assembly provided by an embodiment of the present application;
[0031] Figure 3 It is a schematic structural diagram of a cross-domain image registration network in a three-dimensional inspection method for augmented reality-assisted assembly provided by an embodiment of the present application;
[0032] Figure 4 Schematic diagram of the cross - domain registration dataset of the cross - domain image registration network in a 3D inspection method for augmented reality - assisted assembly provided by an embodiment of the present application;
[0033] Figure 5 Schematic diagram of an implementation manner of the texture registration result in a 3D inspection method for augmented reality - assisted assembly provided by an embodiment of the present application;
[0034] Figure 6 Schematic diagram of an implementation manner of the spatial pose registration in a 3D inspection method for augmented reality - assisted assembly provided by an embodiment of the present application;
[0035] Figure 7 Schematic diagram of an implementation manner of the six - dimensional pose estimation result of the spatial pose registration in a 3D inspection method for augmented reality - assisted assembly provided by an embodiment of the present application;
[0036] Figure 8 Schematic diagram of the visualization interface of an implementation manner of the augmented reality inspection of each component in a 3D inspection method for augmented reality - assisted assembly provided by an embodiment of the present application;
[0037] Figure 9 Schematic diagram of the visualization interface of an implementation manner of the augmented reality inspection of each step component during the assembly process in a 3D inspection method for augmented reality - assisted assembly provided by an embodiment of the present application;
[0038] Figure 10 Three - view drawings of the incorrect assembly situation of the component;
[0039] Figure 11 Three - view drawings of the correct assembly template of the component;
[0040] Figure 12 Schematic diagram of an implementation manner of the incorrect assembly inspection result in a 3D inspection method for augmented reality - assisted assembly provided by an embodiment of the present application;
[0041] Figure 13 Schematic diagram of the visualization interface of an implementation manner of the error reminder in a 3D inspection method for augmented reality - assisted assembly provided by an embodiment of the present application;
[0042] Figure 14 Step - by - step schematic diagram for assisting in assembling the building block fox in a 3D inspection method for augmented reality - assisted assembly provided by an embodiment of the present application. Detailed implementation manners
[0043] The technical solutions in the embodiments of the present invention will be described below with reference to the accompanying drawings in the embodiments of the present invention. Among them, in the description of the embodiments of the present invention, unless otherwise specified, "a plurality of" means two or more than two. "At least one (item)" or its similar expression refers to any combination of these items, including any combination of a single item or plural items. For example, at least one (item) of a, b, or c may represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, c may be single or plural.
[0044] In addition, in order to clearly describe the technical solutions of the embodiments of the present invention, in the embodiments of the present invention, terms such as "first" and "second" are used to distinguish identical or similar items with basically the same functions and roles. Those skilled in the art can understand that terms such as "first" and "second" do not limit the quantity and execution order, and "first", "second", etc. do not necessarily mean different. At the same time, in the embodiments of the present invention, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, using words such as "exemplary" or "for example" aims to present relevant concepts in a specific way for easy understanding.
[0045] The principles and features of the present application will be described below. The examples given are only used to explain the present application and are not used to limit the scope of the present application.
[0046] Before describing the embodiments of the present application, the related concepts will be described.
[0047] Cost Volume (cost volume) is a concept used in computer vision to measure the similarity between left and right views, especially in optical flow estimation. It is a four-dimensional (4D) tensor, usually represented as [B, C, D, H]. Among them, B is the batch dimension, C is the channel dimension, D is the disparity dimension, and H is the height of the feature map. In optical flow estimation, Cost Volume is used to represent the data matching cost between a pixel and its corresponding pixel in the next frame.
[0048] Optical flow, in image registration, is represented as the displacement of corresponding pixels between the images to be registered.
[0049] An embodiment of the present application provides a three-dimensional inspection method for augmented reality-assisted assembly, which is applicable to the field of industrial assembly. This method for three-dimensional inspection of augmented reality-assisted assembly comprehensively examines the assembly situation based on the texture registration result and the spatial pose registration result. The two are not simply spliced and combined. The spatial pose registration is carried out on the basis of the texture registration result. In this way, the accuracy of the inspection result is improved.
[0050] An embodiment of the present application provides a three-dimensional inspection system for augmented reality-assisted assembly, which can execute the three-dimensional inspection method for augmented reality-assisted assembly provided by the embodiment of the present application. Figure 1 It is a schematic structural diagram of a three-dimensional inspection system for augmented reality-assisted assembly provided by an embodiment of the present application.
[0051] As Figure 1 shown, the three-dimensional inspection system 001 for augmented reality-assisted assembly includes at least one processor 011 and a memory 012 communicatively connected to the at least one processor; wherein, the memory 012 stores instructions executable by the at least one processor 011, and the instructions are executed by the at least one processor 011 so that the at least one processor 011 can execute the three-dimensional inspection method for augmented reality-assisted assembly provided by the embodiment of the present application.
[0052] Figure 2 It is a schematic flow diagram of a three-dimensional inspection method for augmented reality-assisted assembly provided by an embodiment of the present application. As Figure 2 shown, in some embodiments, the three-dimensional inspection method for augmented reality-assisted assembly includes the following steps:
[0053] S1, obtain the real camera image and the three-dimensional digital model of the assembly to be inspected in the same state.
[0054] Use a camera to photograph the entire assembly process of the assembly to be inspected, save the real camera image, and strictly calibrate the camera pose.
[0055] Collect and save the real camera image and the three-dimensional digital model in the same state. When the assembly to be inspected has multiple states, record the real camera image and the three-dimensional digital model in each state respectively. In some embodiments, if the assembly to be inspected has multiple components, the assembly steps of the assembly to be inspected can be multiple process assembly steps. Completion of each process assembly step corresponds to one state. Multiple process assembly steps correspond to multiple states. In one implementation, one process assembly step can assemble multiple components, and the number of states of the assembly to be inspected is less than the number of components. In another implementation, one process assembly step can assemble only one component, and the successful assembly of each component corresponds to one state, so the number of states of the assembly to be inspected is equal to the number of components.
[0056] In one implementation, taking assembled building blocks as an example, it includes seven components. The entire process of assembling the building blocks is photographed with a camera, the real camera images are saved, and the camera poses are strictly calibrated.
[0057] In some embodiments, to enrich the training data set of the cross-domain image registration network model, the YCB-Video data set is obtained and used as the training data set together with the assembled body to be inspected. The YCB-Video data set includes six-dimensional (6D) pose parameters and the camera parameters of the real camera.
[0058] S2. According to the three-dimensional digital model, through virtual camera rendering, virtual rendered images are obtained.
[0059] According to the calibrated camera poses, virtual cameras are set, and the building block CAD model is rendered onto a two-dimensional image to obtain virtual rendered images aligned with the real camera images during the process of assembling the building blocks, and the optical flow between them is recorded.
[0060] It can be understood that the real camera images and the virtual rendered images in each state correspond to each other. For example, the second step in the process of assembling corresponds to a virtual rendered image. The third step in the process of assembling corresponds to a virtual rendered image.
[0061] In some embodiments, to enrich the training data set of the cross-domain image registration network model, according to the YCB-Video data set, a three-dimensional digital model (such as a Computer Aided Design (CAD) model) is rendered onto a two-dimensional image to obtain virtual rendered images aligned with the real camera images, and the optical flow between them is recorded.
[0062] In some embodiments, to enrich the training data set of the cross-domain image registration network model, affine transformation, homogeneous transformation, and thin plate spline methods are used to deform the above-aligned YCB data set and the building block assembly data set, and the optical flow between the real camera images and the virtual rendered images after deformation is recorded. In one implementation, an enhanced data set as shown in Figure 4 is obtained. The enhanced data set includes the YCB data set, the data set of the assembled body to be inspected (the real camera images, three-dimensional digital models, and virtual rendered images of the assembled body to be inspected), the real camera images after deformation, and the virtual rendered images after deformation.
[0063] By performing step S1 and step S2, the training data set of the cross-domain image registration network model is obtained.
[0064] S3. Input the real camera images and the virtual rendered images into the cross-domain image registration network model to train the cross-domain image registration network model.
[0065] In some embodiments, the real camera image and the virtual rendering image are pre - processed to obtain feature maps of different sizes. The pre - processing includes image segmentation and size adjustment. In some embodiments, four different - sized feature maps are obtained, namely the first - sized feature map, the second - sized feature map, the third - sized feature map, and the fourth - sized feature map. The size of the first - sized feature map is smaller than that of the second - sized feature map. The size of the second - sized feature map is smaller than that of the third - sized feature map. The size of the third - sized feature map is smaller than that of the fourth - sized feature map. In one implementation, a real camera image with a size of "256×256" is input into the VGG - 16 network to obtain a first - sized feature map with a size of "16×16", a second - sized feature map with a size of "32×32", a third - sized feature map with a size of "64×64", and a fourth - sized feature map with a size of "128×128". A virtual rendering image with a size of "256×256" is input into the VGG - 16 network to obtain a first - sized feature map with a size of "16×16", a second - sized feature map with a size of "32×32", a third - sized feature map with a size of "64×64", and a fourth - sized feature map with a size of "128×128".
[0066] As Figure 3 shown, in some embodiments, the cross - domain image registration network model includes a first matching layer, a second matching layer, a third matching layer, and a fourth matching layer. The first matching layer, the second matching layer, the third matching layer, and the fourth matching layer are in a pyramid structure. The first matching layer, the second matching layer, the third matching layer, and the fourth matching layer can also be referred to as pyramid layers. The first matching layer is the top layer of the pyramid and can also be called the top - layer pyramid.
[0067] It can be understood that the number of types of feature map sizes is the same as the number of layers of the matching layers.
[0068] The first matching layer is used to input the first - sized feature map of the real camera image and the first - sized feature map of the virtual rendering image to obtain a first de - convolution result. In this way, rough registration of the feature maps is performed.
[0069] In some embodiments, the first matching layer includes a coarse correlation layer and a coarse decoder. The coarse correlation layer is built according to the matching layer proposed by Rocco, and the coarse decoder is built according to the correspondence mapping decoder.
[0070] In some embodiments, the three - dimensional inspection method for augmented - reality - assisted assembly further includes:
[0071] S311, input the first - sized feature map of the real camera image and the first - sized feature map of the virtual rendering image into the coarse correlation layer and the coarse decoder in sequence to obtain a first optical flow result.
[0072] In some embodiments, when performing step S311, the three-dimensional inspection method for augmented reality-assisted assembly further includes:
[0073] S3111, in the coarse correlation layer, by means of global search, calculate the first correlation result between each pixel point in the first size feature map of the real camera image and all pixel points in the first size feature map of the virtual rendering image.
[0074] S3112, input the first correlation result into the coarse decoder to output the coarse registration optical flow.
[0075] The coarse registration optical flow, that is, the first optical flow result.
[0076] S312, deconvolve the first optical flow result to obtain the first deconvolution result.
[0077] The second matching layer is used to input the second size feature map of the real camera image, the second size feature map of the virtual rendering image, and the first deconvolution result to obtain the second deconvolution result.
[0078] In some embodiments, the second matching layer includes a warp layer, a fine correlation layer, and a fine decoder. The warp layer is used to deform the feature map by means of bilinear interpolation. The fine correlation layer is built according to the cost volume layer of FlowNet2.0. The fine decoder is built according to the optical flow estimator of FlowNet2.0.
[0079] In some embodiments, the three-dimensional inspection method for augmented reality-assisted assembly further includes:
[0080] S321, input the second size feature map of the virtual rendering image and the first deconvolution result into the warp layer to obtain the first deformation result.
[0081] S322, input the second size feature map of the real camera image and the first deformation result into the fine correlation layer to obtain the first cost volume estimation result.
[0082] In the fine correlation layer, by means of local search, calculate the correlation between each pixel point in the second size feature map of the real camera image and the neighborhood pixel points in the first deformation result to obtain the first cost volume estimation result.
[0083] S323, input the first cost volume estimation result and the first deconvolution result into the fine decoder to obtain the first fine decoding result.
[0084] S324, obtain the second optical flow result according to the first fine decoding result and the first deconvolution result.
[0085] S325, Deconvolve the second optical flow result to obtain the second deconvolution result.
[0086] The third matching layer is used to input the third-size feature map of the real camera image, the third-size feature map of the virtual rendering image, and the second deconvolution result to obtain the third deconvolution result.
[0087] The fourth matching layer is used to input the fourth-size feature map of the real camera image, the fourth-size feature map of the virtual rendering image, and the third deconvolution result to obtain the fourth deconvolution result.
[0088] For the middle structure and data processing methods in the third matching layer and the fourth matching layer, reference can be made to the second matching layer, which will not be elaborated here. After three layers of fine registration layers (the second matching layer, the third matching layer, and the fourth matching layer), the fine registration optical flow result is output.
[0089] Among them, the fourth deconvolution result is used to obtain the estimated optical flow result. In some embodiments, the fourth deconvolution result is upsampled to output the estimated optical flow result.
[0090] In one implementation, the batch size is set to 8, the Adam optimizer is used, the initial learning rate is set to 0.0001, and the weight decay is set to 0.0004. The loss weight α1 of the first matching layer = 0.32. The loss weight α2 of the second matching layer = 0.08. The loss weight α3 of the third matching layer = 0.02. The loss weight α4 of the fourth matching layer = 0.01. The fine correlation layer search value of the second matching layer is set to 81. The fine correlation layer search value of the third matching layer is set to 81. The fine correlation layer search value of the fourth matching layer is set to 81.
[0091] Exemplarily, the experiment can be carried out using Pytorch on a three-dimensional inspection system for augmented reality-assisted assembly equipped with a central processing unit (model: Intel Core i9-13900KF CPU) and a graphics processing unit (model: NVIDIA GeForce RTX 4080 GPU).
[0092] The estimated optical flow result is used for spatial pose registration.
[0093] In some embodiments, the assembly to be inspected includes a first component and a second component, and the three-dimensional inspection method for augmented reality-assisted assembly further includes:
[0094] S331, Perform image segmentation on the real camera image and the virtual rendering image to obtain the image segmentation masks of the first component and the second component.
[0095] In some embodiments, MaskCNN is used for image segmentation to obtain the image segmentation masks (masks) of each component.
[0096] S332. Input the image segmentation masks of the first component and the second component into the cross - domain image registration network model to obtain the estimated optical flow result.
[0097] To better visualize the network results, as Figure 5 shown, using the estimated optical flow result, each pixel in the real camera image (as Figure 5 shown in the first column) is remapped to the corresponding pixel position in the virtual rendered image (as Figure 5 shown in the second column), thereby converting the virtual rendered image to the domain where the real camera image is located, and obtaining the deformed result of the real camera image through the estimated optical flow result (as Figure 5 shown in the fourth column). Among them, the deformed result of the real camera image through the real optical flow result (as Figure 5 shown in the third column) is pre - obtained after calibrating the camera pose and is the real (relative to the estimated) deformed result.
[0098] In some embodiments, to obtain a better model, the endpoint error (EPE) is used as the evaluation index of the optical flow, and a multi - scale training loss is adopted to train the network to obtain a trained model.
[0099] By executing steps S1 to S3, the training of the cross - domain image registration network model is completed.
[0100] S4. Input the image to be inspected of the assembly to be inspected into the cross - domain image registration network model to match the first virtual image.
[0101] In one implementation, as Figure 9 shown, taking the assembled building blocks as an example, when the image to be inspected of the assembly to be inspected is the assembly image of the second step (step 2) in the process assembly steps, the first virtual image is the virtual rendered image of the second step (step 2) in the process assembly steps obtained according to the aforementioned steps S1 and S2.
[0102] S5. Obtain the texture registration result according to the image to be inspected and the first virtual image; and, obtain the texture similarity score according to the texture registration result.
[0103] Input the image to be inspected and the first virtual image into the cross - domain image registration network model. The processing of the image to be inspected and the first virtual image in the cross - domain image registration network model can refer to step S3 and will not be elaborated here. After the image to be inspected and the first virtual image are calculated by the first matching layer, the second matching layer, the third matching layer, and the fourth matching layer, the texture registration result is obtained.
[0104] The texture registration result includes the estimated optical flow result.
[0105] In some embodiments, according to the estimated optical flow result, the real camera image is deformed to align with the virtual rendered image.
[0106] In one implementation, the calculation formula of texture similarity includes:
[0107]
[0108] Wherein, for each component i, I matching point pairs are randomly selected and the similarity within their L neighborhoods is calculated. N and respectively represent the pixel points within the neighborhood range of the matching point pairs in the original real camera image and the deformed real camera image.
[0109] S6. According to the texture registration result, perform spatial pose registration on the assembly to be inspected and the 3D digital model to obtain the spatial pose similarity score.
[0110] Such as Figure 6 shown, in some embodiments, in a 3D inspection system for augmented reality assisted assembly, an object coordinate system is established and numbered for each component of the 3D digital model. In one implementation, taking an assembly block as an example, there are a total of seven components, and an object coordinate system is established and numbered for each block component. For the sake of simplicity in calculation, exemplarily, it is ensured that the axis directions of each component are consistent, which means that the relative rotation between each component is "0".
[0111] To obtain the spatial pose similarity score, before performing spatial pose registration on the assembly to be inspected and the 3D digital model, first obtain the relative displacement template values between each component. In some embodiments, based on the object coordinate system of the 3D digital model, the relative displacement template values are obtained.
[0112] As shown in Table 1, where |ΔT x |, |ΔT y |, and |ΔT z | respectively represent the relative translation amounts of two components on the X, Y, and Z axes. Combined with Figure 6 and Table 1, the relative displacement template values (mm) of each component of the block car.
[0113] Table 1 Relative displacement template values (mm) of each component of the block car
[0114] ①② ①③ ①④ ①⑤ ①⑥ ①⑦ ②③ <![CDATA[|ΔT x |]]> 60.0 0 40.0 40.0 30.0 90.0 60.0 <![CDATA[|ΔT y |]]> 0 29.0 29.0 43.0 5.0 5.0 29.0 <![CDATA[|ΔT z |]]> 0 0 0 0 0 0 0 ②④ ②⑤ ②⑥ ②⑦ ③④ ③⑤ ③⑥ <![CDATA[|ΔT x |]]> 20.0 20.0 30.0 30.0 40.0 40.0 30.0 <![CDATA[|ΔT y |]]> 29.0 43.0 5.0 5.0 0 14.0 24.0 <![CDATA[|ΔT z |]]> 0 0 0 0 0 0 0 ③⑦ ④⑤ ④⑥ ④⑦ ⑤⑥ ⑤⑦ ⑥⑦ <![CDATA[|ΔT x |]]> 90.0 0 10.0 50.0 10.0 50.0 60.0 <![CDATA[|ΔT y |]]> 24.0 14.0 24.0 24.0 38.0 38.0 0 <![CDATA[|ΔT z |]]> 0 0 0 0 0 0 0
[0115] In some embodiments, the 3D inspection method for augmented reality assisted assembly further includes:
[0116] S611. According to the estimated optical flow result, obtain the first dense registration point pairs.
[0117] The first dense registration point pair is a two-dimensional to two-dimensional (2D-2D) dense registration point pair corresponding to the image to be inspected and the first virtual image.
[0118] S612. Obtain a second dense registration point pair according to the preset rendering parameters and the first dense registration point pair.
[0119] The second dense registration point pair is a two-dimensional to three-dimensional (2D-3D) dense registration point pair corresponding to the image to be inspected and the three-dimensional digital model. The second dense registration point pair includes multiple pairs.
[0120] S613. Randomly select some of the second dense registration point pairs and iteratively calculate the six-dimensional pose.
[0121] In some embodiments, using the Random Sample Consensus (RANSAC) framework, randomly select some 2D-3D matching point pairs and input them into the Perspective-n-Point (PNP) algorithm to recover the pose of a three-dimensional object from a two-dimensional image, and iteratively calculate the six-dimensional (6D) pose of each component of the assembly. As Figure 7 shown, in one implementation, taking an assembled building block as an example, obtain the six-dimensional (6D) pose of each component.
[0122] S614. Based on the six-dimensional pose, obtain the displacement similarity and the rotation similarity to obtain the spatial pose similarity score.
[0123] In some embodiments, based on the difference between the six-dimensional pose and the relative displacement template value, obtain the relative displacement. In one implementation, as Figure 7 and shown in Table 1, subtract the six-dimensional pose from the relative displacement template value to obtain the relative displacement shown in Table 2.
[0124] Table 2 Estimated relative displacements of components of the building block car
[0125] ①② ①③ ①④ ①⑤ ①⑥ ①⑦ ②③ <![CDATA[|ΔT x | est > 59.8 4.11 39.957 40.52 33.249 85.598 55.69 <![CDATA[|ΔT y | est > 4.779 26.833 33.356 44.936 9.994 9.939 22.054 <![CDATA[|ΔT z | est > 0.744 6.767 1.478 6.879 3.537 4.068 6.023 ②④ ②⑤ ②⑥ ②⑦ ③④ ③⑤ ③⑥ <![CDATA[|ΔT x | est > 19.843 19.28 26.551 25.798 35.847 36.41 29.139 <![CDATA[|ΔT y | est > 28.577 40.157 5.215 5.16 6.523 18.103 16.839 <![CDATA[|ΔT z | est > 0.734 6.135 2.793 3.324 5.289 0.112 3.23 ③⑦ ④⑤ ④⑥ ④⑦ ⑤⑥ ⑤⑦ ⑥⑦ <![CDATA[|ΔT x | est > 81.488 0.563 6.708 45.641 7.271 45.078 52.349 <![CDATA[|ΔT y | est > 16.894 11.58 23.362 23.417 34.942 34.997 0.055 <![CDATA[|ΔT z | est > 2.699 5.401 2.059 2.59 3.342 2.811 0.531
[0126] Compare the estimated relative displacements between components with the relative displacement template values between components in the CAD model to obtain the average displacement error of each component during the assembly process, as shown in Table 3.
[0127] Table 3 Estimated average displacement errors of components of the building block car
[0128] ① ② ③ ④ ⑤ ⑥ ⑦ <![CDATA[ΔT est > 3.287 2.675 4.648 2.504 3.060 2.936 3.662
[0129] In one implementation, the calculation formula for displacement similarity includes:
[0130]
[0131] where N represents the number of components; Represents the estimated translation value of component i on the τ(X / Y / Z) axis; Represents the relative translation value of components i and j on the τ(X / Y / Z) axis in the 3D digital model; ΔT i Represents the average displacement error of each component i, i.e., displacement similarity, which is determined by the average value of the displacement errors on the X, Y, and Z axes.
[0132] In one implementation, the calculation formula for rotation similarity includes:
[0133]
[0134] where N represents the number of components, Represents the estimated rotation value of the component on the axis, Represents the estimated rotation value of component i on the τ(X / Y / Z) axis, Represents the relative rotation value of components i and j on the τ(X / Y / Z) axis in the 3D digital model; ΔR i Represents the average rotation error of component i, i.e., rotation similarity, which is determined by the average value of the rotation errors on the X, Y, and Z axes.
[0135] In this way, by combining texture registration and spatial pose registration, adopting the architecture of "image texture registration + PnP 6D pose estimation", the texture registration result and the spatial pose registration result are integrated. They are not simply spliced together, and the spatial pose registration is carried out on the basis of the texture registration result.
[0136] S7. Calculate the assembly score according to the texture similarity score and the spatial pose similarity score.
[0137] In one implementation, the calculation formulas for the texture similarity score and the spatial pose similarity score include:
[0138] TS i , RS i , AMS i =Z_Score(ΔT i , ΔR i , AM i );
[0139] By integrating the texture similarity and the spatial pose similarity (including displacement similarity and rotation similarity) results of each component i, the texture similarity score and the spatial pose similarity score AMS are obtained using the Z-score normalization method. i , TS i , RS i .
[0140] In one implementation, the calculation formula for the assembly inspection score TAI includes:
[0141] TAI i = αTS i + βRS i +(1 - α - β)AMS i ;
[0142] The weighted average texture similarity score and the spatial pose similarity score are used to obtain the assembly inspection score TAI. When TAI is less than the threshold, it is determined that the assembly is successful; otherwise, it is determined that the assembly fails.
[0143] There are various implementation forms for the visual display of the assembly score.
[0144] For example,
[0145] In some embodiments, after the assembly is completed, all components of the building block car are inspected simultaneously. The effect is as Figure 8 shown, which shows the augmented reality inspection visual interface. The numbers in the figure represent the TAI scores of each component. Combining Figure 6 and Figure 8 , the TAI score of component ① is 2.928. The TAI score of component ② is 2.762. The TAI score of component ③ is 4.335. The TAI score of component ④ is 2.774. The TAI score of component ⑤ is 2.977. The TAI score of component ⑥ is 3.002. The TAI score of component ⑦ is 3.445.
[0146] Also, for example,
[0147] In other embodiments, or, the method provided by the embodiments of the present application can inspect components separately for each step of the assembly. The effect is as Figure 9 shown, which shows the augmented reality inspection visual interface of components at each assembly stage.
[0148] In addition, according to the TAI score, the method provided by the embodiments of the present application can find the misassembly situation and display an error reminder in the visual interface. As Figure 10 shown, the building block component ② is misaligned by 16 mm along the z-axis, while the correct CAD assembly template is as Figure 11 shown. The method provided by the embodiments of the present application can obtain the error inspection result by combining texture registration and spatial pose registration, as Figure 12 shown, while the visual inspection interface is as Figure 13 shown.
[0149] In addition to assembling the building block car, the method provided by the embodiments of the present application assembles the building block fox according to the virtual guidance provided by the augmented reality assisted assembly system, as Figure 14As shown. (a)-(b) Augmented reality virtual guidance, where the augmented reality technology superimposes the virtual CAD models of the components required for each assembly step onto the real model. (c) The real model assembled according to the augmented reality virtual guidance. (d) CAD assembly template. During this process, 100 images of correct assembly and 100 images of incorrect assembly are randomly taken from any angle to test and check the accuracy. After testing, the method provided by the embodiment of the present application can add an inspection function to the augmented reality assisted assembly system to determine whether the assembly is correctly installed, and the correct inspection rate is 92%.
[0150] Compared with the related technology that extracts Speeded-Up Robust Features (SURF) from the video stream, matches the SURF features with the local features extracted from the reference image, and determines whether the assembly meets the standard according to the matching result. In this way, only two-dimensional features are considered, lacking the three-dimensional feature information of the assembly, resulting in low accuracy of the inspection result. The three-dimensional inspection method for augmented reality assisted assembly provided by the embodiment of the present application comprehensively checks the assembly situation based on the texture registration result and the spatial pose registration result. The two are not simply spliced and combined. The spatial pose registration is carried out on the basis of the texture registration result. In this way, the accuracy of the inspection result is improved.
[0151] Compared with the related technology that performs multi-view joint inspection through two-dimensional features, the complexity is relatively high and the time consumption is relatively long. The three-dimensional inspection method for augmented reality assisted assembly provided by the embodiment of the present application takes both two-dimensional features and three-dimensional features into account, improving the convenience of actual operation.
[0152] Compared with the related technology that registers the three-dimensional digital model with the actual assembly components by using 3D point cloud. By aligning the assembly base, the point cloud of the remaining components can be spatially registered for inspection. Using 3D point cloud relies on 3D vision technology, with relatively high algorithm complexity and long time consumption, and it is difficult to be applied to the actual industrial assembly scenario. The three-dimensional inspection method for augmented reality assisted assembly provided by the embodiment of the present application not only has high accuracy, does not rely on 3D technology, has short time consumption, and is applicable to various industrial assembly scenarios.
[0153] Compared with the related technology where the pose estimation error is in the order of 5 centimeters and 5 degrees, it is difficult to be directly applied to the industrial assembly field. The three-dimensional inspection method for augmented reality assisted assembly provided by the embodiment of the present application has a small pose estimation error and can be applied to the industrial assembly field.
[0154] Based on the same application concept, an embodiment of the present application further provides a three-dimensional inspection system for augmented reality-assisted assembly. The method corresponding to the three-dimensional inspection system for augmented reality-assisted assembly may be the three-dimensional inspection method for augmented reality-assisted assembly in the foregoing embodiments, and the principle of solving problems is similar to that of this method. The three-dimensional inspection system for augmented reality-assisted assembly provided by the embodiment of the present application includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the methods and / or technical solutions of multiple foregoing embodiments of the present application.
[0155] Another embodiment of the present application further provides a computer-readable storage medium, on which computer program instructions are stored, and the computer program instructions can be executed by a processor to implement the methods and / or technical solutions of any one or more foregoing embodiments of the present application.
[0156] Specifically, this embodiment can adopt any combination of one or more computer-readable media. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of a computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device.
[0157] A computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including - but not limited to - electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device.
[0158] The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including - but not limited to - wireless, wire, optical fiber cable, RF, and the like, or any suitable combination of the foregoing.
[0159] The computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0160] The flowcharts or block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of devices, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a portion of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0161] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0162] In several embodiments provided by this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or page components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical, or other forms.
[0163] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0164] In addition, in each embodiment of this application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of a combination of hardware and software functional units.
[0165] The above integrated units implemented in the form of software functional units can be stored in a computer-readable storage medium. The above software functional units stored in a storage medium include several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute some steps of the methods described in each embodiment of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0166] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of this application.
[0167] In addition, it is obvious that the term "comprising" does not exclude other elements or steps, and the singular does not exclude the plural. A plurality of elements or devices recited in a device claim may also be implemented by one element or device through software or hardware. Terms such as first, second, etc. are used to denote names and do not denote any particular order.
Claims
1. A three-dimensional inspection method for augmented reality assisted assembly, characterized in that: include: Obtain the real camera image and 3D digital model of the assembly to be inspected in the same state; According to the three-dimensional digital model, a virtual rendering image is obtained by rendering with a virtual camera; The real camera image and the virtual rendered image are input into a cross-domain image registration network model to train the cross-domain image registration network model; wherein the cross-domain image registration network model includes a first matching layer, a second matching layer, a third matching layer and a fourth matching layer; the first matching layer is used to input a first size feature map of the real camera image and a first size feature map of the virtual rendered image to obtain a first deconvolution result; the second matching layer is used to input a second size feature map of the real camera image, a second size feature map of the virtual rendered image and the first deconvolution result to obtain a second deconvolution result; the third matching layer is used to input a third size feature map of the real camera image , a third size feature map of the virtual rendered image and the second deconvolution result to obtain a third deconvolution result; the fourth matching layer is used to input the fourth size feature map of the real camera image, the fourth size feature map of the virtual rendered image and the third deconvolution result to obtain a fourth deconvolution result; the fourth deconvolution result is used to obtain an estimated optical flow result; the estimated optical flow result is an optical flow result of all pixels, which is used for spatial pose registration; the size of the first size feature map is smaller than the size of the second size feature map; the size of the second size feature map is smaller than the size of the third size feature map; the size of the third size feature map is smaller than the size of the fourth size feature map; Inputting the image to be inspected of the assembly to be inspected into the cross-domain image registration network model to match the first virtual image; Obtaining a texture registration result according to the image to be inspected and the first virtual image; and obtaining a texture similarity score according to the texture registration result; the texture registration result includes the estimated optical flow result; According to the texture registration result, the assembly to be inspected and the three-dimensional digital model are spatially registered to obtain a spatial posture similarity score; Calculating an assembly score according to the texture similarity score and the spatial pose similarity score; The three-dimensional inspection method for augmented reality assisted assembly also includes: According to the estimated optical flow result, a first dense registration point pair is obtained; the first dense registration point pair is a two-dimensional pair of two-dimensional dense registration point pairs corresponding to the image to be inspected and the first virtual image; According to the preset rendering parameters and the first dense registration point pair, a second dense registration point pair is obtained; the second dense registration point pair is a two-dimensional pair of three-dimensional dense registration point pairs corresponding to the image to be inspected and the three-dimensional digital model; the second dense registration point pair includes a plurality of pairs; Randomly selecting some of the second dense registration point pairs, and iteratively calculating the six-dimensional pose; Based on the six-dimensional posture, displacement similarity and rotation similarity are obtained to obtain the spatial posture similarity score; Wherein, the calculation formula of the displacement similarity includes: Where N represents the number of components; represents the estimated translation value of component i on the τ(X / Y / Z) axis; Indicates the relative translation value of components i and j on the τ(X / Y / Z) axis in the 3D digital model; ΔT i represents the average displacement error of each component i, i.e., the displacement similarity, which is determined by the average value of its displacement errors in the X, Y, and Z axes; The calculation formula of the rotation similarity includes: Where N represents the number of components, Represents the estimated rotation of the component about its axis. represents the estimated rotation value of component i on the τ(X / Y / Z) axis, Represents the relative rotation value of components i and j on the τ(X / Y / Z) axis in the 3D digital model, ΔR i The rotational mean error, i.e., the rotational similarity, of part i is determined by the average of its rotational errors in the X, Y, and Z axes.
2. The three-dimensional inspection method for augmented reality assisted assembly according to claim 1, characterized in that: The first matching layer includes a coarse correlation layer and a coarse decoder; the coarse correlation layer is constructed according to the matching layer; The coarse decoder is constructed according to the mapping decoder; The three-dimensional inspection method for augmented reality assisted assembly also includes: Inputting the first size feature map of the real camera image and the first size feature map of the virtual rendering image into the coarse correlation layer and the coarse decoder in sequence to obtain a first optical flow result; Deconvolute the first optical flow result to obtain the first deconvolution result.
3. The three-dimensional inspection method for augmented reality assisted assembly according to claim 1 or 2, characterized in that: The second matching layer includes a deformation layer, a fine correlation layer and a fine decoder; the deformation layer is used to deform the feature map by a bilinear interpolation method; the fine correlation layer is constructed according to the cost volume layer; The fine decoder is built based on the optical flow estimator; The three-dimensional inspection method for augmented reality assisted assembly also includes: Inputting the second size feature map of the virtual rendering image and the first deconvolution result into the deformation layer to obtain a first deformation result; Inputting the second size feature map of the real camera image and the first deformation result into the fine correlation layer to obtain a first cost volume estimation result; Inputting the first cost volume estimation result and the first deconvolution result into the fine decoder to obtain a first fine decoding result; Obtaining a second optical flow result according to the first precise decoding result and the first deconvolution result; Deconvolute the second optical flow result to obtain the second deconvolution result.
4. The three-dimensional inspection method for augmented reality assisted assembly according to claim 1 or 2, characterized in that: The assembly to be inspected includes a first component and a second component; The three-dimensional inspection method for augmented reality assisted assembly also includes: Performing image segmentation on the real camera image and the virtual rendering image to obtain an image segmentation mask of the first component and an image segmentation mask of the second component; The image segmentation mask of the first component and the image segmentation mask of the second component are input into a cross-domain image registration network model to obtain the estimated optical flow result.
5. The three-dimensional inspection method for augmented reality assisted assembly according to claim 1 or 2, characterized in that: The calculation formula for texture similarity includes: For each component i, I matching point pairs are randomly selected and their similarities within L neighborhoods are calculated, N and They respectively represent the pixel points in the neighborhood of the matching point pairs in the original real camera image and the deformed real camera image.
6. A three-dimensional inspection system for augmented reality assisted assembly, characterized in that: include: at least one processor; a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 to 5.
7. A computer readable medium having computer program instructions stored thereon, wherein the computer program instructions can be executed by a processor to implement the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Outdoor augmented reality application method based on cross-source image matching
CN111260794A