Model training method, device and data processing device based on monocular images
Through the model training method based on monocular images, the optical flow prediction results are used as proxy marks for proxy learning, which solves the problems of high training cost and poor model recognition capabilities in binocular image alignment, and realizes self-supervised learning and efficient recognition of binocular image stereo matching.
Patent Information
- Application Number
- CN201910753810.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-08-15
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2039-08-15
AI Technical Summary
The prior art requires a large number of corrected binocular image samples in binocular image alignment, which is expensive to train, and the models trained using synthetic simulated images have poor recognition ability of real images.
A model training method based on monocular images is adopted, two monocular images collected at different time points are used as training samples, and the results obtained through optical flow prediction are used as proxy marks, and proxy learning is carried out to guide the model to learn again optical flow prediction.
Without relying on corrected binocular image samples, self-supervised learning of binocular image stereo matching is realized, and the same model is used to predict optical flow and stereo matching, which reduces training costs and improves the recognition ability of the model.
Smart Images

Figure CN112396074B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision technology, and in particular, to a model training method, device, and data processing device based on a monocular image. Background Art
[0002] Stereo matching is a basic computer vision problem and is widely used in fields such as 3D digital scene reconstruction and autonomous driving. The goal of stereo matching is to predict the displacement of pixels, that is, the stereo disparity map between two stereo images.
[0003] When dealing with the stereo matching problem, a Convolutional Neural Networks (CNN) model is often used. The CNN model is trained with a large number of samples, and then the trained model is used to achieve stereo matching of stereo images. Since the cost of obtaining stereo image training samples with correct annotations is very high, in some implementations, synthetic simulation images are used for training instead. However, the model trained in this way has poor recognition ability for real images. In other implementations, unlabeled stereo images are used. The right image is distorted to the left image according to the predicted disparity map, and then the photometric loss is used to measure the difference between the distorted right image and the left image. However, this method still requires a large number of calibrated stereo images, and the training cost is high. Summary of the Invention
[0004] To overcome at least one of the above deficiencies, in a first aspect, the present application provides a model training method based on a monocular image, which is applied to training an image matching model. The method includes:
[0005] Obtain a first training image and a second training image collected by a monocular image acquisition device at different time points;
[0006] Obtain a first optical flow prediction result from the first training image to the second training image according to the photometric loss between the first training image and the second training image;
[0007] Use the first optical flow prediction result as a proxy label, and perform proxy learning for optical flow prediction using the first training image and the second training image.
[0008] In a second aspect, the present application provides a model training device based on a monocular image, which is applied to training an image matching model. The device includes:
[0009] An image acquisition unit, configured to obtain a first training image and a second training image collected by a monocular image acquisition device at different time points;
[0010] The first optical flow prediction module is configured to obtain a first optical flow prediction result from the first training image to the second training image according to the photometric loss between the first training image and the second training image;
[0011] The second optical flow prediction module is configured to use the first optical flow prediction result as a proxy label and perform proxy learning for optical flow prediction using the first training image and the second training image.
[0012] In a third aspect, the present application provides a data processing device, including a machine-readable storage medium and a processor. The machine-readable storage medium stores machine-executable instructions, and when the machine-executable instructions are executed by one or more of the processors, the data processing device is caused to implement the model training method based on a monocular image provided by the present application.
[0013] Compared with the prior art, the present application has the following beneficial effects:
[0014] The model training method, device, and image processing device based on a monocular image provided by the present application regard binocular image matching as a special case of optical flow prediction, and adopt a proxy learning method. The optical flow prediction result obtained by using two monocular images collected at different time points as training samples is used as a proxy label to guide the model to perform learning for optical flow prediction again. In this way, self-supervised learning of binocular image stereo matching can be achieved without relying on calibrated binocular image samples, and the same model can be used to predict optical flow and perform stereo matching. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] For clearly illustrating the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments. It should be understood that the following drawings only include illustrations of some optional embodiments of the present application and should not be regarded as a limitation on the scope of the present application. Without creative efforts, those skilled in the art can also obtain other relevant drawings based on these drawings.
[0016] Figure 1 It is a schematic block diagram of the data processing device provided by the embodiment of the present application;
[0017] Figure 2 It is a schematic flowchart of the steps of the model training method based on a monocular image provided by the embodiment of the present application;
[0018] Figure 3 It is a schematic diagram of the binocular image alignment principle provided by the embodiment of the present application;
[0019] Figure 4 It is a schematic diagram of the binocular image alignment principle provided by the embodiment of the present application;
[0020] Figure 5Schematic diagram of image matching model processing provided by an embodiment of this application;
[0021] Figure 6 Schematic diagram of comparison of optical flow prediction test results on the same data set;
[0022] Figure 7 Schematic diagram of comparison of binocular image alignment test results on the same data set;
[0023] Figure 8 Module schematic diagram of a model training device based on a monocular image provided by an embodiment of this application. Detailed implementation manners
[0024] To introduce the purpose, technical solution, and advantages of the embodiments of this application more clearly, the technical solutions in the embodiments of this application will be described below with reference to the accompanying drawings.
[0025] Please refer to Figure 1 , Figure 1 , which is a schematic hardware structure diagram of a data processing device 100 provided in this embodiment. The data processing device 100 may include a processor 130 and a machine-readable storage medium 120. The processor 130 and the machine-readable storage medium 120 may communicate via a system bus. Moreover, the machine-readable storage medium 120 stores machine-executable instructions (such as code instructions related to the image model training device 110). By reading and executing the machine-executable instructions corresponding to the image model training logic in the machine-readable storage medium 120, the processor 130 may execute the model training method based on a monocular image described above.
[0026] The machine-readable storage medium 120 mentioned in this article may be any electronic, magnetic, optical, or other physical storage device that can contain or store information, such as executable instructions, data, and so on. For example, the machine-readable storage medium may be: RAM (Radom Access Memory, random access memory), volatile memory, non-volatile memory, flash memory, storage drives (such as hard disk drives), solid-state drives, any type of storage disk (such as optical discs, DVDs, etc.), or similar storage media, or a combination thereof.
[0027] Please refer to Figure 2 , which is a flowchart of a model training method based on a monocular image provided in this embodiment. The following will elaborate on each step included in the method in detail.
[0028] Step S210, obtain a first training image and a second training image collected by a monocular image acquisition device at different time points.
[0029] Step S220: Obtain a first optical flow prediction result from the first training image to the second training image according to the photometric loss between the first training image and the second training image.
[0030] Step S230: Use the first optical flow prediction result as a proxy label, and perform proxy learning for optical flow prediction using the first training image and the second training image.
[0031] Specifically, since binocular image alignment is a computer vision task of determining the same object from two binocular images with horizontal stereo inspection.
[0032] Optical flow prediction is a technique for determining the motion of the same object in different frame images based on the photometric properties of pixels under the assumptions of brightness constancy and spatial smoothness.
[0033] Through the research of the inventor, it is found that binocular image alignment and optical flow prediction can be regarded as a kind of problem, that is, the matching problem of corresponding pixel points in the image. The main difference between the two is that binocular image alignment is a one-dimensional search problem. On the rectified binocular images, the corresponding pixels are located on the epipolar line. While optical flow prediction does not have such a constraint and can be regarded as a two-dimensional search problem. Therefore, binocular image alignment can be regarded as a special case of optical flow. If a pixel matching model that can perform well in a two-dimensional scene is trained, it can also well implement the pixel matching task in a one-dimensional scene. Therefore, in step S210 of this embodiment, two images collected at different time points by a monocular image acquisition device can be obtained as training samples to train the image matching model.
[0034] Specifically, for binocular image alignment, the left and right cameras of the binocular camera capture images simultaneously, and the relative positions of the two cameras are fixed. Therefore, according to this geometric characteristic, in the process of binocular image alignment, for the pixels on the epipolar line of the left image, their corresponding pixels should be located on the epipolar line of the right image, that is, this is a one-dimensional image matching problem.
[0035] Please refer to Figure 3 , the projection point of point P in the three-dimensional scene on the left image of the binocular image is pixel P l , and the projection point on the right image is pixel P r . When P l is determined, the epipolar line passes through the left image pole e l , and P l is located on the epipolar line, then the corresponding pixel P l on the right image corresponding to P r is also always located on the epipolar line, and the epipolar line passes through the right image pole e r . Among them, O l and O r are the centers of the left and right cameras respectively, and e land e r are poles. Please refer to Figure 4 , Figure 4 which shows an example of binocular stereo image rectification. The left and right cameras are parallel, and the epipolar lines are horizontal, that is, the binocular image alignment finds matching pixels along the horizontal line.
[0036] Optical flow describes the dense motion between two adjacent frames. The two images are taken at different times, and the camera position and pose can change between these two frames. The scenarios predicted by optical flow are rigid or non-rigid scenarios. For a rigid scenario, the objects in the scene do not move, and the difference in the images is only due to the movement (rotation or translation) of the camera. Then, the optical flow prediction can also become a one-dimensional image matching problem along the epipolar line. Binocular images are pictures taken at different angles at the same time. The binocular image alignment problem can be regarded as a problem of predicting the optical flow of two images after the camera moves from one position to another and takes pictures at these two positions in a rigid scene.
[0037] Since estimating the ego-motion itself will cause additional errors and the scene is not always rigid, the problem of camera ego-motion is not considered in this embodiment, and only the binocular image alignment is regarded as a special case of optical flow prediction. That is to say, if the image matching model can achieve good optical flow prediction in two-dimensional space, it should also be able to achieve good binocular image alignment in one-dimensional space.
[0038] Therefore, in step S220 of this embodiment, during the optical flow prediction process, the target image is warped to the reference image according to the predicted optical flow, and the photometric loss is constructed by measuring the difference between the warped target image and the reference image. However, for the pixels corresponding to the objects occluded by the foreground in the scene, the brightness constancy assumption no longer holds. Therefore, for the occluded pixels, the photometric loss may lead to incorrect training supervision. Therefore, in this embodiment, it is necessary to pre-determine and exclude the occluded pixels when using the photometric loss to predict the optical flow.
[0039] Specifically, in this embodiment, according to the photometric loss between the first training image and the second training image, an initial optical flow map and an initial confidence map from the first training image to the second training image are obtained, and then according to the initial optical flow map and the initial confidence map, the first optical flow prediction result after excluding the occluded pixels is obtained.
[0040] For example, the confidence of the occluded pixels in the initial confidence map is set to 0, and the confidence of the non-occluded pixels is set to 1. Then, according to the initial optical flow map and the initial confidence map, the first optical flow prediction result is obtained.
[0041] Since the confidence of the occluded pixels is 0, when the initial optical flow map is multiplied by the initial confidence map, the data of the occluded pixels is removed from the initial optical flow map, thereby obtaining an optical flow map with high confidence composed of non-occluded pixels.
[0042] Furthermore, in this embodiment, forward-backward photometric detection can be used to process the initial optical flow map, and the confidence of each pixel point is determined according to the photometric difference to obtain the confidence map. Among them, the confidence of the pixels with a photometric difference exceeding the preset threshold is set to 0 as the occluded pixels; the confidence of the pixels with a photometric difference not exceeding the preset threshold is set to 1 as the non-occluded pixels.
[0043] When performing forward-backward photometric detection, the forward optical flow F t to the second training map I t+1 of the pixel p on the initial optical flow map of the first training map I t→t+1 and the backward optical flow F' t→t+1 (p) are obtained, where F' t→t+1 (p) = F t+1 → t (p + F t→t+1 (p)), and F t+1→t is the initial optical flow from the second training map to the first training map.
[0044] The confidence map M t→t+1 (p) of the pixel p is obtained according to the following formula based on the forward optical flow and backward optical flow of the pixel p
[0045]
[0046] where p represents the pixel point, and δ(p) = 0.1(|F t→t+1 (p) + F' t→t+1 (p)|) + 0.05.
[0047] In addition, in this embodiment, the first training map and the second training map can also be exchanged for training to obtain the reverse optical flow map from the second training map to the first training map.
[0048] In step S220, the optical flow prediction from the first training map to the second training map can be performed according to the preset photometric loss function and smoothness loss function to obtain the first optical flow prediction result.
[0049] Specifically, the photometric loss function L p is:
[0050]
[0051] where p represents the pixel point, To obtain the first training image I t The image obtained after using Census transformation According to the forward optical flow from the first training image to the second training image Warped to The obtained warped image, where Hamming(x) is the Hamming distance.
[0052] The smoothness loss function L m Is in the form of:
[0053]
[0054] Where I(p) is the pixel point on the first training image or the second training image, N is the total number of pixels on the first training image or the second training image, Denotes the gradient, T denotes the transpose, I(p) is the pixel point on the first training image or the second training image, and F(p) is the point on the currently processed optical flow map.
[0055] In step S220, use L p + λL m As the loss function to train the image matching model, where λ = 0.1.
[0056] In addition, in the above step S230, since even with only sparse correct labels, the CNN can learn quite good optical flow predictions on the KITTI dataset. Therefore, in this embodiment, first obtain sparse and highly confident optical flow predictions from step S220, and then use them as proxy labels to guide the learning of image matching predictions.
[0057] Please refer to Figure 5 , in this embodiment, the first optical flow prediction result can be used as the proxy label, and using a preset proxy self-supervised loss function and smoothness loss function, perform optical flow prediction from the first training image to the second training image.
[0058] Specifically, the proxy self-supervised loss function L s Is in the form of:
[0059]
[0060] Where p represents the pixel point, F py Is the initial optical flow map, M py Is the initial confidence map, and F is the currently processed optical flow map.
[0061] In step S230, use L S + λL m As the loss function to train the image matching model, where λ = 0.1.
[0062] It should be noted that, different from the training process of step S220, in step S230, the action of removing unoccluded pixels is no longer performed, so that the model can predict the optical flow of the occluded area.
[0063] Optionally, in this embodiment, during the training process of step S230, the first training image and the second training image can be randomly cropped at the same position and with the same size and / or randomly downsampled in the same way first, and then the cropped and / or downsampled first training image and second training image are used for the training of step S230, so as to improve the accuracy of optical flow prediction for both occluded points and occluded areas.
[0064] Optionally, in this embodiment, during the training process of step S230, the first training image and the second training image can also be randomly scaled by the same coefficient or randomly rotated by the same angle first, and then the processed first training image and second training image are used for the training of step S230.
[0065] It should be noted that other methods can also be used to obtain high-confidence optical flow prediction. For example, traditional methods are used to calculate reliable disparity. In this embodiment, the model finally needs to perform optical flow prediction. Therefore, the optical flow prediction results and confidence maps obtained in step S220 are used, and then the high-confidence optical flow prediction is used as a proxy ground truth in step S230 to guide the neural network to learn image matching, and the above training process can be completed in one model.
[0066] In this embodiment, after proxy learning, the number of high-confidence pixels will increase. Therefore, after step S230, the second optical flow prediction result obtained by proxy learning can also be used for iterative training to further improve the recognition ability of the image matching model.
[0067] It should be noted that the image matching model trained by the method provided in this embodiment can be used for both optical flow prediction and binocular image alignment. When the trained image matching model performs optical flow prediction, the first training image I collected at different time points t to the second training image I t+1 can be used as inputs, and the optical flow map from I t to I t+1 is output. When the trained image matching model is used for binocular image alignment, the images I collected by the left and right cameras in the binocular images l and I r can be used as inputs, and the stereo disparity map of the output images from I l to I r is obtained as the matching result.
[0068] In this embodiment, the Adam optimizer can be used to build the image matching model on the TensorFlow system, and the batch size of the model is set to 4, the initial learning rate is 1e-4, and it is decayed by half every 60k iterations. During training, the normalized images can be used as inputs and data augmentation such as random cropping, scaling, or rotation can be performed. In particular, the cropping size can be set to [256, 640] pixel sizes, and the random scaling factor range can be set to [0.75, 1.25].
[0069] In step S220, the photometric loss can be applied to all pixels, and the image matching model can be trained using the photometric loss for 100k iterations from scratch. It should be noted that at the beginning, high-confidence pixels and low-confidence pixels are not distinguished, because directly applying the photometric loss only to high-confidence pixels may result in an obvious solution where all pixels are regarded as low-confidence pixels. After that, the photometric loss function L p and the smoothness loss function L m are used to perform 400k iterations for the image matching model. In step S230, the proxy self-supervised loss function L s and the smoothness loss function L m are used to perform 400k iterations to train the image matching model.
[0070] Figure 6 shows the test results of optical flow prediction using the existing model and the image matching model trained by the method provided in this embodiment on the KITTI 2012 dataset and the KITTI 2015 dataset. From Figure 6 it can be seen that the recognition ability of the image matching model ("Our+proxy" item) trained by the monocular image-based model training method provided in this embodiment is significantly better than existing state-of-the-art unsupervised methods such as MultiFrameOccFlow and DDFlow.
[0071] Figure 7 shows the test results of binocular image alignment using the existing model and the image matching model trained by the method provided in this embodiment on the KITTI 2012 dataset and the KITTI 2015 dataset. From Figure 7 it can be seen that the recognition ability of the image matching model ("Our+proxy+ft" item) trained by the monocular image-based model training method provided in this embodiment is significantly better than other existing state-of-the-art unsupervised methods.
[0072] Please refer to Figure 8, this embodiment also provides a model training device 110 based on a monocular image. The device includes an image acquisition module 111, a first optical flow prediction module 112, and a second optical flow prediction module 113.
[0073] The image acquisition unit 111 is configured to acquire a first training image and a second training image collected by a monocular image acquisition device at different time points.
[0074] The first optical flow prediction module 112 is configured to obtain a first optical flow prediction result from the first training image to the second training image according to the photometric loss between the first training image and the second training image.
[0075] The second optical flow prediction module 113 is configured to use the first optical flow prediction result as a proxy label and perform proxy learning for optical flow prediction using the first training image and the second training image.
[0076] In summary, for the model training method, device, and image processing device based on a monocular image provided in this application, by regarding binocular image matching as a special case of optical flow prediction and adopting the method of proxy learning, the first optical flow prediction result obtained by using two monocular images collected at different time points as training samples is used as a proxy label to guide the model to perform learning for optical flow prediction again. In this way, self-supervised learning of binocular image stereo matching can be performed without relying on calibrated binocular image samples, and the same model can be used to predict optical flow and perform stereo matching.
[0077] In the embodiments provided in this application, it should be understood that the disclosed device and method can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code includes one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0078] In addition, the functional modules in the various embodiments of the present application may be integrated together to form an independent part, or each module may exist alone, or two or more modules may be integrated to form an independent part.
[0079] If the above-mentioned function is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0080] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitations, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.
Claims
1. A model training method based on monocular images, characterized in that, applied to train an image matching model, the method comprising: Obtaining a first training image and a second training image collected by a monocular image acquisition device at different time points; Obtaining a first optical flow prediction result from the first training image to the second training image according to the photometric loss between the first training image and the second training image; Using the first optical flow prediction result as a proxy label and performing proxy learning of optical flow prediction using the first training image and the second training image; Using the trained image matching model to perform binocular image alignment and optical flow prediction; When the image matching model is used to perform binocular image alignment, inputting the binocular image to be processed into the trained image matching model; obtaining a stereo disparity map output by the image matching model for the binocular image to be processed.
2. The method according to claim 1, characterized in that, The step of obtaining the first optical flow prediction result from the first training image to the second training image includes: Obtaining an initial optical flow map and an initial confidence map from the first training image to the second training image according to the photometric loss between the first training image and the second training image; Obtaining the first optical flow prediction result after excluding occluded pixels according to the initial optical flow map and the initial confidence map.
3. The method according to claim 2, characterized in that, The manner of obtaining the initial confidence map includes: Processing the initial optical flow map using forward-backward photometric detection and determining the confidence corresponding to each pixel point according to the photometric difference to obtain the confidence map; Wherein, setting the confidence of pixels with a photometric difference exceeding a preset threshold to 0 as occluded pixels; setting the confidence of pixels with a photometric difference not exceeding the preset threshold to 1 as non-occluded pixels.
4. The method according to claim 3, characterized in that, The step of processing the initial optical flow map using forward-backward photometric detection and determining the confidence corresponding to each pixel point according to the photometric difference to obtain the confidence map includes: Obtain the first training graph I t to the second training graph I t+1 The forward optical flow F of pixel p on the initial optical flow graph t→t+1 (p) and the backward optical flow F′ t→t+1 (p), where F′ t→t+1 (p) = F t+1→t (p + F t→t+1 (p)), F t+1→t is the initial optical flow from the second training graph to the first training graph; Obtain the confidence map M t→t+1 (p) of the pixel p according to the forward optical flow and backward optical flow of the pixel p according to the following formula t→t+1 (p), where δ(p) = 0.1(|F t→t+1 (p) + F' t→t+1 (p)|) + 0.
05.
5. The method according to claim 4, characterized in that, The step of obtaining the first optical flow prediction result according to the initial optical flow map and the initial confidence map includes: Performing optical flow prediction from the first training image to the second training image according to a preset photometric loss function and smoothness loss function to obtain the first optical flow prediction result; The photometric loss function L p is in the form of: Among them, For the first training image I t The image obtained after using Census change, According to the forward optical flow from the first training image to the second training image Warped to The obtained warped image, where Hamming(x) is the Hamming distance; The smoothness loss function L m is in the form of: Among them, I(p) is the pixel point on the first training graph or the second training graph, and N is the total number of pixels of the first training graph or the second training graph. denotes the gradient, T denotes the transpose, I(p) is the pixel point on the first training graph or the second training graph, and F(p) is the point on the currently processed optical flow graph.
6. The method according to claim 4, characterized in that, The step of using the first optical flow prediction result as a proxy label and performing proxy learning of optical flow prediction using the first training image and the second training image includes: Using the first optical flow prediction result as a proxy label and performing optical flow prediction from the first training image to the second training image using a preset proxy self-supervised loss function and smoothness loss function.
7. The method according to claim 6, characterized in that, The proxy self-supervised loss function L s is in the form of: Among them, F py is the initial optical flow map, M py is the initial confidence map, and F is the optical flow map being currently processed.
8. The method according to claim 6, characterized in that, The step of using the first optical flow prediction result as a proxy label and performing optical flow prediction training from the first training image to the second training image using a preset proxy self-supervised loss function and a smoothness loss function includes: Performing the same random cropping and / or the same random downsampling on the first training image and the second training image; Using the first optical flow prediction result as a proxy label and performing machine learning training for image element matching using the cropped and / or downsampled first training image and second training image.
9. The method according to claim 1, wherein, after the step of using the first optical flow prediction result as a proxy label and performing proxy learning of optical flow prediction using the first training image and the second training image, the method further includes: Performing iterative training using the second optical flow prediction result obtained by proxy learning.
10. A model training device based on a monocular image, wherein, applied to train an image matching model, the device includes: An image acquisition unit, configured to acquire a first training image and a second training image collected by a monocular image acquisition device at different time points; A first optical flow prediction module, configured to obtain a first optical flow prediction result from the first training image to the second training image according to the photometric loss between the first training image and the second training image; A second optical flow prediction module, configured to use the first optical flow prediction result as a proxy label and perform proxy learning of optical flow prediction using the first training image and the second training image; The image matching model is configured to, when performing binocular image alignment, input the binocular image to be processed into the trained image matching model and obtain a stereo disparity map output by the image matching model for the binocular image to be processed.
11. A data processing device, wherein, comprising a machine-readable storage medium and a processor, the machine-readable storage medium stores machine-executable instructions, and when the machine-executable instructions are executed by the processor, the method according to any one of claims 1-9 is implemented.
Citation Information
Patent Citations
Method for detecting long-distance barrier
CN101701818A
End-to-end semantic instant-positioning and graph building method based deep learning
CN108665496A