Optical flow estimation method and device with comparative learning

By adopting a semi-supervised framework and feature-level contrast loss function in optical flow estimation, using real-world data to automatically label features and adjust pseudo-labels, the limitations of existing optical flow estimation methods in complex scenarios and real-world data are solved, and more accurate optical flow analysis and reduce model overfitting is achieved.

CN120019406APending Publication Date: 2025-05-16创峰科技
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202380069901.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-12
Filing Date
2023-06-29
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Existing optical flow estimation methods have limitations in handling occlusions, small fast moving objects, and capturing global motions, and deep learning models do not perform well on real-world data and are prone to overfitting synthetic datasets.

Method used

Using a semi-supervised framework, the features are automatically marked using real-world data sets to form false real value data, and the pseudo-label is adjusted through feature-level comparison loss function to improve optical flow prediction. The framework combines a gated cycle unit (GRU) and a contrast loss unit to improve the accuracy of optical flow analysis.

Benefits of technology

Improves the accuracy of optical flow analysis, especially when dealing with complex scenarios and real-world data, reduces the model's overfitting of synthetic data sets, and enhances performance on real-world data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120019406A_ABST
    Figure CN120019406A_ABST
Patent Text Reader

Abstract

A method for a computing system includes determining a first feature map and a second feature map in response to a first image and a second image, implementing a gated loop unit (GRU) to determine a pixel-level flow prediction in response to the first feature map and the second feature map, determining a warped feature map in response to the second image map, implementing a feature-level contrast loss function, the processor is configured to determine a feature-level loss in response to a first feature in the first image map and a second feature in the warped feature map, determine a pixel-level flow loss in response to the pixel-level flow prediction and the pixel-level truth data, and modify a parameter of the GRU in response to the pixel-level flow loss and the feature-level loss.
Need to check novelty before this filing date? Find Prior Art

Description

CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] The present invention is non-provisional and claims priority to U.S. Application No. 63 / 415,358 filed on October 12, 2022. This application is incorporated herein by reference for all purposes. Background Art

[0002] The present invention relates to object motion prediction. More particularly, the present invention relates to methods and apparatus for more accurate optical flow analysis.

[0003] Optical flow analysis / estimation is a key component of several high-level vision problems such as action recognition, video segmentation, autonomous driving, editing, etc. Traditional optical flow estimation methods involve formulating the problem as an optimization problem using hand-crafted features, which may be empirical. Optical flow estimation is usually performed by trying to maximize the visual similarity between adjacent frames through energy minimization. Deep neural networks have been shown to be effective in terms of the accuracy of optical flow estimation. However, these methods still have limitations, such as difficulty in handling occlusions, small fast-moving objects and capturing global motion, as well as correcting and recovering from early errors.

[0004] In order to improve the accuracy of end-to-end optical flow networks, synthetic datasets are often used for pre-training rather than real-world datasets. Synthetic datasets are usually easier to generate and include labels or identifiers of all features appearing on the image, such as vehicles 1, pedestrians 2, buildings 5, etc., i.e., ground truth data. In contrast, real-world data usually includes few labels and therefore requires manual labeling. Since large-scale data is required to train deep learning networks, synthetic datasets (i.e., computer-generated scenes and images) are often used for model training. One disadvantage of this strategy is that the generated deep learning models tend to overfit the synthetic training datasets, which subsequently show degraded performance on real-world data.

[0005] In view of the above, a solution is needed that can address the above challenges and reduce the disadvantages. Summary of the invention

[0006] The present invention relates to object motion prediction. More particularly, the present invention relates to methods and apparatus for more accurate optical flow analysis.

[0007] Embodiments disclose a semi-supervised framework for improving optical flow determination. More specifically, a real-world dataset is used to determine optical flow estimates using automatically labeled features on true real-world data. Features can span groups of consecutive pixels or individual pixels on an image. In operation, features within the true data (real-world data) are automatically labeled with pseudo-feature labels to form pseudo-true data. Subsequently, when the pseudo-feature label helps to reduce the optical flow loss, the pseudo-label of the feature is usually retained, and when the pseudo-feature label increases the optical flow loss, the pseudo-label of the feature is usually removed or deleted. More specifically, pseudo-labels can be assigned to features to form pseudo-true data. The pseudo-true data is then used to predict the feature-level flow of the feature in the first time feature map. The predicted feature-level flow is referred to as a warped feature map in this article. Then, the embodiment uses a contrast flow loss based on the warped feature map and the second time feature map to determine the correspondence or lack of correspondence between features. In some cases, the feature-level contrast flow loss can then be fed back to help determine whether the pseudo-label should be modified or retained.

[0008] In some embodiments, the semi-supervised framework described herein provides optical flow feedback based on feature-to-feature comparisons of feature maps. This optical flow feedback can be combined with other optical flow systems, such as the recurrent all-pairsfield transforms (RAFT) (Teed2020) architecture. Traditional optical flow systems (such as RAFT) typically perform pixel-by-pixel optical flow analysis. By combining this contrast loss function with RAFT optical flow analysis, optical flow prediction can be improved.

[0009] According to one aspect, a method of a computing system for estimating optical flow is disclosed. A technique may include determining a first feature map in a computing system in response to a first image, and determining a second feature map in response to a second image map, and implementing a gated recurrent unit (GRU) in the computing system to determine a pixel-level flow prediction in response to the first feature map and the second feature map. A process may include determining a warped feature map in a computing system in response to a second image map, and implementing a feature-level contrast loss function in the computing system to determine a feature-level loss in response to a first feature in the first image map and a second feature in the warped feature map. A method may include determining a pixel-level flow loss in a computing system in response to a pixel-level flow prediction and in response to pixel-level true value data, and modifying parameters of the GRU in response to the pixel-level flow loss and the feature-level loss in the computing system.

[0010] According to another aspect, a computing system for estimating optical flow is disclosed. A device may include a pixel-based analysis system, the pixel-based analysis system is used to determine a first feature map in response to a first image, and to determine a second feature map in response to a second image map, wherein the pixel-based analysis system includes a gated recurrent unit (GRU), and the GRU is used to determine pixel-level flow predictions in response to the first feature map and the second feature map. An apparatus may include a feature-based analysis system, the feature-based analysis system is coupled to a pixel-based optical flow analysis system, wherein the feature-based analysis system is used to determine a warp feature map in response to the second image map, wherein the feature-based analysis system includes a contrast loss unit, and the contrast loss unit is used to determine a feature-level loss in response to a first feature in the first image map and a second feature in the warp feature map. In some systems, the pixel-based analysis system is used to modify parameters of the GRU in response to the pixel-level flow loss and the feature-level loss.

[0011] According to another aspect, a method is disclosed. A process may include running an optical flow prediction system including a contrast loss function in response to a synthetic data set to determine a first optical flow loss, and adjusting parameters of the optical flow prediction system in response to the first optical flow loss. A technique may include running an optical flow prediction system in response to a real-world data set to determine a second optical flow loss, and adjusting parameters of the optical flow prediction system in response to the second optical flow loss. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] For a more complete understanding of the present invention, reference is made to the accompanying drawings. It should be understood that these drawings should not be considered as limiting the scope of the present invention, and the presently described embodiments and the best mode of the present invention currently understood are described in more detail through the use of the accompanying drawings, in which:

[0013] Figure 1 shows a functional block diagram of some embodiments of the present invention;

[0014] FIG. 2A to FIG. 2C shows a flow chart according to various embodiments of the present invention; and

[0015] Figure 3 A system diagram according to various embodiments of the present invention is shown. DETAILED DESCRIPTION

[0016] Figure 1A logical block diagram is shown, an embodiment of the present invention. More specifically, the system 100 includes a contrast loss portion 102 and a cyclic full-field transform (RAFT) portion 104. The input of the system 100 includes a series of images 106. In some embodiments, the input image 106 may include a synthetic image, such as a data set including labeled features, where the true value flow data 108 is known. In other embodiments, the input image 106 may include a real-world image, such as a data set in which some features are manually labeled and many features are unlabeled. In this case, the true value flow data 108 of the unlabeled features is unknown. The system 100 may be referred to as RAFT-CF in this article.

[0017] In various embodiments, the RAFT portion 104 includes three main components: a feature encoder 110 that determines a feature map 112 that stores a feature vector for each pixel in the input image 106; a correlation portion 114 that determines a four-dimensional (4D) correlation volume 116 that includes a displacement based on the feature map 112; and a gated recurrent unit (GRU) 118 that determines an optical flow prediction 120 for the feature map 122. The optical flow prediction 120 is compared to the true flow data 108 to determine a flow loss 124.

[0018] In operation, the feature encoder 110 receives an image 106 having a height, width, and color depth (e.g., HxWx3). Next, the location of the pixel (e.g., x-location and y-location) is determined and stored in the form of a coordinate frame 126 (e.g., HxWx5). Next, the encoder 128 is used to identify features, such as pixels of a specific color, from the coordinate frame 126. In various embodiments, the features can be stored as feature vectors in feature maps 112 of different resolutions.

[0019] In various embodiments, the correlation layer 114 receives the feature map 112 and determines a multi-dimensional (e.g., four-dimensional 4D) correlation volume 116. The 4D correlation volume 116 is typically determined by comparing the visual similarity between pixels in the feature map 112 and then determining the displacement of the pixels in the feature map 112. The correlation volume 116 is then input to a gated recurrent unit (GRU) 118, which warps the input feature map 146 and outputs an optical flow prediction 120. Therefore, the optical flow prediction 120 represents the predicted optical flow.

[0020] In various embodiments, as shown, the optical flow prediction 120 is then compared to the true value flow data 108 to determine a flow loss 124. The flow loss 124 generally indicates how well the GRU 118 can predict the pixel optical flow. The flow loss 124 can be fed back 140 into the GRU 118 to change its parameters. In some cases, the change in parameters when the flow loss 122 is small can be smaller than the change when the flow loss 122 is large.

[0021] RAFT portion 104 is typically limited to determining predicted optical flow when the true value flow is known, i.e., with synthetic images and synthetic datasets. In addition, it is typically limited to determining optical flow vectors pixel by pixel. As described above, RAFT portion 104 may sometimes overfit the synthetic dataset, so when training, RAFT portion 104 may have difficulty determining accurate optical flow vectors when provided with a real-world dataset.

[0022] In various embodiments, the contrastive loss portion 102 is used to supplement the RAFT portion 104. The contrastive loss portion 102 includes a feature warping process 144 that receives the feature map 128 and the truth stream 130 and outputs a warped feature map 142. The contrastive loss process 132 receives the warped feature map 142 and the feature map 134 and provides contrastive loss feedback 136. In some cases, the feature map 128 can be based on a synthetic dataset, while in other cases, the feature map 128 can be based on real-world data. In the case of real-world data, the contrastive loss feedback 136 can be used to fine-tune the truth stream 130, as described below.

[0023] In operation, the truth stream 130 may specify an optical flow of manually labeled features from real-world data, such as houses, cars, bicycles, etc. In addition, in some cases, the truth stream 130 may specify a flow of features predicted or guessed from an image, for example, a first block of pixels may be guessed and labeled as a pedestrian, a second block of similarly colored pixels may be guessed and labeled as a ball, etc. Such features are referred to as pseudo-labels, and the truth stream 108 with pseudo-labels is referred to as a pseudo-truth stream.

[0024] In various embodiments, the optical flow of the labeled features and pseudo-labeled features in the true value stream 130 is used to process the feature map 128. More specifically, in the feature warping process 144, the true value stream 130 warps the feature map 128 and outputs a warped feature map 142. Next, a contrast loss process 132 is performed by comparing the warped feature map 142 with the feature map 134. In various embodiments, the comparison is based on features (e.g., groups of pixels) rather than individual pixels. The contrast loss process 132 provides feedback for the pseudo-true value stream 130. For example, if the (pseudo-labeled) features of the warped feature map 142 are aligned or matched with the features of the feature map 134, the pseudo-label can be retained because it is correct. In addition, if the (pseudo-labeled) features of the warped feature map 142 are not aligned or matched with the features of the feature map 134, the pseudo-label can be removed because it is an incorrect label. Again, the optical flow of the pseudo-labeled features that substantially match the optical flow of the features in the feature map 134 can be retained in the pseudo-truth 130, and the optical flow of the pseudo-labeled features that do not match the optical flow of the features in the feature map 134 is removed from the pseudo-truth 130. In this way, unlabeled features in the real-world data can be labeled in this iterative process. In some embodiments, the process can then be repeated (136) using different pseudo-labels for the features.

[0025] In some embodiments, after certain conditions, feedback of the contrastive loss portion 102 data can be used as feedback 148 to adjust the parameters of the GRU 118. More specifically, as described above, the GRU 118 can predict optical flow pixel by pixel, but the contrastive loss portion 102 predicts optical flow feature by feature. Therefore, by combining these optical flow predictions, the optical flow prediction of the GRU 118 is generally improved. Experimental data results collected by the inventors confirm this improvement.

[0026] FIG. 2A to FIG. 2C A more complete process according to some embodiments is shown. FIG. 2A to FIG. 2C As shown, the process includes three stages, stage 202 is a synthetic data set training stage, stage 204 is a real-world data set training stage, and stage 206 is a real-world use stage.

[0027] First in Figure 2A In the embodiment of the present invention, a synthetic data image 208 is provided to a system 210 similar to the system disclosed above. As shown, a predicted flow 212 is determined and a flow loss 214 is determined. Feedback 216 is then provided to the training system 210. This process continues until the flow loss decreases, in which case a system 210' is formed.

[0028] Next, in Figure 2BIn this example, real-world images 218 are provided to the training system 210'. In this example, N^2 images are used and are logically arranged in an NxN grid 216. Next, a K-fold cross-validation process is performed to select a training subset of real-world images 218 as input to the system 210'. Figure 1 As shown, system 210' uses a contrastive loss function for features and pseudo-features to train system 210' based on a training subset of real-world images. The trained system 210' is then tested using real-world images 214 that are not in the training subset (test subset) to determine a prediction stream 218. The prediction stream 218 is compared to the pseudo-truth stream 220 of the test subset to determine an error 222, as shown. The next training subset and test subset from images 218 are used to determine an error 224, and so on. In various embodiments, after all K-fold cross validations, the errors can be combined or averaged to determine an average error or feedback. The feedback can be used as feedback 226 to change one or more parameters of the GRU within system 210', as described above. This process can be repeated until the error feedback decreases, in which case system 210" is formed.

[0029] exist Figure 2C In, system 210" has been used Figure 2A The synthetic dataset is trained on Figure 2B Thus, a new image 228 may be provided to the system 210" and the system 210" may predict an optical flow 230. In some cases, if the true optical flow 232 is known, the flow loss 234 may be determined again and fed back 236 into the system 210".

[0030] In some embodiments, the predicted optical flow 230 may be used as an input to other processes such as a driver assistance system, an autonomous driving system, an area mapping system, and the like.

[0031] Figure 3 Functional block diagrams of various embodiments of the present invention are shown. More specifically, it is contemplated that a computer (eg, a server, a laptop, a streaming server, a virtual machine, etc.) may be implemented with a subset or superset of the components shown below.

[0032] exist Figure 3In the embodiment, the computing device 300 may include some of the following components, but not necessarily all of the following components: application processor / microprocessor 302, memory 304, display 306, image acquisition device 310, audio input / output device 312, etc. Data and communication from and to the computing device 300 may be provided via: wired interface 314 (e.g., Ethernet, dock, plug, controller interface for peripheral devices); various radio frequency receivers, such as GPS / Wi-Fi / Bluetooth interface / UWB 316; NFC interface (e.g., antenna or coil) and driver 318; radio frequency interface and driver 320, etc. In some embodiments, physical sensors 322 (e.g., (MEMS-based) accelerometers, gyroscopes, magnetometers, pressure sensors, temperature sensors, bio-imaging sensors, etc.) are also included.

[0033] In various embodiments, the computing device 300 may be a computing device (e.g., Apple iPad, Microsoft Surface, Samsung Galaxy Note, Android Tablet); a smartphone (e.g., Apple iPhone, Google Pixel, Samsung Galaxy S); a computer (e.g., a netbook, a laptop, a convertible computer), a media player (e.g., Apple iPod); and the like. Typically, the computing device 300 may include one or more processors 302. Such a processor 302 may also be referred to as an application processor, and may include a processor core, a video / graphics core, and other cores. The processor 302 may include processors from Apple (A14 Bionic, A15 Bionic), NVidia (Tegra), Intel (Core), Qualcomm (Snapdragon), Samsung (Exynos), ARM (Cortex), MIPS Technology, microcontrollers, and the like. In some embodiments, a processing accelerator may also be included, such as an AI accelerator, Google (Tensor Processing Unit), a GPU, and the like. It is contemplated that other existing and / or later developed processors / microcontrollers may also be used in various embodiments of the present invention.

[0034] In various embodiments, memory 304 may include different types of memory (including memory controllers), such as flash memory (e.g., NOR, NAND), SRAM, DDR SDRAM, etc. Memory 304 may be fixed within computing device 300, and may also include removable memory (e.g., SD, SDHC, MMC, MINI SD, MICRO SD, SIM). The above are examples of computer-readable tangible media that can be used to store embodiments of the present invention, such as computer executable software code (e.g., firmware, application programs), security applications, application data, operating system data, databases, etc. In addition, in some embodiments, a security device including a secure memory and / or a secure processor is provided. It is contemplated that other existing and / or later developed memories and memory technologies may be used in various embodiments of the present invention.

[0035] In various embodiments, the display 306 can be based on various later developed or current display technologies, including LED or OLED displays and / or status lights; touch screen technology (e.g., resistive displays, capacitive displays, optical sensor displays, electromagnetic resonance, etc.); and the like. In addition, the display 306 can include single-touch or multi-touch sensing capabilities. Any later developed or conventional output display technology can be used for embodiments of the output display, such as LED IPS, OLED, plasma, electronic ink (e.g., electrophoresis, electrowetting, interferometric modulation), and the like. In various embodiments, the resolution of such a display and the resolution of such a touch sensor can be set based on engineering or non-engineering factors (e.g., sales, marketing). In some embodiments, the display 306 can be integrated into the computing device 300 or can be separate. In some embodiments, the display 306 can have almost any size or resolution, such as a 3K resolution display, a microdisplay, one or more separate status or communication lights (e.g., LEDs), and the like.

[0036] In some embodiments of the present invention, the acquisition device 310 may include one or more sensors, drivers, lenses, etc. The sensors may be visible light, infrared and / or UV sensitive sensors, ultrasonic sensors, etc., which are based on any later developed or traditional sensor technology, such as CMOS, CCD, etc. In some embodiments of the present invention, image recognition algorithms, image processing algorithms, or other software programs are run on the processor 302 to process the acquired data. For example, such software can be paired with enabled hardware to provide functions such as: facial recognition (e.g., Face ID, head tracking, camera parameter control, etc.); fingerprint capture / analysis; blood vessel capture / analysis; iris scan capture / analysis; otoacoustic emission (OAE) analysis and matching; and so on.

[0037] In various embodiments, audio input / output 312 may include a microphone / speaker. In various embodiments, speech processing and / or recognition software may be provided to application processor 302 to enable a user to operate computing device 300 by issuing voice commands. In various embodiments of the invention, audio input 312 may provide user input data in the form of spoken words or phrases, etc., as described above. In some embodiments, audio input / output 312 may be integrated into computing device 300 or may be separate.

[0038] In various embodiments, the wired interface 314 can be used to provide data or instruction transmission between the computing device 300 and an external source such as a computer, a remote server, a POS server, a local secure server, a storage network, another computing device 300, an IMU, a camera, etc. Embodiments may include any later developed or traditional physical interface / protocol, such as: USB, micro USB, mini USB, USB-C, Firewire, Apple Lightning connector, Ethernet, POTS, custom interface or base, etc. In some embodiments, the wired interface 314 can also provide power to the power supply 324, etc. In other embodiments, the interface 314 can utilize the close physical contact between the device 300 and the base to transmit data, magnetic energy, thermal energy, light energy, laser energy, etc. In addition, software that allows communication over such a network is generally provided.

[0039] In various embodiments, a wireless interface 316 may also be provided to provide wireless data transmission between the computing device 300 and an external source such as a computer, storage network, headset, microphone, camera, IMU, etc. Figure 3 As shown, the wireless protocols may include Wi-Fi (e.g., IEEE 802.11a / b / g / n, WiMAX), Bluetooth, Bluetooth low energy (BLE), IR, near field communication (NFC), ZigBee, ultra-wide band (UWB), Wi-Fi, mesh communication, etc.

[0040] Various embodiments of the present invention may also include GNSS (e.g., GPS) reception capabilities. Figure 3 The GPS functionality is shown as part of the wireless interface 316 for convenience only, but in implementation, such functionality may be performed by circuitry other than Wi-Fi circuitry, Bluetooth circuitry, etc. In various embodiments of the invention, the GPS receiving hardware may provide user input data in the form of current GPS coordinates, etc., as described above.

[0041] In various embodiments, additional wireless communications may be provided through an RF interface. In various embodiments, the RF interface 320 may support any future developed or legacy RF communication protocols, such as CDMA-based protocols (e.g., WCDMA), GSM-based protocols, HSUPA-based protocols, G4, G5, etc. In some embodiments, various functions are provided on a single IC package such as the Marvel PXA330 processor. As described above, data transmission between the smart device and the service may be performed via Wi-Fi, mesh networks, 4G, 5G, etc.

[0042] Although Figure 3 The functional blocks in the embodiment are shown as independent, but it should be understood that various functions can be recombined into different physical devices. For example, some processors 302 may include Bluetooth functions. In addition, some functions do not need to be included in some blocks, for example, GPS functions do not need to be provided in the provider server.

[0043] In various embodiments, any number of future developed, current operating systems or custom operating systems can be supported, such as iPhone OS (e.g., iOS), Google Android, Linux, Windows, MacOS, etc. In various embodiments of the present invention, the operating system can be a multi-threaded multi-tasking operating system. Therefore, inputs and / or outputs from and to display 306 and inputs and / or outputs to physical sensor 322 can be processed in parallel processing threads. In other embodiments, such events or outputs can be processed serially. In other embodiments of the present invention, inputs and outputs from other functional blocks such as acquisition device 310 and physical sensor 322 can also be processed in parallel or serially.

[0044] In some embodiments of the present invention, physical sensors 322 (e.g., MEMS-based) may include accelerometers, gyroscopes, magnetometers, pressure sensors, temperature sensors, imaging sensors (e.g., blood oxygen, heartbeat, blood vessels, iris data, etc.), thermometers, otoacoustic emission (OAE) testing hardware, etc. Data from such sensors may be used to capture data associated with the device 300 and the user of the device 300. Such data may include physical motion data, pressure data, direction data, etc. The data captured by the sensor 322 may be processed by software running on the processor 302 to determine characteristics of the user, such as gait, gesture performance data, etc., and used for user authentication. In some embodiments, the sensor 322 may also include physical output data, such as vibration, pressure, etc.

[0045] In some embodiments, the power source 324 may be implemented with a battery (e.g., LiPo), a supercapacitor, etc., that provides operating power for the device 300. In various embodiments, the power source 324 may be supplemented or even replaced by any number of power generation technologies such as solar energy, liquid metal power generation, thermoelectric engines, radio frequency collection (e.g., NFC), etc.

[0046] Figure 3 The components that may be used by the processing device are represented. It is easy for a person skilled in the art to understand that many other hardware and software configurations are also applicable to the present invention. Embodiments of the present invention may include Figure 3 At least some of the functional blocks shown, but not necessarily included Figure 3 For example, a processing unit may include Figure 3 Some functional blocks in, but not necessarily including, an accelerometer or other physical sensor 322, an acquisition device 310, an accelerometer 322, an internal power supply 324, etc.

[0047] In view of the above, other variations and modifications may be envisioned by a person skilled in the art. For example, the output of the embodiment may be provided to an autonomous driving system that may manipulate a vehicle (e.g., a car, a drone) based on the predicted flow data; for example, if a feature in the field of view is identified or marked as a pedestrian, the output may be used to provide auditory, visual, or tactile feedback to the user; for example, if the determined product optical flow does not match a predefined standard, the product being manufactured may be identified for further inspection; for example, if the determined optical flow does not match a predefined standard, the robot may be identified as requiring repair; and so on. In addition, in other embodiments, in addition to K-fold cross-validation, other methods may be used to segment real-world data.

[0048] For ease of understanding, the block diagrams and flow charts of the architecture are grouped. However, it should be understood that in alternative embodiments of the present invention, combinations of blocks, addition of new blocks, rearrangement of blocks, etc. may be considered. Therefore, the description and drawings should be regarded as illustrative rather than restrictive. However, it is obvious that various modifications and changes may be made without departing from the broader spirit and scope of the present invention as described in the claims.

Claims

1. A method for a computing system for estimating optical flow, comprising: determining, in a computing system, a first feature map in response to the first image, and determining a second feature map in response to the second image map; implementing a gated recurrent unit (GRU) in the computing system to determine a pixel-level flow prediction in response to the first feature map and the second feature map; determining, in the computing system, a distortion feature map in response to the second image map; implementing a feature level contrast loss function in the computing system to determine a feature level loss in response to a first feature in the first image map and a second feature in the warped feature map; determining, in the computing system, a pixel-level flow loss in response to the pixel-level flow prediction and in response to pixel-level truth data; as well as Parameters of the GRU are modified in the computing system in response to the pixel-level flow loss and the feature-level loss.

2. The method according to claim 1 in, Determining the distortion feature map in the computing system includes: determining the distortion feature map in the computing system also in response to feature-level truth data; and The feature-level true value data includes pre-identified features.

3. The method according to claim 2, further comprising: marking, in the computing system, pseudo features in the feature-level truth data; determining, in the computing system, feature-level pseudo-truth data in response to the marking of the pseudo-feature; determining, in the computing system, a modified warped feature map in response to the feature-level pseudo-truth data; as well as The feature-level contrastive loss function is implemented in the computing system to determine a modified feature-level loss in response to a first feature in the first image map and a third feature in the modified warped feature map.

4. The method according to claim 3, further comprising: determining in the computing system whether the modified feature level loss is less than the feature level loss; as well as Wherein, modifying the parameters of the GRU in the computing system includes: modifying the parameters of the GRU in the computing system in response to the pixel-level flow loss and the modified feature-level loss and in response to determining that the modified feature-level loss is less than the feature-level loss.

5. The method according to claim 3, further comprising: determining in the computing system whether the feature level loss is less than the modified feature level loss; as well as In response to determining in the computing system that the feature level loss is less than the modified feature level loss, a label of the pseudo feature in the feature level truth data is removed.

6. The method according to claim 1, further comprising: determining a correlation volume in response to the first feature map and the second feature map; as well as Wherein, implementing the gated recurrent unit (GRU) in the computing system includes: implementing the gated recurrent unit (GRU) in the computing system to determine the pixel-level flow prediction in response to the first feature map and the correlation body.

7. The method according to claim 6, wherein: The correlation volume includes parameters selected from the group consisting of: a location of a region on an image, a size of a region within an image, a correlation parameter between images, and temporal data.

8. The method according to claim 1 in, The feature-level loss is associated with a plurality of pixels; and The pixel-level flow loss is associated with a pixel in the plurality of pixels.

9. A computing system for estimating optical flow, comprising: a pixel-based analysis system for determining a first feature map in response to a first image and determining a second feature map in response to a second image map, wherein the pixel-based analysis system comprises a gated recurrent unit (GRU) for determining pixel-level flow predictions in response to the first feature map and the second feature map; and a feature-based analysis system coupled to the pixel-based optical flow analysis system, wherein the feature-based analysis system is configured to determine a warp feature map in response to the second image map, wherein the feature-based analysis system includes a contrast loss unit configured to determine a feature-level loss in response to a first feature in the first image map and a second feature in the warp feature map; Wherein, the pixel-based analysis system is used to modify the parameters of the GRU in response to the pixel-level flow loss and the feature-level loss.

10. The computing system according to claim 9 in, The feature-based analysis system is for determining the distortion feature map in response to the second feature map and feature-level truth data; and The feature-level true value data includes pre-identified features.

11. The computing system according to claim 10 in, The feature-based analysis system is used to mark pseudo features in the feature-level truth data; wherein the feature-based analysis system is used to determine feature-level pseudo-truth data in response to the marking of the pseudo-feature; wherein the feature-based analysis system is used to determine a modified distortion feature map in response to the feature-level pseudo-truth data; and The contrast loss unit is used to determine a modified feature level loss in response to a first feature in the first image map and a third feature in the modified distortion feature map.

12. The computing system of claim 11, further comprising: in, The feature-based analysis system is used to determine whether the modified feature level loss is less than the feature level loss; as well as Wherein, the pixel-based analysis system is used to modify the parameters of the GRU in response to the pixel-level flow loss and the modified feature-level loss and in response to determining that the modified feature-level loss is less than the feature-level loss.

13. The computing system of claim 11, further comprising: in, The feature-based analysis system is used to determine whether the feature level loss is less than the corrected feature level loss; as well as The feature-based analysis system is configured to remove the label of the pseudo-feature in the feature-level truth data in response to determining that the feature-level loss is less than the modified feature-level loss.

14. The computing system according to claim 9 in, The pixel-based analysis system is used to determine a correlation volume in response to the first feature map and the second feature map; as well as The gated recurrent unit (GRU) is used to determine the pixel-level flow prediction in response to the first feature map, the second feature map, and the correlation body.

15. The computing system of claim 14, wherein: The correlation volume includes parameters selected from the group consisting of: a location of a region on an image, a size of a region within an image, a correlation parameter between images, and temporal data.

16. The computing system according to claim 9 in, The feature-level loss is associated with a plurality of pixels; and The pixel-level flow loss is associated with a pixel in the plurality of pixels.

17. The computing system of claim 9, wherein: The pixel-based analysis system includes a recurrent total field transformation (RAFT) system.

18. A method comprising: executing an optical flow prediction system including a contrastive loss function in response to the synthetic data set to determine a first optical flow loss; adjusting parameters of the optical flow prediction system in response to the first optical flow loss; Afterwards running the optical flow prediction system in response to a real-world dataset to determine a second optical flow loss; and Parameters of the optical flow prediction system are adjusted in response to the second optical flow loss.

19. The method according to claim 18, wherein: Running the optical flow prediction system in response to the real-world dataset includes: determining a first subset of real-world data from the real-world data set; operating the optical flow prediction system in response to the first subset of the real-world data to determine a third optical flow loss; determining a second subset of real-world data from the real-world data set; operating the optical flow prediction system in response to a second subset of the real-world data to determine a fourth optical flow loss; and The second optical flow loss is determined in response to the third optical flow loss and the fourth optical flow loss.

20. The method according to claim 19 in, The first subset of real-world data includes a plurality of real-world images, the plurality of real-world images including a first real-world image; wherein the first real-world image comprises a plurality of manually labeled features; and Wherein, running the optical flow prediction system in response to the first subset of the real-world data to determine comprises: automatically marking pseudo features in the first real-world image.

Citation Information

Cited By

  • Optical flow calculation method, device and equipment based on comparative learning

    CN121837318A