Parallax correction method and device for fusion of monocular model and binocular model, medium, program product and terminal

By fusing monocular and binocular models, relative disparity labels are generated through feature map extraction and model distillation. Combined with the binocular model, a corrected real disparity map is generated, which solves the problem of poor accuracy of monocular models and improves the accuracy of disparity estimation and the depth information recovery effect.

CN120997054APending Publication Date: 2025-11-21HEILONGJIANG HUIDA TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510892975.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing monocular depth estimation models are not accurate in disparity estimation, cannot directly obtain absolute scale information of the scene, and deep learning binocular disparity datasets are scarce and expensive.

Method used

Feature maps are extracted from monocular images, and relative disparity labels are generated by combining the initial monocular model and the distilled monocular model. These labels are then fused with the binocular model to generate a corrected real disparity map. Model distillation and multimodal feature fusion are used to improve the accuracy of disparity estimation.

Benefits of technology

The scale ambiguity problem of monocular models was solved. By distilling the monocular model to assist the binocular model, the disparity estimation accuracy was improved. This overcame the problems of insufficient labeled data and scale uncertainty, and achieved more accurate disparity estimation and depth information recovery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997054A_ABST
    Figure CN120997054A_ABST
Patent Text Reader

Abstract

The invention provides a monocular model and binocular model fused parallax correction method and device, a medium, a program product and a terminal, and the method comprises the steps: obtaining a first monocular image and a second monocular image, carrying out the feature map extraction operation, and generating a corresponding first feature map and a corresponding second feature map; executing model distillation operation through the initial monocular model to obtain a distillation monocular model; inputting the first feature map to a distillation monocular model, and generating a relative parallax label; inputting the first feature map and the second feature map into a binocular parallax estimation model at the same time to generate a preliminary parallax map; and finally, generating a corrected real disparity map by fusing the initial disparity map and the relative disparity label. The inherent problem of scale blurring of a monocular model is solved, and the parallax estimation precision is improved by assisting a binocular model through a distillation monocular model. Through model distillation and multi-modal feature fusion, the defects of insufficient annotation data and scale uncertainty are overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision, and in particular to methods, devices, media, program products and terminals for disparity correction by fusing monocular and binocular models. Background Technology

[0002] Parallax is a core concept in stereo vision, referring to the horizontal displacement difference between corresponding points in the left and right images. This displacement difference is fundamental for calculating depth information and reconstructing 3D scenes. Parallax accuracy directly impacts the performance of fields such as autonomous driving, robot navigation, augmented reality, and 3D reconstruction. Currently, deep learning-based binocular disparity estimation algorithms heavily rely on supervised learning.

[0003] However, publicly available binocular parallax datasets are scarce and have limitations. Acquiring labeled data through LiDAR or virtual engines is costly. To address the limitations of labeled data, the Depth Anything model was developed. As the current state-of-the-art monocular depth estimation algorithm, the Depth Anything model significantly improves its robustness by training with a combination of massive amounts of unlabeled images and partially labeled images, making it adaptable to various real-world scenarios. Despite the significant progress made in monocular depth estimation models, monocular depth estimation is inherently limited by its input information and cannot directly obtain the absolute scale information of the scene. Summary of the Invention

[0004] In view of the shortcomings of the prior art described above, the purpose of this application is to provide a disparity correction method, device, medium, program product and terminal that fuses monocular and binocular models, in order to solve the problem of poor accuracy in existing disparity estimation methods, such as the monocular depth estimation model of Depth Anything.

[0005] To achieve the above and other related objectives, a first aspect of this application provides a disparity correction method for fusing monocular and binocular models, comprising: acquiring a first monocular image and a second monocular image; performing a feature map extraction operation on the first monocular image to generate a first feature map; performing a feature map extraction operation on the second monocular image to generate a second feature map; acquiring a preset initial monocular model; performing a model distillation operation on the initial monocular model to generate a distilled monocular model; inputting the first feature map into the distilled monocular model to generate a relative disparity label; inputting the first feature map and the second feature map into a binocular model to generate a preliminary disparity map; and fusing the preliminary disparity map and the relative disparity label to generate a corrected true disparity map.

[0006] In some embodiments of the first aspect of this application, the process of inputting the first feature map into the distillation monocular model to generate a relative disparity label includes: inputting the first feature map into an initial monocular model to generate a first relative disparity label; inputting the first feature map into the distillation monocular model to generate a second relative disparity label; constructing a loss function based on the first relative disparity label and the second relative disparity label; and optimizing the distillation monocular model based on the loss function to generate an optimized second relative disparity label.

[0007] In some embodiments of the first aspect of this application, the process of fusing the preliminary disparity map and the relative disparity label to generate a corrected true disparity map includes: inputting the first feature map and the second feature map into a stereo model to generate a scale factor; and multiplying the second relative disparity label by the scale factor to generate the corrected true disparity map.

[0008] In some embodiments of the first aspect of this application, the process of inputting the first feature map into the distillation monocular model to generate a relative disparity label includes: inputting the first feature map into the distillation monocular model to generate a third relative disparity label.

[0009] In some embodiments of the first aspect of this application, the process of fusing the preliminary disparity map and the relative disparity label to generate a corrected true disparity map includes: performing a cost body construction operation on the first feature map and the second feature map to generate a matching cost body; performing a cost aggregation operation on the matching cost body to generate an aggregated cost body; and performing a disparity tuning operation on the aggregated cost body and the third relative disparity label to generate a corrected true disparity map.

[0010] In some embodiments of the first aspect of this application, the process of performing a model distillation operation on the initial monocular model to generate a distilled monocular model includes: obtaining a preset initial student model; inputting the first feature map into the initial monocular model to generate teacher output; inputting the first feature map into the initial student model to generate student output; calculating the distillation loss between the teacher output and the student output; and optimizing the student model based on the distillation loss to generate a distilled monocular model.

[0011] To achieve the above and other related objectives, a second aspect of this application provides a disparity correction device for fusing monocular and binocular models, comprising: a feature extraction module for acquiring a first monocular image and a second monocular image, performing a feature map extraction operation on the first monocular image to generate a first feature map, and performing a feature map extraction operation on the second monocular image to generate a second feature map; a disparity calculation module for acquiring a preset initial monocular model, performing a model distillation operation on the initial monocular model to generate a distilled monocular model, and inputting the first feature map into the distilled monocular model to generate a relative disparity label; a preliminary disparity module for inputting the first feature map and the second feature map into a binocular model to generate a preliminary disparity map; and a disparity optimization module for fusing the preliminary disparity map and the relative disparity label to generate a corrected true disparity map.

[0012] To achieve the above and other related objectives, a third aspect of this application provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements a disparity correction method for fusing the monocular and binocular models.

[0013] To achieve the above and other related objectives, a fourth aspect of this application provides a computer program product comprising computer program code that, when executed on a computer, enables the computer to implement the disparity correction method for fusing the monocular and binocular models.

[0014] To achieve the above and other related objectives, a fifth aspect of this application provides an electronic terminal, including a memory, a processor, and a computer program stored in the memory; the processor executes the computer program to implement the method.

[0015] As described above, the disparity correction method, apparatus, medium, program product, and terminal for fusing monocular and binocular models in this application have the following beneficial effects: By acquiring a first monocular image and a second monocular image, feature map extraction operations are performed respectively to generate corresponding first and second feature maps. Further, a model distillation operation is performed on a preset initial monocular model to obtain a distilled monocular model. The first feature map is input into the distilled monocular model to generate a relative disparity label; simultaneously, the first and second feature maps are input into a binocular disparity estimation model to generate a preliminary disparity map. Finally, by fusing the preliminary disparity map and the relative disparity label, a corrected true disparity map is generated, achieving more accurate disparity estimation and depth information recovery. This solves the inherent scale ambiguity problem of monocular models and improves disparity estimation accuracy by assisting binocular models through distilled monocular models. Through model distillation and multimodal feature fusion, the defects of insufficient labeled data and scale uncertainty are overcome. Attached Figure Description

[0016] Figure 1 The diagram shows a flowchart of an embodiment of the disparity correction method for fusing monocular and binocular models according to this application.

[0017] Figure 2 The diagram shows a schematic representation of the model in one embodiment of the disparity correction method for fusing monocular and binocular models according to this application.

[0018] Figure 3 This illustration shows a schematic diagram of the model structure in another embodiment of the disparity correction method for fusing monocular and binocular models according to this application.

[0019] Figure 4 This diagram shows a structural schematic of an embodiment of the disparity correction device for fusing monocular and binocular models according to this application.

[0020] Figure 5 This diagram shows a structural schematic of an embodiment of a disparity correction terminal that fuses monocular and binocular models according to this application. Detailed Implementation

[0021] The following specific examples illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. This application can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. It should be noted that, unless otherwise specified, the following embodiments and features in the embodiments can be combined with each other.

[0022] Before providing a further detailed description of the present invention, the nouns and terms used in the embodiments of the present invention are explained, and the nouns and terms used in the embodiments of the present invention are subject to the following interpretations:

[0023] <1> Monocular Image: A monocular image is an image captured from a single viewpoint or by a single camera. Such images do not contain depth information; depth inference requires other cues (such as perspective, occlusion, texture gradients, etc.) or prior knowledge.

[0024] <2> Binocular images: Binocular images refer to two images (left-eye image and right-eye image) of the same scene captured simultaneously by two cameras at a fixed distance (baseline) in the horizontal direction. By analyzing the parallax between the two images, the depth information of the scene can be recovered.

[0025] <3> Monocular models: Monocular models refer to models that use only monocular images as input for feature extraction, semantic understanding, or task prediction (such as depth estimation, object detection, etc.). Due to the limited input information, monocular models are more challenging than binocular models when handling depth-related tasks, but they have a wider range of applications.

[0026] <4> Binocular model: A binocular model refers to a model that uses binocular images as input to process tasks requiring depth information (such as binocular depth estimation, 3D reconstruction, etc.). Binocular models can directly utilize the disparity information of the left and right images to obtain the depth of a scene.

[0027] <5> Distillation models: Within the knowledge distillation framework, distillation models are small, efficient models trained by learning from the output (usually soft labels or features) of a "teacher model." Their goal is to significantly reduce model size and inference time while maintaining high performance.

[0028] <6> Teacher Model: In knowledge distillation, the teacher model refers to a pre-trained model with a large number of parameters and superior performance. It is responsible for generating "knowledge" (such as predicted probability distributions, intermediate features, etc.) to guide the learning of the "student model".

[0029] <7> Student model: In knowledge distillation, the student model refers to a model with fewer parameters and higher computational efficiency. It learns knowledge by imitating the behavior of the teacher model in order to achieve performance close to that of the teacher model.

[0030] <8> Disparity Label: In binocular vision, the disparity label refers to the true disparity value corresponding to each pixel in a binocular image. Disparity is the difference in pixel position of the same point in two images, and it is inversely proportional to depth. It is an important supervisory signal for training binocular depth estimation algorithms.

[0031] <9> Loss function: In machine learning and deep learning, the loss function is used to quantify the difference between the model's predicted values ​​and the true values. The goal of model training is to minimize the value of the loss function through optimization algorithms (such as gradient descent), thereby enabling the model to learn the correct mapping relationship.

[0032] <10> Cost aggregation operation: In stereo matching algorithms, cost aggregation refers to the process of summarizing or propagating local matching cost information within a certain region. Its purpose is to resolve the ambiguity and uncertainty of local matching, making the matching cost more spatially smooth and robust, thereby obtaining a more accurate disparity map. Common aggregation methods include box filtering, tree aggregation, and graph cut-based aggregation.

[0033] <11> Aggregated Cost Volume: In stereo matching, the aggregated cost volume refers to the set of matching costs obtained after cost aggregation. It is a three-dimensional data structure represented as width × height × disparity range, where each element represents the aggregated matching cost at a specific pixel location and a specific disparity value.

[0034] <12> Matching cost volume: In stereo matching, the matching cost volume refers to the set formed by calculating the initial matching cost of each pixel under different disparity assumptions. It is a three-dimensional data structure, represented as width × height × disparity range, where each element quantifies the similarity or dissimilarity between a pixel in the left image and its corresponding disparity pixel in the right image.

[0035] To facilitate understanding of the embodiments of this application, firstly, in conjunction with Figure 1 Detailed explanation. Figure 1 This document illustrates a flowchart of a disparity correction method for fusing monocular and binocular models according to an embodiment of the present invention. The disparity correction method for fusing monocular and binocular models in this embodiment mainly includes the following steps:

[0036] Step S11: Obtain a first monocular image and a second monocular image; perform a feature map extraction operation on the first monocular image to generate a first feature map; and perform a feature map extraction operation on the second monocular image to generate a second feature map.

[0037] In one embodiment of this application, the first monocular image refers to the left-side image captured by a single camera, and the second monocular image refers to the right-side image captured by a single camera. These two images are used to obtain a two-dimensional representation of the same scene. Devices for acquiring monocular images include digital cameras, smartphones, etc., and the images are generally stored in formats such as JPEG and PNG for subsequent processing in this application. Feature extraction operations in this embodiment include, but are not limited to, edge detection (such as the Canny operator), corner detection (such as Harris corners), and convolutional neural networks based on deep learning. This application itself does not limit the feature extraction algorithm; it is only used here for illustrative purposes.

[0038] Step S12: Obtain a preset initial monocular model, perform a model distillation operation on the initial monocular model to generate a distilled monocular model; input the first feature map into the distilled monocular model to generate a relative disparity label.

[0039] In one embodiment of this application, a preset initial monocular model is first obtained. This initial monocular model can be any pre-trained monocular depth estimation model to provide a benchmark for depth estimation. Subsequently, a model distillation operation is performed on the initial monocular model to generate a distilled monocular model. It should be noted that model distillation is a knowledge transfer technique aimed at transferring knowledge from a larger, higher-performing model to a smaller, more easily deployed model. In this embodiment, through the distillation operation, knowledge from the initial monocular model is transferred to the distilled monocular model, enabling the distilled monocular model to reduce model complexity and improve computational efficiency while maintaining or even improving accuracy. Specific distillation operations can employ various methods, such as probability distribution-based distillation methods or feature map-based distillation methods.

[0040] Subsequently, after generating the distillation monocular model, the first feature map is input into it. The first feature map is an image feature representation obtained by passing the input image through a feature extraction network. The distillation monocular model receives the first feature map as input and outputs a relative disparity label. The relative disparity label is an encoding of the depth information of each pixel in the input image relative to the camera position. The format of the relative disparity label includes, but is not limited to, representing the disparity value using a floating-point number format or using a discretized disparity level. The relative disparity label is used to provide depth information for each pixel in the image, thereby enabling applications such as 3D scene reconstruction, virtual reality, and augmented reality. By using the distillation monocular model to generate relative disparity labels, computational resource consumption can be reduced and the real-time performance of depth estimation can be improved while ensuring the accuracy of depth estimation.

[0041] In one embodiment of this application, the process of inputting the first feature map into the distillation monocular model to generate a relative disparity label includes: inputting the first feature map into an initial monocular model to generate a first relative disparity label; inputting the first feature map into the distillation monocular model to generate a second relative disparity label; constructing a loss function based on the first relative disparity label and the second relative disparity label; and optimizing the distillation monocular model based on the loss function to generate an optimized second relative disparity label.

[0042] In this embodiment, the relative distance between objects and the observer is accurately reflected by analyzing the displacement of different objects in the image. The relative disparity label is calculated based on disparity from different viewing angles to assess the depth of objects in the scene. This step includes: inputting the image feature map into an initial monocular model to generate a first relative disparity label. This label contains the model's depth understanding of various objects in the scene. Subsequently, the same feature map is input into a distilled monocular model to generate a second relative disparity label. In the construction of the loss function, the model performance is evaluated by comparing the difference between the first and second relative disparity labels and applying the loss function.

[0043] Specifically, the loss function quantifies the prediction discrepancy by calculating the mean squared error or other error metric between the two values. This loss function guides the tuning and optimization of model parameters. Using the backpropagation algorithm, the value of the loss function is calculated in each iteration, and the model updates its parameters based on the gradient of the loss value to reduce prediction error. After each iteration, the model's prediction updates gradually reduce the loss, and the model's performance continuously improves. After multiple rounds of iteration and adjustment, the distilled monocular model can generate more accurate second relative disparity labels, significantly improving prediction accuracy in depth estimation tasks.

[0044] In one embodiment of this application, the process of inputting the first feature map into the distillation monocular model to generate a relative disparity label includes: inputting the first feature map into the distillation monocular model to generate a third relative disparity label.

[0045] In this embodiment, once the distillation monocular model has been trained and possesses inference capabilities, the first input feature map is directly fed into the trained distillation monocular model. After forward inference processing, the model outputs a third relative disparity label corresponding to this input feature map. The third relative disparity label is used to characterize the relative depth or distance distribution of each pixel in the current input feature map, providing basic data support for downstream 3D reconstruction, scene understanding, or other visual analysis tasks.

[0046] Step S13: Input the first feature map and the second feature map into the stereo model to generate a preliminary disparity map.

[0047] In this embodiment, the first monocular image refers to the image captured by the left camera, while the second monocular image is the image captured by the right camera. These two images provide different perspectives of the same scene, and the feature maps generated after feature extraction contain key depth information of the scene. The first and second feature maps originate from the left and right perspectives of the same scene, and there is a pixel-level correspondence between them, exhibiting high correlation and comparability. In this embodiment, the stereo model includes, but is not limited to: traditional stereo algorithms based on stereo matching (such as SAD, SSD), deep convolutional neural networks (such as DispNetC, PSMNet), attention-based stereo networks (such as GANet), Transformer-based stereo models, and hybrid architecture depth estimation models. When these two feature maps are input into the stereo model, the model analyzes the disparity between them to calculate depth information and generate a preliminary disparity map.

[0048] Step S14: Fuse the preliminary disparity map and the relative disparity label to generate a corrected true disparity map.

[0049] In one embodiment of this application, the disparity label generation process involves feeding the input feature maps into a distillation monocular model and another distillation monocular model to obtain their respective disparity labels. Then, the performance of the distillation monocular model is evaluated by calculating the loss between these two labels, and the model is optimized based on this evaluation to generate more accurate and optimized relative disparity labels. At this point, the process of fusing the preliminary disparity map and the relative disparity labels to generate a corrected true disparity map includes: inputting the first feature map and the second feature map into a stereo model to generate a scale factor; and multiplying the second relative disparity label by the scale factor to generate the corrected true disparity map.

[0050] The model in this embodiment is as follows Figure 2 As shown, by feeding the input feature maps into two distilled monocular models, each model independently generates disparity labels. The model performance is then evaluated and optimized by calculating the loss between these two labels. During the model optimization phase, the monocular model is fine-tuned based on the loss function to generate more accurate and optimized relative disparity labels. The scale factor is S = [s1, s2, ..., s]. n A multidimensional vector, where each s i The scaling factor represents a scaling factor for a specific image region or feature point. The generation and optimization of the scaling factor involves in-depth analysis of the structural features of the first and second feature maps using a binocular model, to calculate the depth information and spatial relationships of different regions with finer granularity.

[0051]

[0052] Among them, D trueD represents the true depth map. pred This represents the predicted depth map. The core task of the disparity optimization module is to iteratively optimize the generated scale factor through backpropagation algorithm, gradually adjusting it to approximate the optimal scaling factor. This scale factor serves as an adjustment parameter used to weight the relative disparity labels. During the disparity map correction stage, the second relative disparity label is multiplied by this scale factor to generate the corrected final disparity result. It is worth emphasizing that the optimization module itself does not directly output the final disparity map, but rather outputs the scale factor. The dynamic adjustment of this factor enables fine-grained correction of the disparity map, thereby effectively improving the accuracy and robustness of the overall disparity estimation. The unique feature of the multi-model collaborative method proposed in this application is that, through iterative optimization of the monocular model and fine-grained adjustment of the binocular model, the accuracy and reliability of the disparity map can be gradually improved, effectively overcoming the performance limitations of traditional single models in complex scenes, making the entire disparity map generation process more intelligent and dynamic.

[0053] In one embodiment of this application, the process of inputting the first feature map into the distillation monocular model to generate a relative disparity label includes: after inputting the first feature map into the distillation monocular model to generate a third relative disparity label, fusing the preliminary disparity map and the relative disparity label to generate a corrected true disparity map includes: performing a cost body construction operation on the first feature map and the second feature map to generate a matching cost body; performing a cost aggregation operation on the matching cost body to generate an aggregated cost body; and performing a disparity tuning operation on the aggregated cost body and the third relative disparity label to generate a corrected true disparity map.

[0054] The model structure in this embodiment is as follows: Figure 3 As shown, in the feature extraction and monocular model distillation stage, relative disparity labels are generated by feeding the input first feature map into the monocular model distillation. This stage can employ models such as the ResNet series (ResNet-50 or ResNet-101), EfficientNet, or Vision Transformer (ViT) with a self-attention mechanism. The monocular model for distillation can be a model with good monocular depth estimation performance, such as MiDaS, MonoDepth2, or BTS. Next, in the cost volume construction and matching stage, cost volume construction is performed on the first and second feature maps to generate matching cost volumes, capturing the correspondences and subtle differences between images. This stage uses stereo matching networks based on cost volumes, such as GwcNet, GANet, and PSMNet. These networks achieve high-quality cost volume generation through multi-scale feature matching and normalization processing.

[0055] Subsequently, the cost body undergoes cost aggregation to produce a more stable and comprehensive aggregated cost body, which helps improve matching accuracy. Finally, in the disparity tuning stage, the aggregated cost body and the previously generated relative disparity labels are input into the disparity tuning module. This module can utilize optimized networks such as DeepLabV3+, U-Net variants, or networks with fused attention mechanisms to finely refine the disparity map through model architectures such as multi-scale feature fusion, residual connections, and attention weight adjustment, generating high-precision final disparity results. This process employs an end-to-end learnable multi-model collaborative approach, combining model distillation and multi-stage fine-grained optimization strategies to achieve dynamic adjustment and adaptive feature extraction of the disparity map. This effectively overcomes the sensitivity of traditional methods to occlusion and texture changes in complex scenes, thereby significantly improving the accuracy and robustness of disparity estimation.

[0056] In the embodiments of this application, terms such as "first" and "second" are used to distinguish identical or similar items with essentially the same function and effect. For example, "first monocular image" and "second monocular image" are used only to distinguish different monocular images and do not limit their order. Those skilled in the art will understand that terms such as "first" and "second" do not limit the quantity or execution order, and that "first" and "second" do not necessarily imply that they are different.

[0057] It should be noted that, in the embodiments of this application, the words "exemplary" or "for example" indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0058] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.

[0059] Figure 4 This is a schematic block diagram of the disparity correction device 400 for fusing monocular and binocular models provided in an embodiment of this application. Figure 4As shown, the device includes a feature extraction module 401, a disparity calculation module 402, a preliminary disparity module 403, and a disparity optimization module 404.

[0060] Feature extraction module 401: used to acquire a first monocular image and a second monocular image, perform feature map extraction operation on the first monocular image to generate a first feature map, and perform feature map extraction operation on the second monocular image to generate a second feature map;

[0061] The disparity calculation module 402 is used to obtain a preset initial monocular model, perform a model distillation operation on the initial monocular model to generate a distilled monocular model, and input the first feature map into the distilled monocular model to generate a relative disparity label.

[0062] Preliminary disparity module 403: used to input the first feature map and the second feature map into the stereo model to generate a preliminary disparity map;

[0063] Parallax optimization module 404: used to fuse the preliminary parallax map and the relative parallax label to generate a corrected real parallax map.

[0064] It should be understood that the specific process of each module performing the above-mentioned steps has been described in detail in the above method embodiments, and will not be repeated here for the sake of brevity.

[0065] It should also be understood that the module division in the embodiments of this application is illustrative and only represents a logical functional division; in actual implementation, there may be other division methods. Furthermore, the functional modules in the various embodiments of this application can be integrated into a single processor, exist as separate physical entities, or be integrated into a single module. The integrated modules described above can be implemented in hardware or as software functional modules.

[0066] Figure 5 This is a schematic block diagram of the electronic terminal provided in an embodiment of this application. Figure 5 As shown, the electronic terminal includes at least one processor 501, a memory 502, at least one network interface 503, and a user interface 505. The various components in the device are coupled together via a bus system 504. It is understood that the bus system 504 is used to implement communication between these components. In addition to a data bus, the bus system 504 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in… Figure 5 The general will label all buses as bus systems.

[0067] The user interface 505 may include a monitor, keyboard, mouse, trackball, clicker, button, touchpad, or touch screen.

[0068] It is understood that memory 502 can be volatile memory or non-volatile memory, or both. Non-volatile memory can be read-only memory (ROM) or programmable read-only memory (PROM), used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM) and synchronous static random access memory (SSRAM). The memories described in the embodiments of this invention are intended to include, but are not limited to, these and any other suitable categories of memory.

[0069] In this embodiment of the invention, the memory 502 is used to store various types of data to support the operation of the electronic terminal 500. Examples of this data include: any executable program for operation on the electronic terminal 500, such as the operating system 5021 and application programs 5022; the operating system 5021 contains various system programs, such as the framework layer, core library layer, driver layer, etc., for implementing various basic services and handling hardware-based tasks. The application program 5022 may contain various applications, such as a media player, browser, etc., for implementing various application services. The disparity correction method for the fusion of monocular and binocular models provided in this embodiment of the invention can be included in the application program 5022.

[0070] The methods disclosed in the above embodiments of the present invention can be applied to or implemented by processor 501. Processor 501 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in processor 501 or by instructions in the form of software. The processor 501 may be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Processor 501 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. General-purpose processor 501 may be a microprocessor or any conventional processor, etc. The steps of the accessory optimization method provided in the embodiments of the present invention can be directly reflected as being executed by a hardware decoding processor, or being executed by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium, which is located in a memory. The processor reads the information in the memory and combines it with its hardware to complete the steps of the aforementioned method.

[0071] In an exemplary embodiment, the electronic terminal 500 may be used by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), or complex programmable logic devices (CPLDs) to execute the aforementioned method.

[0072] According to the method provided in the embodiments of this application, this application also provides a computer program product, which includes: computer program code, which, when run on a computer, causes the computer to execute the disparity correction method for fusing monocular and binocular models as described in any of the embodiments above.

[0073] According to the method provided in the embodiments of this application, this application also provides a computer-readable storage medium storing program code. When the program code is run on a computer, it causes the computer to execute the disparity correction method for fusing monocular and binocular models as described in any of the embodiments above.

[0074] As used in this specification, the terms "component," "module," "system," etc., are used to refer to computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or a computer. As illustrated, applications running on computing devices and computing devices can both be components. One or more components may reside in a process and / or an execution thread, and components may be located on a single computer and / or distributed among two or more computers. Furthermore, these components can be executed from various computer-readable media on which various data structures are stored. Components can communicate, for example, via local and / or remote processes based on signals having one or more data packets (e.g., data from two components interacting with another component between a local system, a distributed system, and / or a network, such as the Internet interacting with other systems via signals).

[0075] Those skilled in the art will recognize that the various illustrative logical blocks and steps described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.

[0076] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0077] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0078] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0079] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0080] In the above embodiments, the functions of each functional unit can be implemented entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. A computer program product includes one or more computer instructions (programs). When the computer program instructions (programs) are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, Digital Subscriber Line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., high-density digital video discs (DVDs)), or semiconductor media (e.g., solid-state drives (SSDs)).

[0081] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0082] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0083] In summary, this application provides a disparity correction method, apparatus, medium, program product, and terminal for fusing monocular and binocular models. It acquires a first monocular image and a second monocular image, performs feature map extraction operations on each, and generates corresponding first and second feature maps. Further, it performs model distillation on a preset initial monocular model to obtain a distilled monocular model. The first feature map is input into the distilled monocular model to generate a relative disparity label; simultaneously, the first and second feature maps are input into a binocular disparity estimation model to generate a preliminary disparity map. Finally, by fusing the preliminary disparity map and the relative disparity label, a corrected true disparity map is generated, achieving more accurate disparity estimation and depth information recovery. This solves the inherent scale ambiguity problem of monocular models and improves disparity estimation accuracy by using distilled monocular models to assist binocular models. Through model distillation and multimodal feature fusion, it overcomes the shortcomings of insufficient labeled data and scale uncertainty. Therefore, this application effectively overcomes various shortcomings of existing technologies and has high industrial application value.

[0084] The above embodiments are merely illustrative of the principles and effects of this application and are not intended to limit this application. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in this application should still be covered by the claims of this application.

Claims

1. A disparity correction method of fusing a monocular model and a binocular model, characterized in that, The method comprises the following steps: obtaining a first monocular image and a second monocular image, performing a feature map extraction operation on the first monocular image to generate a first feature map, and performing a feature map extraction operation on the second monocular image to generate a second feature map; obtaining a preset initial monocular model, performing a model distillation operation on the initial monocular model to generate a distilled monocular model, inputting the first feature map into the distilled monocular model to generate a relative disparity label, and inputting the first feature map and the second feature map into a binocular model to generate a preliminary disparity map; fusing the preliminary disparity map and the relative disparity label to generate a corrected real disparity map. The process of inputting the first feature map into the distilled monocular model to generate a relative disparity label comprises:

2. The method according to claim 1, wherein, inputting the first feature map into an initial monocular model to generate a first relative disparity label; inputting the first feature map into the distilled monocular model to generate a second relative disparity label; constructing a loss function based on the first relative disparity label and the second relative disparity label; optimizing the distilled monocular model based on the loss function to generate an optimized second relative disparity label. The process of fusing the preliminary disparity map and the relative disparity label to generate a corrected real disparity map comprises:

3. The method according to claim 2, wherein, inputting the first feature map and the second feature map into a binocular model to generate a scale factor; multiplying the second relative disparity label by the scale factor to generate a corrected real disparity map. The process of inputting the first feature map into the distilled monocular model to generate a relative disparity label comprises: inputting the first feature map into the distilled monocular model to generate a third relative disparity label.

4. The method of claim 1, wherein the monocular model and the binocular model are fused. The process of fusing the preliminary disparity map and the relative disparity label to generate a corrected real disparity map comprises:

5. The method according to claim 4, wherein, performing a cost volume construction operation on the first feature map and the second feature map to generate a matching cost volume; performing a cost aggregation operation on the matching cost volume to generate an aggregated cost volume; performing a disparity fine-tuning operation on the aggregated cost volume and the third relative disparity label to generate a corrected real disparity map. The process of performing a model distillation operation on the initial monocular model to generate a distilled monocular model comprises:

6. The method according to any one of claims 1 to 5, wherein, obtaining a preset initial student model; inputting the first feature map into the initial monocular model to generate a teacher output; inputting the first feature map into the initial student model to generate a student output; calculating a distillation loss between the teacher output and the student output; optimizing the student model based on the distillation loss to generate a distilled monocular model. The method comprises the following steps:

7. A disparity correction device of monocular model and binocular model fusion, characterized in that, a feature extraction module for obtaining a first monocular image and a second monocular image, performing a feature map extraction operation on the first monocular image to generate a first feature map, and performing a feature map extraction operation on the second monocular image to generate a second feature map; a disparity calculation module for obtaining a preset initial monocular model, performing a model distillation operation on the initial monocular model to generate a distilled monocular model, inputting the first feature map into the distilled monocular model to generate a relative disparity label, and inputting the first feature map and the second feature map into a binocular model to generate a preliminary disparity map; ​ a preliminary disparity module configured to input the first feature map and the second feature map into a binocular model to generate a preliminary disparity map; a disparity fine-tuning module configured to fuse the preliminary disparity map and the relative disparity label to generate a corrected true disparity map.

8. A computer-readable storage medium having stored thereon a computer program, characterized in that The computer program, when executed by a processor, implements the disparity correction method of fusing the monocular model and the binocular model according to any one of claims 1 to 6.

9. A computer program product, characterised in that, The computer program product comprises computer program code, and when the computer program code is executed on a computer, the computer is caused to implement the disparity correction method of fusing the monocular model and the binocular model according to any one of claims 1 to 6.

10. An electronic terminal comprising a memory, a processor and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the disparity correction method of fusing the monocular model and the binocular model according to any one of claims 1 to 6.