Binocular data generation method and system, electronic equipment and storage medium
By performing depth estimation, parallax conversion and optimization processing on the monocular data set, binocular data with higher diversity and efficiency are generated, and the problems of insufficient data diversity and poor generalization in the prior art are solved.
Patent Information
- Application Number
- CN202510021396.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-30
AI Technical Summary
In the prior art, binocular data is generated based on fixed camera parameters, resulting in insufficient data diversity and poor generalization, which affects the generation efficiency of binocular data.
By acquiring the monocular data set, depth estimation processing, parallax conversion processing, parallax optimization processing are performed, and binocular data is generated based on the target parallax, improving the diversity and generation efficiency of the data.
The generated binocular data has higher diversity and parallax labeling quality, which improves the efficiency of binocular data generation.
Smart Images

Figure CN120070733A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technologies, and in particular, to a binocular data generation method, system, electronic device, and storage medium. Background Art
[0002] In related technologies, a pair of binocular data is obtained through a set of fixed camera parameters, that is, binocular data is generated based on a binocular camera, and binocular data with disparity annotations is obtained by calibrating point cloud data. However, in practical applications, it is found that in related technologies, obtaining a pair of binocular data based on a certain set of fixed camera parameters will, to a certain extent, result in insufficient data diversity, causing the network to rely on camera parameters, and having poor generalization performance when the camera parameters are inconsistent during testing, which affects the generation efficiency of binocular data.
[0003] In summary, the technical problems existing in related technologies need to be improved. Summary of the Invention
[0004] The main purpose of the embodiments of this application is to propose a binocular data generation method, system, electronic device, and storage medium, which can improve the generation efficiency of binocular data.
[0005] To achieve the above object, on the one hand, an embodiment of this application proposes a binocular data generation method, and the method includes:
[0006] Obtain a monocular data set, where the monocular data set includes a plurality of monocular data;
[0007] Perform depth estimation processing on the monocular data to obtain an estimated depth;
[0008] Perform disparity conversion processing on the estimated depth to obtain an initial disparity;
[0009] Perform disparity optimization processing on the initial disparity to obtain a target disparity;
[0010] Perform projection processing on the monocular data according to the target disparity to obtain a projection image;
[0011] Generate binocular data according to the target disparity, the monocular data, and the projection image.
[0012] In some embodiments, the performing disparity conversion processing on the estimated depth to obtain an initial disparity includes the following steps:
[0013] Perform maximum normalization processing on the estimated depth to obtain a normalized disparity;
[0014] Multiply the normalized disparity according to a scaling factor to obtain the initial disparity, where the scaling factor is obtained by randomly and uniformly sampling a preset range.
[0015] In some embodiments, the process of performing parallax optimization on the initial parallax to obtain the target parallax includes the following steps:
[0016] Calculate the spatial Gaussian weight and the intensity Gaussian weight based on the initial parallax;
[0017] Perform weighted processing on the initial parallax according to the spatial Gaussian weight and the intensity Gaussian weight to obtain the target parallax.
[0018] In some embodiments, the process of calculating the spatial Gaussian weight and the intensity Gaussian weight based on the initial parallax includes the following steps:
[0019] Determine the current pixel and neighboring pixels based on the initial parallax;
[0020] Perform distance calculation processing on the neighboring pixels according to the current pixel to obtain the spatial distance;
[0021] Determine the neighboring parallax according to the neighboring pixels, and perform difference calculation processing on the neighboring parallax according to the initial parallax to obtain the parallax difference;
[0022] Perform normalization processing on the spatial distance and the parallax difference respectively according to the Gaussian function to obtain the spatial Gaussian weight and the intensity Gaussian weight.
[0023] In some embodiments, the process of performing projection processing on the monocular data according to the target parallax to obtain the projected image includes the following steps:
[0024] Determine the pixel distance difference according to the target parallax;
[0025] Perform color value mapping processing on the pixels of the monocular data according to the pixel distance difference to obtain the initial image;
[0026] Perform perspective difference optimization processing on the initial image to obtain the projected image.
[0027] In some embodiments, the process of performing perspective difference optimization processing on the initial image to obtain the projected image includes the following steps:
[0028] Extract multiple mapping values of the same pixel point from the initial image;
[0029] Perform parallax comparison processing on the multiple mapping values, and select and retain the mapping value with the largest parallax to obtain the projected image.
[0030] In some embodiments, the process of performing perspective difference optimization processing on the initial image to obtain the projected image further includes the following steps:
[0031] Perform region extraction processing on the initial image to obtain a non-filled region;
[0032] Perform mask calculation processing on the non-filled region to obtain a region mask;
[0033] Input the initial image and the region mask into an image completion large model for prediction and generation processing to obtain the projection image.
[0034] To achieve the above object, another aspect of the embodiments of the present application proposes a binocular data generation system, the system includes:
[0035] A first module, configured to obtain a monocular data set, the monocular data set includes a plurality of monocular data;
[0036] A second module, configured to perform depth estimation processing on the monocular data to obtain an estimated depth;
[0037] A third module, configured to perform parallax conversion processing on the estimated depth to obtain an initial parallax;
[0038] A fourth module, configured to perform parallax optimization processing on the initial parallax to obtain a target parallax;
[0039] A fifth module, configured to perform projection processing on the monocular data according to the target parallax to obtain a projection image;
[0040] A sixth module, configured to generate binocular data according to the target parallax, the monocular data and the projection image.
[0041] To achieve the above object, another aspect of the embodiments of the present application proposes an electronic device, the electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the foregoing method is implemented.
[0042] To achieve the above object, another aspect of the embodiments of the present application proposes a computer-readable storage medium, the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the foregoing method is implemented.
[0043] The embodiments of the present application at least include the following beneficial effects: The present application provides a binocular data generation method, system, electronic device, and storage medium. This solution obtains a monocular data set, which includes multiple monocular data. Then, depth estimation processing is performed on the monocular data to obtain the estimated depth. By performing disparity conversion processing on the estimated depth, an initial disparity can be obtained, and different disparity results can be generated, enriching the diversity of binocular data. This solution also performs disparity optimization processing on the initial disparity to obtain the target disparity, performs projection processing on the monocular data according to the target disparity to obtain the projection image, and finally generates binocular data based on the target disparity, monocular data, and projection image, which can optimize the initial disparity, improve the quality of the generated disparity annotation, and the efficiency of binocular data generation. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 is a flowchart of a binocular data generation method provided by an embodiment of the present application;
[0045] Figure 2 is Figure 1 a flowchart of step S103 in
[0046] Figure 3 is Figure 1 a flowchart of step S104 in
[0047] Figure 4 is Figure 1 a flowchart of step S105 in
[0048] Figure 5 is a schematic structural diagram of a binocular data generation system provided by an embodiment of the present application;
[0049] Figure 6 is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are only examples of systems and methods that are consistent with some aspects of the embodiments of the present application as detailed in the appended claims.
[0051] It can be understood that the terms "first", "second", etc. used in this application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if", "when" as used herein may be interpreted as "when...", "while...", or "in response to determining".
[0052] The terms "at least one", "a plurality of", "each", "any one", etc. used in this application, at least one includes one, two, or more than two, a plurality of includes two or more than two, each refers to each of the corresponding plurality, and any one refers to any one of the plurality.
[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0054] Before elaborating on the embodiments of this application in detail, some nouns and terms involved in the embodiments of this application are first explained, and the nouns and terms involved in the embodiments of this application are applicable to the following explanations.
[0055] 1) Computer vision technology is a technology that enables a computer to understand and process visual information. It uses an image sensor to acquire an image signal and, through an image processing system, identifies, classifies, detects, and tracks objects, scenes, or activities in an image or video, thereby extracting, interpreting, and analyzing useful information, becoming an extension of the human visual system in the digital world.
[0056] 2) Deep learning technology is a machine learning technology. It automatically extracts features from raw data and learns by constructing a network structure of multiple layers of non-linear processing units (neurons), without the need for manual feature design. It can process complex, high-dimensional data and capture non-linear relationships in the data, and is widely used in fields such as computer vision, natural language processing, and speech recognition.
[0057] 3) Binocular data generally refers to data obtained through a binocular vision system. This system uses two cameras to simultaneously capture images of the same scene from different angles. Based on the principle of parallax and triangulation, it can calculate the three-dimensional geometric information of an object, such as position, shape, and depth. These data have wide application value in fields such as machine vision, three-dimensional reconstruction, and augmented reality.
[0058] In the related art, there is a method of using a binocular camera to generate binocular data, that is, generating binocular data based on a binocular camera or simulating a real scene through a rendering engine. However, it is found in actual applications that these methods obtain binocular data pairs based on a fixed set of camera parameters, which will to a certain extent result in insufficient data diversity, causing the network to rely on camera parameters and having poor generalization performance when the camera parameters are inconsistent during testing.
[0059] Exemplarily, for example, by mounting a radar (to obtain point clouds) and a binocular camera (to obtain left and right eyes), the relationship between the coordinate axes of the three devices is obtained through complex calibration. In order to obtain the disparity annotation of the frame of interest, the data of the ten consecutive frames before and after are registered and denoised through the iterative closest point algorithm (ICP), and then the point clouds are projected back into the left and right eye images of this frame. After manually removing the ambiguous areas, the disparity annotation is finally obtained according to the corresponding relationship. The defect of this method is that the radar device is relatively expensive, and at the same time, the point clouds scanned by the radar have poor quality in areas such as glass and need to be filtered manually, which brings labor costs and affects the annotation efficiency. The final annotation obtained is also sparse, with only 10%-30% of the area being effective, and finally only a small number of effective binocular data are obtained.
[0060] In addition, a scene can also be quickly built according to one's own ideas in a physics engine, and data pairs are obtained through rendering. Since all physical quantities in the physics engine are known or derivable, dense and accurate disparity annotations can be obtained. Although the rendering engine has developed rapidly, there are still certain differences from the data collected in the real scene. For example, the rendering engine cannot simulate the imaging noise of camera hardware, etc. This results in the deep learning model trained based on this data having poor performance in the real scene and being unable to be actually implemented.
[0061] In view of this, in the embodiments of the present application, a binocular data generation method, system, electronic device and storage medium are provided. This solution obtains a monocular data set, where the monocular data set includes multiple monocular data, and then performs depth estimation processing on the monocular data to obtain an estimated depth. By performing disparity conversion processing on the estimated depth, an initial disparity is obtained, and different disparity results can be generated, enriching the diversity of binocular data; this solution also performs disparity optimization processing on the initial disparity to obtain a target disparity, performs projection processing on the monocular data according to the target disparity to obtain a projection image, and finally generates binocular data according to the target disparity, monocular data and projection image, which can optimize the initial disparity, improve the quality of the generated disparity annotation and the efficiency of binocular data generation.
[0062] The binocular data generation method provided by the embodiments of this application relates to the field of computer technology. The binocular data generation method provided by the embodiments of this application can be applied to a terminal, or to a server, or can also be software running on a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, etc., but is not limited thereto; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application that implements the binocular data generation method, etc., but is not limited to the above forms.
[0063] This application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0064] Figure 1 is an optional flowchart of the binocular data generation method provided by the embodiments of this application, Figure 1 The method in may include but is not limited to steps S101 to S106.
[0065] Step S101, obtain a monocular data set, where the monocular data set includes a plurality of monocular data;
[0066] Step S102, perform depth estimation processing on the monocular data to obtain an estimated depth;
[0067] Step S103, perform parallax conversion processing on the estimated depth to obtain an initial parallax;
[0068] Step S104, perform parallax optimization processing on the initial parallax to obtain a target parallax;
[0069] Step S105: Project the monocular data according to the target disparity to obtain a projected image;
[0070] Step S106: Generate binocular data based on the target disparity, the monocular data, and the projected image.
[0071] Steps S101 to S106 illustrated in the embodiments of the present application can generate a corresponding binocular data set by obtaining a monocular data set, and can efficiently and low - cost generate binocular data for network training from the monocular data set, providing a data basis for the binocular depth estimation large model. In the embodiments of the present application, the estimated depth is obtained by performing depth estimation on the monocular data, and different disparity results can be obtained through disparity conversion of the estimated depth, further expanding the scope of the data set. Then, disparity optimization is performed on the initial disparity. By performing disparity optimization such as edge sharpening and denoising on the initial disparity, the final target disparity is obtained, making the quality of the generated disparity annotation higher. Then, pixel projection is performed on the monocular data according to the target disparity. The monocular data can be used as the left image in the binocular data, and the color value of each pixel in the left image is mapped to the corresponding position in the right image to obtain a projected image. Finally, the monocular data and the projected image are used as the image pair of the binocular data, and the target disparity is used as the disparity annotation to obtain the final binocular data, which can be directly used to train the binocular disparity estimation network.
[0072] In step S101 of some embodiments, a monocular data set is obtained, and the monocular data set includes multiple monocular data.
[0073] In the embodiments of the present application, monocular data sets of various scenarios can be collected, including indoor scenes, outdoor open scenes, autonomous driving scenes, industrial scenes, etc., which can improve the richness of the generated binocular data set. The monocular data set can be obtained by accessing a database, or by other means, such as collecting the monocular data set from the network through web crawling technology, which is not limited thereto. Among them, the monocular data set includes multiple monocular data, and the monocular data is the obtained image data.
[0074] In step S102 of some embodiments, depth estimation processing is performed on the monocular data to obtain an estimated depth.
[0075] In the embodiments of the present application, a monocular depth estimation large model can be used to perform depth estimation processing on each monocular data, that is, image data, in the monocular data set to obtain an estimated depth. In a feasible embodiment, the model structure and pre - trained weights of DepthAnythingV2 - Large are used as the monocular depth estimation large model. By inputting the monocular data into the monocular depth estimation large model for estimation, the estimated depth is obtained through the output of the large model.
[0076] Please refer to Figure 2 , in step S103 of some embodiments, the processing of performing parallax conversion on the estimated depth to obtain an initial parallax includes the following steps:
[0077] Step S201, perform maximum normalization processing on the estimated depth to obtain a normalized parallax;
[0078] Step S202, perform multiplication processing on the normalized parallax according to a scaling factor to obtain the initial parallax, where the scaling factor is obtained by randomly and uniformly sampling a preset range.
[0079] In the embodiments of the present application, since the camera parameters of the monocular data collected on the network, such as the focal length, etc., cannot be known, and there is no concept of a baseline for a monocular camera, the depth is converted into parallax by constructing a virtual right-eye camera and forging camera parameters. The specific method is to perform maximum normalization processing on the estimated depth to obtain a normalized parallax, and then perform multiplication processing on the normalized parallax according to a scaling factor to obtain the initial parallax. Among them, the parallax and the depth satisfy an inverse proportional relationship, and the relationship formula between the parallax and the depth is as follows:
[0080]
[0081] Among them, Disp is the parallax, representing the pixel position difference of the corresponding points of the left and right eyes in the row direction, fx represents the focal length of the camera, baseline represents the baseline of the binocular camera, that is, the distance between the optical centers of the left and right eyes, and Depth represents the depth.
[0082] In the embodiments of the present application, by performing maximum normalization processing on the estimated depth, the estimated depth is normalized by dividing it by the absolute value of the maximum value in the depth dataset, so as to limit the data range within a fixed interval, and then taking the reciprocal to obtain the normalized parallax, and then multiplying the normalized parallax by the scaling factor to obtain the initial parallax. Usually, this fixed interval is [0, 1], and the depth dataset is a dataset obtained by performing depth estimation on each monocular data in the monocular dataset. The scaling factor is randomly and uniformly sampled from a preset range, and this preset range can be set according to the actual situation, aiming to ensure that the initial parallax will finally fall within a reasonable range, and the scaling factor will be randomly sampled again for each sample. Therefore, the generation formula of the initial parallax is as follows:
[0083]
[0084] In the formula, Disp init represents the initial parallax, Depth maxrepresents the absolute value of the maximum value in the disparity dataset, scale represents the scaling factor, and Depth represents the estimated depth. Comparing the relationship formula between disparity and depth with the generation formula of the initial disparity, it can be seen that Depth max ·scale is equivalent to fx·baseline. When different scaling factors are selected for the same sample, it can be understood that disparity results under different combinations of camera parameters are generated.
[0085] One of the technical solutions in the above technical solutions has the following advantages or beneficial effects: In the embodiment of the present application, the initial disparity is obtained by performing disparity conversion processing on the estimated depth, which can generate different disparity results, further expanding the range of camera parameters and enriching the generated binocular dataset.
[0086] Please refer to Figure 3 , in step S104 of some embodiments, the disparity optimization process for the initial disparity to obtain the target disparity includes the following steps:
[0087] Step S301, calculating the spatial Gaussian weight and the intensity Gaussian weight according to the initial disparity;
[0088] Step S302, performing weighted processing on the initial disparity according to the spatial Gaussian weight and the intensity Gaussian weight to obtain the target disparity.
[0089] In the embodiment of the present application, by performing disparity optimization such as edge sharpening and denoising on the initial disparity to obtain the final target disparity, the spatial Gaussian weight and the intensity Gaussian weight can be calculated according to the initial disparity, and the target disparity can be obtained by performing weighted processing on the initial disparity. Compared with the method of directly filling the disparity at the edge with the nearest neighbor disparity in the embodiment of the present application, the embodiment of the present application additionally considers the similarity of the disparity, and is smoother and more robust in a weighted manner, which can make the generated annotation quality higher.
[0090] In step S301 of some embodiments, the calculating the spatial Gaussian weight and the intensity Gaussian weight according to the initial disparity includes the following steps:
[0091] Determining the current pixel and the neighborhood pixels according to the initial disparity;
[0092] Performing distance calculation processing on the neighborhood pixels according to the current pixel to obtain the spatial distance;
[0093] Determining the neighborhood disparity according to the neighborhood pixels, and performing difference calculation processing on the neighborhood disparity according to the initial disparity to obtain the disparity difference;
[0094] Normalize the spatial distance and the parallax difference respectively according to the Gaussian function to obtain the spatial Gaussian weight and the intensity Gaussian weight.
[0095] In the embodiment of the present application, for the parallax of each pixel position, the current pixel and neighboring pixels can be obtained. The neighboring pixels are other pixels within the filtering window. By calculating the spatial distance and the parallax difference between this pixel and other pixels within the neighborhood, the spatial Gaussian weight and the intensity Gaussian weight are respectively obtained through normalization using the Gaussian function, and finally weighted to obtain the target parallax. The generation formula of the target parallax is as follows:
[0096] Disp gt (p)=∑ q∈Ω G s (||p - q||)·G d (|Disp(p)-Disp(q)|)·Disp(q);
[0097] In the formula, Disp gt represents the target parallax, p represents the current pixel position, p represents the pixel position in the filtering window, Ω represents the filtering window, G s (||p - q||) is the spatial Gaussian weight function, G d (|Disp(p)-Disp(q)|) is the intensity Gaussian weight function, and Disp represents the initial parallax.
[0098] Please refer to Figure 4 , in step S105 of some embodiments, the projecting the monocular data according to the target parallax to obtain a projected image includes the following steps:
[0099] Step S401, determining the pixel distance difference according to the target parallax;
[0100] Step S402, performing color value mapping processing on the pixels of the monocular data according to the pixel distance difference to obtain an initial image;
[0101] Step S403, performing perspective difference optimization processing on the initial image to obtain the projected image.
[0102] In the embodiment of the present application, the monocular data can be used as the left-eye image in the binocular data, and the left-eye image is projected according to the parallax of the pixels to obtain an initial right-eye image. According to the definition of parallax, assuming that the left and right-eye images have been rectified epipolar, that is, the left and right-eye imaging planes are in the same plane, and the rows are aligned, and the two optical axes are strictly aligned. When there is a point p in space, its pixel positions on the left and right-eye images are p left =(y l ,x l ), pright =(y r , x r ), y represents the column coordinate, x represents the row coordinate, and the parallax Disp represents the pixel distance difference of the corresponding points (or matching points) along the row direction in the left and right eye images, that is, Disp = x l - x r . Therefore, according to this correspondence, the color value of each pixel in the left image can be mapped to the corresponding position in the right image in the embodiments of the present application. Specifically, for each pixel p of the left image I left , its coordinate is (y l , x l ), the color value is I left (p), and the corresponding parallax is Disp(p). It can be obtained that I right (q) = I left (p), where the coordinate of q is (y l , x l - Disp(p)). Therefore, the color value of the pixels of the monocular data is mapped through the pixel distance difference to obtain the initial image.
[0103] In step S403 of some embodiments, the step of performing perspective difference optimization processing on the initial image to obtain the projected image includes the following steps:
[0104] Extract multiple mapping values of the same pixel point from the initial image;
[0105] Perform parallax comparison processing on the multiple mapping values, select and retain the mapping value with the largest parallax to obtain the projected image.
[0106] In the embodiments of the present application, due to the perspective difference between the left and right eye cameras, there are two special cases in the mapping process. One is that different pixel points in the left eye image are mapped to the same pixel point in the right eye image, and the other is that some pixel points in the right eye image are not visible in the left eye image, so there is no corresponding mapping value. Therefore, in the embodiments of the present application, multiple mapping values of the same pixel point are extracted from the initial image, and one of them is selected and retained. According to the imaging principle, along the optical path direction, the object that is close to the camera and unobstructed will finally be imaged on the image plane. Therefore, for the points where multiple left image pixels are mapped to the same position in the right image, only the points that are close to the camera, that is, the points with smaller depth, are retained. According to the relationship formula between depth and parallax, the smaller the depth, the larger the parallax. Finally, for this case, the mapping value with a larger parallax is retained to obtain the projected image.
[0107] In step S403 of some embodiments, the step of performing perspective difference optimization processing on the initial image to obtain the projected image further includes the following steps:
[0108] Perform region extraction processing on the initial image to obtain a non-filled region;
[0109] Perform mask calculation processing on the non-filled region to obtain a region mask;
[0110] Input the initial image and the region mask into an image completion large model for prediction and generation processing to obtain the projection image.
[0111] In the embodiments of the present application, an image completion large model can be used to complete the non-filled region in the initial image. The non-filled region is obtained by performing region extraction processing on the initial image. Specifically, the image completion large model Lama uses neighborhood information and powerful prior knowledge of large-scale training data for completion. Therefore, the embodiments of the present application directly adopt the official provided model and pre-trained weights. First, according to whether the right image pixels have no mapped values, by extracting the mask Mask of the region to be filled, the initial image is used as the right image I right and input into the image completion large model together. The image completion large model predicts the corresponding color values for the non-filled region to complete the image and obtains the projection image.
[0112] Next, in combination with specific application examples, the solutions of the embodiments of the present application will be introduced and described in detail:
[0113] The embodiments of the present application can be applied to the field of computer technology and can be widely applied to computer vision scenarios, deep learning scenarios, etc. It can generate binocular data based on a large model and real monocular data, mainly used to generate an annotated near-real binocular data set to alleviate the problem of inter-domain differences in synthetic data and the lack of real data, and provide a data basis for the binocular depth estimation large model. The training of the embodiments of the present application does not require collecting real depth data as training annotations, but uses a large model to efficiently convert monocular data into high-quality annotated binocular data pairs, providing a data basis for binocular depth estimation algorithms based on deep learning, greatly reducing the data collection difficulty and cost. Compared with the model trained based on synthetic data, the data generated by using the binocular data generation method provided by the embodiments of the present application performs better in the real scene. In the following table, the embodiments of the present application train the binocular depth estimation large model on the data generated by the SceneFlow method and the present method for the same duration and then test the generalization effect on the training set of KITTI2015 as shown in Table 1 below. Table 1 is a comparison table of effects.
[0114]
[0115] Table 1 Comparison table of effects
[0116] As can be seen from the data in the table, when the network structure and the training duration are the same, the metrics of the model trained with the data based on the present method have been significantly improved on real data. Among them, the SceneFlow method is a technique for understanding and describing three-dimensional dynamic scenes, the EPE evaluation metric is used to measure the difference between the disparity value predicted by the algorithm and the true disparity value, and the D1 evaluation metric refers to the minimum distance between different pixel points on the surface of the same object in images from two different viewpoints. This metric directly reflects the accuracy and effect of stereo matching.
[0117] Please refer to Figure 5 , the embodiment of the present application further provides a binocular data generation system, which can implement the above-mentioned binocular data generation method. The system includes:
[0118] The first module 501 is used to obtain a monocular data set, and the monocular data set includes a plurality of monocular data;
[0119] The second module 502 is used to perform depth estimation processing on the monocular data to obtain an estimated depth;
[0120] The third module 503 is used to perform disparity conversion processing on the estimated depth to obtain an initial disparity;
[0121] The fourth module 504 is used to perform disparity optimization processing on the initial disparity to obtain a target disparity;
[0122] The fifth module 505 is used to perform projection processing on the monocular data according to the target disparity to obtain a projection image;
[0123] The sixth module 506 is used to generate binocular data according to the target disparity, the monocular data, and the projection image.
[0124] It can be understood that the content in the above method embodiments is applicable to the system embodiments of the present application. The functions specifically implemented by the system embodiments of the present application are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0125] The embodiment of the present application further provides an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above-mentioned binocular data generation method is implemented. The electronic device can be any intelligent terminal including a tablet computer, a vehicle-mounted computer, etc.
[0126] It can be understood that the content in the above method embodiments is applicable to the device embodiments of the present application. The functions specifically implemented by the device embodiments of the present application are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0127] Please refer to Figure 6 , Figure 6 which schematically shows the hardware structure of an electronic device according to another embodiment. The electronic device includes:
[0128] A processor 601, which can be implemented in the form of a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;
[0129] A memory 602, which can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 602 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 602 and are called by the processor 601 to execute the binocular data generation method of the embodiments of the present application;
[0130] An input / output interface 603, which is used to implement information input and output;
[0131] A communication interface 604, which is used to implement communication interaction between this device and other devices, and can implement communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as mobile network, WIFI, Bluetooth, etc.);
[0132] A bus 605, which transmits information between various components of the device (such as the processor 601, the memory 602, the input / output interface 603, and the communication interface 604);
[0133] Among them, the processor 601, the memory 602, the input / output interface 603, and the communication interface 604 are communicatively connected to each other inside the device through the bus 605.
[0134] The embodiments of the present application also provide a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned binocular data generation method is implemented.
[0135] It can be understood that the content in the above method embodiments is applicable to the embodiments of this storage medium. The functions specifically implemented by the embodiments of this storage medium are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0136] A memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0137] A binocular data generation method, system, electronic device, and storage medium provided by an embodiment of the present application. This solution obtains a monocular data set, which includes a plurality of monocular data, and then performs depth estimation processing on the monocular data to obtain an estimated depth. By performing disparity conversion processing on the estimated depth, an initial disparity can be obtained, and different disparity results can be generated, enriching the diversity of binocular data. This solution also performs disparity optimization processing on the initial disparity to obtain a target disparity, performs projection processing on the monocular data according to the target disparity to obtain a projection image, and finally generates binocular data according to the target disparity, monocular data, and projection image, which can optimize the initial disparity, improve the quality of the generated disparity annotation, and the efficiency of binocular data generation.
[0138] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0139] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown, or combine certain steps, or different steps.
[0140] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0141] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices can be implemented as software, firmware, hardware, and appropriate combinations thereof.
[0142] In the description of this application and the above-mentioned drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0143] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression refers to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0144] In several embodiments provided by this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are only illustrative. For example, the above-mentioned division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of systems or units can be in electrical, mechanical or other forms.
[0145] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0146] In addition, in each embodiment of the present application, the functional units may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0147] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in the various embodiments of the present application. The foregoing storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.
[0148] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings. However, this does not limit the scope of the rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall fall within the scope of the rights of the embodiments of the present application.
Claims
1. A binocular data generation method, characterized in that: The method comprises the following steps: Acquire a monocular data set, where the monocular data set includes a plurality of monocular data; Performing depth estimation processing on the monocular data to obtain an estimated depth; Performing a disparity conversion process on the estimated depth to obtain an initial disparity; Performing disparity optimization processing on the initial disparity to obtain a target disparity; Performing projection processing on the monocular data according to the target parallax to obtain a projected image; Binocular data is generated according to the target disparity, the monocular data and the projected image.
2. The method according to claim 1, characterized in that The performing disparity conversion processing on the estimated depth to obtain an initial disparity comprises the following steps: Performing maximum value normalization processing on the estimated depth to obtain normalized disparity; The normalized disparity is multiplied according to a scaling factor to obtain the initial disparity, and the scaling factor is obtained by performing random uniform sampling on a preset range.
3. The method according to claim 1, characterized in that The performing disparity optimization processing on the initial disparity to obtain the target disparity comprises the following steps: Calculate the spatial Gaussian weight and the intensity Gaussian weight according to the initial disparity; The initial disparity is weighted according to the spatial Gaussian weight and the intensity Gaussian weight to obtain the target disparity.
4. The method according to claim 3, characterized in that The step of calculating the spatial Gaussian weight and the intensity Gaussian weight according to the initial disparity comprises the following steps: Determine a current pixel and neighboring pixels according to the initial disparity; Performing distance calculation processing on neighboring pixels according to the current pixel to obtain a spatial distance; Determine a neighborhood disparity according to the neighborhood pixels, and perform a disparity calculation process on the neighborhood disparity according to the initial disparity to obtain a disparity disparity; The spatial distance and the parallax difference are respectively normalized according to Gaussian functions to obtain the spatial Gaussian weight and the intensity Gaussian weight.
5. The method according to claim 1, characterized in that The process of projecting the monocular data according to the target parallax to obtain a projected image comprises the following steps: determining a pixel distance difference according to the target disparity; Performing color value mapping processing on the pixels of the monocular data according to the pixel distance difference to obtain an initial image; The initial image is subjected to viewing angle difference optimization processing to obtain the projection image.
6. The method according to claim 5, characterized in that The step of performing viewing angle difference optimization processing on the initial image to obtain the projection image comprises the following steps: Extracting multiple mapping values of the same pixel from the initial image; The plurality of mapping values are subjected to parallax comparison processing, and a mapping value having a maximum parallax is selected for retention to obtain the projected image.
7. The method according to claim 5, characterized in that The performing of viewing angle difference optimization processing on the initial image to obtain the projection image further comprises the following steps: Performing region extraction processing on the initial image to obtain an unfilled region; Performing mask calculation processing on the unfilled area to obtain a regional mask; The initial image and the regional mask are input into the image completion model for prediction and generation processing to obtain the projected image.
8. A binocular data generation system, characterized in that: The system comprises: The first module is used to obtain a monocular data set, where the monocular data set includes a plurality of monocular data; The second module is used to perform depth estimation processing on the monocular data to obtain an estimated depth; A third module is used to perform a disparity conversion process on the estimated depth to obtain an initial disparity; A fourth module is used to perform disparity optimization processing on the initial disparity to obtain a target disparity; A fifth module is used to perform projection processing on the monocular data according to the target parallax to obtain a projection image; The sixth module is used to generate binocular data according to the target disparity, the monocular data and the projection image.
9. An electronic device, characterized in that: The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Depth information generation method, head-mounted display device, storage medium and product
CN120614444A