Parallax Map Processing, Deep Learning Model Training Method and Related Devices

By inputting the initial disparity map into the trained deep learning model for processing, the problem of long and poor results in the prior art disparity map calculation is solved, and a high-quality disparity map is obtained in a short time.

CN114255268BActive Publication Date: 2025-05-27WUHAN TCL CORP RES CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011018154.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-24
Publication Date
2025-05-27
Estimated Expiration
2040-09-24

AI Technical Summary

Technical Problem

In the prior art, it is difficult to obtain high-quality parallax maps due to repeated textures, weak textures, overexposures, noises and other reasons when calculating parallax maps, and commonly used filters take a long time and have poor results.

Method used

By obtaining the initial disparity map of the image pair to be processed and inputting it into the trained deep learning model for processing, the output image quality is higher than the second disparity map with the initial disparity map. The deep learning model trains the untrained model by using the initial and final disparity maps of multiple sample image pairs as training samples.

Benefits of technology

Obtaining a parallax map with higher image quality in a short time solves the problem that the parallax map calculation takes a long time and is poor in the prior art.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114255268B_ABST
    Figure CN114255268B_ABST
Patent Text Reader

Abstract

The present application discloses a method for processing a disparity map, training a deep learning model, and related devices, belonging to the technical field of image processing. The method includes: obtaining a first disparity map, which is determined according to an image pair to be processed; inputting the first disparity map into a trained deep learning model for processing to output a second disparity map, and the image quality of the second disparity map is higher than that of the first disparity map. In the present application, a disparity map with relatively high image quality can be obtained through a trained deep learning model in a relatively short time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and particularly to a method for processing a disparity map, training a deep learning model, and related devices. Background Art

[0002] The application of dual-camera configurations in terminals such as mobile phones is becoming increasingly widespread. Currently, in order to improve the shooting effect, dual-camera bokeh can be used to shallow the depth of field and focus on the subject. To achieve dual-camera bokeh, it is necessary to calculate the disparity map of the main camera image and the secondary camera image, perform depth estimation through this disparity map, and accordingly perform background bokeh.

[0003] However, due to reasons such as repetitive textures, weak textures, overexposure, noise, etc., the image quality of the disparity map directly calculated by the stereo matching algorithm is difficult to meet the requirements, and filters often need to be used for optimization. For example, filters such as least squares filters, joint bilateral filters, median filters, etc. can be used for optimization. These filters have problems such as long processing time and poor filtering effect. Summary of the Invention

[0004] This application provides a method for processing a disparity map, training a deep learning model, and related devices, which is used to obtain a disparity map with higher image quality in a shorter time.

[0005] In a first aspect, a method for processing a disparity map is provided, including:

[0006] Obtain a first disparity map, where the first disparity map is determined according to an image pair to be processed;

[0007] Input the first disparity map into a trained deep learning model for processing, and output a second disparity map, where the image quality of the second disparity map is higher than that of the first disparity map.

[0008] In this application, after obtaining the first disparity map of the image pair to be processed and inputting the first disparity map into a trained deep learning model, a second disparity map with higher image quality can be obtained in a shorter time.

[0009] In a second aspect, a method for training a deep learning model is provided, including:

[0010] Obtain a plurality of sample image pairs, where the image quality of the final disparity map in each sample image pair is higher than that of the initial disparity map;

[0011] Use the plurality of initial disparity maps as input data in the training sample, and use the plurality of final disparity maps as sample labels in the training sample;

[0012] Use the training sample to train an untrained deep learning model to obtain a trained deep learning model.

[0013] In this application, after obtaining a plurality of sample image pairs, the plurality of sample image pairs are determined as training samples, and the image quality of the final disparity map in each sample image pair is higher than that of the initial disparity map. Then, the un-trained deep learning model is trained using the training samples to obtain a trained deep learning model. The trained deep learning model can obtain a disparity map with high image quality in a short time.

[0014] In a third aspect, a disparity map processing device is provided, including:

[0015] A first disparity map obtaining module, configured to obtain a first disparity map, where the first disparity map is determined according to an image pair to be processed;

[0016] A disparity map processing module, configured to input the first disparity map into the trained deep learning model for processing, and output a second disparity map, where the image quality of the second disparity map is higher than that of the first disparity map.

[0017] In a fourth aspect, a deep learning model training device is provided, including:

[0018] A second disparity map obtaining module, configured to obtain a plurality of sample image pairs, where the image quality of the final disparity map in each sample image pair is higher than that of the initial disparity map;

[0019] A training sample obtaining module, configured to use a plurality of initial disparity maps as input data in the training samples, and use a plurality of final disparity maps as sample labels in the training samples;

[0020] A training module, configured to train the un-trained deep learning model using the training samples to obtain a trained deep learning model.

[0021] In a fifth aspect, a computer device is provided, where the computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the computer program is executed by the processor, the above-mentioned disparity map processing method is implemented.

[0022] In a sixth aspect, a computer device is provided, where the computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the computer program is executed by the processor, the above-mentioned deep learning model training method is implemented.

[0023] In a seventh aspect, a computer-readable storage medium is provided, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned disparity map processing method is implemented.

[0024] In an eighth aspect, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the above-mentioned deep learning model training method is implemented.

[0025] In a ninth aspect, a computer program product containing instructions is provided. When it runs on a computer, the computer is caused to execute the steps of the above-mentioned parallax map processing method.

[0026] In a tenth aspect, a computer program product containing instructions is provided. When it runs on a computer, the computer is caused to execute the steps of the above-mentioned deep learning model training method.

[0027] It can be understood that for the beneficial effects of the above-mentioned third aspect, fifth aspect, seventh aspect, and ninth aspect, reference can be made to the relevant descriptions in the first aspect above, which will not be elaborated here. For the beneficial effects of the above-mentioned fourth aspect, sixth aspect, eighth aspect, and tenth aspect, reference can be made to the relevant descriptions in the second aspect above, which will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0029] Figure 1 is a flowchart of a deep learning model training method provided by an embodiment of the present application;

[0030] Figure 2 is a schematic structural diagram of a deep learning model provided by an embodiment of the present application;

[0031] Figure 3 is a flowchart of a parallax map processing method provided by an embodiment of the present application;

[0032] Figure 4 is a schematic diagram of a parallax map provided by an embodiment of the present application;

[0033] Figure 5 is a schematic structural diagram of a deep learning model training device provided by an embodiment of the present application;

[0034] Figure 6 is a schematic structural diagram of a parallax map processing device provided by an embodiment of the present application;

[0035] Figure 7 is a schematic structural diagram of a computer device provided by an embodiment of the present application;

[0036] Figure 8 It is a schematic structural diagram of another computer device provided by an embodiment of the present application. Detailed implementation manners

[0037] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0038] It should be understood that the "multiple" mentioned in the present application refers to two or more. In the description of the present application, unless otherwise specified, " / " means "or". For example, A / B may represent A or B. The "and / or" herein is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in order to clearly describe the technical solutions of the present application, terms such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and effects. Those skilled in the art can understand that the terms "first" and "second" do not limit the quantity and execution order, and the terms "first" and "second" do not necessarily mean different.

[0039] Before explaining the embodiments of the present application in detail, the application scenarios of the embodiments of the present application will be described first.

[0040] To improve the shooting effect, dual-camera defocusing can be used to shallow the depth of field and focus on the subject. Dual-camera defocusing needs to be realized based on the depth-of-field images of the main and secondary camera images. Currently, in the process of depth estimation of the main and secondary camera images, the disparity map of the main and secondary camera images can be calculated through a stereo matching algorithm. According to the optical principle, the disparity map is inversely proportional to the actual distance of the object. Therefore, the corresponding depth-of-field image can be obtained according to the disparity map. However, due to reasons such as repetitive textures, weak textures, overexposure, and noise, the image quality of the disparity map directly calculated through the stereo matching algorithm is difficult to meet the requirements, and filters often need to be used for optimization. For example, filters such as least squares filters, joint bilateral filters, and median filters can be used for optimization. These filters have problems such as long time consumption and poor filtering effects.

[0041] For this reason, the embodiments of the present application provide a method for processing a disparity map, which can obtain a high-quality disparity map of an image pair through a trained deep learning model. The trained deep learning model is obtained by determining the initial disparity map of a sample image pair and the final disparity map with an image quality higher than that of the initial disparity map as training samples, and then using the training samples to train an untrained deep learning model. Therefore, a disparity map with relatively high image quality can be obtained through the trained deep learning model in a relatively short time.

[0042] Next, the training process of the deep learning model will be described in detail.

[0043] Figure 1 It is a flowchart of a deep learning model training method provided by an embodiment of the present application. Refer to Figure 1 , the method includes the following steps:

[0044] Step 101: The server obtains multiple sample image pairs, and the image quality of the final disparity map in each sample image pair is higher than that of the initial disparity map.

[0045] The two sample images in each sample image pair among the multiple sample image pairs are the initial disparity map and the final disparity map. The multiple sample image pairs correspond one-to-one with multiple first image pairs. The initial disparity map and the final disparity map in each sample image pair are the disparity maps of the corresponding first image pair.

[0046] The two images in the first image pair are two images captured at different angles in the same scene, that is, the first image pair is obtained by shooting the same target with two different cameras. The first image pair can be an image captured by a dual camera in an electronic device, that is, it can be a binocular image pair. For example, the first image pair can be an image captured by the camera with the highest pixel among multiple cameras of a mobile phone and an image captured by a camera other than the camera with the highest pixel among multiple cameras of the mobile phone.

[0047] The disparity map of the first image pair is an image with the size of a reference image (i.e., the main image) in the first image pair and pixel values as disparity values. The image quality of the disparity map can be measured from one or more aspects such as edge fineness, accuracy, noise, and sharpness.

[0048] The fact that the image quality of the final disparity map is higher than that of the initial disparity map means that the final disparity map is superior to the initial disparity map in one or more aspects such as edge fineness, accuracy, noise, and sharpness. Compared with the initial disparity map, the final disparity map can better meet the implementation requirements of dual-camera bokeh.

[0049] For any one of the multiple sample image pairs, this sample image pair can be called sample image pair A. The operations of the server to obtain the initial disparity map and the final disparity map in sample image pair A are as follows:

[0050] Among them, the initial disparity map in sample image pair A can be directly calculated by a relatively simple method. Optionally, the server can obtain the first image pair and determine the initial disparity map in sample image pair A according to the first image pair and a pre-stored stereo matching algorithm.

[0051] The stereo matching algorithm can find the corresponding pixel points of each pixel point in one image in the image from another perspective, calculate the disparity map of these two images, and estimate the depth-of-field image based on this. For example, the stereo matching algorithm can be the SGBM (semi-global block matching) algorithm, the BM (block matching) algorithm, the GC (graph cuts) algorithm, etc., and the embodiments of this application do not make a unique limitation on this.

[0052] In a possible situation, the operation of the server to determine the initial disparity map in the sample image pair A according to the first image pair and the pre-stored stereo matching algorithm can be: The server directly obtains the disparity map of the first image pair through the pre-stored stereo matching algorithm as the initial disparity map in the sample image pair A.

[0053] In another possible situation, the operation of the server to determine the initial disparity map in the sample image pair A according to the first image pair and the pre-stored stereo matching algorithm can be: The server determines the initial disparity map in the sample image pair A according to the first image pair and the multi-scale stereo matching algorithm.

[0054] The multi-scale stereo matching algorithm first performs multi-scale decomposition on the image, then combines the information of each scale, and obtains the disparity map through the stereo matching algorithm. The disparity map obtained through this multi-scale stereo matching algorithm can integrate the advantages of the fineness of large sizes and the accuracy of small sizes.

[0055] Among them, the operation of the server to determine the initial disparity map in the sample image pair A according to the first image pair and the multi-scale stereo matching algorithm can be:

[0056] The server determines the third disparity map according to the first image pair and the pre-stored first stereo matching algorithm;

[0057] The server determines the fourth disparity map according to the second image pair and the pre-stored second stereo matching algorithm. The second image pair is obtained by reducing the size of each image in the first image pair by k times, where k is an integer greater than or equal to 1;

[0058] The server enlarges the size of the fourth disparity map by k times to obtain the fifth disparity map;

[0059] The server determines the initial disparity map in the sample image pair A according to the pixel values of the pixel points in the third disparity map and the fifth disparity map.

[0060] Among them, when the server determines the third disparity map according to the first image pair and the pre-stored first stereo matching algorithm, the server can obtain the disparity map of the first image pair through the first stereo matching algorithm as the third disparity map.

[0061] Among them, when the server determines the fourth disparity map based on the second image pair and the pre-stored second stereo matching algorithm, the server can obtain the disparity map of the second image pair through the second stereo matching algorithm as the fourth disparity map. The first stereo matching algorithm and the second stereo matching algorithm can be the same or different.

[0062] k can be set in advance. For example, k can be 2. In this case, the second image pair is the image pair of the first image pair at 1 / 2 size.

[0063] The size of each image in the first image pair is reduced by k times to obtain the second image pair. The size of the disparity map of the second image pair (i.e., the fourth disparity map) is enlarged by k times to obtain the fifth disparity map. Therefore, the size of the disparity map of the first image pair (i.e., the third disparity map) is the same as the size of the fifth disparity map.

[0064] Since the third disparity map is the disparity map of the first image pair and the fifth disparity map is the disparity map obtained by enlarging the size of the disparity map of the second image pair by k times, the third disparity map and the fifth disparity map are disparity maps at different scales. Therefore, the pixel values of the pixel points in the third disparity map and the pixel values of the pixel points in the fifth disparity map can be combined to determine the initial disparity map in the sample image pair A.

[0065] Optionally, the operation of the server to determine the initial disparity map in the sample image pair A based on the pixel values of the pixel points in the third disparity map and the pixel values of the pixel points in the fifth disparity map can be:

[0066] The server determines the first pixel points at each position in the third disparity map and the second pixel points at the corresponding positions in the fifth disparity map;

[0067] If the difference between the pixel value of the first pixel point and the pixel value of the second pixel point is less than the reference difference, the server uses the pixel value of the first pixel point as the pixel value of the pixel point at this position in the initial disparity map of the sample image pair A; or,

[0068] If the difference between the pixel value of the first pixel point and the pixel value of the second pixel point is greater than the reference difference, the server uses the pixel value of the second pixel point as the pixel value of the pixel point at this position in the initial disparity map of the sample image pair A; or,

[0069] If the difference between the pixel value of the first pixel point and the pixel value of the second pixel point is equal to the reference difference, the server uses the pixel value of the first pixel point or the pixel value of the second pixel point as the pixel value of the pixel point at this position in the initial disparity map of the sample image pair A.

[0070] For any position in the disparity map, the server determines the pixel value of the pixel at this position in the initial disparity map in the sample image pair A based on the pixel value of the first pixel at this position in the third disparity map and the pixel value of the second pixel at this position in the fifth disparity map.

[0071] The reference difference can be set in advance, and the reference difference can be set to a relatively large value. For example, the reference difference can be 100, 120, etc. The embodiments of the present application do not make a unique limitation on this.

[0072] If the difference between the pixel value of the first pixel and the pixel value of the second pixel is less than the reference difference, it indicates that the difference between the pixel value of the first pixel and the pixel value of the second pixel is relatively small. In this case, the pixel value of the first pixel in the third disparity map corresponding to the original size image pair can be taken as the pixel value of the pixel in the initial disparity map in the sample image pair A, so that the initial disparity map in the sample image pair A can obtain high - resolution of the large size.

[0073] If the difference between the pixel value of the first pixel and the pixel value of the second pixel is greater than the reference difference, it indicates that the difference between the pixel value of the first pixel and the pixel value of the second pixel is relatively large. In this case, the pixel value of the second pixel in the fifth disparity map corresponding to the reduced image pair of the original size image pair can be taken as the pixel value of the pixel in the initial disparity map in the sample image pair A, so that the initial disparity map in the sample image pair A can obtain high - accuracy of the small size.

[0074] If the difference between the pixel value of the first pixel and the pixel value of the second pixel is equal to the reference difference, it indicates that the difference between the pixel value of the first pixel and the pixel value of the second pixel is not large. In this case, whether to take the pixel value of the first pixel in the third disparity map corresponding to the original size image pair or the pixel value of the second pixel in the fifth disparity map corresponding to the reduced image pair of the original size image pair, the effects are not much different. Therefore, either one of them can be taken as the pixel value of the pixel in the initial disparity map in the sample image pair A.

[0075] Among them, the final disparity map in the sample image pair A can be obtained in a relatively complex way. Two possible ways are described below. Through the following two possible ways, a final disparity map with better image quality can be obtained.

[0076] In the first possible way, the server processes the initial disparity map in the sample image pair A through an edge - preserving filtering algorithm to obtain the final disparity map in the sample image pair A.

[0077] Through this edge-preserving filtering algorithm, a disparity map that is relatively smooth as a whole and has relatively fine edges can be obtained. For example, this edge-preserving filtering algorithm may include one or more of the WLS (weighted least squares) filtering algorithm, the FBS (fast bilateral filter) algorithm, etc., and the embodiments of the present application do not make a unique limitation thereto.

[0078] For example, if this edge-preserving filtering algorithm includes the WLS filtering algorithm and the FBS algorithm, the server may perform WLS filtering on the initial disparity map in the sample image pair A to obtain a first disparity map, and the overall of the first disparity map is relatively smooth but the edges remain unchanged. Then, the server performs FBS processing on the first disparity map according to the main image in the first image pair to obtain a second disparity map as the final disparity map in the sample image pair A, and the edges of the final disparity map are relatively fine. Among them, the main image in the first image pair is the image captured by the camera with the highest pixel among multiple cameras.

[0079] In the second possible manner, the server processes the initial disparity map in the sample image pair A through an edge-preserving filtering algorithm to obtain a second disparity map, and then performs refinement processing on the second disparity map according to the received image processing instruction to obtain the final disparity map in the sample image pair A.

[0080] Through this edge-preserving filtering algorithm, a disparity map that is relatively smooth as a whole and has relatively fine edges can be obtained. For example, this edge-preserving filtering algorithm may include one or more of the WLS (weighted least squares) filtering algorithm, the FBS (fast bilateral filter) algorithm, etc., and the embodiments of the present application do not make a unique limitation thereto.

[0081] For example, if this edge-preserving filtering algorithm includes the WLS filtering algorithm and the FBS algorithm, the server may perform WLS filtering on the initial disparity map in the sample image pair A to obtain a first disparity map, and the overall of the first disparity map is relatively smooth but the edges remain unchanged. Then, the server performs FBS processing on the first disparity map according to the main image in the first image pair to obtain a second disparity map. Among them, the main image in the first image pair is the image captured by the camera with the highest pixel among multiple cameras.

[0082] This image processing instruction is used to indicate processing of the second disparity map, and may indicate performing image smoothing processing or image sharpening processing. For example, this image processing instruction may indicate performing smoothing processing on the image main body and the background and performing sharpening processing on the image edges. This image processing instruction may be triggered by a technician, and the technician may trigger it through operations such as click operation, slide operation, voice operation, and somatosensory operation, and the embodiments of the present application do not make a unique limitation thereto.

[0083] Optionally, a technician may trigger the image processing instruction in image processing software (such as Photoshop) to process the second parallax map, so as to obtain a parallax map with better image quality (such as higher edge fineness) as the final parallax map.

[0084] Step 102: The server uses multiple initial parallax maps as input data in the training sample, and uses multiple final parallax maps as sample labels in the training sample.

[0085] The training sample is a sample for model training, and the training sample includes input data and sample labels. The training sample includes the multiple sample image pairs. The initial parallax map in each sample image pair is the input data, and the final parallax map in each sample image pair is the sample label. That is, the input data in the training sample is all the initial parallax maps in the multiple sample image pairs, and the sample labels in the training sample are all the final parallax maps in the multiple sample image pairs.

[0086] Step 103: The server uses the training sample to train an untrained deep learning model to obtain a trained deep learning model.

[0087] The deep learning model in the embodiments of the present application may include multiple network layers, and the multiple network layers include an input layer, multiple hidden layers, and an output layer. The input layer is responsible for receiving input data; the output layer is responsible for outputting processed data; the multiple hidden layers are located between the input layer and the output layer and are responsible for processing data, and the multiple hidden layers are not visible to the outside. For example, the deep learning model may be a convolutional neural network or the like.

[0088] In a possible case, the structure of the deep learning model may be as Figure 2 shown. The deep learning model includes an input layer, multiple hidden layers, and an output layer. The multiple hidden layers sequentially include a convolutional layer, one or more dilated convolutional layers ( Figure 2 illustrated by taking 4 dilated convolutional layers as an example), a transposed convolutional layer, a softmax layer, and an output layer.

[0089] The input layer is used to receive input data. For example, the input layer may receive input data with a size of 640x480 and a channel number of 1. Among them, 640 is the width, 480 is the height, and the channel number of the input data refers to the number of input data.

[0090] The convolutional layer is used to perform a convolution operation on the input data to obtain multiple first feature maps. For example, the convolutional layer can first perform a convolution on the input data (including but not limited to a convolution with a kernel size of 3x3 and a stride of 2), then perform batch normalization, and then process it through an activation function (including but not limited to the leaky_relu activation function) to obtain a first feature map with a size of 320x240 and 32 channels, that is, 32 first feature maps with a size of 320x240.

[0091] One or more dilated convolutional layers are used to perform dilated convolution operations on the multiple first feature maps to obtain multiple second feature maps. For example, in order to expand the receptive field to obtain global semantic information, each of the four dilated convolutional layers first performs dilated convolution on the multiple first feature maps, then performs batch normalization, and then processes it through an activation function (including but not limited to the leaky_relu activation function) to obtain a second feature map with a size of 320x240 and 32 channels, that is, 32 second feature maps with a size of 320x240. In one possible implementation, the first dilated convolutional layer among the four dilated convolutional layers performs a dilated convolution with a kernel size of 3x3, a stride of 1, and a dilation rate of 1, then performs batch normalization, and then processes it through the leaky_relu activation function; the second dilated convolutional layer performs a dilated convolution with a kernel size of 3x3, a stride of 1, and a dilation rate of 2, then performs batch normalization, and then processes it through the leaky_relu activation function; the third dilated convolutional layer performs a dilated convolution with a kernel size of 3x3, a stride of 1, and a dilation rate of 4, then performs batch normalization, and then processes it through the leaky_relu activation function; the fourth dilated convolutional layer performs a dilated convolution with a kernel size of 3x3, a stride of 1, and a dilation rate of 1, then performs batch normalization, and then processes it through the leaky_relu activation function.

[0092] The transposed convolutional layer is used to perform a transposed convolution operation on the multiple second feature maps to obtain n third feature maps. For example, the transposed convolutional layer can perform a transposed convolution with a kernel size of 3x3 and a stride of 2 on the multiple second feature maps, and then process it through an activation function (including but not limited to the leaky_relu activation function) to obtain a third feature map with a size of 640x480 and 16 channels, that is, 16 third feature maps with a size of 640x480.

[0093] The softmax layer is used to calculate the scores of each of the n third feature maps, obtaining n scores, where n is an integer greater than or equal to 4. When the softmax layer calculates the scores of each of the n third feature maps, for any one of the n third feature maps, the softmax function can be used to calculate the score of this third feature map. Thus, after the softmax layer calculates the scores of each of the n third feature maps, n scores are obtained, and these n scores are equivalent to a matrix with 1 row and n columns. For example, for a third feature map with a size of 640x480 and 16 channels, the softmax layer performs a softmax operation on 16 channels to calculate the scores of each channel, obtaining 16 scores, and these 16 scores are equivalent to a matrix with 1 row and 16 columns.

[0094] The output layer is used to reshape the n scores into a filter matrix with m rows and m columns, and multiply the pixel value of each pixel point in the input data by this filter matrix one by one to obtain the processed data. m is an integer, and n is the square value of m. When the output layer reshapes the n scores into a filter matrix with m rows and m columns, the reshape function can be used to reshape the n scores into a filter matrix with m rows and m columns. Then, for each pixel point at each position in the input data, the output layer can multiply the pixel value of this pixel point by this filter matrix to obtain the pixel value of the pixel point at this position in the processed data. In this way, the edge fineness of the processed data of the deep learning model can be improved. For example, after the softmax layer calculates the scores of 16 channels, the output layer performs a reshape operation on the scores of these 16 channels to obtain a filter matrix with 4 rows and 4 columns, and then multiplies the pixel value of each pixel point in the input data by this filter matrix one by one to obtain the processed data.

[0095] Among them, the operation of the server using the training sample to train the untrained deep learning model to obtain the trained deep learning model can be:

[0096] The server inputs the input data in the training sample into the untrained deep learning model for processing and outputs the processed data;

[0097] The server determines the loss value between the processed data and the sample label in the training sample through a pre-stored edge loss function;

[0098] The server adjusts the parameters in the untrained deep learning model according to the loss value to obtain the trained deep learning model.

[0099] The edge loss function is: Loss = L 2 (I g ,I o )+L2 (I g ',I o ')

[0100] Let Loss be the loss value, I o be the processed data, I o ' be the data obtained by subtracting the first data from the processed data, where the first data is the data obtained by performing Gaussian smoothing on the processed data, I g be the sample label, I g ' be the data obtained by subtracting the second data from the sample label, where the second data is the data obtained by performing Gaussian smoothing on the sample label, L 2 () is the L2 norm loss function.

[0101] In this edge loss function, I o ' is the high-frequency information (i.e., edge) of the processed data, I g ' is the high-frequency information of the sample label. Thus, by calculating the loss value of the L2 norm loss function for the processed data and the sample label, and calculating the loss value of the L2 norm loss function for the high-frequency information of the processed data and the high-frequency information of the sample label, and taking the sum of these two loss values as the loss value between the processed data and the sample label in the training sample, the finally obtained loss value can be made more accurate. Thus, by adopting this edge loss function, the edge fineness of the processed data of the deep learning model can be made higher.

[0102] Among them, the operation of the server adjusting the parameters in the untrained deep learning model according to the loss value can refer to related technologies, and the embodiments of this application will not elaborate on this in detail. For example, for any parameter in the untrained deep learning model, the server can obtain the partial derivative of this edge loss function with respect to this parameter according to the loss value and this parameter; subtract the product of the learning rate and the partial derivative of this parameter from this parameter to obtain the adjusted parameter. The learning rate can be set in advance, such as the learning rate can be 0.001, 0.000001, etc.

[0103] In the embodiments of this application, the server obtains multiple pairs of sample images, and the image quality of the final disparity map in each pair of sample images is higher than that of the initial disparity map. The server uses multiple initial disparity maps as the input data in the training sample, and uses multiple final disparity maps as the sample labels in the training sample. Then, the server uses this training sample to train the untrained deep learning model to obtain a trained deep learning model. This trained deep learning model can obtain a disparity map with higher image quality in a shorter time.

[0104] Next, for the above Figure 1The usage process of the deep learning model obtained through example training is described.

[0105] Figure 3 It is a flowchart of a parallax map processing method provided by an embodiment of the present application. Refer to Figure 3 This method includes the following steps.

[0106] Step 301: The terminal obtains a first parallax map, and the first parallax map is determined according to an image pair to be processed.

[0107] The two images in the image pair to be processed are two images taken at different angles in the same scene, that is, the image pair to be processed is obtained by shooting the same target with two different cameras. This image pair can be an image taken by the dual cameras in an electronic device, that is, it can be a binocular image pair. For example, this image pair can be an image taken by the camera with the highest pixel among the multiple cameras of a mobile phone and an image taken by a camera other than the camera with the highest pixel among the multiple cameras of the mobile phone. Optionally, this image pair can be an image pair that needs to perform dual-camera blurring.

[0108] The first parallax map can be directly calculated in a relatively simple manner. Optionally, the operation for the terminal to obtain the first parallax map can be: the terminal determines the first parallax map according to the image pair to be processed and a pre-stored stereo matching algorithm.

[0109] This stereo matching algorithm can find the corresponding pixel points of each pixel point in one image in the image of the other perspective, calculate the parallax map of these two images, and accordingly estimate the depth-of-field image. For example, this stereo matching algorithm can be the SGBM algorithm, the BM algorithm, the GC algorithm, etc., and the embodiments of the present application do not make a unique limitation on this.

[0110] The operation for the terminal to determine the first parallax map according to the image pair to be processed and the pre-stored stereo matching algorithm is similar to the Figure 1 operation in the above embodiment where the server determines the initial parallax map in the sample image pair A according to the first image pair and the pre-stored stereo matching algorithm, and the embodiments of the present application will not elaborate on this again.

[0111] Step 302: The terminal inputs the first parallax map into the trained deep learning model for processing, and outputs a second parallax map, and the image quality of the second parallax map is higher than that of the first parallax map.

[0112] The trained deep learning model is the one obtained through the above Figure 1A deep learning model trained by the deep learning model training method described in the embodiment. That is, the trained deep learning model is obtained by training an untrained deep learning model using training samples, and the training samples include multiple sample image pairs, and the image quality of the final disparity map in each sample image pair is higher than that of the initial disparity map.

[0113] In practical applications, the terminal can first calculate the first disparity map of the image pair in a simple manner, and then input the first disparity map into the trained deep learning model, and a second disparity map with higher image quality can be obtained in a short time.

[0114] For example, referring to Figure 4 , the first disparity map of the image pair can be as shown in the left figure in Figure 4 . After inputting the first disparity map into the trained deep learning model, a second disparity map as shown in the right figure in Figure 4 can be obtained. It can be seen that after being processed by the trained deep learning model, the edge error of the disparity map is significantly improved, and the processing time of this process on the mobile phone side is about 400 milliseconds, and the time consumption is also very short.

[0115] In the embodiment of the present application, after the terminal obtains the first disparity map of the image pair to be processed and inputs the first disparity map into the trained deep learning model, a second disparity map with higher image quality can be obtained in a short time.

[0116] Figure 5 is a schematic structural diagram of a deep learning model training device provided by an embodiment of the present application. Referring to Figure 5 , the device includes a second disparity map acquisition module 501, a training sample acquisition module 502, and a training module 503, where:

[0117] The second disparity map acquisition module 501 is configured to acquire multiple sample image pairs, and the image quality of the final disparity map in each sample image pair is higher than that of the initial disparity map;

[0118] The training sample acquisition module 502 is configured to use multiple initial disparity maps as input data in the training samples, and use multiple final disparity maps as sample labels in the training samples;

[0119] The training module 503 is configured to train an untrained deep learning model using the training samples to obtain a trained deep learning model.

[0120] Optionally, the sample image pair A is any one of the multiple sample image pairs, and the second disparity map acquisition module 501 is configured to:

[0121] Obtain a first image pair, which is obtained by two different cameras photographing the same target;

[0122] Determine an initial disparity map in the sample image pair A according to the first image pair and a pre-stored stereo matching algorithm.

[0123] Optionally, the second disparity map acquisition module 501 is used for:

[0124] Perform weighted least squares filtering on the initial disparity map in the sample image pair A to obtain a first disparity map;

[0125] Perform fast bilateral filtering on the first disparity map according to the images in the first image pair to obtain a second disparity map;

[0126] Perform refinement processing on the second disparity map according to the received image processing instruction to obtain the final disparity map in the sample image pair A, and the image processing instruction is used to indicate image smoothing processing or image sharpening processing.

[0127] Optionally, the second disparity map acquisition module 501 is used for:

[0128] Determine a third disparity map according to the first image pair and a pre-stored first stereo matching algorithm;

[0129] Determine a fourth disparity map according to the second image pair and a pre-stored second stereo matching algorithm, and the second image pair is obtained by reducing the size of each image in the first image pair by k times, where k is an integer greater than or equal to 1;

[0130] Enlarge the size of the fourth disparity map by k times to obtain a fifth disparity map;

[0131] Determine the initial disparity map in the sample image pair A according to the pixel values of the pixel points in the third disparity map and the pixel values of the pixel points in the fifth disparity map.

[0132] Optionally, the second disparity map acquisition module 501 is used for:

[0133] Determine a first pixel point at each position in the third disparity map and a second pixel point at the corresponding position in the fifth disparity map;

[0134] If the difference between the pixel value of the first pixel point and the pixel value of the second pixel point is less than the reference difference, then use the pixel value of the first pixel point as the pixel value of the pixel point at this position in the initial disparity map of the sample image pair A; or,

[0135] If the difference between the pixel value of the first pixel point and the pixel value of the second pixel point is greater than the reference difference, then use the pixel value of the second pixel point as the pixel value of the pixel point at this position in the initial disparity map of the sample image pair A; or,

[0136] If the difference between the pixel value of the first pixel point and the pixel value of the second pixel point is equal to the reference difference, then use the pixel value of the first pixel point or the pixel value of the second pixel point as the pixel value of the pixel point at this position in the initial disparity map in the sample image pair A.

[0137] Optionally, the training module 503 is configured to:

[0138] Input the input data in the training sample into an untrained deep learning model for processing, and output the processed data;

[0139] Determine the loss value between the processed data and the sample label in the training sample through a pre-stored edge loss function;

[0140] Adjust the parameters in the untrained deep learning model according to the loss value to obtain a trained deep learning model;

[0141] Wherein, the edge loss function is: Loss = L 2 (I g , I o ) + L 2 (I g ', I o )

[0142] Loss is the loss value, I o is the processed data, I o ' is the data obtained by subtracting the first data from the processed data, and the first data is the data obtained by performing a Gaussian smoothing operation on the processed data, I g is the sample label, I g ' is the data obtained by subtracting the second data from the sample label, and the second data is the data obtained by performing a Gaussian smoothing operation on the sample label, L 2 () is the L2 norm loss function.

[0143] In the embodiments of the present application, multiple sample image pairs are obtained, and the image quality of the final disparity map in each sample image pair is higher than the image quality of the initial disparity map. Use multiple initial disparity maps as the input data in the training sample, and use multiple final disparity maps as the sample labels in the training sample. Then, use the training sample to train an untrained deep learning model to obtain a trained deep learning model. The trained deep learning model can obtain a disparity map with higher image quality in a shorter time.

[0144] It should be noted that: when the deep learning model training device provided in the above embodiment is training the deep learning model, only the division of the above functional modules is used for illustration. In actual applications, the above functions can be assigned to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the deep learning model training device provided in the above embodiment and the embodiment of the deep learning model training method belong to the same concept. For the specific implementation process, please refer to the method embodiment and will not be elaborated here.

[0145] Figure 6 FIG. is a schematic structural diagram of a parallax map processing device provided by an embodiment of the present application. Refer to Figure 6 The device includes a first parallax map acquisition module 601 and a parallax map processing module 602, where:

[0146] The first parallax map acquisition module 601 is configured to acquire a first parallax map, and the first parallax map is determined according to an image pair to be processed;

[0147] The parallax map processing module 602 is configured to input the first parallax map into a trained deep learning model for processing and output a second parallax map, and the image quality of the second parallax map is higher than that of the first parallax map.

[0148] Optionally, the first parallax map acquisition module 601 is configured to:

[0149] Determine the first parallax map according to the image pair to be processed and a pre-stored stereo matching algorithm.

[0150] In the embodiment of the present application, after acquiring the first parallax map of the image pair to be processed and inputting the first parallax map into a trained deep learning model, a second parallax map with higher image quality can be obtained in a short time.

[0151] Optionally, the trained deep learning model is obtained by training an untrained deep learning model with training samples, and the training samples include multiple sample image pairs, and the image quality of the final parallax map in each sample image pair is higher than that of the initial parallax map.

[0152] It should be noted that: when the parallax map processing device provided in the above embodiment is processing the parallax map, only the division of the above functional modules is used for illustration. In actual applications, the above functions can be assigned to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the parallax map processing device provided in the above embodiment and the embodiment of the parallax map processing method belong to the same concept. For the specific implementation process, please refer to the method embodiment and will not be elaborated here.

[0153] Figure 7 The structural schematic diagram of a computer device provided by an embodiment of the present application. As Figure 7 shown, the computer device 7 includes: at least one processor 70 ( Figure 7 only one processor is shown in the figure), a memory 71, and a computer program 72 stored in the memory 71 and executable on at least one processor 70. When the processor 70 executes the computer program 72, the steps in the deep learning model training method in the above embodiment are implemented.

[0154] The computer device 7 can be a single server or a server cluster composed of multiple servers. Those skilled in the art can understand that Figure 7 this is only an example of the computer device 7 and does not constitute a limitation on the computer device 7. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0155] The processor 70 can be a central processing unit (CPU), and the processor 70 can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.

[0156] In some embodiments, the memory 71 can be an internal storage unit of the computer device 7, such as the hard disk or memory of the computer device 7. In other embodiments, the memory 71 can also be an external storage device of the computer device 7, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device 7. Further, the memory 71 can also include both the internal storage unit and the external storage device of the computer device 7. The memory 71 is used to store an operating system, application programs, a boot loader (BootLoader), data, and other programs, such as the program code of the computer program. The memory 71 can also be used to temporarily store data that has been output or will be output.

[0157] Figure 8A structural schematic diagram of a computer device provided by an embodiment of the present application. As Figure 8 shown, the computer device 8 includes: at least one processor 80 ( Figure 8 only one processor is shown in the figure), a memory 81, and a computer program 82 stored in the memory 81 and executable on at least one processor 80. When the processor 80 executes the computer program 82, the steps in the parallax map processing method in the above embodiment are implemented.

[0158] The computer device 8 may be a desktop computer, a notebook, a palm computer, or other terminals. Those skilled in the art can understand that Figure 8 merely examples of the computer device 8, which do not constitute a limitation on the computer device 8, may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0159] The processor 80 may be a CPU, and the processor 80 may also be other general-purpose processors, DSPs, ASICs, FPGAs, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0160] The memory 81 may be an internal storage unit of the computer device 8 in some embodiments, such as the hard disk or memory of the computer device 8. The memory 81 may also be an external storage device of the computer device 8 in other embodiments, such as a plug-in hard disk, SMC, SD card, flash card, etc. equipped on the computer device 8. Further, the memory 81 may also include both the internal storage unit and the external storage device of the computer device 8. The memory 81 is used to store an operating system, application programs, a boot loader, data, and other programs, such as the program code of the computer program. The memory 81 may also be used to temporarily store data that has been output or will be output.

[0161] In some embodiments, a computer-readable storage medium is also provided. The storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the deep learning model training method or the parallax map processing method in the above embodiment are implemented. For example, the computer-readable storage medium may be a ROM (Read-Only Memory), a RAM (Random Access Memory), a CD-ROM (Compact Disc Read-Only Memory), a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0162] It should be noted that the computer-readable storage medium mentioned in this application can be a non-volatile storage medium, in other words, it can be a non-transitory storage medium.

[0163] It should be understood that all or part of the steps of implementing the above embodiments can be achieved through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. These computer instructions can be stored in the above-mentioned computer-readable storage medium.

[0164] That is to say, in some embodiments, there is also provided a computer program product containing instructions, which, when running on a computer, causes the computer to execute the steps of the deep learning model training method or the disparity map processing method in the above embodiments.

[0165] The above are the optional embodiments provided by this application, which are not intended to limit this application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this application shall be included within the protection scope of this application.

Claims

1. A method for processing a disparity map, characterized in that, it includes: obtaining a first disparity map, where the first disparity map is determined according to an image pair to be processed; inputting the first disparity map into a trained deep learning model for processing, and outputting a second disparity map, where the image quality of the second disparity map is higher than that of the first disparity map; wherein, the deep learning model includes an input layer, multiple hidden layers and an output layer, and the multiple hidden layers sequentially include a convolutional layer, one or more dilated convolutional layers, a deconvolutional layer, a softmax layer, and an output layer; the input layer is used to receive input data; the convolutional layer is used to perform a convolution operation on the input data to obtain multiple first feature maps; the one or more dilated convolutional layers are used to perform a dilated convolution operation on the multiple first feature maps to obtain multiple second feature maps; the deconvolutional layer is used to perform a deconvolution operation on the multiple second feature maps to obtain n third feature maps; the softmax layer is used to calculate the scores of each of the n third feature maps among the n third feature maps to obtain n scores, where n is an integer greater than or equal to 4; the output layer is used to transform the n scores into a filtering matrix of m rows and m columns, and multiply the pixel value of each pixel point in the input data by the filtering matrix one by one to obtain processed data.

2. The method according to claim 1, characterized in that, the obtaining of the first disparity map includes: determining the first disparity map according to the image pair to be processed and a pre-stored stereo matching algorithm.

3. The method according to claim 1 or 2, characterized in that, the trained deep learning model is obtained by training an untrained deep learning model with training samples, and the training samples include multiple sample image pairs, and the image quality of the final disparity map in each sample image pair is higher than that of the initial disparity map.

4. A method for training a deep learning model, characterized in that, it includes: obtaining multiple sample image pairs, where the image quality of the final disparity map in each sample image pair is higher than that of the initial disparity map; using multiple initial disparity maps as input data in the training samples, and using multiple final disparity maps as sample labels in the training samples; training an untrained deep learning model with the training samples to obtain a trained deep learning model; wherein, the deep learning model includes an input layer, multiple hidden layers and an output layer, and the multiple hidden layers sequentially include a convolutional layer, one or more dilated convolutional layers, a deconvolutional layer, a softmax layer, and an output layer; the input layer is used to receive input data; the convolutional layer is used to perform a convolution operation on the input data to obtain multiple first feature maps; the one or more dilated convolutional layers are used to perform a dilated convolution operation on the multiple first feature maps to obtain multiple second feature maps; the deconvolutional layer is used to perform a deconvolution operation on the multiple second feature maps to obtain n third feature maps; the softmax layer is used to calculate the scores of each of the n third feature maps among the n third feature maps to obtain n scores, where n is an integer greater than or equal to 4; The output layer is used to transform the n scores into a filtering matrix of m rows and m columns, and multiply the pixel value of each pixel point in the input data by the filtering matrix one by one to obtain the processed data.

5. The method according to claim 4, wherein, Sample image pair A is any one of the multiple sample image pairs. The obtaining of the initial disparity map in sample image pair A includes: Obtaining a first image pair, which is obtained by shooting the same target with two different cameras; Determining the initial disparity map in sample image pair A according to the first image pair and a pre-stored stereo matching algorithm.

6. The method according to claim 5, wherein, The obtaining of the final disparity map in sample image pair A includes: Performing weighted least squares filtering on the initial disparity map in sample image pair A to obtain a first disparity map; Performing fast bilateral filtering on the first disparity map according to the images in the first image pair to obtain a second disparity map; Performing refinement processing on the second disparity map according to the received image processing instruction to obtain the final disparity map in sample image pair A, and the image processing instruction is used to indicate image smoothing processing or image sharpening processing.

7. The method according to claim 6, wherein, The determining of the initial disparity map in sample image pair A according to the first image pair and a pre-stored stereo matching algorithm includes: Determining a third disparity map according to the first image pair and a pre-stored first stereo matching algorithm; Determining a fourth disparity map according to a second image pair and a pre-stored second stereo matching algorithm, where the second image pair is obtained by reducing the size of each image in the first image pair by k times, and k is an integer greater than or equal to 1; Enlarging the size of the fourth disparity map by k times to obtain a fifth disparity map; Determining the initial disparity map in sample image pair A according to the pixel value of the pixel point in the third disparity map and the pixel value of the pixel point in the fifth disparity map.

8. The method according to claim 7, wherein, The determining of the initial disparity map in sample image pair A according to the pixel value of the pixel point in the third disparity map and the pixel value of the pixel point in the fifth disparity map includes: Determining a first pixel point at each position in the third disparity map and a second pixel point at the corresponding position in the fifth disparity map; If the difference between the pixel value of the first pixel point and the pixel value of the second pixel point is less than a reference difference, then using the pixel value of the first pixel point as the pixel value of the pixel point at the position in the initial disparity map in sample image pair A; or, If the difference between the pixel value of the first pixel point and the pixel value of the second pixel point is greater than the reference difference, then using the pixel value of the second pixel point as the pixel value of the pixel point at the position in the initial disparity map in sample image pair A; or, If the difference between the pixel value of the first pixel point and the pixel value of the second pixel point is equal to the reference difference, then use the pixel value of the first pixel point or the pixel value of the second pixel point as the pixel value of the pixel point at the position in the initial disparity map in the sample image pair A.

9. The method according to any one of claims 4-8, characterized in that the training of the untrained deep learning model using the training samples to obtain a trained deep learning model includes: inputting the input data in the training samples into the untrained deep learning model for processing, and outputting the processed data; determining the loss value between the processed data and the sample labels in the training samples through a pre-stored edge loss function; adjusting the parameters in the untrained deep learning model according to the loss value to obtain a trained deep learning model; wherein, the edge loss function is: Loss = L 2 (I g , I o ) + L 2 (I g ', I o ') The Loss is the loss value, the I o is the processed data, the I o ' is the data obtained by subtracting the first data from the processed data, and the first data is the data obtained by performing a Gaussian smoothing operation on the processed data, the I g is the sample label, the I g ' is the data obtained by subtracting the second data from the sample label, and the second data is the data obtained by performing a Gaussian smoothing operation on the sample label, the L 2 () is the L2 norm loss function.

10. A disparity map processing device, characterized in that it includes: A first disparity map acquisition module for acquiring a first disparity map, where the first disparity map is determined according to an image pair to be processed; A disparity map processing module for inputting the first disparity map into a trained deep learning model for processing and outputting a second disparity map, where the image quality of the second disparity map is higher than the image quality of the first disparity map; wherein, the deep learning model includes an input layer, a plurality of hidden layers and an output layer, and the plurality of hidden layers sequentially include a convolutional layer, one or more dilated convolutional layers, a transposed convolutional layer, a softmax layer, and an output layer; The input layer is used to receive input data; The convolutional layer is used to perform a convolution operation on the input data to obtain a plurality of first feature maps; The one or more dilated convolutional layers are used to perform dilated convolution operations on the plurality of first feature maps to obtain a plurality of second feature maps; The transposed convolutional layer is used to perform a transposed convolution operation on the plurality of second feature maps to obtain n third feature maps; The softmax layer is used to calculate the scores of each of the n third feature maps in the n third feature maps to obtain n scores, where n is an integer greater than or equal to 4; The output layer is used to transform the n scores into a filtering matrix of m rows and m columns, and multiply the pixel value of each pixel point in the input data by the filtering matrix one by one to obtain the processed data.

11. A deep learning model training device, characterized in that it includes: A second disparity map acquisition module for acquiring a plurality of sample image pairs, where the image quality of the final disparity map in each sample image pair is higher than the image quality of the initial disparity map; A training sample acquisition module for using a plurality of initial disparity maps as the input data in the training samples and using a plurality of final disparity maps as the sample labels in the training samples; A training module for training an untrained deep learning model using the training samples to obtain a trained deep learning model; wherein, the deep learning model includes an input layer, a plurality of hidden layers and an output layer, and the plurality of hidden layers sequentially include a convolutional layer, one or more dilated convolutional layers, a transposed convolutional layer, a softmax layer, and an output layer; The input layer is used to receive input data; The convolutional layer is used to perform a convolution operation on the input data to obtain a plurality of first feature maps; The one or more dilated convolutional layers are used to perform a dilated convolution operation on the plurality of first feature maps to obtain a plurality of second feature maps; The transposed convolutional layer is used to perform a transposed convolution operation on the plurality of second feature maps to obtain n third feature maps; The softmax layer is used to calculate the scores of each of the n third feature maps among the n third feature maps to obtain n scores, where n is an integer greater than or equal to 4; The output layer is used to transform the n scores into a filtering matrix of m rows and m columns, and multiply the pixel value of each pixel point in the input data by the filtering matrix one by one to obtain the processed data.

12. A computer device, characterized in that, the computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the computer program is executed by the processor, it implements the method according to any one of claims 1-3 or any one of claims 4-9.

13. A computer-readable storage medium, characterized in that, the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the method according to any one of claims 1-3 or any one of claims 4-9.

Citation Information

Patent Citations

  • A binocular disparity estimation method based on cascaded geometric context neural network

    CN109472819A

  • Image optimization network training method and device, image optimization method and device and apparatus

    CN110675357A