Target depth estimation model training method and target depth estimation method
By acquiring prior depth information before training the target depth prediction model and adjusting the model parameters in conjunction with image labels, the problem of insufficient training accuracy in existing technologies is solved, and more accurate and stable target depth estimation is achieved.
Patent Information
- Application Number
- CN202211214517.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-30
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2042-09-30
AI Technical Summary
Existing training methods for depth regression estimation models are unable to meet pre-set regression validation metrics, resulting in insufficient training accuracy and inaccurate target depth estimation.
Before training the target depth prediction model, the depth prior information of the sample images is obtained, and the model is trained by combining the image labels and sample images. The depth prior information is used as a reference to adjust the model node parameters.
It improves the accuracy and stability of model training, enhances depth perception capabilities, and shortens training convergence time.
Smart Images

Figure CN115526926B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of target perception, and in particular to a target depth estimation model training method and a target depth estimation method. BACKGROUND
[0002] In the technical field of assisted driving, robot control, etc., existing environment perception methods include image processing-based environment perception methods. The image processing-based environment perception method uses a pre-trained depth regression estimation model to estimate the depth of a target in a captured image according to the pixel grayscale of each pixel in the captured image. The training accuracy and generalization quality of the depth regression estimation model determine the accuracy of target depth estimation in actual use.
[0003] However, when training a depth regression estimation model using training samples, the training accuracy is difficult to reach the pre-set regression verification indicators. That is, it is difficult to train a good quality depth regression estimation model using existing depth regression estimation model training methods. SUMMARY
[0004] To solve the above technical problems, the present disclosure provides a target depth estimation model training method and a target depth estimation method.
[0005] In a first aspect, the present disclosure provides a target depth estimation model training method, comprising:
[0006] Obtaining a plurality of groups of training samples, each group of training samples comprising a sample image and an image label, the image label comprising a labeled depth of a target of interest in the sample image;
[0007] Obtaining depth prior information corresponding to the sample image, the depth prior information comprising a prior depth of a pixel in the sample image;
[0008] Training a target depth prediction model using the sample image, the depth prior information, and the image label.
[0009] In a second aspect, the present disclosure provides a target depth prediction method, comprising:
[0010] Obtaining a captured image, and obtaining depth prior information corresponding to the captured image, the depth prior information comprising a prior depth of a pixel in the sample image;
[0011] Inputting the sample image and the depth prior information into a target depth prediction model to obtain a predicted depth of a target of interest in the captured image;
[0012] The target depth prediction model is trained using the target depth prediction model training method described above.
[0013] In a third aspect, the embodiments of the present disclosure provide a target depth estimation model training apparatus, comprising:
[0014] a training sample acquisition unit configured to acquire a plurality of groups of training samples, each group of training samples comprising a sample image and an image label, the image label comprising a labeled depth of a target of interest in the sample image;
[0015] a depth prior information acquisition unit configured to acquire depth prior information corresponding to the sample image, the depth prior information comprising a prior depth of a pixel in the sample image;
[0016] a model training unit configured to train a target depth prediction model using the sample image, the depth prior information, and the image label.
[0017] In a fourth aspect, the embodiments of the present disclosure provide a target depth prediction apparatus, comprising:
[0018] an input data acquisition unit configured to acquire a photographed image and acquire depth prior information corresponding to the photographed image, the depth prior information comprising a prior depth of a pixel in the sample image;
[0019] a prediction unit configured to input the sample image and the depth prior information into a target depth prediction model to obtain a predicted depth of a target of interest in the photographed image;
[0020] wherein the target depth prediction model is trained by the target depth prediction model training apparatus as described above.
[0021] In a fifth aspect, the embodiments of the present disclosure provide a computing device comprising a processor and a memory, the memory being configured to store a computer program; the computer program, when loaded by the processor, causes the processor to execute the target depth estimation model training method or the target depth prediction method as described above.
[0022] In a sixth aspect, the embodiments of the present disclosure provide a computer readable storage medium, the storage medium storing a computer program, the computer program, when executed by a processor, causing the processor to implement the target depth estimation model training method or the target depth prediction method as described above.
[0023] The technical solutions provided by the embodiments of the present disclosure have the following advantages compared with the prior art:
[0024] The model training method provided in the embodiments of the present disclosure first acquires the depth prior information of the sample image before training the target depth prediction model, and then trains the target depth prediction model by using the depth prior information and the sample image and the image label in the training sample. During the model training, the prior depth in the depth prior information can be used as a reference for the estimated depth of the target, and the model node parameters are adjusted in combination with the pixel information of the sample image and the labeled depth in the image label. By using the prior depth as a reference for the estimated depth of the target, and using the reference when training the node parameters of the target depth prediction model, the model node can implicitly use the reference as the basis for adjusting the node parameters. Since the reference can be implicitly used as the basis for adjusting the node parameters, the target depth prediction model training method provided in the present solution has more effective information amount of the training sample than the model training method of the prior art, and thus the target depth prediction model obtained by training is more accurate and has more stable depth perception ability. BRIEF DESCRIPTION OF DRAWINGS
[0025] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and serve to explain the principles of the present disclosure together with the specification.
[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the accompanying drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, for those skilled in the art, other drawings can also be obtained from these drawings without creative labor, and the drawings are as follows:
[0027] Figure 1 is a flowchart of the target depth prediction model training method provided in the embodiments of the present disclosure;
[0028] Figure 2 is a flowchart of the method for acquiring the depth prior information provided in some embodiments of the present disclosure;
[0029] Figure 3 is a flowchart of the method for acquiring the depth prior information provided in some embodiments of the present disclosure;
[0030] Figure 4 is a flowchart of the target depth prediction method provided in some embodiments of the present disclosure;
[0031] Figure 5 is a flowchart of the method for acquiring the depth prior information corresponding to the photographed image provided in some embodiments of the present disclosure;
[0032] Figure 6 is a structural schematic diagram of the target depth estimation model training device provided in some embodiments of the present disclosure;
[0033] Figure 7 is a structural schematic diagram of a target depth prediction device provided by some embodiments of the present disclosure.
[0034] Figure 8 is a structural schematic diagram of a computing device provided by some embodiments of the present disclosure. DETAILED DESCRIPTION
[0035] Embodiments of the present disclosure will be described in more detail with reference to the drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms, and should not be interpreted as being limited to the embodiments set forth herein, but rather, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes, and are not intended to limit the scope of protection of the present disclosure.
[0036] The term "comprising" and variations thereof as used herein are open-ended, that is, "comprising but not limited to". The term "based on" is "based, at least in part, on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Related definitions will be given in the following description. It should be noted that the concepts of "first", "second", etc. mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not intended to limit the order or interdependence of the functions performed by these devices, modules or units.
[0037] It should be noted that the modification of "one" or "multiple" mentioned in the present disclosure is illustrative and not limiting, and those skilled in the art should understand that unless the context clearly indicates otherwise, it should be understood as "one or more".
[0038] The embodiments of the present disclosure provide a target depth prediction model training method and a target depth prediction method based on the aforementioned model training method. The method provided by the embodiments of the present disclosure realizes the training of the target depth prediction model or the estimation of the target depth by obtaining prior depth information and comprehensively processing the prior depth information and image information.
[0039] It should be noted that the target depth prediction model training method and the target depth prediction method provided by the embodiments of the present disclosure are executed by a computing device, which can be a server or a terminal computing device such as a smart phone, a tablet computer, a notebook computer, etc.
[0040] Figure 1 is a flowchart of a target depth prediction model training method provided by the embodiments of the present disclosure. As shown in Figure 1As shown, the target depth prediction model training method includes S110-S130.
[0041] S110: Obtain a plurality of sets of training samples, each set of training samples including a sample image and an image label.
[0042] The training sample is a sample used to implement training of the target depth prediction model. According to the function definition (estimating the depth of the target object in the image) of the target depth prediction model to be trained, the input data and the output data, the training sample should include the sample image and the image label.
[0043] The image label is a label used to label the sample image. In the embodiments of the present disclosure, the image label includes the labeled depth of the target of interest in the sample image. The labeled depth is the depth of the target of interest to the reference coordinate point in the sample image determined by the labeling method. It should be noted here that the depth refers to the distance in the vertical direction of the target object relative to the reference coordinate.
[0044] It should be noted that, in addition to including the labeled depth, the image label of the sample object in subsequent model training should also include the position information of the target of interest in the sample image. The foregoing position information can be represented in the form of a rectangular detection box, and can specifically be represented in the form of rectangular detection box diagonal point coordinates.
[0045] It should also be noted that the number of the plurality of sets of training samples should meet the model training requirements, which is determined by the number of nodes of the model and the accuracy requirements of the model. The number of sets of training samples required is not expanded in the present embodiment, and can be referred to in related technical literature.
[0046] S120: Obtain depth prior information corresponding to the sample image, the depth prior information including the prior depth of the pixels in the sample image.
[0047] In the embodiments of the present disclosure, the depth prior information corresponding to the sample image includes the prior depth of the pixels in the sample image. The prior depth is the depth of the pixels in the sample image estimated by some depth estimation method. The "prior" in the prior depth is relative to the estimated depth determined by using the target depth prediction model in the prior art.
[0048] How the embodiments of the present disclosure obtain the depth prior information corresponding to the sample image, that is, how to obtain the prior depth of the pixels in the sample image, will be described in detail later.
[0049] It should be noted that the prior depth information includes the prior depth of the pixels in the sample image, which can be the prior depth of all pixels in the sample image, or can be the prior depth of only part of the pixels in the sample image. The aforementioned part of the pixels refers to the pixels in the sample image representing the target of interest, that is, the pixels in the aforementioned rectangular detection frame.
[0050] It should also be noted that the depth prior information includes not only the prior depth of the pixels in the sample image, but also the one-to-one correspondence between each prior depth and the pixels in the sample image.
[0051] In order to reflect the one-to-one correspondence between the prior depth and the pixels in the sample image, the prior depth information can be represented in the form of a two-dimensional matrix in the embodiment of the present disclosure. The number of rows of the aforementioned two-dimensional matrix is the number of pixels in the length direction of the sample image, and the number of columns of the two-dimensional matrix is the number of pixels in the height direction of the sample image.
[0052] In the case of representing the prior depth information in the form of a two-dimensional matrix, the prior depth information can be understood as a layer or an independent channel provided by the target image. In combination with the foregoing description, in the case where the prior depth information includes the prior depth of all pixels in the image, each element in the two-dimensional matrix is the prior depth of the corresponding pixel. In the case where the prior depth information only includes the prior depth of part of the pixels in the image, the prior depth of the part of the pixels is filled into the corresponding element position of the two-dimensional matrix, and 0 or a blank symbol can be filled into the other element positions.
[0053] S130: training the target depth prediction model using the sample image, the depth prior information, and the image label.
[0054] After the computing device obtains the training sample and the depth prior information using the foregoing S110-S120, the computing device can train the target depth prediction model using the sample image and the image label in the training sample, and the depth prior information.
[0055] Training the target depth prediction model using the sample image and the image label in the training sample, and the depth prior information is to input the sample image and the depth prior information into the model to obtain an output result, the output result including an output depth estimate value of a specific target of interest, and then comparing the depth estimate value with the labeled depth in the image label to determine a difference value. After determining the difference value, the difference value can be used to adjust the node parameters of the target depth prediction model using methods such as the backpropagation algorithm, until the target depth prediction model reaches a set accuracy standard.
[0056] The target depth prediction model can adopt various existing model architectures that can be used for image processing, and the embodiments of the present disclosure do not make special limitations. Specifically, which model architecture can be adopted and how to train the target depth prediction model using sample images, depth prior information, and image labels can be referred to relevant technical literature, and the embodiments of the present disclosure will not be described in detail.
[0057] The target depth prediction model training method provided by the embodiments of the present disclosure acquires the depth prior information of the sample image before training the target depth prediction model, and then trains the target depth prediction model using the depth prior information and the sample images and image labels in the training samples. During model training, the prior depth in the depth prior information can be used as a reference for the estimated depth of the target of interest (more specifically, the pixel corresponding to the target of interest), and the model node parameters are adjusted in combination with the pixel information of the sample image and the labeled depth in the image label. Using the prior depth as a reference for the estimated depth of the target of interest (of course, this reference may not be accurate, but in typical scenarios, this reference is likely to be accurate), and using this reference when training the node parameters of the target depth prediction model, the model node can implicitly use the reference as the basis for adjusting the node parameters.
[0058] Because the reference can be implicitly used as the basis for adjusting the node parameters, the target depth prediction model training method provided by the present solution has more effective information than the training samples used by the model training method of the prior art, and thus the target depth prediction model obtained by training is more accurate and has more stable depth perception capability.
[0059] In addition, in some implementations, because the depth prior information is used during model training, the convergence speed of the model is faster than the convergence speed of the existing model training, that is, the model training overhead is reduced.
[0060] The foregoing S120 mentioned that the depth prior information corresponding to the sample image needs to be acquired. Figure 2 is a method flowchart for acquiring depth prior information provided by some embodiments of the present disclosure. As shown in Figure 2 In some embodiments, the method of acquiring depth prior information by the computing device includes S210-S220.
[0061] S210: Perform a transmission transformation on the pixels in the sample image to obtain the prior depth corresponding to the pixels.
[0062] Transmission transformation is also called perspective transformation, which is a transformation method of converting an image from two-dimensional coordinates to three-dimensional space coordinates, and then converting the three-dimensional coordinates to another two-dimensional coordinates.
[0063] In the embodiments of the present disclosure, the corresponding prior depth of the pixel in the sample image is obtained by performing the transmission transformation on the pixel in the sample image, that is, the pixel in the sample image is converted to a ground plane, and the corresponding prior depth of the pixel is determined by the coordinate of the ground plane. How to perform the transmission transformation on the pixel in the sample image to obtain the corresponding prior depth will be analyzed by example in the following.
[0064] S220: Obtain the depth prior information by combining the corresponding prior depth according to the coordinates of the pixels in the sample image.
[0065] After obtaining the corresponding prior depth of each pixel, the computing device can place the corresponding prior depth at a specific position according to the coordinates of the pixels in the sample image, that is, the corresponding prior depth is arranged and combined according to the coordinates of the pixels in the sample image to obtain the depth prior information.
[0066] As mentioned above, in the case that the depth prior information is a two-dimensional matrix, the combination of the corresponding prior depth according to the coordinates of the pixels in the sample image to obtain the depth prior information is to fill the coordinates of the pixels in the sample image into the corresponding positions in the two-dimensional matrix.
[0067] It should be noted that in the case of obtaining the depth prior information by using S210-S220, the same transmission transformation method and the same transmission transformation parameter are used for each sample image, so the obtained prior depth and prior depth information are the same. However, the same depth prior information does not hinder the implementation of the foregoing target depth prediction model training method, and does not hinder the use of the foregoing target depth prediction model training method to obtain a more accurate target depth prediction model.
[0068] In actual application, because the depth prior information corresponding to different sample images is the same, the computing device can determine the depth prior information by using the foregoing S210-S220 and store it, and load the stored depth prior information as input data during model training.
[0069] Optionally, in some embodiments of the present disclosure, before performing the foregoing S210-S220, the computing device can further perform S230-S240.
[0070] S230: Obtain the internal and external parameters of the shooting camera and the internal and external parameters of the virtual camera, and the virtual camera is a virtual camera that is virtually set to shoot the road image at a vertical downward angle.
[0071] The shooting camera is at least one camera that is actually used to shoot the sample image and has the same internal and external parameters. The virtual camera is a virtual camera that is virtually set to shoot the road image at a vertical downward angle, that is, the imaging plane of the virtual camera is parallel to the road plane.
[0072] In some embodiments of the present disclosure, for the convenience of the transmission transformation, the projection of the optical axis of the virtual camera on the ground plane can coincide with the projection of the reference point of the imaging plane of the shooting camera on the ground plane. Of course, in other embodiments, the projection of the optical axis of the virtual camera on the ground plane can also not coincide with the projection of the reference point of the imaging plane of the shooting camera on the ground plane, in which case some pixel translation transformation needs to be performed while performing the transmission transformation.
[0073] The aforementioned intrinsic and extrinsic parameters include intrinsic parameters and extrinsic parameters. The intrinsic parameters of the camera can include focal length and distortion parameters, and the extrinsic parameters can include rotation parameters and translation parameters. In a specific implementation, the intrinsic parameters of the camera can be represented by an intrinsic parameter matrix, and the extrinsic parameters of the camera can be represented by a rotation matrix and a translation vector.
[0074] S240: According to the intrinsic and extrinsic parameters of the shooting camera and the intrinsic and extrinsic parameters of the virtual camera, a homography matrix of the imaging plane of the shooting camera to the road plane is constructed.
[0075] The homography matrix is a matrix that maps and converts the points on the same plane in the three-dimensional space to the imaging plane coordinates of the two cameras. The homography matrix in the embodiments of the present disclosure is the matrix of the imaging plane of the shooting camera to the road plane. The homography matrix can be determined according to the intrinsic and extrinsic parameters of the two cameras, and in the embodiments of the present disclosure, the homography matrix can be constructed according to the intrinsic and extrinsic parameters of the shooting camera and the intrinsic and extrinsic parameters of the virtual camera.
[0076] In some embodiments of the present disclosure, the intrinsic parameter matrix constructed according to the intrinsic parameters of the virtual camera is The rotation matrix constructed according to the rotation parameters of the extrinsic parameters of the virtual camera is The translation vector constructed according to the translation parameters of the extrinsic parameters of the virtual camera is In this case, the virtual camera is a camera virtually set at 1 m above the ground. The rotation matrix of the extrinsic parameters of the real camera is R, the translation vector of the real camera is t, and the intrinsic parameter matrix of the real camera is K. According to the coordinate conversion relationship, the rotation matrix between the real camera and the virtual camera can be obtained as And the translation matrix between the real camera and the virtual camera is
[0077] According to the homography relationship between the real camera and the virtual camera, the homography matrix of the real camera to the virtual camera can be obtained as Where n is the normal vector of the road plane n = [0 0 1], and the aforementioned homography matrix is the homography matrix of the shooting camera to the road plane. Note that t'n is the product of the column vector and the row vector, and the homography matrix calculated by the aforementioned formula is a 3x3 matrix. For the sake of convenience, in the embodiments of the present disclosure, the homography matrix H is represented as .
[0078] In the case of determining the homography matrix by using the foregoing S230-S240, the transmission transformation of each pixel in the sample image in S210 can include S231.
[0079] S231: performing transmission transformation on the pixels of the sample image based on the homography matrix to obtain the prior depth corresponding to the pixels.
[0080] The transmission transformation on the pixels of the sample image based on the homography matrix is to calculate the corresponding prior depth by using the parameters in the homography matrix.
[0081] pixel coordinates of a certain point pixel in the sample image corresponding road ground point coordinates According to the operation rule of the homography matrix, Because y represents the depth, only y, i.e., the prior depth, is considered in the following.
[0082] According to the foregoing description, the transmission transformation on the pixels of the sample image based on the homography matrix in S231 to obtain the prior depth corresponding to the pixels can be specifically S2311-S2312.
[0083] S2311: constructing a depth calculation formula according to the parameters in the homography matrix;
[0084] S2312: obtaining the prior depth corresponding to the pixels according to the pixel coordinates of each pixel in the sample image and the depth calculation formula.
[0085] However, in actual applications, the foregoing formula is directly used to calculate the prior depth, which is not accurate for some points in the sample image. The foregoing inaccurate problem is analyzed first, and then the method for solving the inaccurate problem is described.
[0086] For the formula divide the numerator and the denominator by e8 to obtain and make Because and are very small quantities, a can be regarded as a constant, and b is a variable only related to v. In the case of , y is a maximum value, but this is not consistent with the actual situation.
[0087] To solve the foregoing problem, the computing device in the embodiment of the disclosure can construct a depth calculation formula according to the parameters in the homography matrix by using S2311A-S2311B as follows.
[0088] S2311A: determining the split longitudinal coordinate according to the parameters in the homography matrix.
[0089] Specifically, the computing device determines the split vertical coordinate according to the parameters in the homography matrix, and the split vertical coordinate is determined according to e8 and e9 as described above. In specific embodiments, the split vertical coordinate can be determined according to The vanishing point vertical coordinate of the image is determined, and then the split vertical coordinate is obtained by adding a preset value to the vanishing point vertical coordinate. For example, the split vertical coordinate can be determined as
[0090] It should be noted that the foregoing operations are performed on the premise that the vertical coordinates of the sample image gradually increase from top to bottom.
[0091] S2311B: constructing a first depth calculation formula for the pixels in the sample image whose vertical coordinates are less than the split vertical coordinate, and constructing a second depth calculation formula for the pixels in the sample image whose vertical coordinates are greater than the first vertical coordinate.
[0092] In the embodiments of the present disclosure, the first depth calculation formula is a formula obtained by performing multiple Taylor expansions based on the split vertical coordinate. Specifically, the first calculation formula can be a formula obtained by performing three Taylor expansions based on the split vertical coordinate, and specifically is The foregoing According to the foregoing analysis, the second calculation formula is
[0093] In the case of determining the first depth calculation formula and the second depth calculation formula by using the foregoing S2311A-S2311B, the computing device performing S2312 can specifically include: for the first pixel in the sample image whose vertical coordinate is less than or equal to the split vertical coordinate, calculating the prior depth by using the vertical coordinate of the first pixel and the first calculation formula; and for the second pixel in the sample image whose vertical coordinate is greater than the split vertical coordinate, calculating the prior depth by using the vertical coordinate of the second pixel and the second calculation formula.
[0094] As described above, the computing device can only perform once when performing S210-S220, and the same computing device can also only perform once when performing the foregoing S230-S240. Of course, in some actual applications, the computing device can also obtain the depth prior information of the sample image determined by other devices using the foregoing method for performing S130.
[0095] It is mentioned in the foregoing S120 that the depth prior information corresponding to the sample image needs to be obtained. Figure 3 is a method flowchart for obtaining the depth prior information provided by some embodiments of the present disclosure. As Figure 3 shown, in some embodiments, the method for obtaining the depth prior information by the computing device includes S310-S350.
[0096] S310: performing pre-recognition processing on the sample image to determine the detection frame of the target of interest in the sample image and the type of the target of interest.
[0097] The pre-recognition processing of the sample image by the computing device is to process the sample image by using a pre-trained target recognition model, to determine a detection box of the target of interest in the sample image, and to determine the type of the target of interest according to the features of the pixels in the detection box. The foregoing target recognition model can be various models already used in the art, and the embodiments of the present disclosure are not limited thereto. However, it should be noted that the target recognition model does not have the ability to determine the estimated depth of the target of interest.
[0098] S320: determining a typical width of the target of interest according to the type of the target of interest.
[0099] In the case of determining the type of the target of interest, the computing device can look up a pre-set table according to the type of the target of interest to determine the typical width of the target of interest. The typical width of the target of interest is a pre-set width. For example, for a domestic passenger vehicle, the corresponding typical width can be set to 2 m; for a bicycle, the corresponding typical width is 0.8 m (taking into account the width of the person riding the bicycle); and for a pedestrian, the corresponding typical width is 0.6 m.
[0100] S330: calculating the estimated depth of the target of interest to the camera according to the typical width and the pixel width of the detection box.
[0101] According to the principle of pinhole imaging, in the case of determining the width of the target of interest, the closer the distance between the target of interest and the camera, the more the number of pixels occupied in the lateral direction on the camera; on the contrary, the farther the distance between the target of interest and the camera, the less the number of pixels occupied in the lateral direction on the camera. Therefore, according to the typical distance and the pixel width of the detection box (i.e., the pixel width of the target of interest), the estimated depth of the target of interest to the camera can be generally calculated.
[0102] S340: taking the estimated depth as the prior depth of each pixel in the detection box.
[0103] S350: combining the corresponding prior depth according to the coordinates of the pixels in the sample image to obtain the depth prior information.
[0104] After obtaining the corresponding prior depth of each pixel, the computing device can place the corresponding prior depth at a specific position according to the coordinates of the pixels in the sample image, that is, arrange and combine the corresponding prior depth according to the coordinates of the pixels in the sample image to obtain the depth prior information.
[0105] As mentioned before, in the case that the depth prior information is a two-dimensional matrix, the combination of the corresponding prior depth according to the coordinates of the pixels in the sample image to obtain the depth prior information is to fill the coordinates of the pixels in the sample image into the corresponding positions in the two-dimensional matrix. It should be noted that because some regions in the sample image are not selected as the detection frame of the target of interest, the pixels in these regions do not have the corresponding prior depth, and the prior depth of these regions should be set to 0 or a null value when performing S350.
[0106] By adopting the foregoing S310-S350, the computing device determines the corresponding specific depth prior information for each sample image, so that the prior depth information and the sample image have a one-to-one correspondence, and has uniqueness according to the characteristics of the sample image. Training the target depth prediction model using the depth prior information with uniqueness can make the input data more distinctive, and thus make the trained model more accurate.
[0107] As mentioned before, when performing S220 or S350, the combination of the corresponding prior depth according to the coordinates of the pixels in the sample image to obtain the depth prior information. In some embodiments of the present disclosure, the computing device can also perform S100 before performing S220 or S350.
[0108] S100: encode the prior depth using a pre-set encoding function to obtain the prior depth encoding corresponding to the pixel.
[0109] In specific implementations, the computing device can encode the prior depth using functions such as Log function, Sigmoid function, etc. to obtain the prior depth encoding corresponding to the pixel.
[0110] In the case of performing the foregoing S410, the combination of the corresponding prior depth according to the coordinates of the pixels in the sample image in S220 or S350 is specifically: combination of the corresponding prior depth encoding according to the coordinates of the pixels in the sample image.
[0111] In some embodiments of the present disclosure, the target depth prediction model only has one input interface. In this case, the computing device can combine the sample image and the depth prior information and input them into the foregoing input interface to realize the training of the target depth prediction model. In some specific implementations, the computing device can combine the target image and the depth prior information in sequence to obtain combined data, and then input the combined data into the target depth prediction model. In some other implementations, the computing device can combine a certain pixel of the target image and the corresponding depth prior information into a data pair, then combine the data pairs into an input matrix according to the sequence of the pixels, and then input the input matrix into the target depth prediction model.
[0112] In some other embodiments of the present disclosure, the target depth prediction model comprises two data input interfaces and two processing sub-modules. The two processing sub-modules are a pre-processing sub-module and a post-processing sub-module, respectively. The pre-processing sub-module corresponds to the front data input interface, and the post-processing module corresponds to the rear data input interface. The front data input interface is an interface for inputting sample images, and the rear data input interface is an interface for inputting depth prior information. In this case, the S130 performed by the computing device can specifically include S131-S134.
[0113] S131: processing the sample image using the pre-processing sub-module to obtain an intermediate output tensor.
[0114] S132: splicing the intermediate output tensor and the depth prior information to obtain a spliced tensor.
[0115] In specific implementations, splicing the intermediate output tensor and the depth prior information to obtain the spliced tensor can have the following cases: (1) in the case where the dimensions of the intermediate output tensor and the depth prior information are not the same, the two can be directly spliced to obtain the spliced tensor; (2) in the case where the dimensions of the intermediate output tensor and the depth prior information are the same, the corresponding data in the intermediate output tensor and the depth prior information can be combined into a data pair, and the data pair can be combined into the spliced tensor; (3) in the case where the dimensions of the intermediate output tensor and the depth prior information are the same, the two can be directly spliced to obtain the spliced tensor.
[0116] S133: processing the spliced tensor using the post-processing sub-module to obtain a prediction result.
[0117] S134: training the pre-processing sub-module and the post-processing sub-module according to the prediction result and the image label.
[0118] The present disclosure also provides a target depth prediction method using the aforementioned target depth estimation model. Figure 4 is a flowchart of a target depth prediction method provided by some embodiments of the present disclosure. As shown in Figure 4 the target depth prediction method comprises S410-S420.
[0119] S410: obtaining a photographed image, and obtaining depth prior information corresponding to the photographed image, the depth prior information comprising prior depths of pixels in the sample image.
[0120] In the embodiments of the present disclosure, the computing device obtaining the photographed image can be obtaining the image photographed by the photographing camera in real time, or reading the stored photographed image from the memory.
[0121] The method of the computing device obtaining the depth prior information corresponding to the photographed image will be analyzed later.
[0122] S420: input the sample image and the depth prior information into the target depth prediction model to obtain the predicted depth of the target of interest in the photographed image.
[0123] The target depth prediction model used in step S420 is trained by the target depth prediction model training method of the foregoing embodiment. Since the target depth prediction model training method is used in the depth information prediction method provided in the present solution, and the input data includes the depth prior information corresponding to the photographed image, and the target depth prediction model has high accuracy, the predicted depth of the target of interest obtained by the present solution is more accurate.
[0124] In some embodiments of the present disclosure, the computing device performs S410 to obtain the depth prior information corresponding to the photographed image, which specifically can include S411-S412.
[0125] S411: perform a transmission transformation on the pixels in the photographed image to obtain the corresponding prior depth;
[0126] S412: combine the corresponding prior depth according to the coordinates of the pixels in the photographed image to obtain the depth prior information.
[0127] In some embodiments of the present disclosure, before performing the foregoing S411, the computing device can further perform S413-S414.
[0128] S413: obtain the internal and external parameters of the photographing camera and the internal and external parameters of the virtual camera, the virtual camera being a virtual camera that is virtually set to photograph a road image at a vertical downward angle;
[0129] S414: construct a homography matrix from the imaging plane of the photographing camera to the road plane according to the internal and external parameters of the photographing camera and the internal and external parameters of the virtual camera.
[0130] In the case of performing S413-S414, the foregoing S411 specifically includes: performing a transmission transformation on the pixels of the photographed image based on the homography matrix to obtain the prior depth corresponding to each pixel.
[0131] In some embodiments, S414 includes S414A-S414B.
[0132] S414A: construct a depth calculation formula according to the parameters in the homography matrix.
[0133] S414B: obtain the prior depth corresponding to each pixel according to the pixel coordinates of each pixel of the sample image and the depth calculation formula.
[0134] Wherein S414A includes SA1-SA2.
[0135] SA1: determining the split longitudinal coordinate according to the parameters in the homography matrix.
[0136] SA2: constructing a first depth calculation formula for the pixels in the photographed image with a longitudinal coordinate less than the split longitudinal coordinate, and constructing a second depth calculation formula for the pixels in the photographed image with a longitudinal coordinate greater than the first longitudinal coordinate, wherein the first depth calculation formula is obtained based on the split longitudinal coordinate by multiple Taylor expansions.
[0137] In the case of performing the aforementioned SA1-SA2, S414B specifically comprises: for a first pixel in the photographed image with a longitudinal coordinate less than or equal to the split longitudinal coordinate, calculating the prior depth by using the longitudinal coordinate of the first pixel and the first calculation formula; and for a second pixel in the photographed image with a longitudinal coordinate greater than the split longitudinal coordinate, calculating the prior depth by using the longitudinal coordinate of the second pixel and the second calculation formula.
[0138] The specific implementation process of S410-S420 of this embodiment is the same as that of the foregoing S210-S240, except that the processed image is modified from the sample image to the photographed image. The specific implementation process of S410-S420 will not be described here, and the specific content can be referred to the foregoing description.
[0139] Figure 5 is a flowchart of a process of obtaining depth prior information corresponding to a photographed image provided by some embodiments of the present disclosure. As shown in Figure 5 the method of obtaining depth prior information corresponding to a photographed image comprises S510-S550.
[0140] S510: performing pre-recognition processing on the photographed image to determine the detection frame of the target of interest in the photographed image and the type of the target of interest.
[0141] S520: determining the typical width of the target of interest according to the type of the target of interest.
[0142] S530: calculating the estimated depth of the target of interest to the photographed camera according to the typical width and the pixel width of the detection frame.
[0143] S540: taking the estimated depth as the prior depth of the pixels in the detection frame.
[0144] S550: combining the corresponding prior depth according to the coordinates of the pixels in the photographed image to obtain the depth prior information.
[0145] The specific implementation process of S510-S550 of this embodiment is the same as that of the foregoing S310-S350, except that the processed image is modified from the sample image to the photographed image. The specific implementation process of S510-S550 will not be described here, and the specific content can be referred to the foregoing description.
[0146] In some embodiments of the present disclosure, before performing S412 and S550, the computing device can further perform a step of encoding the prior depth by using a preset encoding function to obtain the prior depth encoding corresponding to each pixel.
[0147] In the case of performing the foregoing steps, the combining of the corresponding prior depth according to the coordinates of the pixels in the photographed image in S412 and S550 can specifically be: combining the corresponding prior depth encoding according to the coordinates of the pixels in the sample image.
[0148] In some embodiments of the present disclosure, the target depth prediction model includes a pre-processing sub-model and a post-processing sub-model; correspondingly, the computing device can include S421-S423 when performing S420.
[0149] S421: processing the sample image by using the pre-processing sub-model to obtain an intermediate output tensor;
[0150] S422: splicing the intermediate output tensor and the depth prior information to obtain a spliced tensor;
[0151] S423: processing the spliced tensor by using the post-processing sub-model to obtain the predicted depth.
[0152] The specific execution process of S421-S423 is the same as that of S131-S133, which will not be repeated here, and please refer to the foregoing description.
[0153] In addition to providing the foregoing target depth prediction model training method, the embodiments of the present disclosure also provide a target depth estimation model training apparatus 600. Figure 6 is a structural schematic diagram of the target depth estimation model training apparatus 600 provided by some embodiments of the present disclosure. As shown in Figure 6 The target depth estimation model training apparatus 600 includes a training sample acquisition unit 601, a depth prior information acquisition unit 602, and a model training unit 603.
[0154] The training sample acquisition unit 601 is configured to acquire a plurality of groups of training samples, each group of training samples including a sample image and an image label, and the image label including the labeled depth of the target of interest in the sample image. The depth prior information acquisition unit 602 is configured to acquire the depth prior information corresponding to the sample image, and the depth prior information including the prior depth of the pixel in the sample image. The model training unit 603 is configured to train the target depth prediction model by using the sample image, the depth prior information, and the image label.
[0155] In some embodiments, the depth prior information acquisition unit 602 comprises a transmission transformation subunit and a prior depth combination subunit. The transmission transformation subunit is configured to perform transmission transformation on pixels in the sample image to obtain prior depth corresponding to the pixels; and the prior depth combination subunit is configured to combine the prior depth corresponding to the pixels according to the coordinates of the pixels in the sample image to obtain the depth prior information.
[0156] In some embodiments of the present disclosure, the target depth estimation model training apparatus 600 further comprises a parameter acquisition unit and a homography matrix determination unit. The parameter acquisition unit is configured to acquire the intrinsic and extrinsic parameters of the shooting camera and the intrinsic and extrinsic parameters of the virtual camera, the virtual camera being a virtual camera set virtually to shoot the road image at a vertical downward angle; and the homography matrix determination unit is configured to construct a homography matrix of the imaging plane of the shooting camera to the road plane according to the intrinsic and extrinsic parameters of the shooting camera and the intrinsic and extrinsic parameters of the virtual camera; and the transmission transformation subunit performs transmission transformation on the pixels of the sample image based on the homography matrix to obtain the prior depth corresponding to the pixels.
[0157] In some embodiments of the present disclosure, the transmission transformation subunit comprises a formula construction module and a prior depth calculation module. The formula construction module is configured to construct a depth calculation formula according to the parameters in the homography matrix; and the prior depth calculation module is configured to obtain the prior depth corresponding to the pixels according to the pixel coordinates of each pixel in the sample image and the depth calculation formula.
[0158] In some embodiments of the present disclosure, the formula construction module determines a split longitudinal coordinate according to the parameters in the homography matrix; and constructs a first depth calculation formula for the pixels in the sample image whose longitudinal coordinates are less than the split longitudinal coordinate, and constructs a second depth calculation formula for the pixels in the sample image whose longitudinal coordinates are greater than the first longitudinal coordinate, wherein the first depth calculation formula is obtained based on multiple Taylor expansions of the split longitudinal coordinate; the prior depth calculation module calculates the prior depth of the first pixels in the sample image whose longitudinal coordinates are less than or equal to the split longitudinal coordinate by using the longitudinal coordinates of the first pixels and the first calculation formula; and calculates the prior depth of the second pixels in the sample image whose longitudinal coordinates are greater than the split longitudinal coordinate by using the longitudinal coordinates of the second pixels and the second calculation formula.
[0159] In some embodiments of this disclosure, the depth prior information acquisition unit 602 includes a preprocessing subunit, a typical width determination subunit, an estimated depth determination subunit, a prior depth determination subunit, and a prior depth combination subunit. The preprocessing subunit performs pre-identification processing on the sample image to determine the detection box of the target of interest and the type of the target of interest in the sample image; the typical width determination subunit determines the typical width of the target of interest based on its type; the estimated depth determination subunit calculates the estimated depth from the target of interest to the camera based on the typical width and the pixel width of the detection box; the prior depth determination subunit uses the estimated depth as the prior depth of each pixel within the detection box; and the prior depth combination subunit combines the corresponding prior depths according to the coordinates of the pixels in the sample image to obtain depth prior information.
[0160] In some embodiments of this disclosure, the target depth estimation model training device 600 further includes an encoding unit. The encoding unit is used to encode the prior depth using a pre-defined encoding function to obtain the prior depth code corresponding to the pixel; the prior depth combination subunit combines the corresponding prior depth codes according to the coordinates of the pixels in the sample image.
[0161] In some embodiments of this disclosure, the target depth prediction model includes a pre-processing sub-model and a post-processing sub-model; the model training unit 603 includes a pre-computation sub-unit, a stitching sub-unit, a post-computation sub-unit, and a training sub-unit. The pre-computation sub-unit is used to process the sample image using the pre-processing sub-model to obtain an intermediate output tensor; the stitching sub-unit is used to stitch the intermediate output tensor and depth prior information to obtain a stitched tensor; the post-computation sub-unit is used to process the stitched tensor using the post-processing sub-model to obtain a prediction result; and the training sub-unit is used to train the pre-processing sub-model and the post-processing sub-model based on the prediction result and image labels.
[0162] This disclosure also provides a target depth prediction device. Figure 7 This is a schematic diagram of the target depth prediction device 700 provided in some embodiments of this disclosure. For example... Figure 7 As shown, the target depth prediction device 700 includes an input data acquisition unit 701 and a prediction unit 702. The input data acquisition unit 701 is used to acquire captured images and acquire depth prior information corresponding to the captured images. The depth prior information includes the prior depth of pixels in the sample image. The prediction unit 702 is used to input the sample image and the depth prior information into the target depth prediction model to obtain the predicted depth of the target of interest in the captured image. The target depth prediction model is trained using the target depth prediction model training device described above.
[0163] In some embodiments of the present disclosure, the input data acquisition unit 701 comprises a transmission transformation subunit and a prior depth combination subunit. The transmission transformation subunit is configured to perform transmission transformation on pixels in the photographed image to obtain corresponding prior depth; and the prior depth combination subunit is configured to combine the corresponding prior depth according to the coordinates of the pixels in the photographed image to obtain the depth prior information.
[0164] In some embodiments of the present disclosure, the target depth prediction device 700 further comprises a parameter acquisition unit and a homography matrix determination unit. The parameter acquisition unit is configured to acquire the intrinsic and extrinsic parameters of the photographed camera and the intrinsic and extrinsic parameters of a virtual camera, the virtual camera being a virtual camera set virtually to photograph a road image at a vertical downward angle; and the homography matrix determination unit is configured to construct a homography matrix of an imaging plane of the photographed camera to a road plane according to the intrinsic and extrinsic parameters of the photographed camera and the intrinsic and extrinsic parameters of the virtual camera; and the transmission transformation subunit performs transmission transformation on the pixels of the photographed image based on the homography matrix to obtain the corresponding prior depth of each pixel.
[0165] In some embodiments of the present disclosure, the transmission transformation subunit comprises a formula construction module and a prior depth calculation module. The formula construction module is configured to construct a depth calculation formula according to the parameters in the homography matrix; and the prior depth calculation module is configured to obtain the corresponding prior depth of each pixel according to the pixel coordinates of each pixel of the sample image and the depth calculation formula.
[0166] In some embodiments of the present disclosure, the formula construction module determines a split longitudinal coordinate according to the parameters in the homography matrix; and constructs a first depth calculation formula for the pixels in the photographed image whose longitudinal coordinates are less than the split longitudinal coordinate, and constructs a second depth calculation formula for the pixels in the photographed image whose longitudinal coordinates are greater than the first longitudinal coordinate, wherein the first depth calculation formula is obtained based on multiple Taylor expansions of the split longitudinal coordinate; and the prior depth calculation module calculates the prior depth of the first pixel in the photographed image whose longitudinal coordinate is less than or equal to the split longitudinal coordinate by using the longitudinal coordinate of the first pixel and the first calculation formula; and,
[0167] The prior depth calculation module calculates the prior depth of the second pixel in the photographed image whose longitudinal coordinate is greater than the split longitudinal coordinate by using the longitudinal coordinate of the second pixel and the second calculation formula.
[0168] In some embodiments of the present disclosure, the depth prior information acquisition unit comprises a preprocessing subunit, a typical width determination subunit, an estimated depth determination subunit, a prior depth determination subunit, and a prior depth combination subunit.
[0169] The preprocessing subunit is configured to perform pre-recognition processing on the photographed image, determine a detection box of the target of interest in the photographed image and a type of the target of interest, the typical width determination subunit is configured to determine a typical width of the target of interest according to the type of the target of interest, the estimated depth determination subunit is configured to calculate an estimated depth of the target of interest to the photographing camera according to the typical width and a pixel width of the detection box, the prior depth determination subunit is configured to take the estimated depth as a prior depth of a pixel in the detection box, and the prior depth combination subunit is configured to combine corresponding prior depths according to coordinates of the pixels in the photographed image to obtain the depth prior information.
[0170] In some embodiments of the present disclosure, the target depth prediction device 700 further comprises an encoding unit and a prior depth combination subunit. The encoding unit is configured to encode the prior depth by using a preset encoding function to obtain a prior depth encoding corresponding to each pixel, and the prior depth combination subunit is configured to combine corresponding prior depth encodings according to coordinates of the pixels in the sample image.
[0171] In some embodiments of the present disclosure, the target of interest depth prediction model comprises a pre-processing sub-model and a post-processing sub-model, and the prediction unit 702 comprises a pre-computation subunit, a splicing subunit and a post-computation subunit. The pre-computation subunit is configured to process the sample image by using the pre-processing sub-model to obtain an intermediate output tensor, the splicing subunit is configured to splice the intermediate output tensor and the depth prior information to obtain a spliced tensor, and the post-computation subunit is configured to process the spliced tensor by using the post-processing sub-model to obtain the predicted depth.
[0172] 1. A target depth estimation model training method, comprising:
[0173] obtaining a plurality of groups of training samples, each group of training samples comprising a sample image and an image label, the image label comprising a labeled depth of a target of interest in the sample image;
[0174] obtaining depth prior information corresponding to the sample image, the depth prior information comprising a prior depth of a pixel in the sample image;
[0175] training a target depth prediction model by using the sample image, the depth prior information and the image label.
[0176] 2. The method of 1, wherein the obtaining of the depth prior information corresponding to the sample image comprises:
[0177] performing a transmission transformation on the pixel in the sample image to obtain a prior depth corresponding to the pixel;
[0178] combining corresponding prior depths according to coordinates of the pixels in the sample image to obtain the depth prior information.
[0179] 3. The method of 2, before performing the perspective transformation on the pixels of the sample image, the method further comprising:
[0180] obtaining intrinsic and extrinsic parameters of a shooting camera and intrinsic and extrinsic parameters of a virtual camera, the virtual camera being a virtual camera set virtually to shoot a road image at a vertical downward angle;
[0181] constructing a homography matrix from an imaging plane of the shooting camera to a road plane according to the intrinsic and extrinsic parameters of the shooting camera and the intrinsic and extrinsic parameters of the virtual camera;
[0182] the perspective transformation on each pixel in the sample image comprises: performing the perspective transformation on the pixels of the sample image based on the homography matrix to obtain the prior depth corresponding to the pixels.
[0183] 4. The method of 3, the perspective transformation on the pixels of the sample image based on the homography matrix to obtain the prior depth corresponding to the pixels comprises:
[0184] constructing a depth calculation formula according to parameters in the homography matrix;
[0185] obtaining the prior depth corresponding to the pixels according to the pixel coordinates of each pixel in the sample image and the depth calculation formula.
[0186] 5. The method of 4, constructing a depth calculation formula according to parameters in the homography matrix comprises:
[0187] determining a split longitudinal coordinate according to parameters in the homography matrix;
[0188] constructing a first depth calculation formula for pixels in the sample image with a longitudinal coordinate less than the split longitudinal coordinate, and constructing a second depth calculation formula for pixels in the sample image with a longitudinal coordinate greater than the first longitudinal coordinate, wherein the first depth calculation formula is a formula obtained by multiple Taylor expansions based on the split longitudinal coordinate;
[0189] the obtaining the prior depth corresponding to each pixel according to the pixel coordinates of each pixel in the sample image and the depth calculation formula comprises:
[0190] for a first pixel in the sample image with a longitudinal coordinate less than or equal to the split longitudinal coordinate, calculating the prior depth using the longitudinal coordinate of the first pixel and the first calculation formula; and,
[0191] for a second pixel in the sample image with a longitudinal coordinate greater than the split longitudinal coordinate, calculating the prior depth using the longitudinal coordinate of the second pixel and the second calculation formula.
[0192] 6. The method of claim 1, wherein the obtaining the depth prior information corresponding to the sample image comprises:
[0193] performing pre-recognition processing on the sample image to determine a bounding box of a target of interest in the sample image and a type of the target of interest;
[0194] determining a typical width of the target of interest according to the type of the target of interest;
[0195] calculating an estimated depth of the target of interest to the camera according to the typical width and a pixel width of the bounding box;
[0196] taking the estimated depth as a prior depth of each pixel in the bounding box;
[0197] combining the prior depths corresponding to the pixels in the sample image according to coordinates of the pixels to obtain the depth prior information.
[0198] 7. The method of any one of claims 2-6, wherein before the combining the prior depths corresponding to the pixels in the sample image according to coordinates of the pixels, the method further comprises:
[0199] encoding the prior depths using a pre-set encoding function to obtain prior depth encodings corresponding to the pixels;
[0200] the combining the prior depths corresponding to the pixels in the sample image according to coordinates of the pixels comprises:
[0201] combining the prior depth encodings corresponding to the pixels in the sample image according to coordinates of the pixels.
[0202] 8. The method of claim 1, wherein the depth prediction model of the target of interest comprises a pre-processing sub-model and a post-processing sub-model;
[0203] training the depth prediction model of the target of interest using a sample image, the depth prior information, and an image label comprises:
[0204] processing the sample image using the pre-processing sub-model to obtain an intermediate output tensor;
[0205] concatenating the intermediate output tensor and the depth prior information to obtain a concatenated tensor;
[0206] processing the concatenated tensor using the post-processing sub-model to obtain a prediction result;
[0207] training the pre-processing sub-model and the post-processing sub-model according to the prediction result and the image label.
[0208] 9. A method for predicting a depth of a target, comprising:
[0209] obtaining a photographed image, and obtaining depth prior information corresponding to the photographed image, the depth prior information including prior depths of pixels in the sample image;
[0210] inputting the sample image and the depth prior information into a target depth prediction model to obtain predicted depths of a target of interest in the photographed image;
[0211] The target depth prediction model is trained by the target depth prediction model training method in any one of 1-8.
[0212] 10. The method of 9, wherein the obtaining the depth prior information corresponding to the photographed image comprises:
[0213] performing a transmission transformation on the pixels in the photographed image to obtain corresponding prior depths;
[0214] combining the corresponding prior depths according to the coordinates of the pixels in the photographed image to obtain the depth prior information.
[0215] 11. The method of 10, wherein before the performing the transmission transformation on the pixels in the photographed image, the method further comprises:
[0216] obtaining internal and external parameters of a photographing camera and internal and external parameters of a virtual camera, the virtual camera being a virtual camera that is virtually set to photograph a road image at a vertical downward angle;
[0217] constructing a homography matrix from an imaging plane of the photographing camera to a road plane according to the internal and external parameters of the photographing camera and the internal and external parameters of the virtual camera;
[0218] The performing the transmission transformation on the pixels in the photographed image comprises performing a transmission transformation on the pixels in the photographed image based on the homography matrix to obtain the prior depths corresponding to the pixels.
[0219] 12. The method of 11, wherein the performing the transmission transformation on the pixels in the photographed image based on the homography matrix to obtain the prior depths corresponding to the pixels comprises:
[0220] constructing a depth calculation formula according to parameters in the homography matrix;
[0221] obtaining the prior depths corresponding to the pixels according to the pixel coordinates of the pixels in the sample image and the depth calculation formula.
[0222] 13. The method of 12, wherein the constructing the depth calculation formula according to the parameters in the homography matrix comprises:
[0223] determining a split longitudinal coordinate according to parameters in the homography matrix;
[0224] constructing a first depth calculation formula for pixels in the photographed image with longitudinal coordinates less than the split longitudinal coordinate, and constructing a second depth calculation formula for pixels in the photographed image with longitudinal coordinates greater than the first longitudinal coordinate, wherein the first depth calculation formula is obtained based on split longitudinal coordinate and multiple Taylor expansions;
[0225] the obtaining of the prior depth corresponding to each pixel according to the pixel coordinates of each pixel in the photographed image and the depth calculation formula includes:
[0226] for the first pixel in the photographed image with a longitudinal coordinate less than or equal to the split longitudinal coordinate, calculating the prior depth using the longitudinal coordinate of the first pixel and the first calculation formula; and,
[0227] for the second pixel in the photographed image with a longitudinal coordinate greater than the split longitudinal coordinate, calculating the prior depth using the longitudinal coordinate of the second pixel and the second calculation formula.
[0228] 14. The method of 9, the obtaining of the depth prior information corresponding to the photographed image includes:
[0229] performing pre-recognition processing on the photographed image to determine a detection frame of a target of interest in the photographed image and a type of the target of interest;
[0230] determining a typical width of the target of interest according to the type of the target of interest;
[0231] calculating an estimated depth of the target of interest to the photographed camera according to the typical width and a pixel width of the detection frame;
[0232] taking the estimated depth as the prior depth of pixels in the detection frame;
[0233] combining the corresponding prior depths according to the coordinates of the pixels in the photographed image to obtain the depth prior information.
[0234] 15. The method of any one of 10-14, before the combining of the corresponding prior depths according to the coordinates of the pixels in the photographed image, the method further includes:
[0235] encoding the prior depths using a pre-set encoding function to obtain prior depth encodings corresponding to the pixels;
[0236] the combining of the corresponding prior depths according to the coordinates of the pixels in the photographed image includes:
[0237] The corresponding prior depth coding is combined according to the coordinates of the pixels in the sample image.
[0238] 16. The method of 8, wherein the target depth prediction model comprises a pre-processing sub-model and a post-processing sub-model.
[0239] The sample image and the depth prior information are input into a target depth prediction model to obtain a predicted depth of a target of interest in the photographed image.
[0240] The sample image is processed using the pre-processing sub-model to obtain an intermediate output tensor.
[0241] The intermediate output tensor and the depth prior information are spliced to obtain a spliced tensor.
[0242] The spliced tensor is processed using the post-processing sub-model to obtain a predicted depth.
[0243] 17. A device for training a target depth estimation model,
[0244] a training sample acquisition unit configured to acquire a plurality of groups of training samples, each group of training samples comprising a sample image and an image label, the image label comprising a labeled depth of a target of interest in the sample image;
[0245] a depth prior information acquisition unit configured to acquire depth prior information corresponding to the sample image, the depth prior information comprising a prior depth of a pixel in the sample image;
[0246] a model training unit configured to train a target depth prediction model using the sample image, the depth prior information, and the image label.
[0247] 18. The device of 17, wherein the depth prior information acquisition unit comprises:
[0248] a transmission transformation sub-unit configured to perform transmission transformation on a pixel in the sample image to obtain a prior depth corresponding to the pixel;
[0249] a prior depth combination sub-unit configured to combine the corresponding prior depths according to the coordinates of the pixels in the sample image to obtain the depth prior information.
[0250] 19. The device of 18, further comprising:
[0251] a parameter acquisition unit configured to acquire internal and external parameters of a photographing camera and internal and external parameters of a virtual camera, the virtual camera being a virtual camera that is virtually set to photograph a road image at a vertically downward angle;
[0252] A homography matrix determination unit is configured to construct a homography matrix of an imaging plane of the shooting camera to a road plane according to intrinsic and extrinsic parameters of the shooting camera and intrinsic and extrinsic parameters of the virtual camera.
[0253] The transmission transformation sub-unit performs transmission transformation on pixels of the sample image based on the homography matrix to obtain an a priori depth corresponding to the pixels.
[0254] 20. The apparatus according to 19, wherein the transmission transformation sub-unit comprises:
[0255] A formula construction module is configured to construct a depth calculation formula according to parameters in the homography matrix.
[0256] An a priori depth calculation module is configured to obtain an a priori depth corresponding to a pixel according to a pixel coordinate of the pixel in the sample image and the depth calculation formula.
[0257] 21. The apparatus according to 20,
[0258] The formula construction module determines a split longitudinal coordinate according to parameters in the homography matrix, and constructs a first depth calculation formula for pixels in the sample image with a longitudinal coordinate less than the split longitudinal coordinate and a second depth calculation formula for pixels in the sample image with a longitudinal coordinate greater than the first longitudinal coordinate, wherein the first depth calculation formula is obtained based on split longitudinal coordinate and multiple Taylor expansions.
[0259] The a priori depth calculation module calculates the a priori depth for a first pixel in the sample image with a longitudinal coordinate less than or equal to the split longitudinal coordinate using the longitudinal coordinate of the first pixel and the first calculation formula, and calculates the a priori depth for a second pixel in the sample image with a longitudinal coordinate greater than the split longitudinal coordinate using the longitudinal coordinate of the second pixel and the second calculation formula.
[0260] 22. The apparatus according to 17, wherein the depth a priori information acquisition unit comprises:
[0261] A preprocessing sub-unit is configured to perform pre-recognition processing on the sample image to determine a detection frame of a target of interest in the sample image and a type of the target of interest.
[0262] A typical width determination sub-unit is configured to determine a typical width of the target of interest according to the type of the target of interest.
[0263] An estimated depth determination sub-unit is configured to calculate an estimated depth of the target of interest to the shooting camera according to the typical width and a pixel width of the detection frame.
[0264] a priori depth determination subunit, configured to take the estimated depth as a priori depth of each pixel in the bounding box;
[0265] a priori depth combination subunit, configured to combine the corresponding a priori depths according to the coordinates of the pixels in the sample image to obtain the depth prior information.
[0266] 23. The apparatus according to any one of claims 18-22, further comprising:
[0267] an encoding unit, configured to encode the a priori depth by using a preset encoding function to obtain a priori depth encoding corresponding to the pixel;
[0268] the a priori depth combination subunit combines the corresponding a priori depth encoding according to the coordinates of the pixels in the sample image.
[0269] 24. The apparatus according to claim 17, wherein the target depth prediction model comprises a pre-processing sub-model and a post-processing sub-model; and the model training unit comprises:
[0270] a pre-computation subunit, configured to process the sample image by using the pre-processing sub-model to obtain an intermediate output tensor;
[0271] a concatenation subunit, configured to concatenate the intermediate output tensor and the depth prior information to obtain a concatenated tensor;
[0272] a post-computation subunit, configured to process the concatenated tensor by using the post-processing sub-model to obtain a prediction result;
[0273] a training subunit, configured to train the pre-processing sub-model and the post-processing sub-model according to the prediction result and the image label.
[0274] 25. A target depth prediction apparatus, comprising:
[0275] an input data acquisition unit, configured to acquire a photographed image and acquire depth prior information corresponding to the photographed image, the depth prior information comprising a priori depth of a pixel in the sample image;
[0276] a prediction unit, configured to input the sample image and the depth prior information into a target depth prediction model to obtain a predicted depth of a target of interest in the photographed image;
[0277] wherein the target depth prediction model is trained by using the target depth prediction model training apparatus according to any one of claims 1-8.
[0278] 26. The apparatus according to claim 25, wherein the input data acquisition unit comprises:
[0279] a transmission transformation subunit, configured to perform transmission transformation on pixels in the photographed image to obtain corresponding prior depth;
[0280] a prior depth combination subunit, configured to combine the corresponding prior depth according to the coordinates of the pixels in the photographed image to obtain the depth prior information.
[0281] 27. The apparatus according to 26, further comprising:
[0282] a parameter acquisition unit, configured to acquire the intrinsic and extrinsic parameters of the photographed camera and the intrinsic and extrinsic parameters of the virtual camera, the virtual camera being a virtual camera set virtually to photograph a road image at a vertical downward angle;
[0283] a homography matrix determination unit, configured to construct a homography matrix of an imaging plane of the photographed camera to a road plane according to the intrinsic and extrinsic parameters of the photographed camera and the intrinsic and extrinsic parameters of the virtual camera;
[0284] the transmission transformation subunit performs transmission transformation on the pixels of the photographed image based on the homography matrix to obtain the corresponding prior depth of each pixel.
[0285] 28. The apparatus according to 27, the transmission transformation subunit comprising:
[0286] a formula construction module, configured to construct a depth calculation formula according to parameters in the homography matrix;
[0287] a prior depth calculation module, configured to obtain the corresponding prior depth of each pixel according to the pixel coordinates of each pixel of the sample image and the depth calculation formula.
[0288] 29. The apparatus according to 28,
[0289] the formula construction module determines a split longitudinal coordinate according to parameters in the homography matrix, and constructs a first depth calculation formula for pixels in the photographed image with longitudinal coordinates less than the split longitudinal coordinate and a second depth calculation formula for pixels in the photographed image with longitudinal coordinates greater than the first longitudinal coordinate, wherein the first depth calculation formula is a formula obtained by multiple Taylor expansions based on the split longitudinal coordinate;
[0290] the prior depth calculation module calculates the prior depth of a first pixel in the photographed image with a longitudinal coordinate less than or equal to the split longitudinal coordinate by using the longitudinal coordinate of the first pixel and the first calculation formula; and
[0291] calculates the prior depth of a second pixel in the photographed image with a longitudinal coordinate greater than the split longitudinal coordinate by using the longitudinal coordinate of the second pixel and the second calculation formula.
[0292] 30. The apparatus of claim 25, wherein the depth prior information obtaining unit comprises:
[0293] a preprocessing subunit configured to perform pre-recognition processing on the photographed image to determine a bounding box of a target of interest in the photographed image and a type of the target of interest;
[0294] a typical width determining subunit configured to determine a typical width of the target of interest according to the type of the target of interest;
[0295] an estimated depth determining subunit configured to calculate an estimated depth of the target of interest to the photographing camera according to the typical width and a pixel width of the bounding box;
[0296] a prior depth determining subunit configured to take the estimated depth as a prior depth of a pixel in the bounding box;
[0297] a prior depth combining subunit configured to combine corresponding prior depths according to coordinates of pixels in the photographed image to obtain the depth prior information.
[0298] 31. The apparatus of any one of claims 26-30, further comprising:
[0299] an encoding unit configured to encode the prior depth using a pre-set encoding function to obtain a prior depth encoding corresponding to each pixel;
[0300] the prior depth combining subunit combines corresponding prior depth encodings according to coordinates of pixels in the sample image.
[0301] 32. The apparatus of claim 25, wherein the target of interest depth prediction model comprises a pre-processing sub-model and a post-processing sub-model, and the prediction unit comprises:
[0302] a pre-computing subunit configured to process the sample image using the pre-processing sub-model to obtain an intermediate output tensor;
[0303] a concatenating subunit configured to concatenate the intermediate output tensor and the depth prior information to obtain a concatenated tensor;
[0304] a post-computing subunit configured to process the concatenated tensor using the post-processing sub-model to obtain a predicted depth.
[0305] Figure 8 is a structural schematic diagram of a computing device provided by some embodiments of the present disclosure. Specific reference is made below to Figure 8 which shows a structural schematic diagram of a computing device 800 suitable for use to implement the computing device in embodiments of the present disclosure. Figure 8The computing device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0306] like Figure 8 As shown, the computing device 800 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory ROM 802 or a program loaded from a storage device 808 into a random access memory RAM 803. The RAM 803 also stores various programs and data required for the operation of the computing device 800. The processing unit 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0307] Typically, the following devices can be connected to I / O interface 805: input devices 805 including, for example, touchscreens, touchpads, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 807 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 808 including, for example, magnetic tapes, hard disks, etc.; and communication devices 809. Communication device 809 allows computing device 800 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 A computing device 800 with various devices is shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.
[0308] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 809, or installed from a storage device 808, or installed from a ROM 802. When the computer program is executed by a processing device 801, it performs the functions defined in the methods of embodiments of this disclosure.
[0309] It is noted that the aforementioned computer-readable medium of the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium can be, for example and without limitation, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a computer-readable program code transmitted by a computer-readable medium or a carrier wave transmits, propagates, or transfers a program used by or in connection with an instruction execution system, apparatus, or device. The program code contained in the computer-readable medium can be transmitted using any suitable medium, including but not limited to wire, cable, optical fiber, RF (radio frequency), or the like, or any suitable combination of the foregoing.
[0310] In some embodiments, the client, server, or both can communicate using any current known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet, and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any current known or future developed networks.
[0311] The aforementioned computer-readable medium can be contained in the aforementioned computing device; or can exist separately from the computing device without being incorporated into the computing device.
[0312] The computer readable medium carries one or more programs, when the one or more programs are executed by the computing device, cause the computing device to: obtain a plurality of sets of training samples, each set of training samples including a sample image and an image label, the image label including a labeled depth of a target object in the sample image; obtain depth prior information corresponding to the sample image, the depth prior information including a prior depth of a pixel in the sample image; and train a target depth prediction model using the sample image, the depth prior information, and the image label.
[0313] Computer program code for carrying out operations of the present disclosure can be written in any one or more programming languages or combinations of languages including object oriented programming languages such as Java, Smalltalk, C++ or conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0314] The flow diagrams and the block diagrams in the drawings are illustrations of architectures, functionalities, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a procedure, or a part of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur in a different order than that noted in the figures. For example, two blocks noted in succession can in fact be executed substantially concurrently or in the opposite order, depending on the functionality involved. It is also noted that each block in the block diagrams and / or flow diagrams, and combinations of blocks in the block diagrams and / or flow diagrams, can be implemented by dedicated hardware-based systems that perform the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0315] The units described in the embodiments of the present disclosure can be implemented by software, or by hardware. In some cases, the names of the units do not constitute a limitation on the units themselves.
[0316] The functionality described herein above can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0317] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0318] The present disclosure also provides a computer readable storage medium, wherein the storage medium stores a computer program. When the computer program is executed by a processor, the method of any one of the above method embodiments can be implemented, and the execution manner and beneficial effects are similar, which will not be described here again.
[0319] The present disclosure also provides a vehicle, which comprises the above computing device. The specific vehicle can be a fuel vehicle, or a pure electric vehicle, etc., and the present disclosure does not make any limitation.
[0320] It should be noted that, in this document, relational terms such as "first" and "second", and the like, are used solely to distinguish one entity or action from another entity or action, without necessarily requiring or implying any actual such relationship or order between such entities or actions. Moreover, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without limitation, an element preceded by "comprises... a" does not, without more constraints, foreclose the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.
[0321] The foregoing is merely illustrative of the various implementations of the present disclosure and the general principles thereof. Numerous modifications can be made to these illustrations, and equivalents can be substituted therefor, without departing from the scope of the present disclosure. The specific embodiments commensurate with the specific application are intended to be illustrative only and not limiting of the scope of the application as set forth in the following claims.
Claims
1. A method for training a target depth estimation model, characterized in that, include: Multiple sets of training samples are obtained. Each set of training samples includes a sample image and an image label. The image label includes the annotation depth of the target of interest in the sample image. Obtain the depth prior information corresponding to the sample image, wherein the depth prior information includes the prior depth of the pixels in the sample image; The target depth prediction model is trained using sample images, the depth prior information, and the image labels; The step of obtaining the depth prior information corresponding to the sample image includes: Perform a transmission transformation on the pixels in the sample image to obtain the prior depth corresponding to the pixels; The prior depth information is obtained by combining the corresponding prior depths according to the coordinates of the pixels in the sample image. Before performing transmission transformation on the pixels of the sample image, the method further includes: Acquire the intrinsic and extrinsic parameters of the shooting camera and the intrinsic and extrinsic parameters of the virtual camera, wherein the virtual camera is a virtual camera set up to shoot road images at a vertically downward angle; Based on the intrinsic and extrinsic parameters of the shooting camera and the intrinsic and extrinsic parameters of the virtual camera, construct the homography matrix from the imaging plane of the shooting camera to the road plane; The transmission transformation of each pixel in the sample image includes: Based on the homography matrix, a transmission transformation is performed on the pixels of the sample image to obtain the prior depth corresponding to the pixel.
2. The method according to claim 1, characterized in that, The step of performing a transmission transformation on the pixels of the sample image based on the homography matrix to obtain the prior depth corresponding to the pixels includes: A depth calculation formula is constructed based on the parameters in the homography matrix; Based on the pixel coordinates of each pixel in the sample image and the depth calculation formula, the prior depth corresponding to the pixel is obtained.
3. The method according to claim 2, characterized in that, The depth calculation formula is constructed based on the parameters in the homography matrix, including: The segmentation ordinate is determined based on the parameters in the homography matrix; A first depth calculation formula is constructed for pixels in the sample image whose ordinate is less than or equal to the segmentation ordinate, and a second depth calculation formula is constructed for pixels in the sample image whose ordinate is greater than the segmentation ordinate, wherein the first depth calculation formula is a formula obtained by performing multiple Taylor expansions based on the segmentation ordinate. The step of obtaining the prior depth corresponding to each pixel based on the pixel coordinates of each pixel in the sample image and the depth calculation formula includes: For the first pixel in the sample image whose ordinate is less than or equal to the segmentation ordinate, the prior depth is calculated using the ordinate of the first pixel and the first depth calculation formula; and, For the second pixel in the sample image whose ordinate is greater than the segmentation ordinate, the prior depth is calculated using the ordinate of the second pixel and the second depth calculation formula.
4. A target depth prediction method, characterized in that, include: Acquire captured images and acquire depth prior information corresponding to the captured images, wherein the depth prior information includes the prior depth of pixels in the sample image; The sample image and the depth prior information are input into the target depth prediction model to obtain the predicted depth of the target of interest in the captured image; The step of obtaining the depth prior information corresponding to the captured image includes: Perform a transmission transformation on the pixels in the captured image to obtain the corresponding prior depth; The prior depth information is obtained by combining the coordinates of the pixels in the captured image with the corresponding prior depth information. Before performing transmission transformation on the pixels in the captured image, the method further includes: Acquire the intrinsic and extrinsic parameters of the shooting camera and the intrinsic and extrinsic parameters of the virtual camera, wherein the virtual camera is a virtual camera set up to shoot road images at a vertically downward angle; Based on the intrinsic and extrinsic parameters of the shooting camera and the intrinsic and extrinsic parameters of the virtual camera, construct the homography matrix from the imaging plane of the shooting camera to the road plane; The transmission transformation of the pixels in the captured image includes: performing transmission transformation on the pixels of the captured image based on the homography matrix to obtain the prior depth corresponding to each pixel; The target depth prediction model is trained using the target depth prediction model training method as described in any one of claims 1-3.
5. A target depth estimation model training device, characterized in that, The training sample acquisition unit is used to acquire multiple sets of training samples. Each set of training samples includes a sample image and an image label. The image label includes the annotation depth of the target of interest in the sample image. A depth prior information acquisition unit is used to acquire the depth prior information corresponding to the sample image, wherein the depth prior information includes the prior depth of pixels in the sample image; The model training unit is used to train the target depth prediction model using sample images, the depth prior information, and the image labels; The depth prior information acquisition unit includes: The transmission transformation subunit is used to perform transmission transformation on the pixels in the sample image to obtain the prior depth corresponding to the pixels. The prior depth combination subunit is used to combine the corresponding prior depths according to the coordinates of the pixels in the sample image to obtain the prior depth information. The device further includes: The parameter acquisition unit is used to acquire the intrinsic and extrinsic parameters of the shooting camera and the intrinsic and extrinsic parameters of the virtual camera. The virtual camera is a virtual camera that is set up to shoot road images at a vertically downward angle. The homography matrix determination unit is used to construct a homography matrix from the imaging plane of the camera to the road plane based on the intrinsic and extrinsic parameters of the camera and the virtual camera. The transmission transformation subunit is used to perform transmission transformation on the pixels of the sample image based on the homography matrix to obtain the prior depth corresponding to the pixel.
6. A target depth prediction device, characterized in that, include: An input data acquisition unit is used to acquire captured images and acquire depth prior information corresponding to the captured images, wherein the depth prior information includes the prior depth of pixels in the sample image; The prediction unit is used to input the sample image and the depth prior information into the target depth prediction model to obtain the predicted depth of the target of interest in the captured image; The input data acquisition unit includes: A transmission transformation subunit is used to perform transmission transformation on pixels in the captured image to obtain the corresponding prior depth. The prior depth combination subunit is used to combine the corresponding prior depths according to the coordinates of the pixels in the captured image to obtain the prior depth information; The device further includes: The parameter acquisition unit is used to acquire the intrinsic and extrinsic parameters of the shooting camera and the intrinsic and extrinsic parameters of the virtual camera. The virtual camera is a virtual camera that is set up to shoot road images at a vertically downward angle. The homography matrix determination unit is used to construct a homography matrix from the imaging plane of the camera to the road plane based on the intrinsic and extrinsic parameters of the camera and the virtual camera. The transmission transformation subunit is used to perform transmission transformation on the pixels of the captured image based on the homography matrix to obtain the prior depth corresponding to each pixel. The target depth prediction model is trained using the target depth prediction model training method as described in any one of claims 1-3.
7. A computing device, characterized in that, Includes a processor and a memory, the memory being used to store computer programs; When the computer program is loaded by the processor, it causes the processor to execute the target depth estimation model training method as described in any one of claims 1-3 or the target depth prediction method as described in claim 4.
8. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, causes the processor to implement the target depth estimation model training method as described in any one of claims 1-3 or the target depth prediction method as described in claim 4.
Citation Information
Patent Citations
Model training method and device
CN112241976A