Depth estimation method based on continuous time flow field
By using a continuous time flow field-based method in vehicle depth estimation, and using the characteristic change trend of traffic image information for depth prediction, the problem of low accuracy of depth estimation in the prior art is solved, and high-precision target object recognition in complex scenarios is achieved.
Patent Information
- Application Number
- CN202510204869.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-24
AI Technical Summary
The prior art is difficult to ensure accuracy in vehicle depth estimation, especially in complex application scenarios. Due to the assumption that the target is located on the ground, ground flatness and vehicle motion stability, it is difficult to achieve reliable depth estimation under diversified and complex operating conditions.
By adopting a depth estimation method based on a continuous time flow field, the traffic image information of the vehicle during driving is obtained, feature processing and feature space determination are performed, depth prediction is performed based on feature change trends at adjacent moments, and target objects are identified in combination with feature space at the current moment.
Through the analysis of feature change trends, the deep feature information of the target object can be accurately calculated, the accuracy of target object recognition can be improved, the accuracy of target object recognition can be adapted to complex application scenarios, and the robustness and generalization ability of the model can be enhanced.
Smart Images

Figure CN119992478A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image data processing, and in particular to a depth estimation method based on a continuous-time flow field in the technical field of image data processing. Background Art
[0002] Although trusted artificial intelligence (AI) technology is developing rapidly at this stage, its implementation faces many challenges, including conflicts of principles, large amounts of resources and time consumption, and complexity of working conditions (such as the diversity of application scenarios and differences in supervision and regulations).
[0003] In the related art, there are two methods for automobile depth estimation: traditional depth estimation and depth estimation using neural networks. However, in the implementation process of both traditional depth estimation and depth estimation using neural networks, the reliable target recognition conditions are relatively harsh, so it is difficult to ensure the accuracy of depth estimation. Summary of the invention
[0004] The purpose of the present invention is to provide a depth estimation method based on a continuous time flow field, and the technical solution adopted is as follows:
[0005] In a first aspect, an embodiment of the present invention provides a depth estimation method based on a continuous-time flow field, the method comprising:
[0006] Obtain traffic image information of vehicles while they are traveling;
[0007] Performing feature processing on the traffic image information to obtain a feature space;
[0008] Based on the feature space corresponding to two adjacent moments, the feature change trend is determined in the time series;
[0009] Performing depth prediction on the target object during the driving process based on the characteristic change trend to obtain depth characteristic information of the target object;
[0010] The target object is identified based on the depth feature information and the feature space at the current moment.
[0011] In a second aspect, a depth estimation system based on a continuous time flow field is provided, the system comprising:
[0012] An acquisition module is used to acquire traffic image information of a vehicle during driving;
[0013] A processing module, used for performing feature processing on the traffic image information to obtain a feature space;
[0014] A determination module, used to determine the feature change trend in the time series based on the feature space corresponding to two adjacent moments;
[0015] A prediction module, configured to perform depth prediction on a target object during the driving process based on the characteristic change trend, and obtain depth characteristic information of the target object;
[0016] The recognition module is used to recognize the target object based on the depth feature information and the feature space at the current moment.
[0017] According to a third aspect, a computer program product is provided, comprising: a computer program code, which, when executed on a computer, enables the computer to execute the method in the first aspect or any possible implementation of the first aspect.
[0018] In a fourth aspect, a computer-readable storage medium is provided, which stores a computer program code. When the computer program code runs on a computer, the computer executes the method in the first aspect or any possible implementation manner described in the first aspect.
[0019] The present invention has the following beneficial effects: after obtaining the traffic image information of the vehicle during driving, the feature space is obtained by performing feature processing on the traffic image information, and the feature change trend is determined in the time series based on the feature space corresponding to two adjacent moments; in this way, the change of the traffic image information in the time series can be accurately reflected through the feature change trend. Afterwards, the depth of the target object in the driving process is predicted based on the feature change trend to obtain the depth feature information of the target object; in this way, the displacement of the target object at different time points during driving can be characterized in combination with the feature change trend, so that the depth feature information of the target object can be accurately calculated; it is convenient to identify the target object through the depth feature information and the feature space at the current moment, so as to improve the accuracy of target object identification. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0021] Figure 1 It is a schematic diagram of the implementation principle of the vehicle depth estimation method in the related art;
[0022] Figure 2It is another schematic diagram of the implementation principle of the vehicle depth estimation method in the related art;
[0023] Figure 3 It is another schematic diagram of the implementation principle of the vehicle depth estimation method in the related art;
[0024] Figure 4 It is a schematic diagram of an implementation flow of a depth estimation method based on a continuous-time flow field provided by an embodiment of the present invention;
[0025] Figure 5 It is a schematic diagram of the implementation principle of a depth estimation method based on a continuous-time flow field provided by an embodiment of the present invention;
[0026] Figure 6 It is another schematic diagram of the implementation principle of a depth estimation method based on a continuous-time flow field provided by an embodiment of the present invention;
[0027] Figure 7 is another schematic diagram of the implementation principle of a depth estimation method based on a continuous-time flow field provided by an embodiment of the present invention;
[0028] Figure 8 Schematic diagram of an application scenario of a depth estimation method based on a continuous-time flow field provided by an embodiment of the present invention;
[0029] Fig. 9 1 is a schematic diagram of the composition structure of a depth estimation system based on a continuous-time flow field provided by an embodiment of the present invention;
[0030] Fig.10 It is a structural schematic diagram of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0031] In order to further explain the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following is a detailed description of the depth estimation method based on the continuous time flow field proposed by the present invention, its specific implementation method, structure, characteristics and effects, in combination with the accompanying drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" does not necessarily refer to the same embodiment. In addition, specific features, structures or characteristics in one or more embodiments may be combined in any suitable form as described.
[0032] Among them, in the description of the embodiments of the present invention, unless otherwise specified, " / " means or, for example, A / B can mean A or B: "and / or" in the text is only a way to describe the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of the present invention, "multiple" refers to two or more than two.
[0033] In the following, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as suggesting or implying relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features.
[0034] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0035] Trusted AI technology is developing rapidly, but the road to practical application is full of thorns. On the one hand, it is not easy to achieve multiple principles such as robustness and analyticity, and conflicts between principles frequently occur. Taking deep learning as an example, adversarial training aimed at improving robustness is very likely to sacrifice fairness. At the same time, the realization of trustworthy AI requires massive computing resources and a long time investment. On the one hand, reliable AI is like a big tree that relies on high-quality soil and is extremely hungry for high-quality data. Once the data has quality problems and is full of low-quality or biased information, the performance and accuracy of the model will be greatly reduced. In actual application, there are many loopholes and it may even output misleading results. Therefore, deep learning requires data collection, cleaning, labeling, and subsequent rigorous testing and verification, all of which are time-consuming and laborious. On the other hand, as the scale and complexity of deep learning models continue to rise, the computing resources required for training and inference are also rising, bringing cost pressure beyond imagination. Its high cost can make many companies shy away. Furthermore, the complexity of working conditions is like another difficulty ahead. The working conditions are complicated and changeable, and there are many factors such as equipment working status, environmental conditions, load changes, AI with different intelligence levels and various applications, endless combinations, running time, and interaction with other systems. It is undoubtedly a fantasy to collect all working conditions. It makes the analyzability of trusted AI even more difficult.
[0036] In the related art, the depth estimation of the vehicle can be done by Figure 1 The depth estimation method shown is implemented as follows: first obtain the camera installation height H (camera installation Height), focal length f, and height h of the target object in the image; where: Right now In this depth estimation method, it is necessary to ensure that the target is on the ground, the ground is flat, and the vehicle moves stably without pitching motion; thus, the depth estimation method has the following limitations: 1. The limitation of the target being on the ground. Ground condition restrictions: The material, hardness, humidity and other conditions of different ground surfaces may affect the accuracy of the detection results. For example, when detecting on soft muddy ground, the sinking of the wheels and the deformation of the ground may not be able to ensure that the target is on the ground; 2. The limitation of the flatness of the ground, it is difficult to ensure absolute flatness: In actual detection, it is difficult to find a completely flat ground with no slope. Even if the ground is leveled, there may be slight slopes or uneven areas, which will have a certain impact; 3. The limitation of the vehicle moving stably without pitching motion, it is difficult to completely eliminate the pitching motion: In actual detection, due to the weight distribution of the vehicle itself, the elasticity of the suspension system and the influence of external excitations (such as wind resistance, uneven road surface, etc.), it is difficult to completely eliminate the pitching motion of the vehicle.
[0037] In the related art, the depth estimation of the vehicle can be done by Figure 2 The depth estimation implementation of the neural network shown in the figure: Depth distribution refers to estimating the depth information of each image feature point in three-dimensional space and obtaining the probability distribution of these points at different depths. The LSS (Lift, Splat, Shoot) algorithm is a deep learning model, especially using convolutional neural networks (CNN), etc., which relies on large-scale data sets to learn the mapping relationship from input images to output depth maps or other related information (such as Figure 2 The mapping relationship architecture described above). Therefore, in order to train a high-performance LSS model, it is usually necessary to collect a large amount of labeled data and ensure that the data has sufficient diversity and accuracy. This is also the limitation of LSS; a large amount of labeled data is required to obtain the real-world depth distribution, and its reliability depends on the labeled data it sees.
[0038] The basic idea of human depth estimation, such as Figure 3 As shown: Right now Where B is the baseline distance, i.e. the horizontal distance between the two eyes. D: Depth is the vertical distance from the target object to the camera. F is the focal length; Δy is y c and e The difference between c It represents the distance between the intersection point of line 31 and line 32, and the intersection point of line 33 and line 32; it represents the distance between the intersection point of line 34 and line 32, and the intersection point of line 35 and line 32.
[0039] The embodiments of the present invention use different displacement times to predict depth. In mobile scenarios, such as autonomous vehicles or robot navigation, the depth information of the target object will change dynamically as the observation position changes. By recording and analyzing the displacement of the target object at different time points, we can use this displacement information to predict its depth. This method can more accurately estimate the depth of the target object by comparing images or depth data at different time points, thereby improving the car's resolvable target recognition.
[0040] The following is a detailed description of a method for depth estimation based on a continuous time flow field provided by the present invention in conjunction with the accompanying drawings. Figure 4 , which shows a schematic diagram of an implementation flow of a depth estimation method based on a continuous-time flow field provided by an embodiment of the present invention, the method comprising:
[0041] 401, obtaining traffic image information of a vehicle during driving.
[0042] Here, the traffic image information may be a continuous video stream collected by a vehicle-mounted camera during the driving of the vehicle.
[0043] 402 , performing feature processing on the traffic image information to obtain a feature space.
[0044] Here, the feature space refers to the vector space composed of all possible feature vectors. The feature space is obtained by using a convolutional residual network (CNN+ResNet) to extract two-dimensional (2D) features of traffic image information and performing feature fusion on the extracted features.
[0045] In some possible implementations, first, the traffic image features are obtained by extracting features from the traffic image information; for example, the edges and contours of the traffic image information are extracted by the 2D feature extraction network in the CNN+ResNet network to obtain the traffic image features. The 2D feature extraction network is pre-trained based on the traditional back-propagation method, and the CNN+ResNet network uses the U-Net architecture to reproduce and enhance the original information. In order to reduce the amount of calculation and training, CNN+ResNet uses fewer layers than ordinary networks and shares parameter information when necessary.
[0046] Then, the traffic image features are subjected to feature fusion to obtain the feature space. Here, the 2D feature information network performs feature fusion, and the features of different levels in the feature fusion are matched respectively, and the features of the same level are optimized and combined. For example, the features of different levels are fused by top-down and lateral connection, and then prediction is performed. For example, lateral connection: at each layer of the top-down path, the upsampled feature map is added or spliced with the feature map of the corresponding level in the bottom-up path, and the feature information of different levels is fused to obtain a feature map with higher quality. In this way, feature fusion can improve the performance of the model, adapt to complex scenes, enhance the generalization ability of the model, etc. Feature fusion refers to the optimization combination of different feature vectors extracted from the same mode to form a more descriptive and discriminative feature set. These features can come from different sensors, different time periods, different processing levels or different data types. By fusing these features, richer information can be captured, thereby improving the performance, accuracy and robustness of the model, and then improving the performance of the model, adapting to complex scenes, enhancing the generalization ability of the model, etc.
[0047] 403 , based on the feature space corresponding to two adjacent moments, determine the feature change trend in the time series.
[0048] Here, after the traffic image features are fused, the feature space obtained is first regressed, and then the feature change trend is determined in the time series based on the spatial features after feature regression at two adjacent moments. The 2D feature information network performs feature regression, extracts the features that have an impact on the target variable, and standardizes or normalizes the features to ensure that they are on the same scale. The purpose of feature regression is to predict the value of the target variable based on the existing feature data.
[0049] like Figure 5 As shown, for the input image space 51 (i.e., traffic image information), feature extraction is performed, and then feature fusion is performed to obtain a feature space; then, feature regression is performed on the obtained feature space to obtain a feature space 52 at time t and a feature space 53 at time (t-1) (i.e., feature spaces at two adjacent moments). By performing feature differentiation on the feature space 52 and the feature space 53 at two adjacent moments, the feature change trend in the time series can be obtained.
[0050] In some possible implementations, first, feature differences are performed on the feature spaces corresponding to each two adjacent moments to obtain a plurality of differential features; and based on the plurality of differential features, the feature change trend is determined in a time series. Figure 5As shown, feature space 52 and feature space 53 are feature differentiated to obtain multiple differential features, and multiple differential features are concatenated to obtain differential space 54. Feature space 54 can be used to characterize the changing trend of feature information over time, that is, the changing trend of features in the time series is obtained. Here, a differential network is constructed based on the time flow, and a differential operation is performed on the feature space of every two adjacent moments. After all feature spaces are differentiated, multiple differential features are obtained, and a differential network is formed by multiple differential features. The purpose of differential operation is to capture the changing trend or pattern of feature information over time, so as to enhance the information.
[0051] 404 , performing depth prediction on the target object during the driving process based on the feature change trend to obtain depth feature information of the target object.
[0052] Here, by analyzing the feature change trend in the time series, combining the imaging height difference of the target object in the image at adjacent moments and the displacement of the vehicle between two adjacent moments, the depth feature information of the target object can be accurately calculated.
[0053] In some possible implementations, the above step 404 may be implemented by the following steps 441 to 443 (not shown):
[0054] 441, performing time characteristic analysis on the characteristic change trend to obtain the characteristic change trend at different moments.
[0055] Here, in the time series, by extracting the feature change trend at the feature moment, the feature change trend at different moments can be obtained. Since the feature change trend is represented by the differential space, the feature change trend at different moments can be represented by the differential features at different moments. In time series analysis, version differences (shift) are used to predict depth. In time series analysis, shift refers to moving the time series data forward or backward by a certain time step. This operation can be used to generate lagged variables to analyze the autocorrelation of time series data. Shift can be used to construct features, which can then be used to train models to predict future depth values.
[0056] 442, based on the feature change trend at the current moment and the feature change trend at the previous moment, respectively determine the imaging height difference of the target object between the current moment and the previous moment, and the moving distance of the vehicle between the current moment and the previous moment.
[0057] In some possible implementations, the difference between the first imaging height and the second imaging height in the traffic image information corresponding to the target object at adjacent moments is analyzed to obtain the imaging height difference; thus, the above step 442 can be implemented by the following process:
[0058] Firstly, in the characteristic change trend at the previous moment, a first imaging height of the target object in the traffic image information corresponding to the previous moment is determined.
[0059] Here, since the feature change trend is obtained through multiple differential features, the imaging height in the traffic image information at each moment can be obtained from the feature change trend at each moment. Figure 6 As shown in the figure, in the characteristic change trend at the previous moment (t+1), the height of the target object in the image is obtained, that is, the imaging height of the object on the camera imaging plane at that moment, that is, the first imaging height h is obtained. t+1 .exist Figure 6 In the equation, H represents the target height of the target object, i.e., the height of the object in the real scene. f represents the focal length, i.e., the distance from the center of the camera lens to the imaging plane. D represents the distance from the object to the camera, i.e., the depth distance. W represents the width of the object in the real scene (the width is used as a scene dimension). s represents the distance moved by the vehicle between time t and t+1.
[0060] Secondly, in the characteristic change trend at the current moment, a second imaging height of the target object in the traffic image information corresponding to the current moment is determined.
[0061] like Figure 7 As shown, in the characteristic change trend at the current time (t), the height of the target object in the image is obtained, that is, the second imaging height h is obtained. t . Among them, H represents the target height of the target object, that is, the height of the object in the real scene. f represents the focal length, that is, the distance from the center of the camera lens to the imaging plane. D represents the distance from the object to the camera, that is, the depth distance. W represents the width of the object in the real scene (the width is used as a scene dimension).
[0062] Finally, an imaging height difference between the first imaging height and the second imaging height is determined.
[0063] In this way, by subtracting the first imaging height from the second imaging height, the imaging height difference between adjacent moments can be accurately obtained, so that the depth feature information can be calculated by combining the imaging height difference with the moving distance between adjacent moments.
[0064] 443. Obtain depth feature information of the target object based on the imaging height difference and the moving distance.
[0065] Here, by dividing the imaging height difference by the imaging height at the previous moment and combining it with the moving distance, the depth feature information of the target object can be calculated. Figure 5As shown, a deep network 55 is constructed by using a plurality of differential features, and then the deep feature information of the target object is output through the deep network.
[0066] In some possible implementations, the above step 443 may be implemented by the following process:
[0067] First, a first imaging height of the target object in the traffic image information corresponding to the previous moment is obtained.
[0068] Here, the imaging height of the target object in the traffic image information at the previous moment (ie, t+1) is the first imaging height.
[0069] Next, a ratio between the first imaging height and the imaging height difference is determined.
[0070] Here, the ratio is obtained by calculating the quotient between the first imaging height and the imaging height difference.
[0071] Finally, the ratio and the moving distance are fused to obtain the depth feature information of the target object at the current moment.
[0072] Here, by multiplying the ratio and the moving distance, the depth feature information of the target object at the current moment can be obtained.
[0073] Here, the real height of the target object and the focal length corresponding to the current moment are first obtained; then based on the real height of the target object and the focal length corresponding to the current moment, the ratio and the moving distance are fused to obtain the depth feature information of the target object at the current moment. Figure 6 and 7 As shown, the real height H and focal length f of the target object are obtained, and the depth feature information D can be calculated by the following formulas (1) and (2):
[0074]
[0075] Based on the above formulas (1) and (2), we can get: Thus we get
[0076]
[0077] 405 , identifying the target object based on the depth feature information and the feature space at the current moment.
[0078] Here, multiple target recognition prediction heads are constructed through deep feature information and combined with the feature space at the current moment, so as to achieve accurate recognition of the target object.
[0079] In some possible implementations, the above step 405 may be implemented by the following steps 451 to 453 (not shown):
[0080] 451, constructing a target recognition prediction head based on the deep feature information.
[0081] For example, through deep feature information, a target recognition prediction head is constructed to achieve target recognition, obstacle recognition, motion feature prediction, operational space recognition, semantic information, lane line recognition, traffic sign recognition, etc.
[0082] 452, transferring the feature space at the current moment to the target recognition prediction head.
[0083] Here, the feature space at the current moment is fed back to the target recognition prediction head so that the target recognition prediction head can combine the feature space at the current moment with the deep feature information to more accurately recognize the target object.
[0084] 453, using the target recognition prediction head to recognize the target object.
[0085] Here, after forming a target recognition prediction head through deep feature information, the target recognition prediction head can accurately identify the target object in the traffic image information.
[0086] In an embodiment of the present invention, after obtaining the traffic image information of the vehicle during driving, feature processing is performed on the traffic image information to obtain a feature space, and based on the feature space corresponding to two adjacent moments, the feature change trend is determined in the time series; in this way, the change of the traffic image information in the time series can be accurately reflected through the feature change trend. Afterwards, the depth of the target object in the driving process is predicted based on the feature change trend to obtain the depth feature information of the target object; in this way, the displacement of the target object at different time points during driving can be characterized in combination with the feature change trend, so that the depth feature information of the target object can be accurately calculated; it is convenient to identify the target object through the depth feature information and the feature space at the current moment, so as to improve the accuracy of target object identification.
[0087] In some possible implementations, the present invention provides a depth estimation method based on a continuous time flow field, which can be achieved by Figure 8 To achieve this, first obtain image data 801; secondly, use the model to perform depth estimation 802; then calculate the depth information of each object in the scene 803; and perform multi-target recognition 804, such as obstacle recognition, lane line recognition, traffic sign recognition, etc.; finally, adjust the vehicle driving policy 805 in real time (such as obstacle avoidance, lane change, parking, etc.).
[0088] The embodiment of the present invention provides a depth estimation system based on continuous time flow field, please refer to Fig. 9 , which shows a schematic diagram of the composition structure of a depth estimation system based on a continuous-time flow field provided by an embodiment of the present invention. The system 900 includes:
[0089] The acquisition module 901 is used to acquire traffic image information of the vehicle during driving;
[0090] A processing module 802 is used to perform feature processing on the traffic image information to obtain a feature space;
[0091] A determination module 903 is used to determine a feature change trend in a time series based on feature spaces corresponding to two adjacent moments;
[0092] A prediction module 904 is used to predict the depth of the target object in the driving process based on the characteristic change trend to obtain the depth characteristic information of the target object;
[0093] The recognition module 905 is used to recognize the target object based on the depth feature information and the feature space at the current moment.
[0094] In some possible implementations, the prediction module is also used to perform time feature analysis on the feature change trend to obtain the feature change trend at different moments; based on the feature change trend at the current moment and the feature change trend at the previous moment, determine the imaging height difference of the target object between the current moment and the previous moment, and the moving distance of the vehicle between the current moment and the previous moment; based on the imaging height difference and the moving distance, obtain the depth feature information of the target object.
[0095] In some possible implementations, the prediction module is further used to determine, in the characteristic change trend at the previous moment, a first imaging height of the target object in the traffic image information corresponding to the previous moment; and in the characteristic change trend at the current moment, determine a second imaging height of the target object in the traffic image information corresponding to the current moment;
[0096] An imaging height difference between the first imaging height and the second imaging height is determined.
[0097] In some possible implementations, the prediction module is also used to obtain a first imaging height in the traffic image information corresponding to the target object at the previous moment; determine a ratio between the first imaging height and the imaging height difference; and fuse the ratio with the moving distance to obtain depth feature information of the target object at the current moment.
[0098] In some possible implementations, the prediction module is also used to obtain the true height of the target object and the focal length corresponding to the current moment; based on the true height of the target object and the focal length corresponding to the current moment, the ratio and the moving distance are fused to obtain the depth feature information of the target object at the current moment.
[0099] In some possible implementations, the determination module is further used to perform feature differentiation on the feature space corresponding to every two adjacent moments to obtain a plurality of differential features; and based on the plurality of differential features, determine the feature change trend in a time series.
[0100] In some possible implementations, the recognition module is further used to construct a target recognition prediction head based on the deep feature information; transfer the feature space at the current moment to the target recognition prediction head; and use the target recognition prediction head to identify the target object.
[0101] In some possible implementations, the processing module is further used to perform feature extraction on the traffic image information to obtain traffic image features; and perform feature fusion on the traffic image features to obtain the feature space.
[0102] Optionally, the transmission medium can be a wired link (for example, but not limited to, coaxial cable, optical fiber and digital subscriber line (DSL), etc.) or a wireless link (for example, but not limited to, wireless Fidelity (WIFI), Bluetooth and mobile device network, etc.). It should be noted that: the system provided in the above embodiment is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device is divided into different functional modules to complete all or part of the functions described above. In addition, the method embodiments provided in the above embodiments belong to the same concept. The specific implementation process is detailed in the method embodiments, which will not be repeated here.
[0103] Fig.10 is a schematic diagram of the structure of a computer device provided by an embodiment of the present invention. Fig.10 As shown, the computer device 1000 includes: a memory 1001, a processor 1002, and a computer program 1003 stored in the memory 1001 and running on the processor 1002, wherein when the processor 1002 executes the computer program 1003, the computer device can execute any one of the depth estimation methods based on the continuous-time flow field introduced above.
[0104] In addition, an embodiment of the present invention also protects a system, which may include a memory and a processor, wherein an executable program code is stored in the memory, and the processor is used to call and execute the executable program code to perform a depth estimation method based on a continuous-time flow field provided in an embodiment of the present invention. In this embodiment, the system can be divided into functional modules according to the above method example. For example, it can correspond to each functional module, or two or more functions can be integrated into a processing module, and the above integrated module can be implemented in the form of hardware. It should be noted that the division of modules in this embodiment is schematic, which is only a logical function division, and there may be other division methods in actual implementation. It should be noted that all relevant contents of each step involved in the above method embodiment can be referred to the functional description of the corresponding functional module, which will not be repeated here.
[0105] It should be understood that the system provided in this embodiment is used to perform the above-mentioned depth estimation method based on continuous time flow field, so the same effect as the above-mentioned implementation method can be achieved. In the case of an integrated unit, the system may include a processing module and a storage module. Among them, when the system is applied to a device, the processing module can be used to control and manage the actions of the device. The storage module can be used to support the device to execute mutual program codes, etc. Among them, the processing module can be a processor or a controller, which can implement or execute various exemplary logic boxes, modules and circuits described in conjunction with the disclosure of the present invention. The processor can also be a combination that implements a computing function, such as a combination of one or more microprocessors, a combination of digital signal processing (DSP) and a microprocessor, etc., and the storage module can be a memory.
[0106] In addition, the system provided by the embodiment of the present invention may be a chip, a component or a module, and the chip may include a connected processor and a memory; wherein the memory is used to store instructions, and when the processor calls and executes the instructions, the chip can execute a depth estimation method based on a continuous-time flow field provided in the above embodiment. This embodiment also provides a computer-readable storage medium, in which a computer program code is stored, and when the computer program code is run on a computer, the computer executes the above-mentioned related method steps to implement a depth estimation method based on a continuous-time flow field provided in the above embodiment.
[0107] This embodiment also provides a computer program product, when the computer program product is run on a computer, the computer executes the above-mentioned related steps to implement a depth estimation method based on a continuous time flow field provided in the above embodiment. Among them, the system, computer-readable storage medium, computer program product or chip provided in this embodiment are all used to execute the corresponding method provided above, so the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method provided above, and will not be repeated here. Through the description of the above implementation mode, the technicians in the relevant field can understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In practical applications, the above-mentioned function allocation can be completed by different functional modules as needed, that is, the internal structure of the system is divided into different functional modules to complete all or part of the functions described above. In the embodiments provided by the present invention, it should be understood that the disclosed system and method can be implemented in other ways. For example, the system embodiment described above is only schematic, for example, the division of modules or units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interface, system or unit, which may be electrical, mechanical or other forms.
[0108] It should be noted that the sequence of the above-mentioned embodiments of the present invention is only for description and does not represent the advantages and disadvantages of the embodiments. The process depicted in the accompanying drawings does not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are also possible or may be advantageous. The various embodiments in this specification are described in a progressive manner, and the same and similar parts between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from other embodiments. The above content is only a specific implementation method of the present invention, but the protection scope of the present invention is not limited to this. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be covered within the protection scope of the present invention.
Claims
1. A depth estimation method based on continuous time flow field, characterized in that: The depth estimation method based on continuous time flow field includes: Obtain traffic image information of vehicles while they are traveling; Performing feature processing on the traffic image information to obtain a feature space; Based on the feature space corresponding to two adjacent moments, the feature change trend is determined in the time series; Performing depth prediction on the target object during the driving process based on the characteristic change trend to obtain depth characteristic information of the target object; The target object is identified based on the depth feature information and the feature space at the current moment.
2. A depth estimation method based on continuous time flow field according to claim 1, characterized in that: The performing depth prediction on the target object in the driving process based on the feature change trend to obtain depth feature information of the target object includes: Performing time characteristic analysis on the characteristic change trend to obtain the characteristic change trend at different times; Based on the feature change trend at the current moment and the feature change trend at the previous moment, respectively determine the imaging height difference of the target object between the current moment and the previous moment, and the moving distance of the vehicle between the current moment and the previous moment; Based on the imaging height difference and the moving distance, depth feature information of the target object is obtained.
3. A depth estimation method based on continuous time flow field according to claim 2, characterized in that: The determining, based on the feature change trend at the current moment and the feature change trend at the previous moment, respectively the imaging height difference of the target object between the current moment and the previous moment, and the moving distance of the vehicle between the current moment and the previous moment, comprises: In the characteristic change trend at the previous moment, determining a first imaging height of the target object in the traffic image information corresponding to the previous moment; Determining, in the characteristic change trend at the current moment, a second imaging height of the target object in the traffic image information corresponding to the current moment; An imaging height difference between the first imaging height and the second imaging height is determined.
4. The depth estimation method based on continuous time flow field according to claim 2, characterized in that: The obtaining the depth feature information of the target object based on the imaging height difference and the moving distance includes: Acquire a first imaging height of the target object in the traffic image information corresponding to the previous moment; determining a ratio between the first imaging height and the imaging height difference; The ratio and the moving distance are fused to obtain the depth feature information of the target object at the current moment.
5. A depth estimation method based on continuous time flow field according to claim 4, characterized in that: Before fusing the ratio and the moving distance to obtain the depth feature information of the target object at the current moment, the method further includes: Obtaining the real height of the target object and the focal length corresponding to the current moment; Based on the real height of the target object and the focal length corresponding to the current moment, the ratio and the moving distance are fused to obtain the depth feature information of the target object at the current moment.
6. A depth estimation method based on continuous time flow field according to claim 1, characterized in that: The determining of the feature change trend in the time series based on the feature space corresponding to two adjacent moments includes: Perform feature difference on the feature space corresponding to each two adjacent moments to obtain multiple differential features; Based on the multiple differential features, the feature change trend is determined in time series.
7. The depth estimation method based on continuous time flow field according to claim 1, characterized in that: The identifying the target object based on the depth feature information and the feature space at the current moment includes: Based on the deep feature information, construct a target recognition prediction head; Transferring the feature space at the current moment to the target recognition prediction head; The target object is identified using the target recognition prediction head.
8. The depth estimation method based on continuous time flow field according to claim 1, characterized in that: The performing feature processing on the traffic image information to obtain a feature space includes: Extracting features from the traffic image information to obtain traffic image features; The traffic image features are subjected to feature fusion to obtain the feature space.
9. A depth estimation system based on continuous time flow field, characterized in that: The system comprises: An acquisition module is used to acquire traffic image information of a vehicle during driving; A processing module, used for performing feature processing on the traffic image information to obtain a feature space; A determination module, used to determine the feature change trend in the time series based on the feature space corresponding to two adjacent moments; A prediction module, configured to perform depth prediction on a target object during the driving process based on the characteristic change trend, and obtain depth characteristic information of the target object; The recognition module is used to recognize the target object based on the depth feature information and the feature space at the current moment.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program code, and when the computer program code is executed on a computer, the computer is enabled to execute the depth estimation method based on a continuous-time flow field according to any one of claims 1 to 8.
Citation Information
Patent Citations
Target tracking method based on particle filtering and depth distance metric learning
CN113128605A
Scene depth and camera motion prediction method and device, electronic equipment and medium
CN113822918A
Road surface detection method and device, vehicle, storage medium and chip
CN114782447A
Vehicle entering and leaving detection method, device and equipment and storage medium
CN117496387A
Image processing method, apparatus and device, and computer-readable storage medium
US20230316742A1