Target identification method based on online depth feedback alignment

Through the target recognition method based on online depth feedback alignment, deep learning needs for a large amount of training data in automotive depth estimation, and efficient depth estimation in a data scarce environment is achieved, and recognition accuracy and training efficiency are improved.

CN119992479APending Publication Date: 2025-05-13XIAMEN UNIV OF TECH

Patent Information

Application Number
CN202510204870.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Deep learning methods have an urgent need for a large number of reliable training data in the field of automotive depth estimation. Especially in the environment of scarce data, traditional regression methods are difficult to effectively solve the problem of continuous distance prediction, resulting in low-quality depth estimation results.

Method used

The target recognition method based on online depth feedback alignment is adopted. By obtaining the traffic image information of the vehicle during driving, feature processing and predictive feature generation, the difference information between the predicted feature and the true value feature is calculated, and feedback to the depth estimation network, and network parameters are adjusted to improve the recognition accuracy.

Benefits of technology

By enriching the training data of the deep estimation network, the dependence on high-quality labeled data is reduced, training efficiency is improved, and the generalization ability and recognition accuracy of the model are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992479A_ABST
    Figure CN119992479A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image data processing, in particular to a target recognition method based on online depth feedback alignment, and the method comprises the steps: obtaining the traffic image information of a vehicle in a driving process; performing feature processing on the traffic image information by adopting a depth estimation network to obtain a feature space; based on the feature space of the previous moment, generating a prediction feature of a next moment corresponding to the previous moment; determining difference information between the prediction feature at the next moment and the feature truth value at the current moment, and feeding back the difference information to the depth estimation network; adjusting network parameters of the depth estimation network based on the difference information to obtain a target identification network; and identifying a target object in the traffic image information based on the target identification network. According to the method, through instant feedback and adjustment, the training process of the depth estimation network is accelerated, and the training efficiency is improved, so that the recognition accuracy of the target recognition network on the target object in the traffic image information can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image data processing, and in particular to a target recognition method based on online depth feedback alignment in the technical field of image data processing. Background Art

[0002] Although trusted artificial intelligence (AI) technology is developing rapidly at this stage, its implementation faces many challenges, including conflicts of principles, large amounts of resources and time consumption, and complexity of working conditions (such as the diversity of application scenarios and differences in supervision and regulations).

[0003] In the field of automotive depth estimation technology, although traditional monocular and binocular ranging and deep learning-based methods have developed, deep learning methods have an urgent need for a large amount of reliable training data, especially in data-scarce environments. Most deep learning methods treat depth estimation as a regression problem. However, depth estimation is actually a continuous distance prediction problem that is more complex than classification, and the solution using traditional regression methods is often not ideal. As the number of layers in the deep convolutional depth estimation network increases, erroneous information is easily accumulated, resulting in low-quality depth estimation results. In addition, depth estimation algorithms require a lot of training time and high-quality labeled data, which limits their practicality in some application scenarios. Summary of the invention

[0004] The purpose of the present invention is to provide a target recognition method based on online depth feedback alignment, and the technical solution adopted is as follows:

[0005] In a first aspect, an embodiment of the present invention provides a target recognition method based on online depth feedback alignment, the method comprising:

[0006] Obtain traffic image information of vehicles while they are traveling;

[0007] Using a depth estimation network to perform feature processing on the traffic image information to obtain a feature space;

[0008] Based on the feature space of the previous moment, generating a prediction feature of the next moment corresponding to the previous moment;

[0009] Determine the difference information between the predicted feature at the next moment and the true value of the feature at the current moment, and feed it back to the depth estimation network;

[0010] Adjusting the network parameters of the depth estimation network based on the difference information to obtain a target recognition network;

[0011] Based on the target recognition network, the target object in the traffic image information is recognized.

[0012] In a second aspect, a target recognition system based on online deep feedback alignment is provided, the system comprising:

[0013] An acquisition module is used to acquire traffic image information of a vehicle during driving;

[0014] A processing module, used for performing feature processing on the traffic image information using a depth estimation network to obtain a feature space;

[0015] A generation module, used to generate a prediction feature of a next moment corresponding to the previous moment based on the feature space of the previous moment;

[0016] A determination module, used to determine the difference information between the predicted feature at the next moment and the feature true value at the current moment, and feed it back to the depth estimation network;

[0017] An adjustment module, configured to adjust network parameters of the depth estimation network based on the difference information to obtain a target recognition network;

[0018] The recognition module is used to recognize the target object in the traffic image information based on the target recognition network.

[0019] According to a third aspect, a computer program product is provided, comprising: a computer program code, which, when executed on a computer, enables the computer to execute the method in the first aspect or any possible implementation of the first aspect.

[0020] In a fourth aspect, a computer-readable storage medium is provided, which stores a computer program code. When the computer program code runs on a computer, the computer executes the method in the first aspect or any possible implementation manner described in the first aspect.

[0021] The present invention has the following beneficial effects: after obtaining the traffic image information of the vehicle during driving, the feature space is obtained by feature processing the traffic image information, and based on the feature space of the previous moment, the prediction feature of the next moment corresponding to the previous moment is generated; in this way, the prediction feature of the next moment is produced by the feature space of the previous moment, which can enrich the training data of the depth estimation network. Afterwards, the difference information between the prediction feature of the next moment and the feature true value of the current moment is determined, and fed back to the depth estimation network; so as to adjust the network parameters of the depth estimation network based on the difference information to obtain the target recognition network; in this way, after calculating the difference information between the prediction feature of the next moment and the feature true value of the current moment, unlabeled or weakly labeled data can be used for training through feedback learning, thereby reducing the dependence of the depth estimation network on a large amount of high-quality labeled data. At the same time, through immediate feedback and adjustment, the training process of the depth estimation network can be accelerated, the training efficiency can be improved, and then the recognition accuracy of the target recognition network for the target object in the traffic image information can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0023] Figure 1 It is a schematic diagram of the implementation flow of the target recognition method based on online deep feedback alignment provided by an embodiment of the present invention;

[0024] Figure 2 Schematic diagram of the implementation principle of the target recognition method based on online deep feedback alignment provided by an embodiment of the present invention;

[0025] Figure 3 is another implementation flow diagram of the target recognition method based on online depth feedback alignment provided by an embodiment of the present invention;

[0026] Figure 4 It is another schematic diagram of the implementation principle of the target recognition method based on online depth feedback alignment provided by an embodiment of the present invention;

[0027] Figure 5 is another schematic diagram of the implementation principle of the target recognition method based on online depth feedback alignment provided by an embodiment of the present invention;

[0028] Figure 6 It is a schematic diagram of the composition structure of the target recognition system based on online deep feedback alignment provided by an embodiment of the present invention;

[0029] Figure 7 It is a structural schematic diagram of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0030] In order to further explain the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following is a detailed description of the target recognition method based on online deep feedback alignment proposed by the present invention, its specific implementation method, structure, features and effects, in combination with the accompanying drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" does not necessarily refer to the same embodiment. In addition, specific features, structures or characteristics in one or more embodiments may be combined in any suitable form as described.

[0031] Among them, in the description of the embodiments of the present invention, unless otherwise specified, " / " means or, for example, A / B can mean A or B: "and / or" in the text is only a way to describe the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of the present invention, "multiple" refers to two or more than two.

[0032] In the following, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as suggesting or implying relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features.

[0033] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0034] At present, trusted AI technology is in a stage of rapid development, but it faces many challenges in its process of practical application. On the one hand, it is not easy to achieve multiple principles such as robustness and analyticity, and there are often conflicts between these principles. Taking deep learning as an example, adversarial training used to enhance robustness may come at the expense of fairness. On the other hand, building trusted AI requires huge computing resources and long-term investment. Reliable AI is like a big tree that needs fertile soil to nourish it, and has extremely high requirements for data quality. Once the data is flawed and contains low-quality or biased information, the performance and accuracy of the model will be greatly reduced, and loopholes may appear in practical applications, and even misleading results may be produced. Therefore, deep learning is extremely time-consuming and laborious from data collection, cleaning, labeling to subsequent rigorous testing and verification. In addition, as the scale and complexity of deep learning models continue to increase, the computing resources required for training and reasoning have also increased dramatically, and the cost is high, which makes it difficult for many companies to bear. At the same time, the complexity of working conditions is also a major problem facing trusted AI. Working conditions are diverse and changeable, involving many factors such as equipment status, environmental conditions, load changes, AI intelligence and application combination, running time, and interaction with other systems. In addition, the regulatory standards and policies and regulations in different countries and regions are different and constantly changing, making it almost impossible to collect data covering all working conditions, further exacerbating the difficulty of analyzability of trusted AI.

[0035] In the field of automotive depth estimation technology, although traditional monocular and binocular ranging and deep learning-based methods have developed, deep learning methods have an urgent need for a large amount of reliable training data, especially in data-scarce environments. In order to improve the reliability and analyzability of automotive depth estimation, the embodiment of the present invention will propose an online network based on feedback learning, which is of great significance for promoting the development of automotive intelligent driving technology and its industrial application.

[0036] Online network systems based on feedback learning have shown the potential to break through traditional limitations. They no longer rely heavily on large amounts of pre-collected data, nor do they need to traverse all possible working conditions. This system has the ability to learn and adapt in real time. When encountering new working conditions, it can quickly and online adjust its parameters to achieve accurate response and optimization processing to the new environment. This mechanism greatly improves the flexibility and generalization ability of the system, and opens up new paths for the practical application of trusted AI technology, especially in scenarios where data is scarce or working conditions are complex and changeable, showing great application prospects.

[0037] The following is a detailed description of the specific scheme of the target recognition method based on online depth feedback alignment provided by the present invention in conjunction with the accompanying drawings. Figure 1, which shows a schematic diagram of an implementation flow of a target recognition method based on online deep feedback alignment provided by an embodiment of the present invention, the method comprising:

[0038] 101, obtaining traffic image information of a vehicle during driving.

[0039] Here, the traffic image information may be a continuous video stream collected by a vehicle-mounted camera during the driving of the vehicle.

[0040] 102 , using a depth estimation network to perform feature processing on the traffic image information to obtain a feature space.

[0041] Here, the depth estimation network can be a convolutional depth estimation network or a residual depth estimation network. The feature space refers to a vector space composed of all possible feature vectors. The convolutional residual network (CNN+ResNet) is used to extract two-dimensional (2D) features of traffic image information, and the extracted features are fused to obtain the feature space.

[0042] In some possible implementations, first, the traffic image features are obtained by extracting features from the traffic image information; for example, the edges and contours of the traffic image information are extracted by the 2D feature extraction network in the CNN+ResNet network to obtain the traffic image features. The 2D feature extraction network is pre-trained based on the traditional back-propagation method, and the CNN+ResNet network uses the U-Net architecture to reproduce and enhance the original information. In order to reduce the amount of calculation and training, CNN+ResNet uses fewer layers than ordinary networks and shares parameter information when necessary.

[0043] Then, the traffic image features are subjected to feature fusion to obtain the feature space. Here, the 2D feature information network performs feature fusion, and the features of different levels in the feature fusion are matched respectively, and the features of the same level are optimized and combined. For example, the features of different levels are fused by top-down and lateral connection, and then prediction is performed. For example, lateral connection: at each layer of the top-down path, the upsampled feature map is added or spliced ​​with the feature map of the corresponding level in the bottom-up path, and the feature information of different levels is fused to obtain a feature map with higher quality. In this way, feature fusion can improve the performance of the model, adapt to complex scenes, enhance the generalization ability of the model, etc. Feature fusion refers to the optimization combination of different feature vectors extracted from the same mode to form a more descriptive and discriminative feature set. These features can come from different sensors, different time periods, different processing levels or different data types. Before feature fusion, the features that affect the variables are extracted, and the features are standardized or normalized to ensure that the features to be fused are on the same scale. By fusing these features, richer information can be captured, thereby improving the performance, accuracy, and robustness of the model, which in turn can improve the performance of the model, adapt to complex scenarios, enhance the generalization ability of the model, and so on.

[0044] 103. Generate prediction features of a next moment corresponding to the previous moment based on the feature space of the previous moment.

[0045] Here, after the feature space is extracted, the feature space is cached in the database. After that, the feature space of the previous moment is read from the database, and the prediction features of the next moment corresponding to the previous moment are obtained by reconstructing the features of the feature space of the previous moment. The feature space after feature regression of the previous moment is saved; the feature space of the current moment is passed to the Bird's Eye View (BEV) feature network to better perform feature fusion, thereby constructing multiple prediction heads. Among them, the BEV feature network is a network structure that represents the vehicle's surrounding environment information in the form of a bird's-eye view and extracts key features from it. BEV can fuse data from different sensors (such as lidar, cameras, etc.). This data fusion capability improves the accuracy and robustness of the autonomous driving system's perception of the environment, that is, it can better fuse data.

[0046] In some possible implementations, step 103 can be implemented through the following process: first, feature regression is performed on the feature space of the previous moment in the time series to obtain the regressed features of the previous moment; then, feature reconstruction is performed on the regressed features to obtain the predicted features of the next moment. Here, the predicted features of the next moment are predicted by reconstructing the regressed features. The convolutional depth estimation network + residual depth estimation network (CNN+ResNet) network uses a deep learning network (Convolutional Networks for Biomedical Image Segmentation, U-Net) architecture based on image segmentation to reproduce and enhance the original information to obtain the predicted features of the next moment. In order to reduce the amount of calculation and training, CNN+ResNet uses fewer layers than ordinary networks and shares parameter information when necessary. Figure 2 As shown, the CNN network performs feature extraction on the input image space 21 (i.e., traffic image information) to extract semantic features, wherein the CNN network is pre-trained based on the back-propagation method; thereafter, feature fusion is performed to obtain a feature space; thereafter, feature regression is performed on the obtained feature space to obtain a feature space 22 at time t and a feature space 23 at time (t-1), respectively; and then feature reconstruction is performed through the feature space 23 at time (t-1) to obtain a feature space 24 corresponding to the predicted feature at time t; at the same time, the feature space at the current moment (i.e., the feature space 22 at time t) is passed to the BEV feature network 25. The BEV feature network can be formed by a 2D feature information network; the function of the BEV feature network is to convert and fuse data from multiple sensors (such as cameras, radars, etc.) into feature representations from a bird's-eye view perspective. As shown Figure 2 As described above, multiple prediction heads are formed based on the BEV feature network, such as target recognition, obstacle recognition, motion feature prediction, operable space recognition, semantic information, lane line recognition, traffic sign recognition, etc. In this way, by reconstructing the feature space of the previous moment, the prediction features of the next moment are obtained, which can not only enrich the training data of the depth estimation network but also reduce the dependence of the depth estimation network on a large amount of labeled data.

[0047] In some possible implementations, after the feature space is extracted, the feature space is cached in a preset database; wherein the preset database is used to store the feature space at a historical moment; and in the preset database, the feature space at the previous moment is obtained. Figure 2 As shown, the feature space 23 at time (t-1) is read from the database so as to reconstruct the feature space at the next moment through the cached feature space at the previous moment, thereby obtaining the feature space at the next moment reconstructed based on the feature space at the previous moment.

[0048] 104, determining the difference information between the predicted feature at the next moment and the feature true value at the current moment, and feeding back the difference information to the depth estimation network.

[0049] Here, the difference information is obtained by differentiating the predicted features at the next moment and the feature true value at the current moment. The difference information can characterize the feature change trend of the predicted features at the next moment and the feature true value at the current moment in the time series. In this way, the predicted features are compared with the true value features, the difference is calculated, and the difference information is obtained. The difference information can be a loss function (such as mean square error); if the difference information represents that there is an error between the predicted features and the true value features, the gradient is calculated by the back propagation algorithm, and the network parameters (such as weights and biases) of the depth estimation network are optimized to reduce the future prediction error of the depth estimation network. In this way, through real-time feedback and online learning, errors in depth estimation can be discovered and corrected in time, avoiding low-quality results caused by the accumulation of erroneous information. This helps to improve the reliability and robustness of depth estimation. Moreover, feedback learning can use unlabeled or weakly labeled data for training, thereby reducing the dependence on a large amount of high-quality labeled data. At the same time, through instant feedback and adjustment, the training process of the model can be accelerated and the training efficiency can be improved.

[0050] In some possible implementations, the above step 104 can be performed by Figure 3 The steps shown achieve:

[0051] 301, performing differential processing on the predicted feature at the next moment and the feature true value at the current moment to obtain a differential result.

[0052] Here, the feature information of every two adjacent moments is compared to obtain a differential result, which is used to characterize whether the predicted feature is correct. Figure 2 As shown, feature difference 26 is performed between the predicted feature and the feature space at time t to obtain a difference result.

[0053] 302. Determine the difference information based on the difference result.

[0054] Here, the difference result is used as the difference information. In this way, by performing difference processing on the predicted feature and the feature true value at the next moment, it is possible to accurately analyze whether the predicted feature and the feature true value at the next moment are equal, thereby analyzing the prediction accuracy of the predicted feature.

[0055] 105. Adjust the network parameters of the depth estimation network based on the difference information to obtain a target recognition network.

[0056] Here, by feeding back the difference information to the feature extraction stage or feature regression stage of the depth estimation network, the depth estimation network re-extracts features or re-regresses features through the difference information to adjust the network parameters of the depth estimation network, thereby obtaining a target recognition network.

[0057] In some possible implementations, the above step 105 may be implemented by the following steps 151 and 152 (not shown):

[0058] 151, feeding back the difference information to the depth estimation network to update the feature space to obtain an updated feature space.

[0059] Here, if Figure 2 As shown, the feature difference is fed back to the feature extraction stage to re-extract features, and an updated feature space is obtained, so as to timely adjust the network parameters of the depth estimation network.

[0060] 152. Based on the predicted features at the next moment corresponding to the updated feature space and the true value features at the current moment, adjust the network parameters of the depth estimation network to obtain a target recognition network.

[0061] Here, after re-extracting the features, the difference information between the predicted features at the next moment and the true features at the current moment corresponding to the updated feature space is calculated again, so as to continue to adjust the network parameters of the depth estimation network and obtain the target recognition network. In this way, the regression parameters in the depth estimation network can be adjusted and optimized through feedback learning, so that it can better adapt to the continuous distance prediction problem, which helps to improve the accuracy and stability of depth estimation.

[0062] In some embodiments, the difference information can also be fed back to the feature regression stage of the depth estimation network to re-perform feature regression on the feature space of the previous moment to obtain updated regression features; then, the updated regression features are feature reconstructed to obtain the updated prediction features of the current moment; finally, based on the difference information between the updated prediction features at the current moment and the corresponding true value features, the network parameters of the depth estimation network are adjusted to obtain the target recognition network.

[0063] In some possible implementations, the above step 152 may be implemented by the following process:

[0064] First, the predicted features at the next moment corresponding to the updated feature space and the true value features at the current moment are differentially processed to obtain updated difference information.

[0065] Here, if Figure 2As shown, after the difference information is fed back to the feature extraction stage, feature extraction and feature fusion are performed again, and the difference calculation is performed on the predicted features of the next moment corresponding to the updated feature space and the true value of the current moment to obtain updated difference information, and then the updated difference information is fed back to the feature extraction stage of the depth estimation network again until the updated difference information meets the preset conditions.

[0066] Then, the updated difference information is fed back to the depth estimation network, the feature space is re-extracted and the network parameters are adjusted until the updated difference information meets the preset conditions to obtain the target recognition network.

[0067] Here, the preset condition may be that the predicted features of the next moment corresponding to the updated feature space represented by the updated difference information are equal to the true features of the current moment. When the predicted features and true features of two adjacent moments are equal, it means that the predicted features of the moment are correct. Afterwards, the predicted features of the moment and the difference results are fed back to the feature space of the previous moment, so that the depth estimation network knows that the predicted features of the moment are correct, and the depth estimation network of the moment is used as the target recognition network.

[0068] In some possible implementations, the first height information of the target object corresponding to the current moment and the second height information of the target object corresponding to the current moment are determined respectively by using the predicted features of the next moment corresponding to the updated feature space and the true features of the current moment; wherein the first height information is the height of the target object in the real scene calculated by the predicted features of the next moment; and the second height information is the height of the target object in the real scene calculated by the predicted features of the current moment. If the first height information and the second height information are the same, it is determined that the updated difference information meets the preset condition. The updated difference information ΔH can be implemented by formulas (1), (2) and (3):

[0069]

[0070] Only when hour, The depth estimation will be more accurate. Denotes the depth value of the depth estimation. Using ΔH to adjust D so that ΔH = 0 is the principle of time alignment, which means minimizing ΔH by adjusting the value of D until ΔH reaches 0 (or a threshold close to 0), which indicates the completion of time alignment. When it is successful, it can be determined is correct. Combined Figure 4 and 5 The following instructions are given: Figure 4As shown, in the prediction feature at the previous time (t+1), the height of the target object in the image is obtained, that is, the first height information H is obtained. t+1 .exist Figure 4 Middle, H t+1 represents the target height of the target object, that is, the height of the object in the real scene. f represents the focal length, that is, the distance from the center of the camera lens to the imaging plane; h represents the focal length, that is, the distance from the center of the camera lens to the imaging plane; t+1 W represents the imaging height of the target object on the camera imaging plane at that moment. D represents the distance from the object to the camera, that is, the depth distance. t+1 Represents the width of the object in the real scene (width is used as a scene dimension). s represents the distance the vehicle moves between time t and t+1.

[0071] exist Figure 5 In the example, the height of the target object in the image is obtained through the true value feature at the current time (t), that is, the second height information H is obtained. t Among them, H t represents the target height of the target object, that is, the height of the object in the real scene. f represents the focal length, that is, the distance from the center of the camera lens to the imaging plane. D represents the distance from the object to the camera, that is, the depth distance. W represents the width of the object in the real scene (the width is used as a scene dimension).

[0072] 106. Based on the target recognition network, identify the target object in the traffic image information.

[0073] Here, multiple target recognition prediction heads are constructed through the depth information output by the target recognition network, and combined with the feature space at the current moment, so as to achieve accurate recognition of the target object. In some possible implementations, Figure 2 The deep feature information output by the BEV feature network shown is used to construct a target recognition prediction head. For example, through the deep feature information, a target recognition prediction head is constructed for realizing target recognition, obstacle recognition, motion feature prediction, operational space recognition, semantic information, lane line recognition, traffic sign recognition, etc. Afterwards, the feature space at the current moment is passed to the target recognition prediction head. Here, the feature space at the current moment is fed back to the target recognition prediction head so that the target recognition prediction head can more accurately recognize the target object by combining the feature space at the current moment with the deep feature information.

[0074] Finally, the target object is identified by the target recognition prediction head. Here, after the target recognition prediction head is formed by the deep feature information, the target recognition prediction head can accurately identify the target object in the traffic image information.

[0075] In an embodiment of the present invention, feature space is obtained by processing the traffic image information, and based on the feature space of the previous moment, the prediction feature of the next moment corresponding to the previous moment is generated; in this way, the prediction feature of the next moment is produced by the feature space of the previous moment, which can enrich the training data of the depth estimation network. Afterwards, the difference information between the prediction feature of the next moment and the feature true value of the current moment is determined, and fed back to the depth estimation network; so that the network parameters of the depth estimation network are adjusted by the difference information to obtain the target recognition network; in this way, after calculating the difference information between the prediction feature of the next moment and the feature true value of the current moment, unlabeled or weakly labeled data can be used for training through feedback learning, thereby reducing the dependence of the depth estimation network on a large amount of high-quality labeled data. At the same time, through immediate feedback and adjustment, the training process of the depth estimation network can be accelerated, the training efficiency can be improved, and then the recognition accuracy of the target recognition network for the target object in the traffic image information can be improved.

[0076] The embodiment of the present invention provides an object recognition system based on online depth feedback alignment, please refer to Figure 6 , which shows a schematic diagram of the composition structure of an object recognition system based on online deep feedback alignment provided by an embodiment of the present invention, the system 600 includes:

[0077] The acquisition module 601 is used to acquire traffic image information of the vehicle during driving;

[0078] A processing module 602 is used to perform feature processing on the traffic image information using a depth estimation network to obtain a feature space;

[0079] A generating module 603, configured to generate a prediction feature of a next moment corresponding to the previous moment based on the feature space of the previous moment;

[0080] A determination module 604 is used to determine the difference information between the predicted feature at the next moment and the feature true value at the current moment, and feed it back to the depth estimation network;

[0081] An adjustment module 605 is used to adjust the network parameters of the depth estimation network based on the difference information to obtain a target recognition network;

[0082] The recognition module 606 is used to recognize the target object in the traffic image information based on the target recognition network.

[0083] In some possible implementations, the generation module 603 is further used to perform feature regression on the feature space of the previous moment in time series to obtain the regressed features of the previous moment; and perform feature reconstruction on the regressed features to obtain the predicted features of the next moment.

[0084] In some possible implementations, the generation module 603 is further used to cache the feature space to a preset database; wherein the preset database is used to store feature spaces at historical moments; and in the preset database, the feature space at the previous moment is obtained.

[0085] In some possible implementations, the determination module 604 is further used to perform differential processing on the predicted features at the next moment and the feature true value at the current moment to obtain a differential result; and determine the difference information based on the differential result.

[0086] In some possible implementations, the adjustment module 605 is further configured to feed back the difference information to the depth estimation network to update the feature space to obtain an updated feature space;

[0087] Based on the predicted features at the next moment corresponding to the updated feature space and the true value features at the current moment, the network parameters of the depth estimation network are adjusted to obtain a target recognition network.

[0088] In some possible implementations, the adjustment module 605 is also used to perform differential processing on the predicted features at the next moment corresponding to the updated feature space and the true value features at the current moment to obtain updated difference information; feed back the updated difference information to the depth estimation network, re-extract the feature space and adjust the network parameters until the updated difference information meets the preset conditions to obtain the target recognition network.

[0089] In some possible implementations, the adjustment module 605 is also used to determine the first height information of the target object corresponding to the current moment and the second height information of the target object corresponding to the current moment based on the predicted features of the next moment corresponding to the updated feature space and the true value features of the current moment; when the first height information and the second height information are the same, determine that the updated difference information meets the preset condition.

[0090] In some possible implementations, the processing module 602 is further used to perform feature extraction on the traffic image information to obtain traffic image features during the feature extraction stage of the depth estimation network; and perform feature fusion on the traffic image features to obtain the feature space.

[0091] Optionally, the transmission medium can be a wired link (for example, but not limited to, coaxial cable, optical fiber and digital subscriber line (DSL), etc.) or a wireless link (for example, but not limited to, wireless Fidelity (WIFI), Bluetooth and mobile device network, etc.). It should be noted that: the system provided in the above embodiment is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device is divided into different functional modules to complete all or part of the functions described above. In addition, the method embodiments provided in the above embodiments belong to the same concept. The specific implementation process is detailed in the method embodiments, which will not be repeated here.

[0092] Figure 7 is a schematic diagram of the structure of a computer device provided by an embodiment of the present invention. Figure 7 As shown, the computer device 700 includes: a memory 701, a processor 702, and a computer program 703 stored in the memory 701 and running on the processor 702, wherein when the processor 702 executes the computer program 703, the computer device can execute any of the depth estimation methods based on the continuous-time flow field introduced above.

[0093] In addition, an embodiment of the present invention also protects a system, which may include a memory and a processor, wherein an executable program code is stored in the memory, and the processor is used to call and execute the executable program code to perform the depth estimation method based on the continuous time flow field provided by the embodiment of the present invention. In this embodiment, the system can be divided into functional modules according to the above method example. For example, it can correspond to each functional module, or two or more functions can be integrated into one processing module, and the above integrated module can be implemented in the form of hardware. It should be noted that the division of modules in this embodiment is schematic, which is only a logical function division, and there may be other division methods in actual implementation. It should be noted that all relevant contents of each step involved in the above method embodiment can be referred to the functional description of the corresponding functional module, which will not be repeated here.

[0094] It should be understood that the system provided in this embodiment is used to perform the above-mentioned depth estimation method based on continuous-time flow field, so the same effect as the above-mentioned implementation method can be achieved. In the case of an integrated unit, the system may include a processing module and a storage module. Among them, when the system is applied to a device, the processing module can be used to control and manage the actions of the device. The storage module can be used to support the device to execute mutual program codes, etc. Among them, the processing module can be a processor or a controller, which can implement or execute various exemplary logic boxes, modules and circuits described in conjunction with the disclosure of the present invention. The processor can also be a combination that implements a computing function, such as a combination of one or more microprocessors, a combination of digital signal processing (DSP) and a microprocessor, etc., and the storage module can be a memory.

[0095] In addition, the system provided by the embodiment of the present invention may be specifically a chip, a component or a module, and the chip may include a connected processor and a memory; wherein the memory is used to store instructions, and when the processor calls and executes the instructions, the chip can execute the depth estimation method based on the continuous time flow field provided in the above embodiment. This embodiment also provides a computer-readable storage medium, in which a computer program code is stored, and when the computer program code is run on a computer, the computer executes the above-mentioned related method steps to implement the depth estimation method based on the continuous time flow field provided in the above embodiment.

[0096] This embodiment also provides a computer program product, when the computer program product is run on a computer, the computer executes the above-mentioned related steps to implement the depth estimation method based on the continuous time flow field provided in the above embodiment. Among them, the system, computer-readable storage medium, computer program product or chip provided in this embodiment are all used to execute the corresponding method provided above, so the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method provided above, and will not be repeated here. Through the description of the above implementation mode, the technicians in the relevant field can understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In practical applications, the above-mentioned function allocation can be completed by different functional modules as needed, that is, the internal structure of the system is divided into different functional modules to complete all or part of the functions described above. In the embodiments provided by the present invention, it should be understood that the disclosed system and method can be implemented in other ways. For example, the system embodiment described above is only schematic, for example, the division of modules or units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interface, system or unit, which may be electrical, mechanical or other forms.

[0097] It should be noted that the sequence of the above-mentioned embodiments of the present invention is only for description and does not represent the advantages and disadvantages of the embodiments. The process depicted in the accompanying drawings does not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are also possible or may be advantageous. The various embodiments in this specification are described in a progressive manner, and the same and similar parts between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from other embodiments. The above content is only a specific implementation method of the present invention, but the protection scope of the present invention is not limited to this. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be covered within the protection scope of the present invention.

Claims

1. An object recognition method based on online deep feedback alignment, characterized in that: The target recognition method based on online deep feedback alignment includes: Obtain traffic image information of vehicles while they are traveling; Using a depth estimation network to perform feature processing on the traffic image information to obtain a feature space; Based on the feature space of the previous moment, generating a prediction feature of the next moment corresponding to the previous moment; Determine the difference information between the predicted feature at the next moment and the true value of the feature at the current moment, and feed it back to the depth estimation network; Adjusting the network parameters of the depth estimation network based on the difference information to obtain a target recognition network; Based on the target recognition network, the target object in the traffic image information is recognized.

2. The target recognition method based on online deep feedback alignment according to claim 1, characterized in that: The generating, based on the feature space of the previous moment, a prediction feature of the next moment corresponding to the previous moment includes: Performing feature regression on the feature space of the previous moment in the time series to obtain the regressed features of the previous moment; The regressed features are reconstructed to obtain the predicted features at the next moment.

3. The target recognition method based on online depth feedback alignment according to claim 1, characterized in that: The method further comprises: Cache the feature space in a preset database; wherein the preset database is used to store the feature space of historical moments; In the preset database, the feature space at the previous moment is obtained.

4. The target recognition method based on online depth feedback alignment according to claim 1, characterized in that: The determining of the difference information between the predicted feature at the next moment and the feature true value at the current moment includes: Performing differential processing on the predicted feature at the next moment and the feature true value at the current moment to obtain a differential result; Based on the difference result, the difference information is determined.

5. The target recognition method based on online deep feedback alignment according to claim 1, characterized in that: The adjusting the network parameters of the depth estimation network based on the difference information to obtain the target recognition network includes: Feeding back the difference information to the depth estimation network to update the feature space to obtain an updated feature space; Based on the predicted features at the next moment corresponding to the updated feature space and the true value features at the current moment, the network parameters of the depth estimation network are adjusted to obtain a target recognition network.

6. The target recognition method based on online depth feedback alignment according to claim 5, characterized in that: The method of adjusting the network parameters of the depth estimation network based on the predicted features at the next moment corresponding to the updated feature space and the true features at the current moment to obtain the target recognition network includes: Performing differential processing on the predicted features at the next moment corresponding to the updated feature space and the true value features at the current moment to obtain updated difference information; The updated difference information is fed back to the depth estimation network, the feature space is re-extracted and the network parameters are adjusted until the updated difference information meets the preset conditions to obtain the target recognition network.

7. The target recognition method based on online depth feedback alignment according to claim 6, characterized in that: The method further comprises: Based on the predicted features at the next moment corresponding to the updated feature space and the true value features at the current moment, respectively determine the first height information of the target object corresponding to the current moment and the second height information of the target object corresponding to the current moment; In a case where the first height information and the second height information are the same, it is determined that the updated difference information meets the preset condition.

8. The target recognition method based on online depth feedback alignment according to claim 1, characterized in that: The method of using a depth estimation network to perform feature processing on the traffic image information to obtain a feature space includes: In the feature extraction stage of the depth estimation network, feature extraction is performed on the traffic image information to obtain traffic image features; The traffic image features are subjected to feature fusion to obtain the feature space.

9. Object recognition system based on online deep feedback alignment, characterized in that: The system comprises: An acquisition module is used to acquire traffic image information of a vehicle during driving; A processing module, used for performing feature processing on the traffic image information using a depth estimation network to obtain a feature space; A generation module, used to generate a prediction feature of a next moment corresponding to the previous moment based on the feature space of the previous moment; A determination module, used to determine the difference information between the predicted feature at the next moment and the feature true value at the current moment, and feed it back to the depth estimation network; An adjustment module, configured to adjust network parameters of the depth estimation network based on the difference information to obtain a target recognition network; The recognition module is used to recognize the target object in the traffic image information based on the target recognition network.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program code, and when the computer program code is executed on a computer, the computer is enabled to execute the object recognition method based on online deep feedback alignment according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Unsupervised learning techniques for temporal difference models

    CN110073369A

  • Crowd safety abnormal event identification method

    CN116229347A

  • Vehicle pixel height ratio detection method and device and computer equipment

    CN118052848A

  • Pure vision automatic driving environment sensing method and system based on variable depth mechanism

    CN119048998A

  • Object motion detection system based on combining 3D warping techniques and a proper object motion detection

    US20100315505A1

Cited By

  • Target sensing method based on feature difference and error propagation

    CN122157207A