Vehicle attitude estimation method and system under insufficient illumination condition
By constructing complementary teacher network models and pseudo-label training, the problem of vehicle attitude estimation under insufficient lighting is solved, and accurate vehicle attitude estimation in dark light scenarios is achieved, reducing the cost and difficulty in obtaining data.
Patent Information
- Application Number
- CN202411958160.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2044-12-30
AI Technical Summary
Under the condition of insufficient light, it is difficult for the prior art to effectively carry out vehicle attitude estimation, especially in dark light scenarios, with low accuracy, and high cost of existing methods or difficult to obtain data.
A complementary teacher network model is constructed, and the teacher network model is generated by annotating and enhancing the normal light image data, and a pseudo-label training student network is used to realize vehicle attitude estimation in dark light scenarios.
Good vehicle attitude estimation is achieved under extremely low light conditions, reducing the difficulty of vehicle attitude estimation under insufficient light conditions, and avoiding the need for installation of high-cost equipment and paired data acquisition.
Smart Images

Figure CN120236262A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent parking management, and particularly to a vehicle pose estimation method and system under insufficient lighting conditions. Background Art
[0002] In recent years, the high-position video technology has developed rapidly. By installing high-position video cameras on the roadside, it is possible to judge and manage the parking of vehicles. Through vehicle detection and vehicle body pose detection, and then data analysis with known parking space positions, the state of the vehicle can be judged, such as whether the vehicle is within the parking space, whether the vehicle is parked across the line, whether the vehicle is parked in a no-parking area, etc. In addition, judging whether the vehicle is illegally parked according to the vehicle body pose has a positive promoting effect on various aspects such as urban traffic management and driving safety.
[0003] For the judgment of the vehicle body pose, one is to use devices such as lidar, binocular or multi-camera to obtain the three-dimensional information of the vehicle, so as to estimate the vehicle pose and judge whether the parking state of the vehicle is illegally parked. For the roadside parking scenario, the cost of installing devices such as lidar is relatively high. The other is to directly use two-dimensional images for vehicle pose estimation methods, without the need to install other devices, only using the video image data captured by the camera to achieve vehicle pose estimation at a relatively low cost. However, in low-light scenarios, the image quality captured by the camera is low, the visibility of vehicle targets is low, and there is too much noise information in the image, resulting in low accuracy of vehicle pose estimation in low-light scenarios. However, in night scenarios, judging the parking state of the vehicle according to the vehicle body pose is crucial for traffic management, etc.
[0004] Currently, for vehicle pose estimation methods in low-light scenarios, one is to combine the RGB images captured by a visible light camera and the infrared images captured by an infrared camera, and use the advantage that infrared images are not affected by low-light conditions to perform vehicle pose estimation in low-light scenarios and improve the accuracy of vehicle pose estimation in low-light scenarios. However, the installation of thermal infrared cameras also limits their application in actual scenarios. Another method is to use paired normal-light and low-light image data for model learning to achieve low-light image enhancement. However, the acquisition of paired normal-light and low-light data is very difficult, and in actual applications, it is found that the enhanced low-light data often shows unrealistic image traces. Therefore, vehicle pose estimation based on the enhanced low-light images also has certain difficulties. Summary of the Invention
[0005] To solve the above technical problems, the present invention provides a vehicle pose estimation method and system under insufficient lighting conditions, which can solve the limitations and great implementation difficulties of existing vehicle pose estimation under insufficient lighting conditions.
[0006] To achieve the above object, the present invention provides a vehicle pose estimation method under insufficient light conditions, and the method includes:
[0007] According to the traffic scene video image data, data annotation is performed on the image data with normal light to obtain the two-dimensional detection frame and three-dimensional bounding box annotation information of the vehicle in the image;
[0008] According to the light-normal image data after data augmentation, model training is performed to construct a teacher network model;
[0009] Perform data augmentation on the annotated normal light image data, and train the teacher network model according to the augmented normal light image data to generate a trained teacher network model;
[0010] According to the trained teacher network model and the unannotated low-light scene image data, construct a student network and generate pseudo-labels corresponding to the low-light scene data;
[0011] Estimate the vehicle pose under insufficient light conditions according to the student network and the pseudo-labels corresponding to the low-light scene data.
[0012] Further, the step of constructing the teacher network model by performing model training according to the light-normal image data after data augmentation includes:
[0013] According to the light-normal image data after data augmentation, perform superposition operations on the convolutional layer, normalization layer, activation function layer, and convolutional combination layer of the backbone network model in the teacher network model;
[0014] Perform data aggregation on the features between different layers of the feature aggregation network model in the teacher network model;
[0015] According to the two-dimensional detection frame and three-dimensional bounding box annotation information of the vehicle in the image and the key point prediction network model in the teacher network model, learn the center points of eight vehicle key points, the offset of each vehicle key point from the center point, and the heat map.
[0016] Further, the step of performing data augmentation on the annotated normal light image data and training the teacher network model according to the augmented normal light image data to generate a trained teacher network model includes:
[0017] According to the loss function L main = L C + α main L O and L com = L HTrain the teacher network model to generate a trained teacher network model, where L C represents the training loss function for the center point of the vehicle key points, and the MSE loss function is used; α main is the coefficient parameter, and L O is the loss function for the offset, and the L1 loss function is used. L H represents the key point heat map loss function, and the MSE loss function is used.
[0018] Furthermore, the teacher network model includes a main teacher network model and an auxiliary teacher network model. The step of constructing a student network and generating pseudo-labels corresponding to the low-light scene data according to the trained teacher network model and unlabeled low-light scene image data includes:
[0019] According to the formula P all = NMS(Concat(P main [C main > s main , P com [C com > s com )), construct a student network and generate pseudo-labels corresponding to the low-light scene data, where P main , P com represent the prediction results of the main teacher model and the auxiliary teacher model respectively; s main , s com represent the score thresholds of the two teacher models respectively, C main , C com represent the prediction result scores of the two teacher models, Concat() means to splice and combine the prediction results greater than the threshold in the prediction results of the two models, and then use NMS() for filtering to remove duplicate prediction results.
[0020] Furthermore, before the step of estimating the vehicle pose under insufficient lighting conditions according to the student network and the pseudo-labels corresponding to the low-light scene data, the method includes:
[0021] According to the formula L S = α sup L sup + α unsup l unsup construct the loss function of the student network and train and update the student network, where α sup , α unsup represent the weight coefficients of supervised learning and unsupervised learning in the student network respectively, L sup , L unsuprespectively represent the supervised loss function and the unsupervised loss function in the student network, which are the same as the loss function trained by the main teacher model.
[0022] Furthermore, the present invention provides a vehicle pose estimation system under insufficient lighting conditions, and the system includes:
[0023] An annotation module, configured to perform data annotation on the image data with normal lighting according to the traffic scene video image data, and obtain the two-dimensional detection frame and three-dimensional bounding box annotation information of the vehicle in the image;
[0024] A construction module, configured to perform model training according to the image data with normal lighting after data augmentation, and construct a teacher network model;
[0025] A generation module, configured to perform data augmentation on the annotated normal lighting image data, and train the teacher network model according to the augmented normal lighting image data to generate a trained teacher network model;
[0026] The construction module is further configured to construct a student network according to the trained teacher network model and the unannotated low-light scene image data, and generate pseudo-labels corresponding to the low-light scene data;
[0027] An estimation module, configured to estimate the vehicle pose under insufficient lighting conditions according to the student network and the pseudo-labels corresponding to the low-light scene data.
[0028] Furthermore, the construction module is specifically configured to perform superposition operations on the convolutional layer, normalization layer, activation function layer, and convolutional combination layer of the backbone network model in the teacher network model according to the image data with normal lighting after data augmentation; perform data aggregation on the features between different layers of the network in the feature aggregation network model of the teacher network model; and learn the center points of eight vehicle key points, the offsets of each vehicle key point from the center point, and the heat map according to the two-dimensional detection frame and three-dimensional bounding box annotation information of the vehicle in the image and the key point prediction network model in the teacher network model.
[0029] Furthermore, the generation module is specifically configured to train the teacher network model according to the loss function L main =L C +α main L O and L com =L H to generate a trained teacher network model, where L C represents the training loss function for the center points of vehicle key points, and the MSE loss function is adopted; α main is a coefficient parameter, and L OThe loss function for the offset uses the L1 loss function, L H represents the key point heat map loss function and uses the MSE loss function.
[0030] Furthermore, the building block is also used to calculate according to the formula P all = NMS(Concat(P main [C main > s main , P com [C com > s com )) to construct the student network and generate the pseudo-labels corresponding to the low-light scene data. Here, P main , P com represent the prediction results of the main teacher model and the auxiliary teacher model respectively; s main , s com represent the score thresholds of the two teacher models respectively, C main , C com represent the prediction result scores of the two teacher models. Concat() means to splice and combine the prediction results greater than the threshold in the prediction results of the two models, and then use NMS() for filtering to remove duplicate prediction results.
[0031] Furthermore, the building block is also used to calculate according to the formula L S = α sup L sup + α unsup L unsup to construct the loss function of the student network and train and update the student network. Here, α sup , α unsup represent the weight coefficients of supervised learning and unsupervised learning in the student network respectively, and L sup , L unsup represent the supervised loss function and the unsupervised loss function in the student network respectively, which are the same as the loss function for training the main teacher model.
[0032] A vehicle pose estimation method and system under insufficient lighting conditions provided by the present invention generate more reliable pseudo-labels by constructing a set of complementary teacher network models, enabling the student model to achieve better vehicle pose estimation on extremely low-light images; realizing vehicle pose estimation learning in low-light scenes, without the need to obtain paired normal-light and low-light scene data, only normal-light images are required, greatly reducing the implementation difficulty of vehicle pose estimation under insufficient lighting conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 is a flowchart of a vehicle pose estimation method under insufficient lighting conditions provided by the present invention;
[0034] Figure 2 It is a schematic diagram of a vehicle attitude estimation system under insufficient lighting conditions provided by the present invention. Specific embodiments
[0035] The following further describes in detail the device structure and implementation manner of the present invention through the accompanying drawings and embodiments.
[0036] The present invention provides a vehicle attitude estimation method under insufficient lighting conditions, as Figure 1 shown, specifically including the following steps:
[0037] 101. According to the traffic scene video image data, data annotation is performed on the image data with normal lighting to obtain the two-dimensional detection box and three-dimensional bounding box annotation information of the vehicles in the image.
[0038] Specifically, in the traffic scene, vehicle parking behaviors include the process of a vehicle driving into a parking space, the vehicle in the parking space, and the process of the vehicle driving out of the parking space, and there are targets such as parking lines and roadside green belts; it includes vehicle data at different camera perspectives, different camera installation heights, different seasons, different lighting conditions, and different time periods; the two-dimensional detection box of the vehicle refers to the x-coordinates and y-coordinates of the upper left corner and the lower right corner of the vehicle rectangle, represented by (x_min, y_min, x_max, y_max); the three-dimensional bounding box information of the vehicle refers to the type of key points of the vehicle describing the three-dimensional bounding box of the vehicle, the coordinates of the key points, and the visibility attribute of the key points. Usually, the three-dimensional bounding box of the vehicle is used to describe the attitude of the vehicle; the types of eight key points are divided according to the following rules: the four key points where the vehicle contacts the ground and the four key points of the vehicle roof in the air.
[0039] Among them, according to the orientation of the vehicle head, the key point on the left side of the vehicle head in contact with the ground is used as a reference and denoted as point 1, bottom-front-1. Then, it is rotated clockwise, and the other three grounding points are respectively denoted as bottom-front-2, bottom-back-3, and bottom-back-4. Then, the key point on the left side of the vehicle head on the roof is denoted as point 5, top-front-5. After rotating clockwise, the other three roof key points are respectively denoted as top-front-6, top-back-7, and top-back-8, that is, eight types of key points describing the body attitude are obtained. The coordinates of the key points are the x-coordinates and y-coordinates in the image coordinate system. The visibility attribute includes two types: visible and invisible. The visible point attribute is denoted as 1, and the invisible point attribute is denoted as 0; the category, coordinate position, and visibility attribute of the above key points are used as label data.
[0040] 102. Train a model based on the light-normalized image data after data augmentation to construct a teacher network model.
[0041] For the embodiments of the present invention, step 102 may specifically include: performing a stacking operation of a convolutional layer, a normalization layer, an activation function layer, and a convolutional combination layer on the backbone network model in the teacher network model according to the light-normalized image data after data augmentation; aggregating data of features between different layers in the feature aggregation network model in the teacher network model; and learning the center points of eight vehicle key points, the offset of each vehicle key point from the center point, and the heat map according to the two-dimensional detection box and three-dimensional bounding box annotation information of the vehicle in the image and the key point prediction network model in the teacher network model.
[0042] Among them, there are two teacher networks to achieve complementary functions. One main teacher network learns the vehicle center point and key point offset through model learning, and the other auxiliary teacher network learns the key point heat map through model learning. For both the main teacher network and the auxiliary teacher network, they all include three parts, namely the backbone network, the feature aggregation network, and the key point prediction network. Among them, the network structures of the backbone network and the feature aggregation network are the same, only the learning methods for key points are slightly different in the key point prediction network part.
[0043] Specifically, the role of the backbone network is to extract image features. The input image size is H*W*C, where H and W respectively represent the height and width of the image, and C represents the number of channels. At this time, C is 3, representing an RGB three-channel image. The backbone network mainly adopts a convolutional combination method of convolutional layer - normalization layer - activation function layer to perform a stacking operation of multiple convolutional combination layers. In each convolutional combination operation, downsampling is performed once, including but not limited to using backbone networks such as ResNet, VGG, and MobileNet, and the downsampling multiple is R. The normalization layer includes but is not limited to instance normalization layer, adaptive instance normalization layer, etc. The non-linear activation layer includes but is not limited to non-linear activation functions such as ReLU and Leaky ReLU. The role of the feature aggregation network is to aggregate the high-level and low-level features extracted from different layers in the backbone network to provide more feature representations for subsequent key point detection. In a convolutional neural network, the high-level semantic features are close to the output end of the network but have a lower resolution, while the high-resolution features are close to the input end but have fewer semantic features. Therefore, aggregating the features between different layers in the network can achieve the fusion of high-level and low-level features, thereby improving the detection accuracy of subsequent key point detection tasks.
[0044] For the key point prediction part of the main teacher network, the goal is to learn the center points of eight vehicle key points and the offset of each vehicle key point from the center point.
[0045] The definition of the center point is as follows: Among them, represents the two-dimensional coordinates of the k-th key point of the i-th vehicle; the offset of each vehicle key point relative to the center point is defined as:
[0046] Furthermore, for the key point prediction part of the auxiliary teacher network, the goal is to learn the heatmaps of eight vehicle key points. Since for the main teacher network, the goal is to learn the center point and the offset, and it has a large dependence on the center point. When the vehicle center point is not detected due to occlusion or other reasons, the pose of the vehicle cannot be estimated. Therefore, an auxiliary teacher network is added to supplement the shortcomings of the main teacher network by learning the heatmaps of the key points.
[0047] 103. Perform data augmentation on the labeled normal illumination image data, and train the teacher network model according to the augmented normal illumination image data to generate a trained teacher network model.
[0048] For the embodiments of the present invention, step 103 may specifically include: according to the loss function L main = L C + α main L O and L com = L H train the teacher network model to generate a trained teacher network model, where L C represents the training loss function for the center point of the vehicle key points, and the MSE loss function is adopted; α main is a coefficient parameter, L O is the loss function for the offset, and the L1 loss function is adopted, L H represents the key point heatmap loss function, and the MSE loss function is adopted.
[0049] It should be noted that for the training data, the data collection and annotation were described in step 101. Due to the difficulty of obtaining paired data and the difficulty of annotating low-light data, only the data with normal illumination was annotated. In order to improve the generalization ability of the teacher model for low-light scene data, first perform data augmentation on the annotated normal illumination data; the purpose of the data augmentation is to perform data processing on the normal illumination image to simulate the characteristics of low-light scene data, including but not limited to data processing methods such as gamma correction, brightness adjustment, contrast adjustment, and adding Gaussian noise. Among them, the functions of gamma correction and brightness adjustment are to increase the low-light characteristics of the image, the function of contrast adjustment is to reduce the contrast of the image, and the function of adding Gaussian noise is to introduce noise to simulate the noise characteristics of low-light images;
[0050] Specifically, the above data augmentation method is used to perform data augmentation on the labeled normal illumination images, and the proportion of data introduced for data augmentation can be set to 0.5; based on the above augmented training data, a loss function is constructed to train the two teacher models; for the main teacher model, the loss function is defined as: L main = L C + α main L O , where L C represents the training loss function for the center point of the vehicle key points, and the MSE loss function is adopted; α main is the coefficient parameter, and L O is the loss function for the offset, and the L1 loss function is adopted; for the auxiliary loss function, the loss function is defined as: L com = L H
[0051] where L H represents the key point heat map loss function, and the MSE loss function is adopted; according to the above definitions of the data, loss function, and network structure of the teacher model, the teacher model is trained to obtain a trained teacher model for generating pseudo-labels in the subsequent training process of the student model.
[0052] 104. According to the trained teacher network model and the unlabeled dark light scene image data, construct a student network and generate pseudo-labels corresponding to the dark light scene data.
[0053] For the embodiment of the present invention, step 104 may specifically include: constructing a student network and generating pseudo-labels corresponding to the dark light scene data according to the formula P all = NMS(Concat(P main [C main > s main , P com [C com > s com ))), where P main , P com respectively represent the prediction results of the main teacher model and the auxiliary teacher model; s main , s com respectively represent the score thresholds of the two teacher models, C main , C com represent the prediction result scores of the two teacher models, Concat() represents concatenating and combining the prediction results greater than the threshold in the prediction results of the two models, and then filtering using NMS() to remove duplicate prediction results.
[0054] It should be noted that the network structure of the student model is the same as that of the main teacher model, including the backbone network, the feature aggregation network, and the key point prediction network. The key point prediction network uses the method of model learning for the vehicle center point and the key point offset. To improve the vehicle pose estimation learning ability of the student model in low-light scenarios and surpass the teacher model, the learning of the student model is divided into two parts, namely the supervised learning part and the unsupervised learning part. For the supervised learning part, it is similar to the training process of the main teacher model, only increasing the proportion of low-light data augmentation in the training data to improve the learning of the student model for low-light images. For the unsupervised learning part, the training data is unlabeled real low-light scene image data. The annotation generation for this data uses the trained teacher model for prediction, and the vehicle pose estimation results of the two teacher models for the unlabeled real low-light scene images are selected and fused as pseudo-labels to guide the training of the student model.
[0055] Specifically, for selecting and fusing the vehicle pose estimation results of the two teacher models to obtain the final pseudo-labels, the selection and fusion process is defined as:
[0056] P all = NMS(Concat(P main [C main > s main , P com [C com > s com ))
[0057] Among them, P main , P com respectively represent the prediction results of the main teacher model and the auxiliary teacher model; s main , s com respectively represent the score thresholds of the two teacher models, C main , C com represent the prediction result scores of the two teacher models, Concat() means concatenating and combining the prediction results greater than the threshold in the prediction results of the two models, and then using NMS() for filtering to remove duplicate prediction results.
[0058] 105. Estimate the vehicle pose under insufficient lighting conditions according to the student network and the pseudo-labels corresponding to the low-light scene data.
[0059] For the embodiments of the present invention, before step 105, it may further include: According to the formula L S = α sup L sup + α unsup L unsupConstruct a loss function for the student network and train and update the student network, where α sup and α unsup respectively represent the weight coefficients of supervised learning and unsupervised learning in the student network, and L sup and L unsup respectively represent the supervised loss function and the unsupervised loss function in the student network, which are the same as the loss function for training the main teacher model.
[0060] An vehicle pose estimation method under insufficient illumination conditions provided by an embodiment of the present invention generates more reliable pseudo-labels by constructing a set of complementary teacher network models, enabling the student model to achieve good vehicle pose estimation on images with extremely low illumination; realizes vehicle pose estimation learning in low-light scenarios without obtaining paired normal-illumination and low-light scenario data, only requiring normal-illumination images, greatly reducing the implementation difficulty of vehicle pose estimation under insufficient illumination conditions.
[0061] As Figure 1 a specific implementation manner of the method shown, an embodiment of the present invention provides a vehicle pose estimation system under insufficient illumination conditions, as Figure 2 shown, the system includes: an annotation module 21 for data annotating the image data with normal illumination according to traffic scene video image data to obtain two-dimensional detection box and three-dimensional bounding box annotation information of the vehicles in the images;
[0062] a construction module 22 for model training according to the data-augmented image data with normal illumination to construct a teacher network model;
[0063] a generation module 23 for data augmenting the annotated normal-illumination image data and training the teacher network model according to the augmented normal-illumination image data to generate a trained teacher network model;
[0064] the construction module 22 is further configured to construct a student network and generate pseudo-labels corresponding to the low-light scene data according to the trained teacher network model and unannotated low-light scene image data;
[0065] an estimation module 24 for estimating the vehicle pose under insufficient illumination conditions according to the student network and the pseudo-labels corresponding to the low-light scene data.
[0066] Further, the construction module 22 is specifically configured to perform superposition operations on the convolutional layer, normalization layer, activation function layer, and convolutional combination layer of the backbone network model in the teacher network model according to the light-normalized image data after data augmentation; perform data aggregation on the features between different layers in the feature aggregation network model in the teacher network model; and learn the center points of eight vehicle key points, the offsets of each vehicle key point from the center point, and the heat map according to the two-dimensional detection box and three-dimensional bounding box annotation information of the vehicle in the image and the key point prediction network model in the teacher network model.
[0067] Further, the generation module 23 is specifically configured to train the teacher network model according to the loss function L main = L C + α main L O and L com = L H to generate a trained teacher network model, where L C represents the training loss function for the center points of vehicle key points, and the MSE loss function is adopted; α main is a coefficient parameter, L O is the loss function for the offset, and the L1 loss function is adopted, and L H represents the key point heat map loss function, and the MSE loss function is adopted.
[0068] Further, the construction module 22 is further configured to construct a student network and generate pseudo-labels corresponding to the low-light scene data according to the formula P all = NMS(Concat(P main [C main > s main , P com [C com > s com ))), where P main , P com respectively represent the prediction results of the main teacher model and the auxiliary teacher model; s main , s com respectively represent the score thresholds of the two teacher models, C main , C com represent the prediction result scores of the two teacher models, Concat() represents concatenating and combining the prediction results greater than the threshold in the prediction results of the two models, and then filtering using NMS() to remove duplicate prediction results.
[0069] Further, the construction module 22 is further configured to calculate according to the formula L S = α sup L sup + αunsup L unsup Construct a loss function for the student network and train and update the student network, where α sup 、α unsup respectively represent the weight coefficients of supervised learning and unsupervised learning in the student network, and L sup 、L unsup respectively represent the supervised loss function and the unsupervised loss function in the student network, which are the same as the loss function trained by the main teacher model.
[0070] A vehicle pose estimation system under insufficient illumination conditions provided by an embodiment of the present invention generates more reliable pseudo-labels by constructing a set of complementary teacher network models, enabling the student model to achieve better vehicle pose estimation on images with extremely low illumination; realizing vehicle pose estimation learning in low-light scenarios, without the need to obtain paired normal illumination and low-light scenario data, only normal illumination images need to be obtained, greatly reducing the implementation difficulty of vehicle pose estimation under insufficient illumination conditions.
[0071] It should be understood that the specific order or hierarchy of steps in the disclosed process is an example of an exemplary method. Based on design preferences, it should be understood that the specific order or hierarchy of steps in the process can be rearranged without departing from the scope of the present disclosure. The appended method claims present the elements of the various steps in an exemplary order and are not intended to be limited to the specific order or hierarchy.
[0072] In the above detailed description, various features are combined in a single embodiment to simplify the present disclosure. This disclosure method should not be construed as reflecting an intention that the embodiments of the claimed subject matter require more features than those clearly stated in each claim. On the contrary, as reflected by the appended claims, the present invention is in a state with fewer features than all the features of the disclosed single embodiment. Therefore, the appended claims are hereby expressly incorporated into the detailed description, where each claim stands alone as a separate preferred embodiment of the present invention.
[0073] In order to enable any person skilled in the art to implement or use the present invention, the above-described disclosed embodiments have been described. For those skilled in the art; various modification methods of these embodiments are obvious, and the general principles defined herein can also be applied to other embodiments without departing from the spirit and scope of the present disclosure. Therefore, the present disclosure is not limited to the embodiments given herein, but is consistent with the broadest scope of the principles and novel features disclosed in this application.
[0074] The foregoing description includes examples of one or more embodiments. Of course, it is not possible to describe all possible combinations of components or methods for the purpose of describing the above embodiments, but those of ordinary skill in the art should recognize that the various embodiments can be further combined and arranged. Accordingly, the embodiments described herein are intended to cover all such changes, modifications, and variations that fall within the scope of the appended claims. In addition, with respect to the term "comprising" used in the specification or claims, this term is inclusive in a manner similar to the term "including", as that term is interpreted when used as a transitional word in a claim. Further, any use of the term "or" in the specification or claims is to be construed as "non-exclusive or".
[0075] Those skilled in the art will also appreciate that the various illustrative logical blocks, units, and steps listed in the embodiments of the present invention can be implemented by electronic hardware, computer software, or a combination of both. To clearly show the interchangeability of hardware and software, the various illustrative components, units, and steps have been generally described in terms of their functions. Whether such functions are implemented by hardware or software depends on the specific application and the design requirements of the overall system. Those skilled in the art can use various methods to implement the described functions for each specific application, but such implementation should not be construed as exceeding the scope of protection of the embodiments of the present invention.
[0076] The various illustrative logical blocks or units described in the embodiments of the present invention can be implemented or operated to perform the described functions by a general-purpose processor, a digital signal processor, an application specific integrated circuit (ASIC), a field programmable gate array or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination of the above designs. The general-purpose processor can be a microprocessor, and optionally, the general-purpose processor can also be any conventional processor, controller, microcontroller, or state machine. The processor can also be implemented by a combination of computing devices, such as a digital signal processor and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a digital signal processor core, or any other similar configuration.
[0077] In the embodiments of the present invention, the steps of the methods or algorithms described may be directly embedded in hardware, software modules executed by a processor, or a combination of the two. The software modules may be stored in a RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium in the art. Exemplarily, the storage medium may be connected to the processor so that the processor can read information from the storage medium and write information to the storage medium. Optionally, the storage medium may also be integrated into the processor. The processor and the storage medium may be provided in an ASIC, and the ASIC may be provided in a user terminal. Optionally, the processor and the storage medium may also be provided in different components of the user terminal.
[0078] In one or more exemplary designs, the above-described functions of the embodiments of the present invention may be implemented in hardware, software, firmware, or any combination of the three. If implemented in software, these functions may be stored on a computer-readable medium or transmitted on a computer-readable medium in the form of one or more instructions or codes. A computer-readable medium includes a computer storage medium and a communication medium that facilitates the transfer of a computer program from one place to another. The storage medium may be any available medium accessible by a general or special computer. For example, such a computer-readable medium may include, but is not limited to, RAM, ROM, EEPROM, CD-ROM, or other optical disk storage, magnetic disk storage, or other magnetic storage devices, or any other medium that can be used to carry or store program code in the form of instructions or data structures and other forms readable by a general or special computer, or a general or special processor. In addition, any connection may be appropriately defined as a computer-readable medium. For example, if software is transmitted from a website, server, or other remote resource via a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless means such as infrared, wireless, and microwave, it is also included in the defined computer-readable medium. The disks and discs include compact disks, laser disks, optical discs, DVDs, floppy disks, and Blu-ray discs. Disks usually reproduce data magnetically, while discs usually reproduce data optically by laser. The above combinations may also be included in the computer-readable medium.
[0079] The above-described specific embodiments have further elaborated on the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.
Claims
1. A vehicle posture estimation method under insufficient lighting conditions, characterized in that: The method comprises: Based on the traffic scene video image data, the image data with normal lighting is annotated to obtain the two-dimensional detection frame and three-dimensional bounding box annotation information of the vehicle in the image; Performing model training based on the normal illumination image data after data enhancement to construct a teacher network model; Performing data enhancement on the labeled normal illumination image data, and training the teacher network model according to the enhanced normal illumination image data to generate a trained teacher network model; According to the trained teacher network model and the unlabeled dark-light scene image data, a student network is constructed and a pseudo label corresponding to the dark-light scene data is generated; The vehicle posture under insufficient lighting conditions is estimated according to the student network and the pseudo labels corresponding to the dark light scene data.
2. The method for estimating vehicle posture under insufficient lighting conditions according to claim 1, characterized in that: The step of performing model training based on the normal illumination image data after data enhancement to construct a teacher network model comprises: According to the normal illumination image data after data enhancement, the backbone network model in the teacher network model is subjected to a superposition operation of a convolution layer, a normalization layer, an activation function layer and a convolution combination layer; Performing data aggregation on features between different layers of the network in the feature aggregation network model in the teacher network model; Based on the two-dimensional detection box and three-dimensional bounding box annotation information of the vehicle in the image and the key point prediction network model in the teacher network model, the center points of eight vehicle key points, the offset of each vehicle key point with respect to the center point, and the heat map are learned.
3. The method for estimating vehicle posture under insufficient lighting conditions according to claim 1 or 2, characterized in that: The step of performing data enhancement on the labeled normal illumination image data, and training the teacher network model according to the enhanced normal illumination image data to generate the trained teacher network model comprises: According to the main loss function And auxiliary loss function The teacher network model is trained to generate a trained teacher network model, wherein: represents the training loss function for the center of the vehicle key points, is the coefficient parameter, is the loss function of the offset, Represents the key point heat map loss function.
4. The method for estimating vehicle posture under insufficient lighting conditions according to claim 1, characterized in that: The teacher network model includes a main teacher network model and an auxiliary teacher network model. The step of constructing a student network and generating a pseudo label corresponding to the dark-light scene data according to the trained teacher network model and unlabeled dark-light scene image data includes: According to the formula Construct a student network and generate pseudo labels corresponding to the dark light scene data, where: is the pseudo label corresponding to the dark light scene data, , Represent the prediction results of the main teacher model and the auxiliary teacher model respectively; , Represent the score thresholds of the two teacher models, , represents the prediction result scores of the two teacher models, () indicates that the prediction results of the two models that are greater than the threshold are concatenated and combined, and NMS () indicates filtering to remove duplicate prediction results.
5. The method for estimating vehicle posture under insufficient lighting conditions according to claim 4, characterized in that: Before the step of estimating the vehicle posture under insufficient lighting conditions according to the student network and the pseudo labels corresponding to the dark light scene data, the method includes: According to the formula Construct a loss function of the student network and train and update the student network, where: is the loss function of the student model, , Respectively represent the weight coefficients of supervised learning and unsupervised learning in the student network, , denote the supervised and unsupervised loss functions in the student network, respectively, which are the same as the loss functions used in training the main teacher model.
6. A vehicle posture estimation system under insufficient lighting conditions, characterized in that: The system comprises: The annotation module is used to annotate the image data with normal illumination according to the traffic scene video image data, and obtain the two-dimensional detection frame and three-dimensional bounding box annotation information of the vehicle in the image; A construction module, used for performing model training according to the normal illumination image data after data enhancement to construct a teacher network model; A generation module, used to perform data enhancement on the labeled normal illumination image data, and train the teacher network model according to the enhanced normal illumination image data to generate a trained teacher network model; The construction module is further used to construct a student network and generate pseudo labels corresponding to the dark-light scene data based on the trained teacher network model and unlabeled dark-light scene image data; An estimation module is used to estimate the vehicle posture under insufficient lighting conditions based on the student network and the pseudo labels corresponding to the dark light scene data.
7. The vehicle posture estimation system under insufficient lighting conditions according to claim 6, characterized in that: The construction module is specifically used to perform superposition operations of convolution layer, normalization layer, activation function layer and convolution combination layer on the backbone network model in the teacher network model according to the normal illumination image data after data enhancement; perform data aggregation on the features between different network layers in the feature aggregation network model in the teacher network model; and learn the center points of eight vehicle key points, the offset of each vehicle key point relative to the center point, and the heat map according to the two-dimensional detection box and three-dimensional bounding box annotation information of the vehicle in the image and the key point prediction network model in the teacher network model.
8. A vehicle posture estimation system under insufficient lighting conditions according to claim 6 or 7, characterized in that: The generation module is specifically used to generate a loss function based on the main And auxiliary loss function The teacher network model is trained to generate a trained teacher network model, wherein: represents the training loss function for the center of the vehicle key points, is the coefficient parameter, is the loss function of the offset, Represents the key point heat map loss function.
9. The vehicle posture estimation system under insufficient lighting conditions according to claim 6, characterized in that: The building blocks are also used according to the formula Construct a student network and generate pseudo labels corresponding to the dark light scene data, where: is the pseudo label corresponding to the dark light scene data, , Represent the prediction results of the main teacher model and the auxiliary teacher model respectively; , Represent the score thresholds of the two teacher models, , represents the prediction result scores of the two teacher models, () indicates that the prediction results of the two models that are greater than the threshold are concatenated and combined, and NMS () indicates filtering to remove duplicate prediction results.
10. The vehicle posture estimation system under insufficient lighting conditions according to claim 9, characterized in that: The building blocks are also used according to the formula Construct a loss function of the student network and train and update the student network, where: is the loss function of the student model, , Respectively represent the weight coefficients of supervised learning and unsupervised learning in the student network, , denote the supervised and unsupervised loss functions in the student network, respectively, which are the same as the loss functions used in training the main teacher model.
Citation Information
Patent Citations
Roadside parking management method and system based on course angle attitude
CN115908558A
Roadside parking management method and system based on key point moving posture
CN115908559A
Roadside parking management method and system based on three-dimensional vehicle attitude
CN115909228A
High-order video scene analysis method and system
CN116091964A
Vehicle boundary positioning method for night scene
CN117152513A