Night bridge traffic detection data enhancement method based on CycleGAN
Through the CycleGAN-based night bridge traffic detection data enhancement method, the problems of low night detection accuracy and high data labeling cost are solved, and high-precision night traffic detection and the effect of reducing labeling cost is achieved.
Patent Information
- Application Number
- CN202510110314.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-01-23
AI Technical Summary
The prior art has low detection accuracy and high data labeling cost in night bridge traffic detection. Traditional data enhancement methods have failed to effectively solve the problems of complex lighting and color changes at night.
The night bridge traffic detection data enhancement method based on CycleGAN is adopted, and the night bridge traffic detection data enhancement model is constructed, real day and night traffic data are used for training, and combined with visual semantic information comparison, loss function model and multi-scale screening of FID values, false night traffic image data with consistent semantic information is generated.
The recognition accuracy and generalization ability of the YOLOv5 target detection model in night scenes is improved, the labeling cost of night traffic data is reduced, and the effect of generative data is enhanced.
Smart Images

Figure CN120071269A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of bridge traffic load identification, and particularly to a method for enhancing night bridge traffic detection data based on CycleGAN. Background Art
[0002] At present, some progress has been made in the method of bridge traffic load identification based on computer vision. However, the training and verification of most detection models mainly rely on data in daytime scenes, and their applicability in night scenes is significantly insufficient. The traffic flow on the bridge has different characteristics at different time periods. Especially at night, the lighting conditions are poor, and the interference of vehicle lights on the camera further exacerbates the difficulty of detection. Since the object detection technology is data-driven, this leads to the accuracy of the detection model highly depending on the training data set. When the training data set mainly covers daytime scenes, its detection performance in night scenes often significantly decreases. The most direct way to improve the accuracy of night traffic flow identification is to increase the proportion of night data in the training set, but this also causes the problem of the labor cost of data annotation.
[0003] Although some open-source data sets provide a certain amount of night data in the traffic field, these data sets are not fully applicable to the bridge scene. In order to obtain a more accurate bridge traffic flow object detection model, it is often necessary to additionally collect and annotate the traffic flow data of the night bridge scene. However, traditional data augmentation methods such as random cropping and flipping can increase the diversity of training data to a certain extent, but do not consider the complex lighting and color changes at night. Therefore, their effects on the pre-trained model in daytime for night scenes are limited. And annotating data is both time-consuming and expensive. Especially in the night environment, due to factors such as poor lighting conditions and camera interference, the annotation difficulty is further increased.
[0004] At present, although there are already some automatic annotation technologies that use machine learning for feature extraction, the annotation quality of these technologies is usually not as good as manual annotation. The increase in the cost of manual annotation also indirectly affects the above accuracy problem. Even if the personnel in this field try to fine-tune using the weights pre-trained on daytime images in an attempt to improve the night detection accuracy without additionally increasing the night marked data, the effect is not stable. In summary, how to enhance the night bridge traffic detection data while effectively reducing the cost and difficulty of data annotation has become one of the problems that need to be solved urgently in this field. Summary of the Invention
[0005] In view of the above deficiencies in the prior art, the present invention provides a method for enhancing night bridge traffic detection data based on CycleGAN.
[0006] In order to achieve the above invention purpose, the technical solution adopted by the present invention is as follows:
[0007] A method for enhancing night-time bridge traffic detection data based on CycleGAN, comprising the following steps:
[0008] S1. Collect the daytime vehicle flow picture data and night-time vehicle flow picture data of bridge traffic, fully annotate the daytime vehicle flow picture data to obtain the true daytime vehicle flow picture data with complete labels, and partially annotate the night-time vehicle flow picture data to obtain the true night-time vehicle flow picture data with partial labels;
[0009] S2. Build a night-time bridge traffic detection data enhancement model based on CycleGAN, and use the true daytime vehicle flow picture data with complete labels and the true night-time vehicle flow picture data with partial labels to train the night-time bridge traffic detection data enhancement model to obtain the trained night-time bridge traffic detection data enhancement model;
[0010] S3. Perform multi-scale screening on the trained night-time bridge traffic detection data enhancement model to obtain the final night-time bridge traffic detection data enhancement model;
[0011] S4. According to the true daytime vehicle flow picture data with complete labels and the final night-time bridge traffic detection data enhancement model, obtain the enhanced night-time bridge traffic detection data.
[0012] Furthermore, in step S2, the night-time bridge traffic detection data enhancement model includes a first generator, a second generator, a first discriminator, and a second discriminator; the first generator is used to convert the true daytime vehicle flow picture data in the first independent training domain into fake night-time vehicle flow picture data, and convert the fake daytime vehicle flow picture data generated by the second generator into fake night-time vehicle flow picture data; the second generator is used to convert the true night-time vehicle flow picture data in the second independent training domain into fake daytime vehicle flow picture data, and convert the fake night-time vehicle flow picture data generated by the first generator into fake daytime vehicle flow picture data; the first discriminator is used to discriminate the true night-time vehicle flow picture data in the second independent training domain and the fake night-time vehicle flow picture data generated by the first generator; the second discriminator is used to discriminate the true daytime vehicle flow picture data in the first independent training domain and the fake daytime vehicle flow picture data generated by the second generator.
[0013] Furthermore, using the true daytime vehicle flow picture data with complete labels and the true night-time vehicle flow picture data with partial labels to train the night-time bridge traffic detection data enhancement model to obtain the trained night-time bridge traffic detection data enhancement model, comprising the following steps:
[0014] S21. Build a first independent training domain and a second independent training domain;
[0015] S22. Input the real daytime traffic flow picture data with complete labels into the first independent training domain, and input the real nighttime traffic flow picture data with partial labels into the second independent training domain to train the nighttime bridge traffic detection data enhancement model, and obtain the trained nighttime bridge traffic detection data enhancement model.
[0016] Further, step S3 includes the following steps:
[0017] S31. Use the visual semantic information comparison method to preliminarily screen the trained nighttime bridge traffic detection data enhancement model to obtain an initial combination of nighttime bridge traffic detection data enhancement models;
[0018] S32. Construct a loss function model for nighttime bridge traffic detection data enhancement;
[0019] S33. Use the FID method and the loss function model for nighttime bridge traffic detection data enhancement to finally screen the initial combination of nighttime bridge traffic detection data enhancement models to obtain the final nighttime bridge traffic detection data enhancement model.
[0020] Further, in step S31, when using the visual semantic information comparison method to preliminarily screen the trained nighttime bridge traffic detection data enhancement model, the specific process is as follows: Determine the visual semantic information, including vehicle appearance type, relative distribution of traffic flow position, street lights, and isolation belts, and compare the consistency of semantic information between the fake nighttime traffic flow picture data generated by the trained nighttime bridge traffic detection data enhancement model and the real daytime traffic flow picture data before generation to preliminarily screen the trained nighttime bridge traffic detection data enhancement model.
[0021] Further, in step S32, the loss function model for nighttime bridge traffic detection data enhancement includes the least squares loss function for nighttime bridge traffic detection data enhancement, the cycle consistency loss function, and the identity mapping function.
[0022] Further, the least squares loss function for nighttime bridge traffic detection data enhancement is expressed as:
[0023]
[0024] Where: G is the first generator, D Y is the first discriminator, X is the first independent training domain, Y is the second independent training domain, E is the expectation for a certain distribution, y ∼ p data (y) means that y is sampled from the real nighttime traffic flow picture data distribution p data (y), x ∼ p data (x) means that x is sampled from the real daytime traffic flow picture data distribution p data (x), DY (y) is the output value of the first discriminator for the real nighttime traffic image data y. The closer the value is to 1, the better the first discriminator D Y The higher the probability that y is a true sample, the higher the probability that D Y (G(x)) is the output value of the first discriminator for the fake nighttime traffic image data G(x), F is the second generator, D X is the second discriminator, D x (x) is the output value of the second discriminator for the real daytime traffic image data x, D x (F(y)) is the output value of the second discriminator for the holiday traffic image data F(y).
[0025] Furthermore, the cycle consistency loss function for nighttime bridge traffic detection data enhancement is expressed as:
[0026]
[0027] Among them: G is the first generator, F is the second generator, E is the expectation of a certain distribution, x~p data (x) is the distribution of x from real daytime traffic image data p data (x) is sampled, || || 1 is the L1 norm, y~p data (y) is the distribution of y from real nighttime traffic image data p data (y), F(G(x)) is the output value of the second generator for the fake nighttime traffic picture data G(x), x is the real daytime traffic picture data, G(F(y)) is the output value of the first generator for the holiday traffic picture data F(y), y is the real nighttime traffic picture data.
[0028] Furthermore, the identity mapping function for nighttime bridge traffic detection data enhancement is expressed as:
[0029]
[0030] Among them: G is the first generator, F is the second generator, E is the expectation of a certain distribution, y~p data (y) is the distribution of y from real nighttime traffic image data p data (y) is sampled, || || 1 is the L1 norm, x~p data (x) is the distribution of x from real daytime traffic image data p data (x), G(y) is the output value of the first generator for the real nighttime traffic picture data y, y is the real nighttime traffic picture data, F(x) is the output value of the second generator for the real daytime traffic picture data x, x is the real daytime traffic picture data.
[0031] Further, in step S33, the initial combination of night-time bridge traffic detection data enhancement models is finally screened using the loss function model for night-time bridge traffic detection data enhancement. The specific process is as follows: Based on the loss function model for night-time bridge traffic detection data enhancement, the loss function curves of each model in the initial combination of night-time bridge traffic detection data enhancement models are plotted, and models with the four loss function values in the range tending to stable values and no fluctuations exceeding 0.3 are screened to finally screen the initial combination of night-time bridge traffic detection data enhancement models.
[0032] The present invention has the following beneficial effects:
[0033] The present invention constructs a night-time bridge traffic detection data enhancement model based on CycleGAN, trains it using unpaired real daytime and night-time traffic flow data, and adopts a multi-scale screening strategy combining visual training results, the loss function model for night-time bridge traffic detection data enhancement, and the FID value for comprehensive analysis from qualitative to quantitative to obtain the final night-time bridge traffic detection data enhancement model. Then, the final night-time bridge traffic detection data enhancement model is used to perform domain conversion on the real daytime traffic flow picture data with labels, and the label data of the daytime traffic flow is migrated as a whole to generate a false night-time traffic flow picture training set with consistent semantic information to obtain enhanced night-time bridge traffic detection data, which not only improves the recognition accuracy and generalization ability of the YOLOv5 object detection model, achieving the effect of generative data enhancement, but also the overall migration of the labeled data reduces the labor cost of night-time traffic flow data annotation. Description of the Drawings
[0034] Figure 1 It is a schematic flow chart of a night-time bridge traffic detection data enhancement method based on CycleGAN;
[0035] Figure 2 It is a schematic structural diagram of the night-time bridge traffic detection data enhancement model of the present invention;
[0036] Figure 3 It is a schematic flow chart of a simulation experiment using the method provided by the present invention;
[0037] Figure 4 It is a loss function curve diagram of the present invention;
[0038] Figure 5 It is an FID value result diagram of the present invention;
[0039] Figure 6 It is a simulation result diagram with validation set A as the validation set;
[0040] Figure 7It is a simulation result diagram with validation set B as the validation set;
[0041] Figure 8 It is a simulation result diagram with validation set C as the validation set. Specific implementation manners
[0042] The following describes the specific implementation manners of the present invention to facilitate those skilled in the art of this technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific implementation manners. For those of ordinary skill in the art of this technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.
[0043] As Figure 1 shown, a method for enhancing night bridge traffic detection data based on CycleGAN includes steps S1 - S4, specifically as follows:
[0044] S1. Collect daytime traffic flow picture data and night traffic flow picture data of bridge traffic, fully annotate the daytime traffic flow picture data to obtain true daytime traffic flow picture data with complete labels, and partially annotate the night traffic flow picture data to obtain true night traffic flow picture data with partial labels.
[0045] In an optional embodiment of the present invention, the present invention collects daytime traffic flow picture data and night traffic flow picture data of bridge traffic, and fully annotates the daytime traffic flow picture data to obtain true daytime traffic flow picture data with complete labels. In reality, due to the difficulty of night data annotation, the present invention simulates the real situation to partially annotate the night traffic flow picture data to obtain true night traffic flow picture data with partial labels.
[0046] S2. Build a night bridge traffic detection data enhancement model based on CycleGAN, and use the true daytime traffic flow picture data with complete labels and the true night traffic flow picture data with partial labels to train the night bridge traffic detection data enhancement model to obtain the trained night bridge traffic detection data enhancement model.
[0047] In an optional embodiment of the present invention, the night bridge traffic detection data enhancement model includes a first generator, a second generator, a first discriminator, and a second discriminator. As Figure 2As shown in the figure. A first generator is used to convert the real daytime traffic flow picture data in the first independent training domain into fake night-time traffic flow picture data, and convert the fake daytime traffic flow picture data generated by the second generator into fake night-time traffic flow picture data; the second generator is used to convert the real night-time traffic flow picture data in the second independent training domain into fake daytime traffic flow picture data, and convert the fake night-time traffic flow picture data generated by the first generator into fake daytime traffic flow picture data; the first discriminator is used to discriminate the real night-time traffic flow picture data in the second independent training domain and the fake night-time traffic flow picture data generated by the first generator; the second discriminator is used to discriminate the real daytime traffic flow picture data in the first independent training domain and the fake daytime traffic flow picture data generated by the second generator.
[0048] The present invention uses the real daytime traffic flow picture data with complete labels and the real night-time traffic flow picture data with partial labels to train the night-time bridge traffic detection data enhancement model, and obtains the trained night-time bridge traffic detection data enhancement model, including the following steps:
[0049] S21. Construct a first independent training domain X and a second independent training domain Y.
[0050] In the present invention, the first independent training domain X and the second independent training domain Y are constructed, then the real daytime traffic flow picture data x ∈ X, and the real night-time traffic flow picture data y ∈ Y.
[0051] S22. Input the real daytime traffic flow picture data with complete labels into the first independent training domain, and input the real night-time traffic flow picture data with partial labels into the second independent training domain to train the night-time bridge traffic detection data enhancement model, and obtain the trained night-time bridge traffic detection data enhancement model.
[0052] S3. Perform multi-scale screening on the trained night-time bridge traffic detection data enhancement model to obtain the final night-time bridge traffic detection data enhancement model.
[0053] In an optional embodiment of the present invention, step S3 includes the following steps:
[0054] S31. Use the visual semantic information comparison method to perform preliminary screening on the trained night-time bridge traffic detection data enhancement model to obtain an initial combination of night-time bridge traffic detection data enhancement models.
[0055] The present invention uses a visual semantic information comparison method to preliminarily screen the trained night bridge traffic detection data enhancement model. The specific process is as follows: Determine the visual semantic information, including vehicle appearance types, relative distributions of traffic flow positions, street lights, and isolation belts, and compare the consistency of semantic information between the fake night traffic flow picture data generated by the trained night bridge traffic detection data enhancement model and the real day traffic flow picture data before generation, so as to preliminarily screen the trained night bridge traffic detection data enhancement model.
[0056] S32. Construct a loss function model for night bridge traffic detection data enhancement.
[0057] The loss function model for night bridge traffic detection data enhancement includes the least squares loss function, cycle consistency loss function, and identity mapping function for night bridge traffic detection data enhancement.
[0058] When using traditional cross-entropy loss for training, the situation of gradient disappearance may occur. Especially when the samples generated by the generator are close to the real samples, the feedback of the discriminator becomes weak, thus affecting the learning of the generator. Therefore, for Figure 2 the adversarial loss, the present invention uses the least squares loss function that is the same as the loss of LSGAN and has obvious advantages.
[0059] The present invention considers that only adversarial loss is insufficient because the adversarial loss is only used to train the generator and the discriminator, making the image distribution generated by the generator as close as possible to the real data distribution in the target domain. This optimization path only requires the generated pictures to be more real, but cannot guarantee a clear correspondence between the input and output. For example, for an input image x in domain X with several cars, the generation result G(x) of the generator may be mapped to multiple different samples in domain Y. As long as these samples conform to the statistical distribution of domain Y, even if there is not a single car in G(x), it still meets the training expectation of the adversarial loss. In this case, the training of the generators G and F is unstable, which will lead to mode collapse or semantic inconsistency between the templates. Therefore, the present invention introduces cycle consistency loss. As Figure 2 shown, after a series of generation processes, the pictures x and y are respectively output as F(G(x)) and G(F(y)). The present invention hopes to achieve the following effects during the training process: x≈F(G(x)), y≈G(F(y)), that is, the present invention hopes that the pictures only change the domain distribution but not the content and semantic information of the samples after going through a cycle. The present invention realizes the above constraints through the cycle consistency loss function formula to force the generators G and F to retain the key features of the input samples during the training process, ensuring semantic consistency and mapping uniqueness.
[0060] In addition, the present invention introduces an identity mapping function to ensure that the generators G and F do not perform unnecessary conversions on pictures that already belong to the target domain.
[0061] The least - square loss function for data augmentation in night - time bridge traffic detection is expressed as:
[0062]
[0063] Where: G is the first generator, D Y is the first discriminator, X is the first independent training domain, Y is the second independent training domain, E is the expectation for a certain distribution, y ∼ p data (y) means that y is sampled from the true night - time traffic flow picture data distribution p data (y), x ∼ p data (x) means that x is sampled from the true day - time traffic flow picture data distribution p data (x), D Y (y) is the output value of the first discriminator for the true night - time traffic flow picture data y. The closer the value is to 1, the higher the probability that the first discriminator D Y considers y to be a real sample. D Y (G(x)) is the output value of the first discriminator for the fake night - time traffic flow picture data G(x). F is the second generator, D X is the second discriminator, D x (x) is the output value of the second discriminator for the true day - time traffic flow picture data x. D x (F(y)) is the output value of the second discriminator for the fake day - time traffic flow picture data F(y).
[0064] Compared with the traditional cross - entropy loss, the above - mentioned least - square loss for data augmentation in night - time bridge traffic detection can make the fine - tuning of discriminator parameters smoother, provide a more reasonable optimization path, and thus avoid the problem of gradient disappearance to a certain extent.
[0065] The cycle - consistency loss function for data augmentation in night - time bridge traffic detection is expressed as:
[0066]
[0067] Where: G is the first generator, F is the second generator, E is the expectation for a certain distribution, x ∼ p data (x) means that x is sampled from the true day - time traffic flow picture data distribution p data (x), || || 1 is the L1 norm, y ∼ p data (y) means that y is sampled from the true night - time traffic flow picture data distribution p dataSampling is performed in (y), F(G(x)) is the output value of the second generator for the fake night traffic flow picture data G(x), x is the real day traffic flow picture data, G(F(y)) is the output value of the first generator for the fake day traffic flow picture data F(y), and y is the real night traffic flow picture data.
[0068] The identity mapping function for enhancing night bridge traffic detection data is expressed as:
[0069]
[0070] Where: G is the first generator, F is the second generator, E is the expectation for a certain distribution, and y ~ p data (y) means that y is sampled from the real night traffic flow picture data distribution p data (y), and || || 1 is the L1 norm, and x ~ p data (x) means that x is sampled from the real day traffic flow picture data distribution p data (x), G(y) is the output value of the first generator for the real night traffic flow picture data y, y is the real night traffic flow picture data, F(x) is the output value of the second generator for the real day traffic flow picture data x, and x is the real day traffic flow picture data.
[0071] The present invention makes the mapping of the generator smoother in the two domains by constraining G(y) ≈ y and F(x) ≈ x, ensuring that only the data that needs to be transformed is transformed, while directly outputting the data that is already in the target domain, and avoiding meaningless mapping.
[0072] S33. Use the FID method and the loss function model for enhancing night bridge traffic detection data to finally screen the initial combined model for enhancing night bridge traffic detection data, so as to obtain the final model for enhancing night bridge traffic detection data.
[0073] FID (Fréchet Inception Distance, the evaluation criterion for the diffusion model) calculates the Fréchet distance between two multi-dimensional Gaussian distributions, and these two distributions represent the feature distributions of real images and generated images respectively. The present invention calculates the FID value, which is expressed as:
[0074] FID = ‖μ real - μ fake ‖ 2 + Tr(Σ real + Σ fake - 2(Σ real Σ fake ) 1 / 2 )
[0075] Where: μ real and μfake are the means of the real image and the generated image features, respectively, ∑ real and ∑ fake are the covariance matrices of the real image and the generated image features, respectively, and Tr is the trace of the matrix.
[0076] The smaller the FID value calculated by the present invention, the smaller the Fréchet distance between the generated data and the real data can be determined, and further, the higher the generation quality of the model can be determined. The present invention screens the initial combination of night bridge traffic detection data enhancement models by this method.
[0077] The present invention uses the loss function model for night bridge traffic detection data enhancement to finally screen the initial combination of night bridge traffic detection data enhancement models. The specific process is as follows: Based on the loss function model for night bridge traffic detection data enhancement, draw the loss function curves of each model in the initial combination of night bridge traffic detection data enhancement models, and screen the models whose four loss function values in the loss function curves tend to be stable values in this interval range and have no fluctuations exceeding 0.3 to finally screen the initial combination of night bridge traffic detection data enhancement models.
[0078] S4. Obtain the enhanced night bridge traffic detection data according to the real daytime vehicle flow picture data with complete labels and the final night bridge traffic detection data enhancement model.
[0079] In an optional embodiment of the present invention, the present invention uses the finally screened night bridge traffic detection data enhancement model to convert the real daytime vehicle flow picture data with complete labels, and on the premise of not changing the vehicle shape, position and relative layout of the vehicle flow, convert the image background into a night environment. On this basis, the generated false night vehicle flow image data can be matched one by one with the label file of the daytime source data to complete the overall migration of the labels, so as to obtain the enhanced night bridge traffic detection data.
[0080] Simulation experiment:
[0081] As Figure 3 shown, the present invention uses the YOLOv5 object detection model to evaluate the effect of the method of the present invention on the enhancement of night bridge traffic detection data. Specifically, the present invention uses the pre-trained model based on the open-source YOLOv5s model, and then uses the real daytime vehicle flow picture data and the false night vehicle flow picture data to make different training sets, trains the YOLOv5 object detection model, and calculates the recognition accuracy of each training set on different validation sets to verify the effectiveness of data enhancement.
[0082] The present invention first screens the finally night bridge traffic detection data enhancement model by the above method. As Figure 4As shown, the present invention records the specific values of various loss functions during the training process. As can be seen from the above, the night bridge traffic detection data enhancement model constructed by the present invention presents a cyclic structure. Specifically, there are mainly two directions during the training process of the night bridge traffic detection data enhancement model, which respectively correspond to Figure 3 the left and right parts of Figure 3 . The present invention pays more attention to the conversion from daytime vehicle flow picture data to night-time vehicle flow picture data, because the present invention only focuses on analyzing the Figure 7 left half of Figure 7 . Here, the present invention is denoted as the X-Y part. Through the calculation of the loss function model of night bridge traffic detection data enhancement, the present invention respectively obtains four loss function values, namely the discriminator loss DX-Y, the generator loss GX-Y in the least squares loss function, the cycle consistency loss cycleX-Y in the cycle consistency loss function, and the identity mapping loss idtX-Y in the identity mapping function. It can be seen from Figure 7 that all four loss functions show a gradually decreasing trend before the epoch (number of iterations) is 50. At this time, the model has not converged, which corresponds to the poor effect of the generated data when the epoch is 10. When the epoch is in the range of 50 to 130, the identity mapping loss value idtX-Y is in the range of 0.1 to 0.2, and the function value has no obvious mutation and shows a slightly decreasing trend, which is already at a relatively good level. However, the other three loss function curves do not show an obvious decreasing trend, but instead show fluctuations around a certain value. Among them, the generator loss GX-Y and the cycle consistency loss cycleX-Y fluctuate the most violently, and even show a mutation exceeding 0.5 in value. Combining the results of the visual semantic information comparison method, the present invention can draw the conclusion that in this epoch interval, the fluctuations of the model around a certain reference value reflect better stability than the model with an epoch less than 50, but still do not reach the best level. Such fluctuations are extremely common in deep learning, and the most likely reason is that the model has difficulty learning the complex details of the samples. When the epoch is in the range of 130 to 135, the present invention can find that all four loss functions show a uniform change trend, which is very similar to the change characteristics of the loss function before the common model converges. When the epoch is greater than 135, the present invention observes through Figure 7It is found that the values of the four loss functions have basically stabilized. Among them, the identity mapping loss value, discriminator loss value, and cycle consistency loss value stabilize at around the levels of 0.1, 0.2, and 0.4 respectively, and there are no significant mutations among the three. Although the generator loss value has a stable trend, there are still small-scale mutations. Even so, the magnitude of the mutations is much smaller than that during the period when the epoch is between 50 and 130. Therefore, the present invention can conclude that when the model is trained to epoch = 135, it has basically converged. Therefore, the optimal generation model should be selected from the range of epoch from 135 to 200. Even though we believe that the model has nearly converged at this time, there are still slight differences in the generated data of the models with different weights. Therefore, the present invention then uses the FID index dedicated to evaluating the generation quality of the generation model to further screen out the final night bridge traffic detection data enhancement model. The present invention uses 2000 real daytime vehicle flow picture data as input data, performs domain conversion on the real daytime vehicle flow picture data using the model weights in the range of epoch from 135 to 200, and then calculates the FID values between the generated data of each epoch and 2000 real night vehicle flow picture data. The results are as Figure 5 shown. It can be Figure 5 seen that when the epoch is in the range of 135 to 165, the FID value fluctuates slightly. When the epoch is greater than 165, the FID value has basically stabilized, and the maximum difference in FID values is only 1.751. Looking back at Figure 4 the generator loss function value GX-Y, the present invention finds that the mutations in the range of epoch from 165 to 200 are also slightly smaller than those in the range of epoch from 135 to 165, further reflecting that the model has converged at this stage. Therefore, the present invention selects the weight model of the 200th epoch with the lowest FID value (18.264) as the final night bridge traffic detection data enhancement model.
[0083] Next, the present invention utilized the enhanced model weights of the final night-time bridge traffic detection data to perform domain conversion on the true daytime vehicle flow picture data with complete labels, obtaining 2,000 pieces of false night-time vehicle flow picture data with complete labels. Since the semantic information such as the vehicle models, quantities, and relative position distributions within the data remains unchanged after the style transfer of the original daytime vehicle flow data by the domain conversion model based on the CycleGAN algorithm, the present invention directly transferred the labels of the true daytime vehicle flow picture data as a whole. After simply correcting some invalid labels that were not conducive to YOLOv5 training, they were directly used as false night-time vehicle flow picture data with complete labels, greatly saving the manual annotation cost of night-time vehicle flow data. Based on the above data, the present invention produced three training sets, named the true daytime training data set, the false night-time training set, and the mixed set of the true daytime training data set and the false night-time training set, with the total amounts of data for the three being 2,000, 2,000, and 4,000 respectively.
[0084] The present invention randomly selected 200 pieces of vehicle flow data from the true daytime vehicle flow picture data and the true night-time vehicle flow picture data respectively, and produced three validation sets (Validation Set A, Validation Set B, and Validation Set C) for verifying the accuracy of the YOLOv5 model with these. Among them, Validation Set A was 200 pieces of true night-time vehicle flow picture data, Validation Set B was 400 pieces of mixed true day and night vehicle flow picture data, and Validation Set C was 200 pieces of true daytime vehicle flow picture data, and the true daytime vehicle flow picture data in Validation Set C were all from Validation Set B. And the maximum value of mAP50 during the training process was recorded in real time. When the mAP50 values of the last 100 epochs did not exceed the current maximum mAP50 value, the present invention determined that YOLOv5 had converged, and recorded the accuracy evaluation index values of each working condition at this time, such as Figure 6 、 Figure 7 and Figure 8 as shown.
[0085] The present invention focused on the F1 value and the mAP50 value of each training set under different validation sets. In the comparison, the present invention considered that when both of the above were higher than the other, it was determined that the model was more excellent. From Figure 6It can be seen that when using the real daytime traffic flow picture dataset as the training set and validation set A as the validation set, the F1 value and mAP50 value are 0.707 and 0.721 respectively, which are at a relatively low level. This reflects the low recognition performance of the pre-trained model using real daytime traffic flow data in the night scene, further demonstrating the necessity of data augmentation. When using the fake night traffic flow dataset and the mixed set of real daytime training dataset and fake night training set as the training set, the F1 value and mAP50 value are 0.772 / 0.805 and 0.811 / 0.835 respectively, which are 9.2% / 13.9% and 12.5% / 15.8% higher than those when using the real daytime traffic flow picture dataset as the training set. Therefore, we conclude that the fake night traffic flow picture data after domain conversion using the method provided by the present invention can provide more valuable feature information that conforms to the real night traffic flow scene compared with the real daytime traffic flow data, and can effectively improve the recognition accuracy of the target detection model in the night scene. In addition, when training the model by mixing the real daytime traffic flow picture data and the fake night traffic flow picture data, the detection accuracy can be further improved.
[0086] However, in reality, it is impossible to only focus on the recognition accuracy of night traffic flow. Therefore, the present invention also uses validation set B, which is more challenging for the target detection model, as the validation set, and trains YOLOv5 using the above three training sets respectively. Figure 7 It can be seen that when using the mixed set of real daytime training dataset and fake night training set as the training set, the F1 value and mAP50 value are 0.835 and 0.859 respectively, which are 1.7% and 2.3% higher than those when using the real daytime training dataset as the training set, indicating that there is still a data augmentation effect when using the mixed set of real daytime training dataset and fake night training set as the training set. However, when using the fake night training set as the training set, the F1 value and mAP50 value are 0.779 and 0.796 respectively, which are 5.1% and 5.2% lower than those when using the real daytime training dataset as the training set. Since when using validation set A as the validation set, the recognition accuracy of the model under the fake night training set has been significantly improved (F1: 9.2% and mAP50: 13.9%), but it shows a performance decline on validation set B. This is because when using the fake night training set as the training set, the recognition accuracy of the real daytime training dataset is greatly sacrificed while the accuracy of the night traffic flow scene is improved. Such a phenomenon is normal because there are huge differences in features between day and night scenes. When using real daytime picture data as the training set to verify the night scene and when using fake night picture data as the training set to verify the day scene, the decline in the recognition accuracy of the target detection model is inevitable.
[0087] To further verify the above conjecture, the present invention extracted 200 pieces of daytime traffic flow data from the validation set B to create a validation set C consisting entirely of daytime scenes, and trained YOLOv5 using the three training sets again. As Figure 8 can be seen, when using the fake night training set as the training set, the F1 value and the mAP50 value are 0.793 and 0.814 respectively, which are 8.4% and 8.8% lower than when using the real daytime training data set as the training set, verifying the above conjecture. However, when using the mixed set of the real daytime training data set and the fake night training set as the training set, its F1 value and mAP50 value are 0.866 and 0.895 respectively, showing almost no change compared with when using the real daytime training data set as the training set. This indicates that when using the mixed set of the real daytime training data set and the fake night training set as the training set, it will not affect the accuracy of the model in the daytime traffic flow scenario.
[0088] In summary, in terms of the feasibility of data augmentation, the present invention draws the following conclusions: (1) The method provided by the present invention for generating fake night traffic flow picture data successfully preserves key semantic information such as vehicle type, quantity, and relative position. This ensures that the annotations of the original daytime data set can be directly adjusted with only a small number of corrections, thus greatly reducing the manual labeling cost of night traffic flow picture data. (2) When training the YOLOv5 model on the fake night training set or the mixed set and validating on the validation set A, both F1 and mAP50 are significantly improved. Specifically, the mixed set achieves a maximum improvement of 13.9% in F1 and 15.8% in mAP50, highlighting the value of the augmented data obtained by the method of the present invention in improving the accuracy of the night scene model. (3) When the target domain involved contains the daytime traffic flow scenario, using only the fake night training set for training will reduce the recognition accuracy of object detection, while using the mixed set for training can not only ensure that the recognition accuracy is not affected in the daytime traffic flow scenario, but also, under the validation set B, can increase the F1 value and the mAP50 value by 1.7% and 2.3% respectively, showing strong generalization ability. (4) Therefore, although the augmented data obtained by the present invention performs excellently in the augmentation (night) of a specific domain, the mixed set can ensure robust performance in different scenarios.
[0089] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a means for implementing the functions specified in one flow Figure 1 or more flows and / or blocks Figure 1 or a means for implementing the functions specified in one block or more blocks.
[0090] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction means that implements the functions specified in one flow Figure 1 or more flows and / or blocks Figure 1 or a means for implementing the functions specified in one block or more blocks.
[0091] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 or more flows and / or blocks Figure 1 or a means for implementing the functions specified in one block or more blocks.
[0092] Specific embodiments are used in the present invention to elaborate on the principles and implementation manners of the present invention. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
[0093] Those of ordinary skill in the art will realize that the embodiments described herein are for helping the reader understand the principles of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not depart from the essence of the present invention according to the technical revelations disclosed in the present invention, and these deformations and combinations are still within the protection scope of the present invention.
Claims
1. A nighttime bridge traffic detection data enhancement method based on CycleGAN, characterized in that: The following steps are involved: S1. Collect daytime traffic image data and nighttime traffic image data of bridge traffic, fully annotate the daytime traffic image data to obtain true daytime traffic image data with complete labels, and partially annotate the nighttime traffic image data to obtain true nighttime traffic image data with partial labels; S2. Building a nighttime bridge traffic detection data enhancement model based on CycleGAN, using real daytime traffic image data with complete labels and real nighttime traffic image data with partial labels to train the nighttime bridge traffic detection data enhancement model, and obtaining the trained nighttime bridge traffic detection data enhancement model; S3, performing multi-scale screening on the trained nighttime bridge traffic detection data enhancement model to obtain the final nighttime bridge traffic detection data enhancement model; S4. Obtain enhanced nighttime bridge traffic detection data based on the real daytime traffic image data with complete labels and the final nighttime bridge traffic detection data enhancement model.
2. The method for nighttime bridge traffic detection data enhancement based on CycleGAN according to claim 1 is characterized in that: In step S2, the nighttime bridge traffic detection data enhancement model includes a first generator, a second generator, a first discriminator, and a second discriminator; the first generator is used to convert the real daytime traffic image data of the first independent training domain into fake nighttime traffic image data, and convert the holiday traffic image data generated by the second generator into fake nighttime traffic image data; The second generator is used to convert the real nighttime traffic image data of the second independent training domain into holiday traffic image data, and convert the fake nighttime traffic image data generated by the first generator into holiday traffic image data; the first discriminator is used to discriminate the real nighttime traffic image data of the second independent training domain and the fake nighttime traffic image data generated by the first generator; The second discriminator is used to discriminate between the real daytime traffic flow picture data of the first independent training domain and the holiday traffic flow picture data generated by the second generator.
3. The method for nighttime bridge traffic detection data enhancement based on CycleGAN according to claim 1 is characterized in that: The nighttime bridge traffic detection data enhancement model is trained using real daytime traffic image data with complete labels and real nighttime traffic image data with partial labels to obtain the trained nighttime bridge traffic detection data enhancement model, including the following steps: S21, constructing a first independent training domain and a second independent training domain; S22. Input the real daytime traffic image data with complete labels into the first independent training domain, and input the real nighttime traffic image data with partial labels into the second independent training domain, so as to train the nighttime bridge traffic detection data enhancement model and obtain the trained nighttime bridge traffic detection data enhancement model.
4. The method for nighttime bridge traffic detection data enhancement based on CycleGAN according to claim 1 is characterized in that: Step S3 includes the following steps: S31. Preliminarily screen the trained nighttime bridge traffic detection data enhancement model using a visualized semantic information comparison method to obtain an initial nighttime bridge traffic detection data enhancement model combination; S32. Construct a loss function model for nighttime bridge traffic detection data enhancement; S33. Use the FID method and the loss function model of nighttime bridge traffic detection data enhancement to perform final screening on the initial nighttime bridge traffic detection data enhancement model combination to obtain the final nighttime bridge traffic detection data enhancement model.
5. The method for nighttime bridge traffic detection data enhancement based on CycleGAN according to claim 4 is characterized in that: In step S31, a visualized semantic information comparison method is used to perform a preliminary screening on the trained nighttime bridge traffic detection data enhancement model. The specific process is: determining visualized semantic information, including vehicle appearance type, relative distribution of traffic flow positions, street lights and isolation belts, and performing a semantic information consistency comparison between the fake nighttime traffic flow picture data generated by the trained nighttime bridge traffic detection data enhancement model and the real daytime traffic flow picture data before generation, so as to perform a preliminary screening on the trained nighttime bridge traffic detection data enhancement model.
6. The method for nighttime bridge traffic detection data enhancement based on CycleGAN according to claim 4 is characterized in that: In step S32, the loss function model for nighttime bridge traffic detection data enhancement includes a least squares loss function, a cycle consistency loss function, and an identity mapping function for nighttime bridge traffic detection data enhancement.
7. The method for nighttime bridge traffic detection data enhancement based on CycleGAN according to claim 6 is characterized in that: The least squares loss function for nighttime bridge traffic detection data enhancement is expressed as: Among them: G is the first generator, D Y is the first discriminator, X is the first independent training domain, Y is the second independent training domain, E is the expectation of a certain distribution, y~p data (y) is the distribution of y from real nighttime traffic image data p data (y) is sampled, x~p data (x) is the distribution of x from real daytime traffic image data p data (x) is sampled, D Y (y) is the output value of the first discriminator for the real nighttime traffic image data y. The closer the value is to 1, the better the first discriminator D Y The higher the probability that y is a true sample, the higher the probability that D Y (G(x)) is the output value of the first discriminator for the fake nighttime traffic image data G(x), F is the second generator, D X is the second discriminator, D x (x) is the output value of the second discriminator for the real daytime traffic image data x, D x (F(y)) is the output value of the second discriminator for the holiday traffic image data F(y).
8. The method for nighttime bridge traffic detection data enhancement based on CycleGAN according to claim 6 is characterized in that: The cycle consistency loss function for nighttime bridge traffic detection data enhancement is expressed as: Among them: G is the first generator, F is the second generator, E is the expectation of a certain distribution, x~p data (x) is the distribution of x from real daytime traffic image data p data (x) is sampled, || ||1 is the L1 norm, y~p data (y) is the distribution of y from real nighttime traffic image data p data (y), F(G(x)) is the output value of the second generator for the fake nighttime traffic picture data G(x), x is the real daytime traffic picture data, G(F(y)) is the output value of the first generator for the holiday traffic picture data F(y), y is the real nighttime traffic picture data.
9. The method for nighttime bridge traffic detection data enhancement based on CycleGAN according to claim 6, characterized in that: The identity mapping function for nighttime bridge traffic detection data enhancement is expressed as: Among them: G is the first generator, F is the second generator, E is the expectation of a certain distribution, y~p data (y) is the distribution of y from real nighttime traffic image data p data (y) is sampled, || ||1 is the L1 norm, x~p data (x) is the distribution of x from real daytime traffic image data p data (x), G(y) is the output value of the first generator for the real nighttime traffic picture data y, y is the real nighttime traffic picture data, F(x) is the output value of the second generator for the real daytime traffic picture data x, x is the real daytime traffic picture data.
10. The method for nighttime bridge traffic detection data enhancement based on CycleGAN according to claim 4, characterized in that: In step S33, the loss function model of night bridge traffic detection data enhancement is used to perform a final screening on the initial night bridge traffic detection data enhancement model combination. The specific process is: based on the loss function model of night bridge traffic detection data enhancement, the loss function curve of each model in the initial night bridge traffic detection data enhancement model combination is drawn, and the models whose four loss function values in the loss function curve tend to be stable values within the range and have no fluctuations exceeding 0.3 are screened, so as to perform a final screening on the initial night bridge traffic detection data enhancement model combination.
Citation Information
Patent Citations
Unmanned vehicle reinforcement learning training environment construction method and training system thereof
CN111795700A
Cross-domain vehicle detection method based on CycleGAN generative adversarial network
CN113822248A
Method for quickly calculating mask near field by using cyclic consistent adversarial network
CN115047721A
Cited By
A computer vision-based strong anti-interference bridge deformation detection method
CN122636626A