Method and apparatus for testing the robustness of a vehicle detection model based on deep learning
By establishing a parameterized texture generation network model and an adversarial generation network, and generating adversarial sample pictures, the inadequate robustness of the deep learning vehicle detection model in data set dependence and overfitting problems is solved, and the robustness and accuracy of the model are improved.
Patent Information
- Application Number
- CN202011615690.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-30
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2040-12-30
AI Technical Summary
Deep learning based on vehicle detection model has insufficient robustness in dataset dependence and overfitting problems, resulting in a high rate of error detection when the test data and training data are not independently distributed.
By establishing a parameterized texture generation network model, combining deep neural network, triangular mesh model and deep convolution generation network, parametric texture pictures are generated, and adversarial sample pictures are generated through adversarial generation network, which is used to test the robustness of vehicle detection models.
It improves the robustness and accuracy of the vehicle detection model, and can more effectively handle the non-independent distribution of test data and training data.
Smart Images

Figure CN112766311B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning training, and particularly to a method and device for testing the robustness of a vehicle detection model based on deep learning. Background Art
[0002] Deep learning is the mainstream method in current computer vision. In vehicle detection and recognition tasks, most algorithms adopt deep learning algorithms based on convolutional neural networks. This type of method is highly dependent on the data set and is prone to overfitting. Therefore, if the test data and the training data are not independently distributed, the classifier is likely to misdetect. Therefore, in the training process, in order to improve the accuracy of the deep learning algorithm, it is necessary to add some synthetic images to the training set, and further improve the robustness of the vehicle detection model through the discrimination and feedback of the synthetic images.
[0003] In view of this, it is very meaningful to establish a method and device for testing the robustness of a vehicle detection model based on deep learning. Summary of the Invention
[0004] In view of the problems such as overfitting existing in the vehicle detection model in the above-mentioned prior art that depends on the deep learning data set. The purpose of the embodiments of the present application is to propose a method and device for testing the robustness of a vehicle detection model based on deep learning to solve the technical problems mentioned in the above background art section.
[0005] In a first aspect, the embodiments of the present application provide a method for testing the robustness of a vehicle detection model based on deep learning, including the following steps:
[0006] A data acquisition step of acquiring vehicle pictures in a road traffic scene and establishing a corresponding CAD model according to the vehicle pictures;
[0007] A training step of a parametric texture generation network model of establishing and training a parametric texture generation network model, where the parametric texture generation network model includes a deep neural network, a triangular mesh model, and a deep convolutional generation network. Feature extraction is performed on the vehicle pictures through the deep neural network to obtain a first feature, feature extraction is performed on the CAD model through the triangular mesh model to obtain a second feature, the first feature and the second feature are connected as the input of the deep convolutional generation network, and the output of the deep convolutional generation network is a parametric texture picture;
[0008] An appearance rendering step of collecting two-dimensional pictures corresponding to the CAD models of vehicles of a specified model, inputting the two-dimensional pictures and their corresponding CAD models into the parametric texture generation network model to obtain corresponding parametric texture pictures, calculating the image mean of the parametric texture pictures, adding noise to the image mean and combining with the CAD model to obtain synthetic vehicle pictures;
[0009] Discrimination step: Input the synthesized vehicle images into the vehicle classification network to obtain discrimination results. Record the parametric texture images corresponding to all the images classified as non-vehicles in the discrimination results, and model the parametric texture images corresponding to the non-vehicle images to obtain the final adversarial sample images; and
[0010] Testing step: Input the vehicle images with the adversarial sample images pasted on the vehicle body into the vehicle detection model to be tested, obtain the number of vehicles detected by the vehicle detection model, and calculate the robustness representing the robustness of the vehicle detection model based on the number.
[0011] In some embodiments, the vehicle images in the data acquisition step are obtained by analyzing the positional relationship between the monitoring camera, the road, and the vehicle in the road traffic scene. Collecting a large number of vehicle images in the road traffic scene can obtain implicit features related to the road traffic scene.
[0012] In some embodiments, the parametric texture generation network model training step further includes: Combining the parametric texture images with the CAD model for texture mapping to obtain the vehicle images after texture mapping, performing backpropagation update according to the reconstruction error loss between the input vehicle images and the vehicle images after texture mapping, and adjusting the parameters of the parametric texture generation network model to realize the training of the parametric texture generation network model. Training the parametric texture generation network model with the images after texture mapping to obtain the trained parametric texture generation network model.
[0013] In some embodiments, the modeling in the discrimination step uses a mixture Gaussian model with k degrees of freedom to model the parametric texture images corresponding to the non-vehicle images. Modeling the parametric texture images can obtain the final adversarial samples.
[0014] In some embodiments, the means of the k mixture Gaussian models obtained after modeling are used as k adversarial sample images.
[0015] In some embodiments, the CAD model is a three-dimensional model. It is convenient to map the parametric texture images on the three-dimensional CAD models of each vehicle model.
[0016] In some embodiments, the noise refers to Gaussian noise, the mean of the Gaussian noise is 0, and the covariance matrix is the identity matrix. The appearance of the vehicle is gradually added with noise from the existing non-vehicle images.
[0017] In some embodiments, the formula for calculating the robustness r of the vehicle detection model is
[0018] r = m / k;
[0019] Among them, m is the number of vehicles detected by the vehicle detection model, and k is the number of adversarial sample images. The larger the value of r, the stronger the robustness of the tested vehicle detection model.
[0020] In a second aspect, an embodiment of the present application further provides a device for testing the robustness of a vehicle detection model based on deep learning, including:
[0021] A data acquisition module, configured to acquire vehicle images in a road traffic scene and establish corresponding CAD models according to the vehicle images;
[0022] A parameterized texture generation network model training module, configured to establish and train a parameterized texture generation network model. The parameterized texture generation network model includes a deep neural network, a triangular mesh model, and a deep convolutional generative network. The first feature is obtained by extracting features from the vehicle images through the deep neural network, and the second feature is obtained by extracting features from the CAD model through the triangular mesh model. The first feature and the second feature are connected as the input of the deep convolutional generative network, and the output of the deep convolutional generative network is a parameterized texture image;
[0023] An appearance rendering module, configured to collect two-dimensional images corresponding to the CAD models of vehicles of a specified model, input the two-dimensional images and their corresponding CAD models into the parameterized texture generation network model to obtain corresponding parameterized texture images, calculate the image mean value of the parameterized texture images, add noise to the image mean value and combine it with the CAD model to obtain synthetic vehicle images;
[0024] A discrimination module, configured to input the synthetic vehicle images into a vehicle classification network to obtain discrimination results, record the parameterized texture images corresponding to all the images classified as non-vehicles in the discrimination results, and model the parameterized texture images corresponding to the non-vehicle images to obtain final adversarial sample images; and
[0025] A testing module, configured to input vehicle images with adversarial sample images pasted on the vehicle body into the tested vehicle detection model to obtain the number of vehicles detected by the vehicle detection model, and calculate the robustness degree representing the robustness of the vehicle detection model according to the number.
[0026] In a third aspect, an embodiment of the present application provides an electronic device, including: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner in the first aspect.
[0027] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in any implementation manner in the first aspect is implemented.
[0028] The present invention discloses a method and device for testing the robustness of a vehicle detection model based on deep learning. By acquiring vehicle pictures in a road traffic scene and establishing a corresponding CAD model according to the vehicle pictures; establishing and training a parametric texture generation network model, which includes a deep neural network, a triangular mesh model, and a deep convolutional generation network. The deep neural network extracts features from the vehicle pictures to obtain a first feature, and the triangular mesh model extracts features from the CAD model to obtain a second feature. The first feature and the second feature are connected as the input of the deep convolutional generation network, and the output of the deep convolutional generation network is a parametric texture picture; collecting two-dimensional pictures corresponding to the CAD models of vehicles of a specified model, and inputting the two-dimensional pictures and their corresponding CAD models into the parametric texture generation network model to obtain corresponding parametric texture pictures, and calculating the image mean of the parametric texture pictures. Adding noise to the image mean and combining it with the CAD model to obtain a synthetic vehicle picture; inputting the synthetic vehicle picture into a vehicle classification network to obtain a discrimination result, recording the parametric texture pictures corresponding to all the pictures classified as non-vehicles in the discrimination result, and modeling the parametric texture pictures corresponding to the non-vehicle pictures to obtain a final adversarial sample picture; inputting the vehicle picture with the adversarial sample picture pasted on the vehicle body into the vehicle detection model to be tested, obtaining the number of vehicles detected by the vehicle detection model, and calculating the robustness representing the robustness of the vehicle detection model according to the number. By using the parametric texture generation network model to obtain parametric texture pictures, further obtaining synthetic vehicle pictures, classifying the synthetic pictures through a vehicle classification network nested in an adversarial generation network to obtain adversarial sample pictures, and using the adversarial sample pictures to calculate the robustness of the vehicle detection model, the accuracy of the vehicle detection model is improved. Description of the Drawings
[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0030] Figure 1 It is an exemplary device architecture diagram to which an embodiment of the present application can be applied;
[0031] Figure 2 It is a flowchart of the method for testing the robustness of a vehicle detection model based on deep learning according to the embodiment of the present invention;
[0032] Figure 3Schematic diagram of an apparatus for testing the robustness of a deep learning-based vehicle detection model according to an embodiment of the present invention;
[0033] Figure 4 It is a schematic structural diagram of a computer device of an electronic device suitable for implementing the embodiments of the present application. Detailed implementation manners
[0034] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0035] Figure 1 An exemplary device architecture 100 is shown that can apply the method for testing the robustness of a deep learning-based vehicle detection model or the apparatus for testing the robustness of a deep learning-based vehicle detection model according to the embodiments of the present application.
[0036] As Figure 1 shown, the device architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0037] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various applications may be installed on the terminal devices 101, 102, 103, such as data processing applications, file processing applications, etc.
[0038] The terminal devices 101, 102, 103 may be hardware or software. When the terminal devices 101, 102, 103 are hardware, they may be various electronic devices, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc. When the terminal devices 101, 102, 103 are software, they may be installed in the above-listed electronic devices. It may be implemented as multiple software or software modules (such as software or software modules for providing distributed services), or may be implemented as a single software or software module. No specific limitation is made here.
[0039] The server 105 may be a server that provides various services, such as a background data processing server that processes files or data uploaded by the terminal devices 101, 102, and 103. The background data processing server may process the acquired files or data to generate a processing result.
[0040] It should be noted that the method for testing the robustness of a vehicle detection model based on deep learning provided by the embodiments of the present application may be executed by the server 105, or may be executed by the terminal devices 101, 102, and 103. Correspondingly, the device for testing the robustness of a vehicle detection model based on deep learning may be set in the server 105, or may be set in the terminal devices 101, 102, and 103.
[0041] It should be understood that Figure 1 the numbers of the terminal devices, the network, and the server in
[0042] Figure 2 The embodiments of the present application disclose a method for testing the robustness of a vehicle detection model based on deep learning, including the following steps:
[0043] Step S1, obtain vehicle pictures in a road traffic scene, and establish a corresponding CAD model according to the vehicle pictures.
[0044] In a specific embodiment, the vehicle pictures in step S1 are obtained by analyzing the positional relationship among the monitoring cameras, the road, and the vehicles in the road traffic scene. Collecting a large number of vehicle pictures in the road traffic scene can obtain implicit features related to the road traffic scene. And at the same time, establish a three-dimensional CAD model corresponding to each vehicle model. It is required that each vehicle picture has a corresponding CAD model, denoted as <image_i, CAD_i>, where i represents the i-th vehicle picture, and CAD_i represents the CAD model corresponding to the i-th vehicle picture. In a specific embodiment, the CAD model is a three-dimensional model. It is convenient to map the parameterized texture pictures onto the three-dimensional CAD models of each vehicle model.
[0045] Step S2, establish a parameterized texture generation network model and train it. The parameterized texture generation network model includes a deep neural network, a triangular mesh model, and a deep convolutional generative network. Extract features from the vehicle pictures through the deep neural network to obtain a first feature, extract features from the CAD model through the triangular mesh model to obtain a second feature, connect the first feature and the second feature as the input of the deep convolutional generative network, and the output of the deep convolutional generative network is a parameterized texture picture.
[0046] In a specific embodiment, the working mode among the three sub-networks in the parametric texture generation network model is as follows:
[0047] For the input vehicle image image_i, the corresponding implicit feature fi, i.e., the first feature, is obtained through the deep neural network, and the feature Fi, i.e., the second feature, is obtained after the CAD_i passes through the triangular mesh model. The Fi and fi features are concatenated as the input of the deep convolutional generation network, and the output of the deep convolutional generation network is the parametric texture image.
[0048] In a specific embodiment, the training steps of the parametric texture generation network model further include: combining the parametric texture image with the CAD model for texture mapping to obtain the vehicle image after texture mapping, performing backpropagation update according to the reconstruction error loss between the input vehicle image and the vehicle image after texture mapping, and adjusting the parameters of the parametric texture generation network model to realize the training of the parametric texture generation network model. The parametric texture generation network model is trained with the image after texture mapping to obtain the trained parametric texture generation network model.
[0049] Step S3: Collect the two-dimensional images corresponding to the CAD models of the vehicles of the specified model, and input the two-dimensional images and their corresponding CAD models into the parametric texture generation network model to obtain the corresponding parametric texture images, calculate the image mean of the parametric texture images, add noise to the image mean, and combine it with the CAD model to obtain the synthetic vehicle images.
[0050] In a specific embodiment, after the parametric texture generation network model is trained, for the vehicles of a certain model, first collect the two-dimensional images corresponding to its CAD model, denoted as CAD-IMAGES-OLD. Input the images of CAD-IMAGES-OLD into the deep neural network for feature extraction to obtain the first feature, input the CAD model into the triangular mesh model for feature extraction to obtain the second feature, connect the first feature and the second feature as the input of the deep convolutional generation network to obtain the corresponding parametric texture image, and calculate the mean of the parametric texture image, denoted as PARA-IMAGE-MEAN. Add noise to PARA-IMAGE-MEAN and generate n synthetic vehicle images based on its three-dimensional CAD model. In a specific embodiment, the noise refers to Gaussian noise, the mean of the Gaussian noise is 0, and the covariance matrix is the identity matrix. The appearance of the vehicle is gradually obtained by adding noise to the existing non-vehicle images.
[0051] Step S4: Input the synthesized vehicle images into the vehicle classification network to obtain the discrimination results. Record the parametric texture images corresponding to all the images classified as non-vehicles in the discrimination results, and model the parametric texture images corresponding to the non-vehicle images to obtain the final adversarial sample images.
[0052] In a specific embodiment, input the n synthesized vehicle images generated in step S3 into a known vehicle classification network, record the parametric texture images corresponding to all the images classified as non-vehicles, and model the parametric texture images corresponding to the non-vehicle images using a Gaussian mixture model with k degrees of freedom. The means of the k Gaussian mixture models after modeling are used as the final k adversarial sample images. After pasting the adversarial sample images on the vehicle surface, the known vehicle classification network will classify them as negative samples with a high probability.
[0053] Step S5: Input the vehicle images with the adversarial sample images pasted on the vehicle body into the vehicle detection model to be tested, obtain the number of vehicles detected by the vehicle detection model, and calculate the robustness representing the robustness of the vehicle detection model according to the number.
[0054] In a specific embodiment, the calculation formula for the robustness r of the vehicle detection model is
[0055] r = m / k;
[0056] where m is the number of vehicles detected by the vehicle detection model, and k is the number of adversarial sample images. The larger the value of r, the stronger the robustness of the tested vehicle detection model.
[0057] Further referring to Figure 3 , as an implementation of the methods shown in the above figures, an embodiment of a device for testing the robustness of a vehicle detection model based on deep learning is provided in the present application. This device embodiment corresponds to the Figure 2 shown method embodiment, and this device can be specifically applied to various electronic devices.
[0058] An embodiment of a device for testing the robustness of a vehicle detection model based on deep learning proposed in the embodiments of the present application includes:
[0059] A data acquisition module 1, configured to acquire vehicle images in a road traffic scene and establish a corresponding CAD model according to the vehicle images;
[0060] The parametric texture generation network model training module 2 is configured to establish and train a parametric texture generation network model. The parametric texture generation network model includes a deep neural network, a triangular mesh model, and a deep convolutional generative network. The deep neural network extracts features from vehicle pictures to obtain a first feature, and the triangular mesh model extracts features from the CAD model to obtain a second feature. The first feature and the second feature are concatenated as the input of the deep convolutional generative network, and the output of the deep convolutional generative network is a parametric texture picture;
[0061] The appearance rendering module 3 is configured to collect two-dimensional pictures corresponding to the CAD models of vehicles of a specified model, and input the two-dimensional pictures and their corresponding CAD models into the parametric texture generation network model to obtain corresponding parametric texture pictures, and calculate the image mean value of the parametric texture pictures. Noise is added to the image mean value and combined with the CAD model to obtain a synthetic vehicle picture;
[0062] The discrimination module 4 is configured to input the synthetic vehicle picture into a vehicle classification network to obtain a discrimination result, record the parametric texture pictures corresponding to all the pictures classified as non-vehicles in the discrimination result, and model the parametric texture pictures corresponding to the non-vehicle pictures to obtain a final adversarial sample picture; and
[0063] The testing module 5 is configured to input a vehicle picture with an adversarial sample picture pasted on the vehicle body into the vehicle detection model to be tested, obtain the number of vehicles detected by the vehicle detection model, and calculate the robustness representing the robustness of the vehicle detection model according to the number.
[0064] The present invention discloses a method and device for testing the robustness of a vehicle detection model based on deep learning. By obtaining vehicle pictures in a road traffic scene and establishing a corresponding CAD model according to the vehicle pictures; establishing and training a parametric texture generation network model, which includes a deep neural network, a triangular mesh model, and a deep convolutional generation network. The deep neural network extracts features from the vehicle pictures to obtain a first feature, and the triangular mesh model extracts features from the CAD model to obtain a second feature. The first feature and the second feature are connected as the input of the deep convolutional generation network, and the output of the deep convolutional generation network is a parametric texture picture; collecting two-dimensional pictures corresponding to the CAD models of vehicles of a specified model, and inputting the two-dimensional pictures and their corresponding CAD models into the parametric texture generation network model to obtain corresponding parametric texture pictures, and calculating the image mean of the parametric texture pictures, adding noise to the image mean and combining with the CAD model to obtain a synthetic vehicle picture; inputting the synthetic vehicle picture into a vehicle classification network to obtain a discrimination result, recording the parametric texture pictures corresponding to all the pictures classified as non-vehicles in the discrimination result, and modeling the parametric texture pictures corresponding to the non-vehicle pictures to obtain a final adversarial sample picture; inputting the vehicle picture with the adversarial sample picture pasted on the vehicle body into the vehicle detection model to be tested to obtain the number of vehicles detected by the vehicle detection model, and calculating the robustness degree representing the robustness of the vehicle detection model according to the number. By using the parametric texture generation network model to obtain parametric texture pictures, further obtaining synthetic vehicle pictures, classifying the synthetic pictures through the vehicle classification network nested in the adversarial generation network to obtain adversarial sample pictures, and using the adversarial sample pictures to calculate the robustness of the vehicle detection model, the accuracy of the vehicle detection model is improved.
[0065] Reference is made below to Figure 4 , which shows a schematic structural diagram of a computer device 400 of an electronic device (such as Figure 1 the server or terminal device shown) suitable for use in implementing the embodiments of the present application. Figure 4 The electronic device shown is merely an example and should not impose any limitation on the functions and scope of use of the embodiments of the present application.
[0066] As Figure 4As shown, the computer device 400 includes a central processing unit (CPU) 401 and a graphics processing unit (GPU) 402, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 403 or a program loaded from a storage section 409 into a random access memory (RAM) 404. In the RAM 404, various programs and data required for the operation of the device 400 are also stored. The CPU 401, GPU 402, ROM 403, and RAM 404 are connected to each other via a bus 405. An input / output (I / O) interface 406 is also connected to the bus 405.
[0067] The following components are connected to the I / O interface 406: an input section 407 including a keyboard, a mouse, etc.; an output section 408 including, for example, a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 409 including a hard disk, etc.; and a communication section 410 including a network interface card such as a LAN card, a modem, etc. The communication section 410 performs communication processing via a network such as the Internet. A drive 411 may also be connected to the I / O interface 406 as needed. A removable medium 412, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 411 as needed so that a computer program read from it can be installed into the storage section 409 as needed.
[0068] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 410, and / or installed from the removable medium 412. When the computer program is executed by the central processing unit (CPU) 401 and the graphics processing unit (GPU) 402, the above functions defined in the method of the present application are executed.
[0069] It should be noted that the computer-readable medium described in this application can be a computer-readable signal medium, a computer-readable medium, or any combination of the two. The computer-readable medium can be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor devices, apparatuses, or components, or any combination of the above. More specific examples of the computer-readable medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, the computer-readable medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution device, apparatus, or component. And in this application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution device, apparatus, or component. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0070] The computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages - such as Java, Smalltalk, C++, and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or, alternatively, can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0071] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of devices, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based device that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0072] The modules described in the embodiments of the present application can be implemented in software or in hardware. The described modules can also be provided in a processor.
[0073] As another aspect, the present application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or may exist separately without being assembled into the electronic device. The above computer-readable medium carries one or more programs. When the above one or more programs are executed by the electronic device, the electronic device is enabled to: obtain vehicle pictures in a road traffic scenario and establish corresponding CAD models according to the vehicle pictures; establish and train a parametric texture generation network model, which includes a deep neural network, a triangular mesh model, and a deep convolutional generation network. The deep neural network extracts features from the vehicle pictures to obtain a first feature, and the triangular mesh model extracts features from the CAD models to obtain a second feature. The first feature and the second feature are connected as the input of the deep convolutional generation network, and the output of the deep convolutional generation network is a parametric texture picture; collect two-dimensional pictures corresponding to the CAD models of vehicles of a specified model, and input the two-dimensional pictures and their corresponding CAD models into the parametric texture generation network model to obtain corresponding parametric texture pictures, and calculate the image mean of the parametric texture pictures, add noise to the image mean and combine it with the CAD model to obtain a synthetic vehicle picture; input the synthetic vehicle picture into a vehicle classification network to obtain a discrimination result, record the parametric texture pictures corresponding to all the pictures classified as non-vehicles in the discrimination result, and model the parametric texture pictures corresponding to the non-vehicle pictures to obtain a final adversarial sample picture; input the vehicle picture with the adversarial sample picture pasted on the vehicle body into the vehicle detection model to be tested, obtain the number of vehicles detected by the vehicle detection model, and calculate the robustness representing the robustness of the vehicle detection model according to the number.
[0074] The above description is only a preferred embodiment of the present application and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features having similar functions disclosed in the present application.
Claims
1. A method for testing the robustness of a vehicle detection model based on deep learning, characterized in that, it includes the following steps: A data acquisition step of acquiring vehicle pictures in a road traffic scenario and establishing a corresponding CAD model according to the vehicle pictures; A training step of a parametric texture generation network model, establishing and training a parametric texture generation network model, the parametric texture generation network model includes a deep neural network, a triangular mesh model and a deep convolutional generation network, extracting features from the vehicle pictures through the deep neural network to obtain a first feature, extracting features from the CAD model through the triangular mesh model to obtain a second feature, connecting the first feature and the second feature as the input of the deep convolutional generation network, and the output of the deep convolutional generation network is a parametric texture picture; An appearance rendering step of collecting two-dimensional pictures corresponding to the CAD models of vehicles of a specified model, inputting the two-dimensional pictures and their corresponding CAD models into the parametric texture generation network model to obtain the corresponding parametric texture pictures, calculating the image mean value of the parametric texture pictures, adding noise to the image mean value and combining with the CAD model to obtain synthetic vehicle pictures; A discrimination step of inputting the synthetic vehicle pictures into a vehicle classification network to obtain discrimination results, recording the parametric texture pictures corresponding to all the pictures classified as non-vehicles in the discrimination results, and modeling the parametric texture pictures corresponding to the non-vehicle pictures to obtain final adversarial sample pictures; and A testing step of inputting vehicle pictures with the adversarial sample pictures pasted on the vehicle body into the vehicle detection model to be tested, obtaining the number of vehicles detected by the vehicle detection model, and calculating the robustness degree representing the robustness of the vehicle detection model according to the number; The calculation formula for the robustness degree r of the vehicle detection model is: r = m / k, where m is the number of vehicles detected by the vehicle detection model and k is the number of the adversarial sample pictures.
2. The method for testing the robustness of a vehicle detection model based on deep learning according to claim 1, characterized in that, in the data acquisition step, the vehicle pictures are obtained by analyzing the positional relationship between the monitoring camera, the road and the vehicle in the road traffic scenario.
3. The method for testing the robustness of a vehicle detection model based on deep learning according to claim 1, characterized in that, the training step of the parametric texture generation network model further includes: combining the parametric texture pictures with the CAD model for texture mapping to obtain texture-mapped vehicle pictures, performing backpropagation update according to the reconstruction error loss between the input vehicle pictures and the texture-mapped vehicle pictures, and adjusting the parameters of the parametric texture generation network model to realize the training of the parametric texture generation network model.
4. The method for testing the robustness of a vehicle detection model based on deep learning according to claim 1, characterized in that, In the modeling in the discrimination step, a Gaussian mixture model with k degrees of freedom is used to model the parameterized texture image corresponding to the image of the non-vehicle.
5. The method for testing the robustness of a vehicle detection model based on deep learning according to claim 4, wherein, the means of the k Gaussian mixture models obtained after modeling are used as the k adversarial sample images.
6. The method for testing the robustness of a vehicle detection model based on deep learning according to any one of claims 1-5, wherein, the CAD model is a three-dimensional model.
7. The method for testing the robustness of a vehicle detection model based on deep learning according to any one of claims 1-5, wherein, the noise refers to Gaussian noise, the mean of the Gaussian noise is 0, and the covariance matrix is the identity matrix.
8. An apparatus for testing the robustness of a vehicle detection model based on deep learning, wherein, comprising: a data acquisition module configured to acquire vehicle images in a road traffic scene and establish a corresponding CAD model according to the vehicle images; a parameterized texture generation network model training module configured to establish and train a parameterized texture generation network model, the parameterized texture generation network model including a deep neural network, a triangular mesh model, and a deep convolutional generative network, extracting features from the vehicle images through the deep neural network to obtain first features, extracting features from the CAD model through the triangular mesh model to obtain second features, connecting the first features and the second features as the input of the deep convolutional generative network, and the output of the deep convolutional generative network being a parameterized texture image; an appearance rendering module configured to collect two-dimensional images corresponding to the CAD models of vehicles of a specified model, input the two-dimensional images and their corresponding CAD models into the parameterized texture generation network model to obtain the corresponding parameterized texture images, calculate the image mean of the parameterized texture images, add noise to the image mean and combine it with the CAD model to obtain synthetic vehicle images; a discrimination module configured to input the synthetic vehicle images into a vehicle classification network to obtain discrimination results, record the parameterized texture images corresponding to all the images classified as non-vehicles in the discrimination results, and model the parameterized texture images corresponding to the non-vehicle images to obtain final adversarial sample images; and a testing module configured to input vehicle images with the adversarial sample images pasted on the vehicle body into the vehicle detection model to be tested, obtain the number of vehicles detected by the vehicle detection model, and calculate the robustness degree characterizing the robustness of the vehicle detection model according to the number; The calculation formula for the robustness degree r of the vehicle detection model is: r = m / k, where m is the number of vehicles detected by the vehicle detection model and k is the number of adversarial sample images.
9. An electronic device, comprising: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, wherein, when the program is executed by a processor, the method according to any one of claims 1-7 is implemented.
Citation Information
Patent Citations
Multidirectional vehicle model detection recognition system based on deep learning
CN105975941A
Traffic image multi-type vehicle detection method based on deep study
CN106096531A