Dangerous source identification method and device, electronic equipment and storage medium

By using a classification discriminant model trained by deep autoencoder and classifier, preprocessing and feature extraction of vehicle body 360 images has been solved, and the problem of difficulty in identifying dangerous sources in complex driving environments in the prior art is solved, and the safety of autonomous driving technology is improved.

CN120014601APending Publication Date: 2025-05-16CHONGQING CHANGAN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510094566.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art is difficult to effectively identify and judge different hazard sources in the vehicle body 360 image in complex driving environments, resulting in weak model judgment ability and lack of targeting.

Method used

The classification discriminant model trained by deep autoencoder and classifier is used to identify the hazard source for the grayscale images to be identified, and image features are extracted through the densely connected convolutional network architecture to judge the number and risk of hazards.

Benefits of technology

It improves the recognition ability of the hazard source identification model, can more accurately judge the characteristics of the hazard source in the image, and enhances the safety of autonomous driving technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014601A_ABST
    Figure CN120014601A_ABST
Patent Text Reader

Abstract

The invention relates to a dangerous source identification method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring to-be-identified image data of a to-be-identified dangerous source; preprocessing the to-be-recognized image data to obtain a to-be-recognized grayscale image; performing hazard source identification on the to-be-identified grayscale image based on a pre-trained classification and discrimination model to obtain the number of hazard sources and the risk degree of each hazard source; wherein the classification discrimination model is obtained by training a depth auto-encoder and a classifier; an encoder of the deep auto-encoder is a convolutional network architecture based on dense connection superposed on a channel. According to the method, an encoder of a deep auto-encoder for training a classification and discrimination model is based on a densely connected convolutional network architecture superposed on a channel, so that the classification and discrimination model can judge the shape, size, number, risk degree and other characteristics of a hazard source in an image, the recognition capability of the classification and discrimination model is improved, and the classification and discrimination efficiency is improved. Therefore, data support is provided for the safety of the automatic driving technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image recognition technology, and in particular to a hazard source identification method, device, electronic device and storage medium. Background Art

[0002] With the rapid development of the automotive industry, intelligent driving and autonomous driving technologies have become the core direction of future transportation systems. With the continuous advancement of these technologies, vehicle safety issues have become increasingly prominent, especially in complex driving environments. How to ensure the safety of drivers and passengers has become a key research topic. In the past, although ordinary artificial intelligence algorithms could determine the source of danger to a certain extent, they did not conduct in-depth research on the characteristics of the 360-degree image of the vehicle body, and the algorithms could not well mine the characteristics of different sources of danger in the image, resulting in the model's relatively weak judgment ability and lack of pertinence to the problem. Summary of the invention

[0003] In order to solve the above technical problems or at least partially solve the above technical problems, the present application provides a hazard source identification method, device, electronic device and storage medium.

[0004] In a first aspect, the present application provides a method for identifying a hazard source, the method comprising:

[0005] Acquire image data of a hazard source to be identified;

[0006] Preprocessing the image data to be identified to obtain a grayscale image to be identified;

[0007] Based on a pre-trained classification and discrimination model, the grayscale image to be identified is used to identify the hazard sources, and the number of hazard sources and the hazard level of each hazard source are obtained; wherein the classification and discrimination model is trained by a deep autoencoder and a classifier; the encoder of the deep autoencoder is a convolutional network architecture based on dense connections superimposed on channels.

[0008] Optionally, before performing risk source identification on the grayscale image to be identified based on a pre-trained classification and discrimination model to obtain the number of risk sources and the risk level of each risk source, the method further comprises: acquiring the classification and discrimination model;

[0009] The training process of the classification and discrimination model includes:

[0010] Acquire training sample data; wherein the training sample data includes a first training sample and a second training sample; the first training sample includes a plurality of non-dangerous vehicle body panoramic images without dangerous sources, and the second training sample includes a plurality of target training sample pairs, and the target training sample pairs include a dangerous vehicle body panoramic image and a non-dangerous vehicle body panoramic image;

[0011] Obtain an initial deep autoencoder and an initial classifier; wherein the encoder of the initial deep autoencoder is built based on the network convolution part of the dense convolutional network; and the initial classifier is built based on the classification head part of the dense convolutional network;

[0012] Training the initial deep autoencoder based on the first training sample to obtain a deep autoencoder;

[0013] Training the initial classifier based on the second training sample to obtain a classifier;

[0014] The encoder of the deep autoencoder and the classifier are used as the classification discrimination model.

[0015] Optionally, obtaining training sample data includes:

[0016] Acquire a panoramic image of the vehicle body;

[0017] Preprocessing the vehicle body panoramic image to obtain a panoramic grayscale image; wherein the panoramic grayscale image includes M non-dangerous vehicle body panoramic images without dangerous sources and N dangerous vehicle body panoramic images with dangerous sources; the M is greater than the N;

[0018] Determine a first image set and a second image set from the M non-dangerous vehicle body panoramic images; wherein the first image set includes N non-dangerous vehicle body panoramic images corresponding to the N dangerous vehicle body panoramic images respectively, and the second image set includes other images of the M non-dangerous vehicle body panoramic images that do not belong to the first image set;

[0019] Using the other images as the first training samples;

[0020] The first image set and N dangerous vehicle body panoramic images are used as the second training samples.

[0021] Optionally, training the initial deep autoencoder based on the first training sample to obtain a deep autoencoder includes:

[0022] Get the reconstruction error function;

[0023] The following training process is performed on each sample data in the first training sample respectively:

[0024] Inputting the sample data into the encoder of the initial deep autoencoder to obtain a latent vector for reducing the dimension of the sample data;

[0025] Inputting the latent vector into the decoder of the initial deep autoencoder to obtain restored data of the latent vector by increasing its dimension;

[0026] Calculating a reconstruction error between the sample data and the restored data based on the reconstruction error function;

[0027] After optimizing the parameters of the initial deep autoencoder, the next sample data is obtained from the first training sample, and the training process is repeated until the reconstruction error tends to be stable, and the initial deep autoencoder is used as the final deep autoencoder.

[0028] Optionally, training the initial classifier based on the second training sample to obtain a classifier includes:

[0029] The following training process is performed on each group of sample images in the second training sample:

[0030] The following processing is performed on each sample image in a group of sample images respectively: the sample image is input into the deep autoencoder for encoding, X network layers are used in sequence to extract features of the sample image, features of the X network layers are obtained, and the features of the X network layers are integrated to obtain features of the hazard source in the sample image; wherein, when feature extraction is performed on each target network layer in the X network layers, the outputs of all network layers before the target network layer are superimposed on the input of the target network layer;

[0031] Inputting the features of the danger sources in the sample image into the initial classifier to classify the danger sources and determine the danger level of the danger sources;

[0032] The loss function is calculated based on the hazard source classification results and the hazard level judgment results, as well as the hazard source category identification and hazard probability of the group of sample images. After optimizing the parameters of the initial classifier according to the loss function, the next group of sample images is obtained from the second training sample, and the training process is repeated until the loss function tends to be stable, and the initial classifier is used as the final classifier.

[0033] Optionally, obtain an initial deep autoencoder and an initial classifier, including:

[0034] Obtain a first deep autoencoder and a first classifier;

[0035] Initializing weights of the decoder of the first deep autoencoder and the first classifier;

[0036] Initializing the encoder of the first deep autoencoder based on the visualization database training parameters;

[0037] Using the initialized encoder and decoder as the initial deep autoencoder;

[0038] The initialized first classifier is used as the initial classifier.

[0039] Optionally, preprocessing the vehicle body panoramic image to obtain a panoramic grayscale image includes:

[0040] Get the weight values ​​corresponding to the preset three color channels and the color value of each pixel;

[0041] A weighted average calculation is performed on each pixel in the vehicle body panoramic image based on the weight value and the color value to obtain the panoramic grayscale image.

[0042] In a second aspect, the present application provides a hazard source identification device, the device comprising:

[0043] An acquisition module, used for acquiring image data to be identified of a hazard source to be identified;

[0044] A preprocessing module, used for preprocessing the image data to be identified to obtain a grayscale image to be identified;

[0045] An identification module is used to identify hazardous sources in the grayscale image to be identified based on a pre-trained classification and discrimination model to obtain the number of hazardous sources and the degree of hazard of each hazardous source; wherein the classification and discrimination model is obtained by training a deep autoencoder and a classifier; the encoder part of the deep autoencoder is a convolutional network architecture based on dense connections superimposed on channels.

[0046] In a third aspect, the present application provides an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus;

[0047] Memory, used to store computer programs;

[0048] The processor is used to implement the steps of the hazard source identification method described in any embodiment of the first aspect when executing the program stored in the memory.

[0049] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the hazard source identification method as described in any embodiment of the first aspect are implemented.

[0050] Beneficial effects of this application:

[0051] The method provided in the embodiment of the present application obtains image data to be identified of a hazardous source to be identified; pre-processes the image data to be identified to obtain a grayscale image to be identified; identifies hazardous sources on the grayscale image to be identified based on a pre-trained classification and discrimination model to obtain the number of hazardous sources and the degree of hazard of each hazardous source; wherein the classification and discrimination model is obtained by training a deep autoencoder and a classifier; the encoder of the deep autoencoder is a convolutional network architecture based on dense connections superimposed on channels. In this method, the encoder of the deep autoencoder for training the classification and discrimination model is a convolutional network architecture based on dense connections superimposed on channels, so the classification and discrimination model can judge features such as the number and degree of hazard of hazardous sources in the image, thereby improving the recognition ability of the classification and discrimination model, thereby providing data support for the safety of autonomous driving technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0053] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0054] Figure 1 A system architecture diagram of a hazard source identification method provided in one embodiment of the present application;

[0055] Figure 2 A schematic diagram of a flow chart of a method for identifying a hazard source provided in one embodiment of the present application;

[0056] Figure 3 A structural diagram of a deep autoencoder provided in one embodiment of the present application;

[0057] Figure 4 A DenseNet network structure diagram provided for one embodiment of the present application;

[0058] Figure 5 An overall flow chart of a hazard source identification method provided in one embodiment of the present application;

[0059] Figure 6 An overall network structure diagram of a hazard source identification method provided by an embodiment of the present application;

[0060] Figure 7 A schematic diagram of the structure of a hazard source identification device provided in one embodiment of the present application;

[0061] Figure 8 A schematic diagram of the structure of an electronic device provided in accordance with an embodiment of the present application. DETAILED DESCRIPTION

[0062] The following will describe the implementation methods of the present application with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present application from the contents disclosed in this specification. The present application can also be implemented or applied through other different specific implementation methods, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present application. It should be understood that the preferred embodiments are only for illustrating the present application, not for limiting the scope of protection of the present application.

[0063] The first embodiment of the present application provides a method for identifying a hazard source, which can be applied to Figure 1 The system architecture shown in the figure includes at least an image input module 101 and an image recognition module 102, and the image input module 101 and the image recognition module 102 establish a communication connection. Specifically, the system architecture can be a vehicle with automatic driving capability, wherein the type of vehicle is not limited, for example, it can be a fuel vehicle, a pure electric vehicle, a hybrid vehicle or a fuel cell vehicle, etc., without limitation.

[0064] Next, based on the system architecture, the hazard source identification method is described in detail, such as Figure 2 , the hazard source identification method includes:

[0065] Step 201, obtaining image data of a hazard source to be identified.

[0066] The image data to be identified of the hazard source to be identified may be a panoramic image of the vehicle body, such as a 360-degree image of the vehicle body. The 360-degree image of the vehicle body may be acquired and fused by a plurality of image acquisition devices arranged in the front, rear, left, and right directions of the vehicle.

[0067] Step 202: pre-process the image data to be identified to obtain a grayscale image to be identified.

[0068] When identifying hazardous sources, color features are unnecessary. Therefore, the grayscale image to be identified can be obtained by performing grayscale preprocessing on the image data to be identified. The grayscale image to be identified not only removes unnecessary color features but also retains the shape, size and other features of each object in the image to the greatest extent. It can simplify the model output while removing noise features and convert the three-channel form of the model RGB into a single-channel form of a grayscale image, thereby improving the model accuracy and computing performance.

[0069] Step 203, based on the pre-trained classification and discrimination model, the hazard source is identified on the grayscale image to be identified, and the number of hazard sources and the hazard level of each hazard source are obtained; wherein the classification and discrimination model is trained by a deep autoencoder and a classifier; the encoder of the deep autoencoder is a convolutional network architecture based on dense connections superimposed on channels.

[0070] First, the structure of the deep autoencoder is introduced. Figure 3 is a structural diagram of a deep autoencoder. The autoencoder uses the data x itself as a mapping signal to guide the training of the neural network, that is, the mapping relationship of the network is It divides the network into two parts. The first half of the sub-network learns g θ1 : The mapping relationship of x→z, the second half of the sub-network learning The mapping relationship of g θ1 It is considered as a process of encoding input data, that is, encoding high-dimensional input information x into relatively low-dimensional latent vectors or latent variables z. This part of the network is called the encoder. θ2 It can be regarded as the process of decoding the latent vector, that is, decoding the encoded latent vector z back to the high-dimensional x. This part of the network is called the decoder. θ It is the encoder g θ1 and decoder h θ2 In order to make the dimension reduction complex and nonlinear to achieve better results, the encoder g θ1 and decoder h θ2 This is usually done using a neural network, and such an autoencoder is called a deep autoencoder.

[0071] In this method, the encoder of the deep autoencoder for training the classification discrimination model is a convolutional network architecture based on dense connections superimposed on channels. Therefore, the classification discrimination model can judge the characteristics of the danger sources in the image such as shape, size, quantity and degree of danger, thereby improving the recognition ability of the classification discrimination model and providing data support for the safety of autonomous driving technology.

[0072] In one embodiment, before performing risk source identification on the grayscale image to be identified based on a pre-trained classification and discrimination model and obtaining the number of risk sources and the risk level of each risk source, the method further includes: acquiring the classification and discrimination model.

[0073] Among them, the training process of the classification and discrimination model includes:

[0074] Step S1, obtaining training sample data; wherein the training sample data includes a first training sample and a second training sample; the first training sample includes a plurality of non-dangerous vehicle body panoramic images without dangerous sources, and the second training sample includes a plurality of target training sample pairs, and the target training sample pairs include a dangerous vehicle body panoramic image and a non-dangerous vehicle body panoramic image.

[0075] Since most of the images collected during the image acquisition process are non-dangerous panoramic images of the vehicle body in a non-dangerous state, the data set used for training will bring about the problem of data imbalance while simulating this ratio in reality. At the same time, it is very difficult to collect image data for training, so a method is also needed to reasonably utilize these large amounts of data in a non-dangerous state to further improve the performance of the model and make the model more generalizable. In this embodiment, firstly, training sample data is obtained from a panoramic image of the vehicle body (hereinafter also referred to as a 360-degree image of the vehicle body), and further, the first training sample includes a plurality of non-dangerous panoramic images of the vehicle body without dangerous sources, and the second training sample includes a plurality of target training sample pairs, and the target training sample pair includes a dangerous panoramic image of the vehicle body and a non-dangerous panoramic image of the vehicle body. Among them, the first training sample is used to train a deep autoencoder, and the second training sample is a training sample including a plurality of target training sample pairs, which is used to train a classifier.

[0076] In this embodiment, using an autoencoder as a feature extractor can not only help alleviate the problem of imbalance in the number of dangerous vehicle body panoramic images and non-dangerous vehicle body panoramic images in the training sample data, but also effectively utilize a large amount of non-dangerous vehicle body panoramic image data.

[0077] In one embodiment, obtaining training sample data includes: obtaining a panoramic image of a vehicle body; preprocessing the panoramic image of the vehicle body to obtain a panoramic grayscale image; wherein the panoramic grayscale image includes M panoramic images of a non-dangerous vehicle body without a hazard source and N panoramic images of a dangerous vehicle body with a hazard source; M is greater than N; determining a first image set and a second image set from the M non-dangerous vehicle body panoramic images; wherein the first image set includes N non-dangerous vehicle body panoramic images corresponding to the N dangerous vehicle body panoramic images respectively, and the second image set includes other images of the M non-dangerous vehicle body panoramic images that do not belong to the first image set; using the other images as the first training sample; and using the first image set and the N dangerous vehicle body panoramic images as the second training samples.

[0078] In this embodiment, the panoramic image of the vehicle body is preprocessed to obtain a panoramic grayscale image, including: obtaining weight values ​​corresponding to preset three-color channels and the color value of each pixel; and performing weighted average calculation on each pixel in the panoramic image of the vehicle body based on the weight value and the color value to obtain a panoramic grayscale image.

[0079] When identifying hazardous sources, color features are unnecessary. Therefore, the panoramic grayscale image can be obtained by grayscale preprocessing the panoramic image of the vehicle body. The panoramic image of the vehicle body removes unnecessary color features while retaining the shape, size and other features of each object in the image to the greatest extent. It can simplify the model output while removing noise features and convert the three-channel form of the model RGB into a single-channel form of a grayscale image, thereby improving the model accuracy and computing performance.

[0080] In this embodiment, before formally starting to train the model, the 360-degree image of the vehicle body is preprocessed. Here, the weighted average method can be used to grayscale the image. Since the human eye is most sensitive to green and least sensitive to blue, the weighted average of the three components of RGB can be performed according to the following formula to obtain a grayscale image that is closer to the human eye, thereby removing unnecessary color features while retaining the shape and size features of the object to the greatest extent.

[0081] Gray(i,j)=0.299R(i,j)+0.578G(i,j)+0.114B(i,j)

[0082] Among them, i and j represent the positions of pixels, and R, G, and B represent the color values ​​of the red, green, and blue channels at (i, j). After grayscale processing, the image is resized and normalized by scaling and cropping without affecting the key content of the image, so that it can be used as a standard input for the neural network model.

[0083] In this embodiment, a first image set and a second image set are determined from M non-dangerous vehicle body panoramic images; wherein the first image set includes N non-dangerous vehicle body panoramic images corresponding to the N dangerous vehicle body panoramic images respectively, and the second image set includes other images in the M non-dangerous vehicle body panoramic images that do not belong to the first image set; the other images are used as the first training samples; the first image set and the N dangerous vehicle body panoramic images are used as the second training samples. Generally speaking, most of the collected images are non-dangerous vehicle body panoramic images in a non-dangerous state. For example, M is 10000 and N may be only 100. At this time, 100 non-dangerous vehicle body panoramic images corresponding to the dangerous vehicle body panoramic images are determined from the M non-dangerous vehicle body panoramic images. These 100 non-dangerous vehicle body panoramic images and the 100 dangerous vehicle body panoramic images are used as the first image set, and the remaining 9900 non-dangerous vehicle body panoramic images are used as the second image set.

[0084] Step S2, obtaining an initial deep autoencoder and an initial classifier; wherein the encoder of the initial deep autoencoder is built based on the network convolution part of the dense convolutional network; and the initial classifier is built based on the classification head part of the dense convolutional network.

[0085] Dense convolutional networks are DenseNet (Densely Connected Convolutional Networks). Since the DenseNet network itself only has the function of classifying image data and does not have the characteristics of an autoencoder, it is necessary to modify the DenseNet network structure to integrate the idea of ​​autoencoding into the network. Here, the DenseNet121 network is used as an example. The specific method is to branch when the DenseNet121 network is forward propagated to the last convolution layer: one part is used as the decoder in the autoencoder to reconstruct the image. At this time, the convolution part of the DenseNet121 network is equivalent to the encoder part in the autoencoder, and one part is globally averaged pooled for classification. Among them, the decoder uses a deconvolution operation, which is an incomplete inverse process of the convolution operation, and expands the feature map with a step size of 2.

[0086] Since pedestrians, vehicles and other dangerous sources have great differences in size, the number of pixels represented in the 360-degree vehicle body surround camera image also varies greatly, so the DenseNet network is used in this embodiment. Figure 4 , DenseNet can build rich and complex feature representations from the ultra-deep layers of the neural network. This architecture has a recursive combination connection method called dense connection. Each layer of DenseNet directly superimposes the feature maps of all previous layers on the channel dimension, reuses the previous detailed features, and ensures that more attention is paid to the details of the image in the subsequent feature extraction, so that larger and smaller features can work together on the result, fit the characteristics of the car body 360 image, and improve the network's ability to perform this task. For example, the input of h4 not only contains x3 from h3, but also includes x0, x1, and x2 from the previous layers, so that they are superimposed on the channel dimension. This network structure allows the output of one layer to be directly connected to the input of the next layer across multiple layers, so that the learned features can be reused throughout the network, taking into account both large and small features, so that their weights can be trained accordingly and work together on the result.

[0087] In one embodiment, obtaining an initial deep autoencoder and an initial classifier includes: obtaining a first deep autoencoder and a first classifier; initializing weights of a decoder and a first classifier of the first deep autoencoder; initializing the encoder of the first deep autoencoder based on training parameters of a visualization database; using the initialized encoder and decoder as the initial deep autoencoder; and using the initialized first classifier as the initial classifier.

[0088] In this embodiment, the parameters of the classification head (i.e., the first classifier) ​​and the decoder part are set to perform weight initialization, such as Glorot initialization, which is a method for initializing the weights of a neural network. At the same time, transfer learning is used to initialize the parameters of the convolution part of the DenseNet121 network to the visualization database training parameters, such as the parameters of ImageNet training, so that it can obtain certain feature extraction capabilities. However, since the features of the images in the ImageNet dataset are quite different from those of the 360-degree images of the vehicle body, the feature extraction capabilities obtained by transfer learning are limited and require further fine-tuning.

[0089] Step S3: training an initial deep autoencoder based on the first training sample to obtain a deep autoencoder.

[0090] In one embodiment, an initial deep autoencoder is trained based on a first training sample to obtain a deep autoencoder, including: obtaining a reconstruction error function; performing the following training process on each sample data in the first training sample respectively: inputting the sample data into the encoder of the initial deep autoencoder to obtain a latent vector that reduces the dimension of the sample data; inputting the latent vector into the decoder of the initial deep autoencoder to obtain restored data that increases the dimension of the latent vector; calculating the reconstruction error between the sample data and the restored data based on the reconstruction error function; after optimizing the parameters of the initial deep autoencoder, obtaining the next sample data from the first training sample, repeating the training process until the reconstruction error tends to be stable, and using the initial deep autoencoder as the final deep autoencoder.

[0091] In this embodiment, the autoencoder first transforms the input into a low-dimensional latent vector z through the encoder, and reconstructs the output based on the original input x through the decoder. Right now This is the optimization goal of the autoencoder: minimize the Distance Right now in,

[0092] here That is the reconstruction error function. For the entire autoencoder, the reconstruction error function can be regarded as the loss function of the neural network, because there is no essential difference between the autoencoder and the neural network, except that the supervisory signal used for training is changed from the data label to the data input x itself. Generally speaking, the reconstruction error function can be designed directly using the mean square error method:

[0093]

[0094] where x i represents the i-th feature of the current data x, Represents the output of the current data The i-th feature of . Thanks to the nonlinear transformation ability of deep neural networks, autoencoders can obtain good data dimensionality reduction expressions. Using autoencoders as feature extractors can not only help alleviate the problem of data imbalance, but also effectively utilize a large amount of non-dangerous image data. After training the autoencoder, use the encoder part of the autoencoder as a data feature extractor, or feature reducer, for the next step of classifying the degree of danger in the classifier.

[0095] In this embodiment, using an autoencoder as a feature extractor can not only help alleviate the problem of data imbalance, but also effectively utilize a large number of non-dangerous vehicle body panoramic images. When the reconstruction error tends to be stable, the obtained deep autoencoder can make the reconstructed image as similar as possible to the original image, that is, after the training is completed, the encoder part of the model further improves the feature extraction ability.

[0096] Step S4: training an initial classifier based on the second training sample to obtain a classifier.

[0097] In one embodiment, an initial classifier is trained based on a second training sample to obtain a classifier, including: performing the following training process on each group of sample images in the second training sample respectively: performing the following processing on each sample image in a group of sample images respectively, inputting the sample image into a deep autoencoder for encoding, sequentially using X network layers to extract features of the sample image, obtaining features of the X network layers, and integrating the features of the X network layers to obtain features of the hazard source in the sample image; wherein, when extracting features at each target network layer in the X network layers, the outputs of all network layers before the target network layer are superimposed on the input of the target network layer; inputting the features of the hazard source in the sample image into the initial classifier to classify the hazard source and judge the degree of hazard of the hazard source; calculating the loss function based on the hazard source classification result and the hazard degree judgment result, as well as the hazard source category identification and hazard probability of a group of sample images, optimizing the parameters of the initial classifier based on the loss function, obtaining the next group of sample images from the second training sample, and repeating the training process until the loss function becomes stable, and then using the initial classifier as the final classifier.

[0098] In this embodiment, the convolution part and the classification head part of the DenseNet121 network can be regarded as a classifier as a whole, and training data containing panoramic images of non-dangerous vehicle bodies and panoramic images of dangerous vehicle bodies can be trained. The strong feature extraction ability obtained by the previous two convolution layers is used to obtain the final classification network. However, since there is still data imbalance at this time, the loss function can use the focal loss function to let the network focus on difficult samples and further alleviate the data imbalance problem.

[0099] Step S5, using the encoder and classifier of the deep autoencoder as a classification discriminant model.

[0100] After the network training is completed, the encoder + classifier parts can be combined into a whole and saved in the car computer. The 360-degree image of the driving car body can be input into the model, and the output neuron meaning determined before the network training can be obtained (in this embodiment, it can be "number of dangers" and "degree of danger"), and the information can be fed back to the car computer or hardware to guide the car computer, hardware and driver to make corresponding judgments and reactions to ensure driving safety.

[0101] In a specific embodiment, the overall flow chart of the hazard source identification method is as follows: Figure 5 ,include:

[0102] Image preprocessing;

[0103] Model initialization and transfer learning;

[0104] Training of autoencoder models;

[0105] Training of classification and discrimination models;

[0106] Classification and discrimination model loading and use.

[0107] In this embodiment, before the formal training model is started, the 360-degree image of the vehicle body is preprocessed. Here, the weighted average method is used to grayscale the image. Since the human eye is most sensitive to green and least sensitive to blue, the weighted average of the three components of RGB according to the following formula can obtain a grayscale image that is closer to the human eye, thereby removing unnecessary color features while retaining the shape and size characteristics of the object to the greatest extent. The purpose of this embodiment is to identify the degree of danger. The number of output neurons can be freely set to train the network. Here, the output of two node neurons is taken as an example, assuming that they represent "number of dangers" and "degree of danger". Since the DenseNet network itself only has the function of classifying image data and does not have the characteristics of an autoencoder, it is necessary to modify the network structure to some extent and integrate the idea of ​​autoencoding into the network. Here, the DenseNet121 network is used as an example. The specific method is to branch when the DenseNet121 network is forward propagated to the last convolution layer: one part is used as a decoder in the autoencoder to reconstruct the image. At this time, the convolution part of the DenseNet121 network is equivalent to the encoder part in the autoencoder; one part is globally averaged pooled for classification. Among them, the decoder uses a deconvolution operation, which is an incomplete inverse process of the convolution operation, and expands the feature map with a step size of 2.

[0108] The model is trained below. The overall network structure of the hazard source identification method is as follows: Figure 6 shown.

[0109] In this embodiment, first, after all data are uniformly preprocessed, the parameters of the classification head and decoder are set to Glorot initialization, and the parameters of the convolution part of the DenseNet121 network are initialized to the parameters of the training ImageNet using transfer learning, so that it can obtain certain feature extraction capabilities. However, since the features of the images in the ImageNet dataset are quite different from those of the 360-degree images of the vehicle body, the feature extraction capabilities obtained by transfer learning are limited and require further fine-tuning.

[0110] Then, the autoencoder part of the network was trained using excess non-dangerous 360-degree image data, that is, fine-tuning the convolution part of the DenseNet121 network after transfer learning, and simultaneously training the decoder in the second half, with the training goal of making the reconstructed image as similar as possible to the original image. The loss function used was the mean square error function at the pixel level. After the training was completed, the encoder part of the model further improved its feature extraction capabilities.

[0111] Finally, the convolution part and the classification head of the DenseNet121 network are considered as a classifier, and the training data containing 360-degree images of dangerous vehicles and 360-degree images of non-dangerous vehicles are trained. The strong feature extraction capability obtained by the previous two convolution layers is used to obtain the final classification network. However, since there is still data imbalance at this time, the loss function uses the Focal Loss function to let the network focus on difficult samples and further alleviate the data imbalance problem.

[0112] After the network training is completed, the encoder + classifier parts can be combined into a whole and saved in the car computer. The 360-degree image of the driving car body can be input into the model, and the output neuron meaning determined before the network training can be obtained (in this case, "number of dangers" and "degree of danger"), and the information can be fed back to the car computer or hardware to guide the car computer, hardware and driver to make corresponding judgments and reactions to ensure driving safety.

[0113] Based on the same technical concept, the second embodiment of the present application provides a hazard source identification device, such as Figure 7 , the device comprises:

[0114] The acquisition module 701 is used to acquire the image data to be identified of the hazard source to be identified;

[0115] A preprocessing module 702 is used to preprocess the image data to be identified to obtain a grayscale image to be identified;

[0116] The identification module 703 is used to identify the hazard sources of the grayscale image to be identified based on a pre-trained classification and discrimination model to obtain the number of hazard sources and the hazard level of each hazard source; wherein the classification and discrimination model is obtained by training a deep autoencoder and a classifier; the encoder part of the deep autoencoder is a convolutional network architecture based on dense connections superimposed on channels.

[0117] In this device, the encoder of the deep autoencoder for training the classification discrimination model is a convolutional network architecture based on dense connections superimposed on channels. Therefore, the classification discrimination model can judge the characteristics of the danger sources in the image such as shape, size, quantity and danger level, thereby improving the recognition ability of the classification discrimination model and providing data support for the safety of autonomous driving technology.

[0118] like Figure 8 As shown, the third embodiment of the present application provides an electronic device, including a processor 111, a communication interface 112, a memory 113 and a communication bus 114, wherein the processor 111, the communication interface 112, and the memory 113 communicate with each other through the communication bus 114.

[0119] Memory 113, used for storing computer programs;

[0120] In one embodiment, the processor 111 is used to implement the hazard source identification method provided by any one of the aforementioned method embodiments when executing the program stored in the memory 113.

[0121] The memory and processor in the above electronic device communicate through the communication bus and the communication interface. The communication bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The communication bus can be divided into an address bus, a data bus, a control bus, etc.

[0122] The memory may include a random access memory (RAM) or a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.

[0123] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0124] A fourth embodiment of the present application provides a computer-readable medium having a non-volatile program code executable by a processor.

[0125] Optionally, in an embodiment of the present application, a computer-readable medium is configured to store program code for a processor to execute the above method.

[0126] Optionally, the specific examples in this embodiment may refer to the examples described in the above embodiments, and this embodiment will not be described in detail here.

[0127] When the embodiments of the present application are specifically implemented, reference may be made to the above-mentioned embodiments, which have corresponding technical effects.

[0128] It is understood that the embodiments described herein may be implemented in hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit may be implemented in one or more application specific integrated circuits (ASIC), digital signal processors (DSP), digital signal processing devices (DSPD), programmable logic devices (PLD), field programmable gate arrays (FPGA), general purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions of the present application, or a combination thereof.

[0129] For software implementation, the technology of this article can be implemented by a unit that performs the functions of this article. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.

[0130] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0131] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0132] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of modules is only a logical function division. There may be other division methods in actual implementation, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0133] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0134] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0135] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods of each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard drives, ROM, RAM, magnetic disks, or optical disks.

[0136] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0137] The above embodiments are only preferred embodiments for fully illustrating the present application, and the protection scope of the present application is not limited thereto. Any equivalent substitution or change made by a person skilled in the art based on the present application is within the protection scope of the present application.

Claims

1. A method for identifying a hazard source, characterized in that: The method comprises: Acquire image data of a hazard source to be identified; Preprocessing the image data to be identified to obtain a grayscale image to be identified; Based on a pre-trained classification and discrimination model, the grayscale image to be identified is used to identify the hazard sources, and the number of hazard sources and the hazard level of each hazard source are obtained; wherein the classification and discrimination model is trained by a deep autoencoder and a classifier; the encoder of the deep autoencoder is a convolutional network architecture based on dense connections superimposed on channels.

2. The method according to claim 1, characterized in that Before identifying the hazard sources of the grayscale image to be identified based on the pre-trained classification and discrimination model and obtaining the number of hazard sources and the hazard level of each hazard source, the method further includes: acquiring the classification and discrimination model; The training process of the classification and discrimination model includes: Acquire training sample data; wherein the training sample data includes a first training sample and a second training sample; the first training sample includes a plurality of non-dangerous vehicle body panoramic images without dangerous sources, and the second training sample includes a plurality of target training sample pairs, and the target training sample pairs include a dangerous vehicle body panoramic image and a non-dangerous vehicle body panoramic image; Obtain an initial deep autoencoder and an initial classifier; wherein the encoder of the initial deep autoencoder is built based on the network convolution part of the dense convolutional network; and the initial classifier is built based on the classification head part of the dense convolutional network; Training the initial deep autoencoder based on the first training sample to obtain a deep autoencoder; Training the initial classifier based on the second training sample to obtain a classifier; The encoder of the deep autoencoder and the classifier are used as the classification discrimination model.

3. The method according to claim 2, characterized in that Get training sample data, including: Acquire a panoramic image of the vehicle body; Preprocessing the vehicle body panoramic image to obtain a panoramic grayscale image; wherein the panoramic grayscale image includes M non-dangerous vehicle body panoramic images without dangerous sources and N dangerous vehicle body panoramic images with dangerous sources; the M is greater than the N; Determine a first image set and a second image set from the M non-dangerous vehicle body panoramic images; wherein the first image set includes N non-dangerous vehicle body panoramic images corresponding to the N dangerous vehicle body panoramic images respectively, and the second image set includes other images of the M non-dangerous vehicle body panoramic images that do not belong to the first image set; Using the other images as the first training samples; The first image set and N dangerous vehicle body panoramic images are used as the second training samples.

4. The method according to claim 3, characterized in that Training the initial deep autoencoder based on the first training sample to obtain a deep autoencoder includes: Get the reconstruction error function; The following training process is performed on each sample data in the first training sample respectively: Inputting the sample data into the encoder of the initial deep autoencoder to obtain a latent vector for reducing the dimension of the sample data; Inputting the latent vector into the decoder of the initial deep autoencoder to obtain restored data of the latent vector by increasing its dimension; Calculating a reconstruction error between the sample data and the restored data based on the reconstruction error function; After optimizing the parameters of the initial deep autoencoder, the next sample data is obtained from the first training sample, and the training process is repeated until the reconstruction error tends to be stable, and the initial deep autoencoder is used as the final deep autoencoder.

5. The method according to claim 4, characterized in that Training the initial classifier based on the second training sample to obtain a classifier includes: The following training process is performed on each group of sample images in the second training sample: The following processing is performed on each sample image in a group of sample images respectively: the sample image is input into the deep autoencoder for encoding, X network layers are used in sequence to extract features of the sample image, features of the X network layers are obtained, and the features of the X network layers are integrated to obtain features of the hazard source in the sample image; wherein, when feature extraction is performed on each target network layer in the X network layers, the outputs of all network layers before the target network layer are superimposed on the input of the target network layer; Inputting the features of the danger sources in the sample image into the initial classifier to classify the danger sources and determine the danger level of the danger sources; The loss function is calculated based on the hazard source classification results and the hazard level judgment results, as well as the hazard source category identification and hazard probability of the group of sample images. After optimizing the parameters of the initial classifier according to the loss function, the next group of sample images is obtained from the second training sample, and the training process is repeated until the loss function tends to be stable, and the initial classifier is used as the final classifier.

6. The method according to claim 2, characterized in that Get the initial deep autoencoder and initial classifier, including: Obtain a first deep autoencoder and a first classifier; Initializing weights of the decoder of the first deep autoencoder and the first classifier; Initializing the encoder of the first deep autoencoder based on the visualization database training parameters; Using the initialized encoder and decoder as the initial deep autoencoder; The initialized first classifier is used as the initial classifier.

7. The method according to claim 3, characterized in that Preprocessing the vehicle body panoramic image to obtain a panoramic grayscale image includes: Get the weight values ​​corresponding to the preset three color channels and the color value of each pixel; A weighted average calculation is performed on each pixel in the vehicle body panoramic image based on the weight value and the color value to obtain the panoramic grayscale image.

8. A hazard source identification device, characterized in that: The device comprises: An acquisition module, used for acquiring image data to be identified of a hazard source to be identified; A preprocessing module, used for preprocessing the image data to be identified to obtain a grayscale image to be identified; An identification module is used to identify hazardous sources in the grayscale image to be identified based on a pre-trained classification and discrimination model to obtain the number of hazardous sources and the degree of hazard of each hazardous source; wherein the classification and discrimination model is obtained by training a deep autoencoder and a classifier; the encoder part of the deep autoencoder is a convolutional network architecture based on dense connections superimposed on channels.

9. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, for implementing the method according to any one of claims 1 to 7 when executing a program stored in a memory.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.