Intelligent perception method, device and medium for visual failure mode in fusion parking scenario

By using a spatiotemporal fusion detection model based on four onboard cameras and an AI classification network, the problem of image quality degradation caused by camera contamination in intelligent driving has been solved, achieving efficient and real-time visual failure detection and improving the safety and reliability of intelligent parking.

CN116030268BActive Publication Date: 2026-01-02HEFEI UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310033005.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-10
Publication Date
2026-01-02
Estimated Expiration
2043-01-10

AI Technical Summary

Technical Problem

Existing technologies lack effective methods for identifying and addressing image quality degradation caused by pollution in vehicle cameras during intelligent driving, leading to erroneous information transmission and safety hazards. Furthermore, deep learning algorithms have low timeliness in edge detection with limited resources.

Method used

By capturing overlapping image data using four vehicle-mounted cameras, a spatiotemporal fusion detection model is established. Combining inter-frame correlation, gradient bias, and information entropy features, an AI-based classification network model is constructed using an SVM classifier and multi-view vision fusion to achieve accurate detection and classification of visual failure modes.

Benefits of technology

It improves the accuracy and real-time performance of visual failure detection, alleviates the problems of insufficient detection accuracy of traditional algorithms and excessive resource consumption of deep learning algorithms, and enhances the reliability and safety of intelligent parking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116030268B_ABST
    Figure CN116030268B_ABST
Patent Text Reader

Abstract

The application discloses a kind of intelligent perception method, equipment and medium in fusion parking scene visual failure mode, the steps of the method include:1 obtains the fisheye image data with continuous or equal-interval frame;2 establish space-time fusion detection model;3 establish the classification network model of multi-scene data based on AI;4 realize prediction using the model established, to achieve the purpose of failure mode detection.The application can overcome the problem that the robustness of single traditional algorithm is lower and the timeliness of AI prediction is poorer under the condition of limited hardware resources on the end side, and the data is fused and detected at pixel level, feature level and decision level in time and space, so that real-time accurate intelligent perception of visual failure mode can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer vision and intelligent driving, and particularly relates to a method for intelligent perception of fusion of visual failure modes in a parking scene. BACKGROUND

[0002] The on-board lens is an important sensor for environment perception in the assisted driving scene, and its failure will cause the functions developed based on vision to fail to work normally, so it is necessary to monitor its failure condition. The failure detection algorithm can identify the condition of image quality degradation caused by the camera due to various reasons described above, and perform function degradation, function shutdown or active reminder and the like.

[0003] Most of the sensors of intelligent vehicles are exposed outside the vehicle, and will be disturbed or even damaged to varying degrees when driving on some extreme road conditions. For example, when the vehicle drives on a muddy road, mud on the road surface is easy to splash onto the camera surface, causing the camera to be contaminated. The image captured by the contaminated camera has reduced clarity, making it difficult to reflect the real road conditions. After these erroneous information is transmitted to the control unit, the vehicle will make an error judgment, which is likely to cause a traffic accident. Therefore, the image quality captured by the camera is particularly important for the safe driving of intelligent vehicles, but there is currently no specific identification method for disturbed images in intelligent driving image recognition algorithms, relying only on the subjective judgment of the driver. During the interference of the camera, the image captured by the camera changes from clear to blurred, so a blurred image recognition method can be used to determine whether the camera is disturbed. Current methods for identifying blurred images are mainly divided into two categories. One is a subjective evaluation method, in which an observer evaluates the blurring degree of an image. This method has high stability, but it is inefficient and requires an observer to make a judgment, which cannot be embedded in other systems, so it is less used in practical applications. The other method is an objective evaluation method, which evaluates image quality through a specific algorithm. This method is fast and efficient, and can be combined with other systems, so it is the focus of current research. According to whether there is a reference image, the objective evaluation method can be divided into full-reference method, partial-reference method and no-reference method. The full-reference method requires a reference image to be selected first, and then some statistical error values between the target image and the reference image are obtained based on statistical principles, such as the mean square error of the corresponding pixel values and the peak signal-to-noise ratio of the entire image. Then, the blurring degree of the target image is judged according to these error values. This method has strict requirements for the reference image, so it has low stability. The no-reference method does not require a reference image, but only needs to extract some characteristic value information from the target image, such as image edge information, image frequency energy and image information entropy, to judge the blurring degree of the image according to the size of these characteristic values. The semi-reference method is between the full-reference method and the no-reference method, and a reference image is needed in the evaluation process, but only part of the information in the reference image is used to compare with the target image. In practical applications, the images captured by the vehicle-mounted camera are constantly changing, and the road conditions are complex and variable, so it is difficult to find a suitable reference image, so the no-reference method is usually used.

[0004] The no-reference image method can be further divided into traditional algorithms and deep learning-based algorithms. The representative traditional algorithms mainly focus on the extraction of pixel statistical distribution features, such as mean square error, correlation coefficient and probability distribution model, etc. Such algorithms have the advantage of fast speed, but due to the difficulty of a single distribution feature to completely represent various types of failure, the detection accuracy is relatively low. With the rapid development of the field of deep learning, more and more attention has been paid to image quality evaluation methods based on deep learning. In the field of deep learning, the end-to-end NR-IQA (no-reference image quality evaluation) method can integrate feature extraction and fitting / regression into a unified framework and optimize simultaneously. This has become the mainstream NR-IQA scheme. Under this framework, some IQA methods have appeared in recent years, and the mainstream ideas include the following categories: the first category is to use GAN to restore the distorted image, and use the generated restored image and the distorted image to calculate the loss and measure the distance to generate a quality score; the second category is to use the Rank learning idea, which can solve the problem of no large-scale data set in IQA, thereby improving the model accuracy. In addition, some research ideas use attention mechanisms to improve the weight of the region of interest to evaluate image quality; some methods use meta-learning to learn prior knowledge to solve the problem of too small scale of the no-reference image data set. The deep learning-based method extracts deep features of the image, which can significantly improve the detection accuracy, but the deep learning algorithm consumes a large amount of computing power, and the timeliness of detection is low, which is difficult to meet the real-time detection requirements under the condition of limited hardware resources on the end side. SUMMARY

[0005] The present application is to overcome the deficiencies of the prior art, and proposes an intelligent perception method, device and medium that fuse the visual failure mode in the parking scene, so as to realize accurate perception of the visual failure state, thereby improving the reliability and safety of intelligent parking.

[0006] To achieve the above-mentioned application purposes, the following technical solutions are adopted:

[0007] The intelligent perception method of the present application fuses the visual failure mode in the parking scene, which is used for detecting the visual failure mode. The visual failure mode refers to the image appearing as blur or all black in a local area or the whole image, and the local area or the whole image that is blurred or all black is the failure area. The intelligent perception method is performed according to the following steps:

[0008] Step 1: Four-way vehicle-mounted cameras are used to shoot overlapping surround view image data and calibrate to obtain a pixel-level overlapping area of four-way surround view image data.

[0009] Collecting N frames of fisheye image sequence from the same surround view image data for preprocessing, the preprocessing adopts grayscale and median filter algorithm, and a preprocessed sample set is denoted as S={S1, S2,..., S n ,...,S N}; S n represents the n-th preprocessed image.

[0010] Step 2, establishing a detection model of space-time fusion, comprising: an inter-frame correlation feature extraction module, an inter-frame gradient deviation extraction module, an inter-frame information entropy extraction module, a spatial domain multi-feature fusion module and a multi-view vision fusion module; wherein the inter-frame correlation feature extraction module is composed of a dynamic mean calculation layer and a dynamic variance calculation layer; the inter-frame gradient deviation extraction module comprises a gradient convolution layer;

[0011] Step 2.1, the inter-frame correlation feature extraction module models the features of the failure area by traversing the discrete degree of the pixels of the multi-frame images in the time dimension, and obtains a discrete feature vector T(i,j) at the pixel coordinate (i,j) in the image S n .

[0012] Step 2.2, the inter-frame gradient deviation extraction module generates a dynamic background image sequence through time sampling operation, and then completes static gradient saliency modeling by the gradient convolution layer, and obtains a multi-feature vector H(i,j);

[0013] Step 2.3, based on the inter-frame information entropy extraction module, a dynamic background image sequence is generated through time sampling operation, and then entropy increase saliency modeling is completed by an information entropy increase graph, and an information entropy increase feature inf n (i,j) at the pixel coordinate (i,j) in the n-th image S n is obtained.

[0014] Step 2.4, the spatial domain multi-feature fusion module inputs the multi-feature vector H(i,j) and the information entropy increase feature Inf(i,j) into a SVM classifier for training, and generates a binary classification model, so as to output a pixel-level binary image by using the binary classification model, and after morphological processing, a detection result of the failure area A is obtained;

[0015] Step 2.5, if the detection result of the failure area A is true, step 2.6 is executed; if the detection result of the failure area A is false, it indicates that the corresponding one-way vehicle-mounted camera is working normally, and the next surround view image data is processed;

[0016] Step 2.6, the multi-view vision fusion module judges whether the failure area is in the overlap area or intersects with the overlap area, if yes, step 2.7 is executed; otherwise, it indicates that the one-way vehicle-mounted camera corresponding to the failure area A is working abnormally, and step 3 is executed.

[0017] Step 2.7, another road overlapping area is called and preprocessed image data is input into the detection model of spatio-temporal fusion for processing to obtain another road failure area B detection result;

[0018] If the detection result of another road failure area B is true, it indicates that the failure area A is caused by the external scene, the failure area A corresponds to a normal working road camera, and the next road around view image data is processed;

[0019] If the detection result of another road failure area B is false, it indicates that the failure area A is caused by lens failure, the failure area A corresponds to an abnormal working road camera, and step 3 is executed;

[0020] Step 3, an AI-based multi-scene data classification network model is established to confirm the image corresponding to the failure area output by the spatio-temporal fusion detection model;

[0021] Step 3.1, the classification network model is constructed, including: a first convolution layer, a second maximum pooling layer, a third first down-sampling module and a first basic module layer; a fourth second down-sampling module and three basic module layers, a fifth third down-sampling module and a fifth basic module layer, a sixth convolution layer, a seventh global pooling layer, and a last full connection layer; wherein the basic module outputs a scale-invariant feature map through a residual convolution layer, and the down-sampling module is used to output a two-dimensional feature map with a size reduced by half;

[0022] Step 3.2, the failure image is obtained and input into each layer of the classification network model for processing in sequence to obtain a prediction probability for constructing a cross-entropy loss function;

[0023] Step 3.3, the classification network model is trained by using gradient descent method, and the cross-entropy loss function is calculated to update the model parameters until the cross-entropy loss function converges, so as to obtain the trained classification network model;

[0024] Step 3.4, the trained classification network model is used to predict the Nth preprocessed image S N of the failure area A to obtain a predicted failure probability;

[0025] Step 3.5, the predicted failure probability is compared with a threshold M, if greater than the threshold M, it indicates that the failure area A corresponds to an abnormal working road camera, and a prompt is given, otherwise, it indicates that the failure area A corresponds to a normal working road camera.

[0026] The intelligent perception method for fusing the parking scene visual failure mode has the characteristics that the step 2.1 comprises:

[0027] Step 2.1.1, initialize n = 1;

[0028] Let S n (i,j) denote the pre-processed image S n at pixel coordinate (i,j) in the n-th iteration;

[0029] Initialize the dynamic mean layer at pixel coordinate (i,j) in the image S n-1 n-1 to be the mean value μ n-1 (i,j) = 0; Initialize the dynamic variance layer at pixel coordinate (i,j) in the image S n-1 n-1 to be the variance σ n-1 (i,j) = 0;

[0030] Step 2.1.2, calculate the mean value μ n (i,j) of the dynamic mean layer at pixel coordinate (i,j) in the image S n n; calculate the variance σ n (i,j) of the dynamic variance layer at pixel coordinate (i,j) in the image S n-1 n; if n ≥ N, then the discrete feature vector T(i,j) = (μ n (i,j), σ n (i,j)) at pixel coordinate (i,j) in the image S n-1 n is composed of the mean value μ n (i,j) of the dynamic mean layer at pixel coordinate (i,j) in the image S n n and the variance σ n (i,j) of the dynamic variance layer at pixel coordinate (i,j) in the image S n n, otherwise, assign n + 1 to n, and execute Step 2.1.4;

[0032] Step 2.1.4, calculate the mean value μ n (i,j) of the dynamic mean layer at pixel coordinate (i,j) in the image S n n; calculate the variance σ n (i,j) of the dynamic variance layer at pixel coordinate (i,j) in the image S n n; if n ≥ N, then the discrete feature vector T(i,j) = (μ n (i,j), σ n (i,j)) at pixel coordinate (i,j) in the image S n-1 n is composed of the mean value μ n (i,j) of the dynamic mean layer at pixel coordinate (i,j) in the image S n n and the variance σ n (i,j) of the dynamic variance layer at pixel coordinate (i,j) in the image S n n, otherwise, assign n + 1 to n, and execute Step 2.1.4;n (i,j)-μ n (i,j)|+|μ n-1 (i,j)-μ n (i,j)|) / 2;

[0033] Step 2.1.4, then return to step 2.1.3 and execute sequentially.

[0034] Step 2.2 includes:

[0035] Step 2.2.1: Initialize n = 1;

[0036] Let the gradient convolutional layer of the (n-1)th iteration be in the image S n-1 The gradient g at pixel coordinates (i, j) n-1 (i,j)=0;

[0037] Step 2.2.2: Calculate image S using the Sobel operator. n The global gradient G at pixel coordinates (i, j) n (i,j); thus, the gradient of the nth iteration convolutional layer in image S is calculated. n The gradient g at pixel coordinates (i, j) n (i,j)=G n (i,j)+g n-1 (i,j);

[0038] Step 2.2.3: Determine if n≥N holds true. If true, then change the gradient g. n The combination of (i,j) and T(i,j) forms a multi-feature vector, denoted as H(i,j)=(μ n (i,j), σ n (i,j), g n (i,j)) and output; otherwise, assign n+1 to n and return to step 2.2.2 to execute sequentially.

[0039] Step 2.3 includes:

[0040] Step 2.3.1: Initialize n = 1; Let the (n-1)th frame image S n-1 The information entropy increase feature at pixel coordinates (i, j) inf n-1 (i,j)=0;

[0041] Let the nth frame image S n The information entropy increase feature inf at pixel coordinates (i, j) n (i,j)=S n (i,j)+inf n-1 (i,j);

[0042] Step 2.3.2, judging whether n>=N is established, if yes, the information entropy increment feature of the n-th frame image S n is outputted, otherwise, n+1 is assigned to n, and step 2.3.3 is executed.

[0043] Step 2.3.3, calculating the information entropy increment feature inf n of the n-th frame image S n (i,j)=-S n-1 (i,j)log2(S n-1 (i,j))+inf n-1 (i,j) at pixel coordinate (i,j) is outputted.

[0044] Step 2.3.4, returning to step 2.3.2 for sequential execution.

[0045] The electronic device of the present application comprises a memory and a processor, and the feature lies in that the memory is used for storing a program supporting the processor to execute any of the intelligent perception methods, and the processor is configured to execute the program stored in the memory.

[0046] The computer readable storage medium of the present application stores a computer program, and the feature lies in that the computer program is executed by the processor to execute any of the steps of the intelligent perception method.

[0047] Compared with the prior art, the present application has the following beneficial effects:

[0048] 1. The present application constructs a solution combining the advantages of traditional algorithms and intelligent algorithms, realizes efficient fusion of visual failure modes at pixel level, feature level and decision level in different space-time dimensions, and overcomes the problem of low detection accuracy of traditional algorithms, and also alleviates the problem of excessive occupation of embedded resources of deep learning algorithms, thereby improving the overall detection performance.

[0049] 2. In the time dimension, based on the time invariance of failure modes, the present application obtains pixel-level continuous frame correlation features, frame gradient deviation features and frame information entropy features; in the space dimension, the above features maintain structural invariance, and constitute a spatial fusion feature set, which is trained to obtain a SVM detection model on a large amount of labeled failure data, and is used for detecting failure areas; in addition, based on the multi-view visual characteristics of vehicle-mounted lenses, under the condition that the vision sensor is calibrated, the failure areas of the current lens can be fused and verified with the failure results of another adjacent lens, and finally the detection result of the traditional algorithm is obtained. The result drives the intelligent algorithm to make further classification.

[0050] 3. This invention designs a network that can still perform real-time high-precision classification under the condition of no GPU on the ARM side. It optimizes the weights on massive self-collected failed data to obtain highly reliable classification information and finally gives a decision-level fusion result. Attached Figure Description

[0051] Figure 1 This is a schematic diagram of the method flow of the present invention;

[0052] Figure 2 This is a structural diagram of the efficient classification network of the present invention;

[0053] Figure 3 This is a diagram illustrating the preprocessing results of the present invention;

[0054] Figure 4 This is a partial dataset used for training an SVM classifier using multi-feature fusion in this invention.

[0055] Figure 5 This is the training and testing dataset for the classification network in this invention. Detailed Implementation

[0056] In this embodiment, an intelligent perception method that integrates visual failure modes in parking scenarios is used for the detection of image failure modes. These visual failure modes refer to images appearing blurred or completely black in local areas or throughout the entire image, and the blurred or black local areas or the entire image are considered failure areas. Figure 1 As shown, this intelligent sensing method proceeds according to the following steps:

[0057] Step 1: Capture overlapping surround view image data using four vehicle-mounted cameras and calibrate them to obtain the pixel-level overlapping area of ​​the four surround view image data.

[0058] N frames of fisheye image sequences were collected from the same surround view image data and preprocessed using grayscale conversion and median filtering algorithms. The resulting preprocessed sample set is denoted as S = {S1, S2, ..., S...}. n ,...,S N};S n This represents the preprocessed image of the nth frame; in this embodiment, grayscale and medianization processing is performed using video frame data captured from the actual vehicle, and the result is as follows. Figure 3 As shown. In this experiment, the frame length N of the short video sequence was 10, but it is not limited to this value.

[0059] Step 2, the detection model of spatio-temporal fusion is established, including: an inter-frame correlation feature extraction module, an inter-frame gradient deviation extraction module, an inter-frame information entropy extraction module, a spatial domain multi-feature fusion module and a multi-view vision fusion module; wherein the inter-frame correlation feature extraction module is composed of a dynamic mean calculation layer and a dynamic variance calculation layer; the inter-frame gradient deviation extraction module includes a gradient convolution layer.

[0060] Step 2.1, the inter-frame correlation feature extraction module models the features of the failure area by traversing the discrete degree of the pixels of the multi-frame images in the time dimension:

[0061] Step 2.1.1, initializing n = 1;

[0062] Let S n (i,j) represent the preprocessed image S n (i,j) of the n-th frame.

[0063] Initialize the mean value μ n-1 (i,j) of the dynamic mean layer of the n-1-th iteration at the pixel coordinate (i, j) in the image S n-1 (i,j) = 0; initialize the variance σ n-1 (i,j) of the dynamic variance layer of the n-1-th iteration at the pixel coordinate (i, j) in the image S n-1 (i,j) = 0;

[0064] Step 2.1.2, calculate the mean value μ n (i,j) of the dynamic mean layer of the n-th iteration at the pixel coordinate (i, j) in the image S n (i,j) = S n (i,j) + μ n-1 (i,j); calculate the variance σ n (i,j) of the dynamic variance layer of the n-th iteration at the pixel coordinate (i, j) in the image S n (i,j) = σ n-1 (i,j);

[0065] Step 2.1.3, if n ≥ N, then the mean value μ n (i,j) of the dynamic mean layer of the n-th iteration at the pixel coordinate (i, j) in the image S n (i,j) and the variance σ n (i,j) of the dynamic variance layer of the n-th iteration at the pixel coordinate (i, j) in the image S n (i,j) constitute the discrete feature vector at the pixel coordinate (i, j) in the image S n (i,j), denoted as T(i,j) = (μ n (i,j), σ n (i,j)), otherwise, assign n + 1 to n, and execute step 2.1.4.

[0066] Step 2.1.4, calculate the dynamic mean layer of the nth iteration at the pixel coordinate (i, j) in the image S n Step 2.1.5, calculate the dynamic variance layer of the nth iteration at the pixel coordinate (i, j) in the image S n (i,j) = (S n (i,j) + μ n-1 (i,j)) / 2; calculate the dynamic variance layer of the nth iteration at the pixel coordinate (i, j) in the image S n Step 2.1.6, calculate the dynamic gradient layer of the nth iteration at the pixel coordinate (i, j) in the image S n (i,j) = (|S n (i,j) - μ n (i,j)| + |μ n-1 (i,j) - μ n (i,j)|) / 2;

[0067] Step 2.1.4, return to step 2.1.3 for sequential execution;

[0068] In this embodiment, based on the video frame data in step 1, since the background of the inter-frame non-failure region is constantly changing, the local brightness of the failure region is weakened, resulting in that the mean value of the region in the dynamic mean graph is lower than that of other regions, and the gray scale of the region is stable, which is represented as the variance value of the region in the dynamic variance graph being smaller than that of other regions.

[0069] Step 2.2, the inter-frame gradient deviation extraction module generates a dynamic background image sequence through a time sampling operation, and then a gradient convolution layer is used to complete static gradient saliency modeling;

[0070] Step 2.2.1, initialize n = 1;

[0071] Let the gradient g n-1 (i,j) of the gradient convolution layer of the (n-1)th iteration at the pixel coordinate (i, j) in the image S n-1 (i,j) = 0;

[0072] Step 2.2.2, calculate the global gradient G n (i,j) at the pixel coordinate (i, j) in the image S n (i,j); thereby calculating the gradient g n (i,j) of the gradient convolution layer of the nth iteration at the pixel coordinate (i, j) in the image S n (i,j) = G n (i,j) + g n-1 (i,j);

[0073] Step 2.2.3, determine whether n ≥ N is true, if true, combine the gradient g n (i,j) and T(i,j) into a multi-feature vector, denoted as H(i,j) = (μ n(i,j), σ n (i,j), g n (i,j)) and output; otherwise, assign n+1 to n and return to step 2.2.2 to execute sequentially.

[0074] In this embodiment, the Sobel operator convolution kernel size is 3×3, but it is not limited to this value. Due to different sampling times, the sampled image sequence S contains gradient information under different scenarios. Since gradient vanishing or weakening will occur in the failure region, the feature will become more obvious by superimposing the gradient in the time dimension.

[0075] Step 2.3: Based on the inter-frame information entropy extraction module, a dynamic background image sequence is generated through time sampling operation, and then the entropy increase significance model is completed by information entropy augmentation map;

[0076] Step 2.3.1: Initialize n = 1; Let the (n-1)th frame image S n-1 The information entropy increase feature at pixel coordinates (i, j) inf n-1 (i,j)=0;

[0077] Let the nth frame image S n The information entropy increase feature at pixel coordinates (i, j) inf n (i,j)=S n (i,j)+inf n-1 (i,j);

[0078] Step 2.3.2: Determine if n≥N holds true. If true, then set the nth frame image S... n The information entropy increase feature at pixel coordinates (i, j) is denoted as Inf(i, j) and output. Otherwise, n+1 is assigned to n and step 2.3.3 is executed.

[0079] Step 2.3.3: Calculate the image S of the nth frame. n The information entropy increase feature inf at pixel coordinates (i, j) n (i,j)=-S n-1 (i,j)log2(S n-1 (i,j))+inf n-1 (i,j);

[0080] Step 2.3.4, return to step 2.3.2 and execute sequentially;

[0081] In this embodiment, the most intuitive feeling caused by visual failure is that the target cannot be seen clearly, that is, information is lost. By calculating the increase and decrease of information entropy in the same area on different frame images, the failed area can be distinguished.

[0082] Step 2.4, the spatial domain multi-feature fusion module inputs the multi-feature vector H(i,j) and the information entropy increasing feature Inf(i,j) into the SVM classifier for training, and generates a binary classification model, so as to output a pixel-level binary image by using the binary classification model, and the binary image is subjected to morphological processing to obtain the detection result of the failure region A.

[0083] In the embodiment, the training sample directly determines the fitting degree of the model to the failure result. By collecting 10,000 actual measurement data, label 0 represents a positive sample, i.e., a normal scene picture, and label 1 represents a negative sample, i.e., a failure scene, such as Figure 4 Some positive and negative samples are shown.

[0084] Step 2.5, if the detection result of the failure region A is true, step 2.6 is executed; if the detection result of the failure region A is false, it indicates that the corresponding vehicle-mounted camera works normally, and the next surround image data is processed;

[0085] Step 2.6, the multi-view vision fusion module judges whether the failure region is in the overlapping region or intersects with the overlapping region, if yes, step 2.7 is executed; otherwise, it indicates that the corresponding vehicle-mounted camera works abnormally, and step 3 is executed;

[0086] Step 2.7, the surround image data of another overlapping region is retrieved and preprocessed, and then input into the spatio-temporal fusion detection model for processing to obtain the detection result of another failure region B;

[0087] If the detection result of another failure region B is true, it indicates that the failure region A is caused by an external scene, the corresponding vehicle-mounted camera works normally, and the next surround image data is processed;

[0088] If the detection result of another failure region B is false, it indicates that the failure region A is caused by lens failure, the corresponding vehicle-mounted camera works abnormally, and step 3 is executed.

[0089] Step 3, a multi-scene data classification network model based on AI is established, which is used for confirming the image corresponding to the failure region output by the spatio-temporal fusion detection model;

[0090] Step 3.1, constructing a classification network model, comprising: a convolutional layer in the first layer, a maximum pooling layer in the second layer, a first down-sampling module and a first base module layer in the third layer; a second down-sampling module and three base module layers in the fourth layer, a third down-sampling module and a fifth base module layer in the fifth layer, a convolutional layer in the sixth layer, a global pooling layer in the seventh layer, and a fully connected layer in the last layer; wherein the base module outputs a feature map with invariant scale through a residual convolutional layer, and the down-sampling module is used to output a feature map with two-dimensional size reduced by half.

[0091] In this embodiment, Figure 2 The final network structure is shown.

[0092] Step 3.2, obtaining a failure image and inputting each layer in the classification network model for processing in turn to obtain a prediction probability, which is used to construct a cross-entropy loss function;

[0093] Step 3.3, training the classification network model by using the gradient descent method, and calculating the cross-entropy loss function to update the model parameters until the cross-entropy loss function converges, thereby obtaining a trained classification network model;

[0094] Step 3.4, using the trained classification network model to predict the Nth preprocessed image S N corresponding to the failure region A to obtain a predicted failure probability.

[0095] In this embodiment, based on the powerful feature fitting capability of the deep learning network, a large amount of real vehicle test data is used for network model training, the number of data sets is 40,000 normal scenes and 60,000 various failure scene data, and currently six types of data, i.e., stains, splashing, full occlusion, blur, low illumination and normal, are collected, and the data is divided into a training set and a test set according to a ratio of 4:1, Figure 5 The data corresponding to different failure modes is shown.

[0096] Step 3.5, comparing the predicted failure probability with a threshold M, if the predicted failure probability is greater than the threshold M, it indicates that the corresponding vehicle-mounted camera of the failure region A works abnormally, and a prompt is given, otherwise, it indicates that the corresponding vehicle-mounted camera of the failure region A works normally.

Claims

1. An intelligent perception method that integrates visual failure modes in parking scenarios, characterized in that, This method is used to detect visual failure modes, which refer to images appearing blurry or completely black in local areas or throughout the entire image. The blurry or black local areas or the entire image are then identified as failure areas. The intelligent perception method is performed according to the following steps: Step 1: Capture overlapping surround view image data using four vehicle-mounted cameras and calibrate them to obtain the pixel-level overlapping area of ​​the four surround view image data. N frames of fisheye image sequences were collected from the same surround view image data and preprocessed using grayscale conversion and median filtering algorithms. The resulting preprocessed sample set is denoted as S = {S1, S2, ..., S...}. n ,...,S N };S n This represents the preprocessed image of the nth frame; Step 2: Establish a spatiotemporal fusion detection model. It includes: an inter-frame correlation feature extraction module, an inter-frame gradient bias extraction module, an inter-frame information entropy extraction module, a spatial domain multi-feature fusion module, and a multi-view vision fusion module; wherein, the inter-frame correlation feature extraction module consists of a dynamic mean calculation layer and a dynamic variance calculation layer; the inter-frame gradient bias extraction module includes a gradient convolution layer; Step 2.1: The inter-frame correlation feature extraction module models the features of the failure region by traversing the discreteness of pixels in the time dimension across multiple frames of images, thus obtaining image S. n The discrete feature vector T(i,j) at pixel coordinates (i,j); Step 2.2: The inter-frame gradient deviation extraction module generates a dynamic background image sequence through time sampling operation, and then completes static gradient saliency modeling by gradient convolution layer to obtain multiple feature vectors H(i,j); Step 2.3: Based on the inter-frame information entropy extraction module, a dynamic background image sequence is generated through time sampling operation. Then, the entropy increase significance model is completed by the information entropy augmentation map, and the nth frame image S is obtained. n The information entropy increase feature inf at pixel coordinates (i, j) n (i,j); Step 2.4: The spatial domain multi-feature fusion module inputs the multi-feature vector H(i,j) and the information entropy increase feature Inf(i,j) into the SVM classifier for training and generates a binary classification model. The binary classification model is then used to output a pixel-level binary image. After morphological processing, the binary image is used to obtain the detection result of the failure region A. Step 2.5: If the detection result of the failed area A is true, proceed to step 2.6; if the detection result of the failed area A is false, it means that the corresponding vehicle camera is working normally, and the next surround view image data is processed. Step 2.6: The multi-view vision fusion module determines whether the failure area is in the overlapping area or intersects with the overlapping area. If so, proceed to step 2.7; otherwise, it indicates that the vehicle camera corresponding to the failure area A is malfunctioning, and proceed to step 3. Step 2.7: Retrieve the panoramic image data of the other overlapping area, preprocess it, and then input it into the spatiotemporal fusion detection model for processing to obtain the detection result of the other failed area B; If the detection result of the other failure area B is true, it means that the failure area A is caused by the external scene, the vehicle camera corresponding to the failure area A is working normally, and the next surround view image data is processed. If the detection result of the other failure area B is false, it means that the failure area A is caused by lens failure, the vehicle camera corresponding to failure area A is malfunctioning, and step 3 is executed. Step 3: Establish an AI-based classification network model for multi-scene data to confirm the images corresponding to the failure areas output by the spatiotemporal fusion detection model; Step 3.1: Construct a classification network model, including: a convolutional layer in the first layer, a max pooling layer in the second layer, a first downsampling module and a first basic module layer in the third layer; a second downsampling module and three basic module layers in the fourth layer, a third downsampling module and a fifth basic module layer in the fifth layer, a convolutional layer in the sixth layer, a global pooling layer in the seventh layer, and a fully connected layer in the last layer; wherein, the basic module outputs a scale-invariant feature map through a residual convolutional layer, and the downsampling module is used to output a two-dimensional feature map with its size reduced by half; Step 3.2: Obtain the failure image and input it into each layer of the classification network model for sequential processing to obtain the predicted probability, which is used to construct the cross-entropy loss function; Step 3.3: Train the classification network model using gradient descent and calculate the cross entropy loss function to update the model parameters until the cross entropy loss function converges, thereby obtaining the trained classification network model. Step 3.4: Use the trained classification network model to preprocess the Nth frame image S corresponding to the failure region A. N Make a prediction to obtain the predicted failure probability; Step 3.5 compares the predicted failure probability with the threshold M. If it is greater than the threshold M, it means that the vehicle camera corresponding to the failure area A is malfunctioning and a prompt is made. Otherwise, it means that the vehicle camera corresponding to the failure area A is working normally.

2. The intelligent perception method for visual failure modes in a parking scenario according to claim 1, characterized in that, Step 2.1 includes: Step 2.1.1: Initialize n = 1; Let S n (i,j) represents the preprocessed image S of the nth frame. n The gray value at pixel coordinates (i, j); Initialize the dynamic mean layer for the (n-1)th iteration in image S n-1 The mean μ at pixel coordinates (i, j) n-1 (i,j)=0; Initialize the dynamic variance layer for the (n-1)th iteration in image S n-1 The variance σ at pixel coordinates (i, j) n-1 (i,j)=0; Step 2.1.2: Calculate the dynamic mean layer of the nth iteration in image S. n The mean μ at pixel coordinates (i, j) n (i,j)=S n (i,j)+μ n-1 (i,j); Calculate the dynamic variance layer of the nth iteration in image S n The variance σ at pixel coordinates (i, j) n (i,j)=σ n-1 (i,j); Step 2.1.3: If n ≥ N, then the dynamic mean layer of the nth iteration in image S... n The mean μ at pixel coordinates (i, j) n (i,j) and the dynamic variance layer of the nth iteration in image S n The variance σ at pixel coordinates (i, j) n (i,j) constitute the image S n The discrete feature vector at pixel coordinates (i, j) is denoted as T(i, j) = (μ n (i,j), σ n (i,j)), otherwise, assign n+1 to n and then execute step 2.1.4; Step 2.1.4: Calculate the dynamic mean layer of the nth iteration in image S. n The mean μ at pixel coordinates (i, j) n (i,j)=(S n (i,j)+μ n-1 (i,j)) / 2; Calculate the dynamic variance layer of the nth iteration in image S n The variance σ at pixel coordinates (i, j) n (i,j)=(|S n (i,j)-μ n (i,j)|+|μ n-1 (i,j)-μ n (i,j)|) / 2; Step 2.1.4, then return to step 2.1.3 and execute sequentially.

3. The intelligent perception method for visual failure modes in a parking scenario according to claim 2, characterized in that, Step 2.2 includes: Step 2.2.1: Initialize n = 1; Let the gradient convolutional layer of the (n-1)th iteration be in the image S n-1 The gradient g at pixel coordinates (i, j) n-1 (i,j)=0; Step 2.2.2: Calculate image S using the Sobel operator. n The global gradient G at pixel coordinates (i, j) n (i,j); thus, the gradient of the nth iteration convolutional layer in image S is calculated. n The gradient g at pixel coordinates (i, j) n (i,j)=G n (i,j)+g n-1 (i,j); Step 2.2.3: Determine if n≥N holds true. If true, then change the gradient g. n The combination of (i,j) and T(i,j) forms a multi-feature vector, denoted as H(i,j)=(μ n (i,j), σ n (i,j), g n (i,j)) and output; otherwise, assign n+1 to n and return to step 2.2.2 to execute sequentially.

4. The intelligent perception method for visual failure modes in a parking scenario according to claim 3, characterized in that, Step 2.3 includes: Step 2.3.1: Initialize n = 1; Let the (n-1)th frame image S n-1 The information entropy increase feature inf at pixel coordinates (i, j) n-1 (i,j)=0; Let the nth frame image S n The information entropy increase feature inf at pixel coordinates (i, j) n (i,j)=S n (i,j)+inf n-1 (i,j); Step 2.3.2: Determine if n≥N holds true. If true, then set the nth frame image S... n The information entropy increase feature at pixel coordinates (i, j) is denoted as Inf(i, j) and output. Otherwise, n+1 is assigned to n and step 2.3.3 is executed. Step 2.3.3: Calculate the image S of the nth frame. n The information entropy increase feature inf at pixel coordinates (i, j) n (i,j)=-S n-1 (i,j)log2(S n-1 (i,j))+inf n-1 (i,j); Step 2.3.4, return to step 2.3.2 and execute sequentially.

5. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store a program that supports the processor in executing any of the intelligent sensing methods of claims 1-4, and the processor is configured to execute the program stored in the memory.

6. A computer-readable storage medium storing a computer program thereon, characterized in that, The computer program is executed by the processor to perform the steps of any of the intelligent sensing methods described in claims 1-4.