Visual identification method and device of whole vehicle road spectrum, electronic equipment and storage medium

Multi-spectral images are collected through the combination of vehicle-mounted visible light and infrared cameras, combined with image processing and convolutional neural networks, and the problems of visual recognition road spectrum technology being sensitive to light, poor environmental adaptability and high cost are solved, achieving efficient and low-cost recognition in complex environments.

CN120451933APending Publication Date: 2025-08-08CHONGQING TONGWO AUTOMOBILE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510417317.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The existing visual recognition road spectrum technology is sensitive to light, relies on high-quality images, has poor environmental adaptability and high cost, which limits its application in mid- and low-end models.

Method used

Multi-spectral image data is collected by combining vehicle-mounted visible light cameras and infrared cameras. The image denoising, enhancement and normalization process is input to the convolutional neural network for feature extraction and classification recognition, and the recognition results are generated to support driving decisions.

Benefits of technology

It improves the adaptability of visual recognition road spectrum technology in complex lighting and extreme environments, reduces dependence on image quality, and reduces system costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451933A_ABST
    Figure CN120451933A_ABST
Patent Text Reader

Abstract

The invention provides a visual identification method and device for a whole vehicle road spectrum, electronic equipment and a storage medium. The method comprises the steps that multispectral image data are collected through a vehicle-mounted camera, and the multispectral image data comprise images of roads around a vehicle; preprocessing the multispectral image data, and carrying out normalization processing on the preprocessed multispectral image data; inputting the normalized multispectral image data into a convolutional neural network for feature extraction to obtain a feature vector; according to the feature vector, feature information in the road image is classified and recognized, a recognition result is generated, and the recognition result is output to a system of the vehicle so that the vehicle can make a driving decision based on the recognition result. According to the invention, the adaptability of a visual identification road spectrum technology in complex illumination and extreme environments can be improved, the dependence on image quality is reduced, and the system cost is reduced while the accuracy is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of new energy vehicle technology, and in particular to a method, device, electronic device and storage medium for visual recognition of a vehicle's road spectrum. Background Art

[0002] With the development of autonomous driving technology, the demand for vehicles to accurately perceive the road environment in real time is growing. Visual recognition road map technology is one of the core means for autonomous driving systems to perceive road features, identify lanes, traffic signs, and road surface conditions. In existing technologies, visual recognition road maps are usually based on computer vision and image processing technology. High-definition cameras are used to capture road images, and the collected images are pre-processed by operations such as denoising and enhancement to improve image quality. Subsequently, image processing algorithms are used to extract key feature information from the road images, such as lane lines, edges of traffic signs, and road surface textures. The extracted features are then matched with a preset pattern library to identify different types and conditions of the road surface. The recognition results are used for vehicle navigation, control, or decision support.

[0003] However, existing technologies still have the following shortcomings: First, visual recognition technology is sensitive to changes in light. Conditions such as strong light, shadows, or low illumination at night can significantly reduce recognition accuracy and affect the vehicle's real-time perception capabilities. Second, visual recognition road spectrum technology relies on high-definition cameras to obtain high-quality images. Once the image is blurred, obscured, or there are interference factors, the recognition effect will be impaired, reducing the reliability of the system. Third, in extreme environments such as rain, snow, fog, and haze, it is difficult for the camera to clearly capture road images, resulting in a decline in visual recognition effectiveness and a significant decrease in recognition accuracy. Fourth, this technology usually requires the support of high-precision sensors and complex image processing algorithms, which makes the cost of the visual recognition system relatively high and limits its widespread application in mid- and low-end models. Summary of the Invention

[0004] In view of this, the embodiments of the present application provide a method, device, electronic device and storage medium for visual recognition of the road spectrum of an entire vehicle to solve the problems of the prior art such as sensitivity to light, reliance on high-quality images, poor environmental adaptability and high system cost.

[0005] In a first aspect of an embodiment of the present application, a method for visual recognition of a whole-vehicle road spectrum is provided, comprising: collecting multispectral image data using a vehicle-mounted camera, wherein the multispectral image data includes an image of a road surrounding the vehicle; preprocessing the multispectral image data, and normalizing the preprocessed multispectral image data; inputting the normalized multispectral image data into a convolutional neural network for feature extraction to obtain a feature vector, wherein the convolutional neural network uses a detection algorithm to extract features of edge information, texture information, and color information in the road image; classifying and identifying feature information in the road image based on the feature vector, generating a recognition result, and outputting the recognition result to the vehicle's system so that the vehicle makes driving decisions based on the recognition result.

[0006] According to a second aspect of an embodiment of the present application, a visual recognition device for a whole-vehicle road spectrum is provided, comprising: an acquisition module configured to acquire multispectral image data using a vehicle-mounted camera, wherein the multispectral image data includes images of roads around the vehicle; a preprocessing module configured to preprocess the multispectral image data and perform normalization on the preprocessed multispectral image data; an extraction module configured to input the normalized multispectral image data into a convolutional neural network for feature extraction to obtain a feature vector, wherein the convolutional neural network uses a detection algorithm to perform feature extraction on edge information, texture information, and color information in the road image; and a recognition module configured to classify and recognize feature information in the road image according to the feature vector, generate a recognition result, and output the recognition result to the vehicle's system so that the vehicle makes driving decisions based on the recognition result.

[0007] According to a third aspect of an embodiment of the present application, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the steps of the above method are implemented when the processor executes the computer program.

[0008] According to a fourth aspect of an embodiment of the present application, a readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the above method are implemented.

[0009] At least one of the above technical solutions adopted in the embodiments of the present application can achieve the following beneficial effects:

[0010] By using an on-board camera to collect multispectral image data, wherein the multispectral image data includes images of the road around the vehicle; preprocessing the multispectral image data, and normalizing the preprocessed multispectral image data; inputting the normalized multispectral image data into a convolutional neural network for feature extraction to obtain a feature vector, wherein the convolutional neural network uses a detection algorithm to extract features of edge information, texture information, and color information in the road image; based on the feature vector, classifying and identifying the feature information in the road image, generating a recognition result, and outputting the recognition result to the vehicle's system, so that the vehicle can make driving decisions based on the recognition result. This application can improve the adaptability of visual recognition road spectrum technology in complex lighting and extreme environments, reduce dependence on image quality, and reduce system costs while ensuring accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0012] Figure 1 This is a flow chart of a method for visually recognizing a vehicle's road profile provided in an embodiment of the present application;

[0013] Figure 2 This is a schematic diagram of the structure of the visual recognition device for the entire vehicle road spectrum provided in an embodiment of the present application;

[0014] Figure 3 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0015] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.

[0016] Existing visual road recognition technology is primarily based on computer vision and image processing. It uses high-definition cameras to capture road images, then uses a series of algorithms to analyze the road information in the images to provide navigation, control, or decision support for vehicles. Specifically, existing visual road recognition technology typically includes the following main steps:

[0017] 1. Image acquisition: Real-time images of the road are acquired through high-definition cameras.

[0018] 2. Image preprocessing: De-noising, enhancement and other processing are performed on the collected images to improve image clarity and quality.

[0019] 3. Feature extraction: Use image processing algorithms to extract key road features from the image, such as lane lines, edges of traffic signs, and road surface texture.

[0020] 4. Feature matching and recognition: The extracted features are matched with a preset pattern library to identify the different types and states of the road surface. The recognition results will be used for vehicle navigation and decision support.

[0021] Although existing technologies have achieved the recognition of features on the road, they also have significant shortcomings, which affect their effectiveness and widespread application in practical applications. The existing technologies still have the following defects:

[0022] 1. Sensitive to light: Visual recognition relies on image processing, and changes in light can affect image quality and recognition accuracy. For example, strong light, shadows, or low light at night can affect the image quality captured by the camera, which in turn affects the algorithm's feature extraction and judgment.

[0023] 2. Reliance on high-quality images: Image clarity is fundamental to effective visual recognition. If the camera image is blurry, obstructed, or subject to other interference, it will affect the system's ability to effectively identify road features, leading to inaccurate navigation or control decisions.

[0024] 3. Poor environmental adaptability: In extreme weather conditions (such as rain, snow, fog, and haze), cameras struggle to clearly capture road images, causing the visual recognition system's performance to degrade and recognition accuracy to drop significantly. This poses a challenge to the reliability of driver assistance systems in inclement weather or complex road conditions.

[0025] 4. High cost: This technology relies on high-precision cameras and complex computing algorithms, which are costly. Its application is particularly limited in mid- and low-end models, which restricts the popularity of this technology in a wider market.

[0026] Therefore, to sum up, although the existing visual recognition road spectrum solutions have certain road recognition and decision support capabilities, they have obvious shortcomings in terms of lighting adaptability, image quality dependence, adaptability to extreme environments, and cost.

[0027] In view of the problems existing in the above-mentioned prior art, this application proposes a vision-based whole-vehicle road spectrum recognition method. This application captures road images through on-board visible light cameras and infrared cameras, uses image denoising algorithms and image enhancement technologies in artificial intelligence technology to pre-process the input road images, and normalizes the road images through artificial intelligence algorithms to lay the foundation for subsequent feature extraction and recognition. It uses artificial intelligence edge detection algorithms to extract edge information from road images, uses artificial intelligence algorithms to conduct in-depth analysis of the texture and color of road images, extracts features related to road surface characteristics, and classifies and recognizes the extracted features. It can accurately distinguish different categories of objects such as roads, vehicles, pedestrians, etc., and finally outputs the processed recognition results to the autonomous driving system or other related systems to support the vehicle's driving decisions and path planning.

[0028] The contents of the technical solution of this application are described in detail below with reference to the accompanying drawings and specific embodiments.

[0029] Figure 1 This is a flow chart of the visual recognition method of the vehicle road spectrum provided by the embodiment of the present application. Figure 1 As shown, the visual recognition method of the entire vehicle road spectrum may specifically include:

[0030] S101, collecting multispectral image data using a vehicle-mounted camera, wherein the multispectral image data includes an image of a road surrounding the vehicle;

[0031] S102, preprocessing the multispectral image data and normalizing the preprocessed multispectral image data;

[0032] S103, inputting the normalized multispectral image data into a convolutional neural network for feature extraction to obtain a feature vector, wherein the convolutional neural network uses a detection algorithm to extract features from edge information, texture information, and color information in the road image;

[0033] S104 , classifying and identifying feature information in the road image according to the feature vector, generating a recognition result, and outputting the recognition result to the vehicle system so that the vehicle can make driving decisions based on the recognition result.

[0034] In some embodiments, collecting multispectral image data using a vehicle-mounted camera includes:

[0035] The vehicle-mounted visible light camera is used to collect road image data within the visible light range to obtain road feature information under normal lighting conditions;

[0036] Use infrared cameras to collect road image data within the infrared spectrum in low-light or nighttime environments to obtain road feature information under weak light conditions;

[0037] The road image data collected by the vehicle-mounted visible light camera and the infrared camera are synthesized to obtain multispectral image data, wherein the multispectral image data contains road images with different spectral information.

[0038] Specifically, the system uses an onboard visible light camera installed on the vehicle to collect road image data under normal lighting conditions. Capable of capturing images within the visible spectrum, this visible light camera is particularly well-suited for collecting road image information during the day or in well-lit environments. The data collected by this camera clearly captures key road features such as lane markings, traffic signs, and road surface texture, providing essential data support for subsequent recognition.

[0039] Furthermore, the vehicle is equipped with an infrared camera capable of capturing road image data within the infrared spectrum in low-light or nighttime environments. Infrared cameras offer excellent imaging capabilities at night or in low-light environments, and the infrared images they capture contain information about road features that are difficult to capture with visible light. For example, during nighttime driving, infrared cameras can identify lane boundaries, pedestrians, and other obstacles on the road, complementing the limited recognition capabilities of visible light cameras in low-light environments.

[0040] Furthermore, after completing image acquisition, the system synthesizes the road image data captured by the visible light camera and the infrared camera to produce a road image containing multispectral information. Specifically, the synthesis step uses an image processing algorithm to superimpose or fuse the visible light and infrared images at corresponding pixel locations in space to form a comprehensive image with more spectral information. This multispectral image data contains road feature information from both the visible and infrared spectra, thereby improving the image's adaptability and recognition accuracy in complex lighting conditions.

[0041] Through the multispectral image data acquisition and synthesis processing in the above-mentioned embodiment, the system can obtain more complete road information in different lighting environments, providing reliable basic data support for subsequent image recognition and vehicle decision-making.

[0042] In some embodiments, preprocessing the multispectral image data and normalizing the preprocessed multispectral image data includes:

[0043] Image denoising technology is used to filter out noise information in multispectral image data, and image enhancement technology is used to enhance the denoised multispectral image data; the enhanced multispectral image data is normalized to standardize the parameters of the road image.

[0044] Specifically, image denoising techniques are applied to remove noise from multispectral image data. Multispectral image data typically consists of a superposition of visible and infrared images, which can introduce noise due to sensor characteristics or environmental factors, such as light speckles and blurred edges that appear during the imaging process. Denoising uses methods such as Gaussian filtering and median filtering to reduce unnecessary random noise in the image while preserving key road edge information and texture features, thereby improving overall image clarity.

[0045] After denoising, the multispectral image data undergoes image enhancement to improve contrast and detail. Specifically, image enhancement techniques enhance road features by adjusting brightness, contrast, and edge sharpening. This enhancement process sharpens features such as lane markings, traffic signs, and road surface textures, improving visibility in complex lighting conditions. This significantly enhances the recognition accuracy of the enhanced image in subsequent steps.

[0046] The enhanced image data is then normalized. Normalization standardizes the image's size, scale, and color parameters, for example, adjusting the image to a uniform resolution and normalizing the pixel value range for each spectral channel. This normalization ensures consistent input format and dynamic range for all image data, helping to eliminate image differences caused by different cameras, acquisition conditions, and environmental variations, providing standardized data input for subsequent convolutional neural network processing.

[0047] Through the denoising, enhancement and normalization processing in the above embodiment, the system can generate clear and standardized multispectral image data, laying a stable data foundation for subsequent road image feature extraction and recognition.

[0048] In some embodiments, the normalized multispectral image data is input into a convolutional neural network for feature extraction to obtain a feature vector, including:

[0049] The normalized multispectral image data is input into the convolutional neural network and features are extracted through the convolutional layer, pooling layer and fully connected layer in sequence;

[0050] Among them, in the convolution layer, a sliding convolution operation is performed on the multispectral image data using a preset convolution kernel, and an activation function is applied after the convolution layer to perform nonlinear transformation on the feature map;

[0051] In the pooling layer, the feature map output by the convolution layer is pooled to reduce the size of the feature map and retain the main feature information;

[0052] The pooled feature map is flattened in the fully connected layer, and a linear transformation is performed based on the preset weights and bias terms to obtain the feature vector.

[0053] Specifically, the normalized multispectral image data is first input into the first layer of the convolutional neural network, namely the convolution layer. In the convolution layer, a sliding convolution operation is performed on the image data using a preset convolution kernel. Specifically, the convolution kernel slides pixel by pixel on the image according to the set step size to calculate the local features of the area, such as edge information, texture features, and color changes. Through the convolution operation, the system can identify local structural features in the image and generate a preliminary feature map. In addition, after each convolution operation, an activation function (such as ReLU) is applied to the output feature map to introduce a nonlinear transformation to enhance the network's ability to express complex road features.

[0054] The feature maps output by the convolutional layer are then processed in the pooling layer. The pooling layer performs a reduction operation on the feature map using either max pooling or average pooling. The primary purpose of pooling is to reduce the resolution of the feature map, minimizing the amount of data while retaining key feature information. By reducing the feature map, the system can preserve key information while reducing computational complexity, thereby improving the model's operational efficiency and enhancing its robustness to changes such as image displacement and rotation.

[0055] The pooled feature map is then fed into a fully connected layer for further processing. In this layer, the multidimensional feature map output by pooling is flattened into a one-dimensional vector to facilitate connections with each neuron. A linear transformation is then applied to the flattened feature vector using a preset weight matrix and bias term. This linear transformation allows the system to integrate and enhance important features within the feature vector, allowing the extracted features to be used for classification and recognition.

[0056] Through the convolution, pooling, and fully connected operations described in this embodiment, the system ultimately generates a set of feature vectors representing the road environment. These feature vectors contain key features from the road image, such as lane boundaries, obstacle outlines, and other road-related information. These feature vectors serve as input to subsequent classification and recognition modules, supporting the vehicle's driving decisions and path planning.

[0057] In some embodiments, a sliding convolution operation is performed on the multispectral image data using a preset convolution kernel in a convolution layer, including:

[0058] Convert the multispectral image data into a format that can be processed by the convolutional neural network to form a three-dimensional array structure;

[0059] The size, step size and padding of the convolution kernel are set, and the convolution kernel is used to perform sliding convolution on the multispectral image data to extract local feature information from the road image.

[0060] Using the preset edge detection algorithm, the edge information in the road image is extracted, and the texture information and color information in the road image are extracted to generate a convolution feature map related to the road surface characteristics.

[0061] Specifically, the normalized multispectral image data is first converted into a format that can be processed by a convolutional neural network (CNN), forming a three-dimensional array structure. For example, for a multispectral image containing RGB channels for visible light and infrared channels, the multispectral information includes both visible and infrared spectra. This structure enables the CNN to deeply process the multidimensional features of the image.

[0062] Furthermore, the convolutional layer configures the kernel size, stride, and padding. The kernel size determines the size of the image area covered by each convolution operation. For example, commonly used 3*3 or 5*5 kernels are used to capture local image features. The stride determines the interval at which the kernel slides across the image. A smaller stride can extract more detailed features, while padding (such as zero padding) is used to preserve the image boundary information after the convolution operation, preventing the loss of edge features.

[0063] After configuring the convolution kernel, a sliding convolution operation is performed on the multispectral image data. During each convolution operation, the kernel performs a matrix dot product with the local image region, generating an output value that extracts the local features of that region. These local features include information such as road edges, textures, and color variations. The sliding convolution operation within the convolutional layer extracts local features layer by layer, generating a series of preliminary feature maps.

[0064] Furthermore, during the convolution operation, this embodiment utilizes a preset edge detection algorithm, such as the Sobel or Canny algorithm, to accurately extract edge information from road images. This edge detection process identifies the edges of lane lines and other key objects by capturing areas with significant pixel value changes. Simultaneously, the convolution layer analyzes the image's texture and color information, further extracting texture features and color patterns related to road surface characteristics through multi-channel convolution operations, thereby generating a convolution feature map containing various road features.

[0065] The convolution feature map generated by the convolution operation in the above embodiment contains rich road surface information, which can provide depth information support for subsequent pooling layers and fully connected layers, so that subsequent classification and recognition operations can make accurate judgments based on comprehensive road features.

[0066] The following is a detailed description of the structure and processing of the convolutional layer of the convolutional neural network of the present application with reference to a specific example, which may include the following:

[0067] The road images captured by the vehicle's visible light and infrared cameras are converted into a digital form that can be understood by CNNs. This is typically a three-dimensional array (dimensions: length, width, and color). For example, in a three-dimensional array unfolded in a planar diagram, 5*5*3: 5*5 represents the length and width; 3 represents the three channels R, G, and B.

[0068] 1) Define the convolution kernel (filter)

[0069] Set the filter size: that is, the length and width of the filter. Common sizes are 3*3 and 5*5.

[0070] Set the step size: the interval at which the filter slides on the input data.

[0071] Padding: This is the process of adding extra zero-valued pixels at the boundaries of the input data; its purpose is to resolve the image information loss caused by boundary convolution.

[0072] Set the convolution type: types include standard convolution, transposed convolution (for upsampling), and dilated convolution (for expanding the receptive field).

[0073] Select the initialization method: Set weights and bias items, including zero initialization, random initialization, Xavler initialization, and He initialization.

[0074] Convolution function to be considered: generally used after convolution operation.

[0075] 2) Convolution operation

[0076] Explain in detail how the final result 1 is convolved through the filter W0 and the bias term h0.

[0077] R(0*1+0*1+0*-1)+(0*-1+0*0+1*1)+(0*-1+0*-1+1*0)=1

[0078] G(0*-1+0*0+0*-1)+(0*0+1*0+1*-1)+(0*1+0*-1+2*0)=-1

[0079] B(0*0+0*1+0*0)+(0*1+2*0+0*1)+(0*0+0*-1+0*1)=0

[0080] The final result is R (red channel) + G (green channel) + B (blue channel) + b0 (bias term), which is 1 + (-1) + 0 + 1 = 1

[0081] 3) Sliding convolution kernel (filter)

[0082] According to the step size of 2, the convolution kernel is sliding to calculate other values.

[0083] 4) Generate convolutional feature map

[0084] Generate feature map: usually refers to the final result matrix obtained after the convolution operation is completed.

[0085] 5) Apply activation function

[0086] Purpose: To introduce nonlinearity between layers of a neural network so that the network can learn and represent complex functions.

[0087] Common activation functions: ReLU, Sigmold, Tanh. Using the activation function, a nonlinear transformation is applied to each neuron (that is, the value in each input matrix).

[0088] The general location of the activation function: after the convolutional layer and after the fully connected layer.

[0089] Furthermore, the structure and processing of the pooling layer of the convolutional neural network of the present application will be described in detail below in combination with the above examples, which may include the following:

[0090] 1) Define the pooling window

[0091] The pooling window, also known as the pooling kernel, is a small rectangular area used to extract information from the input feature map, usually 2*2 or 3*3.

[0092] 2) Select pooling operation

[0093] The optional pooling operations are: maximum pooling and average pooling.

[0094] 3) Application pooling operation

[0095] Input feature map: refers to the convolution feature map generated by the previous convolution layer.

[0096] 4) Sliding pooling window

[0097] The pooling window moves across the input feature map with a specified stride.

[0098] 5) Generate pooled feature map

[0099] You can use maximum pooling or average pooling.

[0100] 6) Pass the pooled feature map to the next layer

[0101] The next layer may be a convolutional layer or a fully connected layer. The input matrix of the next layer is the feature map after pooling.

[0102] In some embodiments, the pooled feature map is flattened in the fully connected layer, and a linear transformation is performed based on preset weights and bias terms, including:

[0103] Flatten the pooled feature map into a one-dimensional vector, where each channel in the pooled feature map is expanded into a long vector, and the long vectors of all channels are concatenated together to form a one-dimensional vector;

[0104] Assign weights and bias terms to each element in the one-dimensional vector to form a weight matrix and bias vector;

[0105] According to the weight matrix and bias vector, the one-dimensional vector is linearly transformed to obtain the feature vector output by the fully connected layer.

[0106] Specifically, the multi-dimensional feature map output by the pooling layer is flattened into a one-dimensional vector so that it can be input into the fully connected layer for further processing. Specifically, each channel of the pooled feature map is flattened separately, converting the two-dimensional feature map of each channel into a long one-dimensional vector. Then, the long vectors of all channels are concatenated in sequence to form a large, unified one-dimensional vector, ensuring that all key information extracted from the feature map is fully expressed in the one-dimensional vector.

[0107] Furthermore, weights and bias terms are assigned to each element in the one-dimensional vector to form a weight matrix and bias vector. Specifically, each neuron in the fully connected layer (i.e., each value in the one-dimensional vector) is connected to each element of the input vector, and these connections are weighted by the weight matrix. Each value in the weight matrix represents the degree of influence of the corresponding input feature on the neuron output, reflecting the degree of attention paid by the neuron to different input features. In addition, each neuron is also equipped with a bias term to adjust the overall output level to ensure that the system output is within a reasonable numerical range.

[0108] The system then performs a linear transformation on the flattened vector. By performing this linear transformation on the input vector, the net input value of each neuron is calculated, resulting in a feature vector output by the fully connected layer. This feature vector condenses various road features extracted from the multispectral image data, including key information such as lane boundaries, road surface texture, and obstacles.

[0109] Furthermore, after completing the linear transformation, an activation function (such as ReLU or Sigmoid) is applied to the output of the fully connected layer to introduce nonlinear characteristics and ensure the system's ability to effectively discern complex road features. The activation function applies a nonlinear transformation to the output of each neuron, giving the fully connected layer's output a characteristic bias that more accurately reflects the actual road image.

[0110] Through the flattening, weight distribution, linear transformation and activation function processing in the above-mentioned embodiment, the fully connected layer ultimately generates a feature vector containing important features of the road environment, providing comprehensive and reliable feature data support for subsequent road classification and recognition.

[0111] The following is a detailed description of the structure and processing of the fully connected layer (FC) of the convolutional neural network of the present application, with reference to a specific example. The following may be included:

[0112] 1) Receive feature map

[0113] Input matrix: It is the feature map after all pooling and / or activation function processing. It is a multidimensional vector with the following dimensions:

[0114] batch_size: batch size

[0115] heigh: the height of the feature map

[0116] width: width of the feature map

[0117] channels: number of feature channels

[0118] 2) Flatten the feature map

[0119] Purpose: Flatten the multi-dimensional feature map into a one-dimensional feature map.

[0120] Implementation: Each channel in the feature map is "expanded into a long vector" and the vectors of all channels are concatenated together to form a large one-dimensional vector. The length of the flattened vector is batch_size*height*width*channels.

[0121] 3) Weight distribution

[0122] Each neuron in the fully connected layer (i.e., each value in the one-dimensional vector) is connected to every element in the input matrix, and each connection has a corresponding weight. These weights form the weight matrix, which reflects the importance the neuron attaches to the input features. In addition, each neuron has a bias term (Bias) to adjust the overall response level of the neuron.

[0123] 4) Linear transformation

[0124] Perform a linear transformation on the input vector using the following formula: Net Input = W*x+b, where W represents the weight matrix, x represents the input vector, and b represents the bias vector.

[0125] Use this formula to perform a linear transformation on the input vector to obtain the net input of each neuron (that is, the final output result of the fully connected layer).

[0126] 5) Activation function application (again)

[0127] By using the activation function again for final processing, the result may remain unchanged (such as the positive interval of ReLU), or the result may be compressed to a certain range (such as Sigmoid, Tanh), ensuring that the output meets the specifications and has a clear tendency.

[0128] In some embodiments, classifying and identifying feature information in a road image based on a feature vector to generate a recognition result includes:

[0129] The feature vector is input into the classifier, and the classifier is used to analyze the feature vector and classify the feature information into different road object categories based on the edge information, texture information and color information in the road image;

[0130] The recognition results are generated based on the classification results, wherein the recognition results include road type, road surface status, and location and category information of different target objects on the road.

[0131] Specifically, the feature vector output by the convolutional neural network is fed into a classifier, which uses a machine learning algorithm to analyze the information in the feature vector. This analysis is based on information such as image edges, texture, and color contained in the feature vector. Edge information in the feature vector helps identify the outlines of lane lines and obstacles, texture information helps analyze the material properties or roughness of the road surface, and color information helps distinguish different types of road signs and obstacles.

[0132] Furthermore, the classifier analyzes this information and assigns features to different road object categories, such as lane markings, traffic signs, obstacles, and pedestrians. The classification results provide a detailed picture of the actual road image and provide the autonomous driving system with highly accurate environmental perception data. For example, if the classifier detects an area with lane markings, it can classify it as part of a lane; if it detects an irregular shape, it may classify it as an obstacle or pedestrian.

[0133] After completing the classification, the system generates a recognition result based on the classification results. The recognition result includes not only the road type (such as urban road, highway, etc.) and road surface condition (such as flat, slippery), but also the location and category information of different target objects. For example, the recognition result may include information such as "there is a pedestrian 10 meters ahead" or "lane line detected at the edge of the right lane." After these recognition results are formatted, they will be directly output to the autonomous driving system or other relevant control systems of the vehicle.

[0134] Ultimately, the generated recognition results provide essential support for the vehicle's driving decisions and path planning. For example, when the recognition results detect a pedestrian or obstacle, the autonomous driving system can use this information to slow down or avoid the obstacle. Guided by the road type or lane boundary recognition results, the autonomous driving system can plan a safe driving path. Through the classification and recognition process in this embodiment, the system can accurately perceive the surrounding environment based on various road characteristics, providing critical perception data for the vehicle's intelligent driving decisions.

[0135] According to the technical solution provided in the embodiments of the present application, the technical solution of the present application has at least the following advantages:

[0136] This application uses a combination of on-board visible light cameras and infrared cameras to capture road image data. The visible light camera can obtain clear road feature information under normal lighting conditions, while the infrared camera can still capture road images normally at night or in low-light environments, effectively compensating for the impact of insufficient nighttime lighting on image quality and enabling all-weather road spectrum recognition. This spectral complementarity scheme not only improves the overall quality of the input image, but also ensures the continuity and stability of road spectrum recognition under various lighting conditions. Compared with the solution using only binocular visible light cameras, it significantly improves the recognition rate and image clarity at night.

[0137] The technical solution of this application utilizes artificial intelligence algorithms for road spectrum recognition, avoiding the complex processes and insufficient precision of traditional image recognition algorithms. Through deep learning feature extraction and classification techniques, this solution can accurately extract key features such as road edges, road surface texture, and target objects from the input image, greatly improving recognition accuracy. Compared with traditional image processing technologies, this solution not only achieves accurate recognition but also increases the recognition rate, enabling autonomous driving systems to obtain reliable road environment information in a shorter time.

[0138] By replacing the high-cost LiDAR with a combination of onboard visible light and infrared cameras, the technical solution of this application significantly reduces the overall system cost. Visible light and infrared cameras not only offer a cost advantage, but also meet the requirements of autonomous driving in terms of road environment perception resolution and accuracy, thereby promoting the popularization of this road spectrum recognition technology in mid- and low-end vehicles.

[0139] By combining visible light and infrared cameras, this application effectively overcomes the image quality degradation of traditional visible light cameras in low-light conditions. Infrared cameras can continuously provide clear image input in low-light or inclement weather conditions, compensating for the limitations of visible light cameras at night or in poor lighting conditions. This further improves road spectrum recognition in complex lighting environments and significantly enhances nighttime driving safety.

[0140] This application's technical solution uses artificial intelligence to optimize the entire image process, from preprocessing and feature extraction to classification and recognition. This reduces redundant steps in traditional image recognition processes and significantly improves recognition speed and efficiency. Compared with traditional solutions, the AI-based recognition process has higher processing efficiency and accuracy, providing fast and reliable data support for real-time decision-making in autonomous driving systems.

[0141] The following are device embodiments of the present application, which can be used to implement the method embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.

[0142] Figure 2 This is a schematic diagram of the structure of the visual recognition device for the entire vehicle road spectrum provided in the embodiment of the present application. Figure 2 As shown, the visual recognition device for the entire vehicle road spectrum includes:

[0143] The acquisition module 201 is configured to acquire multispectral image data using a vehicle-mounted camera, wherein the multispectral image data includes an image of a road surrounding the vehicle;

[0144] The preprocessing module 202 is configured to preprocess the multispectral image data and perform normalization processing on the preprocessed multispectral image data;

[0145] The extraction module 203 is configured to input the normalized multispectral image data into a convolutional neural network for feature extraction to obtain a feature vector, wherein the convolutional neural network uses a detection algorithm to extract features from edge information, texture information, and color information in the road image;

[0146] The recognition module 204 is configured to classify and recognize feature information in the road image according to the feature vector, generate a recognition result, and output the recognition result to the vehicle system so that the vehicle can make driving decisions based on the recognition result.

[0147] In some embodiments, Figure 2 The acquisition module 201 uses a vehicle-mounted visible light camera to acquire road image data within the visible light range to obtain road feature information under normal lighting conditions; uses an infrared camera to acquire road image data within the infrared spectrum range in low-light or nighttime environments to obtain road feature information under weak lighting conditions; and synthesizes the road image data acquired by the vehicle-mounted visible light camera and the infrared camera to obtain multispectral image data, wherein the multispectral image data includes road images with different spectral information.

[0148] In some embodiments, Figure 2 The pre-processing module 202 uses image denoising technology to filter out noise information in the multispectral image data, and uses image enhancement technology to enhance the denoised multispectral image data; and normalizes the enhanced multispectral image data to standardize the parameters of the road image.

[0149] In some embodiments, Figure 2 The extraction module 203 inputs the normalized multispectral image data into the convolutional neural network, and performs feature extraction in sequence through the convolution layer, the pooling layer and the fully connected layer; wherein, in the convolution layer, a sliding convolution operation is performed on the multispectral image data using a preset convolution kernel, and an activation function is applied after the convolution layer to perform a nonlinear transformation on the feature map; in the pooling layer, a pooling operation is performed on the feature map output by the convolution layer to reduce the size of the feature map and retain the main feature information; in the fully connected layer, the pooled feature map is flattened, and a linear transformation is performed based on the preset weight and bias term to obtain a feature vector.

[0150] In some embodiments, Figure 2 The extraction module 203 converts the multispectral image data into a format that can be processed by the convolutional neural network to form a three-dimensional array structure; sets the size, step size and padding of the convolution kernel, and facilitates the convolution kernel to perform a sliding convolution operation on the multispectral image data to extract local feature information in the road image; uses a preset edge detection algorithm to extract edge information from the road image, and extracts texture information and color information from the road image to generate a convolution feature map related to the road surface characteristics.

[0151] In some embodiments, Figure 2The extraction module 203 flattens the pooled feature map into a one-dimensional vector, wherein each channel in the pooled feature map is expanded into a long vector, and the long vectors of all channels are connected together to form a one-dimensional vector; a weight and a bias term are assigned to each element in the one-dimensional vector to form a weight matrix and a bias vector; and a linear transformation is performed on the one-dimensional vector according to the weight matrix and the bias vector to obtain the feature vector output by the fully connected layer.

[0152] In some embodiments, Figure 2 The recognition module 204 inputs the feature vector into a classifier, uses the classifier to analyze the feature vector, and classifies the feature information into different road object categories based on edge information, texture information, and color information in the road image; generates a recognition result based on the classification result, wherein the recognition result includes road type, road surface status, and location and category information of different target objects on the road.

[0153] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0154] Figure 3 Schematic diagram of the structure of the electronic device 3 provided in the embodiment of the present application. Figure 3 As shown, the electronic device 3 of this embodiment includes: a processor 301, a memory 302, and a computer program 303 stored in the memory 302 and executable on the processor 301. When the processor 301 executes the computer program 303, the steps of the above-mentioned method embodiments are implemented. Alternatively, when the processor 301 executes the computer program 303, the functions of the modules / units in the above-mentioned device embodiments are implemented.

[0155] For example, computer program 303 may be divided into one or more modules / units, which are stored in memory 302 and executed by processor 301 to implement the present application. One or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of computer program 303 in electronic device 3.

[0156] The electronic device 3 may be a desktop computer, a notebook, a PDA, a cloud server or other electronic device. The electronic device 3 may include but is not limited to a processor 301 and a memory 302. Those skilled in the art will understand that Figure 3 It is only an example of electronic device 3 and does not constitute a limitation of electronic device 3. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the electronic device may also include input and output devices, network access devices, buses, etc.

[0157] The processor 301 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0158] The memory 302 can be an internal storage unit of the electronic device 3, such as a hard drive or memory of the electronic device 3. The memory 302 can also be an external storage device of the electronic device 3, such as a plug-in hard drive, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. equipped on the electronic device 3. Furthermore, the memory 302 can include both an internal storage unit of the electronic device 3 and an external storage device. The memory 302 is used to store computer programs and other programs and data required by the electronic device. The memory 302 can also be used to temporarily store data that has been output or is about to be output.

[0159] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0160] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0161] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0162] In the embodiments provided in this application, it should be understood that the disclosed apparatus / computer equipment and methods can be implemented in other ways. For example, the apparatus / computer equipment embodiments described above are merely schematic. For example, the division of modules or units is merely a logical function division. In actual implementation, there may be other division methods. Multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection of the apparatus or unit, which may be electrical, mechanical or other forms.

[0163] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0164] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0165] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. The computer program may include computer program code, which may be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc.

[0166] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A method for visually recognizing a vehicle's road profile, characterized in that: include: Collecting multispectral image data using a vehicle-mounted camera, wherein the multispectral image data includes an image of a road surrounding the vehicle; Preprocessing the multispectral image data and normalizing the preprocessed multispectral image data; Inputting the normalized multispectral image data into a convolutional neural network for feature extraction to obtain a feature vector, wherein the convolutional neural network uses a detection algorithm to extract features from edge information, texture information, and color information in the road image; Based on the feature vector, feature information in the road image is classified and identified, a recognition result is generated, and the recognition result is output to the system of the vehicle so that the vehicle makes a driving decision based on the recognition result.

2. The method according to claim 1, characterized in that The method of collecting multispectral image data using a vehicle-mounted camera includes: The vehicle-mounted visible light camera is used to collect road image data within the visible light range to obtain road feature information under normal lighting conditions; Use infrared cameras to collect road image data within the infrared spectrum in low-light or nighttime environments to obtain road feature information under weak light conditions; The road image data collected by the vehicle-mounted visible light camera and the infrared camera are synthesized to obtain the multispectral image data, wherein the multispectral image data includes road images with different spectral information.

3. The method according to claim 1, characterized in that The preprocessing of the multispectral image data and normalization of the preprocessed multispectral image data includes: The noise information in the multispectral image data is filtered out by using image denoising technology, and the denoised multispectral image data is enhanced by using image enhancement technology; the enhanced multispectral image data is normalized to standardize the parameters of the road image.

4. The method according to claim 1, wherein The normalized multispectral image data is input into a convolutional neural network for feature extraction to obtain a feature vector, including: The normalized multispectral image data is input into a convolutional neural network, and feature extraction is performed through a convolutional layer, a pooling layer, and a fully connected layer in sequence; Wherein, in the convolution layer, a sliding convolution operation is performed on the multispectral image data using a preset convolution kernel, and an activation function is applied after the convolution layer to perform a nonlinear transformation on the feature map; Performing a pooling operation on the feature map output by the convolutional layer in the pooling layer to reduce the size of the feature map and retain main feature information; The pooled feature map is flattened in the fully connected layer, and a linear transformation is performed based on preset weights and bias items to obtain the feature vector.

5. The method according to claim 4, characterized in that The sliding convolution operation is performed on the multispectral image data using a preset convolution kernel in the convolution layer, comprising: Converting the multispectral image data into a format processable by a convolutional neural network to form a three-dimensional array structure; The size, step size, and padding of the convolution kernel are set to facilitate the convolution kernel to perform a sliding convolution operation on the multispectral image data to extract local feature information from the road image; Using a preset edge detection algorithm, edge information in the road image is extracted, and texture information and color information in the road image are extracted to generate a convolution feature map related to road surface characteristics.

6. The method according to claim 4, characterized in that The flattening of the pooled feature map in the fully connected layer and performing a linear transformation based on preset weights and bias terms includes: Flattening the pooled feature map into a one-dimensional vector, wherein each channel in the pooled feature map is expanded into a long vector, and the long vectors of all channels are connected together to form the one-dimensional vector; Assigning a weight and a bias term to each element in the one-dimensional vector to form a weight matrix and a bias vector; Perform a linear transformation on the one-dimensional vector according to the weight matrix and the bias vector to obtain a feature vector output by the fully connected layer.

7. The method according to claim 1, characterized in that The classifying and identifying the feature information in the road image according to the feature vector to generate a recognition result includes: Inputting the feature vector into a classifier, analyzing the feature vector using the classifier, and classifying the feature information into different road object categories based on edge information, texture information, and color information in the road image; A recognition result is generated based on the classification result, wherein the recognition result includes road type, road surface condition, and location and category information of different target objects on the road.

8. A visual recognition device for a vehicle's road profile, characterized in that: include: an acquisition module configured to acquire multispectral image data using a vehicle-mounted camera, wherein the multispectral image data includes an image of a road surrounding the vehicle; a preprocessing module, configured to preprocess the multispectral image data and perform normalization processing on the preprocessed multispectral image data; an extraction module configured to input the normalized multispectral image data into a convolutional neural network for feature extraction to obtain a feature vector, wherein the convolutional neural network uses a detection algorithm to extract features from edge information, texture information, and color information in the road image; The recognition module is configured to classify and recognize feature information in the road image according to the feature vector, generate a recognition result, and output the recognition result to the system of the vehicle so that the vehicle makes driving decisions based on the recognition result.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.