A method, device, and electronic device for road surface condition perception based on road images and tire mechanics.

By combining multimodal feature extraction and cross-modal attention fusion with road surface images and tire acceleration signals, the accuracy and robustness of road surface condition recognition in complex environments are solved, improving the vehicle's recognition capability and safety under different road surface conditions.

CN122493418APending Publication Date: 2026-07-31HANGZHOU DIANZI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU DIANZI UNIV
Filing Date
2026-07-02
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies lack accuracy and robustness in road condition recognition under complex environments. Visual sensors struggle to reflect the true contact state between the tire and the road surface, tire signal methods lose high-frequency information, and visual-tire signal fusion methods suffer from the problem of high-dimensional visual features dominating low-dimensional dynamic information.

Method used

By acquiring multimodal data, including road surface images, tire acceleration signal sequences, and preset state information, feature extraction is performed. Combined with multi-scale one-dimensional temporal feature extraction and cross-modal attention fusion, modulation weights are generated to modulate the semantic features of the images, and the road surface state classification results are output.

Benefits of technology

It improves the accuracy and robustness of identifying conditions such as water accumulation, damage, asphalt, and cement road surfaces, enhances the vehicle's safety and intelligent perception capabilities in complex environments, and provides driving risk warnings and route decision-making basis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122493418A_ABST
    Figure CN122493418A_ABST
Patent Text Reader

Abstract

This application provides a method, device, and electronic device for road surface state perception based on road images and tire mechanics. The method includes: acquiring multimodal data; extracting features from the multimodal data to obtain image semantic features, tire dynamic features, and state features; concatenating the tire dynamic features and state features along the channel dimension; generating modulation weights based on the concatenated physical prior joint features; and using the modulation weights to modulate the image semantic features; and obtaining a road surface state classification result based on the modulated fusion features. This application extracts tire dynamic features based on tire acceleration signal sequences and their corresponding acceleration change sequences, reducing information loss caused by statistical features. Furthermore, by using tire dynamics and state features to modulate image semantic features, it alleviates modality competition and visual dominance problems in multimodal fusion, improving the accuracy and robustness of road surface state recognition in complex scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of autonomous driving technology, specifically to a method, device, and electronic device for perceiving road conditions based on road images and tire mechanics. Background Technology

[0002] Accurate identification of road conditions is crucial for vehicle active safety control, autonomous driving path planning, and advanced driver assistance systems. Different road conditions directly affect the contact characteristics, adhesion, and vehicle dynamics response between the tire and the road surface. Therefore, it is necessary to accurately identify road conditions such as asphalt, cement, damage, and water accumulation.

[0003] In existing solutions, road condition recognition primarily relies on visual sensors such as vehicle-mounted cameras, judging the condition based on features like color, texture, cracks, and reflections in road images. However, visual sensors are non-contact sensing methods, making it difficult to directly reflect the actual contact state between the tire and the road surface. In complex environments such as rain, fog, nighttime, strong light, shadows, water reflections, or camera malfunctions, image features are easily confused, leading to decreased recognition accuracy. While tire sensor-based recognition methods can utilize signals such as tire radial acceleration, tire pressure, and tire temperature to reflect the physical response of the road surface, current methods largely rely on manually statistical features, easily losing local vibration and transient impact information from high-frequency tire signals. Furthermore, existing visual and tire signal fusion methods often employ simple feature stitching, easily leading to high-dimensional visual features dominating and hindering the effective use of low-dimensional tire dynamics information.

[0004] Therefore, existing technologies still suffer from insufficient accuracy and robustness in identifying road conditions under complex environments. Summary of the Invention

[0005] In view of this, this application proposes a method, device, and electronic device for road surface condition perception based on road images and tire mechanics. Specifically, this application is implemented through the following technical solution:

[0006] According to a first aspect of the embodiments of this specification, a method for perceiving road surface conditions based on road images and tire mechanics is provided, the method comprising the following steps:

[0007] Step S1: Acquire multimodal data, which includes road surface images, tire acceleration signal sequences, and preset state information;

[0008] Step S2: Feature extraction is performed on the multimodal data to obtain multimodal features, which include image semantic features, tire dynamics features, and state features; wherein, by performing interpolation processing on the tire acceleration signal sequence, an acceleration signal change sequence is obtained, and by performing multi-scale one-dimensional time-series feature extraction on the multi-channel time-series input composed of the tire acceleration signal sequence and the acceleration signal change sequence, the tire dynamics features are obtained.

[0009] Step S3: The tire dynamics features and the state features are concatenated in the channel dimension to obtain physical prior joint features. Modulation weights are generated based on the physical prior joint features. The image semantic features are modulated using the modulation weights to obtain fused features.

[0010] Step S4: Input the fused features into the classifier and output the road surface state classification result.

[0011] According to a second aspect of the embodiments of this specification, a road surface condition sensing device based on road images and tire mechanics is provided, the device comprising:

[0012] The data acquisition unit is used to acquire multimodal data, which includes road surface images, tire acceleration signal sequences, and preset state information.

[0013] The feature extraction unit is used to extract features from the multimodal data to obtain multimodal features, which include image semantic features, tire dynamics features, and state features. Specifically, by performing interpolation processing on the tire acceleration signal sequence, an acceleration signal change sequence is obtained. The tire dynamics features are obtained by performing multi-scale one-dimensional temporal feature extraction on the multi-channel temporal input composed of the tire acceleration signal sequence and the acceleration signal change sequence.

[0014] The feature modulation unit is used to concatenate the tire dynamics features and the state features in the channel dimension to obtain physical prior joint features, generate modulation weights based on the physical prior joint features, and use the modulation weights to modulate the image semantic features to obtain fused features.

[0015] The road surface recognition unit is used to input the fused features into the classifier and output the road surface state classification result.

[0016] According to a third aspect of the embodiments of this specification, an electronic device is provided, comprising: a processor; and a computer-readable storage medium storing computer program instructions that, when executed by the processor, cause the processor to perform the method as described in the first aspect.

[0017] The embodiments of this application have at least the following technical effects:

[0018] This application embodiment can comprehensively utilize road surface images, tire acceleration signal sequences, and preset state information to identify and judge the current road surface state of the vehicle. Addressing the issue that single visual perception is easily affected by lighting, shadows, reflections, occlusion, and the similarity of road surface appearance, this application embodiment introduces acceleration signals generated by direct contact between the tire and the road surface as supplementary physical information. Furthermore, it combines preset state quantities such as tire pressure, tire temperature, and vehicle speed to dynamically guide and correct visual features. Therefore, this application embodiment can simultaneously perceive the semantic features of the road surface appearance and the dynamic response features during tire contact, improving the ability to distinguish between four different road surface states: flooded roads, damaged roads, asphalt roads, and concrete roads. This application embodiment can output the road surface state category corresponding to the current vehicle's driving area, and this identification result can be provided to the vehicle environmental safety assessment module, driving assistance system, or vehicle control system, providing a basis for driving risk warning, speed adjustment, braking control, and path decision-making, thereby improving the vehicle's safety, stability, and intelligent perception capabilities in complex road environments. Attached Figure Description

[0019] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the accompanying drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Some specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings in an exemplary and non-limiting manner. The same reference numerals in the drawings indicate the same or similar parts or components. Those skilled in the art should understand that these drawings are not necessarily drawn to scale. In the drawings:

[0020] Figure 1 This is a schematic flowchart illustrating an exemplary embodiment of the present application of a road surface condition perception method based on road images and tire mechanics;

[0021] Figure 2 This is a schematic diagram illustrating the workflow of road surface condition perception in an exemplary embodiment of this application;

[0022] Figure 3 This is a schematic diagram illustrating a multimodal data acquisition method according to an exemplary embodiment of this application;

[0023] Figure 4 This is a block diagram illustrating an electronic device according to an exemplary embodiment of this application;

[0024] Figure 5 This is a block diagram illustrating a road surface condition sensing device based on road images and tire mechanics, as shown in an exemplary embodiment of this application. Detailed Implementation

[0025] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0026] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0027] This application proposes a road surface state perception scheme based on road image information and tire dynamic signals. The scheme takes road images, tire acceleration signal sequences, and preset state information such as tire pressure, tire temperature, vehicle speed, and ambient temperature as inputs. Image semantic features, tire dynamic features, and state features are obtained through a visual feature extraction network, a multi-scale one-dimensional temporal feature extraction network, and a state feature mapping network, respectively. Based on this, a cross-modal channel attention mechanism guided by sensor physical priors is constructed, transforming tire dynamic features and state features into modulation weights for modulating the visual feature channel response, thereby achieving dynamic guidance and correction of high-dimensional visual features by low-dimensional physical signals. Furthermore, this application utilizes the original high-frequency dynamic signals of the tire for weight adjustment to avoid the loss of local vibration modes and transient impact information caused by using traditional statistical features. Furthermore, this application alleviates the heterogeneous feature conflict, modal competition, and vision dominance problems existing in traditional multimodal stitching through a physical prior-guided soft fusion mechanism. This can improve the accuracy and robustness of road condition recognition in four easily confused scenarios: damaged road surface, waterlogged road surface, asphalt road surface, and cement road surface. It helps to reduce the model's dependence on a single vision and has good applicability for vehicle-side deployment and engineering application value.

[0028] The embodiments described in this specification will now be described in detail.

[0029] This application provides a road surface condition perception method based on road images and tire mechanics. This method can be deployed in an onboard computing unit, a cloud server, or a roadside unit. Real-time road surface condition recognition can be performed independently by the vehicle, or the perception task can be completed through vehicle-road-cloud collaboration. For example, the vehicle can collect and preprocess multimodal signals, while the cloud or roadside unit performs model inference and road surface condition updates; alternatively, the vehicle can independently perform all feature extraction and classification output, and upload the results to the cloud to construct a road condition map. The aforementioned vehicle-side deployment, cloud deployment, roadside deployment, and vehicle-road-cloud collaborative deployment methods are all variations of the technical solution in this application.

[0030] Figure 1 This is a schematic flowchart illustrating an exemplary embodiment of a road surface condition perception method based on road images and tire mechanics, as shown in this application. Figure 1 As shown, the road surface condition sensing method includes the following steps:

[0031] Step S1: Acquire multimodal data, which includes road surface images, tire acceleration signal sequences, and preset state information.

[0032] The tire acceleration signal sequence includes at least one of a tire radial acceleration signal sequence, a tire tangential acceleration signal sequence, or a tire lateral acceleration signal sequence; the tire radial acceleration signal sequence reflects the acceleration change of the tire after being excited by the road surface in the direction perpendicular to the tread; the tire tangential acceleration signal sequence reflects the acceleration change of the tire along the rolling direction; and the tire lateral acceleration signal sequence reflects the acceleration change of the tire in the lateral direction.

[0033] The preset state information includes vehicle state information and environmental state information. The vehicle state information includes one or more of the following: tire pressure, tire temperature, vehicle speed, tire load, tire wear, wheel speed, slip ratio, steering wheel angle, braking pressure, driving torque, vehicle longitudinal acceleration, and vehicle lateral acceleration. The environmental state information includes ambient temperature, rain sensor signal, etc.

[0034] This embodiment can acquire road surface images along the vehicle's driving direction using an in-vehicle camera; acquire tire acceleration signal sequences using sensors installed on the tires; synchronously acquire preset state information using a state acquisition unit; and assign timestamps to the road surface images, tire acceleration signal sequences, and preset state information using a time synchronization unit, and extract the tire acceleration signal sequences and preset state information within a preset time window based on the timestamps of the road surface images to obtain the multimodal data.

[0035] Step S2: Feature extraction is performed on the multimodal data to obtain multimodal features, which include image semantic features, tire dynamics features, and state features. Specifically, the tire dynamics features are obtained by performing interpolation processing on the tire acceleration signal sequence to obtain an acceleration signal change sequence, and by performing multi-scale one-dimensional time-series feature extraction on the multi-channel time-series input composed of the tire acceleration signal sequence and the acceleration signal change sequence.

[0036] This embodiment preprocesses each tire acceleration signal sequence; calculates the acceleration difference between adjacent sampling points in the preprocessed acceleration signal sequence to obtain the acceleration signal change sequence. Specifically, each tire acceleration signal sequence is low-pass filtered to obtain a filtered acceleration signal sequence; the filtered acceleration signal sequence is then subjected to zero-bias removal and zero-mean normalization to obtain the preprocessed acceleration signal sequence.

[0037] Optionally, the cutoff frequency of the low-pass filter is set to 200Hz to 400Hz to filter out high-frequency noise and retain the effective vibration components generated by the contact between the tire and the road surface.

[0038] Step S3: The tire dynamics features and the state features are concatenated along the channel dimension to obtain physical prior joint features. Modulation weights are generated based on the physical prior joint features, and the image semantic features are modulated using the modulation weights to obtain fused features.

[0039] Step S4: Input the fused features into the classifier and output the road surface state classification result.

[0040] Since the original high-frequency acceleration signal sequence can preserve the temporal order of the sampling points, and the acceleration signal change sequence can characterize the continuous rolling vibration and local impact response of the tire, this embodiment constructs tire dynamic characteristics based on the original high-frequency acceleration signal sequence and the acceleration signal change sequence, rather than based on the statistical characteristics such as the mean, variance, peak value, and energy of the acceleration signal. Therefore, this embodiment can reduce the loss of temporal details caused by statistical feature compression and improve the ability to identify easily confused road surface conditions such as damage and water accumulation.

[0041] Furthermore, this embodiment constructs a sensor-physical prior-guided cross-modal channel attention fusion mechanism. This mechanism does not simply stitch together image features, tire dynamics features, and vehicle state features and then classify them directly. Instead, it first transforms tire dynamics features and vehicle state features into attention weights for modulating the visual feature channel response. This enables the dynamic guidance and correction of high-dimensional visual features by low-dimensional physical signals, alleviating the problems of heterogeneous feature conflict, modal competition, visual dominance, and insufficient role of low-dimensional physical signals in traditional multimodal stitching. As a result, it can improve the robustness of recognition in complex lighting, water reflection, shadow interference, or road surface appearance similar scenarios.

[0042] like Figure 2 As shown, the road condition perception process includes signal acquisition, feature extraction, feature modulation, and road classification. Next, we will combine... Figure 2 This section details the process by which the road surface perception system perceives the road surface condition.

[0043] First, perform step S1 to collect data.

[0044] This step aims to acquire real-time road surface images, high-frequency tire acceleration signals, and preset state information to provide high-quality raw data to support subsequent processing and analysis.

[0045] like Figure 3 As shown, in practical applications, the vehicle-side multi-sensor system synchronously acquires image data, tire acceleration signals, and vehicle status signals in real time. Road surface images reflect the road's texture, color, and macroscopic shape; tire acceleration signals reflect the vibration excitation generated by the tire's contact with the road surface; and vehicle and environmental status information supplements the low-frequency physical parameters of the vehicle's operating status and tire's working status.

[0046] The road perception system in this embodiment can obtain the aforementioned multiple modal data by interacting with a multi-sensor system. That is, this embodiment can simultaneously obtain the macroscopic appearance features of the road surface and the underlying physical response generated by tire-road contact, providing a data foundation for improving the robustness of road condition recognition in complex traffic scenarios.

[0047] Road surface image acquisition primarily relies on vehicle-mounted cameras. These cameras are installed at a stable position at the front of the vehicle to continuously acquire images of the road surface along the vehicle's direction of travel. The raw images acquired by the cameras can have a resolution of 1920×1080, with a frame rate of 25 frames per second. Road surface images can reflect the texture, color, cracks, water accumulation and reflections, damaged areas, and the appearance differences between different paving materials, serving as crucial visual evidence for road surface condition classification. Through this visual acquisition method, the road surface perception system can obtain macroscopic spatial semantic information about the road surface, providing input for the subsequent extraction of high-dimensional image features by the visual backbone network.

[0048] It is worth noting that the road surface image acquisition device is not limited to a single forward-looking camera; it can also use a monocular camera, a binocular camera, a surround-view camera, a fisheye camera, an infrared camera, a thermal imaging camera, an event camera, or a combination of the above-mentioned visual sensors. All of these sensors can be used to acquire visual information such as road surface texture, color, reflection, damaged edges, and water accumulation areas.

[0049] The acquisition of tire acceleration signals is primarily accomplished using a Tire Pressure Monitoring System (TPMS) installed on the inner sidewall of the tire. During vehicle operation, the tire, as the only component directly in contact with the road surface, directly reflects changes in road surface micro-roughness, pothole impacts, damage edges, and contact conditions through its vibration response. When a vehicle traverses smooth asphalt, cement, puddles, or damaged surfaces, the tire acceleration signal sequence exhibits varying amplitude fluctuations, frequency components, and transient impact characteristics. Therefore, this embodiment does not use traditional statistical features such as mean, variance, and peak value as input. Instead, it directly acquires the original one-dimensional acceleration sequence and uses it as the input signal for the subsequent multi-scale one-dimensional temporal feature extraction network, with a sampling frequency of 2000Hz. This method maximizes the preservation of high-frequency dynamic information generated during tire-road contact, avoiding the loss of detail caused by statistical feature compression.

[0050] In other embodiments, the tire acceleration signal acquisition device is not limited to a TPMS installed on the inner sidewall of the tire. Any device that can acquire acceleration signals related to the tire-road contact state can be used as an alternative acquisition device in this embodiment.

[0051] Vehicle state signals are also acquired using a TPMS (Tire Pressure Monitoring System) mounted on the inner sidewall of the tire, with a sampling frequency of 10Hz. In addition to high-frequency tire acceleration signals, the TPMS simultaneously acquires low-frequency scalar signals such as tire pressure, tire temperature, and vehicle speed to characterize the vehicle's operating state and tire performance. Specifically, tire pressure reflects changes in internal tire pressure, tire temperature reflects temperature changes during tire operation, and vehicle speed reflects the vehicle's current speed. While these scalar signals do not contain the high-frequency vibration details of the tire acceleration signals, they provide necessary state references for tire dynamics response. For example, during vehicle operation, changes in tire pressure affect tire stiffness and contact patch, changes in tire temperature affect tire material properties and vibration transmission, and changes in vehicle speed affect the tire excitation frequency and acceleration response amplitude. Therefore, using the tire pressure, tire temperature, and vehicle speed simultaneously acquired by the TPMS as input signals to the vehicle state scalar signal model helps the model interpret tire acceleration signals in conjunction with the current tire performance, providing more complete state information support for subsequent road condition identification.

[0052] To ensure a unified time reference among multimodal data, the signal acquisition module is equipped with a time synchronization unit. This time synchronization unit is used to perform unified clock management or time reference calibration for the vehicle-mounted camera, tire sensor, and vehicle status acquisition unit during the hardware acquisition phase, so that road images, tire acceleration signal sequences, and vehicle status scalars such as tire pressure, tire temperature, and vehicle speed are all assigned corresponding timestamps during acquisition.

[0053] Specifically, the vehicle-mounted camera generates or records an image frame timestamp when acquiring each frame of road surface image; the tire sensor generates or records an acceleration sampling timestamp when acquiring tire acceleration sampling points; and the vehicle status acquisition unit generates or records status quantity timestamps when acquiring tire pressure, tire temperature, and vehicle speed data. These timestamps use a unified time reference to identify the actual acquisition time of different modal data. By completing unified clock management and timestamp calibration during the acquisition phase, road images, tire acceleration signal sequences, and preset status information obtained at different sampling frequencies can have a basis for subsequent pairing and alignment, avoiding data mismatch caused by inconsistent sensor time references. This provides a reliable data foundation for subsequent multimodal sample construction, time window extraction, and fusion recognition.

[0054] Through the above signal acquisition scheme, this embodiment can simultaneously obtain visual appearance information of the road surface, high-frequency physical response generated by tire-road contact, and vehicle operating status information. Compared with a single visual acquisition scheme, this embodiment can provide additional physical discrimination criteria under conditions of changing lighting, shadow occlusion, water reflection, or similar road surface textures, laying the foundation for achieving highly robust, low-visual-dependence, and data-efficient road condition perception in complex traffic scenarios.

[0055] After data acquisition is completed, the acquired data undergoes signal preprocessing.

[0056] Signal preprocessing is a crucial step in road perception systems, especially when dealing with vehicle-mounted image signals, tire acceleration signals, and preset state signals. The quality of preprocessing directly impacts the accuracy and reliability of subsequent feature extraction, cross-modal fusion, and road state classification. The main purpose of signal preprocessing is to remove irrelevant noise and abnormal fluctuations introduced during acquisition by performing operations such as data standardization, noise suppression, time alignment, windowing, and normalization on multi-source signals. This enhances effective features related to road conditions, thereby providing clean, synchronous, and standardized multimodal input data for subsequent road recognition.

[0057] In this embodiment, signal preprocessing mainly includes four operations: road surface image preprocessing, tire acceleration signal sequence preprocessing, preset state signal preprocessing, and multimodal data time alignment. Through the above preprocessing operations, the road perception system can reduce the impact of sensor noise, vehicle vibration interference, image illumination changes, and different sampling frequencies, enabling data from different modalities to participate in subsequent analysis at a unified time scale and a unified feature scale.

[0058] For the preprocessing of road surface images:

[0059] Road surface image preprocessing is mainly used to improve the consistency and stability of image input and reduce the interference of changes in imaging conditions on the visual feature extraction results. Since road surface images captured by cameras are easily affected by factors such as light intensity, shadow occlusion, camera shake, road surface reflection, rain and fog, and changes in shooting angle when vehicles are driving in real road environments, it is necessary to normalize the road surface images before inputting them into the visual feature extraction network.

[0060] To improve the quality of the input image, this embodiment performs size unification processing on the acquired original road surface images, adjusting images of different resolutions to a fixed input size required by the visual feature extraction network. For example, the cropped road surface images are uniformly scaled to 640(H)×640(W) while retaining the RGB three color channels, thus obtaining an input image with a size of 640×640×3. Furthermore, the RGB pixel values ​​are normalized from 0-255 to [0,1], or standardized using the channel mean and standard deviation of the training set images, to obtain the input image for the visual feature extraction network.

[0061] This embodiment can reduce the pixel amplitude differences caused by different shooting devices, different exposure conditions or different lighting environments by preprocessing the road surface image.

[0062] For the preprocessing of tire acceleration signal sequences:

[0063] Preprocessing of the tire acceleration signal sequence is mainly used to suppress high-frequency noise interference introduced during the acquisition process and to retain as much of the effective vibration response generated during tire-road contact as possible. During vehicle operation, sensors installed on or near the tire continuously acquire tire acceleration signals. These signals not only contain effective dynamic information caused by road roughness, pothole impacts, damage edges, and changes in tire contact with the road surface, but may also be affected by sensor noise, mechanical vibration, and environmental disturbances. Therefore, before inputting the signal into a multi-scale one-dimensional temporal feature extraction network, the original acceleration signal sequence needs to be low-pass filtered to improve signal quality. Low-pass filtering is applied to the acquired tire acceleration signal sequence to remove high-frequency noise components above a set cutoff frequency, retaining the main vibration information related to the tire-road contact state. Let the original tire acceleration signal sequence... for:

[0064] (1)

[0065] in, Indicates the number of sampling points in the sequence. This represents the acceleration value corresponding to the Lth sampling point.

[0066] After low-pass filtering, the filtered acceleration signal sequence is obtained. :

[0067] (2)

[0068] in, This indicates a low-pass filter operation.

[0069] In some implementations, the low-pass filter can be set with a cutoff frequency based on the sensor sampling frequency, vehicle speed, and the effective frequency range of the tire vibration signal.

[0070] The low-pass filtered tire acceleration signal sequence was subjected to zero-bias removal and zero-mean normalization to reduce the impact of sensor installation deviation, low-frequency drift, and signal amplitude differences under different operating conditions on subsequent model inputs. (Preprocessed tire acceleration signal sequence) It can be represented as:

[0071] (3)

[0072] in, This indicates normalization processing.

[0073] For example, if the sampling frequency of the tire acceleration sensor is 2000Hz, the time window length is 0.5s, and the length of the acceleration signal sequence input to the multi-scale one-dimensional convolutional neural network is 1000, the road perception system performs the above preprocessing on the 1000-point raw acceleration signal sequence. The low-pass filter cutoff frequency can be set to 200Hz to 400Hz to filter out sensor noise and high-frequency interference, while retaining the effective tire vibration components caused by road surface roughness, pothole impact, water contact, and damaged edges.

[0074] It is worth noting that in this embodiment, the tire acceleration signal sequence is not compressed into statistical features such as mean, variance, peak value, and energy during the preprocessing stage. Instead, it is only subjected to signal quality improvement and scale unification processing. The preprocessed signal sequence still retains the time sequence, amplitude fluctuation, and local mutation information between the original sampling points, so as to facilitate the construction of a multi-channel time series matrix and the extraction of tire continuous rolling vibration features, local impact features, and multi-scale contact response features.

[0075] Regarding the preprocessing of preset state information:

[0076] The preprocessing of the preset state signals is mainly used to eliminate the dimensional differences between different physical quantities, enabling low-frequency state signals such as tire pressure, tire temperature, and vehicle speed to be used collaboratively with image features and high-frequency tire time-series features in the network. Since the numerical range, physical meaning, and frequency of change of tire pressure, tire temperature, and vehicle speed are all different, directly inputting them into the state feature mapping network may result in variables with larger numerical ranges having excessively high weights, thus affecting the accuracy of feature mapping. Therefore, this embodiment performs normalization processing on these state vectors. For each type of state quantity... The min-max normalization method can be used:

[0077] (4)

[0078] in, This represents the preprocessed state variable. , This represents the maximum and minimum values ​​in the state variables. This indicates the preset minimum value to avoid the denominator being zero.

[0079] After preprocessing the modal data through the above embodiments, time alignment processing is performed on the multimodal data.

[0080] Multimodal time alignment is a crucial step in ensuring the correct fusion of image signals, tire acceleration signal sequences, and preset state information. Since the sampling frequencies of the camera, tire sensors, and vehicle state acquisition unit are typically different, if there is a time shift between different modes, the image may show a certain road surface area, while the tire signal corresponds to a different driving position, leading the model to learn incorrect modal correspondences.

[0081] In this embodiment, the image frame timestamp is used as a reference. The tire acceleration signal sequence within a certain time range before and after the image frame is selected as the physical signal window corresponding to the road surface image. The tire pressure, tire temperature and vehicle speed data within the same time range are read as preset status information.

[0082] In this way, the visual information and tire physical response in the same sample can correspond as closely as possible to the same vehicle driving segment. The data structure of a multimodal sample is shown in Table 1 below. Each inference sample includes a road surface image, a tire acceleration signal sequence, and a three-dimensional state vector, ensuring consistency in the time scale and data structure of the multimodal data.

[0083] Table 1: Multimodal Sample Data Structure Table

[0084]

[0085] Then, step S2 is performed to extract multimodal features.

[0086] Multimodal feature extraction extracts discriminative feature representations from road surface images, tire acceleration signal sequences, and preset state information, providing a foundation for subsequent sensor-guided cross-modal attention fusion.

[0087] Because road surface images, tire acceleration signal sequences, and preset state information differ significantly in data format, sampling frequency, physical meaning, and feature dimensions, direct splicing or unified processing can easily lead to information interference between heterogeneous features, resulting in modal competition. Therefore, this embodiment employs three independent front-end feature extraction links to extract features from the high-dimensional visual image, the one-dimensional tire acceleration signal sequence, and the preset state scalar, respectively, to obtain image semantic features, tire dynamic features, and state features.

[0088] Image semantic features are used to describe the texture, color, cracks, water reflection, damage morphology, and differences in paving materials of the road surface; tire dynamic features are used to reflect the high-frequency vibration, local impact, and micro-roughness information generated during the direct contact between the tire and the road surface; state features are used to supplement the influence of tire pressure, tire temperature, and vehicle speed on tire dynamic response.

[0089] This embodiment ensures the complete preservation of visual semantic information and underlying physical response information by independently extracting the three types of features mentioned above. Specifically, tire dynamics features and state features are used to generate visual channel modulation weights, while image semantic features participate in the final fusion classification as the modulated object. Therefore, the generation of modulation weights is primarily based on the physical response generated by tire-road contact, vehicle operating state, and environmental state, rather than being determined by the image features themselves.

[0090] Regarding image semantic features:

[0091] The main purpose of extracting semantic features from images is to obtain high-dimensional spatial features that characterize the road surface condition from road surface images captured by vehicles. During actual driving, cameras can capture information such as the road surface's texture, damaged edges, water accumulation areas, and color changes. This visual information is crucial for determining the road surface category. However, under complex lighting conditions, localized shadows, road surface reflections, or similar textures, relying solely on image information can easily lead to feature confusion.

[0092] Therefore, in this embodiment, the visual branch does not directly complete the final classification, but instead serves as a high-dimensional spatial semantic feature extractor, providing a basic visual representation for subsequent physical prior modulation. For example, the visual backbone network uses a pre-trained residual network, ResNet-18. To preserve the spatial semantic expressiveness of the image and avoid the visual branch prematurely outputting independent classification results, the original fully connected classification layer of ResNet-18 is removed from the visual feature extraction network, retaining only its convolutional layers, residual mapping structure, and global average pooling layer. After the image undergoes multiple layers of two-dimensional convolution and residual feature transformation, the global average pooling layer outputs fixed-dimensional image semantic features. These image semantic features can comprehensively describe the macroscopic texture distribution, local damage morphology, road material differences, and water reflection in the current road surface image.

[0093] For example, the input to the visual feature extraction network is a 640×640×3 RGB road surface image. After processing by convolutional layers, residual blocks, and global average pooling layers of ResNet-18, the output is 512-dimensional image semantic features. Because the original fully connected final classification layer of ResNet-18 has been removed, the visual feature extraction network does not directly output the road surface category, but instead outputs high-dimensional visual semantic features for subsequent cross-modal attention modulation.

[0094] It is worth noting that this embodiment does not limit the structure of the visual feature extraction network. Those skilled in the art can flexibly construct the visual feature extraction network, for example, based on MobileNet, EfficientNet, Transformer or other visual networks to extract semantic features of images.

[0095] Regarding tire dynamics:

[0096] In order to further highlight the short-term abrupt response of the tire when it passes over damaged edges, potholes, cracks or water accumulation areas, this embodiment also constructs an acceleration signal change sequence based on the preprocessed tire acceleration signal sequence. As shown in the following formula (5), this acceleration signal change sequence is obtained by the difference between adjacent sampling points:

[0097] (5)

[0098] in, , This represents a sequence of acceleration signal changes that retains its temporal order and is not compressed into statistics. Therefore, this sequence of acceleration signal changes can enhance information on local shocks, short-term peak abrupt changes, and non-stationary vibration changes.

[0099] Based on this, this embodiment uses the preprocessed tire acceleration signal sequence and its corresponding acceleration signal change sequence as timing input.

[0100] As mentioned above, the tire acceleration signal sequence includes at least one of the following: tire radial acceleration signal sequence, tire tangential acceleration signal sequence, or tire lateral acceleration signal sequence. The tire radial acceleration signal sequence can characterize the vibration response of the tire after being excited by the road surface in the direction perpendicular to the tread. The tire tangential acceleration signal sequence mainly reflects the acceleration change of the tire along the rolling direction and can characterize the tangential dynamic response caused by factors such as driving, braking, changes in road surface adhesion, or water accumulation during the tire-road contact process. The tire lateral acceleration signal sequence mainly reflects the acceleration change of the tire in the lateral direction and can be used to characterize the lateral dynamic response caused by vehicle steering, lateral disturbance, or uneven road surface contact.

[0101] Therefore, when using tire radial acceleration signal sequences, tire tangential acceleration signal sequences, or tire lateral acceleration signal sequences simultaneously, the three acceleration signals can be synchronously captured within the same time window and subjected to zero-bias removal, filtering, and normalization processing respectively to obtain the preprocessed radial acceleration signal sequence. Tangential acceleration signal sequence and lateral acceleration signal sequence Construct corresponding acceleration signal change sequences respectively. , and This is used to characterize local changes and short-time impulse responses between adjacent sampling points. Subsequently, the above sequences are combined according to the channel dimension to form a multi-channel time series matrix. :

[0102] (6)

[0103] This multi-channel time-series matrix simultaneously contains continuous acceleration responses and their variations in the radial, tangential, and lateral directions, enabling it to characterize vertical vibration, rolling disturbances, lateral disturbances, and local impact features during tire-road contact. A multi-scale one-dimensional time-series feature extraction network can extract features from this input matrix along the time dimension, thereby obtaining more complete tire dynamics characteristics.

[0104] In the feature extraction process, this embodiment employs a multi-scale one-dimensional temporal feature extraction network to process the multi-channel temporal matrix. This multi-scale one-dimensional temporal feature extraction network includes a long-term vibration feature extraction branch and a short-term impact feature extraction branch. The long-term vibration feature extraction branch uses a larger convolutional kernel to obtain a wider temporal receptive field and extract features related to road surface roughness, paving materials, and continuous vibration modes during the continuous rolling of the tire; the short-term impact feature extraction branch uses a smaller convolutional kernel to extract local abrupt features caused by damaged edges, pothole impacts, and water contact.

[0105] Specifically, the multi-scale one-dimensional temporal feature extraction network includes:

[0106] The first one-dimensional convolution branch is used to extract long-term vibration features. ;

[0107] The second one-dimensional convolutional branch is used to extract short-term impact features. The kernel size of the first one-dimensional convolution branch is larger than the kernel size of the second one-dimensional convolution branch;

[0108] The fusion layer is used to stitch together the long-term vibration features and the short-term impact features to obtain joint temporal features;

[0109] A mapping layer is used to sequentially perform global average pooling and fully connected mapping on the joint temporal features to obtain the tire dynamics features.

[0110] For example, , , This indicates a one-dimensional convolution operation with a large kernel size. This represents a one-dimensional convolution operation with a small kernel size. In practical applications, It can be set to 7. It can be set to 3. Optionally, after each one-dimensional convolution operation, a batch normalization layer, a ReLU activation function, and a pooling layer can be sequentially connected to improve training stability and reduce redundant information. Subsequently, the long-term vibration features and short-term impact features are concatenated to obtain joint temporal features:

[0111] (7)

[0112] in, This indicates a feature splicing operation.

[0113] The concatenated joint temporal features are further mapped through a global average pooling layer and a fully connected layer to obtain fixed-dimensional tire dynamics features. :

[0114] (8)

[0115] in, This indicates a global average pooling operation. This indicates a fully connected mapping operation.

[0116] Thus, this embodiment, through a multi-scale one-dimensional temporal feature extraction network, can learn the tire rolling vibration mode at a long time scale and the local impact features at a short time scale while preserving the temporal structure of the original acceleration signal, thereby improving the ability to distinguish between asphalt pavement, cement pavement, waterlogged pavement and damaged pavement.

[0117] It is worth noting that in other embodiments, a multi-scale one-dimensional temporal feature extraction network can be constructed based on LSTM, GRU, TCN, or temporal Transformer.

[0118] Regarding the extraction of state features:

[0119] The main purpose of state feature extraction is to map physical quantities such as tire pressure, tire temperature, and vehicle speed into feature representations suitable for co-modeling with other modes. Although tire pressure, tire temperature, and vehicle speed have low dimensionality, they directly affect the tire's contact patch state, material stiffness, vibration transmission characteristics, and acceleration response amplitude. Therefore, when interpreting tire acceleration signals, relying solely on the time-series waveform is insufficient; a comprehensive judgment based on the vehicle's current operating state and the current environmental conditions is also necessary. For example, tire pressure, tire temperature, and vehicle speed can be concatenated into a three-dimensional state vector. :

[0120] (9)

[0121] in, Indicates tire pressure. Indicates fetal temperature. Indicates vehicle speed.

[0122] Considering the significant differences in dimensions and numerical range between the aforementioned state variables and acceleration signals, this embodiment employs a multilayer perceptron for nonlinear mapping, converting low-dimensional state information into high-dimensional latent semantic features. :

[0123] (10)

[0124] in, This refers to a multilayer perceptron, a type of network used to extract state features.

[0125] For example, if a preset state vector is used... Tire pressure Fetal temperature and vehicle speed The state feature mapping network consists of a two-layer fully connected structure. The first fully connected layer maps the 3D state vector to 32-dimensional features, and the second fully connected layer maps it to 32-dimensional high-dimensional latent semantic features. This high-dimensional latent semantic feature It can compensate for the differences in the amplitude and frequency distribution of tire radial acceleration response under different tire pressure, tire temperature and vehicle speed conditions.

[0126] It is worth noting that in this embodiment, a multilayer perceptron is used to map state information such as tire pressure, tire temperature, and vehicle speed into state features. In other embodiments, linear mapping, nonlinear fully connected networks, embedding layers, gating networks, normalized coding modules, conditional coding modules, or attention coding modules can also be used to perform feature mapping on preset state information.

[0127] Next, step S3 is performed to perform feature modulation.

[0128] This step aims to convert image semantic features, tire dynamics features, and state features into fused features for road surface state classification.

[0129] Unlike traditional multimodal models that directly concatenate features from different modalities before inputting them into the classifier, this embodiment does not directly utilize the concatenated features for final classification, nor does it use image features for modulation weight generation. Instead, it first constructs a physical prior joint feature from tire dynamics features and state features, and then generates modulation weights corresponding to the visual feature channels from this physical prior joint feature. Subsequently, the image semantic features are dynamically modulated channel by channel using these modulation weights, and finally, the modulated fused features are input into the classifier. In this way, tire dynamics features and state features are no longer merely additional features participating in classification, but are transformed into guiding information for the visual feature channels, thereby enabling physical sensing information to directly influence the expression of visual features.

[0130] Specifically, this step includes the following three sub-steps:

[0131] Step S3.1 Construction of physical prior joint features.

[0132] After completing the extraction of the three types of features, image semantic features It does not directly participate in the generation of modulation weights, but rather serves as the object to be modulated subsequently. The road perception system first incorporates tire dynamic characteristics. With state characteristics By concatenating the features along the channel dimension, we obtain the physical prior joint features. :

[0133] (11)

[0134] The physical prior joint features are used to generate modulation weights, which are mainly determined by the tire-road contact response and preset state information such as tire pressure, tire temperature, and vehicle speed.

[0135] For example, , , Therefore, the spliced ​​physical prior joint features The 160-dimensional physical prior joint features include tire contact vibration information and preset state information. Image semantic features. It is 512-dimensional and does not participate. Instead of being constructed, it accepts dynamic modulation of modulation weights in subsequent steps.

[0136] Step S3.2 Modulation weight generation.

[0137] After obtaining the physical prior joint features Then, it is input into an attention generator to calculate modulation weights consistent with the number of channels in the image's semantic features. For example, this attention generator consists of two fully connected layers, first performing dimensionality reduction and compression on the physical prior joint features, and then mapping them back to the same dimension as the image's semantic features. Its calculation process can be represented as follows:

[0138] (12)

[0139] in, This represents the generated modulation weights; , These represent the weight matrices of the two fully connected layers, respectively. and These represent the bias terms of the two fully connected layers, respectively. Represents the ReLU activation function; This represents the Sigmoid activation function.

[0140] For example, the attention generator includes a first fully connected layer, a ReLU activation layer, a second fully connected layer, and a Sigmoid activation layer. The first fully connected layer combines 160-dimensional physical prior features. Mapping to 128-dimensional latent features, the second fully connected layer maps the 128-dimensional latent features to 512-dimensional modulation weights. The modulation weight Image semantic features The number of channels is consistent, which is used for subsequent channel-by-channel modulation of image semantic features.

[0141] When short-term impacts, sudden peak changes, or enhanced high-frequency vibrations are present in tire dynamics features, the attention generator increases the weights of channels related to visual features such as cracks, potholes, and damage edges. When tire dynamics features exhibit vibration changes related to water contact, and these are combined with current vehicle speed, tire pressure, and tire temperature, the attention generator generates weights to enhance visual channels related to wetness, water accumulation, or low adhesion. These weights are then applied to image semantic features, enhancing relevant visual responses already extracted from the image and reducing misjudgments caused by relying solely on road surface images.

[0142] It is worth noting that in other embodiments, gating fusion, cross-attention, feature weighting, dynamic weighting, conditional normalization, confidence fusion, or post-fusion methods can also be used to process the physical prior joint features to generate modulation weights.

[0143] Step S3.3 Dynamic modulation of image semantic features.

[0144] After obtaining the modulation weight Then, the modulation weights are used to analyze the semantic features of the image. By performing channel-by-channel multiplication, the modulated fusion features are obtained. :

[0145] (13)

[0146] in, This represents the element-wise multiplication operator.

[0147] Since image semantic features and modulation weights have the same channel dimension, the modulation weights can independently modulate each channel of the image semantic features. When certain visual channels are more relevant to the current tire dynamics response and preset state, their corresponding weights can be relatively increased; when certain visual channels are less relevant to the current tire dynamics response and preset state, their corresponding weights can be relatively decreased. Thus, the fused features can incorporate physical guidance information provided by the tire sensors while preserving the visual spatial semantic information.

[0148] For example, image semantic features It has 512 dimensions and modulation weights. The feature is 512-dimensional; multiplying the two elements element-wise yields a 512-dimensional fused feature. If the modulation weight If a weight in one dimension is close to 1, it indicates that the visual channel is highly correlated with the current tire dynamics response and the preset state, and the features of that channel are preserved or enhanced; if the weight is modulated... If the weight of a certain dimension is close to 0, it means that the visual channel may correspond to interference information such as lighting, shadows, and ordinary texture changes, and the visual features of the channel are suppressed.

[0149] For example, when a vehicle travels over a damaged road surface, short-term peaks and local abrupt changes appear in the tire acceleration signal sequence. After the multi-scale one-dimensional temporal feature extraction network extracts this impact feature, the attention generator increases the weights of channels related to crack edges, pothole contours, and damaged textures in the image semantic features, thus making the classifier more inclined to output the damaged road surface category. Similarly, when a vehicle travels over a flooded road surface, there may be highly reflective areas in the road image, which can easily be confused with light-colored cement roads based solely on the road image. In this case, tire dynamics features and state features such as vehicle speed, tire pressure, and tire temperature jointly participate in generating modulation weights, enhancing the visual channels related to water reflection and slippery road surfaces, thereby reducing the probability of misclassification.

[0150] Finally, step S4 is performed to classify the road surface.

[0151] In this embodiment, after feature modulation is completed, the features will be fused. Input classifier. For example, this classifier consists of a multilayer perceptron and a softmax function, used to output the predicted probability distribution for different road surface categories:

[0152] (14)

[0153] in, This indicates the predicted classification of road surface categories. This represents the fully connected mapping structure in the classifier.

[0154] If the number of road surface categories is Then the output probability can be expressed as:

[0155] (15)

[0156] in, This indicates that the input sample belongs to the first... The probability of a road surface condition.

[0157] Thus, the road perception system ultimately determines the road surface category corresponding to the current input sample based on the probability distribution. :

[0158] (16)

[0159] In some embodiments, the road perception system ultimately outputs a road condition classification result and its probability distribution corresponding to the current vehicle driving area, rather than an independent judgment result of any single modality. This classification result comprehensively reflects four road surface types—flooded road surface, damaged road surface, asphalt road surface, and concrete road surface—and their impact on tire dynamics response. It can be further used for environmental safety assessment, low-adhesion risk identification, and driver assistance decision-making in vehicle driving scenarios. For example, the classifier outputs the probability distribution of the four road surface conditions: asphalt road surface, concrete road surface, flooded road surface, and damaged road surface.

[0160] For example, for a given input sample, if the road perception system obtains the following classification probabilities: asphalt road 0.06, cement road 0.08, waterlogged road 0.81, and damaged road 0.05, then the road perception system classifies the sample as a waterlogged road. Similarly, for another input sample, if the road perception system obtains the following classification probabilities: asphalt road 0.10, cement road 0.07, waterlogged road 0.09, and damaged road 0.74, then the road perception system classifies the sample as a damaged road. When the predicted probability of a waterlogged or damaged road exceeds a preset risk threshold, such as 0.70, the road perception system can output a road risk warning signal to the vehicle's driver assistance system or vehicle control system to assist in speed adjustment, braking control, or route decision-making.

[0161] In a scenario where a vehicle is traveling at 40 km / h on a flooded road after rain, an onboard camera captures images of the road ahead at 25 frames per second. These images are then cropped and scaled to form an RGB road image. A tire acceleration sensor collects tire acceleration signals at 2000 Hz. For an image frame with timestamp t, the road perception system extracts acceleration signals from t-0.25 s to t+0.25 s, forming a 1000-dimensional tire acceleration signal sequence. Simultaneously, the road perception system reads tire pressure, tire temperature, and vehicle speed within this 0.5 s time window and averages these values ​​to form a state vector. This sample is input into the road state perception model. The visual branch outputs 512-dimensional image semantic features, the tire radial acceleration branch outputs 128-dimensional temporal features, and the vehicle state branch outputs 32-dimensional state features. The acceleration branch outputs 128-dimensional tire dynamic features, and the state scalar branch outputs 32-dimensional state features. These two are concatenated to form a 160-dimensional physical prior joint feature, which is then used by an attention generator to generate 512-dimensional modulation weights. The 512-dimensional image semantic features output by the visual branch do not participate in the generation of modulation weights. Instead, these modulation weights are modulated channel by channel to obtain fused features, which are then input into the classifier. The classifier outputs the probability distributions for asphalt pavement, cement pavement, waterlogged pavement, and damaged pavement. If the output results are 0.06 for asphalt pavement, 0.08 for cement pavement, 0.81 for waterlogged pavement, and 0.05 for damaged pavement, then the pavement perception system determines the current pavement condition as waterlogged pavement.

[0162] Figure 4 This is a schematic diagram of an electronic device illustrated in this specification according to an exemplary embodiment. Please refer to... Figure 4 At the hardware level, the device includes a processor 410, an internal bus 420, a network interface 430, memory 440, a hardware acceleration device 450, and non-volatile memory 460, and may also include other hardware required for its functions. One or more embodiments of this application can be implemented in software, for example, the processor 410 reads the corresponding computer program from the non-volatile memory 460 into the memory 440 and then runs it. Of course, in addition to software implementation, one or more embodiments of this application do not exclude other implementation methods, such as logic devices or a combination of hardware and software, etc. That is to say, the execution subject of the above processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0163] Figure 5 This is a structural block diagram illustrating an exemplary embodiment of a road surface condition sensing device based on road images and tire mechanics. The road surface condition sensing device can be applied to, for example... Figure 4The electronic device shown implements the technical solution of this application. The road surface condition sensing device includes: a data acquisition unit 510, a feature extraction unit 520, a feature modulation unit 530, and a road surface recognition unit 540, wherein:

[0164] The data acquisition unit 510 is used to acquire multimodal data, which includes road surface images, tire acceleration signal sequences, and preset state information.

[0165] The feature extraction unit 520 is used to extract features from the multimodal data to obtain multimodal features, which include image semantic features, tire dynamics features, and state features. Specifically, the tire dynamics features are obtained by performing interpolation processing on the tire acceleration signal sequence to obtain an acceleration signal change sequence, and by performing multi-scale one-dimensional time-series feature extraction on the multi-channel time-series input composed of the tire acceleration signal sequence and the acceleration signal change sequence.

[0166] The feature modulation unit 530 is used to concatenate the tire dynamics features and the state features in the channel dimension to obtain physical prior joint features, generate modulation weights based on the physical prior joint features, and use the modulation weights to modulate the image semantic features to obtain fused features.

[0167] The road surface recognition unit 540 is used to input the fused features into the classifier and output the road surface state classification result.

[0168] In some embodiments, the data acquisition unit 510 is used to acquire road surface images in the direction of vehicle travel via an on-board camera; acquire tire acceleration signal sequences via sensors installed on the tires; synchronously acquire preset state information via a state acquisition unit; assign timestamps to the road surface images, tire acceleration signal sequences, and preset state information via a time synchronization unit; and extract the tire acceleration signal sequences and preset state information within a preset time window based on the timestamps of the road surface images to obtain the multimodal data.

[0169] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this application according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0170] Accordingly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the methods described in any of the above embodiments.

[0171] Accordingly, embodiments of this application also provide a computer program product configured to perform the methods described in any of the above embodiments.

[0172] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, which can take the form of a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email sending and receiving device, game console, tablet computer, wearable device, or any combination of these devices.

[0173] In a typical configuration, a computer includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0174] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0175] Computer-readable media, including both permanent and non-permanent, removable and non-removable media, can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random-access memory (SRAM), dynamic random-access memory (DRAM), other types of random-access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage, quantum memory, graphene-based storage media or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0176] While this specification contains numerous specific implementation details, these should not be construed as limiting the scope of any invention or the scope of the claims, but rather are primarily intended to describe features of specific embodiments of a particular invention. Certain features described in the various embodiments herein may also be implemented in combination in a single embodiment. Conversely, various features described in a single embodiment may also be implemented separately in various embodiments or in any suitable sub-combination. Furthermore, while features may function in certain combinations as described above and even initially claimed in this way, one or more features from a claimed combination may be removed from that combination in some cases, and a claimed combination may refer to a sub-combination or a variation thereof.

[0177] Similarly, although the operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring these operations to be performed in the specific order shown or sequentially, or requiring all illustrated operations to be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0178] Thus, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims may be performed in a different order and still achieve the desired result. Furthermore, the processes depicted in the drawings are not necessarily shown in a specific order or sequence to achieve the desired result. In some implementations, multitasking and parallel processing may be advantageous.

[0179] It should be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0180] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A method for perceiving road surface conditions based on road images and tire mechanics, characterized in that, The method includes the following steps: Step S1: Acquire multimodal data, which includes road surface images, tire acceleration signal sequences, and preset state information; Step S2: Feature extraction is performed on the multimodal data to obtain multimodal features, which include image semantic features, tire dynamics features, and state features; wherein, by performing interpolation processing on the tire acceleration signal sequence, an acceleration signal change sequence is obtained, and by performing multi-scale one-dimensional time-series feature extraction on the multi-channel time-series input composed of the tire acceleration signal sequence and the acceleration signal change sequence, the tire dynamics features are obtained. Step S3: The tire dynamics features and the state features are concatenated in the channel dimension to obtain physical prior joint features. Modulation weights are generated based on the physical prior joint features. The image semantic features are modulated using the modulation weights to obtain fused features. Step S4: Input the fused features into the classifier and output the road surface state classification result.

2. The method according to claim 1, characterized in that, The tire acceleration signal sequence includes at least one of the following: tire radial acceleration signal sequence, tire tangential acceleration signal sequence, or tire lateral acceleration signal sequence; The tire radial acceleration signal sequence reflects the acceleration change of the tire after being excited by the road surface in the direction perpendicular to the tread. The tire tangential acceleration signal sequence reflects the acceleration change of the tire along the rolling direction; The tire lateral acceleration signal sequence reflects the changes in tire acceleration in the lateral direction.

3. The method according to claim 1, characterized in that, The acceleration signal change sequence in step S2 is obtained through the following steps: Preprocess each tire acceleration signal sequence; The acceleration difference between adjacent sampling points in the preprocessed acceleration signal sequence is calculated to obtain the acceleration signal change sequence.

4. The method according to claim 3, characterized in that, The preprocessing of each tire acceleration signal sequence includes: Each tire acceleration signal sequence is low-pass filtered to obtain the filtered acceleration signal sequence. The filtered acceleration signal sequence is subjected to zero-bias removal and zero-mean normalization to obtain the preprocessed acceleration signal sequence.

5. The method according to claim 4, characterized in that, The cutoff frequency of the low-pass filter is set to 200Hz to 400Hz to filter out high-frequency noise and retain the effective vibration components generated by the contact between the tire and the road surface.

6. The method according to claim 1, characterized in that, The tire dynamic characteristics in step S2 are obtained through the following steps: All tire acceleration signal sequences and their corresponding acceleration signal change sequences are spliced ​​together along the channel dimension to form a multi-channel time series matrix. The multi-channel time series matrix is ​​input into a multi-scale one-dimensional time series feature extraction network to obtain the tire dynamics features; wherein, the multi-scale one-dimensional time series feature extraction network includes: The first one-dimensional convolutional branch is used to extract long-term vibration features; The second one-dimensional convolutional branch is used to extract short-term impact features. The kernel size of the first one-dimensional convolutional branch is larger than the kernel size of the second one-dimensional convolutional branch. The fusion layer is used to stitch together the long-term vibration features and the short-term impact features to obtain joint temporal features; A mapping layer is used to sequentially perform global average pooling and fully connected mapping on the joint temporal features to obtain the tire dynamics features.

7. The method according to any one of claims 1 to 6, characterized in that, Step S1 includes: The vehicle-mounted camera captures images of the road surface in the direction the vehicle is traveling. Tire acceleration signal sequences are collected by sensors mounted on the tires; Preset status information is collected synchronously through the status acquisition unit; The time synchronization unit assigns timestamps to the road surface image, tire acceleration signal sequence, and preset state information, and uses the timestamp of the road surface image as a reference to extract the tire acceleration signal sequence and preset state information within a preset time window to obtain the multimodal data.

8. A road surface condition sensing device based on road images and tire mechanics, characterized in that, The device includes: The data acquisition unit is used to acquire multimodal data, including road surface images, tire acceleration signal sequences, and preset state information. The feature extraction unit is used to extract features from the multimodal data to obtain multimodal features, which include image semantic features, tire dynamics features, and state features. Specifically, by performing interpolation processing on the tire acceleration signal sequence, an acceleration signal change sequence is obtained. The tire dynamics features are obtained by performing multi-scale one-dimensional temporal feature extraction on the multi-channel temporal input composed of the tire acceleration signal sequence and the acceleration signal change sequence. The feature modulation unit is used to concatenate the tire dynamics features and the state features in the channel dimension to obtain physical prior joint features, generate modulation weights based on the physical prior joint features, and use the modulation weights to modulate the image semantic features to obtain fused features. The road surface recognition unit is used to input the fused features into the classifier and output the road surface state classification result.

9. An electronic device, characterized in that, include: processor; as well as A computer-readable storage medium storing computer program instructions that, when executed by the processor, cause the processor to perform the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which is executed by a processor according to any one of claims 1 to 7.