Load, working condition, modal three-dimensional interpretable feature map construction method

CN122528075APending Publication Date: 2026-08-07NANJING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANJING UNIV OF SCI & TECH
Filing Date
2026-07-10
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

在执法取证、道路监管和桥梁入口筛查等应用场景中,仅给出载荷估计值或载荷等级判断结果,难以说明模型判别依据,无法清晰回答模型主要依赖哪些模态、哪些特征、哪些频带、哪些时间片段或哪些空间区域作出判断,也难以明确该判断结果在何种工况下具有较高可靠性

Benefits of technology

[0014]The beneficial effects of this invention are as follows: This invention constructs a three-dimensional feature organization framework through load axis, working condition axis, and modal axis, realizing the organization and display of multi-source heterogeneous features in a unified coordinate system; through t-SNE or UMAP dimensionality reduction visualization, it can observe the clustering distribution patterns of different load levels under different working conditions and modes; through feature contribution analysis, it can quantify the contribution of each modal feature to the load estimation results; through attention weight, frequency band contribution, and spatial regional saliency analysis, it can locate the specific evidence sources for the model to distinguish load levels. Therefore, this invention can provide interpretable fusion priors and reliability boundaries for load classification inversion, and can effectively support the traceability requirements in overload enforcement, road supervision, and bridge safety monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122528075A_ABST
    Figure CN122528075A_ABST
Patent Text Reader

Abstract

The application provides a load, working condition and modal three-dimensional interpretable feature map construction method, comprising: collecting multi-modal observation data of a target vehicle under different load levels and working conditions, synchronizing, registering, denoising, filtering, cutting and labeling to obtain a multi-modal sample set with unified labels; extracting each modal feature and normalizing and uniformly encoding to splice into a multi-dimensional fusion feature vector; constructing a three-dimensional feature organization framework with load, working condition and modal as the axis, and organizing the vector into the corresponding map unit according to the label; inputting the vector into a dimension reduction algorithm to obtain a low-dimensional embedding result to determine the clustering distribution; inputting a load estimation model, calculating the feature contribution degree by using the SHAP method, screening a key feature set and summarizing the contribution result; performing interpretable analysis on the sample set according to the result to locate the evidence source; and finally fusing the above framework, low-dimensional embedding, contribution and evidence source result to generate a three-dimensional interpretable feature map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to data fusion technology in pattern recognition, and in particular to a method for constructing a three-dimensional interpretable feature map of load, operating condition, and modality. Background Technology

[0002] Vehicle load information is a crucial foundational parameter for road transport supervision, bridge safety assessment, enforcement of overloading regulations, and traffic operation status analysis. Accurate vehicle load data not only helps identify overloading behavior but also aids in analyzing vehicle operational risks, assessing road structure load-bearing capacity, and developing refined traffic management strategies. Current methods for acquiring vehicle load mainly include static weighing, dynamic weighing, and manual sampling. Static weighing requires vehicles to be stopped for inspection, yielding relatively accurate results, but its efficiency is low and it is difficult to adapt to continuous traffic flow scenarios. Dynamic weighing can measure load while the vehicle is in motion, but equipment installation and maintenance costs are high, and measurement results are easily affected by vehicle speed, wheelbase, road surface smoothness, vehicle posture, and environmental factors. Manual sampling suffers from limited coverage, insufficient enforcement efficiency, and high evidence collection costs. With the development of sensor technology, computer vision technology, and artificial intelligence technology, vehicle load estimation methods based on multimodal information such as acoustics, vibration, optics, and thermal data are gradually gaining attention. When a vehicle passes through a road detection area under different load conditions, its tire-road contact noise, vehicle structural vibration response, vehicle posture, tire deformation, suspension compression, and heat distribution of tires and braking components may all change. Therefore, multimodal perception data can reflect the vehicle's load state from different physical perspectives.

[0003] However, in real-world open road environments, vehicle load-related characteristics are influenced not only by the load itself but also by various operating conditions such as vehicle speed, road surface type, weather conditions, ambient temperature, lighting conditions, vehicle type, tire condition, and sensor installation location. The same load may exhibit different acoustic, vibrational, visual, or thermal characteristics under different operating conditions; different loads may also show overlapping or aliasing of features under certain complex operating conditions. Without a unified organization of the relationships between load, operating conditions, and modes, it is difficult to assess the reliability of a particular modal characteristic under specific operating conditions. Many existing load estimation methods based on machine learning or deep learning often focus more on prediction accuracy and lack sufficient interpretation of the model output. In applications such as law enforcement, road monitoring, and bridge entrance screening, simply providing load estimates or load level judgments is insufficient to explain the model's judgment criteria, clearly answer which modes, features, frequency bands, time segments, or spatial regions the model primarily relies on for judgment, and also makes it difficult to determine under what operating conditions the judgment result has high reliability. Summary of the Invention

[0004] The technical problem to be solved by this invention is that, in view of the lack of a method in the prior art that can uniformly describe the relationship between load level, operating conditions and multimodal characteristics and perform interpretable analysis on the judgment basis of load estimation model, this invention provides a method for constructing a three-dimensional interpretable feature map of load, operating conditions and modes.

[0005] To address the above problems, this invention provides a method for constructing a three-dimensional interpretable feature map of load, operating condition, and modality, comprising: Step S100: Collect multimodal observation data of the target vehicle under different load levels and different operating conditions; Step S200: Perform time synchronization, spatial registration, noise removal, outlier filtering, event window segmentation, load labeling, and operating condition labeling on the multimodal observation data to obtain a multimodal sample set with unified time index, load label, operating condition label, and modal label. Step S300: Extract features from each modality data in the multimodal sample set, normalize and uniformly encode the extracted modal features, and concatenate them to form a multidimensional fusion feature vector; Step S400: Using load level as load axis, working condition as working condition axis, and modal category as modal axis, construct a three-dimensional feature organization framework for load, working condition, and modality, and organize the multi-dimensional fused feature vectors into the corresponding atlas units according to the label information. Step S500: Input the multidimensional fused feature vector into the dimensionality reduction algorithm for low-dimensional embedding mapping, map the high-dimensional features to the low-dimensional space to obtain the corresponding low-dimensional embedding results, and display them according to the labels to determine the clustering distribution results; Step S600: Input the multidimensional fusion feature vector into the trained load estimation model, use the SHAP method to calculate the contribution of each modal feature to the load estimation result, filter to obtain the key feature set and summarize the key feature contribution results. Step S700: Based on the key feature set, backtrack to the multimodal sample set, perform interpretable analysis, and locate the evidence source results for load estimation; Step S800: The three-dimensional feature organization framework, the low-dimensional embedding result, the key feature contribution result, and the evidence source result are fused to generate a three-dimensional interpretable feature map of load, working condition, and modality.

[0006] Furthermore, the multimodal observation data in step S100 includes one or more of acoustic data, vibration data, optical image data, and thermal image data; the operating conditions include at least one of vehicle speed, road surface type, weather conditions, ambient temperature, lane position, vehicle type, or lighting conditions.

[0007] Furthermore, the multimodal sample set in step S200 is represented as follows: , in, Indicates the first A multimodal sample, Indicates the first The payload label of each sample, Indicates the first Operating condition labels for each sample Indicates the first The modal label of each sample, where N is the total number of samples; The event window is segmented based on the center time t when a vehicle passes through the detection area. i As a reference time, the truncation length is T. w The time window is used to obtain the i-th event window W. i , .

[0008] Furthermore, the feature extraction for each modality data in step S300 specifically includes: Extracting time-domain, frequency-domain, and time-frequency-domain features from acoustic and vibration data; Extract one or more of the following features from optical image data: vehicle contour features, body posture features, tire deformation features, suspension compression features, axle region features, and chassis region features; Extract one or more of the following features from thermal image data: tire temperature features, wheel hub temperature features, braking area temperature features, chassis thermal distribution features, and local temperature rise gradient features.

[0009] Furthermore, in step S300, the j-th original feature x j Normalization is performed using the following standardization method. , in, Represents the normalized i-th One characteristic, Represents the first digit before normalization. 1 eigenvalue, Indicates the first The mean of each feature in the training samples Indicates the first The standard deviation of each feature in the training samples This indicates an extremely small positive number that prevents the denominator from being zero; Multidimensional fusion feature vector F i With the multimodal sample X i One-to-one correspondence, its concatenation expression is: , in, Indicates the first The acoustic feature vector of each sample Indicates the first Vibration feature vector of each sample, Indicates the first The optical feature vector of each sample, Indicates the first Thermal feature vectors of each sample.

[0010] Furthermore, in step S400, the load shaft is divided according to a fixed tonnage range; Using load level l, operating condition c, and modal category m as three-dimensional indices, the atlas element is defined. For the corresponding multi-dimensional fusion feature vector The set, .

[0011] Furthermore, the dimensionality reduction algorithm in step S500 is either t-SNE or UMAP algorithm; The low-dimensional embedding result contains the coordinates of each sample mapped to the low-dimensional embedding space. ,in Indicates the first The low-dimensional embedding coordinates of each sample. Represents the dimensionality reduction mapping function of t-SNE or UMAP; The processing of cluster distribution results includes calculating the intra-class divergence of the l-th loading level. , , in, Indicates the first Number of samples corresponding to each load level Represents the L2 norm, Indicates the first The center of each load level sample in the low-dimensional embedding space; and... Calculate different working conditions under the same load level l and the same mode m. and Cross-condition drift index , , in, and These represent the load ratings as follows: , mode is At that time, under working conditions and working conditions The embedded center below.

[0012] Furthermore, the load estimation model in step S600 is one or more of the following: random forest, gradient boosting tree, support vector machine, XGBoost, convolutional neural network, recurrent neural network, Transformer network, or multimodal fusion network; For each spectral unit, its key feature contribution result is obtained by averaging the absolute values ​​of the key feature contributions of all samples within the unit, resulting in the spectral unit-level average contribution. The expression is , in, Indicates the first The sample set corresponding to each spectral unit. Indicates the first Number of samples in each spectral unit Indicates the first The set of key feature indices obtained by filtering each spectral unit. This represents the absolute value operation. Indicates the first The feature is related to the first The SHAP contribution value of each sample output result.

[0013] Furthermore, the results of locating the source of evidence in step S700 specifically include: When the samples in the multimodal sample set are acoustic data or vibration data, the contribution of the b-th frequency band is defined. To pinpoint the source frequency band and time period of the payload evidence. , in, Indicates belonging to the first A set of feature indexes for each frequency band Indicates the first The contribution value of each feature, Indicates the total number of features; When the samples in the multimodal sample set are optical image data or thermal image data, the spatial saliency intensity at position (x,y) in the image is defined. To locate the spatial region from which the payload evidence originates. , in, This represents the load estimation results output by the model. This represents the pixel value at position (x, y) of the input image.

[0014] The beneficial effects of this invention are as follows: This invention constructs a three-dimensional feature organization framework through load axis, working condition axis, and modal axis, realizing the organization and display of multi-source heterogeneous features in a unified coordinate system; through t-SNE or UMAP dimensionality reduction visualization, it can observe the clustering distribution patterns of different load levels under different working conditions and modes; through feature contribution analysis, it can quantify the contribution of each modal feature to the load estimation results; through attention weight, frequency band contribution, and spatial regional saliency analysis, it can locate the specific evidence sources for the model to distinguish load levels. Therefore, this invention can provide interpretable fusion priors and reliability boundaries for load classification inversion, and can effectively support the traceability requirements in overload enforcement, road supervision, and bridge safety monitoring. Attached Figure Description

[0015] Figure 1 This is a schematic diagram of the method flow of the present invention.

[0016] Figure 2 This is a schematic diagram of the multimodal feature extraction and unified coding process.

[0017] Figure 3 This is a schematic diagram of the process of generating a three-dimensional interpretable feature map. Detailed Implementation

[0018] Prior to detection, multimodal data acquisition devices are deployed in the vehicle passage detection area, and a unified time reference and spatial correspondence are established among the devices. These devices include microphone arrays or roadside microphones for acquiring acoustic information about vehicle passage, vibration sensors for acquiring structural response information, visible light cameras for acquiring images of the vehicle's exterior, and infrared thermal imaging devices for acquiring the temperature distribution of the vehicle's surface. To ensure that subsequent multimodal data can correspond to the same vehicle event, the sampling clocks of each device are uniformly calibrated, and a spatial correspondence is established between acoustic, vibration, optical, and thermal modes based on the vehicle passage trajectory in the detection area.

[0019] Combination Figure 1 A method for constructing a three-dimensional interpretable feature map of load, working condition, and modality, comprising: Step S100: Collect multimodal observation data of the target vehicle under different load levels, different vehicle speeds, different road surface types, and different environmental conditions. The multimodal observation data includes one or more of acoustic data, vibration data, optical image data, and thermal image data. Step S200: Perform time synchronization, spatial registration, noise removal, outlier filtering, event window segmentation, load labeling, and operating condition labeling on the multimodal observation data to obtain a multimodal sample set with unified time index, load label, operating condition label, and modal label. Step S300: Feature extraction is performed on acoustic data, vibration data, optical image data and thermal image data respectively to obtain acoustic features, vibration features, optical features and thermal features, and each modal feature is normalized and uniformly encoded to form a multi-dimensional fusion feature vector; Step S400: Construct a three-dimensional feature organization framework for load, working condition, and modality, using load level as the load axis, working condition as the working condition axis, and modal category as the modal axis. Step S500: Input the multidimensional fused feature vector into the t-SNE or UMAP dimensionality reduction algorithm to map the high-dimensional features to a two-dimensional plane and display them according to the load level to observe the clustering distribution of different load levels under different working conditions and different modes. Step S600: Based on the load estimation model, the contribution of each modal feature to the load estimation result is calculated using the SHAP method, and a set of key features is obtained by filtering according to the magnitude of the contribution. Step S700: Combine attention weights, frequency band contributions, time period contributions, or spatial region contributions to perform interpretable analysis on the key feature set and locate the source frequency band, source time period, or source spatial region of the payload estimation evidence. Step S800: The load level, operating conditions, modal categories, clustering distribution results, key feature contribution, and load evidence source location results are fused to generate a three-dimensional interpretable feature map of load, operating conditions, and modalities.

[0020] The multimodal data acquisition and data alignment process in step S100 includes: after the target vehicle enters the detection area, the detection area information and working condition information are obtained, the multimodal acquisition device is started, and acoustic data, vibration data, optical images and thermal images are collected respectively; then the acquisition results are synchronized in time, registered in space, noise removed and abnormal samples filtered, and the event window is segmented under the condition that the samples are valid and the labels are complete, and the load labeling and working condition labeling are completed, thereby generating a multimodal sample set.

[0021] Once the target vehicle enters the detection area, the system simultaneously collects multimodal observation data under different load levels and operating conditions. These operating conditions include at least one of the following: vehicle speed, road surface type, weather conditions, ambient temperature, lane position, vehicle type, or lighting conditions. During the data collection process, each vehicle passing through the detection area is considered an event to be analyzed, and the corresponding collection time, vehicle number, load source information, and operating condition information are recorded to form the original multimodal observation dataset.

[0022] In step S100, acoustic data is collected by a microphone array, roadside microphone, or directional acoustic sensor; vibration data is collected by an acceleration sensor, piezoelectric sensor, ground vibration sensor, or bridge structure response sensor; optical image data is collected by a visible light camera; and thermal image data is collected by an infrared thermal imaging device.

[0023] Combination Figure 2 Step S200 performs time synchronization, spatial registration, noise removal, outlier filtering, event window segmentation, load labeling, and operating condition labeling on the raw multimodal observation data obtained in step S100. Specifically, firstly, the data output from different acquisition devices are time-aligned according to a unified time reference; secondly, spatial registration is completed based on the mapping relationship between the device installation location and the detection area; then, outlier samples with missing signals, severely blurred images, or incomplete labels are deleted; subsequently, the continuous observation stream is segmented into event windows using the center time of the vehicle passing through the detection area as a reference time, forming corresponding multimodal sample segments; finally, each sample is assigned a load label based on the results of static weighing, dynamic weighing, or external calibration system, and an operating condition label is assigned based on the information recorded during acquisition, thereby obtaining a multimodal sample set with a unified time index, load label, operating condition label, and modal label. Let the total number of samples be N, then the sample set obtained after processing in step S200 can be denoted as:

[0024] in, Indicates the first A multimodal sample, Indicates the first The payload label of each sample, Indicates the first Operating condition labels for each sample Indicates the first Modal labels for each sample.

[0025] In step S200, the operating condition labels include one or more of the following: vehicle speed range, road surface type, weather condition, ambient temperature, lane position, vehicle type, and lighting conditions. The multimodal observation data is segmented into event windows, using the center moment of a vehicle passing through the detection area as the reference moment, and the truncated length is... The time window is used to obtain the first Event window Its expression is: , in, Indicates the first The reference time for the passage of each vehicle. Indicates the length of the event window. Indicates the first A unified analysis time interval for each sample.

[0026] The multimodal feature extraction and unified encoding process in step S300 includes: after inputting the multimodal sample set, selecting the i-th sample and extracting the time-domain, frequency-domain, and time-frequency-domain features of the acoustic sample, extracting the time-domain, frequency-domain, and time-frequency-domain features of the vibration sample, extracting the vehicle contour features, body posture features, and tire, suspension, and axle area features of the optical image sample, and extracting the tire temperature features, wheel hub temperature features, and braking area and chassis thermal distribution features of the thermal image sample; then normalizing, scaling, and stitching the features of each modality to obtain a fused feature vector and forming a fused feature set.

[0027] Combination Figure 2 In step S300, the acoustic and vibration features include time-domain features, frequency-domain features, and time-frequency-domain features. Time-domain features include mean, variance, root mean square value, peak value, skewness, and kurtosis; frequency-domain features include dominant frequency, spectral peak, band energy, and power spectral density; time-frequency-domain features include short-time Fourier transform features, wavelet transform features, and Mel-frequency spectrum features. When normalizing and uniformly encoding the original features extracted from each mode, the first... One original feature The following standardization method is adopted: , in, Represents the normalized i-th One characteristic, Represents the first digit before normalization. 1 eigenvalue, Indicates the first The mean of each feature in the training samples Indicates the first The standard deviation of each feature in the training samples This represents the smallest positive number that prevents the denominator from being zero.

[0028] The normalized modal features are then concatenated to form the first... Multidimensional fusion feature vector of each sample Its expression is: , in, Indicates the first The acoustic feature vector of each sample Indicates the first Vibration feature vector of each sample, Indicates the first The optical feature vector of each sample, Indicates the first Thermal feature vectors of each sample. With sample One-to-one correspondence. For samples with missing modalities, zero-padding, mean-padding, or missing mask encoding methods can be used to maintain a uniform feature dimension across different samples.

[0029] In step S300, the optical features include one or more of the following: vehicle contour features, vehicle body posture features, tire deformation features, suspension compression features, axle area features, and chassis area features; the thermal features include one or more of the following: tire temperature features, wheel hub temperature features, braking area temperature features, chassis thermal distribution features, and local temperature rise gradient features.

[0030] Step S400 inputs the fused feature vector set and its corresponding label information obtained in step S300 into the 3D map construction module, using load level as the load axis, operating condition as the operating condition axis, and modal category as the modal axis to construct a load × operating condition × modal 3D feature organization framework. Preferably, the load axis is divided into 5-ton intervals, the operating condition axis is layered according to at least one of vehicle speed range, road surface type, weather condition, or ambient temperature, and the modal axis includes acoustic modes, vibration modes, optical modes, and thermal modes.

[0031] In step S400, the load axis is divided into 5-ton intervals, including at least two load levels from 0–5 tons, 5–10 tons, 10–15 tons, 15–20 tons, 20–25 tons, 25–30 tons, and greater than 30 tons. For any map unit, the load level is... Operating conditions and modal categories Define the map unit as a three-dimensional index. for: , in, Indicates load rating as Operating conditions are Modal category is The sample set corresponding to the spectral unit. Indicates the first A multidimensional fusion feature vector of a sample Indicates the first The payload label of each sample, Indicates the first Operating condition labels for each sample Indicates the first Each sample's modality label. Each map unit stores at least the following information: number of samples, feature statistics, inter-class separation, intra-class dispersion, contribution of key features, and distribution of evidence regions.

[0032] Step S500 inputs the fused feature vector set obtained in step S300 into the t-SNE or UMAP dimensionality reduction algorithm to perform low-dimensional embedding mapping on the high-dimensional features, obtaining the corresponding two-dimensional or three-dimensional embedding results. Then, the samples are colored, differentiated by shape, or displayed in layers according to load level, working condition category, or modal category to observe the clustering distribution of different load levels under different working conditions and different modalities.

[0033] In step S500, t-SNE or UMAP dimensionality reduction is applied to single-modal features, dual-modal fusion features, or full-modal fusion features, respectively, to compare the inter-class separation, intra-class compactness, and cross-condition drift of different modal combinations in the load classification task. The first... High-dimensional fusion feature vector of each sample The expression mapped to the low-dimensional embedding space is: , in, Indicates the first The low-dimensional embedding coordinates of each sample. This represents the dimensionality reduction mapping function of t-SNE or UMAP.

[0034] Let the first The center of each load level in the embedded space is Then the intraclass divergence of this load level Defined as: , in, Indicates the first Intraclass divergence for each load level, Indicates the first Number of samples corresponding to each load level Represents the L2 norm, Indicates the first The center of each load level sample in the low-dimensional embedding space; Under the same load level and the same mode Under different working conditions and The resulting cross-condition drift index Defined as: , in, and These represent the load ratings as follows: , mode is At that time, under working conditions and working conditions The embedded center below, The smaller the value, the stronger the stability of the mode under various operating conditions.

[0035] Step S600 inputs the fused feature vector set obtained in step S300 into the trained load estimation model, and uses the SHAP method to calculate the contribution of each modal feature to the load estimation result, obtaining sample-level contribution results, modal-level contribution results, and a global key feature set. After filtering the features with high contribution, key feature sets under different load levels, different working conditions, and different modalities can be obtained for subsequent evidence source analysis.

[0036] In this embodiment, the first The feature contribution vector corresponding to each sample is denoted as:

[0037] in, Represents the total number of features. Indicates the first The feature is related to the first The contribution value of the sample load estimation results is calculated. By summing the contribution vectors of samples within the same spectral unit, the key feature contribution results of the corresponding spectral unit can be obtained.

[0038] In step S600, the load estimation model is one or more of the following: random forest, gradient boosting tree, support vector machine, XGBoost, convolutional neural network, recurrent neural network, Transformer network, or multimodal fusion network.

[0039] For the For each sample, the output of the load estimation model can be expressed as: , in, The model represents the first Load estimation results for each sample, Indicates the reference output. Indicates the first The feature is related to the first The SHAP contribution value of each sample output result. This represents the total number of features.

[0040] The average contribution of key features of all samples within the same spectral unit is obtained by averaging the contribution values ​​of these features. Its expression is: , in, Indicates the first Average contribution of each spectral unit Indicates the first The sample set corresponding to each spectral unit. Indicates the first Number of samples in each spectral unit Indicates the first The set of key feature indices obtained by filtering each spectral unit. This represents absolute value operations.

[0041] Step S700, based on the key feature set obtained in step S600, traces back to the original sample data to locate the key evidence regions in the corresponding modes. When the sample is acoustic or vibration data, the key features are mapped to the corresponding frequency band, time segment, or time-frequency region; when the sample is optical or thermal image data, the key features are mapped to the tire contact area, wheel hub area, suspension area, braking area, or chassis thermal anomaly area. After this step, the load evidence source results for each spectral unit in the corresponding mode can be obtained, which are used to explain the location of the main evidence on which the model makes load judgments.

[0042] In this embodiment, the first The evidence sources corresponding to each spectral unit are denoted as:

[0043] in, Indicates the first The first in the map unit One source of evidence, This indicates the number of evidence sources corresponding to this map unit.

[0044] In step S700, when the modal data is acoustic data or vibration data, the source frequency band and source time period of the load evidence are located by using a frequency band contribution map, a time-frequency heat map, or a frequency sensitivity curve; when the modal data is optical image data or thermal image data, the source spatial region of the load evidence is located by using an attention heat map, a saliency map, a class activation map, or a spatial region response map.

[0045] For the Contribution of each frequency band Defined as: , in, Indicates the first The contribution of each frequency band Indicates belonging to the first A set of feature indexes for each frequency band Indicates the first The contribution value of each feature, This represents the total number of features.

[0046] Spatial saliency at position (x, y) in the image Defined as: , in, This indicates the significance of the load estimation result at the image location (x, y). This represents the load estimation results output by the model. This represents the pixel value at position (x, y) in the input image; The larger the value, the stronger the evidence for load discrimination in the corresponding spatial region.

[0047] Combination Figure 3 The three-dimensional interpretable feature map generation process in step S800 includes: obtaining the three-dimensional map framework G output in step S400, the low-dimensional embedding result Z output in step S500, the key feature contribution result P output in step S600, and the evidence source result R output in step S700, and inputting them into the map fusion generation module; then performing credibility evaluation, consistency analysis, and priority ranking in sequence to generate a load × working condition × modal three-dimensional interpretable feature map M, wherein the map includes at least an organization layer, a distribution layer, an interpretation layer, and an evaluation layer.

[0048] The three-dimensional atlas framework obtained in step S400, the low-dimensional clustering results obtained in step S500, the key feature contribution results obtained in step S600, and the evidence source localization results obtained in step S700 are fused to generate the final three-dimensional interpretable feature atlas of load, operating condition, and modality. The three-dimensional interpretable feature atlas includes at least three parts: the first part is the organizational relationship between load, operating condition, and modality; the second part is the distribution results of samples in low-dimensional space; and the third part is the contribution of key features and their corresponding evidence source results.

[0049] In this embodiment, the final generated three-dimensional interpretable feature map is denoted as: , in, Represents a three-dimensional atlas framework. This represents the low-dimensional embedding result. This represents the set of key feature contribution results. This represents the set of evidence sources. Furthermore, the graph can be supplemented with graph unit confidence coefficients, cross-condition drift indices, and multimodal evidence consistency coefficients to evaluate the stability and reliability of different load-condition-modal combinations.

[0050] In step S800, in order to quantify the reliability of the spectral unit for load discrimination, a spectral unit reliability coefficient is defined. for: , in, Indicates the first The credibility coefficient of each spectral unit. Indicates the first Inter-class separation of each spectral unit Indicates the first Intra-class scatter of each spectral unit Indicates the first Average contribution of each spectral unit and These are the weighting coefficients, and , This indicates an extremely small positive number that prevents the denominator from being zero; The larger the value, the more suitable the spectral unit is as a basis for load discrimination.

[0051] To quantify the consistency among multimodal evidence, a multimodal evidence consistency coefficient is defined. for: , in, Represents the consistency coefficient of multimodal evidence. This indicates the number of modes participating in the fusion. and They represent the first The first mode and the first The normalized evidence vector corresponding to each modality This represents the inner product of two modal evidence vectors; The larger the value, the more consistent the direction of evidence given by different modes for the same load.

[0052] To generate modal-condition combination ranking results that can be used for model deployment optimization, a graph priority score is defined. for: , in, Indicates the first Priority score for each map unit Indicates the first The credibility coefficient of each spectral unit. Indicates the first The multimodal evidence consistency coefficient corresponding to each spectral unit. Indicates the first Average drift across operating conditions for each spectral unit , and These are non-negative weighting coefficients; A higher value indicates that the combination of load, operating condition, and modal characteristics is more suitable for priority use in actual systems.

Claims

1. A method for constructing a three-dimensional interpretable feature map of load, operating condition, and modal characteristics, characterized in that, include: Step S100: Collect multimodal observation data of the target vehicle under different load levels and different operating conditions; Step S200: Perform time synchronization, spatial registration, noise removal, outlier filtering, event window segmentation, load labeling, and operating condition labeling on the multimodal observation data to obtain a multimodal sample set with unified time index, load label, operating condition label, and modal label. Step S300: Extract features from each modality data in the multimodal sample set, normalize and uniformly encode the extracted modal features, and concatenate them to form a multidimensional fusion feature vector; Step S400: Using load level as load axis, working condition as working condition axis, and modal category as modal axis, construct a three-dimensional feature organization framework for load, working condition, and modality, and organize the multi-dimensional fused feature vectors into the corresponding atlas units according to the label information. Step S500: Input the multidimensional fused feature vector into the dimensionality reduction algorithm for low-dimensional embedding mapping, map the high-dimensional features to the low-dimensional space to obtain the corresponding low-dimensional embedding results, and display them according to the labels to determine the clustering distribution results; Step S600: Input the multidimensional fusion feature vector into the trained load estimation model, use the SHAP method to calculate the contribution of each modal feature to the load estimation result, filter to obtain the key feature set and summarize the key feature contribution results. Step S700: Based on the key feature set, backtrack to the multimodal sample set, perform interpretable analysis, and locate the evidence source results for load estimation; Step S800: The three-dimensional feature organization framework, the low-dimensional embedding result, the key feature contribution result, and the evidence source result are fused to generate a three-dimensional interpretable feature map of load, working condition, and modality.

2. The method according to claim 1, characterized in that, The multimodal observation data in step S100 includes one or more of acoustic data, vibration data, optical image data, and thermal image data; the operating conditions include at least one of vehicle speed, road surface type, weather conditions, ambient temperature, lane position, vehicle type, or lighting conditions.

3. The method according to claim 2, characterized in that, The multimodal sample set in step S200 is represented as follows: , in, Indicates the first A multimodal sample, Indicates the first The payload label of each sample, Indicates the first Operating condition labels for each sample Indicates the first The modal label of each sample, where N is the total number of samples; The event window is segmented based on the center time t when a vehicle passes through the detection area. i As a reference time, the truncation length is T. w The time window is used to obtain the i-th event window W. i , 。 4. The method according to claim 3, characterized in that, Step S300, which involves feature extraction for each modality of data, specifically includes: Extracting time-domain, frequency-domain, and time-frequency-domain features from acoustic and vibration data; Extract one or more of the following features from optical image data: vehicle contour features, body posture features, tire deformation features, suspension compression features, axle region features, and chassis region features; Extract one or more of the following features from thermal image data: tire temperature features, wheel hub temperature features, braking area temperature features, chassis thermal distribution features, and local temperature rise gradient features.

5. The method according to claim 4, characterized in that, In step S300, the j-th original feature x j Normalization is performed using the following standardization method. , in, Represents the normalized i-th One characteristic, Represents the first digit before normalization. 1 eigenvalue, Indicates the first The mean of each feature in the training samples Indicates the first The standard deviation of each feature in the training samples This indicates an extremely small positive number that prevents the denominator from being zero; Multidimensional fusion feature vector F i With the multimodal sample X i One-to-one correspondence, its concatenation expression is: , in, Indicates the first The acoustic feature vector of each sample, Indicates the first The vibration feature vector of each sample, Indicates the first The optical feature vector of each sample, Indicates the first Thermal feature vectors of each sample.

6. The method according to claim 5, characterized in that, In step S400, the load shafts are divided according to a fixed tonnage range; Using load level l, operating condition c, and modal category m as three-dimensional indices, the atlas element is defined. For the corresponding multi-dimensional fusion feature vector The set, 。 7. The method according to claim 6, characterized in that, The dimensionality reduction algorithm in step S500 is either t-SNE or UMAP. The low-dimensional embedding result contains the coordinates of each sample mapped to the low-dimensional embedding space. ,in Indicates the first The low-dimensional embedding coordinates of each sample. Represents the dimensionality reduction mapping function of t-SNE or UMAP; The processing of cluster distribution results includes calculating the intra-class divergence of the l-th loading level. , , in, Indicates the first Number of samples corresponding to each load level Represents the L2 norm, Indicates the first The center of each load level sample in the low-dimensional embedding space; and... Calculate different working conditions under the same load level l and the same mode m. and Cross-condition drift index , , in, and These represent the load ratings as follows: , mode is At that time, under working conditions and working conditions The embedded center below.

8. The method according to claim 7, characterized in that, The load estimation model in step S600 is one or more of the following: random forest, gradient boosting tree, support vector machine, XGBoost, convolutional neural network, recurrent neural network, Transformer network, or multimodal fusion network; For each spectral unit, its key feature contribution result is obtained by averaging the absolute values ​​of the key feature contributions of all samples within the unit, resulting in the spectral unit-level average contribution. The expression is , in, Indicates the first The sample set corresponding to each spectral unit. Indicates the first Number of samples in each spectral unit Indicates the first The set of key feature indices obtained by filtering each spectral unit. This represents the absolute value operation. Indicates the first The feature is related to the first The SHAP contribution value of each sample output result.

9. The method according to claim 8, characterized in that, The specific results for locating the source of evidence in step S700 include: When the samples in the multimodal sample set are acoustic data or vibration data, the contribution of the b-th frequency band is defined. To pinpoint the source frequency band and time period of the payload evidence. , in, Indicates belonging to the first A set of feature indices for each frequency band Indicates the first The contribution value of each feature, Indicates the total number of features; When the samples in the multimodal sample set are optical image data or thermal image data, the spatial saliency intensity at position (x,y) in the image is defined. To locate the spatial region from which the payload evidence originates. , in, This represents the load estimation results output by the model. This represents the pixel value at position (x, y) of the input image.