Method and system for identifying and evaluating second-hand car soaked in water based on multi-modal feature fusion

By using multimodal detection terminal data acquisition and cross-modal preprocessing, combined with a deep discriminative model and an interpretable reasoning module, the problem of low efficiency and insufficient accuracy in traditional flood-damaged vehicle detection is solved, enabling comprehensive, accurate positioning and interpretable assessment of the risks of flood-damaged vehicles.

CN121859255AInactive Publication Date: 2026-04-14BEIJING KUCHE YIMEI NETWORK TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-26
Publication Date
2026-04-14
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional methods for detecting flooded vehicles rely on manual experience, which is inefficient and inaccurate. They cannot fully capture the multimodal data of flooded vehicles, lack cross-modal feature fusion mechanisms, and the detection results lack evidence and interpretation.

Method used

The system collects vehicle images, electronic system data, odor data, and mechanical condition data through a multimodal detection terminal. It then performs cross-modal preprocessing and feature fusion to generate a water immersion risk representation vector. Combined with a deep discriminant model and an interpretable reasoning module, it outputs a structured assessment report.

Benefits of technology

It enables comprehensive and accurate positioning and interpretable assessment of the risks of flooded vehicles, improves identification efficiency and accuracy, and provides reliable identification basis and continuous optimization capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121859255A_ABST
    Figure CN121859255A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data evaluation, in particular to a method and a system for identifying and evaluating a water-soaked second-hand car based on multi-modal feature fusion. The method comprises the following steps: collecting vehicle multi-source modal data through a multi-modal detection terminal, and generating an original multi-modal data set; performing cross-modal preprocessing on the original multi-modal data set, respectively extracting core features of each modal, and generating a water soaking risk representation vector; inputting the water soaking risk representation vector into a pre-trained depth discrimination model, and generating an evaluation result corresponding to the water soaking probability value and the risk level; generating a structured evaluation report in combination with the contribution degree of each modal feature; and recording the original multi-modal data set, the standardized multi-modal data set, the soaking risk representation vector and the evaluation result, optimizing parameters corresponding to the depth discrimination model through an adaptive incremental learning framework, and sending an output structured evaluation report to a terminal. The second-hand car soaking identification efficiency can be continuously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data evaluation technology, and in particular to a method and system for identifying and evaluating water-damaged used cars based on multimodal feature fusion. Background Technology

[0002] With the rapid growth of car ownership, the used car market continues to expand, and accurate identification of flood-damaged vehicles has become a core requirement for ensuring transaction security and maintaining market order. Flood-damaged vehicles suffer potential damage to their electronic systems, mechanical components, and interior materials due to water inundation or immersion. This not only leads to performance degradation but also poses safety hazards, seriously threatening the personal and property safety of consumers. Traditional flood-damaged vehicle inspection relies on manual experience, judging by observing carpet color, smelling odors, and checking for wiring corrosion. This method is not only inefficient and time-consuming but also easily influenced by subjective factors, resulting in insufficient accuracy and making it difficult to meet the demands of large-scale used car inspections.

[0003] In addition, a similar patent, CN119541080A, discloses a vehicle data analysis method and system based on an intelligent diagnostic instrument. This method includes the following steps: acquiring multimodal vehicle monitoring data using an intelligent diagnostic instrument; calculating the operating power change of the multimodal vehicle monitoring data and mining its temporal state features to obtain the vehicle's temporal state features; optimizing the multimodal vehicle monitoring data with multi-frequency noise reduction and mining its normalized periodic evolution to generate normalized vibration periodic evolution features; performing deep mining of potential feature associations between the vehicle's temporal state features and normalized vibration periodic evolution features, and then performing association graph modeling to construct a vehicle multimodal feature knowledge graph; and calculating non-normalized parameter deviations and tracing fault locations based on the vehicle's real-time operating parameters using the vehicle multimodal feature knowledge graph to obtain potential abnormal fault location points. This invention achieves efficient and accurate vehicle fault risk prediction and analysis. This invention improves the efficiency and accuracy of vehicle fault analysis, but the solution focuses on the diagnosis and tracing of real-time vehicle operation faults and does not optimize for the core needs of water-damaged vehicle identification scenarios. It lacks the integration and processing of multimodal data specific to water-damaged vehicles, such as images, odors, and electronic system data, and has not built a cross-modal feature fusion mechanism adapted to water-damaged risk identification. Furthermore, it does not design an interpretable output module for the detection results, and cannot provide fault location basis and natural language explanation. Summary of the Invention

[0004] To address the aforementioned technical problems in existing used car flood damage assessment processes, this invention provides a method for identifying and assessing flood-damaged used cars. This method utilizes a multimodal detection terminal to comprehensively collect vehicle images, electronic system data, odor data, and mechanical condition data, covering multimodal information specific to flood-damaged vehicles and overcoming the limitations of single data types. Cross-modal preprocessing standardizes and aligns the data temporally, laying the foundation for feature fusion. By leveraging multimodal feature extraction and cross-modal attention fusion, a unified flood damage risk representation vector is generated, meeting the core requirements of flood damage risk identification. A deep discriminant model outputs the flood damage probability and risk level, complemented by an interpretable reasoning module providing a risk evidence chain, anomaly heatmap, and natural language explanations, addressing the issue of lack of evidence in the results. Simultaneously, the entire process data is recorded, and model parameters are optimized through incremental learning to continuously improve the identification capability. This multimodal feature fusion-based method for identifying and assessing flood-damaged used cars includes the following steps: Multi-source modal data of a vehicle is collected by a multi-modal detection terminal. The multi-source modal data of the vehicle includes image data of key parts of the vehicle body, electronic system detection data, odor sensing data and mechanical condition detection data, and generates an original multi-modal dataset. Cross-modal preprocessing is performed on the original multimodal dataset, including image environment normalization, non-image data standardization, and multimodal timestamp alignment, to generate a standardized multimodal dataset; Based on a standardized multimodal dataset, core features of each modality are extracted through a multimodal feature extraction network, and a unified water immersion risk representation vector is generated through cross-modal attention fusion. Input the water immersion risk representation vector into a pre-trained deep discriminant model to generate an assessment result that corresponds to the water immersion probability value and the risk level. The interpretability reasoning module is invoked, and the contribution of each modality feature is combined to generate a structured evaluation report that includes a risk evidence chain, an anomaly location heatmap, and natural language interpretation. Record the original multimodal dataset, the standardized multimodal dataset, the water immersion risk representation vector and the evaluation results, optimize the parameters of the deep discriminative model through an adaptive incremental learning framework, and send the output structured evaluation report to the terminal.

[0005] This invention comprehensively collects images of key parts of the vehicle body, electronic system data, odor sensing data, and mechanical status data through a multimodal detection terminal, covering multimodal information specific to flooded vehicles. This overcomes the shortcomings of traditional solutions, which rely on single data types and cannot capture the multidimensional features of flooded vehicles. Cross-modal preprocessing, including image environment normalization, non-image data standardization, and multimodal timestamp alignment, ensures the consistency and correlation of different data types, laying a solid foundation for subsequent feature fusion. A multimodal feature extraction network extracts core features from each modality, and cross-modal attention fusion generates a unified flood risk representation vector, adapting to the core needs of flood risk identification and overcoming the limitations of similar patents that lack a dedicated cross-modal fusion mechanism. The representation vector is input into a pre-trained deep discriminant model, outputting flood probability values ​​and risk level assessment results. Combined with an interpretable reasoning module, a structured report containing a risk evidence chain, anomaly location heatmap, and natural language explanation is generated based on the contribution of each modality feature, solving the problem of lack of evidence and explanation for detection results. Simultaneously, the entire process data is recorded, and the model parameters are continuously optimized through an adaptive incremental learning framework, constantly improving the system's discrimination capabilities. The overall solution enables comprehensive assessment, precise positioning, and interpretable presentation of risks associated with flooded vehicles, fully meeting the professional needs of flooded vehicle identification scenarios and providing reliable technical support for fields such as used car evaluation and vehicle inspection.

[0006] Preferably, the process of acquiring multi-source modal data of the vehicle through the multi-modal detection terminal includes the following steps: Control the image acquisition device to capture images of the vehicle's carpet, floor, and trunk, obtain high-definition image data from different angles, and record the shooting environment parameters, including light intensity, shooting angle, and ambient temperature; The electronic system detection module collects impedance anomaly data, short circuit signals, and corrosion-related parameters of the circuit system to generate raw electronic detection data. The odor sensor array is activated to capture odor component data inside the vehicle, extract the correlation information of musty smell and moisture residue characteristics, and generate raw odor sensor data; The mechanical testing unit collects water ingress related status data of the engine and transmission, records the operating parameters of mechanical components, and generates raw mechanical testing data. Source identification and timestamp binding are performed on high-definition image data, electronic inspection raw data, odor sensing raw data, and mechanical inspection raw data to generate a data association index; Based on the data association index, various modal data are integrated, and data acquisition device identification and vehicle basic information are supplemented to generate the original multimodal dataset; The original multimodal dataset is subjected to integrity verification, and the missing modality types and data quality levels are marked to provide a basis for preprocessing.

[0007] This invention effectively solves the problems of insufficient data coverage and lack of specific features for flooded vehicles in traditional solutions through targeted multimodal data acquisition and integration. Focusing on key areas prone to anomalies in flooded vehicles, the invention controls image acquisition equipment to capture images of areas such as carpets and floorboards and records environmental parameters, ensuring that the image data aligns with identification requirements. Specialized modules collect electronic data such as abnormal circuit impedance and short-circuit signals, odor data such as musty smells and residual moisture, and mechanical data related to water ingress into the engine and transmission, comprehensively capturing the multi-dimensional specific features of flooded vehicles. Source identifiers and timestamps are added to various data types, a data association index is constructed, and the data is integrated to generate a raw multimodal dataset. Simultaneously, integrity checks are performed to mark data quality, ensuring the comprehensiveness, relevance, and traceability of the data. This step overcomes the limitations of similar patents that focus on general fault data acquisition, specifically adapting to the flooded vehicle identification scenario. It provides rich and relevant foundational data for subsequent cross-modal processing and feature fusion, completely changing the poor identification results caused by insufficient data specificity in traditional solutions.

[0008] Preferably, the cross-modal preprocessing of the original multimodal dataset includes the following steps: The high-definition image data in the original multimodal dataset is subjected to environmental normalization processing. Based on the shooting environment parameters, the illumination adaptive algorithm is called to adjust the brightness and color temperature. The image quality is optimized through noise suppression and edge enhancement techniques to generate adapted image data. Outlier removal is performed on the raw data from electronic detection, odor sensing, and mechanical detection. The data is then converted into intermediate data with unified dimensions using a domain standardization algorithm, generating non-image standardized data. Extract the timestamp information of each modality data, align the acquisition time sequence of different modal data through a time synchronization algorithm, and generate a time-series aligned dataset; Calculate the quality assessment parameters for each modal data and generate modal quality weights; Based on modal quality weights, weighted fusion preprocessing is performed on adapted image data and non-image standardized data to correct data bias and generate a fusion preprocessed dataset. The cross-modal data consistency verification model is invoked to detect semantic conflicts among modal data in the fused preprocessed dataset, eliminate contradictory data and fill in missing information to generate a standardized multimodal dataset.

[0009] This invention employs a full-process cross-modal preprocessing approach to precisely address the shortcomings of multimodal data, such as significant differences in format, asynchronous timing, and inconsistent quality. High-resolution image data undergoes environmental normalization, adjusting brightness and optimizing quality based on shooting parameters to eliminate environmental interference. Non-image data related to electronics, odors, and mechanics are removed for outliers and standardized to achieve dimensional uniformity. A time synchronization algorithm aligns the temporal sequences of each modality, ensuring the rationality of data association. Weighted fusion preprocessing based on modal quality weights corrects data deviations and improves data consistency. A cross-modal consistency verification model is invoked to remove contradictory data and supplement missing information, generating a standardized multimodal dataset. This step constructs a complete preprocessing system of "image normalization - non-image standardization - temporal alignment - weighted fusion - consistency verification," providing a foundation for collaborative analysis of different data types, avoiding feature extraction biases caused by data differences, and compensating for the lack of a dedicated cross-modal preprocessing mechanism in similar patents. This provides high-quality data support for subsequent feature fusion and risk identification.

[0010] Preferably, the environmental normalization processing of the high-definition image data in the original multimodal dataset includes the following steps: High-resolution image data and corresponding shooting environment parameters are extracted from the original multimodal dataset. Light type identifiers and interference factor parameters, including reflectivity and occlusion ratio, are generated through an environmental parameter analysis model. Based on the illumination type identifier, the corresponding brightness compensation algorithm is invoked, and the image brightness balance is adjusted in combination with the illumination intensity parameter to generate a brightness-adapted image. For brightness-adapted images, a color temperature calibration model is used to correct ambient color temperature deviations, eliminate color distortion caused by different light sources, and generate color temperature-calibrated images. The noise detection algorithm identifies salt-and-pepper noise and Gaussian noise in the color temperature calibration image, and the adaptive noise reduction algorithm is called to process the noise intensity parameter to generate a noise-reduced image. An edge detection and enhancement model is used to extract texture edge features from the denoised image, and the detail contrast is enhanced based on the edge sharpness parameter to generate an edge-enhanced image; The edge-enhanced image is normalized in size and format, and the effective region of the image is marked by combining interference factor parameters to generate adapted image data.

[0011] This invention comprehensively solves the problem of unstable image quality due to the influence of the shooting environment by refined image environment normalization processing. It extracts high-resolution images and corresponding shooting environment parameters, analyzes lighting type and interference factors to provide a basis for targeted processing; adjusts brightness balance based on light intensity, corrects color distortion through a color temperature calibration model, and eliminates the influence of different light sources on image color; identifies and processes salt-and-pepper noise and Gaussian noise to improve image purity; extracts texture edge features and enhances detail contrast, making water-damaged features such as carpet mildew and metal corrosion more prominent; standardizes size and format, marks effective areas, and generates image data suitable for subsequent processing. This step, targeting the image features required for water-damaged vehicle identification, achieves full-process optimization from environment adaptation and noise removal to detail enhancement, allowing image data to more clearly present water-damaged abnormal traces, avoiding feature omissions due to image quality issues, and providing high-quality image input for subsequent multimodal feature extraction and water-damaged risk identification, completely changing the current situation where traditional image preprocessing lacks scene specificity.

[0012] Preferably, the step of invoking the corresponding brightness compensation algorithm based on the illumination type identifier and adjusting the image brightness balance in combination with the illumination intensity parameter includes the following steps: Construct a mapping library of illumination type and compensation algorithm, and match the corresponding target brightness compensation algorithm based on the illumination type identifier, including strong light suppression algorithm, backlight compensation algorithm and dark light enhancement algorithm; Semantic parsing of light intensity parameters generates light level identifiers and brightness deviation coefficients; The brightness deviation coefficient is input into the target brightness compensation algorithm to calculate the brightness adjustment range of each pixel in the image and generate a brightness adjustment matrix. The brightness of the original high-definition image data is corrected pixel by pixel based on the brightness adjustment matrix, while preserving the image texture details, and a preliminary brightness-adapted image is generated. The brightness uniformity detection model is invoked to analyze the regional brightness variance of the preliminary brightness-adapted image and generate brightness uniformity coefficients. The brightness adjustment matrix is ​​adjusted according to the brightness equalization coefficient, and the initial brightness adaptation image is corrected a second time to eliminate local overly bright or dark areas and generate a brightness adaptation image. Record the parameter configuration and brightness equalization coefficient adjusted twice, and generate a brightness compensation log for algorithm iteration.

[0013] This invention effectively solves the problems of uneven brightness and significant influence from lighting environment in the original image through scene-based brightness compensation and multiple rounds of optimization. A lighting type-compensation algorithm mapping library is constructed, matching target algorithms such as strong light suppression and backlight compensation according to the lighting type identifier to ensure the compensation scheme fits the actual shooting scene. Light intensity parameters are analyzed to generate lighting levels and brightness deviation coefficients, providing data support for precise adjustments. The pixel-by-pixel adjustment range is calculated based on the brightness deviation coefficients to generate a brightness adjustment matrix, preserving texture details and avoiding feature loss while correcting brightness. A brightness uniformity detection model is used to analyze the regional brightness variance, further correcting overly bright or dark areas to generate a brightness-adapted image. Adjustment parameters and uniformity coefficients are recorded for algorithm iteration. This step achieves a transformation from "general brightness adjustment" to "scene-based precise compensation," resulting in more uniform image brightness and clearer details. It accurately presents water-soaked features such as carpet mold and metal corrosion, providing high-quality input for subsequent image modal feature extraction and compensating for the lack of targeted lighting compensation mechanisms in similar patents.

[0014] Preferably, the process of extracting core features of each modality based on a standardized multimodal dataset using a multimodal feature extraction network, and generating a unified water-soaking risk representation vector through cross-modal attention fusion includes the following steps: Adapted image data from a standardized multimodal dataset is input into a multi-level texture-color fusion network. The bottom layer extracts local texture detail features, and the top layer captures global color distribution features to generate image modality feature vectors. Non-image normalized data is input into the Transformer encoder, and semantic feature vectors for each non-image modality are generated through semantic encoding. Calculate the cross-modal similarity between the image modal feature vector and the semantic feature vectors of each non-image modal, and generate modal association strength parameters; The cross-modal attention fusion model is invoked, and the feature weights of each modality are assigned based on the modal correlation strength parameter to enhance the high-contribution feature channels and generate a fusion feature matrix. The fusion feature matrix is ​​reduced in dimension and filtered to remove redundant feature components, retain the core discriminative features, and generate a simplified fusion feature vector. By combining a knowledge graph in the field of water immersion detection, semantic enhancement is performed on the simplified and fused feature vectors to supplement potential correlation information between modalities and generate a water immersion risk representation vector.

[0015] This invention comprehensively addresses the shortcomings of traditional solutions, such as the lack of a dedicated cross-modal fusion mechanism for water immersion risk and insufficient feature representation capabilities, through deep fusion of multimodal features. Adapted image data is input into a multi-level network to extract local texture and global color features, generating image modality feature vectors. A Transformer encoder performs semantic encoding on non-image data, generating semantic feature vectors. Cross-modal similarity is calculated to generate association strength parameters, making the fusion weight allocation more equitable. A cross-modal attention fusion model is invoked to strengthen high-contribution feature channels, generating a fusion feature matrix. Redundant components are eliminated through dimensionality reduction and feature filtering, and potential association information is supplemented by a domain knowledge graph to generate a water immersion risk representation vector. This process constructs a complete workflow of "single-modal feature extraction - cross-modal association analysis - attention fusion - semantic enhancement," effectively integrating the core features of multimodal data (image, electronic, odor, mechanical) to form a unified representation vector that meets the needs of water immersion risk identification. This overcomes the limitations of similar patents' general fusion mechanisms that are disconnected from scenario requirements, providing input data with strong representation capabilities for subsequent deep discrimination models.

[0016] Preferably, the step of inputting the adapted image data from the standardized multimodal dataset into the multi-level texture-color fusion network includes the following steps: The adapted image data is input into the multi-scale convolution module at the bottom layer of the network. Fine-grained texture features, including fluff morphology and stain details, are extracted through convolution kernels with different receptive fields to generate local texture feature maps. The local texture feature map is aggregated to calculate the texture density and texture continuity parameters, and a texture feature vector is generated. The adapted image data is input into the global feature extraction module of the higher layer of the network to capture the color distribution, the outline of the discoloration area and the diffusion pattern of the stain, and generate a global color feature map. Semantic parsing is performed on the global color feature map to extract parameters such as color deviation degree and the proportion of discoloration area, and to generate a color feature vector. The feature fusion attention mechanism is invoked to calculate the association weights between the texture feature vector and the color feature vector, and feature fusion coefficients are generated. Based on the feature fusion coefficient, the texture feature vector and the color feature vector are weighted and fused to generate a preliminary image feature vector; The initial image feature vector is input into the batch normalization layer and the activation function layer to optimize the feature distribution and generate image modal feature vectors.

[0017] This invention effectively solves the problems of incomplete image modality feature extraction and lack of water damage-related specificity by employing hierarchical feature extraction and precise fusion. The adapted image is input into the multi-scale convolutional module at the bottom layer of the network to extract fine-grained texture features such as fluff morphology and stain details, generating a texture feature vector. A high-level global feature extraction module captures global features such as color distribution and discoloration region contours, generating a color feature vector. A feature fusion attention mechanism is invoked to calculate the correlation weights between the two types of features, generating a preliminary image feature vector through weighted fusion. Batch normalization and activation function optimization of the feature distribution generate the image modality feature vector. This step is specifically designed for water-damaged vehicle identification scenarios, capable of extracting local texture details from key areas such as carpets and floorboards while capturing global color change features. The two types of features complement each other to form a comprehensive image modality representation, accurately reflecting material changes and appearance anomalies caused by water damage. This provides high-quality image feature support for cross-modal fusion, completely changing the current situation where traditional image feature extraction lacks scenario specificity and fails to capture key features.

[0018] Preferably, the calculation of the association weight between the texture feature vector and the color feature vector includes the following steps: Construct a texture-color association knowledge graph, extract the logical association rules between texture features and color features corresponding to water immersion marks, and generate an association rule set; Based on the association rule set, the semantic matching degree between each dimension of the texture feature vector and each dimension of the color feature vector is calculated to generate a dimensional association matrix. The dimensional correlation matrix is ​​normalized to generate a standardized dimensional correlation matrix; The feature importance assessment model is invoked, and labeled data in the field of water immersion detection is combined to calculate the discrimination contribution of each dimension of the texture feature vector and the color feature vector, and generate dimension contribution parameters. The standardized dimensional correlation matrix and the dimensional contribution parameter are multiplied element-wise to generate the dimensional fusion weight matrix. The row summation and column summation of the dimension fusion weight matrix are performed to obtain the overall weights of texture features and color features, respectively. Based on the overall weights of texture features and color features, a feature fusion coefficient is constructed, which includes the weight ratios of texture features and color features. The model is validated by feature fusion coefficients, and the coefficient values ​​are adjusted by combining historical fusion effect data to generate the final feature fusion coefficients.

[0019] This invention effectively addresses the issues of insufficient basis and one-sided weight allocation in texture and color feature fusion by leveraging knowledge graph support and multi-dimensional weight calculation. It constructs a texture-color association knowledge graph specific to water damage traces, extracts logical association rules, and provides a domain-specific benchmark for weight calculation. The semantic matching degree of each dimension of the two types of features is calculated, generating a dimensional association matrix to ensure that weight allocation aligns with the semantic association of features. The discriminative contribution of each dimension is calculated using domain-annotated data, allowing weight allocation to consider both semantic fit and practical identification value. A dimensional fusion weight matrix is ​​generated through element-level multiplication, and the overall weights of the two types of features are obtained through row and column summations, constructing feature fusion coefficients. The coefficient values ​​are adjusted based on historical fusion effect data to generate the final fusion coefficients. This step achieves scientific and domain-specific calculation of feature fusion weights, making the fusion of texture and color features more suitable for identifying water-damaged vehicles, fully leveraging the complementary effects of the two types of features, improving the representational ability of image modal features, and overcoming the shortcomings of similar patents that lack targeted feature fusion weight calculation mechanisms.

[0020] Preferably, the step of calling the feature importance evaluation model and combining it with labeled data in the field of water immersion detection to calculate the discriminative contribution of each dimension of the texture feature vector and color feature vector includes the following steps: Extract labeled image modal feature samples of water-damaged vehicles and normal vehicles from the labeled dataset in the field of water damage detection, and generate a feature evaluation sample set; The sample features in the feature evaluation sample set are decomposed into texture feature components and color feature components, which correspond to the dimensions of the texture feature vector and the color feature vector, respectively. The importance ranking algorithm is invoked to randomly rearrange the texture feature components and color feature components in each dimension, calculate the change in the model's discrimination accuracy before and after the rearrangement, and generate the dimension influence parameter. Based on the dimension influence parameter, and combined with the physical correlation logic between feature dimensions and water immersion traces, a dimension correlation priority is generated. The dimensional influence parameters are normalized to generate standardized influence parameters; The standardized impact parameter is weighted and calculated according to the priority of the dimension association to obtain the preliminary discrimination contribution of each dimension; The contribution calibration model is invoked, and the dimensional contribution error in the historical evaluation data is combined to correct the initial contribution judgment and generate dimensional contribution parameters. Sort the dimensional contribution parameters and mark the core contribution dimensions and secondary contribution dimensions.

[0021] This invention comprehensively addresses the problems of insufficient basis and one-sided results in judging the importance of feature dimensions through multi-dimensional evaluation and calibration. Samples are extracted from labeled datasets in the field of water immersion detection and decomposed into texture and color feature components, providing domain-specific data support for evaluation. An importance ranking algorithm is invoked to calculate the impact of dimensions on the model's discrimination accuracy through random rearrangement, generating dimension influence parameters. Combining the physical correlation logic between feature dimensions and water immersion traces, dimension association priorities are generated to ensure the evaluation aligns with actual identification scenarios. The influence parameters are normalized and weighted with the association priorities to obtain a preliminary discrimination contribution. A contribution calibration model is invoked, and combined with historical evaluation error correction results, dimension contribution parameters are generated, and core and secondary contributing dimensions are marked. This process constructs a complete workflow of "data support - algorithm evaluation - logic verification - error calibration," scientifically determining the discrimination value of each feature dimension, providing a reliable basis for subsequent feature fusion weight allocation, ensuring that core discrimination features receive sufficient attention, and completely changing the current situation where traditional feature importance evaluation lacks domain specificity and scientific basis.

[0022] As a preferred embodiment, the second technical solution of the present invention is: a multimodal feature fusion-based system for identifying and evaluating flooded used cars, used to execute the multimodal feature fusion-based method for identifying and evaluating flooded used cars as described above. This multimodal feature fusion-based system includes a multimodal data acquisition module, a cross-modal preprocessing module, a feature extraction and fusion module, a deep discrimination module, an interpretability reasoning module, a model optimization module, and a result output module. The multimodal data acquisition module is used to acquire multi-source modal data through image acquisition equipment, electronic detection module, odor sensor array, and mechanical detection unit to generate raw multimodal dataset; The cross-modal preprocessing module is used to perform environmental normalization, standardization, and time-series alignment on the original multimodal dataset to generate a standardized multimodal dataset. The feature extraction and fusion module is used to extract core features of each modality and generate a water immersion risk representation vector through cross-modal attention fusion; The depth discrimination module is used to generate water immersion probability values ​​and risk level assessment results based on the water immersion risk characterization vector; The interpretable reasoning module is used to generate a structured assessment report that includes a risk evidence chain, heatmap, and natural language explanation; The model optimization module is used to record data and optimize model parameters through an adaptive incremental learning framework. The result output module is used to push the structured evaluation report to the terminal device, providing support for identification and evaluation decisions.

[0023] This invention effectively addresses the problems of insufficient adaptability, fragmented functions, and lack of interpretability in traditional systems through modular architecture design and end-to-end collaboration. The multimodal data acquisition module comprehensively captures multi-source data including images, electronics, odors, and mechanics, adapting to the specific needs of water-damaged vehicle identification; the cross-modal preprocessing module completes environmental normalization, standardization, and time-series alignment, ensuring data consistency and usability; the feature extraction and fusion module extracts core features from each modality and generates a water-damaged risk representation vector through cross-modal attention fusion, compensating for the lack of a dedicated fusion mechanism in traditional solutions; the deep discrimination module outputs the water-damaged probability and risk level, and the interpretability reasoning module generates a structured report containing a risk evidence chain, heatmap, and natural language explanation, solving the problem of unsubstantiated and unexplained detection results; the model optimization module continuously optimizes parameters through adaptive incremental learning, improving the system's identification capabilities; and the results output module pushes the report to the terminal, providing decision support. The modules work together to build a complete closed loop of "acquisition-processing-fusion-discrimination-interpretation-optimization", which is specifically designed for the identification of water-damaged vehicles. It not only achieves effective integration and in-depth analysis of multimodal data, but also ensures the interpretability of the test results and the continuous evolution of the system, providing comprehensive and reliable technical support for the identification of water-damaged vehicles.

[0024] It has the following beneficial effects: (1) By collecting targeted multimodal data, the problem of insufficient data coverage and lack of specific features for water-damaged vehicles in traditional solutions is effectively solved. Relying on the multimodal detection terminal, image data of key parts of the vehicle body, electronic system detection data, odor sensing data and mechanical condition detection data are collected in a comprehensive manner. The core dimensions of water-damaged vehicles that are prone to abnormalities are specifically focused on: image data captures the appearance traces of carpets, floor panels and other parts; electronic data reflects the dampness and corrosion of circuits; odor data captures features such as musty smells; and mechanical data is associated with the water ingress status of the engine and transmission, forming a complete multimodal data system specific to water-damaged vehicles. The generated original multimodal dataset covers multi-dimensional information related to water damage risk, breaking through the limitations of similar patents that focus on general fault data and lack scenario-specificity. It provides rich and demand-appropriate basic data for subsequent cross-modal processing and feature fusion, ensuring that the entire identification process can be carried out based on comprehensive specific data and avoiding identification bias caused by data loss.

[0025] (2) Through systematic cross-modal preprocessing, the defects of large differences in multimodal data formats, asynchronous time sequences, and significant environmental interference are accurately addressed. Environmental normalization is performed on image data to eliminate quality fluctuations caused by environmental factors such as lighting and shooting angle; non-image data such as electronic, odor, and mechanical data are standardized and converted to achieve dimensional uniformity for different types of data; and multimodal timestamp alignment ensures the temporal correlation of various types of data, laying the foundation for cross-modal feature fusion. The generated standardized multimodal dataset eliminates redundant interference and format barriers in the original data, enabling collaborative analysis of different modal data and avoiding feature extraction deviations caused by inconsistent data quality. This step makes up for the lack of a dedicated cross-modal preprocessing mechanism in similar patents, specifically adapting to the multimodal data characteristics of the water-damaged vehicle identification scenario, providing high-quality and highly consistent data input for subsequent feature fusion and risk assessment, and ensuring the reliability of subsequent processes.

[0026] (3) By using dedicated cross-modal feature fusion, the problems of traditional solutions lacking specificity for water damage risk and insufficient feature representation capabilities are comprehensively solved. Based on a standardized multimodal dataset, the core features of each modality are mined through a multimodal feature extraction network: the image modality captures appearance anomalies such as texture and color, the electronic modality extracts fault signals such as impedance and short circuit, the odor modality identifies features such as musty smell, and the mechanical modality captures abnormal operating status. Then, through a cross-modal attention fusion mechanism, weights are assigned according to the contribution of each modality to water damage risk identification, strengthening highly correlated features and weakening irrelevant interference to generate a unified water damage risk representation vector. This step constructs a cross-modal fusion mechanism adapted to the water damage vehicle identification scenario, breaking through the limitations of similar patent general fusion schemes that are out of touch with scenario requirements. It effectively integrates the core identification information of multimodal data, forming a unified feature representation that can comprehensively reflect water damage risk, providing strong feature input for subsequent deep discrimination, and making risk identification more in line with scenario requirements.

[0027] (4) By pre-training a deep discriminant model, the problem of traditional identification methods relying on human experience and lacking objectivity is effectively solved. A unified water damage risk representation vector is input into a deep discriminant model pre-trained with data from the field of water-damaged vehicles. The model learns water damage risk identification patterns based on massive labeled data and directly outputs water damage probability values ​​and corresponding risk level assessment results, realizing the quantitative assessment and classification of water damage risk. This model is specifically adapted to the water-damaged vehicle identification scenario and can accurately capture water damage-related patterns in multimodal fusion features, avoiding the shortcomings of general models in specific scenarios. Compared with similar patents that focus on fault tracing and lack quantitative assessment of water damage risk, this step can directly provide clear risk judgment results, providing users with intuitive identification basis, which not only improves the efficiency of the identification process, but also ensures the objectivity and consistency of the assessment results, meeting the core judgment requirements of the water-damaged vehicle identification scenario.

[0028] (5) The interpretable reasoning module comprehensively solves the problems of lack of evidence and lack of traceability in traditional solutions. By calling the interpretable reasoning module and combining the contribution analysis of each modality feature, a structured evaluation report is generated, which includes a risk evidence chain, an anomaly location heatmap, and natural language interpretation. The evidence chain clearly defines the basis for the anomalies in each modality, the heatmap intuitively marks the abnormal areas in the image, and the natural language interpretation explains the causes of risks and the judgment logic in plain language. This step fills the gap of similar patents lacking interpretable output. It is specifically designed for the practical application needs of water-damaged car identification scenarios, so that the detection results are no longer just simple numerical values ​​or levels, but structured information with complete evidence and clear interpretation. This not only enhances the credibility of the identification results and makes it easier for users to understand the source of risks and the judgment logic, but also provides clear guidance for subsequent objection verification and fault tracing, meeting the core needs of traceability of results in scenarios such as used car evaluation and inspection.

[0029] (6) By implementing closed-loop iterative optimization and result push, the problem of performance stagnation and inability to continuously adapt to changes in scenarios in traditional solutions is effectively solved. Key data throughout the entire process is recorded, including raw data, standardized data, feature vectors, and evaluation results. Through an adaptive incremental learning framework, new data is integrated into the model training process to continuously optimize the parameters of the deep discrimination model, enabling the model to continuously adapt to new water immersion scenarios, new vehicle types, and new fault modes, thereby improving long-term identification capabilities. At the same time, structured evaluation reports are pushed to the terminal to provide users with intuitive and usable decision support. This step constructs a closed-loop system of "data recording - model optimization," breaking through the limitations of similar patents that lack iterative mechanisms and experience performance degradation after long-term use, ensuring that the system can continuously evolve with the expansion of application scenarios and the accumulation of data. The push of structured reports enables the efficient implementation of identification results, allowing users to quickly obtain complete risk assessment information, meet the decision-making needs in actual testing work, and improve the practicality and sustainability of the entire system. Attached Figure Description

[0030] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 This is a schematic diagram of the steps in the method for identifying and evaluating water-damaged used cars based on multimodal feature fusion according to the present invention. Figure 2 This is a schematic diagram of the modules of the water-damaged used car identification and evaluation system based on multimodal feature fusion of the present invention. Detailed Implementation

[0031] The present invention will be further described below with reference to the accompanying drawings and embodiments, but this should not be construed as limiting the present invention.

[0032] To achieve the above objectives, please refer to Figure 1 Embodiment 1 of the present invention provides a method for identifying and evaluating water-damaged used cars based on multimodal feature fusion, comprising the following steps: S01: Collect vehicle multi-source modal data through a multi-modal detection terminal. The vehicle multi-source modal data includes image data of key parts of the vehicle body, electronic system detection data, odor sensing data, and mechanical condition detection data to generate an original multi-modal dataset. In this embodiment of the invention, a multimodal detection terminal activates various acquisition devices and modules to collect multimodal data of the vehicle. Image acquisition devices focus on key parts of the vehicle body, including the front and rear carpets, trunk carpet, driver's cabin floor, and trunk floor, capturing high-definition images from different angles to obtain visual information such as carpet pile morphology, floor surface condition, and distribution of stains in crevices. The electronic system detection module connects to the vehicle's OBD interface to collect impedance anomaly signals, interface corrosion-related parameters, and short-circuit trigger records of the circuit system, reflecting the impact of water on the circuit system. After activation, the odor sensor array captures odor components inside the vehicle, extracting odor characteristics related to water immersion, such as musty smells and residual moisture. The mechanical detection unit connects to the detection interfaces of the engine and transmission, collecting component operating status parameters and water ingress-related status data, including oil purity and component sealing. Source identifiers and unified timestamps are added to the four types of data, which are then associated with basic vehicle information (vehicle model, registration date, vehicle identification number) and integrated to form the original multimodal dataset. The dataset is validated for completeness to ensure that there are no missing or faulty data in each modality. Data quality levels are then assigned to provide a basis for data quality in subsequent preprocessing steps.

[0033] S02: Perform cross-modal preprocessing on the original multimodal dataset, including image environment normalization, non-image data standardization, and multimodal timestamp alignment, to generate a standardized multimodal dataset; In this embodiment of the invention, cross-modal preprocessing is performed on the original multimodal dataset. For image data of key parts of the vehicle body, environmental normalization is performed: based on environmental parameters such as illumination and temperature recorded during shooting, an adaptive illumination algorithm is used to adjust image brightness and color temperature, eliminating visual interference from strong light, backlight, and low light; noise suppression technology is used to remove particle noise and compression noise from the image, and an edge enhancement algorithm is used to highlight the edges of carpet pile and details of the floor weld, generating adapted image data. Electronic system detection data, odor sensing data, and mechanical condition detection data are standardized: an outlier removal algorithm is used to remove abnormal data exceeding a reasonable range; a domain-adaptive standardization algorithm converts data of different dimensions into intermediate data within a unified semantic range, eliminating interference from data type differences and generating non-image standardized data. Timestamp information of all modal data is extracted, and a time synchronization algorithm is used to align the acquisition time sequence of different modal data, ensuring that the time nodes of image capture, electronic detection, odor capture, and mechanical detection are consistent, avoiding feature association errors caused by time sequence deviations. Finally, a standardized multimodal dataset is generated, providing high-quality data support for feature extraction.

[0034] S03: Based on a standardized multimodal dataset, core features of each modality are extracted through a multimodal feature extraction network, and a unified water immersion risk representation vector is generated through cross-modal attention fusion. In this embodiment of the invention, core features of each modality are extracted using a multimodal feature extraction network based on a standardized multimodal dataset. Adapted image data is input into a multi-level texture-color fusion network. The bottom layer network extracts local texture details, including the curvature of carpet fibers, degree of flattening, stain adhesion patterns, and distribution of impurities in the floor crevices, through multi-scale convolutional modules. The top layer network captures global color features, including color deviation between the carpet and the floor, the contour of discolored areas, and the pattern of water stain diffusion. Local and global features are integrated to generate image modality feature vectors. Non-image standardized data is input into a Transformer encoder. Electronic system detection data is encoded as a circuit state feature vector (reflecting impedance anomalies and corrosion levels), odor sensing data is encoded as an odor component feature vector (reflecting musty odor concentration and moisture residue), and mechanical state detection data is encoded as a mechanical state feature vector (reflecting the associated states of engine and transmission water ingress). A cross-modal attention fusion model is invoked to calculate the semantic association strength between features of each modality. Features that contribute highly to water immersion identification (such as velvet texture, musty smell features, and circuit corrosion parameters) are assigned higher weights. Image modal feature vectors and non-image modal feature vectors are integrated through weighted fusion technology, redundant feature components are eliminated, core discriminative features are strengthened, and a unified water immersion risk representation vector is generated to achieve deep fusion of multimodal data.

[0035] S04: Input the water immersion risk representation vector into a pre-trained deep discriminant model to generate an assessment result corresponding to the water immersion probability value and risk level; In this embodiment of the invention, the water damage risk representation vector is input into a pre-trained deep discriminant model. This model is trained on a large amount of multimodal data of water-damaged and normal used cars, and possesses accurate feature classification capabilities. The model performs semantic parsing on the water damage risk representation vector through a deep neural network, mining the water damage-related feature combinations contained in the vector (such as "fallen down + residual moldy smell + abnormal circuit impedance" and "water stain texture + decreased sealing of mechanical parts"). Based on the mapping relationship between feature combinations and water damage status, a water damage probability value (representing the likelihood that the vehicle is a water-damaged car) is generated. Combining the risk level classification rules in the field of water damage detection, the risk level is determined according to the water damage probability value, divided into high risk (clearly a water-damaged car), medium risk (suspected water-damaged car, requiring further verification), and low risk (not a water-damaged car). An evaluation result containing the water damage probability value and risk level is generated, and the contribution parameters of each feature to the evaluation result are output, providing a basis for subsequent interpretable reasoning.

[0036] S05: Call the interpretability reasoning module, combine the contribution of each modality feature, and generate a structured evaluation report containing a risk evidence chain, anomaly location heatmap, and natural language interpretation; In this embodiment of the invention, an interpretable reasoning module is invoked to generate a structured assessment report based on the evaluation results and the contribution of each modality's features. First, a risk evidence chain is constructed: core discriminative features of each modality are integrated, and key evidence is listed in order of contribution, such as "Image modality: carpet has flattened pile and localized yellowing (highest contribution); Odor modality: obvious musty odor residue detected (second highest contribution); Electronic modality: slight corrosion of circuit interfaces (medium contribution)," clearly defining the core basis for the assessment results. Second, an anomaly location heatmap is generated: using visualization technology, high-risk areas are marked on the original image, and the degree of risk is distinguished by color intensity. For example, yellowed areas of the carpet and stains in the floorboard seams are presented as highlighted heatmaps, visually showing the specific locations of water-related anomalies. Finally, the natural language generation module transforms the assessment results, risk evidence chain, and anomaly location information into concise and easy-to-understand textual explanations, illustrating the correlation between the anomaly characteristics and the water damage. For example, "The vehicle is determined to be a high-risk water-damaged vehicle based on the fact that the carpet fibers have flattened due to water immersion and are accompanied by localized yellowing and discoloration, the presence of common musty smells after water damage is detected inside the vehicle, and slight corrosion of the circuit interfaces further confirms that the vehicle has been exposed to water." This forms a structured assessment report that includes the risk evidence chain, anomaly location heatmap, and natural language explanations.

[0037] S06: Record the original multimodal dataset, the standardized multimodal dataset, the water immersion risk representation vector and the evaluation results, optimize the parameters of the deep discriminant model through an adaptive incremental learning framework, and send the output structured evaluation report to the terminal.

[0038] In this embodiment of the invention, a data recording template is constructed, clearly defining the storage fields, including the original multimodal dataset, standardized multimodal dataset, water damage risk representation vector, assessment results (water damage probability value, risk level), feature contribution parameters, and structured assessment reports. All data is categorized and labeled before being stored in a distributed storage system, and a data association index is established to ensure the traceability of all types of data. An adaptive incremental learning framework is initiated to filter high-confidence samples (including accurately assessed normal vehicles and water-damaged vehicles) from the stored historical data. Combined with user feedback on misjudgments and manual correction information, an incremental training dataset is generated. A semi-supervised training method is adopted, allowing for fine-tuning of model parameters without manual re-labeling. Knowledge distillation technology preserves the core performance of the model, optimizes the weight parameters of deep neural networks, and improves the model's ability to identify new vehicle models and new water damage patterns. Simultaneously, the structured assessment report is pushed to the terminal devices (computers and tablets) of inspection personnel through a standardized interface, supporting online viewing, downloading, and printing of the report. This provides inspection personnel with intuitive and reliable identification and assessment decision support, achieving a closed-loop operation of "data acquisition - model discrimination - result output - model optimization."

[0039] Furthermore, the process of acquiring multi-source modal data of the vehicle through the multi-modal detection terminal includes the following steps: Control the image acquisition device to capture images of the vehicle's carpet, floor, and trunk, obtain high-definition image data from different angles, and record the shooting environment parameters, including light intensity, shooting angle, and ambient temperature; The electronic system detection module collects impedance anomaly data, short circuit signals, and corrosion-related parameters of the circuit system to generate raw electronic detection data. The odor sensor array is activated to capture odor component data inside the vehicle, extract the correlation information of musty smell and moisture residue characteristics, and generate raw odor sensor data; The mechanical testing unit collects water ingress related status data of the engine and transmission, records the operating parameters of mechanical components, and generates raw mechanical testing data. Source identification and timestamp binding are performed on high-definition image data, electronic inspection raw data, odor sensing raw data, and mechanical inspection raw data to generate a data association index; Based on the data association index, various modal data are integrated, and data acquisition device identification and vehicle basic information are supplemented to generate the original multimodal dataset; The original multimodal dataset is subjected to integrity verification, and the missing modality types and data quality levels are marked to provide a basis for preprocessing.

[0040] In this embodiment of the invention, high-definition image acquisition equipment is used to capture high-definition image data of the vehicle's carpets (front, rear, and trunk carpets), floorboards (cabin floorboards and trunk floorboards), and trunk area from multiple different angles (front, side, and top views) for each area, while simultaneously recording environmental parameters such as light intensity, shooting angle, and ambient temperature. An electronic system detection module is connected to the vehicle's OBD interface to collect impedance anomaly data, short-circuit signals, and corrosion-related parameters of the circuit system, generating raw electronic detection data. An odor sensor array is activated to capture odor component data inside the vehicle, detect specific gas concentrations, and extract musty odor characteristic association information and moisture residue characteristic association information, generating raw odor sensing data. A mechanical detection unit is connected to the engine and transmission detection interfaces to collect water ingress-related status data and record mechanical component operating parameters, generating raw mechanical detection data. Source identifiers (image data, electronic data, odor data, and mechanical data) and unified timestamps are added to the four types of data to generate a data association index (index format: timestamp-source identifier-vehicle VIN code). Based on data association indexing, various modal data are integrated, and data collection device identifiers and basic vehicle information (VIN code, vehicle model, registration time) are added to generate the original multimodal dataset. The dataset is then subjected to integrity verification to confirm that there are no missing data in the four modalities, and the data quality level is marked as A (high data completeness, no abnormal missing parameters), providing a basis for preprocessing.

[0041] Furthermore, the cross-modal preprocessing of the original multimodal dataset includes the following steps: The high-definition image data in the original multimodal dataset is subjected to environmental normalization processing. Based on the shooting environment parameters, the illumination adaptive algorithm is called to adjust the brightness and color temperature. The image quality is optimized through noise suppression and edge enhancement techniques to generate adapted image data. Outlier removal is performed on the raw data from electronic detection, odor sensing, and mechanical detection. The data is then converted into intermediate data with unified dimensions using a domain standardization algorithm, generating non-image standardized data. Extract the timestamp information of each modality data, align the acquisition time sequence of different modal data through a time synchronization algorithm, and generate a time-series aligned dataset; Calculate the quality assessment parameters for each modal data and generate modal quality weights; Based on modal quality weights, weighted fusion preprocessing is performed on adapted image data and non-image standardized data to correct data bias and generate a fusion preprocessed dataset. The cross-modal data consistency verification model is invoked to detect semantic conflicts among modal data in the fused preprocessed dataset, eliminate contradictory data and fill in missing information to generate a standardized multimodal dataset.

[0042] In this embodiment of the invention, high-definition image data from the original multimodal dataset undergoes environmental normalization: based on shooting environment parameters (light intensity, no obvious occlusion), an adaptive lighting algorithm is invoked to adjust the image brightness uniformity to the standard range; color temperature calibration corrects environmental color temperature deviations to ensure the color temperature conforms to the standard value; a median filtering algorithm is used to suppress salt-and-pepper noise; and an edge detection operator is used to enhance image edges (carpet pile texture, base plate weld details) to generate adapted image data. Outlier removal is performed on the original electronic detection data, odor sensing data, and mechanical detection data: the interquartile range method is used to remove outliers exceeding reasonable ranges in the electronic data (no outlier data); a standardization algorithm converts the three types of non-image data into intermediate data with unified dimensions, generating non-image standardized data. Timestamp information for each modality is extracted, and a linear interpolation time synchronization algorithm is used to align the acquisition time sequence of different modalities, ensuring minimal time difference among the four types of data, generating a time-aligned dataset. Calculate quality assessment parameters for each modality of data (image data clarity, electronic data stability, odor data sensitivity, and mechanical data reliability), and generate modal quality weights based on these parameters (each assigned a corresponding weight to image, electronic, odor, and mechanical data). Perform weighted fusion preprocessing on the adapted image data and non-image standardized data based on these modal quality weights to correct data biases (such as the correlation deviation between specific parameters in mechanical data and humidity parameters in odor data), generating a fused preprocessed dataset. Call a cross-modal data consistency verification model (trained with a large amount of multimodal detection data) to detect semantic conflicts between modalities in the fused data (such as the contradiction of an image showing no signs of waterlogging but with abnormal odor humidity). In this case, there is no contradictory data. Complete the missing image angle annotation information to generate a standardized multimodal dataset.

[0043] Furthermore, the environmental normalization processing of the high-resolution image data in the original multimodal dataset includes the following steps: High-resolution image data and corresponding shooting environment parameters are extracted from the original multimodal dataset. Light type identifiers and interference factor parameters, including reflectivity and occlusion ratio, are generated through an environmental parameter analysis model. Based on the illumination type identifier, the corresponding brightness compensation algorithm is invoked, and the image brightness balance is adjusted in combination with the illumination intensity parameter to generate a brightness-adapted image. For brightness-adapted images, a color temperature calibration model is used to correct ambient color temperature deviations, eliminate color distortion caused by different light sources, and generate color temperature-calibrated images. The noise detection algorithm identifies salt-and-pepper noise and Gaussian noise in the color temperature calibration image, and the adaptive noise reduction algorithm is called to process the noise intensity parameter to generate a noise-reduced image. An edge detection and enhancement model is used to extract texture edge features from the denoised image, and the detail contrast is enhanced based on the edge sharpness parameter to generate an edge-enhanced image; The edge-enhanced image is normalized in size and format, and the effective region of the image is marked by combining interference factor parameters to generate adapted image data.

[0044] In this embodiment of the invention, high-definition image data and corresponding shooting environment parameters (light intensity, shooting angle, ambient temperature) are extracted from the original multimodal dataset. An environmental parameter analysis model (trained based on a decision tree algorithm) is used to generate a light type identifier (natural light) and interference factor parameters (reflection intensity, occlusion ratio). Based on the light type identifier, a histogram equalization brightness compensation algorithm is invoked, and the image brightness equalization is adjusted in conjunction with the light intensity parameter: the image grayscale histogram is calculated, and the grayscale value distribution range is stretched to ensure the average brightness reaches the standard value, generating a brightness-adapted image. For the brightness-adapted image, a grayscale world method color temperature calibration model is used to correct the ambient color temperature deviation: the average values ​​of the image's RGB three channels are calculated, and the channel gain is adjusted to make the three channel average values ​​equal, eliminating color distortion caused by the warmness of natural light, generating a color temperature calibration image. A noise detection algorithm (based on statistical thresholding) is used to identify salt-and-pepper noise and Gaussian noise in the color temperature calibration image. An adaptive bilateral filtering algorithm is invoked in conjunction with the noise intensity parameter to process the noise while preserving the carpet pile texture details, generating a noise-reduced image. An edge detection and enhancement model is employed to extract texture edge features (carpet pile outline, floorboard seam lines) from the denoised image. Detail contrast is enhanced based on edge sharpness parameters: gamma correction improves grayscale differences in edge areas, generating an enhanced edge image (clearer edges of pile distortion and water stains). The enhanced edge image is then normalized in size (uniformly scaled to standard resolution) and format (unified image format to ensure compression quality). Combined with interference factor parameters (no occlusion), the effective image area is marked (completely covering the carpet and floorboard detection areas, with no invalid background), generating adapted image data. Processing time per image is short.

[0045] Furthermore, the step of invoking the corresponding brightness compensation algorithm based on the illumination type identifier and adjusting the image brightness balance in conjunction with the illumination intensity parameter includes the following steps: Construct a mapping library of illumination type and compensation algorithm, and match the corresponding target brightness compensation algorithm based on the illumination type identifier, including strong light suppression algorithm, backlight compensation algorithm and dark light enhancement algorithm; Semantic parsing of light intensity parameters generates light level identifiers and brightness deviation coefficients; The brightness deviation coefficient is input into the target brightness compensation algorithm to calculate the brightness adjustment range of each pixel in the image and generate a brightness adjustment matrix. The brightness of the original high-definition image data is corrected pixel by pixel based on the brightness adjustment matrix, while preserving the image texture details, and a preliminary brightness-adapted image is generated. The brightness uniformity detection model is invoked to analyze the regional brightness variance of the preliminary brightness-adapted image and generate brightness uniformity coefficients. The brightness adjustment matrix is ​​adjusted according to the brightness equalization coefficient, and the initial brightness adaptation image is corrected a second time to eliminate local overly bright or dark areas and generate a brightness adaptation image. Record the parameter configuration and brightness equalization coefficient adjusted twice, and generate a brightness compensation log for algorithm iteration.

[0046] In this embodiment of the invention, a lighting type-compensation algorithm mapping library is constructed. This library clearly defines the binding relationship between lighting types such as natural light, strong light, backlight, and low light, and their corresponding compensation algorithms. Strong light corresponds to a strong light suppression algorithm, backlight corresponds to a backlight compensation algorithm, and low light corresponds to a low light enhancement algorithm. Based on the lighting type identifier (natural light), a target brightness compensation algorithm (histogram equalization brightness compensation algorithm) is matched from the mapping library. Semantic parsing of the lighting intensity parameters is performed, and the lighting intensity range is determined using an environmental parameter analysis model. This generates a lighting level identifier (medium lighting) and a brightness deviation coefficient (representing the degree of difference between the current brightness and the standard brightness). The brightness deviation coefficient is input into the target brightness compensation algorithm. The algorithm analyzes the image's grayscale distribution characteristics, calculates the brightness adjustment range for each pixel, determines the degree of brightness increase or decrease needed, and generates a brightness adjustment matrix covering all pixels in the image. The original high-definition image data is corrected pixel-by-pixel based on a brightness adjustment matrix. A texture preservation mechanism is introduced during the adjustment process, using edge detection algorithms to locate texture details such as carpet pile and floor seams. The brightness adjustment range of these areas is constrained to prevent texture information loss, generating a preliminary brightness-adapted image. A brightness uniformity detection model is then used to analyze the differences in brightness distribution across different regions of the preliminary brightness-adapted image, calculating brightness fluctuations between regions and generating a brightness uniformity coefficient (representing the overall brightness uniformity). The brightness adjustment matrix is ​​adjusted based on the brightness uniformity coefficient, performing secondary correction on locally overly bright reflective areas and locally overly dark shadow areas. Local contrast enhancement technology is used to balance regional brightness, eliminating brightness unevenness and generating a brightness-adapted image. Algorithm parameter configurations (such as filter kernel size and adjustment range threshold) and changes in the brightness uniformity coefficient are recorded during both adjustment processes, generating a brightness compensation log to provide data support for subsequent algorithm iteration and optimization.

[0047] Furthermore, the step of extracting core features of each modality based on a standardized multimodal dataset using a multimodal feature extraction network, and generating a unified water-soaking risk representation vector through cross-modal attention fusion includes the following steps: Adapted image data from a standardized multimodal dataset is input into a multi-level texture-color fusion network. The bottom layer extracts local texture detail features, and the top layer captures global color distribution features to generate image modality feature vectors. Non-image normalized data is input into the Transformer encoder, and semantic feature vectors for each non-image modality are generated through semantic encoding. Calculate the cross-modal similarity between the image modal feature vector and the semantic feature vectors of each non-image modal, and generate modal association strength parameters; The cross-modal attention fusion model is invoked, and the feature weights of each modality are assigned based on the modal correlation strength parameter to enhance the high-contribution feature channels and generate a fusion feature matrix. The fusion feature matrix is ​​reduced in dimension and filtered to remove redundant feature components, retain the core discriminative features, and generate a simplified fusion feature vector. By combining a knowledge graph in the field of water immersion detection, semantic enhancement is performed on the simplified and fused feature vectors to supplement potential correlation information between modalities and generate a water immersion risk representation vector.

[0048] In this embodiment of the invention, adapted image data (carpet and floor images with brightness compensation and color temperature calibration) from a standardized multimodal dataset are input into a multi-level texture-color fusion network. The bottom layer of the network extracts local texture detail features through a multi-scale convolution module, focusing on capturing the carpet pile shape (whether it is twisted or flattened), stain adhesion details, and floor weld seam features to generate a local texture feature map. The top layer of the network captures the overall color distribution of the image, the contours of discolored areas (such as the boundaries of yellowed and darkened areas), and the stain diffusion pattern (whether it spreads like water stains) through a global feature extraction module, integrating local and global features to generate an image modality feature vector. Non-image standardized data (electronic detection data, odor sensing data, and mechanical detection data) are input into a Transformer encoder. The encoder uses semantic encoding technology to convert various types of non-image data into semantically related feature vectors. Specifically, electronic detection data is encoded as a circuit state feature vector (reflecting impedance anomalies and corrosion levels), odor sensing data is encoded as an odor component feature vector (reflecting musty smells and residual moisture), and mechanical detection data is encoded as a mechanical state feature vector (reflecting the associated state of water ingress in the engine and transmission), generating semantic feature vectors for each non-image modality. A semantic similarity calculation algorithm is used to compare the semantic correlation between the image modality feature vectors and the semantic feature vectors of each non-image modality, generating a modal correlation strength parameter (characterizing the closeness of the correlation between different modal data and water immersion judgment). A cross-modal attention fusion model is invoked, and weights are assigned to each modality feature based on the modal correlation strength parameter. The weight ratio of feature channels that contribute highly to water immersion judgment (such as image features reflecting water stains and odor features reflecting musty smells) is strengthened, while the weight of feature channels that contribute less is weakened. A fusion feature matrix is ​​generated by weighted summation. A feature selection algorithm is employed to reduce the dimensionality and filter features of the fused feature matrix, eliminating redundant and irrelevant features while retaining core discriminative features directly related to the water immersion state (such as the texture of flattened fibers, musty odor, and circuit corrosion features), generating a simplified fused feature vector. Combining this with a knowledge graph in the water immersion detection domain (containing association rules between water immersion state and features of each modality), the simplified fused feature vector is semantically enhanced, supplementing potential intermodal correlation information (such as the logical association between carpet water stain features and moisture residue features), generating a water immersion risk representation vector with comprehensive discriminative capabilities.

[0049] Furthermore, the step of inputting the adapted image data from the standardized multimodal dataset into the multi-level texture-color fusion network includes the following steps: The adapted image data is input into the multi-scale convolution module at the bottom layer of the network. Fine-grained texture features, including fluff morphology and stain details, are extracted through convolution kernels with different receptive fields to generate local texture feature maps. The local texture feature map is aggregated to calculate the texture density and texture continuity parameters, and a texture feature vector is generated. The adapted image data is input into the global feature extraction module of the higher layer of the network to capture the color distribution, the outline of the discoloration area and the diffusion pattern of the stain, and generate a global color feature map. Semantic parsing is performed on the global color feature map to extract parameters such as color deviation degree and the proportion of discoloration area, and to generate a color feature vector. The feature fusion attention mechanism is invoked to calculate the association weights between the texture feature vector and the color feature vector, and feature fusion coefficients are generated. Based on the feature fusion coefficient, the texture feature vector and the color feature vector are weighted and fused to generate a preliminary image feature vector; The initial image feature vector is input into the batch normalization layer and the activation function layer to optimize the feature distribution and generate image modal feature vectors.

[0050] In this embodiment of the invention, adapted image data (preprocessed high-definition images of carpets and floorboards) is input into the bottom-level multi-scale convolutional module of a multi-level texture-color fusion network. This module is configured with convolutional kernels of different receptive fields. Small receptive field kernels focus on extracting subtle morphologies of carpet fibers (the bending state of individual fibers) and details of tiny stain particles, while large receptive field kernels capture texture variations over a slightly larger area (such as local areas of pile collapse). Local texture feature maps are generated through multi-scale feature integration. Feature aggregation processing is performed on the local texture feature maps to calculate parameters such as texture density (the density of the pile distribution) and texture continuity (whether the pile arrangement is continuous). These parameters are then converted into structured features to generate texture feature vectors. Simultaneously, the adapted image data is input into the high-level global feature extraction module of the network. This module uses global pooling and contour detection techniques to capture the overall color distribution of the image (whether there is local color deviation), the contours of discolored areas (the shape and boundary clarity of the discolored areas), and the stain diffusion pattern (whether it spreads along the water flow direction and whether there is water stain bleeding), generating a global color feature map. Semantic parsing is performed on the global color feature map to extract key parameters such as color deviation (the degree of difference between the color and the normal vehicle carpet and floorboard) and the proportion of discolored areas (the proportion of discolored areas covering the entire image). These parameters are then structured to generate color feature vectors. A feature fusion attention mechanism is invoked to analyze the correlation weights between texture and color feature vectors in water immersion detection. By calculating the contribution of the two types of features to water immersion status recognition, feature fusion coefficients are generated. Based on the feature fusion coefficients, the texture and color feature vectors are weighted and fused, organically integrating texture features such as fluff morphology and stain details with color features such as color deviation and discolored areas to generate preliminary image feature vectors. The preliminary image feature vectors are input into a batch normalization layer to standardize the feature distribution and eliminate differences in feature distribution between different images. Subsequently, they are input into an activation function layer to strengthen the response intensity of core features and suppress interference from invalid features, ultimately generating image modality feature vectors with strong discriminative power, providing high-quality image feature support for subsequent cross-modal fusion.

[0051] Furthermore, the calculation of the association weights between the texture feature vector and the color feature vector includes the following steps: Construct a texture-color association knowledge graph, extract the logical association rules between texture features and color features corresponding to water immersion marks, and generate an association rule set; Based on the association rule set, the semantic matching degree between each dimension of the texture feature vector and each dimension of the color feature vector is calculated to generate a dimensional association matrix. The dimensional correlation matrix is ​​normalized to generate a standardized dimensional correlation matrix; The feature importance assessment model is invoked, and labeled data in the field of water immersion detection is combined to calculate the discrimination contribution of each dimension of the texture feature vector and the color feature vector, and generate dimension contribution parameters. The standardized dimensional correlation matrix and the dimensional contribution parameter are multiplied element-wise to generate the dimensional fusion weight matrix. The row summation and column summation of the dimension fusion weight matrix are performed to obtain the overall weights of texture features and color features, respectively. Based on the overall weights of texture features and color features, a feature fusion coefficient is constructed, which includes the weight ratios of texture features and color features. The model is validated by feature fusion coefficients, and the coefficient values ​​are adjusted by combining historical fusion effect data to generate the final feature fusion coefficients.

[0052] In this embodiment of the invention, a texture-color association knowledge graph is constructed. This graph includes logical association rules between texture and color features corresponding to water stains, such as rules like "carpet pile lying flat texture is often accompanied by localized yellowing color features," "the texture of stains in the weld seam of the base plate corresponds to the surrounding darkening color features," and "water stain diffusion texture and gradient color change features appear simultaneously," generating an association rule set. Based on this rule set, a semantic similarity algorithm is used to calculate the semantic matching degree between each dimension of the texture feature vector (such as pile shape, stain details, weld seam texture) and each dimension of the color feature vector (such as color deviation, color change area outline, stain diffusion color). For example, "pilled pile" has a high matching degree with "localized yellowing," and "weld seam stain" has a high matching degree with "surrounding darkening," generating a dimensional association matrix. The dimensional association matrix is ​​then subjected to extreme value normalization, transforming all elements to a fixed interval to eliminate the dimensional differences in matching degrees across different dimensions, generating a standardized dimensional association matrix. A feature importance assessment model is invoked, inputting labeled data from the water immersion detection domain (containing a large number of texture and color feature labeled samples of both water-damaged and normal vehicles). The model analyzes the influence of each dimension of features on water immersion judgment, calculating the discriminative contribution of each dimension of the texture feature vector (e.g., fluff morphology dimension, stain detail dimension) and each dimension of the color feature vector (e.g., color deviation dimension, discoloration area proportion dimension), generating dimension contribution parameters. The standardized dimension correlation matrix and dimension contribution parameters are then element-wise multiplied, i.e., each element in the matrix is ​​multiplied by the corresponding dimension's contribution parameter, strengthening the weights of highly correlated and high-contribution dimensions, generating a dimension fusion weight matrix. Row summation is performed on the dimension fusion weight matrix to obtain the overall weight of texture features; column summation is performed to obtain the overall weight of color features. Based on these two types of overall weights, feature fusion coefficients are constructed, clarifying the weight proportions of texture features and color features. For example, a slightly higher weight for texture features and a slightly lower weight for color features aligns with the more intuitive nature of texture details in water immersion detection. The feature fusion coefficient verification model is invoked, and historical fusion effect data (detection accuracy and false positive rate using different fusion coefficients in the past) are input. The model adjusts the coefficient values ​​through comparative analysis to ensure that the coefficients can maximize detection accuracy and generate the final feature fusion coefficients.

[0053] Furthermore, the step of calling the feature importance evaluation model and combining it with labeled data from the water immersion detection domain to calculate the discriminative contribution of each dimension of the texture feature vector and color feature vector includes the following steps: Extract labeled image modal feature samples of water-damaged vehicles and normal vehicles from the labeled dataset in the field of water damage detection, and generate a feature evaluation sample set; The sample features in the feature evaluation sample set are decomposed into texture feature components and color feature components, which correspond to the dimensions of the texture feature vector and the color feature vector, respectively. The importance ranking algorithm is invoked to randomly rearrange the texture feature components and color feature components in each dimension, calculate the change in the model's discrimination accuracy before and after the rearrangement, and generate the dimension influence parameter. Based on the dimension influence parameter, and combined with the physical correlation logic between feature dimensions and water immersion traces, a dimension correlation priority is generated. The dimensional influence parameters are normalized to generate standardized influence parameters; The standardized impact parameter is weighted and calculated according to the priority of the dimension association to obtain the preliminary discrimination contribution of each dimension; The contribution calibration model is invoked, and the dimensional contribution error in the historical evaluation data is combined to correct the initial contribution judgment and generate dimensional contribution parameters. Sort the dimensional contribution parameters and mark the core contribution dimensions and secondary contribution dimensions.

[0054] In this embodiment of the invention, image modal feature samples of labeled water-damaged vehicles (including lightly, moderately, and severely water-damaged) and normal vehicles are selected from a labeled dataset in the field of water damage detection. The samples need to cover vehicles of different models and service years to ensure sample diversity, generating a feature evaluation sample set. The image modal features of each sample in the feature evaluation sample set are decomposed into texture feature components and color feature components. The texture feature components correspond to the dimensions of the texture feature vector (such as the fluff morphology dimension, stain detail dimension, and weld texture dimension), and the color feature components correspond to the dimensions of the color feature vector (such as the color deviation dimension, discoloration area outline dimension, and discoloration area proportion dimension). The importance ranking algorithm is called to randomly rearrange the dimensions of the texture feature components and color feature components. For example, the feature values ​​of the "fluff morphology" dimension are randomly shuffled. The change in the model's discrimination accuracy before and after the rearrangement is compared. If the accuracy drops significantly after the rearrangement, it indicates that this dimension is crucial for discrimination, generating a dimension influence parameter. Based on the dimensional influence parameter, and combined with the physical correlation logic between feature dimensions and water immersion traces (e.g., "pile flattening" is a typical physical change in carpets after water immersion, and "color deviation" is a direct manifestation of material oxidation after water immersion), the dimensional correlation priority is determined. For example, "pile morphology" and "color deviation" have higher priority than "weld texture" and "discoloration area outline." The dimensional influence parameter is standardized by converting all parameters to a uniform range to eliminate interference from numerical differences, generating standardized influence parameters. The standardized influence parameters are then weighted and calculated with the dimensional correlation priority, assigning higher weights to dimensions with higher priority, resulting in the preliminary judgment contribution of each dimension. A contribution calibration model is then invoked, inputting the dimensional contribution error from historical evaluation data (the deviation between the contribution of each dimension in past evaluations and the actual detection effect). The model corrects the preliminary judgment contribution through error feedback, reducing the contribution of high-error dimensions and increasing the contribution of low-error dimensions, generating dimensional contribution parameters. The dimensional contribution parameters are sorted by numerical value, and the core contributing dimensions (such as "fluff morphology" and "color deviation") and secondary contributing dimensions (such as "weld texture" and "color change region outline") are marked to provide a priority basis for subsequent feature fusion.

[0055] Furthermore, Embodiment 2 of the present invention also provides a system for identifying and evaluating flood-damaged used cars based on multimodal feature fusion, used to execute the method for identifying and evaluating flood-damaged used cars based on multimodal feature fusion as described above. This system includes a multimodal data acquisition module, a cross-modal preprocessing module, a feature extraction and fusion module, a deep discrimination module, an interpretability reasoning module, a model optimization module, and a result output module. The multimodal data acquisition module is used to acquire multi-source modal data through image acquisition equipment, electronic detection module, odor sensor array, and mechanical detection unit to generate raw multimodal dataset; The cross-modal preprocessing module is used to perform environmental normalization, standardization, and time-series alignment on the original multimodal dataset to generate a standardized multimodal dataset. The feature extraction and fusion module is used to extract core features of each modality and generate a water immersion risk representation vector through cross-modal attention fusion; The depth discrimination module is used to generate water immersion probability values ​​and risk level assessment results based on the water immersion risk characterization vector; The interpretable reasoning module is used to generate a structured assessment report that includes a risk evidence chain, heatmap, and natural language explanation; The model optimization module is used to record data and optimize model parameters through an adaptive incremental learning framework. The result output module is used to push the structured evaluation report to the terminal device, providing support for identification and evaluation decisions.

[0056] In this embodiment of the invention, the water-damaged used car identification and evaluation system based on multimodal feature fusion includes a multimodal data acquisition module, a cross-modal preprocessing module, a feature extraction and fusion module, a deep discrimination module, an interpretability reasoning module, a model optimization module, and a result output module. These modules interact collaboratively through standardized data interfaces. The multimodal data acquisition module controls an image acquisition device to capture high-definition images of the vehicle's carpet, floor, and trunk from different angles; it connects to the vehicle's OBD interface via an electronic detection module to collect parameters related to circuit system impedance anomalies and corrosion; it activates an odor sensor array to capture data on the components of musty and residual moisture odors inside the vehicle; and it connects to the engine and transmission interfaces via a mechanical detection unit to collect water ingress-related states and operating parameters, integrating all data to generate a raw multimodal dataset. The cross-modal preprocessing module performs environmental normalization processing on the high-definition images in the raw multimodal dataset, adjusting brightness and color temperature to standard states and optimizing image quality; it removes outliers and standardizes electronic, odor, and mechanical data, unifying their dimensions; and it aligns the acquisition sequence of each modality using a time synchronization algorithm to generate a standardized multimodal dataset. The feature extraction and fusion module inputs adapted image data into a multi-level texture-color fusion network to extract image modal feature vectors; it inputs non-image data into a Transformer encoder to generate semantic feature vectors for each non-image modality; and it calls a cross-modal attention fusion model to assign weights based on modal association strength, fusing them to generate a water damage risk representation vector. The deep discrimination module inputs the water damage risk representation vector into a pre-trained deep classification model. The model analyzes the mapping relationship between feature vectors and water damage status to generate water damage probability values ​​(e.g., high probability, medium probability, low probability) and risk level assessment results (e.g., high risk, medium risk, low risk). The interpretable reasoning module uses Grad-CAM technology to generate water stain heatmaps of carpet and floor images, marking high-risk areas; it combines the contribution of each modality feature to generate a "risk level + key evidence chain" report, such as "High risk: The carpet has a flattened pile texture (high texture feature contribution) and localized yellowing (high color feature contribution); a musty smell was detected inside the car (high odor feature contribution)." Through the natural language generation module, it uses concise text to describe the cause of the anomaly, generating a structured assessment report. The model optimization module, based on an adaptive incremental learning framework, automatically collects new detection data (including user-reported misjudgments and missed judgments). It uses a pseudo-label generation mechanism to filter high-confidence samples, enabling semi-supervised model retraining without manual annotation. Regular knowledge distillation and parameter fine-tuning are performed to incorporate features from new vehicle models and new water-damage patterns, optimizing model parameters. The results output module pushes a structured evaluation report (including heatmaps, evidence chains, and textual explanations) to the testing personnel's terminal devices (such as computers and tablets). The report can be viewed, downloaded, and printed online, providing testing personnel with decision support for identification and evaluation, ensuring transparent and reliable test results.

[0057] Therefore, the embodiments should be considered as exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of the equivalents of the application are intended to be included within the invention.

[0058] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.

Claims

1. A method for identifying and evaluating water-damaged used cars based on multimodal feature fusion, characterized in that, Includes the following steps: Multi-source modal data of a vehicle is collected by a multi-modal detection terminal. The multi-source modal data of the vehicle includes image data of key parts of the vehicle body, electronic system detection data, odor sensing data and mechanical condition detection data, and generates an original multi-modal dataset. Cross-modal preprocessing is performed on the original multimodal dataset, including image environment normalization, non-image data standardization, and multimodal timestamp alignment, to generate a standardized multimodal dataset; Based on a standardized multimodal dataset, core features of each modality are extracted through a multimodal feature extraction network, and a unified water immersion risk representation vector is generated through cross-modal attention fusion. Input the water immersion risk representation vector into a pre-trained deep discriminant model to generate an assessment result that corresponds to the water immersion probability value and the risk level. The interpretability reasoning module is invoked, and the contribution of each modality feature is combined to generate a structured evaluation report that includes a risk evidence chain, an anomaly location heatmap, and natural language interpretation. Record the original multimodal dataset, the standardized multimodal dataset, the water immersion risk representation vector and the evaluation results, optimize the parameters of the deep discriminative model through an adaptive incremental learning framework, and send the output structured evaluation report to the terminal.

2. The method for identifying and evaluating water-damaged used cars based on multimodal feature fusion according to claim 1, characterized in that, The process of collecting multi-source modal data of the vehicle through a multi-modal detection terminal includes the following steps: Control the image acquisition device to capture images of the vehicle's carpet, floor, and trunk, obtain high-definition image data from different angles, and record the shooting environment parameters, including light intensity, shooting angle, and ambient temperature; The electronic system detection module collects impedance anomaly data, short circuit signals, and corrosion-related parameters of the circuit system to generate raw electronic detection data. The odor sensor array is activated to capture odor component data inside the vehicle, extract the correlation information of musty smell and moisture residue characteristics, and generate raw odor sensor data; The mechanical testing unit collects water ingress related status data of the engine and transmission, records the operating parameters of mechanical components, and generates raw mechanical testing data. Source identification and timestamp binding are performed on high-definition image data, electronic inspection raw data, odor sensing raw data, and mechanical inspection raw data to generate a data association index; Based on the data association index, various modal data are integrated, and data acquisition device identification and vehicle basic information are supplemented to generate the original multimodal dataset; The original multimodal dataset is subjected to integrity verification, and the missing modality types and data quality levels are marked to provide a basis for preprocessing.

3. The method for identifying and evaluating water-damaged used cars based on multimodal feature fusion according to claim 1, characterized in that, The cross-modal preprocessing of the original multimodal dataset includes the following steps: The high-definition image data in the original multimodal dataset is subjected to environmental normalization processing. Based on the shooting environment parameters, the illumination adaptive algorithm is called to adjust the brightness and color temperature. The image quality is optimized through noise suppression and edge enhancement techniques to generate adapted image data. Outlier removal is performed on the raw data from electronic detection, odor sensing, and mechanical detection. The data is then converted into intermediate data with unified dimensions using a domain standardization algorithm, generating non-image standardized data. Extract the timestamp information of each modality data, align the acquisition time sequence of different modal data through a time synchronization algorithm, and generate a time-series aligned dataset; Calculate the quality assessment parameters for each modal data and generate modal quality weights; Based on modal quality weights, weighted fusion preprocessing is performed on adapted image data and non-image standardized data to correct data bias and generate a fusion preprocessed dataset. The cross-modal data consistency verification model is invoked to detect semantic conflicts among modal data in the fused preprocessed dataset, eliminate contradictory data and fill in missing information to generate a standardized multimodal dataset.

4. The method for identifying and evaluating water-damaged used cars based on multimodal feature fusion according to claim 3, characterized in that, The environmental normalization process for the high-resolution image data in the original multimodal dataset includes the following steps: High-resolution image data and corresponding shooting environment parameters are extracted from the original multimodal dataset. Light type identifiers and interference factor parameters, including reflectivity and occlusion ratio, are generated through an environmental parameter analysis model. Based on the illumination type identifier, the corresponding brightness compensation algorithm is invoked, and the image brightness balance is adjusted in combination with the illumination intensity parameter to generate a brightness-adapted image. For brightness-adapted images, a color temperature calibration model is used to correct ambient color temperature deviations, eliminate color distortion caused by different light sources, and generate color temperature-calibrated images. The noise detection algorithm identifies salt-and-pepper noise and Gaussian noise in the color temperature calibration image, and the adaptive noise reduction algorithm is called to process the noise intensity parameter to generate a noise-reduced image. An edge detection and enhancement model is used to extract texture edge features from the denoised image, and the detail contrast is enhanced based on the edge sharpness parameter to generate an edge-enhanced image; The edge-enhanced image is normalized in size and format, and the effective region of the image is marked by combining interference factor parameters to generate adapted image data.

5. The method for identifying and evaluating water-damaged used cars based on multimodal feature fusion according to claim 4, characterized in that, The step of invoking the corresponding brightness compensation algorithm based on the illumination type identifier and adjusting the image brightness balance in combination with the illumination intensity parameter includes the following steps: Construct a mapping library of illumination type and compensation algorithm, and match the corresponding target brightness compensation algorithm based on the illumination type identifier, including strong light suppression algorithm, backlight compensation algorithm and dark light enhancement algorithm; Semantic parsing of light intensity parameters generates light level identifiers and brightness deviation coefficients; The brightness deviation coefficient is input into the target brightness compensation algorithm to calculate the brightness adjustment range of each pixel in the image and generate a brightness adjustment matrix. The brightness of the original high-definition image data is corrected pixel by pixel based on the brightness adjustment matrix, while preserving the image texture details, and a preliminary brightness-adapted image is generated. The brightness uniformity detection model is invoked to analyze the regional brightness variance of the preliminary brightness-adapted image and generate brightness uniformity coefficients. The brightness adjustment matrix is ​​adjusted according to the brightness equalization coefficient, and the initial brightness adaptation image is corrected a second time to eliminate local overly bright or dark areas and generate a brightness adaptation image. Record the parameter configuration and brightness equalization coefficient adjusted twice, and generate a brightness compensation log for algorithm iteration.

6. The method for identifying and evaluating water-damaged used cars based on multimodal feature fusion according to claim 1, characterized in that, The process of extracting core features of each modality based on a standardized multimodal dataset, and generating a unified water-soaking risk representation vector through cross-modal attention fusion, includes the following steps: Adapted image data from a standardized multimodal dataset is input into a multi-level texture-color fusion network. The bottom layer extracts local texture detail features, and the top layer captures global color distribution features to generate image modality feature vectors. Non-image normalized data is input into the Transformer encoder, and semantic feature vectors for each non-image modality are generated through semantic encoding. Calculate the cross-modal similarity between the image modal feature vector and the semantic feature vectors of each non-image modal, and generate modal association strength parameters; The cross-modal attention fusion model is invoked, and the feature weights of each modality are assigned based on the modal correlation strength parameter to enhance the high-contribution feature channels and generate a fusion feature matrix. The fusion feature matrix is ​​reduced in dimension and filtered to remove redundant feature components, retain the core discriminative features, and generate a simplified fusion feature vector. By combining a knowledge graph in the field of water immersion detection, semantic enhancement is performed on the simplified and fused feature vectors to supplement potential correlation information between modalities and generate a water immersion risk representation vector.

7. The method for identifying and evaluating water-damaged used cars based on multimodal feature fusion according to claim 6, characterized in that, The process of inputting adapted image data from a standardized multimodal dataset into a multi-level texture-color fusion network includes the following steps: The adapted image data is input into the multi-scale convolution module at the bottom layer of the network. Fine-grained texture features, including fluff morphology and stain details, are extracted through convolution kernels with different receptive fields to generate local texture feature maps. The local texture feature map is aggregated to calculate the texture density and texture continuity parameters, and a texture feature vector is generated. The adapted image data is input into the global feature extraction module of the higher layer of the network to capture the color distribution, the outline of the discoloration area and the diffusion pattern of the stain, and generate a global color feature map. Semantic parsing is performed on the global color feature map to extract parameters such as color deviation degree and the proportion of discoloration area, and to generate a color feature vector. The feature fusion attention mechanism is invoked to calculate the association weights between the texture feature vector and the color feature vector, and feature fusion coefficients are generated. Based on the feature fusion coefficient, the texture feature vector and the color feature vector are weighted and fused to generate a preliminary image feature vector; The initial image feature vector is input into the batch normalization layer and the activation function layer to optimize the feature distribution and generate image modal feature vectors.

8. The method for identifying and evaluating water-damaged used cars based on multimodal feature fusion according to claim 7, characterized in that, The calculation of the association weight between the texture feature vector and the color feature vector includes the following steps: Construct a texture-color association knowledge graph, extract the logical association rules between texture features and color features corresponding to water immersion marks, and generate an association rule set; Based on the association rule set, the semantic matching degree between each dimension of the texture feature vector and each dimension of the color feature vector is calculated to generate a dimensional association matrix. The dimensional correlation matrix is ​​normalized to generate a standardized dimensional correlation matrix; The feature importance assessment model is invoked, and labeled data in the field of water immersion detection is combined to calculate the discrimination contribution of each dimension of the texture feature vector and the color feature vector, and generate dimension contribution parameters. The standardized dimensional correlation matrix and the dimensional contribution parameter are multiplied element-wise to generate the dimensional fusion weight matrix. The row summation and column summation of the dimension fusion weight matrix are performed to obtain the overall weights of texture features and color features, respectively. Based on the overall weights of texture features and color features, a feature fusion coefficient is constructed, which includes the weight ratios of texture features and color features. The model is validated by feature fusion coefficients, and the coefficient values ​​are adjusted by combining historical fusion effect data to generate the final feature fusion coefficients.

9. The method for identifying and evaluating water-damaged used cars based on multimodal feature fusion according to claim 8, characterized in that, The step of calling the feature importance evaluation model, combined with labeled data in the field of water immersion detection, to calculate the discriminative contribution of each dimension of the texture feature vector and color feature vector includes the following steps: Extract labeled image modal feature samples of water-damaged vehicles and normal vehicles from the labeled dataset in the field of water damage detection, and generate a feature evaluation sample set; The sample features in the feature evaluation sample set are decomposed into texture feature components and color feature components, which correspond to the dimensions of the texture feature vector and the color feature vector, respectively. The importance ranking algorithm is invoked to randomly rearrange the texture feature components and color feature components in each dimension, calculate the change in the model's discrimination accuracy before and after the rearrangement, and generate the dimension influence parameter. Based on the dimension influence parameter, and combined with the physical correlation logic between feature dimensions and water immersion traces, a dimension correlation priority is generated. The dimensional influence parameters are normalized to generate standardized influence parameters; The standardized impact parameter is weighted and calculated according to the priority of the dimension association to obtain the preliminary discrimination contribution of each dimension; The contribution calibration model is invoked, and the dimensional contribution error in the historical evaluation data is combined to correct the initial contribution judgment and generate dimensional contribution parameters. Sort the dimensional contribution parameters and mark the core contribution dimensions and secondary contribution dimensions.

10. A system for identifying and evaluating water-damaged used cars based on multimodal feature fusion, characterized in that, For executing the multimodal feature fusion-based method for identifying and evaluating flooded used cars as described in claim 1, the multimodal feature fusion-based system for identifying and evaluating flooded used cars includes a multimodal data acquisition module, a cross-modal preprocessing module, a feature extraction and fusion module, a deep discrimination module, an interpretable reasoning module, a model optimization module, and a result output module. The multimodal data acquisition module is used to acquire multi-source modal data through image acquisition equipment, electronic detection module, odor sensor array, and mechanical detection unit to generate raw multimodal dataset; The cross-modal preprocessing module is used to perform environmental normalization, standardization, and time-series alignment on the original multimodal dataset to generate a standardized multimodal dataset. The feature extraction and fusion module is used to extract core features of each modality and generate a water immersion risk representation vector through cross-modal attention fusion; The depth discrimination module is used to generate water immersion probability values ​​and risk level assessment results based on the water immersion risk characterization vector; The interpretable reasoning module is used to generate a structured assessment report that includes a risk evidence chain, heatmap, and natural language explanation; The model optimization module is used to record data and optimize model parameters through an adaptive incremental learning framework. The result output module is used to push the structured evaluation report to the terminal device, providing support for identification and evaluation decisions.

Citation Information

Patent Citations

  • Automobile data analysis method and system based on intelligent diagnostic instrument

    CN119541080A