Automobile quality risk assessment system and method based on new data fusion technology
By building a cross-stage and cross-modal data fusion system and graph neural network structure, the hidden defect risk is quantified, and the problems of low hidden defect recognition rate and high misjudgment rate in the existing technology are solved, and accurate assessment and forward-looking management of automobile quality risks are achieved.
Patent Information
- Application Number
- CN202510662091.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-05-22
AI Technical Summary
When handling vehicle quality risk identification and early warning, it is difficult to accurately distinguish hidden defect characteristics from normal states, resulting in omission of fault pattern recognition or risk signal obscuring, affecting user safety and quality management forward-looking responses.
By building a cross-stage and cross-modal data fusion system, using graph neural network structure and feature engineering technology, quantifying the risk of hidden defects, and combining with local amplification mechanism, intelligent identification and classification of hidden defect features is achieved.
It significantly improves the accuracy and prospectiveness of quality defect identification, avoids hidden defects being overwhelmed by pseudo-normal clusters, and realizes a full-process closed-loop management from perception-modeling-evaluation-enhanced, improving the sensitivity and accuracy of the quality evaluation system.
Smart Images

Figure CN120181592B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of automobile manufacturing and quality management, and in particular to an automobile quality risk assessment system and method based on a novel data fusion technology. Background Art
[0002] "Automotive Quality Risk Assessment Based on New Data Fusion Technology" utilizes multi-source heterogeneous data fusion methods to uniformly model and deeply fuse structured and unstructured data (such as sensor data, fault codes, maintenance records, and user feedback) from multiple stages of the automotive manufacturing process, assembly process, factory inspection, after-sales maintenance, and driving status monitoring, thereby constructing a comprehensive quality data profile covering the entire vehicle lifecycle. By introducing intelligent algorithms such as machine learning, deep neural networks, Bayesian reasoning, or graphical models, the fused data is dynamically modeled and risk factors mined to identify patterns and regularities in potential quality defects. This allows for quantitative assessment and early warning of potential quality risks under specific vehicle models, components, or operating conditions, enabling a shift from passive failure response to proactive quality control and improving the precision and foresight of automotive quality management.
[0003] Existing technologies suffer from the following deficiencies: Intelligent assessment models for automotive quality risk identification and early warning typically rely on the fusion and modeling of multi-source heterogeneous data (such as manufacturing parameters, vehicle sensor data, and user feedback) to automatically identify and discover patterns in potential quality defects. However, existing technologies have significant limitations when dealing with some failure modes that exhibit latent defect characteristics. For example, while key failure modes such as progressive brake system failure and precursors to battery thermal runaway exhibit clear evolutionary trends at the physical level, their data-dimensional characteristics often closely resemble natural fluctuations or operating disturbances under normal operating conditions. This results in overlapping or proximity of feature dimensions with "non-risk samples" in the feature space after data fusion, creating a "pseudo-normal clustering" effect. Under these circumstances, existing training models struggle to accurately distinguish these failure evolutionary characteristics from regular fluctuations, easily misclassifying them as normal samples during training or significantly underestimating their risk levels during subsequent inference. This can lead to missed fault mode identification or masking of risk signals. This type of misjudgment problem is not easy to detect in the early stage of system deployment, but once it enters the actual vehicle use stage, it is very easy to cause delayed risk exposure, affecting user safety and the forward-looking response capability of quality management.
[0004] The above information disclosed in this Background section is only for enhancement of understanding of the background of the present disclosure and therefore it may contain information that does not form the prior art that is already known to a person of ordinary skill in the art. Summary of the Invention
[0005] The purpose of the present invention is to provide an automobile quality risk assessment system and method based on new data fusion technology. By fusing structured and unstructured multimodal data, a graph neural network structure is constructed and key indicators are introduced to quantify the risk of hidden defects. The local amplification mechanism is combined to enhance the expression of weak abnormal features, and an integrated closed-loop recognition process of perception, modeling, evaluation, and enhancement is realized, which significantly improves the accuracy and foresight of quality defect identification to solve the problems in the above-mentioned background technology.
[0006] In order to achieve the above object, the present invention provides the following technical solution: a method for automobile quality risk assessment based on a novel data fusion technology, comprising the following steps:
[0007] By building a data collection mechanism covering the entire process of vehicle design, manufacturing, operation and after-sales maintenance, a cross-stage and cross-modal integrated data fusion system is established to comprehensively collect and manage multi-modal feature information including structured and unstructured data;
[0008] After preprocessing the original multimodal data, we build a graph neural network structure for implicit association modeling by combining the semantic associations, temporal alignment, and structural coupling characteristics between the features of each modality.
[0009] Through feature engineering technology, key indicators representing failure modes as "hidden defect characteristics" are mined from the constructed graph neural network structure. A comprehensive analysis of the extracted key indicators is then performed to quantify the risk credibility of the "hidden defect characteristics";
[0010] The key indicators extracted and quantitatively analyzed through feature engineering are constructed into structured feature vectors and input into a pre-trained machine learning model. The model then performs intelligent identification and classification reasoning on the fault mode of the current sample to determine whether it belongs to a fault type with hidden defect characteristics.
[0011] When the fault mode is identified as a latent defect feature, the system enters the local amplification mode, adjusts the feature attention intensity mapping amount in the data space in real time, and implements "local amplification" of the original low-intensity signal in the feature space, so that the latent defect mode feature is more clearly displayed and avoids it being overwhelmed by pseudo-normal clusters.
[0012] Preferably, building a data collection mechanism covering the entire process of vehicle design, manufacturing, operation and after-sales maintenance includes the following steps:
[0013] First, the stages are divided and data sources are identified to clarify the key systems and accessible data sources involved in each stage of the vehicle life cycle.
[0014] Next, the collection interface and protocol are adapted. For different data types and sources, adaptive data collection interfaces and communication protocols are designed to achieve standardized collection of structured and unstructured data.
[0015] Then, multimodal data analysis and unified modeling are used to pre-process multi-source heterogeneous data and construct a multimodal data set in a unified format;
[0016] Finally, cross-stage data fusion and closed-loop management connect data from different life cycle stages, realize full-process quality data association through unified ID and time link, and deploy data synchronization mechanism and incremental update strategy to realize centralized management and dynamic fusion of multimodal quality data, providing a high-completeness and high-consistency input data foundation for subsequent quality risk modeling.
[0017] Preferably, key indicators characterizing the fault mode as "hidden defect characteristics" are mined from the constructed graph neural network structure through feature engineering technology. The extracted indicators include the measurement value of the degree of overlap of the feature distribution in the fault class and the normal class samples and the degree of difference in the characteristics of the adjacent nodes in the graph neural network structure. The measurement value of the degree of overlap of the feature distribution in the fault class and the normal class samples and the degree of difference in the characteristics of the adjacent nodes in the graph neural network structure are comprehensively analyzed under the detection window to generate feature fuzziness reference values and local graph heterogeneity reference values respectively. The risk credibility of the "hidden defect characteristics" is quantified by the feature fuzziness reference values and the local graph heterogeneity reference values.
[0018] Preferably, the specific steps of comprehensively analyzing the measurement values of the degree of overlap of the distribution of features in the fault class and the normal class samples under the detection window to generate the feature fuzziness reference value are as follows:
[0019] Within the detection window, the probability density distribution functions of the target features in the fault sample set and the normal sample set are extracted respectively. The minimum density superposition method is used to construct the fuzzy overlapping area, and the fuzzy kernel function is introduced for weighted calculation to form the fuzzy kernel response integral expression, which is as follows:
[0020] , where is the fuzzy overlapping response value, representing the characteristics In the current detection window, the total amount of fuzzy response in the overlapping area between normal samples and fault samples, is the characteristic probability density function of the fault class samples, is the characteristic probability density function of the normal class sample, is the fuzzy response weight function, which applies weighted function to the overlap density of different segments in the integral, highlighting the core segment of the fuzzy area. is the set of real numbers;
[0021] Blurring overlapping response values Normalization is performed through a nonlinear function to generate a characteristic ambiguity reference value. The generation formula is as follows:
[0022] , where is the ambiguity response coefficient, is the hyperbolic tangent function, is the characteristic ambiguity reference value.
[0023] Preferably, the specific steps of comprehensively analyzing the degree of difference in features of adjacent nodes of a node in a graph neural network structure under a detection window to generate a local graph heterogeneity reference value are as follows:
[0024] The feature similarity tensor is constructed based on the feature difference relationship between the target node and its adjacent nodes. Let the target node be , its adjacent node set is , through With each adjacent node The feature vector differences between them are nonlinearly mapped to generate nonlinear feature similarity values between node pairs. The generation formula is as follows:
[0025] , where is the node similarity value, is the similarity decay coefficient, is a generalized distance power function, representing the node and The distance measure between the feature vectors of and Node With node The characteristic vector of is the power exponent, is the norm parameter, is an exponential decay function, , is the index variable;
[0026] Based on the obtained node similarity value , further quantify the degree of local heterogeneity of the target node and generate a reference value of local graph heterogeneity. The generation formula is as follows:
[0027] , where is the reference value of local graph heterogeneity, is the number of adjacent nodes, is the nonlinear exponential amplification factor, is the numerical stability control factor.
[0028] Preferably, the feature fuzziness reference value and local graph heterogeneity reference value extracted and quantitatively analyzed by feature engineering are constructed into a structured feature vector and input into a pre-trained machine learning model. The latent defect risk coefficient is generated by the model, and the failure mode of the current sample is intelligently identified and classified based on the latent defect risk coefficient to determine whether it belongs to a fault type with latent defect characteristics.
[0029] Preferably, the latent defect risk coefficient generated by intelligently identifying and classifying the fault mode of the current sample through a pre-trained machine learning model is compared with a pre-set reference threshold value of the latent defect risk coefficient to determine whether it belongs to a fault type with latent defect characteristics. The judgment logic is as follows:
[0030] If the hidden defect risk coefficient is greater than the pre-set hidden defect risk coefficient reference threshold, the fault model is judged to be a hidden defect feature; if the hidden defect risk coefficient is less than or equal to the pre-set hidden defect risk coefficient reference threshold, the fault model is judged not to be a hidden defect feature.
[0031] Preferably, when the fault mode is identified as a hidden defect feature, the local amplification mode is entered, and the feature focus intensity mapping amount in the data space is adjusted in real time. The specific steps of implementing "local amplification" of the original low-intensity signal in the feature space are as follows:
[0032] When the hidden defect risk factor When the value is greater than the preset hidden defect risk coefficient reference threshold, the fault mode is determined to have hidden defect characteristics, triggering the local amplification process. The feature attention intensity mapping amount is calculated based on the difference measure between the key node features identified in the graph neural network structure and the features of the node and its adjacent nodes. The calculation expression of the feature attention intensity mapping amount is as follows:
[0033] , where is a node The local feature differences, is a node The local feature differences, is the number of nodes in the local feature space, It is a sensitive factor. is the feature attention intensity mapping amount;
[0034] For the generated feature attention intensity mapping, the original eigenvalue of the node is weighted and amplified, so that the originally low-amplitude implicit signal is enhanced in the feature space. The calculation formula is as follows:
[0035] , where It is the amplified node feature value, which means after the feature attention strength weighting processing, the The new eigenvalues of nodes, is the original node eigenvalue, indicating the The original eigenvalues of the nodes without any enhancement processing, is the average value of the feature attention intensity mapping of the local node group, is the feature enhancement factor;
[0036] To further confirm the effect of the amplification process, the local anomaly significance index of the nodes with potential hidden defect characteristics is calculated based on the degree of deviation between the amplified characteristic value and the normal node baseline characteristics. The calculation formula is as follows:
[0037] , where is the local anomaly significance indicator, is the normal node characteristic benchmark value, is the significant enhancement factor, yes function.
[0038] The automotive quality risk assessment system based on new data fusion technology includes a multimodal quality data acquisition and fusion module, a graph structure modeling and semantic association analysis module, a key risk feature extraction and credibility quantification module, a fault mode intelligent identification and classification decision module, and an abnormal feature local amplification and attention enhancement module:
[0039] The multimodal quality data collection and fusion module builds a data collection mechanism covering the entire process of vehicle design, manufacturing, operation, and after-sales maintenance, establishing a cross-stage and cross-modal comprehensive data fusion system for comprehensive collection and management of multimodal feature information, including structured and unstructured data;
[0040] The graph structure modeling and semantic association analysis module preprocesses the original multimodal data and then combines the semantic associations, temporal alignment relationships, and structural coupling characteristics between the features of each modality to construct a graph neural network structure for implicit association modeling.
[0041] The key risk feature extraction and credibility quantification module uses feature engineering technology to mine key indicators that represent failure modes as "hidden defect features" from the constructed graph neural network structure. It then conducts a comprehensive analysis of the extracted key indicators to quantify the risk credibility of the "hidden defect features";
[0042] The fault mode intelligent identification and classification decision module constructs key indicators extracted through feature engineering and quantitative analysis into structured feature vectors and inputs them into a pre-trained machine learning model. The model then performs intelligent identification and classification reasoning on the fault mode of the current sample to determine whether it belongs to a fault type with hidden defect characteristics;
[0043] The abnormal feature local amplification and attention enhancement module enters the local amplification mode when the fault mode is identified as a hidden defect feature, and adjusts the feature attention intensity mapping amount in the data space in real time, and implements "local amplification" of the original low-intensity signal in the feature space, making the hidden defect mode feature clearer and avoiding it being overwhelmed by pseudo-normal clusters.
[0044] In the above technical solution, the technical effects and advantages provided by the present invention are:
[0045] The present invention integrates structured and unstructured multimodal data to construct a cross-stage graph neural network structure, making full use of the semantic, temporal and structural associations between various types of data to achieve fine modeling of complex quality risk patterns. At the same time, key indicators such as feature fuzziness and graph heterogeneity are introduced to effectively quantify the risk credibility of hidden defects, and automatically enter the local amplification mode after identification to enhance the expression ability of abnormal features in the model and avoid them being overwhelmed by normal sample clusters. Overall, the solution realizes closed-loop control of the entire process from perception-modeling-evaluation-enhancement, significantly improving the sensitivity, accuracy and foresight of the quality assessment system, and effectively solving key problems in the existing technology such as low recognition rate of hidden quality defects, high misjudgment rate, and delayed risk exposure. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction to the drawings required for use in the embodiments will be given below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0047] Figure 1 This is a flow chart of the automobile quality risk assessment method based on the novel data fusion technology of the present invention.
[0048] Figure 2 This is a module diagram of the automobile quality risk assessment system based on the novel data fusion technology of the present invention. DETAILED DESCRIPTION
[0049] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these example embodiments are provided so that the description of this disclosure will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art.
[0050] The present invention provides Figure 1 The automobile quality risk assessment method based on the novel data fusion technology shown includes the following steps:
[0051] By building a data collection mechanism covering the entire process of vehicle design, manufacturing, operation, and after-sales maintenance, a cross-stage, cross-modal integrated data fusion system will be established to comprehensively collect and manage multimodal feature information, including structured data (such as sensor parameters and control signals) and unstructured data (such as user feedback text and maintenance images);
[0052] The types of data collected include, but are not limited to, manufacturing process control parameters (such as welding temperature and assembly torque), real-time data from onboard sensors (such as speed, voltage, temperature, and acceleration), historical fault codes (DTCs), user feedback text, and maintenance and repair information. This mechanism, through unified data interface protocols (such as CAN, UDS, and OTA platforms), edge gateway collectors, and vehicle-cloud synchronization, ensures the temporal consistency and content comparability of heterogeneous data, laying the information foundation for subsequent multidimensional modeling. This ensures the breadth and coverage of data sources, enabling full-process tracking of vehicle operating status and quality events, and serves as the fundamental data supply system for subsequent identification of hidden defects.
[0053] By deploying data collection nodes and interfaces throughout the vehicle lifecycle (including design, manufacturing, operation, and after-sales), a system capable of integrating data from diverse sources and formats is constructed, enabling comprehensive collection and management of all forms of vehicle quality-related information. This data includes not only structured data (such as sensor readings of temperature, voltage, speed, and other sensors, and output signals from electronic control units), but also unstructured data (such as user text feedback in the app, images of faults, and voice recordings of repair technicians). This cross-stage, cross-modal data fusion system not only breaks the limitations of traditional "island-style" data management but also provides multi-dimensional, comprehensive data support for subsequent intelligent analysis and risk identification.
[0054] Its role is to build a complete data portrait foundation for the identification of hidden defect patterns, improve the tracking ability of complex quality problems, feature extraction accuracy and risk prediction reliability, and enable the intelligent risk assessment system to have the ability to accurately respond to "weak signals", "gradual trends" and "potential hidden dangers", thereby realizing the transformation of quality management from passive response to active early warning.
[0055] Constructing a data collection mechanism covering the entire process of vehicle design, manufacturing, operation and after-sales maintenance includes the following steps: First, stage division and data source identification, clarifying the key systems and accessible data sources involved in each stage of the vehicle's life cycle (design, manufacturing, operation, maintenance), such as process parameters in the manufacturing process, CAN bus data in the operation stage, maintenance records and user feedback in the after-sales stage, etc.; Then, the acquisition interface and protocol adaptation, for different data types and sources, designing adapted data acquisition interfaces (such as OBD, T-Box, industrial gateway) and communication protocols (such as UDS, MQTT, HTTP), and implementing the structure Standardized collection of structured and unstructured data; then, multimodal data analysis and unified modeling, pre-processing multi-source heterogeneous data through timestamp alignment, data cleaning, semantic mapping and other means, and constructing multimodal data sets in a unified format (such as tensors, feature tables, graph structures); finally, cross-stage data fusion and closed-loop management, connecting data from different life cycle stages, realizing full-process quality data association through unified ID and time link, and deploying data synchronization mechanism and incremental update strategy to realize centralized management and dynamic fusion of multimodal quality data, providing a high-completeness and high-consistency input data foundation for subsequent quality risk modeling.
[0056] After preprocessing the original multimodal data, we build a graph neural network structure for implicit association modeling by combining the semantic associations, temporal alignment, and structural coupling characteristics between the features of each modality.
[0057] Raw collected data often suffers from inconsistent formats, time series desynchronization, missing values, and noise. Therefore, preprocessing operations such as cleaning, normalization, time series alignment, dimension completion, and anomaly removal are required. This preprocessed data is then structured using a unified modeling approach to construct a graph neural network architecture to support downstream feature extraction and evaluation algorithms. This improves data quality and analytical effectiveness, avoids feature distortion and recognition errors caused by data heterogeneity during subsequent analysis, and ensures the stability and effectiveness of model inputs.
[0058] After completing the preprocessing of the original multimodal data, the system does not immediately enter the model training phase. Instead, it first constructs a graph structure framework to describe the intrinsic connections between multimodal data based on the semantic relevance between each modal data (such as the similarity between fault codes and user feedback content), temporal alignment relationships (such as the synchronous appearance of sensor events and text descriptions), and structural coupling characteristics (such as the functional linkage between brake system temperature and brake pad wear images). In this structure, the features of different modalities are abstracted as "nodes", and the associations between nodes (such as synchronous occurrence, causal dependencies, or structural bindings) are defined as "edges", thus forming a graph neural network structure. This structure itself is not a trainable model, but provides a basic information propagation topology for subsequent graph neural network models. Its role is to establish an expression basis for the implicit associations between multimodal features, so that subsequent models can perform information propagation, feature aggregation, and relationship modeling in this structure, thereby effectively identifying hidden defect association patterns that are ignored by traditional linear or sequential models.
[0059] Through feature engineering technology, key indicators representing failure modes as "hidden defect characteristics" are mined from the constructed graph neural network structure. A comprehensive analysis of the extracted key indicators is then performed to quantify the risk credibility of the "hidden defect characteristics";
[0060] Through feature engineering technology, key indicators representing the fault mode as "hidden defect characteristics" are mined from the constructed graph neural network structure. The extracted indicators include the measurement value of the degree of feature distribution overlap between the fault class and the normal class samples and the degree of feature difference between the adjacent nodes in the graph neural network structure. The measurement value of the degree of feature distribution overlap between the fault class and the normal class samples and the degree of feature difference between the adjacent nodes in the graph neural network structure are comprehensively analyzed under the detection window to generate feature fuzziness reference value and local graph heterogeneity reference value respectively. The risk credibility of the "hidden defect characteristics" is quantified by the feature fuzziness reference value and the local graph heterogeneity reference value.
[0061] The risk credibility of "hidden defect features" is jointly quantified by the feature fuzziness reference value and the local graph heterogeneity reference value. Its core role is to establish a multi-angle discrimination mechanism from the two dimensions of feature distribution and structural context, and to enhance the system's sensitive recognition capability and classification accuracy for weak signal faults. Specifically, the feature fuzziness reference value is used to measure whether there is significant overlap between a feature in the fault class and the normal class samples. A high degree of overlap indicates that the feature has a low contribution to the model classification task, but it may hide the fault trend; while the local graph heterogeneity reference value reflects the degree of deviation between the node and its adjacent nodes in terms of features. The greater the deviation, the more likely it is a potential abnormal node. The joint analysis of these two types of reference values within the detection window can "amplify" hidden features that are difficult to identify with a single modality and assign higher risk weights, thereby achieving a forward-looking quantitative assessment of quality risks and providing accurate risk credibility for subsequent intelligent early warning and intervention mechanisms.
[0062] An increase in the overlap between the distribution of features in faulty and normal samples generally indicates that the fault pattern may be a "hidden defect feature." From the perspective of pattern recognition and feature discrimination, when the numerical distribution of a feature in faulty and normal samples is highly similar, i.e., there is a significant overlap in probability density, this indicates that the feature cannot effectively distinguish between faulty and normal states. Although these features contribute little to traditional classification tasks, in complex industrial scenarios, they may harbor potential fault trends characterized by "weak signals," "gradual behavior," and "low interference." This is particularly true in early stages of evolution, when their performance has not yet deviated from the "normal statistical fluctuation range," leading traditional models to misclassify them as non-risk data. This "fault feature masked by normal conditions" is one of the core characteristics of hidden defects. Therefore, an increase in distribution overlap does not mean that the feature is useless, but rather indicates that it has the risk potential of "weak representation and deep concealment," requiring further enhancement and identification in conjunction with other structural indicators (such as neighborhood heterogeneity). This feature ambiguity is precisely what intelligent quality management systems need to focus on identifying and modeling.
[0063] The specific steps for comprehensively analyzing the distribution overlap of features in fault class and normal class samples under the detection window to generate feature fuzziness reference values are as follows:
[0064] Within the detection window, the probability density distribution functions of the target features in the fault sample set and the normal sample set are extracted respectively. The minimum density superposition method is used to construct the fuzzy overlapping area, and the fuzzy kernel function is introduced for weighted calculation to form the fuzzy kernel response integral expression, which is as follows:
[0065] , where is the fuzzy overlapping response value, representing the characteristics In the current detection window, the total amount of fuzzy response in the overlapping area of the distribution of normal samples and fault samples. The larger the value, the more similar the distribution of the feature in the two types of samples, that is, the lower the discrimination and the higher the fuzziness, the more likely it is to become a hidden defect feature. is the characteristic probability density function of the fault class sample, for the target feature In the fault class sample set The probability density function constructed in It is the characteristic probability density function of the normal class sample, and the target feature In the normal class sample set The probability density function constructed in , the characteristic probability density function selects the corresponding probability density function according to the actual distribution of the data; for example, when the data presents a symmetrical distribution, the Gaussian distribution can be used; if the data is asymmetric and presents a long-tail distribution, the gamma distribution or exponential distribution may be more appropriate, It is a fuzzy response weight function, which applies weighted function to the overlap density of different segments in the integral, highlighting the core segment of the fuzzy area, and is used to strengthen the dominance of the intermediate fuzzy area on the integral result and suppress the misjudgment caused by micro-overlap of the boundary. is a set of real numbers, representing the integral variable The domain of definition, that is, the integral is performed over the entire range of real numbers, represents Can take any real value;
[0066] By constructing the probability density distribution of faulty and normal samples along the feature dimension and performing a weighted integral on their overlapping areas, the fuzzy intersection of the feature in the two categories is quantified to determine whether it has the potential risk of "insufficient discrimination." This fuzzy kernel response value can accurately reveal whether the feature exhibits "fault characteristics hidden in the normal state," providing a critical numerical basis for subsequent latent defect identification and index modeling.
[0067] Blurring overlapping response values Normalization is performed through a nonlinear function to generate a characteristic ambiguity reference value. The generation formula is as follows:
[0068] , where It is the fuzzy response coefficient, which is used to adjust the sensitivity of the fuzzy response in the output curve. The value range is usually set to 0.5-2.0. It is a hyperbolic tangent function with the characteristics of smooth growth and saturation value suppression, which can effectively avoid the bias or amplification distortion of risk scores caused by extreme ambiguity values. It is the feature fuzziness reference value, and its range is [0, 1]. The larger its value is, the weaker the feature's ability to distinguish between faulty and normal samples is, and thus the more likely it is to correspond to a "hidden defect feature".
[0069] The original fuzzy overlap density is mapped to a standardized risk reference value through a nonlinear function, ensuring comparability and a unified scoring scale across different features. The introduction of the fuzzy response coefficient and hyperbolic tangent function further enhances sensitivity to fuzzy features, suppresses extreme value interference, and improves the stability and accuracy of latent defect identification.
[0070] A larger feature fuzziness reference value, generated by comprehensively analyzing the degree of overlap between the distribution of a feature in faulty and normal samples within the detection window, indicates a higher similarity in the probability distribution of the feature between the two classes and a lower discriminative power. This indicates that the feature is difficult to directly identify fault patterns using traditional discriminant criteria. This fuzziness is a typical manifestation of a "hidden defect feature." In other words, although the feature may not exhibit significant numerical anomalies, its graph structure or evolutionary trends may be associated with progressive or latent risks. Conversely, a lower feature fuzziness reference value indicates a strong distribution separation between normal and faulty samples, a stronger discriminative power, and a greater likelihood of corresponding to an "overt defect pattern." Therefore, a higher feature fuzziness reference value indicates a more "hidden" fault feature, worthy of further identification and amplification in conjunction with other contextual information. This serves as an important reference for identifying hidden quality issues in intelligent risk control systems.
[0071] When the feature differences between a node and its neighbors in a graph neural network (GNN) structure increase significantly, it often indicates that the node may have "hidden defect characteristics." This is because in GNNs, nodes typically represent vehicle feature samples from different modalities, while adjacency relationships express their semantic, temporal, or structural connections. If the feature differences between a node and its neighbors suddenly increase, it indicates that the node has become inconsistent with its "similar background samples" in the representational dimension, indicating that it exhibits an anomalous trend that differs from the environmental context. For hidden defects, such differences often do not manifest as sudden failures, but rather accumulate gradually through slow deviations from normal patterns and gradual evolution. Because their overall magnitude is small and does not directly trigger system alarms, they are easily overlooked by traditional monitoring methods. Therefore, observing increased local heterogeneity in the graph structure essentially reflects a node's gradual "departure" from its normal cluster behavior pattern. It is an important clue to identifying potential degradation, the evolution of hidden faults, or early micro-anomalies, and has significant risk warning value in revealing hidden defects.
[0072] The specific steps for comprehensively analyzing the difference in feature levels of adjacent nodes in the graph neural network structure within the detection window to generate a local graph heterogeneity reference value are as follows:
[0073] The feature similarity tensor is constructed based on the feature difference relationship between the target node and its adjacent nodes. Let the target node be , its adjacent node set is , by targeting the node With each adjacent node The feature vector differences between them are nonlinearly mapped to generate nonlinear feature similarity values between node pairs. The generation formula is as follows:
[0074] , where is the node similarity value, indicating the target node With an adjacent node The feature similarity between , the smaller the value, the greater the difference. is the similarity attenuation coefficient, which controls the effect of feature differences on the final similarity value The decay rate of the feature difference is greater, the stronger the suppression of similarity is, that is, the similarity is more sensitive; the smaller the value is, the weaker the suppression is, and the similarity is more tolerant. is a generalized distance power function, representing the node and The distance measure between feature vectors, and Node With node The characteristic vector of Is a power index, used to amplify or compress the nonlinear change trend of the distance value (such as To magnify the difference, for compressing differences), improving modeling sensitivity to weak differences, is the norm parameter, which controls the distance metric’s paradigm (e.g. represents the Manhattan distance, represents the Euclidean distance, Further strengthen high-dimensional discreteness), by changing The value adjusts the model's response to the "sharpness" of the feature distribution. is an exponential decay function that maps the distance value to The similarity score within the range is calculated. This function can perform nonlinear compression on the feature difference value, so that similar nodes have higher similarity, and nodes with slightly larger differences quickly decrease to low similarity. , Is an index variable, indicating traversal of the collection (ie with the node directly connected adjacent nodes);
[0075] The above steps construct a nonlinear feature similarity relationship between nodes to identify subtle differences in the feature space between the target node and its adjacent nodes. This similarity metric provides a highly sensitive and structured input basis for the subsequent calculation of local graph heterogeneity reference values, enhancing the ability to identify hidden defect features.
[0076] Based on the obtained node similarity value , further quantify the degree of local heterogeneity of the target node and generate a reference value of local graph heterogeneity. The generation formula is as follows:
[0077] , where is the reference value of local graph heterogeneity, is the number of adjacent nodes, is the node The number of directly adjacent nodes indicates its “local connectivity” in the graph structure. It is a nonlinear exponential amplification factor, which adjusts the sensitivity to feature similarity and has a value range of 1-5. Is a numerical stability control factor to prevent the node similarity value Too small will cause the value in the logarithmic function to underflow, the value range .
[0078] By calculating the degree of feature difference between the target node and its adjacent nodes, a local graph heterogeneity reference value is generated, quantifying the local heterogeneity of the target node within the graph neural network structure. This process can effectively reveal potential subtle abnormal patterns, especially hidden defect features, and help identify nodes that have significant feature differences from adjacent nodes, thereby determining whether they pose risks.
[0079] The larger the local graph heterogeneity reference value, generated by comprehensively analyzing the degree of feature differences between a node and its neighbors within a graph neural network structure within the detection window, the more likely the data sample corresponding to that node is to harbor a hidden defect. This is because the local graph heterogeneity reference value reflects the degree of difference in feature space between a node and its neighbors. In a graph neural network structure, normal nodes often exhibit highly consistent features with their neighbors (i.e., structural cohesion). However, hidden defect samples, due to their atypical behavior, small anomalies, or slow trends, may exhibit potential microscopic deviations from their "similar" neighbors, manifesting as structural disjunctions. This difference is difficult to capture significantly using conventional algorithms, but can be amplified and identified through heterogeneity modeling within the graph structure. Therefore, a high local graph heterogeneity reference value indicates that the node's "local consistency" within the graph is compromised, potentially indicating an early stage of fault evolution, a key characteristic of a hidden defect. Conversely, a low heterogeneity reference value indicates strong feature consistency with its neighbors, making it more likely to be a normal sample or a typical anomaly rather than a hidden fault.
[0080] The key indicators extracted and quantitatively analyzed through feature engineering are constructed into structured feature vectors and input into a pre-trained machine learning model. The model then performs intelligent identification and classification reasoning on the fault mode of the current sample to determine whether it belongs to a fault type with hidden defect characteristics.
[0081] The feature fuzziness reference value and local graph heterogeneity reference value extracted and quantitatively analyzed by feature engineering are constructed into a structured feature vector and input into a pre-trained machine learning model. The latent defect risk coefficient is generated by the model, and the failure mode of the current sample is intelligently identified and classified based on the latent defect risk coefficient to determine whether it belongs to a fault type with latent defect characteristics.
[0082] A "pre-trained machine learning model" refers to an intelligent recognition model that is built and trained before actual operation based on a large amount of historical vehicle operation data, manufacturing data, and samples with known fault labels. This model can utilize a variety of machine learning algorithms, such as support vector machines (SVMs), random forests (RFs), gradient boosting trees (XGBoost), and neural networks (such as MLPs, CNNs, and GNNs). The choice of algorithm depends on the data type, feature structure, and the model's ability to identify latent defects. During the training phase, the system uses a large number of labeled samples (both normal and faulty) as input and iteratively optimizes model parameters using loss functions (such as cross-entropy loss and focal loss) to learn the mapping relationship between various features and target labels. During the training process, the model may also incorporate cross-validation, regularization, and sample reweighting to improve its ability to learn sparsely distributed and nonlinear features, ensuring that the final model exhibits good generalization performance and robustness to complex samples.
[0083] The greatest advantage of this "pre-trained machine learning model" is that, once trained and deployed in the system, it can quickly and efficiently perform intelligent recognition and classification inference on newly input samples. Specifically, in this system, the input structured feature vector includes two core quantitative metrics: a feature fuzziness reference value and a local graph heterogeneity reference value. Together, they describe the "degree of abnormality" of the current sample in both feature dimensions and structural space. When the model receives this vector, it uses its internally learned parameter weights, decision boundaries, and classification rules to generate an output value, the "hidden defect risk coefficient," to assess the current sample's risk status. This coefficient can be considered a quantitative basis for determining whether the current sample belongs to a hidden fault mode. A higher value indicates a greater similarity in features with existing hidden defect samples, prompting the system to trigger subsequent warning mechanisms, initiate feature amplification processing, or output repair recommendations. Therefore, this machine learning model serves as the core "discriminator" and "intelligent classification engine" in the entire hidden defect identification system, and is a key algorithmic component for achieving automated, intelligent, and precise fault identification.
[0084] The machine learning model is not limited here and can achieve the feature fuzziness reference value and local map heterogeneity reference value Conduct comprehensive analysis to generate hidden defect risk factors The machine learning model can be used. In order to implement the technical solution of the present invention, the present invention provides a specific implementation method:
[0085] Hidden defect risk factor The generation formula is as follows: , where and are the characteristic ambiguity reference values and local map heterogeneity reference value The preset scaling factor of and Both are greater than 0.
[0086] "Preset proportional coefficient" refers to the risk coefficient of generating hidden defects When , the characteristic ambiguity reference value is assigned and local map heterogeneity reference value The weighted coefficients of and These two coefficients essentially represent the system's prior setting of the proportion of these two types of feature dimensions in the hidden defect risk judgment, that is: when generating the final risk assessment value, the feature fuzziness reference value and local map heterogeneity reference value Which one is more worthy of "trust" or "emphasis". By artificially setting these two positive weights (and and are greater than 0), the two types of indicators can be adjusted to adjust the final risk coefficient The impact strength of the risk scoring mechanism can be adjusted flexibly and the strategy can be optimized. For example, if the graph structure features in a certain application scenario can better reflect the risk than the fuzzy features, it can be set , to highlight the risk factor of hidden defects This approach ensures that the generated hidden defect risk factor Adjustable and adaptable to different environments and model configurations.
[0087] It can be seen from the hidden defect risk coefficient that the larger the feature fuzziness reference value generated after comprehensive analysis of the measurement value of the degree of feature distribution overlap in the fault class and normal class samples under the detection window, and the larger the local graph heterogeneity reference value generated after comprehensive analysis of the degree of difference in features of adjacent nodes in the graph neural network structure under the detection window, the larger the hidden defect risk coefficient generated when the pre-trained machine learning model is used to intelligently identify and classify the fault mode of the current sample, indicating that the probability that the fault model is a hidden defect feature is greater, and vice versa, the smaller the probability that the fault model is a hidden defect feature.
[0088] The hidden defect risk coefficient generated by intelligently identifying and classifying the fault mode of the current sample using a pre-trained machine learning model is compared with the pre-set hidden defect risk coefficient reference threshold to determine whether it belongs to a fault type with hidden defect characteristics. The judgment logic is as follows:
[0089] If the hidden defect risk coefficient is greater than the pre-set hidden defect risk coefficient reference threshold, the fault model is judged to be a hidden defect feature; if the hidden defect risk coefficient is less than or equal to the pre-set hidden defect risk coefficient reference threshold, the fault model is judged not to be a hidden defect feature.
[0090] When the fault mode is identified as a hidden defect feature, it enters the local amplification mode and adjusts the feature focus intensity mapping in the data space in real time. The originally low-intensity signal is "locally amplified" in the feature space, making the hidden defect mode feature clearer and preventing it from being overwhelmed by pseudo-normal clusters.
[0091] When a fault pattern is identified as harboring latent defects, the system enters "local amplification mode." This step explicitly enhances abnormal feature signals that appear low-intensity and low-significance in the feature space, effectively overcoming the feature masking, blurred anomaly boundaries, and pseudo-normal clustering that are common in traditional models in high-dimensional space. In graph neural network architectures, weak signals, often due to their small amplitude, slow change trends, and low statistical significance, are easily surrounded and "absorbed" by normal sample clusters in the mainstream feature space. Consequently, they can be misidentified as normal during model training or inference, leading to serious risk omissions.
[0092] Through the "local amplification mode," the system dynamically adjusts the feature attention intensity mapping based on the fuzziness index and graph heterogeneity index results extracted earlier. This means that by increasing the attention weight, activation function response amplitude, or feature channel priority of the corresponding risk feature in the attention mechanism, the expression of such low-intensity signals in the feature space becomes clearer, the boundaries become more discernible, and their influence on the model classifier is enhanced. This processing method not only improves the model's sensitivity to subtle anomalies, but also effectively expands the model's discrimination boundaries, enabling it to "actively discover" rather than "passively misjudge" hidden risks, thereby significantly improving the accuracy, robustness, and early warning capabilities of the entire quality risk assessment system. This system has a high degree of engineering practicality and innovative value.
[0093] When the fault mode is identified as a hidden defect feature, the system enters the local amplification mode and adjusts the feature intensity mapping amount in the data space in real time. The specific steps for implementing "local amplification" of the originally low-intensity signal in the feature space are as follows:
[0094] When the hidden defect risk factor When the value is greater than the preset hidden defect risk coefficient reference threshold, the fault mode is determined to have hidden defect characteristics, triggering the local amplification process. The feature attention intensity mapping amount is calculated based on the difference measure between the key node features identified in the graph neural network structure and the features of the node and its adjacent nodes. The calculation expression of the feature attention intensity mapping amount is as follows:
[0095] , where is a node The local feature difference of the node The feature difference between the node and its adjacent nodes measures the node The degree of abnormality in the local feature space. If the feature difference between a node and its adjacent nodes is large, it means that the abnormal feature of the node may be more obvious and needs to be given more attention in subsequent processing. is a node The local feature difference of the node The feature difference between the node and its adjacent nodes, is the number of nodes in the local feature space, Is a sensitive factor used to control and (or ) determines the risk factor of hidden defects The degree of influence on the feature amplification strength, the value range is 0.1-10, is the feature attention intensity mapping quantity, representing the node Hidden defect risk factor in graph neural networks The calculated feature attention strength mapping reflects the node The intensity of attention of a feature relative to other nodes in multimodal data. A high value means that the node requires more attention and there may be a risk of hidden defects.
[0096] Through the above mechanism, a dynamic significance weight basis can be formed for potential abnormal nodes.
[0097] For the generated feature attention intensity mapping, the original eigenvalue of the node is weighted and amplified, so that the originally low-amplitude implicit signal is enhanced in the feature space. The calculation formula is as follows:
[0098] , where It is the amplified node feature value, which means after the feature attention strength weighting processing, the The new eigenvalues of nodes, is the original node eigenvalue, indicating the The original eigenvalues of the nodes without any enhancement processing, is the average value of the feature attention intensity mapping of the local node group, It is the feature enhancement factor, a hyperparameter used to control the enhancement amplitude in the entire amplification mechanism. Its value range is 0.1-5. The larger the value, the stronger the amplification effect.
[0099] Through this step, the features of nodes with weak characterization characteristics but significant potential risks can be enhanced to ensure that they are not obscured by normal sample clusters in subsequent processing flows.
[0100] To further confirm the effect of the amplification process, the local anomaly significance index of the nodes with potential hidden defect characteristics is calculated based on the degree of deviation between the amplified characteristic value and the normal node baseline characteristics. The calculation formula is as follows:
[0101] , where It is a local anomaly significance indicator, indicating the After the amplification process, the significance score between the characteristic value of each node and the reference value of the normal sample characteristic is The function is compressed between 0 and 1, It is the normal node feature reference value, which represents the feature reference value of the normal sample node in the current local graph. It is usually the statistical mean or median of the local neighborhood or global normal samples. is the significant enhancement factor, controlling the characteristic difference in The slope sensitivity of the mapping in the function adjusts the system's response strength to abnormal differences of different amplitudes, with a value range of 1-10. yes function.
[0102] This step can effectively verify whether the amplification process successfully creates a recognizable difference between the latent defect signal and the normal state characteristics, thereby improving the model recognition accuracy and system warning reliability.
[0103] The present invention integrates structured and unstructured multimodal data to construct a cross-stage graph neural network structure, making full use of the semantic, temporal and structural associations between various types of data to achieve fine modeling of complex quality risk patterns. At the same time, key indicators such as feature fuzziness and graph heterogeneity are introduced to effectively quantify the risk credibility of hidden defects, and automatically enter the local amplification mode after identification to enhance the expression ability of abnormal features in the model and avoid them being overwhelmed by normal sample clusters. Overall, the solution realizes closed-loop control of the entire process from perception-modeling-evaluation-enhancement, significantly improving the sensitivity, accuracy and foresight of the quality assessment system, and effectively solving key problems in the existing technology such as low recognition rate of hidden quality defects, high misjudgment rate, and delayed risk exposure.
[0104] The present invention provides Figure 2 The automobile quality risk assessment system based on the new data fusion technology shown in the figure includes a multimodal quality data acquisition and fusion module, a graph structure modeling and semantic association analysis module, a key risk feature extraction and credibility quantification module, a fault mode intelligent identification and classification decision module, and an abnormal feature local amplification and attention enhancement module:
[0105] The multimodal quality data collection and fusion module builds a data collection mechanism covering the entire process of vehicle design, manufacturing, operation, and after-sales maintenance, establishing a cross-stage and cross-modal comprehensive data fusion system for comprehensive collection and management of multimodal feature information, including structured and unstructured data;
[0106] The graph structure modeling and semantic association analysis module preprocesses the original multimodal data and then combines the semantic associations, temporal alignment relationships, and structural coupling characteristics between the features of each modality to construct a graph neural network structure for implicit association modeling.
[0107] The key risk feature extraction and credibility quantification module uses feature engineering technology to mine key indicators that represent failure modes as "hidden defect features" from the constructed graph neural network structure. It then conducts a comprehensive analysis of the extracted key indicators to quantify the risk credibility of the "hidden defect features";
[0108] The fault mode intelligent identification and classification decision module constructs key indicators extracted through feature engineering and quantitative analysis into structured feature vectors and inputs them into a pre-trained machine learning model. The model then performs intelligent identification and classification reasoning on the fault mode of the current sample to determine whether it belongs to a fault type with hidden defect characteristics;
[0109] The abnormal feature local amplification and attention enhancement module enters the local amplification mode when the fault mode is identified as a hidden defect feature, and adjusts the feature attention intensity mapping amount in the data space in real time, and implements "local amplification" of the original low-intensity signal in the feature space, making the hidden defect mode feature clearer and avoiding it being overwhelmed by pseudo-normal clusters.
[0110] The automobile quality risk assessment method based on the new data fusion technology provided in an embodiment of the present invention is implemented through the above-mentioned automobile quality risk assessment system based on the new data fusion technology. The specific methods and processes of the automobile quality risk assessment system based on the new data fusion technology are detailed in the embodiment of the automobile quality risk assessment method based on the new data fusion technology, which will not be repeated here.
[0111] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters in the formulas are set by technicians in this field according to actual conditions.
[0112] The above description is merely illustrative of certain exemplary embodiments of the present invention. It goes without saying that those skilled in the art will be able to modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the above drawings and description are illustrative in nature and should not be construed as limiting the scope of protection of the claims.
[0113] It should be noted that, in this document, if there are relational terms such as first and second, etc., they are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises", "comprising" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprising a ..." does not exclude the presence of other identical elements in the process, method, article or device that includes the element.
[0114] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0115] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0116] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0117] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0118] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0119] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0120] The above description is merely illustrative of certain exemplary embodiments of the present invention. It goes without saying that those skilled in the art will be able to modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the above drawings and description are illustrative in nature and should not be construed as limiting the scope of protection of the claims.
Claims
1. The automobile quality risk assessment method based on new data fusion technology is characterized by: The following steps are involved: By building a data collection mechanism covering the entire process of vehicle design, manufacturing, operation and after-sales maintenance, a cross-stage and cross-modal integrated data fusion system is established to comprehensively collect and manage multi-modal feature information including structured and unstructured data; After preprocessing the original multimodal data, we build a graph neural network structure for implicit association modeling by combining the semantic associations, temporal alignment, and structural coupling characteristics between the features of each modality. Through feature engineering technology, key indicators that characterize failure modes as hidden defect characteristics are mined from the constructed graph neural network structure. A comprehensive analysis of the extracted key indicators is then performed to quantify the risk credibility of the hidden defect characteristics. The key indicators extracted and quantitatively analyzed through feature engineering are constructed into structured feature vectors and input into a pre-trained machine learning model. The model then performs intelligent identification and classification reasoning on the fault mode of the current sample to determine whether it belongs to a fault type with hidden defect characteristics. When the fault mode is identified as a hidden defect feature, it enters the local amplification mode, adjusts the feature focus intensity mapping amount in the data space in real time, and implements local amplification of the originally low-intensity signal in the feature space to make the hidden defect mode feature clearly displayed; Through feature engineering technology, key indicators representing the failure mode as hidden defect characteristics are mined from the constructed graph neural network structure. The extracted indicators include the measurement value of the degree of feature distribution overlap between faulty and normal class samples and the degree of feature difference between adjacent nodes in the graph neural network structure. The measurement value of the degree of feature distribution overlap between faulty and normal class samples and the degree of feature difference between adjacent nodes in the graph neural network structure are comprehensively analyzed within the detection window to generate feature fuzziness reference value and local graph heterogeneity reference value respectively. The feature fuzziness reference value and local graph heterogeneity reference value are used to quantify the risk credibility of the hidden defect characteristics. The specific steps for comprehensively analyzing the distribution overlap of features in fault class and normal class samples under the detection window to generate feature fuzziness reference values are as follows: Within the detection window, the probability density distribution functions of the target features in the fault sample set and the normal sample set are extracted respectively. The minimum density superposition method is used to construct the fuzzy overlapping area, and the fuzzy kernel function is introduced for weighted calculation to form the fuzzy kernel response integral expression, which is as follows: , where is the fuzzy overlapping response value, representing the characteristics In the current detection window, the total amount of fuzzy response in the overlapping area between normal samples and fault samples, is the characteristic probability density function of the fault class samples, is the characteristic probability density function of the normal class sample, is the fuzzy response weight function, which applies weighted function to the overlap density of different segments in the integral, highlighting the core segment of the fuzzy area. is the set of real numbers; Blurring overlapping response values Normalization is performed through a nonlinear function to generate a characteristic ambiguity reference value. The generation formula is as follows: , where is the ambiguity response coefficient, is the hyperbolic tangent function, is the characteristic ambiguity reference value; The specific steps for comprehensively analyzing the difference in feature levels of adjacent nodes in the graph neural network structure within the detection window to generate a local graph heterogeneity reference value are as follows: The feature similarity tensor is constructed based on the feature difference relationship between the target node and its adjacent nodes. Let the target node be , its adjacent node set is , by targeting the node With each adjacent node The nonlinear feature similarity value between the node pairs is generated by nonlinear mapping of the feature vector differences between them. The generation formula is as follows: , where is the node similarity value, is the similarity decay coefficient, is a generalized distance power function, representing the node and The distance measure between the feature vectors of and Node With node The characteristic vector of is the power exponent, is the norm parameter, is an exponential decay function, , is the index variable; Based on the obtained node similarity value , further quantify the degree of local heterogeneity of the target node and generate a reference value of local graph heterogeneity. The generation formula is as follows: , where is the reference value of local graph heterogeneity, is the number of adjacent nodes, is the nonlinear exponential amplification factor, is the numerical stability control factor.
2. The automobile quality risk assessment method based on the novel data fusion technology according to claim 1 is characterized in that: Building a data collection mechanism covering the entire process of vehicle design, manufacturing, operation, and after-sales maintenance includes the following steps: Phase division and data source identification to clarify the key systems and accessible data sources involved in each stage of the vehicle life cycle; Adapt the acquisition interface and protocol to design adaptive data acquisition interfaces and communication protocols for different data types and sources, and realize the standardized acquisition of structured and unstructured data; Multimodal data analysis and unified modeling: pre-processing multi-source heterogeneous data and constructing multimodal data sets in a unified format; Cross-stage data fusion and closed-loop management connect data from different life cycle stages, realize full-process quality data association through unified ID and time link, and deploy data synchronization mechanism and incremental update strategy to realize centralized management and dynamic fusion of multimodal quality data.
3. The automobile quality risk assessment method based on the novel data fusion technology according to claim 1 is characterized in that: The feature fuzziness reference value and local graph heterogeneity reference value extracted and quantitatively analyzed by feature engineering are constructed into a structured feature vector and input into a pre-trained machine learning model. The latent defect risk coefficient is generated by the model, and the failure mode of the current sample is intelligently identified and classified based on the latent defect risk coefficient to determine whether it belongs to a fault type with latent defect characteristics.
4. The automobile quality risk assessment method based on the novel data fusion technology according to claim 3 is characterized in that: The hidden defect risk coefficient generated by intelligently identifying and classifying the fault mode of the current sample using a pre-trained machine learning model is compared with the pre-set hidden defect risk coefficient reference threshold to determine whether it belongs to a fault type with hidden defect characteristics. The judgment logic is as follows: If the hidden defect risk coefficient is greater than the pre-set hidden defect risk coefficient reference threshold, the fault model is judged to be a hidden defect feature; if the hidden defect risk coefficient is less than or equal to the pre-set hidden defect risk coefficient reference threshold, the fault model is judged not to be a hidden defect feature.
5. The automobile quality risk assessment method based on the novel data fusion technology according to claim 4 is characterized in that: When the fault mode is identified as a hidden defect feature, the system enters the local amplification mode and adjusts the feature intensity mapping amount in the data space in real time. The specific steps for locally amplifying the originally low-intensity signal in the feature space are as follows: When the hidden defect risk factor When the value is greater than the preset hidden defect risk coefficient reference threshold, the fault mode is determined to have hidden defect characteristics, triggering the local amplification process. The feature attention intensity mapping amount is calculated based on the difference measure between the key node features identified in the graph neural network structure and the features of the node and its adjacent nodes. The calculation expression of the feature attention intensity mapping amount is as follows: , where is a node The local feature differences, is a node The local feature differences, is the number of nodes in the local feature space, It is a sensitive factor. is the feature attention intensity mapping amount; For the generated feature attention intensity mapping, the original eigenvalue of the node is weighted and amplified, so that the originally low-amplitude implicit signal is enhanced in the feature space. The calculation formula is as follows: , where It is the amplified node feature value, which means after the feature attention strength weighting processing, the The new eigenvalues of nodes, is the original node eigenvalue, indicating the The original eigenvalues of the nodes without any enhancement processing, is the average value of the feature attention intensity mapping of the local node group, is the feature enhancement factor; Based on the degree of deviation between the amplified eigenvalue and the normal node benchmark characteristics, the local abnormal significance index of the node with potential hidden defect characteristics is calculated. The calculation formula is as follows: , where is the local anomaly significance indicator, is the normal node characteristic benchmark value, is the significant enhancement factor, yes function.
6. An automobile quality risk assessment system based on a novel data fusion technology, used to implement the automobile quality risk assessment method based on a novel data fusion technology as described in any one of claims 1 to 5, characterized in that: It includes multimodal quality data collection and fusion module, graph structure modeling and semantic association analysis module, key risk feature extraction and credibility quantification module, fault mode intelligent identification and classification decision module, and abnormal feature local amplification and attention enhancement module: The multimodal quality data collection and fusion module builds a data collection mechanism covering the entire process of vehicle design, manufacturing, operation, and after-sales maintenance, establishing a cross-stage and cross-modal comprehensive data fusion system for comprehensive collection and management of multimodal feature information, including structured and unstructured data; The graph structure modeling and semantic association analysis module preprocesses the original multimodal data and then combines the semantic associations, temporal alignment relationships, and structural coupling characteristics between the features of each modality to construct a graph neural network structure for implicit association modeling. The key risk feature extraction and credibility quantification module uses feature engineering technology to mine key indicators that characterize failure modes as hidden defect features from the constructed graph neural network structure, and conducts a comprehensive analysis of the extracted key indicators to quantify the risk credibility of the hidden defect features; The fault mode intelligent identification and classification decision module constructs key indicators extracted through feature engineering and quantitative analysis into structured feature vectors and inputs them into a pre-trained machine learning model. The model then performs intelligent identification and classification reasoning on the fault mode of the current sample to determine whether it belongs to a fault type with hidden defect characteristics; The abnormal feature local amplification and attention enhancement module enters the local amplification mode when the fault mode is identified as a hidden defect feature, adjusts the feature attention intensity mapping amount in the data space in real time, and locally amplifies the original low-intensity signal in the feature space to make the hidden defect mode characteristics clearly displayed.
Citation Information
Patent Citations
Automobile part production quality optimization method and system based on machine learning
CN117217627A
Industrial product quality detection method and system based on machine vision
CN119600032A