A multi-source heterogeneous vehicle model data automatic alignment and conflict resolution method

CN122220581BActive Publication Date: 2026-09-04BEIJING KUCHE YIMEI NETWORK TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202610371318.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-03-25
Publication Date
2026-09-04
Estimated Expiration
2046-03-25

AI Technical Summary

Technical Problem

但该方案聚焦于“数据标准定义”的冲突消解,未针对二手车领域车型数据的特殊性(如车型命名多样性、关键参数动态性、数据质量差异性)进行适配,难以直接解决车型数据在名称匹配、参数核验、误差容忍等方面的专项需求,从而降低了车型数据匹配的精度

Benefits of technology

[0084] (1) By comprehensively collecting three types of vehicle model data—text descriptions, vehicle images, and structured parameters—from multiple heterogeneous data sources, differentiated preprocessing and feature extraction methods were adopted for different types of data. This standardized the data format and captured core features, solving the problem of differences in format and naming rules among different data sources, and fully exploring the multi-dimensional value of the data. By unifying dimensions and fusing features to form a comprehensive feature set, the limitations of single feature extraction were broken, and high-value information such as vehicle component status and accident records in the private domain data of institutions such as Cha Doctor were activated. This laid a comprehensive feature foundation for subsequent data matching and alignment, broke down "data silos," improved data utilization, promoted the transformation of vehicle model data from fragmented to standardized, and provided solid feature support for data-driven decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122220581B_ABST
    Figure CN122220581B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of data conflict analysis, and particularly relates to a vehicle model data automatic alignment and conflict resolution method based on multi-source heterogeneity. The method comprises the following steps: obtaining vehicle model data from multiple heterogeneous data sources for multi-modal feature extraction, generating a comprehensive feature set containing text features, visual features and structured features; constructing a dynamic weight calculation model based on the comprehensive feature set and performing matching calculation and automatic alignment to form a preliminary aligned vehicle model data group; performing conflict detection on the preliminary aligned vehicle model data group, identifying conflict items with inconsistent data, resolving the conflict items, and generating standardized vehicle model data without conflicts; incorporating the standardized vehicle model data into a self-optimizing knowledge base to absorb artificial feedback information through a closed-loop learning mechanism, continuously optimizing the feature extraction model and the matching algorithm, and thus forming corresponding closed-loop learning records. The present application can improve the accuracy and adaptability of vehicle model data matching.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data conflict analysis technology, and in particular to a method for automatic alignment and conflict resolution of vehicle model data based on multi-source heterogeneity. Background Technology

[0002] In the fields of used car inspection and digitalization of the automotive aftermarket, vehicle model data serves as the core foundation supporting vehicle valuation, vehicle condition verification, and market analysis. Its completeness, consistency, and relevance directly determine the effectiveness of data-driven decision-making. Leading used car inspection agencies, such as Cha Doctor, have accumulated massive amounts of private domain inspection data over their long-term operations. This data includes high-value information such as vehicle component condition, paint thickness, and accident records. However, due to the lack of an effective alignment and integration mechanism with multi-source industry data, a large amount of private domain data is difficult to transform into productivity, failing to fully realize its collaborative value. Simultaneously, industry-level vehicle model data integration faces severe challenges due to the heterogeneity of multiple sources—vehicle model data from different data sources (such as inspection agencies, e-commerce platforms, and automaker databases) exhibit significant differences in format, naming rules, and field definitions, forming "data silos" that result in low data utilization and significantly diminished data value.

[0003] Furthermore, Chinese patent CN120653720A discloses a method for resolving conflicts in multiple data standard definitions based on large models and knowledge graphs. This invention relates to a method for resolving conflicts in multiple data standard definitions based on large models and knowledge graphs, belonging to the fields of public data governance and data standardization. This invention includes: extraction and preprocessing of multiple data standard definitions; semantic representation and large model parsing: using a large model to perform semantic analysis on the definition of each data item, generating semantic embedding vectors and key features; knowledge graph construction and entity alignment: constructing a knowledge graph containing all data items, and using an entity alignment algorithm to connect semantically similar data item nodes from different standards to form candidate alignment relationships; conflict detection and type identification: identifying definition conflict points and classifying and labeling conflict types based on the aligned data item definitions, comparative attributes, and values; and a globally optimized conflict resolution algorithm for conflict resolution and unified definition generation. This invention solves the problems of data silos and semantic conflicts caused by inconsistent data standard definitions in existing technologies. By combining large models and knowledge graphs, this invention solves the problems of data silos and semantic conflicts caused by inconsistent data standard definitions, providing a technical approach for multi-source data integration. However, this solution focuses on resolving conflicts in "data standard definition" and does not adapt to the special characteristics of vehicle model data in the used car field (such as the diversity of vehicle model names, the dynamic nature of key parameters, and the differences in data quality). It is difficult to directly solve the specific needs of vehicle model data in terms of name matching, parameter verification, and error tolerance, thereby reducing the accuracy of vehicle model data matching. Summary of the Invention

[0004] To address the aforementioned technical problems in existing vehicle model data matching processes, this invention provides a method for automatic alignment and conflict resolution of multi-source heterogeneous vehicle model data. This method leverages multimodal feature extraction to comprehensively capture multi-dimensional features of vehicle models, overcoming the limitations of single-feature extraction and providing comprehensive support for data matching. Dynamic weight matching and hierarchical clustering enable automatic alignment of multi-source data, resolving integration challenges caused by differences in data source formats and naming conventions, and activating the value of private domain data from institutions like ChaDoctor. A three-level conflict resolution process specifically addresses data inconsistency issues, generating standardized vehicle model data to compensate for the shortcomings of similar patents that do not adapt to the specific characteristics of used car model data. A self-optimizing knowledge base and closed-loop learning mechanism continuously absorb human feedback, optimizing models and algorithms to achieve iterative upgrades in data processing capabilities. This provides reliable support for data-driven decision-making such as vehicle valuation and vehicle condition verification, thereby improving data utilization and collaborative value, and promoting the digital upgrade of the industry. The method includes the following steps:

[0005] Acquire vehicle model data from multiple heterogeneous data sources, perform multimodal feature extraction on the vehicle model data, and generate a comprehensive feature set that includes textual features, visual features, and structured features;

[0006] A dynamic weight calculation model is constructed based on a comprehensive feature set, and the vehicle model data from different sources is matched and calculated based on the dynamic weight calculation model to obtain the vehicle model data matching similarity results.

[0007] Based on the similarity results of vehicle model data matching, a hierarchical clustering method is used to automatically align the vehicle model data to form initially aligned vehicle model data groups. Conflict detection is performed on the initially aligned vehicle model data groups to identify conflicting items with inconsistent data. A three-level processing flow is used to resolve the conflicting items and generate conflict-free standardized vehicle model data.

[0008] Standardized vehicle model data is incorporated into a self-optimizing knowledge base to absorb human feedback through a closed-loop learning mechanism, continuously optimizing the feature extraction model and matching algorithm, thereby forming corresponding closed-loop learning records.

[0009] This invention first acquires vehicle model data from multiple heterogeneous data sources. Through multimodal feature extraction, it comprehensively captures three core features: textual, visual, and structured features, generating a comprehensive feature set. This overcomes the limitations of single feature extraction, fully leveraging the value of different types of vehicle model data. It solves the problem of traditional data processing focusing only on single features and insufficient data information mining, laying a comprehensive feature foundation for subsequent data matching. Based on the comprehensive feature set, a dynamic weight calculation model is constructed to perform matching calculations on vehicle model data from different sources. Combined with hierarchical clustering, automatic alignment is achieved, effectively solving the integration difficulties caused by differences in format, naming rules, and field definitions among different data sources. This breaks down "data silos," allowing the massive private domain high-value data accumulated by institutions such as Cha Doctor to effectively connect with multi-source industry data, fully leveraging data synergy and improving data utilization. Conflict detection is performed on the initially aligned data, employing a three-level processing flow to resolve conflict items and generate conflict-free standardized vehicle model data. This is specifically adapted to the unique characteristics of the used car industry, such as diverse naming, dynamic parameters, and significant quality differences in vehicle model data. It compensates for the shortcomings of similar patents that only focus on data standard definition conflicts and cannot adapt to specific industry needs, ensuring the reliability of standardized data. By incorporating standardized data into a self-optimizing knowledge base and absorbing human feedback through a closed-loop learning mechanism, the feature extraction model and matching algorithm are continuously optimized. This enables iterative upgrades of data processing capabilities, ensuring that models and algorithms can continuously adapt to changes in industry data and guaranteeing the long-term effectiveness and applicability of vehicle model data processing. The entire solution achieves closed-loop management of vehicle model data from extraction, matching, alignment, conflict resolution to continuous optimization. It not only activates the value of private domain data but also solves the core challenge of multi-source data integration. This provides high-quality, standardized vehicle model data for vehicle valuation, vehicle condition verification, market analysis, and other businesses, improving the effectiveness of data-driven decision-making and promoting the high-quality development of the used car inspection and automotive aftermarket towards digitalization and intelligence.

[0010] Preferably, the process of acquiring vehicle model data from multiple heterogeneous data sources and extracting multimodal features from the vehicle model data includes the following steps:

[0011] Vehicle model data is collected from multiple heterogeneous data sources through a data interface. This vehicle model data includes text description information, vehicle image information, and structured parameter information.

[0012] The text description information is preprocessed, including removing redundant characters, standardizing the expression format, and converting aliases to standard terms to form standardized text data; a pre-trained text understanding model is used to extract features from the standardized text data, capturing the semantic information corresponding to the brand, model, configuration, and year contained in the text, and generating text feature vectors;

[0013] The vehicle image information is preprocessed, including size normalization, angle correction, and background removal, to obtain standardized vehicle images. A convolutional neural network is used to extract visual features from the standardized vehicle images, focusing on capturing the unique visual identifiers corresponding to the vehicle's exterior outline, grille shape, and wheel style, and generating visual feature vectors.

[0014] The structured parameter information is parsed to extract vehicle identification number, registration time, mileage, and displacement parameters. After format standardization, a structured feature vector is generated.

[0015] By unifying the dimensions and fusing the features of text feature vectors, visual feature vectors, and structured feature vectors, a comprehensive feature set containing multimodal information is formed.

[0016] This invention comprehensively collects vehicle model data in three categories: text, images, and structured data. Differentiated preprocessing and feature extraction methods are employed for each data type. Text information is formatted and semantic features are extracted; vehicle images are standardized and visual identifiers are captured; and structured parameters are parsed, extracted, and formatted uniformly, ensuring the integrity and standardization of all features. Through dimensional unification and feature fusion, a comprehensive feature set is formed, breaking the limitations of single feature extraction and fully exploring the multi-dimensional value of vehicle model data. This solves the feature confusion caused by differences in data source formats and naming, and activates high-value information from private domain data of institutions such as [Institution Name]. This step provides comprehensive and high-quality feature support for subsequent data matching and automatic alignment, improves the effectiveness of multi-source data integration, breaks down "data silos," and promotes the transformation of vehicle model data from fragmented to standardized and integrated, laying a solid foundation for data-driven decision-making.

[0017] Preferably, the step of constructing a dynamic weight calculation model based on a comprehensive feature set and performing matching calculations on vehicle model data from different sources based on the dynamic weight calculation model includes the following steps:

[0018] A dynamic weight calculation model is constructed, which can dynamically adjust the weight values ​​of each feature dimension based on the credibility, data integrity, and historical matching accuracy of different data sources.

[0019] For the text feature vectors in the comprehensive feature set, a semantic similarity algorithm is used to calculate the text matching degree between different vehicle models, focusing on the consistency of brand and model descriptions and the relevance of configuration descriptions;

[0020] For visual feature vectors, the cosine similarity algorithm of feature vectors is used to calculate the visual matching degree between data of different vehicle models, and the unique logos and design features of the vehicle appearance are compared first.

[0021] For structured feature vectors, different similarity calculation methods are used according to parameter type: string edit distance algorithm is used for vehicle identification code, relative error algorithm is used for numerical parameters, and time difference algorithm is used for time parameters.

[0022] The text matching degree, visual matching degree and structured feature matching degree are weighted and fused by a dynamic weight calculation model to calculate the comprehensive matching degree between vehicle models from different sources.

[0023] The overall matching degree is standardized to obtain vehicle model data matching similarity results with a uniform value range.

[0024] This invention constructs a dynamic weight calculation model that can dynamically adjust feature weights based on factors such as data source credibility and data integrity, overcoming the limitations of fixed weights and adapting to the quality differences of different data sources. A differentiated similarity calculation method is adopted for three types of features, aligning with the characteristics of text, visual, and structured parameters, focusing on core information such as brand model, core visual identifiers, and key parameters. This solves the matching difficulties caused by diverse vehicle naming and parameter differences, while ensuring the targeted and reasonable nature of the matching calculation. Through weighted fusion and standardization, a unified range of matching similarity results is obtained, achieving scientific matching of vehicle data from different sources, breaking down "data silos," promoting the effective integration of private domain data from institutions like ChaBoshi with multi-source industry data, improving data utilization and collaborative value, and providing a reliable matching basis for subsequent automatic alignment of vehicle data.

[0025] Preferably, the weighted fusion of text matching degree, visual matching degree, and structured feature matching degree through a dynamic weight calculation model includes the following steps:

[0026] Retrieve historical performance parameters of each data source from the self-optimizing knowledge base, including data accuracy, integrity score, and conflict rate indicators;

[0027] The credibility weight of each data source is calculated based on historical performance parameters. Features corresponding to data sources with high credibility receive higher weight in the matching calculation.

[0028] The basic weight coefficients are set according to the type characteristics of vehicle data, with vehicle identification code related features set as the highest basic weight, visual features and core configuration text features set as the second highest basic weight, and auxiliary parameter features set as the lowest basic weight.

[0029] By combining the credibility weight and the basic weight coefficient, a dynamic weight calculation model is used to generate real-time weight values ​​for each feature dimension.

[0030] The text matching degree, visual matching degree, and structured feature matching degree are multiplied by their respective real-time weight values ​​to obtain the weighted matching scores for each dimension.

[0031] The weighted matching scores of each dimension are summed to obtain the overall matching degree between vehicle models from different sources.

[0032] This invention retrieves historical performance parameters from a self-optimizing knowledge base to calculate credibility weights, fully considering the quality differences between different data sources. Features from high-quality data sources receive higher weights, ensuring the matching results more closely reflect actual data quality. Basic weights are set based on vehicle type data, highlighting the importance of key information such as vehicle identification numbers and core visual features. This adapts to the specific characteristics of used car vehicle data, overcoming the shortcomings of similar patents that fail to differentiate feature priorities and lack targeted matching. A dynamic weight model generates real-time weights, which are then weighted and fused to obtain a comprehensive matching score. This solves the matching bias problem caused by fixed weights while also considering data source quality and feature importance, ensuring the rationality and consistency of the matching results. This step makes data matching more targeted and flexible, adapting to the diverse naming, dynamic parameters, and significant quality differences in vehicle data. It improves the effectiveness of multi-source data matching, promotes accurate integration of private domain data and industry multi-source data, fully leverages the value of data collaboration, and provides more reliable support for subsequent automatic alignment of vehicle data.

[0033] Preferably, the automatic alignment of vehicle model data using a hierarchical clustering method based on the similarity results of the vehicle model data includes the following steps:

[0034] Set an initial clustering threshold, which is determined based on the best accuracy of historical matching data;

[0035] The vehicle model data matching similarity results are compared with the initial clustering threshold, and vehicle model data with similarity results higher than the threshold are marked as potential matching pairs;

[0036] Hierarchical clustering is used to cluster potential matching pairs. First, a first-level clustering is performed based on brand characteristics, grouping vehicle data from the same brand into one major category. Within each major brand category, a second-level clustering is performed based on vehicle series characteristics, grouping vehicle data belonging to the same series into one sub-cluster. Within each sub-cluster, a third-level clustering is performed based on the overall matching degree, grouping vehicle data with an overall matching degree reaching a preset threshold into the same data group.

[0037] Perform a consistency check on each data group to ensure that the vehicle model data within the group are basically consistent in terms of core features, and remove data items that are obviously mismatched;

[0038] The data groups that pass the consistency check are marked as initially aligned vehicle model data groups, and the source information and matching confidence of each data item in the group are recorded.

[0039] This invention determines the initial clustering threshold based on historical data and combines it with a hierarchical clustering method, clustering into three levels based on brand, series, and overall matching degree. This aligns with the classification logic of vehicle model data, solving the alignment difficulties caused by diverse vehicle model names and complex sources, while ensuring the logical consistency and uniformity of the alignment results. Consistency checks eliminate obviously mismatched data items, ensuring the quality of the initially aligned data groups and preventing invalid data from interfering with subsequent processes. Simultaneously, the data source and matching confidence level are recorded, enabling traceability of the alignment process. This step achieves automatic and accurate alignment of vehicle model data from different sources, breaking down "data silos" and allowing high-value private domain data from organizations like Chaboshi to be effectively integrated with multi-source industry data. This fully leverages the collaborative value of data and solves the pain points of difficult integration and low data utilization of multi-source heterogeneous vehicle model data. This step provides high-quality initial aligned data for subsequent conflict detection and data standardization, promoting the standardization and integration of vehicle model data and improving the effectiveness of data-driven decision-making.

[0040] Preferably, the process of performing conflict detection on the initially aligned vehicle data group, identifying conflicting items with inconsistent data, and resolving the conflicting items using a three-level processing flow includes the following steps:

[0041] Traverse the initially aligned vehicle data group, compare the same parameters of each data item in the group, and identify conflicting items with different parameter values;

[0042] Conflict items are classified into critical parameter conflicts and non-critical parameter conflicts based on parameter importance. Among them, vehicle identification number and core configuration are critical parameters, while body color is a non-critical parameter.

[0043] A three-level processing flow is adopted to resolve conflicts. First, rule filtering is performed, and the conflict items are judged by applying the preset conflict resolution rules. For the conflict items that cannot be resolved by rule filtering, statistical decision processing is performed. The confidence score of each parameter value is calculated based on the historical accuracy of each data source and the data collection time. The parameter value with the highest confidence score is selected as the valid data.

[0044] For highly complex conflicts that cannot be resolved by statistical decision-making, a manual review process is triggered, and the conflict details are sent to professionals for judgment, and the results of the manual decision-making are recorded.

[0045] The parameter values, after being resolved through a three-level processing flow, are integrated to form unified, standardized vehicle model data, while the conflict resolution process and its basis are recorded.

[0046] This invention first traverses the data sets to compare identical parameters, accurately identifying conflicting parameters. Then, it categorizes these conflicts into critical and non-critical based on their importance, aligning with the core needs of the used car industry and prioritizing the consistency of key parameters such as vehicle identification numbers and core configurations. A three-tiered conflict resolution process is implemented progressively: rule-based filtering quickly handles common conflicts, statistical decision-making combines historical data performance and collection time to calculate reliability and select optimal parameter values, and manual review addresses challenging conflicts to ensure comprehensive and reasonable resolution. After resolution, parameters are integrated to generate standardized data, while the process and supporting documentation are recorded, enabling traceability of conflict resolution. This solves the data chaos caused by inconsistent parameters from multiple data sources, ensures the reliability of standardized data, activates the value of private domain data from institutions like ChaDoctor, breaks down "data silos," provides high-quality data support for data-driven decision-making, and promotes the standardization and upgrading of vehicle model data.

[0047] Preferably, the step of incorporating standardized vehicle model data into a self-optimizing knowledge base to absorb human feedback through a closed-loop learning mechanism, continuously optimizing the feature extraction model and matching algorithm, and thus forming corresponding closed-loop learning records includes the following steps:

[0048] The generated standardized vehicle data is categorized and stored in a self-optimizing knowledge base according to data type and timestamp, and a standardized vehicle data indexing system is established.

[0049] Collect human feedback information generated during the conflict resolution process, including human review results, conflict judgment criteria, and parameter correction suggestions;

[0050] The human feedback information is structured and processed to extract the conflict patterns, judgment rules and feature importance assessments contained therein, forming feedback training samples;

[0051] The feedback training samples are input into the optimization module of the feature extraction model to adjust the extraction weights of text features and visual features, thereby enhancing the ability to identify conflicting features.

[0052] The dynamic weight calculation model is optimized based on feedback training samples, and the basic weight coefficients and credibility evaluation algorithms for each feature dimension are updated.

[0053] Offline tests were conducted on the optimized feature extraction model and dynamic weight calculation model to verify their effectiveness in processing historical conflict data.

[0054] The optimized model that passes the test is deployed to the actual application process, and the optimization process and results are recorded and stored in the self-optimization knowledge base to form a complete closed-loop learning record.

[0055] This invention achieves long-term data reuse and rapid retrieval by classifying and storing standardized data in a self-optimizing knowledge base and establishing an indexing system. It collects human feedback from conflict resolution processes, transforming it into training samples to fully leverage the value of human experience. Based on sample-optimized feature extraction and dynamic weight calculation models, it adjusts feature extraction weights, updates basic weight coefficients and credibility algorithms, enhancing the model's ability to identify and match conflict-prone features. Offline testing and actual deployment ensure optimization effectiveness, forming a complete closed-loop learning record. This step enables continuous iterative upgrades of the model and algorithm, allowing the system to adapt to the diverse naming conventions, dynamic parameters, and significant quality differences in vehicle model data. It ensures the long-term effectiveness of multi-source data integration, fully leveraging the synergistic value of private domain data and industry multi-source data, solving the pain points of low data utilization and poor model adaptability, and promoting the continuous digital upgrade of used car inspection and the automotive aftermarket.

[0056] Preferably, the structured processing of the human feedback information, extracting the conflict patterns, judgment rules, and feature importance assessment contained therein, includes the following steps:

[0057] Text parsing of human feedback information identifies the types of conflicting parameters, sources of conflicting data, and final decision results involved in the feedback information;

[0058] Based on the analysis results, we summarize the conflict patterns and common conflict manifestations of different parameter types, including character difference patterns of vehicle identification codes and alias difference patterns of configuration descriptions.

[0059] Extract judgment rules from human decision-making data and transform the judgment criteria expressed in natural language into structured conditional judgment statements;

[0060] Based on the manual assessment of the importance of each feature in conflict judgment, the weighting factors of the corresponding features are adjusted.

[0061] The conflict patterns, judgment rules, and weighting factors are associated with the corresponding conflict cases to form a sample structure that includes input features, conflict types, resolution methods, and effect evaluation.

[0062] Standardize the sample structure, unify the data format and representation, and ensure the consistency and usability of the samples;

[0063] The standardized samples are classified according to the type of conflict, and feedback training samples are constructed for different conflict scenarios.

[0064] This invention uses text parsing of human feedback information to accurately identify the types, sources, and decision results of conflicting parameters, summarizes common conflict patterns for different parameters, and aligns with the conflict characteristics of used car model data, such as differences in vehicle identification number characters and configuration aliases. Natural language judgment criteria are transformed into structured conditional statements, judgment rules are extracted, feature weight influencing factors are adjusted, and conflict cases are linked to form a standardized sample structure, ensuring sample consistency and usability. Feedback training samples are constructed according to conflict type. This step achieves effective transformation and reuse of human feedback information, making fragmented human experience the core driving force for model optimization. It provides targeted, high-quality sample support for subsequent optimization of feature extraction and dynamic weight calculation models, improving the efficiency and rationality of model optimization, enabling the model to better adapt to the specific characteristics of used car model data, further enhancing the effectiveness of multi-source data integration, and fully releasing the value of the data.

[0065] Preferably, the step of inputting the feedback training samples into the optimization module of the feature extraction model, adjusting the extraction weights of text features and visual features, and enhancing the ability to identify conflicting features includes the following steps:

[0066] The text description information and corresponding conflict labels in the feedback training samples are input into the optimization module of the text feature extraction model;

[0067] The model's recognition error on conflict-prone text features is analyzed, and the contribution and error rate of each text feature dimension are calculated.

[0068] The corresponding contribution conflict ratio is calculated based on the contribution and error rate of each text feature dimension.

[0069] Adjust the attention weight of the text feature extraction model based on the contribution conflict ratio, and increase the attention to conflict-prone features, including configuration aliases and differences in year representation;

[0070] The vehicle image information and corresponding conflict labels in the feedback training samples are input into the optimization module of the visual feature extraction model; the recognition error of the model on conflict-prone visual features is analyzed, and the visual regions that are crucial to conflict judgment are located.

[0071] Adjust the layer weights of the convolutional neural network to enhance the feature extraction capability for key visual regions, including vehicle body lines and exclusive logos;

[0072] The adjusted feature extraction model is retrained using the backpropagation algorithm and validated using feedback training samples until the model's recognition accuracy on conflict-prone features reaches a preset threshold.

[0073] This invention optimizes the model by inputting feedback training samples and conflict labels into the model optimization module, analyzing the model's recognition error, calculating the contribution and error rate of text feature dimensions, adjusting attention weights, and enhancing the focus on easily conflicting text features such as configuration aliases and differences in year descriptions. For visual features, key visual regions are located, and the weights of the convolutional neural network layers are adjusted to strengthen the extraction capability of unique visual identifiers such as vehicle body lines and exclusive logos. The model is then retrained using a backpropagation algorithm to ensure that the model's recognition performance on easily conflicting features meets the standards. This step optimizes the performance of the feature extraction model, enabling it to more accurately and comprehensively capture multimodal features of vehicle models, reducing matching bias and the increase in conflict items caused by inaccurate recognition of easily conflicting features, improving the effectiveness of multi-source data matching and alignment, activating high-value visual and textual information in the private domain data of institutions such as Cha Doctor, breaking down "data silos," providing more reliable feature support for subsequent conflict detection and data standardization, and promoting the refinement of vehicle model data integration.

[0074] Preferably, the algorithm for optimizing the dynamic weight calculation model based on feedback training samples and updating the basic weight coefficients and credibility evaluation of each feature dimension includes the following steps:

[0075] Extract the actual contribution value of each feature dimension in conflict resolution from the feedback training samples, and calculate the decision influence of different feature dimensions;

[0076] The basic weight coefficients of each feature dimension are readjusted based on the impact of decision-making, and the weight values ​​of feature dimensions that play a key role in conflict resolution are increased.

[0077] Analyze the performance differences of various data sources under different conflict scenarios, and establish a correlation model between data source credibility and conflict type;

[0078] The credibility assessment algorithm is optimized based on the association model, enabling the algorithm to dynamically adjust the level of trust in different data sources according to the specific conflict type.

[0079] The optimized basic weight coefficients and the credibility assessment algorithm are integrated into the dynamic weight calculation model to form the updated model parameters.

[0080] The updated dynamic weight calculation model was tested using historical matching data to calculate the resolution accuracy of the model under different conflict scenarios.

[0081] The model parameters are fine-tuned based on the model's resolution accuracy in different conflict scenarios until the model's performance in all conflict scenarios reaches the preset standard, thereby completing the optimization of the dynamic weight calculation model.

[0082] This invention extracts the decision-making influence of feature dimensions from feedback training samples, readjusts the basic weight coefficients, and highlights the feature dimensions that play a key role in conflict resolution, aligning with the core needs of used car model data. It analyzes the performance differences of data sources in different conflict scenarios, constructs a correlation model between data source credibility and conflict type, and optimizes the credibility assessment algorithm to dynamically adapt to different conflict scenarios, improving the rationality of weight allocation. The optimized model is tested with historical data and fine-tuned to ensure performance meets standards in various conflict scenarios, completing model optimization. This step improves the flexibility and rationality of the dynamic weight calculation model, allowing it to better adapt to the characteristics of diverse vehicle model data naming, dynamic parameters, and large quality differences. It enhances the consistency and effectiveness of multi-source data matching, fully leverages the synergistic value of private domain data and industry multi-source data, solves the pain points of difficult multi-source data integration and low data utilization, and provides more reliable algorithmic support for automatic vehicle model data alignment and conflict resolution.

[0083] The present invention has the following specific beneficial effects:

[0084] (1) By comprehensively collecting three types of vehicle model data—text descriptions, vehicle images, and structured parameters—from multiple heterogeneous data sources, differentiated preprocessing and feature extraction methods were adopted for different types of data. This standardized the data format and captured core features, solving the problem of differences in format and naming rules among different data sources, and fully exploring the multi-dimensional value of the data. By unifying dimensions and fusing features to form a comprehensive feature set, the limitations of single feature extraction were broken, and high-value information such as vehicle component status and accident records in the private domain data of institutions such as Cha Doctor were activated. This laid a comprehensive feature foundation for subsequent data matching and alignment, broke down "data silos," improved data utilization, promoted the transformation of vehicle model data from fragmented to standardized, and provided solid feature support for data-driven decision-making.

[0085] (2) A dynamic weight calculation model is constructed based on a comprehensive feature set, which can dynamically adjust feature weights according to the credibility of the data source, data integrity, and historical matching accuracy, thus overcoming the limitations of fixed weights and adapting to the quality differences of different data sources. A differentiated similarity calculation method is adopted for multimodal features, which fits the characteristics of each type of feature and focuses on core information such as brand model, core visual identity, and key parameters, solving the matching difficulties caused by the diversity of vehicle names and parameter differences. Through weighted fusion and standardization, a unified range of matching similarity results is obtained, realizing the scientific matching of vehicle data from different sources, promoting the effective connection between private domain data of institutions such as Cha Doctor and multi-source data in the industry, giving full play to the value of data collaboration, breaking down "data silos", improving the effectiveness and rationality of multi-source data matching, providing a reliable matching basis for subsequent automatic alignment of vehicle data, and further improving data utilization and data value.

[0086] (3) Based on the matching similarity results, a hierarchical clustering method is used to automatically align vehicle model data by brand, series, and overall matching degree, which conforms to the classification logic of vehicle model data and solves the alignment difficulties caused by the complex sources and diverse naming of vehicle models. Consistency checks ensure the quality of the initially aligned data groups. Conflict detection is performed on the initially aligned data groups, and conflict items are classified according to parameter importance. A three-level resolution process is adopted to quickly handle common conflicts, scientifically determine difficult conflicts, and accurately resolve high-difficulty conflicts. This prioritizes the consistency of key parameters while reasonably handling differences in non-key parameters. Finally, conflict-free standardized vehicle model data is generated, which solves the data chaos problem caused by inconsistent parameters in multi-source data, ensures data consistency and reliability, activates the value of private domain data, breaks down "data silos," promotes the upgrading of vehicle model data towards standardization and integration, provides high-quality standardized data support for data-driven decision-making, and reduces business deviations caused by data inconsistency.

[0087] (4) By classifying and storing standardized vehicle model data in a self-optimizing knowledge base, an indexing system is established to achieve long-term data reuse and rapid retrieval; human feedback information during the conflict resolution process is collected and transformed into feedback training samples to fully explore the value of human experience. Based on the samples, the feature extraction model and matching algorithm are continuously optimized to enhance the model's ability to identify conflict-prone features and the rationality of matching. The optimization effect is ensured through offline testing and actual deployment, forming a complete closed-loop learning record. This step realizes the continuous iterative upgrade of the model and algorithm, enabling the system to continuously adapt to the characteristics of diverse vehicle model data naming, dynamic parameters, and large quality differences, ensuring the long-term effectiveness of multi-source data integration, giving full play to the synergistic value of private domain data and industry multi-source data, completely solving the pain points of "data silos" and low data utilization, promoting the continuous digital upgrade of used car inspection and the automotive aftermarket, and improving the long-term effectiveness and reliability of data-driven decision-making. Attached Figure Description

[0088] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0089] Figure 1 This is a schematic diagram of the steps of the automatic alignment and conflict resolution method for vehicle model data based on multi-source heterogeneity of the present invention;

[0090] Figure 2 for Figure 1 A flowchart illustrating the steps in S01. Detailed Implementation

[0091] The present invention will be further described below with reference to the accompanying drawings and embodiments, but this should not be construed as limiting the present invention.

[0092] To achieve the above objectives, please refer to Figures 1 to 2 This invention provides a method for automatic alignment and conflict resolution of vehicle model data based on multi-source heterogeneity, comprising the following steps:

[0093] S01: Obtain vehicle model data from multiple heterogeneous data sources, perform multimodal feature extraction on the vehicle model data, and generate a comprehensive feature set containing text features, visual features, and structured features;

[0094] In this embodiment of the invention, a data acquisition tool is used to comprehensively collect vehicle model data from multiple heterogeneous data sources through a preset data interface. The collected data explicitly includes text description information, vehicle image information, and structured parameter information. The text description information covers brand, model, configuration, and functions, while the vehicle image information includes visual elements such as exterior outline, grille design, and wheel style. The structured parameter information includes fixed parameters such as vehicle identification number, registration date, mileage, and engine displacement. After collection, a multimodal feature extraction tool is used to extract features from the three types of data. For the text description information, a text preprocessing tool removes redundant characters, standardizes the expression format, and converts aliases to standard terms. Then, a pre-trained text understanding model is used to capture semantic information and generate text feature vectors. For the vehicle image information, an image preprocessing tool performs size normalization, angle correction, and background removal. Then, a convolutional neural network is used to extract visual identifiers and generate visual feature vectors. For the structured parameter information, a parameter parsing tool extracts core parameters and standardizes the format to generate structured feature vectors. The three types of vector dimensions are calibrated using a dimension unification tool, and text feature vectors, visual feature vectors, and structured feature vectors are integrated using a feature fusion tool to form a comprehensive feature set containing multimodal information, providing full support for subsequent matching calculations.

[0095] S02: Construct a dynamic weight calculation model based on a comprehensive feature set, and perform matching calculations on vehicle model data from different sources based on the dynamic weight calculation model to obtain vehicle model data matching similarity results;

[0096] In this embodiment of the invention, a dynamic weight calculation model is built using a model building tool. This model has a built-in data source evaluation module and weight adjustment algorithm, which can dynamically adjust the weight values ​​of each feature dimension in real time based on the credibility, data integrity, and historical matching accuracy of different data sources, ensuring that the weight allocation is consistent with data quality. After the model is built, for each type of feature vector in the comprehensive feature set, the matching degree is calculated using the corresponding similarity algorithm. For text feature vectors, the semantic similarity algorithm is used to compare brand and model consistency and configuration relevance; for visual feature vectors, the feature vector cosine similarity algorithm is used to compare unique appearance identifiers; for structured feature vectors, the corresponding algorithm is used according to parameter type: the vehicle identification code uses the string edit distance algorithm, numerical parameters use the relative error algorithm, and time parameters use the time difference algorithm. The dynamic weight calculation model performs weighted fusion of the three types of matching degrees to calculate the comprehensive matching degree between vehicle model data from different sources. Then, a standardization tool is used to process the comprehensive matching degree into a vehicle model data matching similarity result with a unified value range, which is convenient for subsequent clustering alignment.

[0097] S03: Based on the similarity results of vehicle model data matching, hierarchical clustering method is used to automatically align the vehicle model data to form a preliminary aligned vehicle model data group; conflict detection is performed on the preliminary aligned vehicle model data group to identify conflicting items with inconsistent data, and a three-level processing flow is used to resolve the conflicting items to generate conflict-free standardized vehicle model data;

[0098] In this embodiment of the invention, an initial clustering threshold is set using a threshold setting tool. This threshold is determined based on the optimal accuracy of historical matching data. A threshold comparison tool compares the matching similarity results with the threshold, marking potential matching pairs with similarity higher than the threshold. A hierarchical clustering tool performs three-level clustering on potential matching pairs: first, they are grouped into major categories based on brand characteristics; then, they are grouped into sub-clusters based on vehicle model series; and finally, they are grouped into the same data group based on comprehensive matching degree, forming initially aligned vehicle model data groups. A data traversal tool traverses the data groups, and a parameter comparison tool compares different values ​​of the same parameter to identify conflicting items with differing parameter values. A conflict classification tool categorizes these conflicts into critical parameter conflicts and non-critical parameter conflicts based on parameter importance. A three-level processing flow resolves conflicts: first, a rule filtering tool applies preset rules to resolve simple conflicts; then, a statistical decision tool calculates a credibility score based on the historical accuracy of the data source and the collection time to select the optimal parameter value; finally, high-difficulty conflicts trigger manual review to determine valid parameter values. A data integration tool integrates the resolved parameter values ​​to generate conflict-free standardized vehicle model data.

[0099] S04: Incorporate standardized vehicle model data into a self-optimizing knowledge base to absorb human feedback through a closed-loop learning mechanism, continuously optimize the feature extraction model and matching algorithm, and thus form corresponding closed-loop learning records.

[0100] In this embodiment of the invention, standardized vehicle model data is generated and categorized by data type and timestamp using a data storage tool, and then stored in an orderly manner in a self-optimizing knowledge base. Simultaneously, an index building tool is used to establish an index system centered on brand, model, and vehicle identification number (VIN), facilitating rapid retrieval and querying. A feedback collection tool is used to collect human feedback information during the conflict resolution process, including human review results, judgment criteria, and parameter correction suggestions. A structured processing tool is used to extract conflict patterns, judgment rules, and feature importance assessments, forming feedback training samples. These training samples are input into the feature extraction model optimization module to adjust the weights of text and visual feature extraction, enhancing the ability to identify conflict-prone features. Simultaneously, the dynamic weight calculation model is optimized based on the samples, updating the basic weight coefficients and credibility evaluation algorithm. An offline testing tool is used to verify the processing effect of the optimized model. The tested model is deployed to practical applications, and a recording tool is used to store the optimization process, parameters, and effects in the self-optimizing knowledge base, forming a complete closed-loop learning record and enabling continuous model iteration.

[0101] Furthermore, such as Figure 2 As shown, the process of acquiring vehicle model data from multiple heterogeneous data sources and extracting multimodal features from the vehicle model data includes the following steps:

[0102] S011: Collect vehicle model data from multiple heterogeneous data sources through a data interface. This vehicle model data includes text description information, vehicle image information, and structured parameter information.

[0103] S012: Preprocess the text description information, including removing redundant characters, standardizing the expression format, and converting aliases to standard terms to form standardized text data; use a pre-trained text understanding model to extract features from the standardized text data, capture the semantic information corresponding to the brand, model, configuration, and year contained in the text, and generate text feature vectors;

[0104] S013: Preprocess the vehicle image information, including size normalization, angle correction, and background removal, to obtain standardized vehicle images; use a convolutional neural network to extract visual features from the standardized vehicle images, focusing on capturing the unique visual identifiers corresponding to the vehicle's exterior outline, grille shape, and wheel style, and generating visual feature vectors.

[0105] S014: Parse the structured parameter information to extract the vehicle identification number, registration time, mileage, and displacement parameters, and generate a structured feature vector after format standardization.

[0106] S015: Unify the dimensions and fuse the features of text feature vectors, visual feature vectors and structured feature vectors to form a comprehensive feature set containing multimodal information.

[0107] In this embodiment of the invention, a data acquisition tool is used to comprehensively collect vehicle model data from multiple heterogeneous data sources through a preset data interface. The collected data explicitly includes textual description information, vehicle image information, and structured parameter information. The textual description information covers brand introduction, vehicle configuration, and functional descriptions. The vehicle image information includes exterior front and side views, interior images, and close-up images of core components. The structured parameter information includes fixed parameters such as vehicle identification number, registration date, mileage, and engine displacement. After collection, a text preprocessing tool is used to perform preprocessing operations on the textual description information. A redundant character removal tool removes spaces, garbled characters, and other redundant characters. A format unification tool adjusts text with different expression formats to a unified standard format. An alias conversion tool converts all aliases of vehicle models and components into industry standard terms, forming standardized text data. A pre-trained text understanding model is used to extract features from the standardized text data. The model deeply mines the semantics of the text, capturing the semantic information corresponding to the brand, model, configuration, and year. A vector generation tool converts the semantic information into text feature vectors. For vehicle image information, image preprocessing tools are used to perform preprocessing operations. A size normalization tool adjusts all images to a uniform size, an angle correction tool corrects image tilt deviations, and a background removal tool removes irrelevant backgrounds to obtain standardized vehicle images. A convolutional neural network is then used to extract visual features from the standardized vehicle images. The network focuses on capturing unique visual identifiers corresponding to the vehicle's exterior outline, grille shape, and wheel style, converting these visual features into visual feature vectors. For structured parameter information, a parameter parsing tool is used to accurately extract vehicle identification numbers, registration dates, mileage, and engine displacement parameters. A format standardization tool adjusts parameters from different formats to a unified standard, generating structured feature vectors. A dimensionality unification tool calibrates the dimensions of textual feature vectors, visual feature vectors, and structured feature vectors to ensure consistency. A feature fusion tool integrates the three types of vectors to form a comprehensive feature set containing multimodal information, providing comprehensive feature support for subsequent matching calculations.

[0108] Furthermore, the process of constructing a dynamic weight calculation model based on a comprehensive feature set and performing matching calculations on vehicle model data from different sources based on the dynamic weight calculation model includes the following steps:

[0109] A dynamic weight calculation model is constructed, which can dynamically adjust the weight values ​​of each feature dimension based on the credibility, data integrity, and historical matching accuracy of different data sources.

[0110] For the text feature vectors in the comprehensive feature set, a semantic similarity algorithm is used to calculate the text matching degree between different vehicle models, focusing on the consistency of brand and model descriptions and the relevance of configuration descriptions;

[0111] For visual feature vectors, the cosine similarity algorithm of feature vectors is used to calculate the visual matching degree between data of different vehicle models, and the unique logos and design features of the vehicle appearance are compared first.

[0112] For structured feature vectors, different similarity calculation methods are used according to parameter type: string edit distance algorithm is used for vehicle identification code, relative error algorithm is used for numerical parameters, and time difference algorithm is used for time parameters.

[0113] The text matching degree, visual matching degree and structured feature matching degree are weighted and fused by a dynamic weight calculation model to calculate the comprehensive matching degree between vehicle models from different sources.

[0114] The overall matching degree is standardized to obtain vehicle model data matching similarity results with a uniform value range.

[0115] In this embodiment of the invention, a dynamic weight calculation model is built using a model building tool. This model incorporates a data source evaluation module and a weight adjustment algorithm, which can dynamically adjust the weight values ​​of each feature dimension in real time based on the credibility, data integrity, and historical matching accuracy of different data sources, ensuring that the weight allocation aligns with data quality. For text feature vectors in the comprehensive feature set, a semantic similarity algorithm is used to calculate the text matching degree between vehicle model data from different sources. The algorithm focuses on comparing the consistency of brand and model descriptions and checking the relevance of configuration descriptions to accurately calculate the degree of matching at the text level. For visual feature vectors, a feature vector cosine similarity algorithm is used to calculate the visual matching degree between vehicle model data from different sources. The algorithm prioritizes comparing the unique identifiers and design features of the vehicle's exterior, focusing on checking visual elements that are not easily confused, such as grille shapes and wheel styles, to ensure the accuracy of visual matching. For the structured feature vectors corresponding to structured parameter information, different similarity calculation methods are adopted according to the parameter type. For string parameters such as vehicle identification numbers, the string edit distance algorithm is used to calculate the matching degree and check the consistency of character sequences. For numerical parameters such as mileage and engine displacement, the relative error algorithm is used to calculate the matching degree and check the degree of numerical deviation. For time parameters such as registration time, the time difference algorithm is used to calculate the matching degree and compare the consistency of time information. A dynamic weight calculation model is used to weight and fuse text matching degree, visual matching degree, and structured feature matching degree. The model calculates the three types of matching degree according to the real-time weight allocation rules to obtain the comprehensive matching degree between vehicle model data from different sources, which comprehensively reflects the matching situation of multi-dimensional features. Standardization tools are used to process the comprehensive matching degree and convert it into vehicle model data matching similarity results with a uniform value range, which facilitates subsequent clustering comparison.

[0116] Furthermore, the weighted fusion of text matching degree, visual matching degree, and structured feature matching degree through a dynamic weight calculation model includes the following steps:

[0117] Retrieve historical performance parameters of each data source from the self-optimizing knowledge base, including data accuracy, integrity score, and conflict rate indicators;

[0118] The credibility weight of each data source is calculated based on historical performance parameters. Features corresponding to data sources with high credibility receive higher weight in the matching calculation.

[0119] The basic weight coefficients are set according to the type characteristics of vehicle data, with vehicle identification code related features set as the highest basic weight, visual features and core configuration text features set as the second highest basic weight, and auxiliary parameter features set as the lowest basic weight.

[0120] By combining the credibility weight and the basic weight coefficient, a dynamic weight calculation model is used to generate real-time weight values ​​for each feature dimension.

[0121] The text matching degree, visual matching degree, and structured feature matching degree are multiplied by their respective real-time weight values ​​to obtain the weighted matching scores for each dimension.

[0122] The weighted matching scores of each dimension are summed to obtain the overall matching degree between vehicle models from different sources.

[0123] In this embodiment of the invention, a data retrieval tool is used to retrieve historical performance parameters of each data source from a self-optimizing knowledge base. These historical performance parameters explicitly include data accuracy, integrity score, and conflict rate indicators. Data accuracy reflects the proportion of correct data from the data source in the past; integrity score reflects the completeness of the data source's data; and conflict rate reflects the frequency of conflicts between the data source and other data sources. A weight calculation tool is used to calculate the credibility weight of each data source based on these historical performance parameters. Data sources with high data accuracy, high integrity score, and low conflict rate have higher credibility weights, and their features receive higher weights in the matching calculation. A basic weight setting tool is used to set basic weight coefficients for each feature dimension based on the type characteristics of the vehicle model data. Vehicle identification code-related features, as the core identifier, are set with the highest basic weight; visual features and core configuration text features, as important distinguishing criteria, are set with the second highest basic weight; and auxiliary parameter features such as mileage are set with lower basic weights. Combining the credibility weights of each data source with the basic weight coefficients of each feature dimension, the dynamic weight calculation model generates real-time weight values ​​for each feature dimension through its built-in computational logic. These real-time weight values ​​are dynamically adjusted according to the performance of the data source. A weighted matching tool is used to multiply the text matching degree, visual matching degree, and structured feature matching degree by their corresponding real-time weight values ​​to obtain a weighted matching score for each dimension. The weighted matching score reflects the actual contribution of each dimension's matching degree. A summation tool is then used to sum the weighted matching scores of each dimension to obtain the comprehensive matching degree among vehicle model data from different sources. The comprehensive matching degree accurately reflects the overall matching level of multi-source data.

[0124] Furthermore, the automatic alignment of vehicle model data using a hierarchical clustering method based on the similarity results of the vehicle model data includes the following steps:

[0125] Set an initial clustering threshold, which is determined based on the best accuracy of historical matching data;

[0126] The vehicle model data matching similarity results are compared with the initial clustering threshold, and vehicle model data with similarity results higher than the threshold are marked as potential matching pairs;

[0127] Hierarchical clustering is used to cluster potential matching pairs. First, a first-level clustering is performed based on brand characteristics, grouping vehicle data from the same brand into one major category. Within each major brand category, a second-level clustering is performed based on vehicle series characteristics, grouping vehicle data belonging to the same series into one sub-cluster. Within each sub-cluster, a third-level clustering is performed based on the overall matching degree, grouping vehicle data with an overall matching degree reaching a preset threshold into the same data group.

[0128] Perform a consistency check on each data group to ensure that the vehicle model data within the group are basically consistent in terms of core features, and remove data items that are obviously mismatched;

[0129] The data groups that pass the consistency check are marked as initially aligned vehicle model data groups, and the source information and matching confidence of each data item in the group are recorded.

[0130] In this embodiment of the invention, an initial clustering threshold is set using a threshold setting tool. This initial clustering threshold is determined based on the optimal accuracy of historical matching data, ensuring that the threshold can effectively distinguish between matching and non-matching data. A threshold comparison tool is used to compare the similarity results of all vehicle model data with the initial clustering threshold one by one. Vehicle model data with similarity results higher than the threshold are marked as potential matching pairs using a labeling tool. Potential matching pairs are data combinations that may belong to the same vehicle model. A hierarchical clustering tool is used to cluster these potential matching pairs. The clustering process is executed in three stages: First, a first-level clustering is performed based on brand characteristics. Brand information for each potential matching pair is extracted using a brand identification tool, grouping vehicle model data from the same brand into a single category, achieving brand-level differentiation. Within each brand category, a second-level clustering is performed based on vehicle series characteristics. Vehicle series information for each data point is extracted, grouping vehicle model data belonging to the same series into a sub-cluster, achieving series-level differentiation. Within each sub-cluster, a third-level clustering is performed based on comprehensive matching degree. Vehicle model data with a comprehensive matching degree reaching a preset threshold are grouped into the same data group, achieving aggregation at the specific vehicle model level. A consistency check tool was used to comprehensively check each data group, verifying the consistency of vehicle identification numbers, brands, models, and other core characteristics within the group to ensure that the data within the group essentially belonged to the same vehicle model. A removal tool was used to eliminate obviously mismatched data items within the group, ensuring data group quality. A marking tool was used to mark the data groups that passed the consistency check as initially aligned vehicle model data groups. Simultaneously, a recording tool was used to record the source information and match confidence of each data item within the group. The source information clearly identifies the data source, and the match confidence reflects the reliability of the data group's matching, providing data support for subsequent conflict resolution.

[0131] Furthermore, the process of performing conflict detection on the initially aligned vehicle data groups, identifying conflicting data inconsistencies, and resolving these conflicts using a three-level processing flow includes the following steps:

[0132] Traverse the initially aligned vehicle data group, compare the same parameters of each data item in the group, and identify conflicting items with different parameter values;

[0133] Conflict items are classified into critical parameter conflicts and non-critical parameter conflicts based on parameter importance. Among them, vehicle identification number and core configuration are critical parameters, while body color is a non-critical parameter.

[0134] A three-level processing flow is adopted to resolve conflicts. First, rule filtering is performed, and the conflict items are judged by applying the preset conflict resolution rules. For the conflict items that cannot be resolved by rule filtering, statistical decision processing is performed. The confidence score of each parameter value is calculated based on the historical accuracy of each data source and the data collection time. The parameter value with the highest confidence score is selected as the valid data.

[0135] For highly complex conflicts that cannot be resolved by statistical decision-making, a manual review process is triggered, and the conflict details are sent to professionals for judgment, and the results of the manual decision-making are recorded.

[0136] The parameter values, after being resolved through a three-level processing flow, are integrated to form unified, standardized vehicle model data, while the conflict resolution process and its basis are recorded.

[0137] In this embodiment of the invention, a data traversal tool is used to comprehensively traverse the initially aligned vehicle model data group, extracting the common parameters of each data item within the group one by one. A parameter comparison tool is then used to compare different parameter values ​​of the same parameter one by one, accurately identifying conflicting items with differing parameter values. The parameter type, parameter value, and data source of each conflicting item are recorded. A conflict classification tool is used to classify the identified conflicting items, clearly distinguishing between critical and non-critical parameter conflicts based on parameter importance. Vehicle identification codes and core configurations, as core criteria for vehicle model identification, are classified as critical parameter conflicts; vehicle body color, as a non-core distinguishing parameter, is classified as a non-critical parameter conflict. After classification, the conflict types are labeled for subsequent targeted processing. A three-level processing flow is used to resolve the conflicting items. First, a rule filtering tool is activated, applying preset conflict resolution rules to judge and process the conflicting items. The preset rules cover parameter priority rules, data source credibility rules, etc., which can directly resolve simple conflicting items. For complex conflicts that cannot be resolved by rule filtering, a statistical decision-making tool is activated. Based on the historical accuracy of each data source and the data collection time, a credibility score is calculated for each conflict parameter value. The credibility score comprehensively reflects the reliability of the parameter value, and the parameter value with the highest credibility score is selected as the valid data for that parameter. For highly complex conflicts that still cannot be resolved by statistical decision-making, a manual review trigger tool is activated. The conflict details, parameter values, data source information, and statistical decision results are uniformly pushed to professionals for manual judgment. The valid parameter values ​​are determined through manual decision-making, and the manual decision results and judgment basis are recorded in detail using a recording tool. A data integration tool is used to uniformly integrate all parameter values ​​after the three-level processing and organize them into a unified standardized vehicle model data according to a preset standard format. At the same time, a process recording tool is used to fully record the entire conflict resolution process, including conflict identification results, resolution steps, resolution methods used, and final basis, ensuring that the conflict resolution process is traceable and verifiable.

[0138] Furthermore, the process of incorporating standardized vehicle model data into a self-optimizing knowledge base to absorb human feedback through a closed-loop learning mechanism, continuously optimizing the feature extraction model and matching algorithm, and thus forming corresponding closed-loop learning records includes the following steps:

[0139] The generated standardized vehicle data is categorized and stored in a self-optimizing knowledge base according to data type and timestamp, and a standardized vehicle data indexing system is established.

[0140] Collect human feedback information generated during the conflict resolution process, including human review results, conflict judgment criteria, and parameter correction suggestions;

[0141] The human feedback information is structured and processed to extract the conflict patterns, judgment rules and feature importance assessments contained therein, forming feedback training samples;

[0142] The feedback training samples are input into the optimization module of the feature extraction model to adjust the extraction weights of text features and visual features, thereby enhancing the ability to identify conflicting features.

[0143] The dynamic weight calculation model is optimized based on feedback training samples, and the basic weight coefficients and credibility evaluation algorithms for each feature dimension are updated.

[0144] Offline tests were conducted on the optimized feature extraction model and dynamic weight calculation model to verify their effectiveness in processing historical conflict data.

[0145] The optimized model that passes the test is deployed to the actual application process, and the optimization process and results are recorded and stored in the self-optimization knowledge base to form a complete closed-loop learning record.

[0146] In this embodiment of the invention, standardized vehicle model data is generated and categorized by data type and timestamp using a data storage tool, and then stored in an orderly manner in a self-optimizing knowledge base. Simultaneously, an index building tool is used to establish a standardized vehicle model data index system, with brand, model, and vehicle identification number as core identifiers for easy retrieval and querying. A feedback collection tool is used to collect all manual feedback information generated during the conflict resolution process. This feedback information explicitly includes manual review results, conflict judgment criteria, parameter correction suggestions, and conflict type analysis, ensuring comprehensive and complete feedback. A structured processing tool is used to process the collected manual feedback information, removing redundant statements and extracting conflict patterns, judgment rules, and feature importance assessments. This information is then organized into feedback training samples according to a preset sample format, with each sample associated with a corresponding conflict case and resolution result. A model optimization tool is used to input the feedback training samples into the optimization module of the feature extraction model. The optimization module adjusts the extraction weights of text and visual features based on the conflict features in the samples, focusing on enhancing the ability to identify conflict-prone features and reducing the recurrence of similar conflicts. The dynamic weight calculation model optimization process is initiated based on feedback training samples. The basic weight coefficients and reliability evaluation algorithms for each feature dimension are updated to better align the model's weight allocation with conflict resolution requirements. Offline testing tools are used to test the optimized feature extraction model and dynamic weight calculation model offline, selecting historical conflict data as test samples to verify the model's performance on such data. The optimized model that passes the tests is then deployed to the actual application process using deployment tools. Simultaneously, a recording tool is used to store information such as the model optimization process, optimization parameters, and test results in a self-optimization knowledge base, forming a complete closed-loop learning record to achieve continuous iterative optimization of the model.

[0147] Furthermore, the structured processing of the human feedback information, extracting the conflict patterns, judgment rules, and feature importance assessment contained therein, includes the following steps:

[0148] Text parsing of human feedback information identifies the types of conflicting parameters, sources of conflicting data, and final decision results involved in the feedback information;

[0149] Based on the analysis results, we summarize the conflict patterns and common conflict manifestations of different parameter types, including character difference patterns of vehicle identification codes and alias difference patterns of configuration descriptions.

[0150] Extract judgment rules from human decision-making data and transform the judgment criteria expressed in natural language into structured conditional judgment statements;

[0151] Based on the manual assessment of the importance of each feature in conflict judgment, the weighting factors of the corresponding features are adjusted.

[0152] The conflict patterns, judgment rules, and weighting factors are associated with the corresponding conflict cases to form a sample structure that includes input features, conflict types, resolution methods, and effect evaluation.

[0153] Standardize the sample structure, unify the data format and representation, and ensure the consistency and usability of the samples;

[0154] The standardized samples are classified according to the type of conflict, and feedback training samples are constructed for different conflict scenarios.

[0155] In this embodiment of the invention, a text parsing tool is used to comprehensively analyze the human feedback information, extracting word by word the conflict parameter types, conflict data sources, differences in parameter values, and final decision results, ensuring that the parsed information is complete and unbiased. Based on the parsing results, a pattern summarization tool is used to summarize common conflict manifestations of different parameter types, clearly identifying character difference patterns in vehicle identification codes, alias difference patterns in configuration descriptions, and deviation patterns in numerical parameters, with each conflict pattern associated with a corresponding case example. A rule extraction tool is used to extract implicit judgment rules from the human decision-making basis, transforming the natural language judgment basis into structured conditional judgment statements through a rule conversion tool, clarifying the judgment logic and applicable scenarios, facilitating subsequent integration into the conflict resolution rule base. Based on the human assessment of the importance of each feature in conflict judgment, a weight adjustment tool is used to adjust the weight influence factors of corresponding features, increasing the influence weight of key judgment features and reducing the influence of secondary features. A sample association tool is used to associate the summarized conflict patterns, extracted judgment rules, and adjusted weight influence factors with corresponding conflict cases, constructing a complete sample structure that includes input features, conflict types, resolution methods, and effect evaluation. The constructed sample structure is processed using sample standardization tools to unify data formats, representation methods, and classification standards, ensuring the consistency and usability of all samples. Classification tools are then used to categorize the standardized samples according to conflict type, constructing feedback training sample sets for different conflict scenarios to provide accurate support for subsequent model optimization.

[0156] Furthermore, the step of inputting the feedback training samples into the optimization module of the feature extraction model, adjusting the extraction weights of text features and visual features, and enhancing the ability to identify conflicting features includes the following steps:

[0157] The text description information and corresponding conflict labels in the feedback training samples are input into the optimization module of the text feature extraction model;

[0158] The model's recognition error on conflict-prone text features is analyzed, and the contribution and error rate of each text feature dimension are calculated.

[0159] The corresponding contribution conflict ratio is calculated based on the contribution and error rate of each text feature dimension.

[0160] Adjust the attention weight of the text feature extraction model based on the contribution conflict ratio, and increase the attention to conflict-prone features, including configuration aliases and differences in year representation;

[0161] The vehicle image information and corresponding conflict labels in the feedback training samples are input into the optimization module of the visual feature extraction model; the recognition error of the model on conflict-prone visual features is analyzed, and the visual regions that are crucial to conflict judgment are located.

[0162] Adjust the layer weights of the convolutional neural network to enhance the feature extraction capability for key visual regions, including vehicle body lines and exclusive logos;

[0163] The adjusted feature extraction model is retrained using the backpropagation algorithm and validated using feedback training samples until the model's recognition accuracy on conflict-prone features reaches a preset threshold.

[0164] In this embodiment of the invention, a feature input tool is used to input the text description information and corresponding conflict labels from the feedback training samples into the optimization module of the text feature extraction model. The optimization module activates an error analysis tool to analyze the model's recognition error on easily conflicting text features, accurately locating text feature dimensions with large recognition deviations. A factor calculation tool is used to calculate the contribution and error rate of each text feature dimension, clarifying the actual role and deviation of each dimension in conflict recognition. Based on the contribution and error rate of each text feature dimension, a ratio calculation tool is used to derive the corresponding contribution-to-conflict ratio, which reflects the balance between feature contribution and recognition error. An attention weight adjustment tool is used to adjust the attention weight of the text feature extraction model based on the contribution-to-conflict ratio, focusing on increasing the attention to easily conflicting features, especially text features that are prone to conflict, such as configuration aliases and differences in year descriptions. The vehicle image information and corresponding conflict labels from the feedback training samples are input into the optimization module of the visual feature extraction model. The optimization module uses a visual error analysis tool to analyze the model's recognition error on easily conflicting visual features, accurately locating visual areas crucial for conflict judgment, including vehicle body lines, exclusive logos, and grille shapes. A network weight adjustment tool was used to adjust the layer weights of the convolutional neural network, enhancing its ability to extract features from key visual regions and improving the recognition accuracy of conflict-prone visual features. A model training tool was then used to retrain the adjusted feature extraction model via backpropagation, with feedback training samples used throughout the process for validation. The training was iterated repeatedly until the model's recognition accuracy on conflict-prone features reached a preset threshold, thus completing the optimization of the text and visual feature extraction model.

[0165] Furthermore, the algorithm for optimizing the dynamic weight calculation model based on feedback training samples and updating the basic weight coefficients and credibility evaluation of each feature dimension includes the following steps:

[0166] Extract the actual contribution value of each feature dimension in conflict resolution from the feedback training samples, and calculate the decision influence of different feature dimensions;

[0167] The basic weight coefficients of each feature dimension are readjusted based on the impact of decision-making, and the weight values ​​of feature dimensions that play a key role in conflict resolution are increased.

[0168] Analyze the performance differences of various data sources under different conflict scenarios, and establish a correlation model between data source credibility and conflict type;

[0169] The credibility assessment algorithm is optimized based on the association model, enabling the algorithm to dynamically adjust the level of trust in different data sources according to the specific conflict type.

[0170] The optimized basic weight coefficients and the credibility assessment algorithm are integrated into the dynamic weight calculation model to form the updated model parameters.

[0171] The updated dynamic weight calculation model was tested using historical matching data to calculate the resolution accuracy of the model under different conflict scenarios.

[0172] The model parameters are fine-tuned based on the model's resolution accuracy in different conflict scenarios until the model's performance in all conflict scenarios reaches the preset standard, thereby completing the optimization of the dynamic weight calculation model.

[0173] In this embodiment of the invention, a contribution value extraction tool is used to extract the actual contribution value of each feature dimension in conflict resolution from the feedback training samples. An influence calculation tool is used to calculate the decision influence of different feature dimensions, clarifying the actual role of each feature in conflict resolution. Based on the calculated decision influence, a basic weight adjustment tool is used to readjust the basic weight coefficients of each feature dimension, significantly increasing the weight values ​​of feature dimensions that play a key role in conflict resolution and decreasing the weights of secondary feature dimensions, making the weight allocation more aligned with practical application needs. A scenario analysis tool is used to analyze the performance differences of each data source under different conflict scenarios, recording the accuracy, conflict rate, and other performance indicators of each data source in various conflicts. Based on these indicators, a correlation model construction tool is used to establish a correlation model between data source credibility and conflict type, clarifying the reliability of each data source under different conflict types. Based on the constructed correlation model, an algorithm optimization tool is used to optimize the credibility assessment algorithm, enabling the algorithm to dynamically adjust the trust level of different data sources according to specific conflict types, improving the accuracy of credibility assessment. A parameter integration tool is used to integrate the optimized basic weight coefficients and the credibility assessment algorithm into a dynamic weight calculation model, updating the model's internal parameters and forming a new model parameter configuration. Using historical matching data, a model testing tool was employed to comprehensively test the updated dynamic weight calculation model, calculating its resolution accuracy under different conflict scenarios and recording the test results for each scenario. Based on the model's resolution accuracy under different conflict scenarios, a parameter fine-tuning tool was used to specifically fine-tune the model parameters, repeatedly optimizing and adjusting them until the model's performance in all types of conflict scenarios met the preset standards, ultimately completing the optimization and upgrade of the dynamic weight calculation model.

[0174] Therefore, the embodiments should be considered as exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of the equivalents of the application are intended to be included within the invention.

[0175] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.

Claims

1. A method for automatic alignment and conflict resolution of vehicle model data based on multi-source heterogeneity, characterized in that, Includes the following steps: Acquire vehicle model data from multiple heterogeneous data sources, perform multimodal feature extraction on the vehicle model data, and generate a comprehensive feature set that includes textual features, visual features, and structured features; A dynamic weight calculation model is constructed based on a comprehensive feature set, and the vehicle model data from different sources is matched and calculated based on the dynamic weight calculation model to obtain the vehicle model data matching similarity results. Based on the similarity results of vehicle model data matching, a hierarchical clustering method is used to automatically align the vehicle model data to form initially aligned vehicle model data groups; Conflict detection is performed on the initially aligned vehicle model data groups to identify conflicting items with inconsistent data. A three-level processing flow is used to resolve these conflicts. First, rule filtering is performed, and preset conflict resolution rules are applied to judge the conflicts. For conflicts that cannot be resolved by rule filtering, statistical decision processing is performed. Based on the historical accuracy of each data source and the data collection time, the confidence score of each parameter value is calculated. The parameter value with the highest confidence score is selected as the valid data, and conflict-free standardized vehicle model data is generated. Standardized vehicle model data is incorporated into a self-optimizing knowledge base to absorb human feedback through a closed-loop learning mechanism, continuously optimizing the feature extraction model and matching algorithm, thereby forming corresponding closed-loop learning records. This includes the following steps: The generated standardized vehicle data is categorized and stored in a self-optimizing knowledge base according to data type and timestamp, and a standardized vehicle data indexing system is established. Collect human feedback information generated during the conflict resolution process, including human review results, conflict judgment criteria, and parameter correction suggestions; The human feedback information is structured and processed to extract the conflict patterns, judgment rules and feature importance assessments contained therein, forming feedback training samples; The feedback training samples are input into the optimization module of the feature extraction model to adjust the extraction weights of text features and visual features, thereby enhancing the ability to identify conflicting features. The dynamic weight calculation model is optimized based on feedback training samples, and the basic weight coefficients and credibility evaluation algorithms for each feature dimension are updated. Offline tests were conducted on the optimized feature extraction model and dynamic weight calculation model to verify their effectiveness in processing historical conflict data. The optimized model that passes the test is deployed to the actual application process, and the optimization process and results are recorded and stored in the self-optimization knowledge base to form a complete closed-loop learning record.

2. The method for automatic alignment and conflict resolution of vehicle model data based on multi-source heterogeneity as described in claim 1, characterized in that, The process of acquiring vehicle model data from multiple heterogeneous data sources and extracting multimodal features from the vehicle model data includes the following steps: Vehicle model data is collected from multiple heterogeneous data sources through a data interface. This vehicle model data includes text description information, vehicle image information, and structured parameter information. The text description information is preprocessed, including removing redundant characters, standardizing the expression format, and converting aliases to standard terms to form standardized text data; a pre-trained text understanding model is used to extract features from the standardized text data, capturing the semantic information corresponding to the brand, model, configuration, and year contained in the text, and generating text feature vectors; The vehicle image information is preprocessed, including size normalization, angle correction, and background removal, to obtain standardized vehicle images. A convolutional neural network is used to extract visual features from the standardized vehicle images, focusing on capturing the unique visual identifiers corresponding to the vehicle's exterior outline, grille shape, and wheel style, and generating visual feature vectors. The structured parameter information is parsed to extract vehicle identification number, registration time, mileage, and displacement parameters. After format standardization, a structured feature vector is generated. By unifying the dimensions and fusing the features of text feature vectors, visual feature vectors, and structured feature vectors, a comprehensive feature set containing multimodal information is formed.

3. The method for automatic alignment and conflict resolution of vehicle model data based on multi-source heterogeneity as described in claim 1, characterized in that, The process of constructing a dynamic weight calculation model based on a comprehensive feature set and then performing matching calculations on vehicle model data from different sources based on this dynamic weight calculation model includes the following steps: A dynamic weight calculation model is constructed, which can dynamically adjust the weight values ​​of each feature dimension based on the credibility, data integrity, and historical matching accuracy of different data sources. For the text feature vectors in the comprehensive feature set, a semantic similarity algorithm is used to calculate the text matching degree between different vehicle models, focusing on the consistency of brand and model descriptions and the relevance of configuration descriptions; For visual feature vectors, the cosine similarity algorithm of feature vectors is used to calculate the visual matching degree between data of different vehicle models, and the unique logos and design features of the vehicle appearance are compared first. For structured feature vectors, different similarity calculation methods are used according to parameter type: string edit distance algorithm is used for vehicle identification code, relative error algorithm is used for numerical parameters, and time difference algorithm is used for time parameters. The text matching degree, visual matching degree and structured feature matching degree are weighted and fused by a dynamic weight calculation model to calculate the comprehensive matching degree between vehicle models from different sources. The overall matching degree is standardized to obtain vehicle model data matching similarity results with a uniform value range.

4. The method for automatic alignment and conflict resolution of vehicle model data based on multi-source heterogeneity as described in claim 3, characterized in that, The weighted fusion of text matching degree, visual matching degree, and structured feature matching degree through a dynamic weight calculation model includes the following steps: Retrieve historical performance parameters of each data source from the self-optimizing knowledge base, including data accuracy, integrity score, and conflict rate indicators; The credibility weight of each data source is calculated based on historical performance parameters. Features corresponding to data sources with high credibility receive higher weight in the matching calculation. The basic weight coefficients are set according to the type characteristics of vehicle data, with vehicle identification code related features set as the highest basic weight, visual features and core configuration text features set as the second highest basic weight, and auxiliary parameter features set as the lowest basic weight. By combining the credibility weight and the basic weight coefficient, a dynamic weight calculation model is used to generate real-time weight values ​​for each feature dimension. The text matching degree, visual matching degree, and structured feature matching degree are multiplied by their respective real-time weight values ​​to obtain the weighted matching scores for each dimension. The weighted matching scores of each dimension are summed to obtain the overall matching degree between vehicle models from different sources.

5. The method for automatic alignment and conflict resolution of vehicle model data based on multi-source heterogeneity as described in claim 1, characterized in that, The automatic alignment of vehicle model data using hierarchical clustering based on similarity results includes the following steps: Set an initial clustering threshold, which is determined based on the best accuracy of historical matching data; The vehicle model data matching similarity results are compared with the initial clustering threshold, and vehicle model data with similarity results higher than the threshold are marked as potential matching pairs; Hierarchical clustering is used to cluster potential matching pairs. First, a first-level clustering is performed based on brand characteristics, grouping vehicle data from the same brand into one major category. Within each major brand category, a second-level clustering is performed based on vehicle series characteristics, grouping vehicle data belonging to the same series into one sub-cluster. Within each sub-cluster, a third-level clustering is performed based on the overall matching degree, grouping vehicle data with an overall matching degree reaching a preset threshold into the same data group. Perform a consistency check on each data group to ensure that the vehicle model data within the group are basically consistent in terms of core features, and remove data items that are obviously mismatched; The data groups that pass the consistency check are marked as initially aligned vehicle model data groups, and the source information and matching confidence of each data item in the group are recorded.

6. The method for automatic alignment and conflict resolution of vehicle model data based on multi-source heterogeneity as described in claim 1, characterized in that, The process of performing conflict detection on the initially aligned vehicle data groups, identifying conflicting data inconsistencies, and resolving these conflicts using a three-level processing flow includes the following steps: Traverse the initially aligned vehicle data group, compare the same parameters of each data item in the group, and identify conflicting items with different parameter values; Conflict items are classified into critical parameter conflicts and non-critical parameter conflicts based on parameter importance. Among them, vehicle identification number and core configuration are critical parameters, while body color is a non-critical parameter. A three-level processing flow is adopted to resolve conflicts. First, rule filtering is performed, and the conflict items are judged by applying the preset conflict resolution rules. For the conflict items that cannot be resolved by rule filtering, statistical decision processing is performed. The confidence score of each parameter value is calculated based on the historical accuracy of each data source and the data collection time. The parameter value with the highest confidence score is selected as the valid data. For highly complex conflicts that cannot be resolved by statistical decision-making, a manual review process is triggered, and the conflict details are sent to professionals for judgment, and the results of the manual decision-making are recorded. The parameter values, after being resolved through a three-level processing flow, are integrated to form unified, standardized vehicle model data, while the conflict resolution process and its basis are recorded.

7. The method for automatic alignment and conflict resolution of vehicle model data based on multi-source heterogeneity as described in claim 1, characterized in that, The process of structuring the human feedback information and extracting the conflict patterns, judgment rules, and feature importance assessments includes the following steps: Text parsing of human feedback information identifies the types of conflicting parameters, sources of conflicting data, and final decision results involved in the feedback information; Based on the analysis results, we summarize the conflict patterns and common conflict manifestations of different parameter types, including character difference patterns of vehicle identification codes and alias difference patterns of configuration descriptions. Extract judgment rules from human decision-making data and transform the judgment criteria expressed in natural language into structured conditional judgment statements; Based on the manual assessment of the importance of each feature in conflict judgment, the weighting factors of the corresponding features are adjusted. The conflict patterns, judgment rules, and weighting factors are associated with the corresponding conflict cases to form a sample structure that includes input features, conflict types, resolution methods, and effect evaluation. Standardize the sample structure, unify the data format and representation, and ensure the consistency and usability of the samples; The standardized samples are classified according to the type of conflict, and feedback training samples are constructed for different conflict scenarios.

8. The method for automatic alignment and conflict resolution of vehicle model data based on multi-source heterogeneity as described in claim 7, characterized in that, The step of inputting feedback training samples into the optimization module of the feature extraction model, adjusting the extraction weights of text features and visual features, and enhancing the ability to identify conflicting features includes the following steps: The text description information and corresponding conflict labels in the feedback training samples are input into the optimization module of the text feature extraction model; The model's recognition error on conflict-prone text features is analyzed, and the contribution and error rate of each text feature dimension are calculated. The corresponding contribution conflict ratio is calculated based on the contribution and error rate of each text feature dimension. Adjust the attention weight of the text feature extraction model based on the contribution conflict ratio, and increase the attention to conflict-prone features, including configuration aliases and differences in year representation; The vehicle image information and corresponding conflict labels in the feedback training samples are input into the optimization module of the visual feature extraction model; the recognition error of the model on conflict-prone visual features is analyzed, and the visual regions that are crucial to conflict judgment are located. Adjust the layer weights of the convolutional neural network to enhance the feature extraction capability for key visual regions, including vehicle body lines and exclusive logos; The adjusted feature extraction model is retrained using the backpropagation algorithm and validated using feedback training samples until the model's recognition accuracy on conflict-prone features reaches a preset threshold.

9. The method for automatic alignment and conflict resolution of vehicle model data based on multi-source heterogeneity as described in claim 1, characterized in that, The algorithm for optimizing the dynamic weight calculation model based on feedback training samples and updating the basic weight coefficients and credibility evaluation of each feature dimension includes the following steps: Extract the actual contribution value of each feature dimension in conflict resolution from the feedback training samples, and calculate the decision influence of different feature dimensions; The basic weight coefficients of each feature dimension are readjusted based on the impact of decision-making, and the weight values ​​of feature dimensions that play a key role in conflict resolution are increased. Analyze the performance differences of various data sources under different conflict scenarios, and establish a correlation model between data source credibility and conflict type; The credibility assessment algorithm is optimized based on the association model, enabling the algorithm to dynamically adjust the level of trust in different data sources according to the specific conflict type. The optimized basic weight coefficients and the credibility assessment algorithm are integrated into the dynamic weight calculation model to form the updated model parameters. The updated dynamic weight calculation model was tested using historical matching data to calculate the resolution accuracy of the model under different conflict scenarios. The model parameters are fine-tuned based on the model's resolution accuracy in different conflict scenarios until the model's performance in all conflict scenarios reaches the preset standard, thereby completing the optimization of the dynamic weight calculation model.

Citation Information

Patent Citations

  • Multi-data standard definition conflict resolution method based on large model and knowledge graph

    CN120653720A

  • Cross-modal knowledge reasoning method and device for industrial quality inspection and medium

    CN120069096A

  • Dietary structure evaluation method and system based on digital intelligent analysis

    CN120708814A