Automobile intelligent factor data preprocessing algorithm
Through the automotive intelligent factor data preprocessing algorithm, data is cleaned and standardized, initial weights are calculated and user feedback is adjusted, and the problems of uneven data quality and calculation errors are solved, and the evaluation accuracy and user experience of automotive intelligence are improved.
Patent Information
- Application Number
- CN202510160048.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-06-03
AI Technical Summary
There are noise, error values and missing values in the automotive intelligent factor data, and the data formats are diverse, resulting in uneven data quality and calculation errors, affecting the accurate evaluation of the automotive intelligent level.
A car intelligent factor data preprocessing algorithm is proposed, including data cleaning and standardization, calculation of initial weights (entropy weight method), calculation of intelligent scores, weight adjustment (based on user feedback) and result output.
By cleaning up data to remove noise and error values, standardizing data formats, calculating weights and making adjustments, ensuring the accuracy and consistency of data, improving the evaluation accuracy of the level of automotive intelligence, and better reflecting users' actual needs and market changes.
Smart Images

Figure CN120086501A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of automotive intelligence, and specifically relates to an algorithm for preprocessing automotive intelligent factor data. Background Art
[0002] With the continuous development of automotive intelligent technology, the analysis and processing of automotive intelligent factor data have become crucial in evaluating the automotive intelligent level.
[0003] Automotive intelligence involves numerous complex functions and features, such as intelligent driving assistance systems, intelligent connectivity functions, intelligent voice recognition, and automatic emergency braking. During the process of collecting these intelligent factor data, due to the wide range of data sources, the data presents complex characteristics.
[0004] On the one hand, the quality of the data is uneven. Data from devices such as sensors have problems such as noise, error values, and missing values. For example, in automotive intelligent factor data, sensors may record incorrect readings such as recording the response time of the intelligent driving assistance system as an unreasonable negative value, which will affect the accuracy of the data and further interfere with the subsequent evaluation of the automotive intelligent level.
[0005] On the other hand, the data formats from different sources are diverse. Some data represents numerical values in string form, and such differences in format will cause problems during data calculation. For example, when processing data related to the intelligent connectivity function of a vehicle and data related to the autonomous driving level, if the data types are inconsistent, the calculation cannot be carried out smoothly, thus affecting the accurate evaluation of the overall intelligence of the vehicle. Summary of the Invention
[0006] In view of the above situation, to overcome the defects of the prior art, the present invention provides an algorithm for preprocessing automotive intelligent factor data to at least partially solve the above technical problems.
[0007] The technical solution adopted by the present invention is as follows:
[0008] The present invention proposes an algorithm for preprocessing automotive intelligent factor data, including:
[0009] S1. Data preprocessing: Clean and standardize the data to ensure that all feature columns used for calculation are of numerical type;
[0010] S2. Calculate the initial weight (entropy weight method): Perform normalization processing on each feature, calculate the information entropy of each feature, and calculate the initial weight based on this;
[0011] S3. Calculate the intelligent score: Use the initial weight and feature data to calculate the intelligent score of each vehicle model, and standardize the score to the range of 0 - 1 for comparison;
[0012] S4. Weight Adjustment (Based on User Feedback): Collect user evaluations (positive or negative) of the vehicle models, analyze the relationship between features and user evaluations, adjust the weights to reflect user preferences, and renormalize the weights to ensure the sum is 1;
[0013] S5. Result Output: Output the final scores and adjusted weights to an Excel spreadsheet for further analysis and reporting.
[0014] In one embodiment of the present invention, the feature column names and paths in the code are adjusted according to the actual data structure to ensure that data cleaning and preprocessing adapt to the characteristics of the dataset, especially for handling missing values and outliers.
[0015] In one embodiment of the present invention, the collection and processing of user feedback need to ensure user privacy and data security, and regularly evaluate and adjust the model to ensure that it reflects the latest user preferences and market changes.
[0016] In one embodiment of the present invention, the data preprocessing algorithm for automotive intelligent factors and the user feedback weight adjustment mechanism include the following steps:
[0017] A. Collect User Feedback: After each user evaluation (positive or negative), record the scores of the vehicle model features related to the evaluation. For unevaluated cases, it can be defaulted not to be processed;
[0018] B. Association Analysis: Analyze the associated features between positive and negative evaluations, which features are often associated with positive evaluations and which are often associated with negative evaluations, by calculating the correlation coefficient between the feature scores and positive or negative evaluations;
[0019] C. Weight Adjustment:
[0020] (1) Increase Weight: If a certain feature is highly correlated with positive evaluations, increase the weight of that feature;
[0021] (2) Decrease Weight: If a certain feature is highly correlated with negative evaluations, decrease the weight of that feature;
[0022] D. Normalization Processing: The adjusted weights need to be normalized again to ensure that the sum of all weights is 1;
[0023] E. Feedback Loop: Conducted periodically, continuously adjust the weights according to new user feedback to continuously optimize the model.
[0024] In one embodiment of the present invention, the data preprocessing algorithm for automotive intelligent factors includes the following code implementation:
[0025]
[0026]
[0027]
[0028] In one embodiment of the present invention, the preprocessing algorithm for automotive intelligent factor data further includes: an application layer, a support layer, and a data layer;
[0029] The data layer is established with real-time data for preliminary data processing and massive data storage;
[0030] The support layer is mainly based on data mining, data integration, and data analysis of the monitoring data in the data layer, and supports permission checking and control forwarding functions;
[0031] The application layer develops corresponding information monitoring according to actual business requirements, providing comprehensive and real-time information monitoring and auxiliary decision-making services for the monitoring subject.
[0032] In one embodiment of the present invention, the application layer includes a user operation interface, which displays information on data mining, data integration, and data analysis in the support layer, and the user operation interface is connected to the support layer through an I / O interface.
[0033] In one embodiment of the present invention, the data of the user operation interface is output after being processed by filtering or fuzzy separation, and after the support layer receives the data from the application layer, it performs data mining, data integration, and data analysis after user permission and identity verification processing.
[0034] In one embodiment of the present invention, the data is subjected to multi-source heterogeneous data standardization processing before massive data storage, and after data integration and data mining, application analysis is performed. The data is classified into vehicle data sets and comprehensive scheduling data sets. The comprehensive scheduling data set provides data support for the optimized scheduling of the entire automotive mobile Internet of Things. After the customer passes the user permission check, they can view the vehicle data set and / or comprehensive scheduling data set that can be viewed with the corresponding permissions.
[0035] In one embodiment of the present invention, the vehicle data set includes the position information of the vehicle, the status information of the passengers, drivers, and goods carried, and the surrounding environment information of the vehicle. The position information of the vehicle includes longitude, latitude, altitude, time, speed, direction, and satellite usage.
[0036] The beneficial effects of the technical solution of the present invention are as follows:
[0037] By cleaning the data, noise, error values, and missing values in the data can be removed. For example, in automotive intelligent factor data, if there are incorrect sensor readings (such as recording the response time of the intelligent driving assistance system as an unreasonable negative value), cleaning can ensure the accuracy of the data.
[0038] Data from different sources has different formats. For example, some represent numerical values in string form. After standardizing to numerical types, calculations can proceed smoothly, avoiding calculation errors caused by inconsistent data types. Only preprocessed data can be used for effective feature engineering and model construction. For example, for data related to the intelligent connectivity function of a car, if the appropriate numerical type cannot be ensured, it cannot accurately participate in calculations with other intelligent features (such as data related to the autonomous driving level), thus affecting the assessment of the overall intelligent level of the car.
[0039] By normalizing each feature, the feature data with different dimensions can be converted into comparable numerical values. For example, the accuracy of the intelligent voice recognition of a car is between 0 and 1, while the number of safety interventions in intelligent driving is a relatively large integer. After normalization, they have a fair basis for comparison in weight calculation. Information entropy reflects the amount of information contained in a feature. The smaller the information entropy of a feature, the greater weight it should have in evaluating the intelligence level of a car. For example, for the automatic emergency braking function of a car, whether it works properly (which can be regarded as a feature) has a great impact on the intelligent safety level of the car, and the information entropy is relatively low. Therefore, it should be given a higher weight in the overall intelligent assessment. This weight calculation based on the entropy weight method can objectively reflect the importance of each feature.
[0040] By collecting user evaluations (positive or negative) of a vehicle model, the user experience in actual use can be reflected. The user's perception of the intelligent functions of a car is different from the weights calculated solely based on technical features. For example, users attach more importance to the human-machine interaction experience in the intelligent cockpit, while in the initial technical weight calculation, this feature is underestimated. Adjusting the weights through user feedback can better meet the actual needs of users.
[0041] By renormalizing the weights to ensure that the sum is 1, the rationality of the weight system is guaranteed. The adjusted weights still meet the basic requirements of weight distribution, making the results consistent and reliable when recalculating intelligent scores and other operations. The evaluation results after weight adjustment based on user feedback can better reflect market demand, which is of great significance for car manufacturers to improve products and enhance market competitiveness.
[0042] Additional aspects and advantages of the present invention will be given in part in the following description, will become apparent in part from the following description, or will be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the following description of embodiments in conjunction with the accompanying drawings, wherein:
[0044] Figure 1Schematic diagram of the algorithm steps for the algorithm of preprocessing automotive intelligent factor data proposed in the embodiments of the present invention;
[0045] Figure 2 Schematic diagram of the steps for the user feedback weight adjustment mechanism proposed in the embodiments of the present invention. Detailed implementation manners
[0046] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions from beginning to end. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present invention, but should not be construed as limiting the present invention.
[0047] A preprocessing algorithm for automotive intelligent factor data according to an embodiment of the present invention will be described below with reference to the accompanying drawings.
[0048] As Figures 1 to 2 shown, an embodiment of the present invention provides a preprocessing algorithm for automotive intelligent factor data, including:
[0049] S1. Data preprocessing: Clean and standardize the data to ensure that all feature columns used for calculation are of numerical type;
[0050] S2. Calculate the initial weight (entropy weight method): Perform normalization processing on each feature, calculate the information entropy of each feature, and calculate the initial weight based on this;
[0051] S3. Calculate the intelligent score: Calculate the intelligent score of each vehicle model using the initial weight and feature data, and standardize the score to the range of 0-1 for comparison;
[0052] S4. Weight adjustment (based on user feedback): Collect user evaluations (positive or negative) of vehicle models, analyze the relationship between features and user evaluations, adjust the weights to reflect user preferences, and renormalize the weights to ensure that the sum is 1;
[0053] S5. Result output: Output the final score and the adjusted weights to an Excel table for further analysis and reporting.
[0054] In the specific application of the embodiments of the present invention, in the automotive intelligent factor data, there are some data points with input errors, such as non-numeric characters (mis-input) or missing numerical values (due to sensor failures or incomplete data collection) in a certain feature column. By cleaning the data, the data quality can be improved, making subsequent calculations more accurate. For missing values, filling methods can be used, such as mean filling (if the data distribution is relatively uniform), median filling, or model-based filling methods. For incorrect values, they need to be corrected according to the data range and logic or the data point can be directly deleted.
[0055] Subsequent calculation methods (such as the entropy weight method, etc.) are usually based on numerical operations. In the data of automotive intelligent factors, there are some data originally represented in text form, which need to be converted into numerical form. For example, some characteristics represent the intelligent level as "high, medium, low", and they need to be mapped to corresponding numerical values (such as 3, 2, 1). At the same time, standardization can also adjust data of different magnitudes to the same scale, avoiding the dominant position of characteristics with larger numerical values in the calculation.
[0056] Normalize each characteristic, aiming to map the values of each characteristic into the same interval. For example, convert the characteristic values into values between 0 and 1. This helps make each characteristic comparable when calculating the information entropy. In the data of automotive intelligent factors, different characteristics have different value ranges. For example, characteristics related to speed have relatively large values, while the activation frequencies of some intelligent assistance functions have relatively small values. Normalization processing can eliminate the influence of this magnitude difference on the weight calculation. Information entropy is an index used to measure the uncertainty of data. In this algorithm, by calculating the information entropy of each characteristic, the amount of information of this characteristic in the dataset can be reflected. The larger the information entropy, the smaller the amount of information contained in this characteristic; on the contrary, the smaller the information entropy, the larger the amount of information contained in this characteristic, and the greater the role in distinguishing the intelligent levels of different vehicle models. Calculate the initial weights according to the information entropy. Generally speaking, the smaller the information entropy of a characteristic, the larger its initial weight. The relative importance of each characteristic in evaluating the intelligent level of vehicles can be initially determined according to the characteristic distribution of the data itself.
[0057] Use the initial weights and characteristic data calculated previously to calculate the intelligent scores of each vehicle model. It is equivalent to performing a weighted sum of the various intelligent characteristics of each vehicle model according to their initial weights. For example, if a vehicle model performs well in an intelligent characteristic with a relatively large weight and performs mediocrely in other characteristics with relatively small weights, then its comprehensive intelligent score will be reasonably evaluated according to the weight distribution.
[0058] Standardize the scores to the range of 0 - 1 for easy comparison. The advantage of this is that it can intuitively compare the intelligent levels of different vehicle models. 0 represents the lowest intelligent level, and 1 represents the highest intelligent level. Through standardization, the intelligent scores of different vehicle models can be compared on the same scale, facilitating users or analysts to quickly understand the relative levels of each vehicle model in terms of intelligence. User evaluations are an important source of information reflecting the actual usage experience of vehicles. By collecting positive or negative comments from users, the subjective feelings of users towards the intelligent functions of vehicles can be obtained. For example, if users are particularly satisfied or dissatisfied with the automatic parking function (one of the intelligent characteristics) of a certain vehicle model, the evaluation information is of great significance for adjusting the weights.
[0059] Deeply study the correlation between each feature and user evaluations. If it is found that a certain feature generally performs better in models with positive evaluations and worse in models with negative evaluations, it indicates that this feature has a greater impact on user satisfaction. Conversely, if the performance of a certain feature shows little difference between models with positive and negative evaluations, it means that this feature has a relatively smaller impact on user satisfaction. Adjust the weights according to the relationship between the features and user evaluations. Increase the weights for features that have a greater impact on user satisfaction, and decrease the weights for features with a smaller impact. This can make the weights more in line with the actual concerns of users regarding the degree of vehicle intelligence, so that the final intelligent score can better reflect the needs and preferences of users.
[0060] After adjusting the weights, to ensure the rationality of the weight system, it is necessary to perform normalization again so that the sum of the weights of all features is 1. To maintain the proportional relationship of the weights and ensure that when calculating the intelligent score, the weights of each feature are still in a reasonable distribution state. Output the finally calculated intelligent score of each vehicle model and the adjusted weights to an Excel table. The Excel table has good readability and operability, which is convenient for further analysis and reporting. Analysts can perform operations such as sorting, filtering, and charting on the data in Excel to more intuitively display the degree of vehicle intelligence of different models and the weight distribution of each feature.
[0061] In one implementation, adjust the feature column names and paths in the code according to the actual data structure to ensure that data cleaning and preprocessing adapt to the characteristics of the dataset, especially handling missing values and outliers. The collection and processing of user feedback need to ensure user privacy and data security, and regularly evaluate and adjust the model to ensure that it reflects the latest user preferences and market changes.
[0062] In the specific application of the embodiments of the present invention, in the vehicle intelligence factor data, different datasets have different variable names to represent the same concept. For example, for the factor of vehicle speed, it is named "vehicle_speed" in one dataset and "car_speed" in another dataset. The algorithm needs to recognize these different naming methods. When obtaining a new dataset, by parsing the data dictionary or data structure, finding the actual column names corresponding to the vehicle intelligence factors defined in the algorithm can ensure that the algorithm can accurately operate on the data. Because if the column names do not match, subsequent data cleaning and processing steps will need to adjust the data columns to be processed. This kind of adjustment is dynamic and can adapt to data from different sources. For example, data collected from different automobile manufacturers or data obtained from different test scenarios (such as urban road tests, highway tests) have different column names.
[0063] Data has different file path structures when stored. When processing automotive intelligent factor data, if an algorithm needs to read a specific data file, such as a file storing automotive sensor data, it needs to adjust the reading path according to the actual storage structure. If in a development environment, the data is stored in a specific local folder structure (such as " / data / sensor_data / "), but in an actual application scenario, the data is stored in a specific path in cloud storage (such as "s3: / / car-data / sensor / "). The algorithm needs to be able to modify the data reading path according to the different deployment environments to ensure that the correct data can be obtained for preprocessing.
[0064] There are missing values in automotive intelligent factor data due to various reasons. For example, some sensors malfunction in specific environments. The algorithm first needs to identify which data is missing. It can be determined by counting the number of null values in each data column or a specific missing value identifier (such as -999 indicating missing). Then, select an appropriate method for handling missing values according to the characteristics of the dataset. If it is continuous automotive speed data, for a small number of missing values, the mean filling method can be used, that is, calculate the average of the non-missing data in this column to fill the missing values; if it is categorical data, such as the vehicle model category of a car, the most common category can be used to fill the missing values. Additionally, missing values can also be processed based on the correlation between data. For example, if the fuel consumption data of a car is missing, but the engine speed and vehicle speed data exist, a regression model can be established between these factors to predict the missing fuel consumption value.
[0065] Outliers in automotive intelligent factor data stem from sensor failures, special driving behaviors, or data acquisition errors. For example, the vehicle speed suddenly shows a very large value (such as 1000 km / h, which is obviously not in line with the actual vehicle driving speed). The algorithm needs to detect such outliers. Outliers can be judged by setting reasonable thresholds, such as setting the upper and lower limits of the vehicle speed according to the highest designed vehicle speed and the normal driving speed range. For the detected outliers, multiple processing methods can be adopted. If the outlier is an isolated point caused by data acquisition error, it can be directly deleted; if the outlier contains useful information (such as a short-term speeding in a special driving scenario), data smoothing techniques, such as the moving average method or the median filtering method, can be used to correct the outlier to make it more conform to the overall distribution of the data.
[0066] Different automotive intelligent factor datasets have different characteristics. For example, some datasets focus on the vehicle's powertrain data (such as engine torque, power, etc.), while others are more concerned with the vehicle's comfort factors (such as in-vehicle temperature, seat adjustment, etc.). The algorithm needs to adjust the data preprocessing strategy according to the focus of the dataset. For datasets dominated by powertrain data, more attention is paid to the processing accuracy of continuous data because small changes in powertrain data have a great impact on the performance evaluation of the vehicle; for comfort factor datasets, more categorical data (such as different seat heating levels) needs to be processed, and the relationship between the user's subjective feelings and the data needs to be considered.
[0067] In the automotive intelligent system, user feedback is a very important source of information. For example, users put forward improvement suggestions for the vehicle's autonomous driving assistance function, or feedback that the intelligent factors of the vehicle perform poorly under certain road conditions. User feedback can be collected through various channels, such as the in-vehicle interaction system (users can directly input text or voice feedback), and mobile applications (after connecting to the vehicle system, users can submit feedback on the mobile phone). During the collection process, it is necessary to classify and organize user feedback, for example, classify it according to the function modules of the feedback (such as intelligent navigation, intelligent safety system, etc.) and the severity of the feedback (such as seriously affecting driving safety, slightly affecting the use experience, etc.).
[0068] In one implementation, for the automotive intelligent factor data preprocessing algorithm, the user feedback weight adjustment mechanism includes the following steps:
[0069] A. Collect user feedback: After each user evaluation (positive or negative), record the vehicle model characteristics score related to the evaluation. If there is no evaluation, it can be defaulted not to be processed.
[0070] B. Association analysis: Analyze the association characteristics between positive and negative reviews, which characteristics are often associated with positive reviews and which are often associated with negative reviews, and complete it by calculating the correlation coefficient between the characteristic score and positive or negative reviews.
[0071] C. Weight adjustment:
[0072] (1) Increase weight: If a certain characteristic is highly correlated with positive reviews, increase the weight of this characteristic.
[0073] (2) Decrease weight: If a certain characteristic is highly correlated with negative reviews, decrease the weight of this characteristic.
[0074] D. Normalization processing: The adjusted weights need to be normalized again to ensure that the sum of all weights is 1.
[0075] E. Feedback loop: It is carried out periodically, and the weights are continuously adjusted according to new user feedback, so as to continuously optimize the model.
[0076] In the specific application of the embodiments of the present invention, in business scenarios related to automobiles, when a user makes an evaluation, the system will specifically record the scores of vehicle model characteristics related to the evaluation. For example, if a user gives a good review of the comfort of a certain vehicle, then the score of the comfort characteristic will be recorded. The scores here are based on a previously set evaluation system, such as a 0-10 point scale, etc. For cases where there is no evaluation, to simplify the processing flow, no special treatment is defaulted because it is difficult to determine the direction of its impact on the characteristic weights without user feedback.
[0077] The purpose of this step is to find the internal connection between vehicle model characteristics and user evaluations, which is achieved by calculating the correlation coefficient between the characteristic scores and good reviews or bad reviews. The correlation coefficient is a statistical measure used to measure the strength and direction of the linear relationship between two variables. For the correlation analysis between characteristics and good reviews, for example, if it is found that when the score of the safety configuration characteristic of a vehicle is high, it is often accompanied by good reviews from users, then it indicates that there is a positive correlation between the safety configuration characteristic and good reviews. Similarly, for bad reviews, if it is found that when the score of the fuel consumption characteristic of a vehicle is high (assuming high fuel consumption is a bad situation), it is often accompanied by bad reviews from users, then there is a positive correlation between the fuel consumption characteristic and bad reviews. This kind of correlation analysis can help us determine which characteristics have a greater impact on user evaluations, and whether it is a positive impact (associated with good reviews) or a negative impact (associated with bad reviews).
[0078] When it is determined through correlation analysis that a certain characteristic is highly correlated with good reviews, it is reasonable to increase the weight of this characteristic, as this characteristic makes a greater positive contribution to user satisfaction. For example, if the intelligent driving assistance function characteristic is highly correlated with good reviews, then in subsequent data analysis and model evaluation, increasing its weight can make this characteristic have a greater influence in the overall evaluation, which helps to highlight those characteristics that are attractive to users and can improve user satisfaction. On the contrary, if a certain characteristic is highly correlated with bad reviews, such as the high fuel consumption characteristic mentioned above, reducing its weight can reduce its negative impact on the overall evaluation. Because in the comprehensive evaluation of a vehicle, if a certain characteristic often causes user dissatisfaction, then appropriately reducing its weight in the model can make the evaluation result more in line with the actual feelings of users, and at the same time can guide automobile manufacturers to improve these characteristics that are prone to cause bad reviews.
[0079] After adjusting the weights, since the sum of all weights needs to remain 1, normalization can ensure that the weights of each characteristic are within a reasonable proportion range. For example, if there are three characteristics A, B, and C, with weights of 0.2, 0.3, and 0.5 respectively before adjustment, and they become 0.3, 0.2, and 0.5 after weight adjustment, normalization will re - distribute according to the new weight ratio so that they still add up to 1. This can ensure fair comparison and comprehensive evaluation among different characteristics, and avoid evaluation result deviation caused by the sum of weights not being equal to 1.
[0080] In a possible implementation manner, the data pre - processing algorithm for automotive intelligent factors includes the following code implementation:
[0081]
[0082]
[0083]
[0084]
[0085] In the specific application of the embodiments of the present invention, the weight - assignment method based on information entropy is used for multi - index comprehensive evaluation. By calculating the information entropy of each index, the weight of each index is determined, thus realizing multi - index decision - making. By calculating the information entropy of each index, the influence of subjective factors on weight assignment can be eliminated, and the objectivity and accuracy of the evaluation results can be improved. The lower the entropy value of each characteristic, the smaller the variability of the data, the fewer the suppliers of this component, and the more it tends to a mature market. For example, for intelligent driving chips, all are NVIDIA chips, then the weight occupied is lower, effectively pulling apart the scores of different vehicle models.
[0086] When designing the scoring rules, the importance of this configuration has been taken into account. For a simple configuration option of whether it is equipped or not, the weight occupied is small, such as the number of cameras, the number of lidar sensors, various screens, dash cams, the number of speakers, DMS cameras, application stores, power - assisted doors and other configurations. When we set the scoring rules, we will naturally make them more detailed, such as intelligent driving chips, lidar sensors, cockpit chips, etc.
[0087] In a more competitive market, the option distribution is more, the entropy is higher, and the weight is higher. For example, for intelligent driving algorithm suppliers, the monetization is more uneven, and the weight of the intelligent driving algorithm option is higher. If the competition is smaller and the configurations of most vehicle models tend to be the same (gradually becoming standard configurations), it means that the importance of this option is lower and the weight is lower.
[0088] In an implementation manner, the data pre - processing algorithm for automotive intelligent factors further includes: an application layer, a support layer, and a data layer;
[0089] The data layer is established with real-time data for preliminary data processing and massive data storage;
[0090] The support layer is mainly based on the monitoring data of the data layer for data mining, data integration, and data analysis, and also supports permission checking and control forwarding functions;
[0091] The application layer develops corresponding information monitoring according to actual business requirements, providing comprehensive and real-time information monitoring and auxiliary decision-making services for the monitoring entity.
[0092] The application layer includes a user operation interface. The user operation interface displays the information of data mining, data integration, and data analysis of the support layer, and the user operation interface is connected to the support layer through an I / O interface.
[0093] The data of the user operation interface is output after being separated by filtering or fuzzification processing. After receiving the data from the application layer, the support layer performs data mining, data integration, and data analysis after user permission and identity verification processing.
[0094] In the specific application of the embodiment of the present invention, for preliminary data processing: this step involves performing simple cleaning operations on the original real-time data collected from various sensors and control systems of the vehicle. For example, removing obvious error data (such as values outside the normal range, which are abnormally high or low values caused by sensor failures), preliminarily marking missing data, or performing simple format unification.
[0095] For massive data storage: Since a large amount of data will be continuously generated during the operation of the vehicle, such as the operating parameters of the engine (rotation speed, temperature, pressure, etc.) and vehicle driving state data (speed, acceleration, steering angle, etc.), sufficient storage capacity is required to save this data, which involves using large-capacity storage devices or distributed storage systems to ensure that the data will not be lost due to insufficient storage space and can be conveniently queried and called later.
[0096] The support layer works based on the monitoring data of the data layer, mainly focusing on data mining, data integration, and data analysis, while supporting permission checking and control forwarding functions. Data mining: Discover valuable information patterns or relationships from massive monitoring data. For example, mining the relationship between different driving habits (such as the frequency of hard acceleration and hard braking) and the wear of vehicle parts, or mining the change rules of various performance indicators of the vehicle under specific road conditions (such as mountain roads and highways).
[0097] Data integration: Uniformly organize the data from different data sources (which are the data of different subsystems inside the vehicle). For example, integrating the data of the chassis control system and the in-vehicle entertainment system to comprehensively understand the overall operating state of the vehicle.
[0098] Data analysis: Conduct in-depth analysis on the integrated data, such as statistical analysis to obtain the mean, variance, etc. of various indicators, or trend analysis to predict the service life of automotive components, etc.
[0099] Permission check: Ensure that only authorized users or system components can access and operate relevant data. For example, automotive maintenance personnel have the permission to view certain fault diagnosis data, while ordinary users can only view basic vehicle status data.
[0100] Control and forwarding function: Forward the processed data to the corresponding targets according to different application requirements and permission settings. For example, forward emergency fault data to the remote monitoring center of the automotive manufacturer, or forward data related to the user's driving habits to the intelligent driving assistance system of the vehicle.
[0101] The application layer develops corresponding information monitoring functions according to actual business requirements, aiming to provide comprehensive and real-time information monitoring and auxiliary decision-making services for the monitoring entity. Information monitoring: Develop corresponding functions for different monitoring requirements of the vehicle. For example, monitor the fuel consumption of the vehicle, and display information such as the remaining fuel quantity and fuel consumption rate in real time; or monitor the safety status of the vehicle, such as whether the tire pressure is normal and whether there are potential faults in the braking system.
[0102] Auxiliary decision-making service: Provide support for relevant decisions based on the monitored information. For example, provide the driver with the best driving route suggestion according to road conditions and vehicle performance data; or provide the maintenance personnel with a decision-making basis for whether to replace parts in advance according to the wear condition of automotive components.
[0103] User operation interface: As part of the application layer, the user operation interface displays the result information of data mining, data integration, and data analysis in the support layer. For example, it is presented to the user in an intuitive chart (such as a bar chart showing the fuel consumption comparison in different time periods) or in numerical form (such as an accurate vehicle speed value). It is connected to the support layer through the I / O interface to achieve data interaction and transmission.
[0104] The data of the user operation interface in the application layer will undergo filtering or obfuscation separation processing before output. Filtering processing is to remove some data that is unimportant for the current business requirements or affects system performance. For example, when displaying the basic driving status of the vehicle, filter out some overly detailed sensor debugging data. Obfuscation separation processing is used to protect user privacy or data security. For example, when providing some data to a third-party application, obfuscate sensitive information such as user identity.
[0105] After the support layer receives the data transmitted from the application layer, it first performs user permission and identity authentication processing. Only the user data that passes the verification can be used for subsequent data mining, data integration, and data analysis operations. This ensures the security and compliance of the data and prevents unauthorized data access and abuse.
[0106] In one implementation, after the data is processed by multi-source heterogeneous data standardization, it is stored in a large amount of data. After data integration and data mining, the data is subjected to application analysis, and the data is classified into a vehicle data set and a comprehensive scheduling data set. The comprehensive scheduling data set provides data support for the optimal scheduling of the entire vehicle mobile Internet of Things. After the customer passes the user permission check, they can view the vehicle data set and / or comprehensive scheduling data set that can be viewed with the corresponding permissions.
[0107] The vehicle data set includes the location information of the vehicle, the status information of the passengers, drivers, and goods carried, and the surrounding environment information of the vehicle. The location information of the vehicle includes longitude, latitude, altitude, time, speed, direction, and satellite usage.
[0108] In the specific application of the embodiment of the present invention, in the field of vehicle intelligence, the data sources are extensive and have multi-source heterogeneity. For example, the location information of the vehicle (longitude, latitude, altitude, time, speed, direction, satellite usage) is a structured numerical data, while the status information of the passengers, drivers, and goods carried and the surrounding environment information of the vehicle include various forms such as text, images, or binary data of sensors. The data from different sources differ in format, range, accuracy, etc. For example, the speed is in kilometers per hour, and the value range is from 0 to the maximum speed of the vehicle; while the satellite usage is a discrete status value (such as the number of satellites connected).
[0109] The purpose is to convert these different types of data into a unified format and scale for subsequent storage and processing. For numerical data, a normalization method is adopted to map it to a specific interval (such as [0,1] or [-1,1]). For example, for speed data, if the maximum speed of the vehicle is 200 kilometers per hour, then for a vehicle with a speed of 100 kilometers per hour, the normalized value is 0.5 (assuming it is mapped to the [0,1] interval). For non-numerical data, a coding method is adopted for standardization. For example, if the surrounding environment information of the vehicle includes states such as "sunny", "cloudy", "rainy", etc., "sunny" can be coded as 1, "cloudy" can be coded as 2, "rainy" can be coded as 3, etc.
[0110] The amount of data after standardization is huge because new data is continuously generated during the operation of the vehicle. For example, data such as location information and vehicle status are updated every second. Mass data storage is to save this data for subsequent analysis and application. A distributed storage system is adopted, such as HDFS (Hadoop Distributed FileSystem) based on Hadoop. This system can store data dispersedly on multiple nodes, improving the reliability and scalability of storage.
[0111] At the same time, to improve the query efficiency of data, data warehouse technology is adopted to organize and manage the stored data. For example, data is partitioned and stored according to dimensions such as time and vehicle ID.
[0112] Due to different data sources, it is necessary to integrate the massive data stored. For example, associate the location information of the vehicle with the information of the passengers and drivers on board. If the vehicle location information is stored in one data table and the passenger and driver information is stored in another data table, it is necessary to connect the two through key information such as vehicle ID to form a comprehensive data set containing more information.
[0113] Data mining techniques are used to discover potential patterns and regularities from the integrated data. For example, by analyzing the location information and time information of the vehicle, the travel patterns of the vehicle can be mined, such as which time periods are the peak travel periods of the vehicle; by analyzing the status information of the goods and the surrounding environment information of the vehicle, the risk factors during the goods transportation process can be mined, such as the probability of goods damage in a specific environment. The data is classified into vehicle data sets and comprehensive scheduling data sets according to the nature and use of the data. The vehicle data set mainly focuses on the vehicle itself and the information directly related to it, such as the location of the vehicle, the people and goods carried, etc., aiming to provide data support for vehicle operation management, safety monitoring, etc. The comprehensive scheduling data set focuses on the optimized scheduling of the entire vehicle mobile Internet of Things. It contains more extensive data, which are some macroscopic information after integration and mining, such as the distribution of vehicles in different regions, the overall traffic flow, etc., providing data basis for optimized scheduling.
[0114] In terms of vehicle operation, the vehicle data set can be used for real-time monitoring of the vehicle. For example, according to the location information and speed information of the vehicle, it can be judged whether the vehicle is speeding or deviating from the predetermined route.
[0115] The comprehensive scheduling data set can be used for traffic flow regulation. For example, in urban traffic, if it is found that the vehicle density in a certain area is too high, the dispatching command center can send guiding information to the vehicles in that area to guide the vehicles to divert and relieve traffic congestion.
[0116] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.
[0117] The above describes the present invention and its embodiments. Such description is not restrictive. What is shown in the drawings is only one of the embodiments of the present invention, and the actual structure is not limited thereto. Generally speaking, if those of ordinary skill in the art are inspired by it and, without departing from the gist of the present invention, design similar structural modes and embodiments to this technical solution without creative efforts, they shall fall within the protection scope of the present invention.
Claims
1. A data preprocessing algorithm for automobile intelligentization factors, characterized in that: include: S1. Data preprocessing: clean and standardize the data to ensure that all feature columns used for calculation are of numerical type; S2. Calculate initial weight (entropy weight method): normalize each feature, calculate the information entropy of each feature and calculate the initial weight based on it; S3. Calculate smart score: Calculate the smart score of each vehicle model using the initial weights and feature data, and normalize the score to the range of 0-1 for comparison; S4, weight adjustment (based on user feedback): collect user reviews of the vehicle model (good or bad), analyze the relationship between features and user reviews, adjust weights to reflect user preferences, and renormalize weights to ensure that the sum is 1; S5. Result output: Output the final score and adjusted weights to an Excel spreadsheet for further analysis and reporting.
2. The automobile intelligent factor data preprocessing algorithm according to claim 1 is characterized in that: Adjust the feature column names and paths in the code according to the actual data structure to ensure that data cleaning and preprocessing are adapted to the characteristics of the dataset, especially handling missing values and outliers.
3. The automobile intelligent factor data preprocessing algorithm according to claim 1 is characterized in that: The collection and processing of user feedback needs to ensure user privacy and data security, and regularly evaluate and adjust the model to ensure that it reflects the latest user preferences and market changes.
4. The automobile intelligent factor data preprocessing algorithm according to claim 3 is characterized in that: The user feedback weight adjustment mechanism includes the following steps: A. Collect user feedback: After each user evaluation (good or bad), record the vehicle model feature score related to the evaluation. If there is no evaluation, it can be defaulted and not processed; B. Correlation analysis: Analyze the correlation characteristics between positive and negative reviews, which characteristics are often associated with positive reviews and which characteristics are often associated with negative reviews, by calculating the correlation coefficient between the characteristic score and the positive or negative reviews; C. Weight adjustment: (1) Increase weight: If a feature is highly correlated with positive reviews, increase the weight of the feature; (2) Reduce weight: If a feature is highly correlated with negative reviews, reduce the weight of the feature; D. Normalization: The adjusted weights need to be normalized again to ensure that the sum of all weights is 1; E. Feedback loop: This is done periodically to continuously adjust weights based on new user feedback, thereby continuously optimizing the model.
5. The automobile intelligent factor data preprocessing algorithm according to claim 1 is characterized in that: Includes the following code implementation: 。 6. The automobile intelligent factor data preprocessing algorithm according to claim 1 is characterized in that: Also includes: Application layer, support layer and data layer; The data layer is established based on real-time data, with preliminary data processing and mass data storage; The support layer is mainly based on data mining, data integration and data analysis based on the monitoring data of the data layer, and supports authority checking and control forwarding functions; The application layer develops corresponding information monitoring according to actual business needs, providing comprehensive and real-time information monitoring and auxiliary decision-making services for the monitoring subject.
7. The automobile intelligent factor data preprocessing algorithm according to claim 1 is characterized in that: The application layer includes a user operation interface, which displays information on data mining, data integration, and data analysis of the support layer. The user operation interface is connected to the support layer through an I / O interface.
8. The automobile intelligent factor data preprocessing algorithm according to claim 1 is characterized in that: The data from the user interface is output after filtering or fuzzy separation processing. After the support layer receives the data from the application layer, it processes the user permissions and identity authentication and then performs data mining, data integration, and data analysis.
9. The automobile intelligent factor data preprocessing algorithm according to claim 1 is characterized in that: The data is stored in massive amounts after being standardized by multi-dimensional heterogeneous data processing. The data is analyzed for application after data integration and data mining. The data is classified into vehicle data sets and comprehensive scheduling data sets. The comprehensive scheduling data sets provide data support for the optimized scheduling of the entire automotive mobile Internet of Things. After the user authority check, the customer can view the vehicle data sets and / or comprehensive scheduling data sets that can be viewed by the corresponding authority.
10. The automobile intelligent factor data preprocessing algorithm according to claim 1 is characterized in that: The vehicle dataset contains the vehicle's location information, the status of the passengers, driver, and cargo, and the vehicle's surrounding environment information. The vehicle's location information includes longitude, latitude, altitude, time, speed, direction, and satellite usage.