Automobile industry data risk assessment method and system based on big data
By adopting the choice of diversified data sources and machine learning algorithms in the data risk assessment of the automotive industry, the problem of single data sources and inaccurate evaluation results in the prior art is solved, the accuracy and reliability of risk assessment are improved, and the safety of automobiles is ensured.
Patent Information
- Application Number
- CN202510104137.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing automotive industry data risk assessment methods rely on a single data source, which have biases and limitations, resulting in inaccurate assessment results, increasing automotive safety threats, and inability to select appropriate algorithms based on data characteristics, introduce evaluation bias, and reduce the reliability of risk assessment.
The automotive industry data risk assessment method is adopted based on big data. By obtaining data from multiple data sources, data verification and preprocessing are carried out, features are extracted, and appropriate machine learning algorithms are selected based on features, risk assessment models are constructed, multi-dimensional risk assessments are carried out, and data risk management strategies are formulated.
Through diverse data sources and machine learning algorithm selection, the bias and limitations of evaluation results are reduced, the accuracy and reliability of evaluation results are improved, and potential automotive safety threats are avoided.
Smart Images

Figure CN120046977A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data risk assessment, and particularly to a method and system for data risk assessment in the automotive industry based on big data. Background Art
[0002] With the rapid development of Internet and big data technologies, the amount of data in the automotive industry has increased sharply.
[0003] However, the existing methods for data risk assessment in the automotive industry have a single data source and mainly rely on public data. These data sources may have certain biases and limitations, resulting in inaccurate assessment results, increasing potential automotive safety threats. At the same time, it is impossible to select appropriate algorithms according to data characteristics, which will introduce assessment biases and reduce the reliability of data risk assessment in the automotive industry.
[0004] Therefore, the present invention proposes a method and system for data risk assessment in the automotive industry based on big data. Summary of the Invention
[0005] The present invention provides a method and system for data risk assessment in the automotive industry based on big data to solve the defects in the prior art that the existing methods for data risk assessment in the automotive industry have a single data source and mainly rely on public data. These data sources may have certain biases and limitations, resulting in inaccurate assessment results, increasing potential automotive safety threats. At the same time, it is impossible to select appropriate algorithms according to data characteristics, which will introduce assessment biases and reduce the reliability of data risk assessment in the automotive industry.
[0006] On the one hand, the present invention provides a method for data risk assessment in the automotive industry based on big data, including: Step 1: Obtain various types of data related to the automotive industry, and determine the accuracy and comprehensiveness of various types of data according to verification rules and different inspection techniques; Step 2: Preprocess the various types of data according to data desensitization technology, and extract different features according to the preprocessed data; Step 3: Obtain the complexity of the problem and the characteristics of the data according to the different features, and obtain the corresponding machine learning algorithm according to the complexity of the problem and the characteristics of the data; Step 4: Define data risk assessment indicators, establish a data risk assessment model for the automotive industry in combination with a machine learning algorithm, and train and optimize the model according to historical data and experimental data from multiple aspects; Step 5: Conduct risk assessment on automotive industry data according to the trained and optimized model, and formulate corresponding data risk management strategies according to the assessment results.
[0007] A method for risk assessment of automotive industry data based on big data provided by the present invention, which obtains various types of data related to the automotive industry and determines the accuracy and comprehensiveness of various types of data according to verification rules and different inspection techniques, including: Obtain various types of data related to automobiles according to road monitoring and in-vehicle sensors; Encrypt the various types of data according to encryption technology, and perform sharding processing on the encrypted data; Obtain the anonymity and decentralization of the data according to the sharded data in combination with anonymization technology; Formulate data quality standards, and determine corresponding data verification rules according to the data quality standards; Determine the accuracy and comprehensiveness of various types of data according to the data verification rules and different inspection techniques.
[0008] A method for risk assessment of automotive industry data based on big data provided by the present invention, which preprocesses the various types of data according to data desensitization technology, and extracts different features according to the preprocessed data, including: Process various types of data according to data desensitization technology, and delete and replace data points containing personal privacy information; Perform normalization and binning preprocessing on the data after desensitization processing; Extract different types of features from the preprocessed data according to the preprocessed data in combination with business requirements and scenarios by using corresponding feature extraction methods.
[0009] A method for risk assessment of automotive industry data based on big data provided by the present invention, which obtains the complexity of the problem and the characteristics of the data according to the different features, and obtains corresponding machine learning algorithms according to the complexity of the problem and the characteristics of the data, including: Determine the data type of the data according to the different features, obtain the corresponding processing method according to the data type, and determine the complexity of the problem according to the processing result; Obtain the relevance between the data, obtain the hidden patterns and trends between the data according to the relevance, and determine the characteristics of the data according to the patterns and trends; Determine the scale of the data and the form of the problem according to the complexity of the problem and the characteristics of the data, and obtain the corresponding machine learning algorithm according to the scale of the data and the form of the problem.
[0010] A method for risk assessment of automotive industry data based on big data provided by the present invention, which defines data risk assessment indicators, establishes an automotive industry data risk assessment model in combination with machine learning algorithms, and trains and optimizes the model according to historical data and experimental data and other aspects of data, including: Define data risk assessment indicators according to the business requirements of the automotive industry, and establish a data risk assessment model for the automotive industry in combination with machine learning algorithms; Obtain historical data and experimental data from different data sources, and perform data fusion; Based on the fused data and combined with cross-validation, conduct preliminary training and optimization of the model; Select corresponding key components according to the preliminarily trained and optimized model, and set different hyperparameters to retrain and optimize the model.
[0011] According to a method for data risk assessment in the automotive industry based on big data provided by the present invention, conduct risk assessment on automotive industry data according to the trained and optimized model, and formulate corresponding data risk management strategies according to the assessment results, including: Conduct multi-dimensional risk assessment on automotive industry data according to the trained and optimized model, and obtain the data risk level and risk influencing factors according to the assessment results; Obtain corresponding series of risk vehicle information according to the data risk level, and determine the risk influence scope according to the risk influencing factors; Formulate corresponding data risk management strategies according to the series of risk vehicle information and the risk influence scope.
[0012] According to a method for data risk assessment in the automotive industry based on big data provided by the present invention, after obtaining various data related to the automotive industry, it further includes: Obtain the data attributes of various data, determine the data characteristics of various data based on the data attributes, and determine the public opinion health status of various data according to the data characteristics; Judge the commercial value of various data based on the public opinion health status, eliminate the first type of data with low commercial value, and retain the second type of data; Obtain the industry domain matching entity relationships of the second type of data, and determine the main entity attributes, sub-entity attributes, behavior entity attributes, and business entity attributes of the second type of data according to the industry domain matching entity relationships; Determine the entity logical relationships of the second type of data based on the main entity attributes, sub-entity attributes, behavior entity attributes, and business entity attributes; Determine the data field association parameters according to the entity logical relationships, and generate a data cleaning dictionary according to the data field association parameters; Use the data cleaning dictionary to clean each second type of data, and obtain the cleaned second type of data; Conduct query intention matching on the cleaned second type of data, and determine the query vector of the second type of data according to the matching results; Determine the data relationship path of the second type of data according to the query vector, and construct a domain knowledge graph of the second type of data according to the data relationship path; Determine the data inspection direction and verification direction of the second type of data according to the domain index atlas, and select verification rules and inspection technical means based on the data inspection direction and verification direction.
[0013] A big data-based automotive industry data risk assessment system provided by the present invention includes: Determination module: Obtain various types of data related to the automotive industry, and determine the accuracy and comprehensiveness of various types of data according to verification rules and different inspection technical means; Extraction module: Preprocess the various types of data according to data desensitization technology, and extract different features according to the preprocessed data; Acquisition module: Obtain the complexity of the problem and the characteristics of the data according to the different features, and obtain the corresponding machine learning algorithm according to the complexity of the problem and the characteristics of the data; Training and optimization module: Define data risk assessment indicators, establish an automotive industry data risk assessment model in combination with machine learning algorithms, and train and optimize the model according to historical data and experimental data from multiple aspects; Risk assessment module: Conduct risk assessment on automotive industry data according to the trained and optimized model, and formulate corresponding data risk management strategies according to the assessment results.
[0014] Compared with the prior art, the beneficial effects of the present application are as follows: By processing various types of data in the automotive industry, obtaining data features, and constructing a risk assessment model according to machine learning algorithms, it can ensure the diversification of data sources for the automotive industry data risk assessment method, reduce certain biases and limitations, improve the accuracy of assessment results, avoid potential automotive safety threats, and at the same time, appropriate algorithms can be selected according to data features, without introducing assessment biases, and improve the reliability of risk assessment of automotive industry data. Description of the Drawings
[0015] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0016] Figure 1 is a schematic flowchart of a big data-based automotive industry data risk assessment method provided by an embodiment of the present invention; Figure 2 is a schematic structural diagram of a big data-based automotive industry data risk assessment system provided by an embodiment of the present invention. Detailed Embodiments
[0017] To make the objectives, technical solutions and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts belong to the scope of protection of the present invention.
[0018] Embodiment 1: A method for risk assessment of automotive industry data based on big data provided by an embodiment of the present invention, as Figure 1 shown, the method mainly includes the following steps: Step 1: Obtain various types of data related to the automotive industry, and determine the accuracy and comprehensiveness of various types of data according to verification rules and different inspection techniques; Step 2: Preprocess the various types of data according to data desensitization technology, and extract different features according to the preprocessed data; Step 3: Obtain the complexity of the problem and the characteristics of the data according to the different features, and obtain the corresponding machine learning algorithm according to the complexity of the problem and the characteristics of the data; Step 4: Define data risk assessment indicators, and establish an automotive industry data risk assessment model in combination with a machine learning algorithm, and train and optimize the model according to historical data and experimental data and other aspects of data; Step 5: Conduct risk assessment on automotive industry data according to the trained and optimized model, and formulate corresponding data risk management strategies according to the assessment results.
[0019] In this embodiment, the various types of data include: vehicle operation data, driver behavior data, vehicle sales data, vehicle usage behavior data, and traffic accident data.
[0020] In this embodiment, the data verification rule is a series of regulations or conditions for verifying, inspecting and constraining input data to ensure that the data meets the expected format, range and requirements.
[0021] In this embodiment, the different inspection techniques include: validity check, range check.
[0022] In this embodiment, the data risk assessment indicator is a quantitative indicator used to measure data quality, security and reliability, such as: Data accuracy indicator: It refers to the correctness and consistency of data. For example, whether the data conforms to the actual value, whether there are errors or omissions in the data, including: data accuracy rate, data precision, data consistency index.
[0023] Data integrity metrics: Refer to the integrity and comprehensiveness of data, such as whether the data covers all possible ranges and whether the data covers all key nodes, including: data coverage rate, data integrity ratio, and data integrity check points.
[0024] Data timeliness metrics: Refer to the timeliness and update frequency of data, such as whether the data reflects the latest situation in a timely manner and whether the data lags behind the actual situation, including: data response speed, data refresh frequency, and data periodicity.
[0025] The beneficial effects of the above technical solutions are: By processing various types of data in the automotive industry, obtaining data characteristics, and constructing a risk assessment model based on machine learning algorithms, it can ensure the diversification of data sources for the data risk assessment method in the automotive industry, reduce certain biases and limitations, improve the accuracy of assessment results, avoid potential automotive safety threats, and at the same time, appropriate algorithms can be selected according to data characteristics without introducing assessment biases, improving the reliability of the risk assessment of automotive industry data.
[0026] Embodiment 2: Based on Embodiment 1, the embodiment of the present invention obtains various types of data related to the automotive industry, and determines the accuracy and comprehensiveness of various types of data according to verification rules and different inspection technical means, including: Obtain various types of data related to automobiles according to road monitoring and in-vehicle sensors; Perform encryption processing on the various types of data according to encryption technology, and perform sharding processing on the encrypted data; Obtain the anonymity and decentralization of the data according to the sharded data in combination with anonymization technology; Formulate data quality standards, and determine corresponding data verification rules according to the data quality standards; Determine the accuracy and comprehensiveness of various types of data according to the data verification rules and different inspection technical means.
[0027] In this embodiment, sharding processing is a machine learning preprocessing technology, and its purpose is to divide the data set into multiple non-overlapping parts.
[0028] In this embodiment, anonymization technology is a technology for protecting personal privacy, which can make an individual's identity unable to be accurately identified by encrypting, obfuscating, or replacing personal identity information, etc.
[0029] In this embodiment, the anonymity of data refers to whether the identity information of an individual can be accurately identified during the collection, storage, and use of data.
[0030] In this embodiment, the decentralization of data refers to storing data in multiple locations to avoid single-point failures and data leakage problems.
[0031] In this embodiment, the data verification rules are a series of regulations or conditions for verifying, checking, and constraining input data to ensure that the data conforms to the expected format, range, and requirements.
[0032] The beneficial effects of the above technical solution are as follows: various types of data related to the automotive industry are obtained, and the accuracy of various data algorithms is determined according to the verification rules and different inspection techniques, so as to obtain accurate data on the operating conditions of the automotive industry, improve the accuracy and comprehensiveness of the data, and at the same time, improve the scientific nature of decision-making in the automotive industry.
[0033] Embodiment 3: Based on Embodiment 2, in this embodiment of the present invention, the various types of data are preprocessed according to the data desensitization technology, and different features are extracted from the preprocessed data, including: The various types of data are processed according to the data desensitization technology, and the data points containing personal privacy information are deleted and replaced; The data after desensitization processing is subjected to normalization and binning preprocessing; According to the preprocessed data, combined with business requirements and scenarios, the corresponding feature extraction methods are used to extract different types of features from the preprocessed data.
[0034] In this embodiment, the data desensitization technology is a technology that combines original personal or organizational data with privacy information to protect sensitive personal information from being leaked or misused. During the data desensitization process, the original identity information is replaced with an anonymous identifier, so that the data can be safely shared and used while still retaining its value. For example: randomized encoding, anonymization algorithms.
[0035] In this embodiment, the data points containing personal privacy information refer to the data involving personal privacy information during the data processing, analysis, or use process. For example: the owner of the vehicle, the information of the vehicle owner.
[0036] In this embodiment, the binning preprocessing is a machine learning preprocessing technology that divides the data set into multiple subsets of equal size, and each subset is called a "bucket".
[0037] The beneficial effects of the above technical solution are as follows: the various types of data are preprocessed according to the data desensitization technology, and sensitive information can be replaced with secure data without compromising the authenticity of the data to protect the information from being leaked. Further, different data features are extracted from the preprocessed data, which can better understand the data and lay a foundation for selecting appropriate algorithms later.
[0038] Embodiment 4: Based on Embodiment 3, an embodiment of the present invention obtains the complexity of the problem and the characteristics of the data according to the different features, and obtains the corresponding machine learning algorithm according to the complexity of the problem and the characteristics of the data, including: Determine the data type of the data according to the different features, obtain the corresponding processing method according to the data type, and determine the complexity of the problem according to the processing result; Obtain the correlation between the data, obtain the hidden patterns and trends between the data according to the correlation, and determine the characteristics of the data according to the patterns and trends; Determine the scale of the data and the form of the problem according to the complexity of the problem and the characteristics of the data, and obtain the corresponding machine learning algorithm according to the scale of the data and the form of the problem.
[0039] In this embodiment, the data types include: structured data, unstructured data, or semi-structured data.
[0040] In this embodiment, the processing method can be: structured data can usually be easily processed through SQL queries and statistical analysis, while unstructured data (such as log files) is processed using natural language processing techniques.
[0041] In this embodiment, the hidden patterns and trends between the data refer to the relationships, laws, and patterns between the data obtained through data analysis.
[0042] The beneficial effects of the above technical solutions are: obtaining the patterns and trends between the data through data features, thereby determining the characteristics of the data, and combining with the problem form to obtain the corresponding machine learning algorithm, which can automatically complete the analysis and solution of complex problems without manual processing, improve the evaluation accuracy of the model, and ensure the reliability of the evaluation of automotive industry data.
[0043] Embodiment 5: Based on Embodiment 4, an embodiment of the present invention defines data risk assessment indicators, and establishes a data risk assessment model for the automotive industry in combination with a machine learning algorithm, and trains and optimizes the model according to historical data and experimental data from multiple aspects, including: Define data risk assessment indicators according to the business requirements of the automotive industry, and establish a data risk assessment model for the automotive industry in combination with a machine learning algorithm; Obtain historical data and experimental data from multiple aspects from different data sources, and perform data fusion; Perform preliminary training and optimization on the model according to the fused data combined with cross-validation; Select the corresponding key components according to the preliminarily trained and optimized model, and set different hyperparameters to perform re-training and optimization on the model.
[0044] In this embodiment, the data risk assessment indicators are quantitative indicators used to measure data quality, security, and reliability. For example: Data accuracy indicators: Refer to the correctness and consistency of data. For example, whether the data conforms to the actual value, whether there are errors or omissions in the data, including: data accuracy rate, data precision, data consistency index.
[0045] Data integrity indicators: Refer to the integrity and comprehensiveness of data. For example, whether the data covers all possible ranges, whether the data covers all key nodes, including: data coverage rate, data integrity ratio, data integrity check points.
[0046] Data timeliness indicators: Refer to the timeliness and update frequency of data. For example, whether the data reflects the latest situation in a timely manner, whether the data lags behind the actual situation, including: data response speed, data refresh frequency, data periodicity.
[0047] In this embodiment, the key components include: loss function and optimizer.
[0048] In this embodiment, training and optimization can be: increasing the number of layers of the neural network or adjusting the hyperparameters of the activation function.
[0049] The beneficial effects of the above technical solutions are: defining data risk assessment indicators, establishing a data risk assessment model for the automotive industry in combination with machine learning algorithms, training and optimizing the model, and optimizing the model again in combination with the corresponding components, which can better adjust the model to adapt to different data sets and workloads, and improve the performance and generalization ability of the model.
[0050] Embodiment 6: Based on Embodiment 5, the embodiment of the present invention conducts a risk assessment on automotive industry data according to the trained and optimized model, and formulates corresponding data risk management strategies according to the assessment results, including: Conduct a multi-dimensional risk assessment on automotive industry data according to the trained and optimized model, and obtain the data risk level and risk influencing factors according to the assessment results; Obtain the corresponding series of risk vehicle information according to the data risk level, and determine the risk impact scope according to the risk influencing factors; Formulate corresponding data risk management strategies according to the series of risk vehicle information and risk impact scope.
[0051] In this embodiment, conducting a multi-dimensional risk assessment on automotive industry data means comprehensively and systematically analyzing various factors within the automotive industry, identifying potential risk points, and scoring according to the nature and impact degree of the risks, so as to obtain an assessment result on the overall risk level of the automotive industry, including: technical risks, market risks, regulatory risks.
[0052] In this embodiment, the series of risk vehicle information refers to the specific vehicle information of a certain brand or a certain type of vehicle with a certain type of risk.
[0053] In this embodiment, the data risk management strategy refers to a series of management and control measures taken to reduce the negative impacts that may be brought by data risks.
[0054] The beneficial effects of the above technical solution are as follows: Risk assessment of the automotive industry data according to the trained and optimized model helps to identify potential data security risks and compliance issues. Further, formulating corresponding data risk management strategies according to the assessment results can better and more efficiently manage risk data and improve the security and privacy of data.
[0055] Embodiment 7: Based on Embodiment 6, after obtaining various types of data related to the automotive industry, this embodiment of the present invention further includes: Obtain the data attributes of various types of data, determine the data characteristics of various types of data based on the data attributes, and determine the public opinion health status of various types of data according to the data characteristics; Judge the commercial value of various types of data based on the public opinion health status, eliminate the first type of data with low commercial value, and retain the second type of data; Obtain the industry domain matching entity relationships of the second type of data, and determine the main entity attributes, sub-entity attributes, behavior entity attributes, and business entity attributes of the second type of data according to the industry domain matching entity relationships; Determine the entity logical relationships of the second type of data based on the main entity attributes, sub-entity attributes, behavior entity attributes, and business entity attributes; Determine the data field association parameters according to the entity logical relationships, and generate a data cleaning dictionary according to the data field association parameters; Use the data cleaning dictionary to clean each second type of data to obtain the cleaned second type of data; Perform query intention matching on the cleaned second type of data, and determine the query vector of the second type of data according to the matching results; Determine the data relationship path of the second type of data according to the query vector, and construct a domain knowledge graph of the second type of data according to the data relationship path; Determine the data inspection direction and verification direction of the second type of data according to the domain index graph, and select verification rules and inspection technical means based on the data inspection direction and verification direction.
[0056] In this embodiment, the data attributes of various types of data include: Vehicle information: including information such as vehicle type, brand, model, identification code, engine number, frame number, and vehicle condition.
[0057] User information: including information such as the user's name, gender, age, occupation, address, contact information, etc.
[0058] Sales information: including information such as the sales date of the car, purchase location, purchase price, sales volume, sales amount, etc.
[0059] In this embodiment, the data characteristics of various types of data include: Vehicle information: It has characteristics such as uniqueness, traceability, immutability, and persistence. For example, the vehicle's unique identification code and vehicle identification number are both unique identifiers, and various vehicle information can be obtained by querying the database.
[0060] User information: Usually has characteristics such as diversity, variability, dynamics, and hierarchy. For example, information such as the user's gender, age, and occupation can reflect the user's personality and needs, and these information are constantly changing.
[0061] Sales information: Usually has characteristics such as timeliness, regionality, seasonality, and trendiness. For example, the date and time period of car sales information can reflect the seasonal changes in the market, and the differences in sales regions can also reflect different market demands.
[0062] In this embodiment, the public opinion health status of various types of data refers to the magnitude of the influence of these data on the public and society and the amount of positive and negative evaluations.
[0063] In this embodiment, the main entity attributes include: Vehicle brand and model: Reflects the vehicle's brand and model, which helps to distinguish different cars and position products in the market.
[0064] Vehicle identification code: It is the code that uniquely identifies each car and can be used to track information such as the vehicle's historical records and maintenance conditions.
[0065] In this embodiment, the sub-entity attributes include: Vehicle size and weight: including parameters such as the vehicle's length, width, height, and curb weight, which can affect aspects such as the vehicle's fuel consumption and acceleration performance.
[0066] Engine information and parameters: including important parameters such as the engine type, displacement, number of cylinders, working principle, fuel injection method, ignition system, etc.
[0067] Transmission system information: including information on important components such as the type of transmission, number of gears, drive method, clutch, gear lever / brake pedal, universal joint, etc.
[0068] In this embodiment, the behavioral entity attributes usually refer to behaviors or actions related to the use and maintenance of the car, such as: Vehicle usage: including parameters such as mileage, average speed, maximum speed, acceleration, deceleration rate, number of hard brakes, daily mileage, etc.
[0069] Driver behavior: including aspects such as the driver's gender, age, driving experience, driving habits, and driving routes.
[0070] In this embodiment, the business entity attributes refer to business information related to the production and sales of automobile enterprises, such as: Production plan: including production volume, raw material procurement plan, and production line plan.
[0071] Inventory management: including raw material inventory, semi-finished product inventory, finished product inventory, and spare parts inventory.
[0072] In this embodiment, the entity logical relationship of data refers to the associative and interdependent relationships existing between data, such as: Hierarchical relationship: Automobile products can be regarded as a hierarchical structure, such as sedans, SUVs, MPVs, etc. This relationship can be represented by a tree diagram, where each node represents a product category, and the child nodes of the node represent different levels of sub-classifications.
[0073] Time relationship: The production, sales, and use processes of automobiles are all affected by time factors. For example, the sales record of a car can be associated with the production time and service life of the car.
[0074] Spatial relationship: Automobile enterprises have different sales points and production bases in different regions.
[0075] In this embodiment, the association parameters of data fields refer to the dependence and constraint relationships between various data fields, such as: Inter-table association parameters: If multiple tables need to share certain data fields, inter-table association parameters need to be defined. For example, a vehicle may need to maintain the same vehicle identification code in multiple tables.
[0076] Intra-field association parameters: If a field contains information of other fields, then intra-field association parameters need to be defined. For example, the purchase price of a car may include information such as the manufacturer's suggested retail price, discounts, and rebates.
[0077] In this embodiment, the data cleaning dictionary is a tool used to describe the data cleaning process and rules, which can standardize and automate the data cleaning work.
[0078] In this embodiment, query intent matching is a natural language processing technology used to analyze the semantics of user queries.
[0079] In this embodiment, the query vector is a representation technology that converts a user's query into a set of numerical vectors.
[0080] In this embodiment, the data relationship path refers to a data structure used to describe entities and the relationships between them in a database management system, including: a starting entity, intermediate entities, an ending entity, and a relationship type.
[0081] The beneficial effects of the above technical solution are as follows: By determining the data inspection direction and verification direction of the second type of data based on the domain index map, corresponding verification rules and inspection technical means can be selected, ensuring the accuracy and reliability of automotive industry data.
[0082] Embodiment 8: An embodiment of the present invention provides a big data-based automotive industry data risk assessment system, as Figure 2 shown, including: A determination module: obtaining various types of data related to the automotive industry, and determining the accuracy and comprehensiveness of various types of data according to verification rules and different inspection technical means; An extraction module: preprocessing the various types of data according to data desensitization technology, and extracting different features based on the preprocessed data; An acquisition module: obtaining the complexity of the problem and the characteristics of the data according to the different features, and obtaining a corresponding machine learning algorithm according to the complexity of the problem and the characteristics of the data; A training and optimization module: defining data risk assessment indicators, establishing an automotive industry data risk assessment model in combination with a machine learning algorithm, and training and optimizing the model based on historical data and experimental data from multiple aspects; A risk assessment module: performing a risk assessment on automotive industry data according to the trained and optimized model, and formulating corresponding data risk management strategies according to the assessment results.
[0083] The beneficial effects of the above technical solution are as follows: By processing various types of data in the automotive industry, obtaining data features, and constructing a risk assessment model based on a machine learning algorithm, it is possible to ensure the diversification of data sources for the automotive industry data risk assessment method, reduce certain biases and limitations, improve the accuracy of the assessment results, avoid potential automotive safety threats, and at the same time, select a suitable algorithm according to the data features without introducing assessment biases, improving the reliability of the risk assessment of automotive industry data.
[0084] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0085] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A data risk assessment method for the automotive industry based on big data, characterized in that: include: Step 1: Obtain various data related to the automotive industry and determine the accuracy and comprehensiveness of various data based on verification rules and different inspection techniques; Step 2: Preprocess the various types of data according to the data desensitization technology, and extract different features based on the preprocessed data; Step 3: Obtain the complexity of the problem and the characteristics of the data according to the different features, and obtain the corresponding machine learning algorithm according to the complexity of the problem and the characteristics of the data; Step 4: Define data risk assessment indicators, and use machine learning algorithms to establish a data risk assessment model for the automotive industry. Train and optimize the model based on historical data and experimental data. Step 5: Conduct risk assessment on automotive industry data based on the trained and optimized model, and formulate corresponding data risk management strategies based on the assessment results.
2. The automotive industry data risk assessment method based on big data according to claim 1 is characterized in that: Obtain various types of data related to the automotive industry and determine the accuracy and comprehensiveness of various types of data based on verification rules and different inspection techniques, including: Obtain various types of car-related data based on road monitoring and vehicle-mounted sensors; Encrypting the various types of data according to encryption technology, and performing sharding on the encrypted data; Obtain anonymity and decentralization of data based on sharded data and combined with anonymization technology; Formulate data quality standards and determine corresponding data validation rules based on the data quality standards; Determine the accuracy and comprehensiveness of various types of data based on the data verification rules and different inspection techniques.
3. The automotive industry data risk assessment method based on big data according to claim 1 is characterized in that: The various types of data are preprocessed according to the data desensitization technology, and different features are extracted according to the preprocessed data, including: Process various types of data using data desensitization technology, and delete and replace data points containing personal privacy information; Normalize and bin the desensitized data; According to the preprocessed data combined with business needs and scenarios, the corresponding feature extraction method is obtained to extract different types of features from the preprocessed data.
4. The automotive industry data risk assessment method based on big data according to claim 1 is characterized in that: Obtaining the complexity of the problem and the characteristics of the data according to the different features, and obtaining the corresponding machine learning algorithm according to the complexity of the problem and the characteristics of the data, including: Determine the data type of the data according to the different features, obtain the corresponding processing method according to the data type, and determine the complexity of the problem according to the processing result; Obtaining correlations between data, obtaining hidden patterns and trends between the data based on the correlations, and determining characteristics of the data based on the patterns and trends; The scale of the data and the form of the problem are determined according to the complexity of the problem and the characteristics of the data, and the corresponding machine learning algorithm is obtained according to the scale of the data and the form of the problem.
5. The automotive industry data risk assessment method based on big data according to claim 1 is characterized in that: Define data risk assessment indicators, and use machine learning algorithms to establish a data risk assessment model for the automotive industry. Train and optimize the model based on historical data and experimental data, including: Define data risk assessment indicators based on the business needs of the automotive industry, and establish a data risk assessment model for the automotive industry in combination with machine learning algorithms; Obtain historical data and experimental data from different data sources and perform data fusion; Perform preliminary training and optimization of the model based on the fused data combined with cross-validation; Select the corresponding key components based on the initially trained and optimized model, and set different hyperparameters to train and optimize the model again.
6. The automotive industry data risk assessment method based on big data according to claim 1 is characterized in that: Carry out risk assessment on automotive industry data based on trained and optimized models, and formulate corresponding data risk management strategies based on the assessment results, including: Conduct multi-dimensional risk assessment of automotive industry data based on trained and optimized models, and obtain data risk levels and risk influencing factors based on the assessment results; Obtain the corresponding series of risky vehicle information based on the data risk level, and determine the risk impact scope based on the risk impact factors; Formulate corresponding data risk management strategies based on the series of risky vehicle information and risk impact scope.
7. The automotive industry data risk assessment method based on big data according to claim 1 is characterized in that: After obtaining various data related to the automotive industry, it also includes: Obtain the data attributes of various types of data, determine the data characteristics of various types of data based on the data attributes, and determine the public opinion health status of various types of data based on the data characteristics; Determine the commercial value of each type of data based on the health status of public opinion, remove the first type of data with low commercial value, and retain the second type of data; Acquire the industry field matching entity relationship of the second type of data, and determine the main entity attribute, sub-entity attribute, behavior entity attribute and business entity attribute of the second type of data according to the industry field matching entity relationship; Determine the entity logical relationship of the second type of data based on the main entity attributes, the sub-entity attributes, the behavior entity attributes and the business entity attributes; Determine data field association parameters based on entity logical relationships, and generate a data cleaning dictionary based on the data field association parameters; Using the data cleaning dictionary to clean each second type of data, and obtaining cleaned second type of data; Performing query intent matching on the cleaned second type of data, and determining a query vector for the second type of data according to the matching result; Determine a data relationship path of the second type of data according to the query vector, and construct a domain knowledge graph of the second type of data according to the data relationship path; The data inspection direction and verification direction of the second type of data are determined according to the domain index map, and the verification rules and inspection technical means are selected based on the data inspection direction and verification direction.
8. A data risk assessment system for the automotive industry based on big data, characterized in that: include: Determination module: obtain various data related to the automotive industry, and determine the accuracy and comprehensiveness of various data based on verification rules and different inspection techniques; Extraction module: pre-processing the various types of data according to data desensitization technology, and extracting different features based on the pre-processed data; Acquisition module: acquiring the complexity of the problem and the characteristics of the data according to the different features, and acquiring the corresponding machine learning algorithm according to the complexity of the problem and the characteristics of the data; Training and optimization module: defines data risk assessment indicators, and establishes a data risk assessment model for the automotive industry in combination with machine learning algorithms, and trains and optimizes the model based on historical data and experimental data; Risk assessment module: Conduct risk assessment on automotive industry data based on the trained and optimized model, and formulate corresponding data risk management strategies based on the assessment results.
Citation Information
Patent Citations
Automobile risk assessment method and system for intelligent networked automobile
CN117273453A
Risk identification model based on data quality monitoring and construction method thereof
CN118014373A
AI-assisted big data anomaly detection and risk assessment system
CN118898038A
Automated master data classification and curation using machine learning
US20210209159A1