Standardized processing method and device for vehicle data assets and electronic equipment
By constructing a quantitative model and a dynamically updated quality standard knowledge base, the subjectivity problem of data classification and grading in the new energy vehicle industry has been solved, and the objectivity and accuracy of data sensitivity grading have been achieved, ensuring data security, compliance and efficient utilization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-21
AI Technical Summary
In existing technologies, the new energy vehicle industry relies on manual qualitative judgment in the data classification and grading process, lacking a unified and quantitative grading basis. This leads to a disconnect between the grading results and actual safety needs, potentially resulting in over-protection or under-protection.
By constructing a quantitative model, matching business scenarios based on data sensitivity factors, and combining a dynamically updated quality standard knowledge base, the security score of data fields is calculated and sensitivity classification labels are marked to achieve data quality verification and loading.
It achieves objectivity and accuracy in data sensitivity classification, ensures data security and compliance, reduces the risk of data leakage, improves data utilization efficiency, and ensures that data processing is highly compatible with business scenarios.
Smart Images

Figure CN121901208A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data management technology, and specifically to a standardized processing method, apparatus, and electronic device for vehicle data assets. Background Technology
[0002] As the new energy vehicle industry rapidly evolves towards electrification, intelligence, and connectivity, data has become a core strategic asset driving business innovation (such as battery safety warnings, user driving behavior analysis, and after-sales rights protection services). Through data asset identification and inventory technologies, companies can locate core data assets supporting key business scenarios. However, transforming these core data assets into standardized, compliant, and high-quality trusted resources has become a key bottleneck in the practical implementation of data assetization. Summary of the Invention
[0003] This invention provides a method, device, and vehicle for displaying an in-vehicle intelligent assistant, in order to solve the problem that in-vehicle intelligent assistants cannot more accurately remind users of the status inside and outside the vehicle.
[0004] In a first aspect, the present invention provides a standardized processing method for vehicle data assets, the method comprising: obtaining initial data assets matching the current business scenario from a data asset catalog; matching a quantification model based on the current business scenario, the quantification model being constructed based on a data sensitivity factor; matching data quality standards corresponding to the current business scenario from a quality standard knowledge base, the information in the quality standard knowledge base being dynamically updated when update conditions are met; calculating security scores for data fields in the initial data assets using the quantification model; labeling each data field with a sensitivity level tag based on the security score to obtain a first data field; performing data quality verification on the first data field based on the data quality standard; and loading a second data field that has passed the quality verification into a target database.
[0005] Secondly, the present invention provides a standardized processing device for vehicle data assets. The device includes: a data asset acquisition module for acquiring initial data assets matching the current business scenario from a data asset catalog; a first matching module for matching a quantitative model based on the current business scenario, wherein the quantitative model is constructed based on a data sensitivity factor; a second matching module for matching data quality standards corresponding to the current business scenario from a quality standard knowledge base, wherein the information in the quality standard knowledge base is dynamically updated when update conditions are met; a scoring module for calculating a security score for data fields in the initial data assets using the quantitative model; a sensitivity grading module for labeling each data field with a sensitivity grading tag based on the security score to obtain a first data field; a data quality inspection module for performing data quality verification on the first data field based on the data quality standard; and a data entry module for loading the second data field that has passed the quality verification into a target database.
[0006] Thirdly, the present invention provides an electronic device comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the computer instructions to perform the method provided in the first aspect.
[0007] The technical solution provided by this invention has the following advantages: This invention retrieves initial data assets matching the current business scenario from a data asset catalog, matches them with a quantitative model adapted to that scenario and a dynamically updated quality standard knowledge base. After the quantitative model calculates the security score of data fields and labels them with sensitivity levels, the qualified data is loaded into the target database through data quality verification. This effectively solves the problem of traditional data classification and grading relying on manual qualitative judgment and lacking a unified quantitative basis. The quantitative model is precisely adapted to the business scenario, ensuring more objective and accurate data sensitivity grading and avoiding over-protection or under-protection. The dynamically updated quality standard knowledge base makes data quality verification more in line with actual requirements, ensuring that the data entering the database is compliant and usable. The overall process not only ensures data security and compliance, reduces the risk of data leakage and compliance costs, but also improves data usage efficiency, making data processing highly compatible with business scenarios. This lays the foundation for the standardization and compliant transformation of data assets and helps the safe and efficient utilization of data assets in the new energy vehicle industry. Attached Figure Description
[0008] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0009] Figure 1 This is a flowchart illustrating a standardized processing method for vehicle data assets according to an embodiment of the present invention. Figure 2 This is another flowchart illustrating a standardized processing method for vehicle data assets according to an embodiment of the present invention; Figure 3 This is another flowchart illustrating a standardized processing method for vehicle data assets according to an embodiment of the present invention; Figure 4 This is another flowchart illustrating a standardized processing method for vehicle data assets according to an embodiment of the present invention; Figure 5 This is a schematic diagram of the process for creating a bloodline analysis system according to an embodiment of the present invention; Figure 6This is a first flowchart of a bloodline analysis system according to an embodiment of the present invention; Figure 7 This is a second flowchart of a bloodline analysis system according to an embodiment of the present invention; Figure 8 This is a third flowchart of a bloodline analysis system according to an embodiment of the present invention; Figure 9 This is a schematic diagram of the structure of a standardized processing device for vehicle data assets according to an embodiment of the present invention; Figure 10 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0010] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0011] It is understood that before using the technical solutions disclosed in the various embodiments of the present invention, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in the present invention and their authorization should be obtained in accordance with relevant laws and regulations through appropriate means.
[0012] Exemplary embodiments of the invention will now be described in more detail with reference to the accompanying drawings. Although exemplary embodiments of the invention are shown in the drawings, it should be understood that the invention may be implemented in various forms and should not be limited to the embodiments set forth herein.
[0013] In the rapid evolution of the new energy vehicle industry towards electrification, intelligence, and connectivity, data has become a core strategic asset driving business innovation (such as battery safety warnings, user driving behavior analysis, and after-sales rights protection services). While data asset identification and inventory technologies enable companies to identify core data assets supporting critical business scenarios, transforming these core data assets into standardized, compliant, and high-quality trusted resources remains a key bottleneck in the implementation of data assetization. In the data classification and grading stage, existing technologies still rely on traditional methods of manual qualitative judgment, where personnel subjectively assess data sensitivity and risks based on experience, lacking unified, quantitative grading criteria and scientific algorithmic support. This leads to a disconnect between grading results and actual safety needs, resulting in either over-protection leading to inefficient data use or insufficient protection causing compliance risks and data leakage vulnerabilities.
[0014] Figure 1 A flowchart of a standardized processing method for vehicle data assets according to the present invention is shown, the method comprising the following steps: Step S101: Obtain the initial data assets that match the current business scenario from the data asset catalog; Step S102: Match a quantification model based on the current business scenario. The quantification model is constructed based on the data sensitivity factor. Step S103: Match the data quality standards corresponding to the current business scenario from the quality standard knowledge base. The information in the quality standard knowledge base is dynamically updated when the update conditions are met. Step S104: Calculate the security score of the data fields in the initial data assets using a quantitative model; Step S105: Label each data field with a sensitivity level label based on the security score to obtain the first data field; Step S106: Perform data quality verification on the first data field based on data quality standards; Step S107: Load the second data field that has passed the quality verification into the target database.
[0015] Specifically, the current business scenario refers to the specific application scenario of data assets generated during the actual operation of new energy vehicles (such as battery safety warning, user driving behavior analysis, etc.). The data asset catalog is a centralized management carrier that stores the metadata, classification, and related business scenarios of all data assets of the enterprise. The initial data asset is the original data set generated in this scenario (including various data fields such as voltage, temperature, and vehicle owner location). The quantification model is a mathematical calculation model specifically adapted to the business scenario and built based on data sensitivity factors. The sensitivity grading label is a classification label used to identify the degree of data sensitivity (such as public, internal, secret, top secret, etc.). The quality standard knowledge base is a dedicated database that stores various data quality verification rules corresponding to different business scenarios. Its update conditions include internal data standard updates, national data standard updates, and data distribution changes meeting preset conditions. The target database is a dedicated database that stores data assets after standardization processing.
[0016] In this embodiment of the invention, the method is applied to electronic devices such as computers and servers. After the device obtains data from the sensors of new energy vehicles, vehicle terminals, and the background data warehouse of R&D servers through the data acquisition interface, it accurately selects the initial data assets that match the current business scenario (taking the battery safety warning scenario as an example) based on the business scenario association recorded in the data asset catalog, so as to ensure that the data source is highly consistent with the business scenario and avoid interference from irrelevant data.
[0017] Next, the device first matches the corresponding quantitative model based on the core requirements of the current business scenario (such as the requirements for data security and real-time performance in battery safety warning scenarios). This model has been pre-adapted to the scenario based on data sensitivity factors. At the same time, it retrieves the data quality standards corresponding to the business scenario from the quality standard knowledge base, such as the valid value range of battery voltage data and data format specifications and other verification rules.
[0018] Then, a quantitative model matching the battery safety warning scenario is invoked to calculate a safety score for each data field in the initial data asset, such as voltage and temperature. The model quantifies the sensitivity of the data based on the characteristics of the scenario. For example, battery voltage data is highly sensitive in this scenario because it is directly related to battery safety. The model will calculate the corresponding safety score based on this and then label each data field with a corresponding sensitivity level label based on the score. Compared with traditional manual qualitative grading, the quantitative model used in this embodiment of the invention is more objective and accurate in assessing data sensitivity.
[0019] Next, based on the data quality standards matched from the quality standard knowledge base, automated quality checks are performed on the first labeled data field. The checks include data integrity, accuracy, and format compliance. For example, for the battery voltage data field, the checks verify whether its value is within a preset valid range (e.g., 2.5V-4.2V); for the vehicle VIN code field, the checks verify whether it meets requirements such as being non-empty and having a compliant format.
[0020] Finally, only the second data field that passes quality verification is loaded into the target database according to the preset data storage protocol, achieving standardized storage and secure management of data assets. This embodiment of the invention achieves precise data sensitivity classification through a scenario-adaptive quantitative model, and simultaneously completes data quality verification by combining a dynamically updated quality standard knowledge base. This effectively solves the problems of traditional manual qualitative classification lacking unified standards, being out of touch with business needs, and lacking targeted data quality control. It ensures the security and compliance of data assets while avoiding inefficient data use due to over-protection, and ensures high-quality data entering the database, providing reliable support for the standardized and efficient utilization of data assets in the new energy vehicle industry. In some optional implementations, data sensitivity factors include business scenario sensitivity, legal and regulatory relevance, and data lineage diffusion. Business scenario sensitivity is an indicator of the sensitivity of a data field under different business scenarios; legal and regulatory relevance is an indicator of the degree to which a data field complies with laws and regulations; and data lineage diffusion is an indicator of the breadth and depth of data citation.
[0021] Specifically, the data sensitivity factor is the core set of input parameters for constructing a quantitative model, used to comprehensively measure the sensitivity of data fields. It includes three types of factors: business scenario sensitivity, legal and regulatory relevance, and data lineage diffusion. Among them, "business scenario sensitivity" refers to the quantitative indicator of the sensitivity of data fields under different business scenarios. Its value is directly related to the coreness of the business scenario and security requirements. "Legal and regulatory relevance" refers to the degree of compliance of data fields with relevant laws and regulations such as GDPR and the Data Security Law, used to assess the compliance risks of data use. "Data lineage diffusion" refers to the extent (e.g., the number of downstream entities referencing the data) and depth (e.g., the degree to which the data is used by downstream entities for core business logic, key decision support, or in-depth processing) of data fields in the entire data flow process, reflecting the size of the data's influence.
[0022] In this embodiment of the invention, taking two typical business scenarios of new energy vehicles, "after-sales rights protection" and "battery safety warning," as examples, for the data field of vehicle owner location information, in the "after-sales rights protection" scenario, this information is directly related to user privacy and rights protection dispute handling, and its business scenario sensitivity index value is high; in the "battery safety warning" scenario, vehicle owner location information is not core data, and the business scenario sensitivity index value is relatively low. Regarding the legal and regulatory relevance, taking the vehicle VIN code data field as an example, since the VIN code is a directly identifiable information (PII), it is directly related to the requirements of multiple laws and regulations regarding personal information protection, so its legal and regulatory relevance index value is high. Regarding the data lineage diffusion, taking the battery voltage data field as an example, if this field is only referenced by internal reports of the battery management system, its diffusion index value is low; if it is simultaneously referenced by multiple downstream applications such as the battery safety warning system, vehicle health diagnosis platform, and after-sales maintenance data system, satisfying both the high threshold of wide reference and the high requirement of deep reference, its diffusion index value will significantly increase. This invention uses three types of factors as core inputs, and the quantitative model can comprehensively assess the sensitivity of data fields from three dimensions: business adaptability, compliance risk, and scope of impact. Compared with the traditional single-dimensional classification method, the assessment results are more comprehensive and scientific, effectively supporting the accurate classification of data sensitivity and providing a reliable basis for the secure and compliant management of new energy vehicle data assets.
[0023] In some alternative implementations, the steps of constructing a quantization model based on the data sensitivity factor include: Step a1: Adapt the corresponding legal knowledge base for different categories of data fields; Step a2: Create a data lineage analysis system. The data lineage analysis system is used to extract the data links associated with the business scenario and the data attribute information in the data links. The data links are used to characterize the breadth of data being cited, and the data attribute information is used to characterize the depth of data being cited. The data attribute information includes at least one of the following: sensitivity level of cited data, importance level of business scenario, and number of times it is cited. Step a3: Define the scenario sensitivity level for different categories of data fields in various business scenarios; Step a4: Define the first mapping rule, which is used to map the legality check results of the legal knowledge base on the data fields to the degree of legal and regulatory relevance; Step a5: Define the second mapping rule, which is used to map the data link and data attribute information extracted by the data lineage analysis system from the business scenario to the data lineage diffusion degree; Step a6: Define the third mapping rule. The third mapping rule is used to map the scenario sensitivity level of a data field in any business scenario to the business scenario sensitivity. Step a7: Define the fourth mapping rule. The fourth mapping rule is used to integrate the relevance of laws and regulations, the diffusion of data lineage, and the sensitivity of business scenarios to obtain the security score of the data field. Step a8: Define the fifth mapping rule, which is used to map the security score to the corresponding sensitivity rating label.
[0024] Specifically, in this embodiment of the invention, the data sensitivity factor is the core foundation for constructing the quantitative model, encompassing business scenario sensitivity, legal and regulatory relevance, and data lineage diffusion. The quantitative model is a mathematical analysis model constructed through a series of standardized steps, capable of accurately calculating the security score of data fields and achieving sensitivity grading. The regulatory knowledge base is a professional database storing legal and regulatory clauses and compliance requirements corresponding to different categories of data fields. The data lineage analysis system is a technical system capable of tracking the entire data flow trajectory from generation to use, extracting business scenario-related data links and data attribute information within those links. The data links characterize the breadth of data citation, while the data attribute information characterizes the depth of data citation, including at least one of the following: sensitivity grading of cited data, business scenario importance level, and number of citations. Scenario sensitivity level is a grading standard (e.g., extremely high, high, medium, low) determined based on the importance and security impact of data fields in different business scenarios. The mapping rule is a preset logical rule for transforming one type of data or result into another type of target data or indicator.
[0025] In this embodiment of the invention, data is first classified and then graded according to business scenarios. Specifically, based on the actual classification of new energy vehicle data assets, corresponding regulatory knowledge bases are adapted for different categories of data fields. For example, personal identity data fields such as vehicle owner ID numbers and mobile phone numbers are adapted to a dedicated regulatory knowledge base containing GDPR and personal information protection clauses in the Data Security Law; equipment operation data fields such as vehicle battery voltage and driving range are adapted to a regulatory knowledge base involving compliance requirements for data collection and storage, ensuring that each category of data field corresponds to an accurate compliance judgment basis.
[0026] In addition, a data lineage analysis system was created. This system collects information on the correlation of data throughout its entire lifecycle, including generation, transmission, processing, and application, and builds a network of dependencies between data. When faced with specific business scenarios (such as battery safety warnings and after-sales rights protection), it can not only quickly locate and extract the data flow links directly related to the scenario, but also simultaneously obtain the data attribute information in the links (such as the sensitivity level of the cited data being "top secret", the importance level of the business scenario being "core", and the number of times it has been cited being 12, etc.), clarifying the source, processing process, downstream application scope, and citation depth of the data.
[0027] In addition, based on the core business scenarios of new energy vehicles (such as battery safety warnings, user driving behavior analysis, and after-sales rights protection services), scenario sensitivity levels are defined for different categories of data fields. For example, in the "battery safety warning" scenario, data fields such as battery temperature and voltage are directly related to vehicle driving safety and are defined as "extremely high" sensitivity level; in the "navigation service" scenario, the same battery temperature and voltage data fields are only used as auxiliary information and are defined as "medium" sensitivity level.
[0028] The mapping rules for converting the aforementioned sensitive factor information into numerical indicators are then defined sequentially, namely the first to fifth mapping rules. The first mapping rule clarifies the numerical correspondence between the legality check results of the data fields in the regulatory knowledge base (e.g., fully compliant, partially compliant, non-compliant) and the correlation with laws and regulations. For example, full compliance corresponds to a correlation of 0.2, partial compliance to 0.6, and non-compliance to 1.0. The second mapping rule converts the data link characteristics extracted by the data lineage analysis system (e.g., the number of downstream applications and the scope of influence) and data attribute information (e.g., the sensitivity level of cited data, the importance level of the business scenario, and the number of citations) into data lineage diffusion. For example, if only one downstream application cites the data and the sensitivity level is "public" and the importance level of the business scenario is "general," the diffusion is 0.3; if five or more downstream applications cite the data and the sensitivity level is "top secret," the importance level of the business scenario is "core," and the number of citations is ≥10, the diffusion is 0.9. The third mapping rule establishes a quantitative correspondence between scenario sensitivity levels and business scenario sensitivity, such as "extremely high" level corresponding to 1.0, "high" level corresponding to 0.8, "medium" level corresponding to 0.5, and "low" level corresponding to 0.2. The fourth mapping rule uses weighted summation algorithms to integrate legal and regulatory relevance, data lineage diffusion, and business scenario sensitivity. For example, it sets a weight of 0.4 for business scenario sensitivity, 0.3 for legal and regulatory relevance, and 0.3 for data lineage diffusion. The security score of the data field is calculated using the formula: "Security Score = Business Scenario Sensitivity × 0.4 + Legal and Regulatory Relevance × 0.3 + Data Lineage Diffusion × 0.3". The fifth mapping rule divides the security score into different ranges, corresponding to sensitivity levels such as public, internal, secret, and top secret. For example, a security score of 0-0.3 corresponds to "public", 0.3-0.6 corresponds to "internal", 0.6-0.8 corresponds to "secret", and 0.8-1.0 corresponds to "top secret".
[0029] This invention constructs a quantitative model through systematic and standardized steps, achieving scientific and precise data sensitivity grading. By adapting dedicated regulatory knowledge bases to different categories of data fields, creating a professional data lineage analysis system, and defining refined scenario sensitivity levels, it breaks through the subjective limitations of traditional manual qualitative grading, making the grading basis more objective and professional, and effectively avoiding over-protection or under-protection caused by inaccurate grading. Secondly, it strengthens the compliance and security of data governance. Through the first to third mapping rules, multi-dimensional information such as business scenarios, laws and regulations, and data flow are transformed into quantifiable indicators. Then, through the fourth and fifth mapping rules, security scores are calculated and grading labels are marked, ensuring that the entire data processing process strictly follows compliance requirements, significantly reducing the risks of data leakage and unauthorized use, and fully meeting the strict requirements of relevant laws and regulations. Thirdly, it improves the utilization efficiency and value of data assets. The construction process of the quantitative model fully combines the business scenario characteristics of new energy vehicles, ensuring that the grading results are highly adapted to actual business needs. This not only ensures the security and controllability of sensitive data but also provides support for the efficient circulation and sharing of non-sensitive data, avoiding the waste of data resources caused by excessive compliance restrictions. Fourth, it provides a standardized and replicable technical solution for data asset management. The entire construction process is logically clear and standardized. It can flexibly adjust the regulatory knowledge base, scenario sensitivity level and mapping rules according to the business expansion and new data types in the new energy vehicle industry. It has strong adaptability and scalability, and can support the standardized management of enterprise data assets in the long term, helping the new energy vehicle industry to achieve business innovation and high-quality development driven by data.
[0030] In some alternative implementations, such as Figure 2 As shown, the above steps S104 and S105 include: Step b1: Determine the data category of the current data field; Step b2: Match the corresponding target legal knowledge base according to the data category, and use the target legal knowledge base to check the legality of the current data field to obtain the current legality check result; Step b3: Map the current legality check result to the current legality relevance using the first mapping rule; Step b4: Use the data lineage analysis system to extract the target data links and target data attribute information associated with the current business scenario; Step b5: Map the target data link and target data attribute information to the current data lineage diffusion degree using the second mapping rule; Step b6: Query the target scenario sensitivity level defined for the current data field in the current business scenario; Step b7: Map the target scenario sensitivity level to the current business scenario sensitivity using the third mapping rule; Step b8: Use the fourth mapping rule to integrate the current legal and regulatory relevance, the current data lineage diffusion, and the current business scenario sensitivity to obtain the target security score of the current data field; Step b9: Use the fifth mapping rule to map the target security score to the target sensitivity rating label of the current data field.
[0031] Specifically, data category refers to the classification based on data content attributes (such as personal identity, equipment operation, location information, etc.), target legal knowledge base is a professional database adapted to specific data categories and storing corresponding compliance clauses, legality check result refers to the judgment conclusion on whether the data field meets the requirements of corresponding laws and regulations (such as fully compliant, partially compliant, non-compliant), target data link is the data flow path directly related to the current business scenario extracted by the data lineage analysis system, target data attribute information is information extracted synchronously by the data lineage analysis system to characterize the depth of data citation, including at least one of the following: sensitivity level of cited data, importance level of business scenario, and number of citations, target scenario sensitivity level is the sensitivity level of the data field pre-set in a specific business scenario (such as extremely high, high, medium, low), target security score is a quantitative value obtained by integrating multi-dimensional factors, and target sensitivity level label is the final classification label used to identify the sensitivity of data (such as public, internal, secret, top secret).
[0032] In this embodiment of the invention, battery voltage data and vehicle owner mobile phone number data in the new energy vehicle battery safety early warning business scenario are used as examples for explanation. After the system obtains the initial data assets of the current business scenario from the data asset catalog, it automatically determines the data category of the current data field through the metadata characteristics of the data field (such as field name and description information). Among them, "battery voltage data" is determined to be equipment operation data, and "vehicle owner mobile phone number data" is determined to be personal identity data. Then, according to the determined data category, the corresponding target regulatory knowledge base is matched. Equipment operation data is matched to a regulatory knowledge base containing data collection and storage compliance requirements, and personal identity data is matched to a regulatory knowledge base containing GDPR and personal information protection clauses in the Data Security Law. Subsequently, the legality of the data fields is checked using the corresponding target regulatory knowledge base. For example, "battery voltage data" receives a current legality check result of "fully compliant" because it meets the data collection compliance requirements, while "vehicle owner mobile phone number data" receives a current legality check result of "partially compliant" because it involves personal privacy protection and does not meet the de-identification requirements.
[0033] Then, the legality check results are mapped to the current legality correlation degree through the preset first mapping rule. Assuming that the rule is set to 1.0 for full compliance, 0.6 for partial compliance, and 0.2 for non-compliance, the current legality correlation degree of "battery voltage data" is 1.0, and the current legality correlation degree of "car owner mobile phone number data" is 0.6.
[0034] In this embodiment of the invention, a data lineage analysis system is used to extract target data links and target data attribute information associated with the "battery safety warning" scenario. For example, the target data link for "battery voltage data" is "vehicle sensor acquisition → battery management system processing → safety warning dashboard display", and the target data attribute information extracted simultaneously is "citation data sensitivity level: high, business scenario importance level: core, number of citations: 8". Similarly, the target data link for "vehicle owner mobile phone number data" is "user registration entry → backend database storage → after-sales warning notification call", and the target data attribute information extracted simultaneously is "citation data sensitivity level: medium, business scenario importance level: general, number of citations: 3". Then, the target data link and target data attribute information are jointly mapped to the current data lineage diffusion degree through the second mapping rule. The rule sets the mapping relationship according to the combination of "number of downstream applications + importance level of business scenario + number of references" (e.g., 1-2 downstream applications + general importance level + number of references ≤ 5 times corresponds to 0.3; 3-5 downstream applications + core importance level + number of references ≥ 6 times corresponds to 0.6; more than 5 downstream applications + core importance level + number of references ≥ 10 times corresponds to 0.9). "Battery voltage data" meets the condition of "3 downstream applications + core importance level + number of references 8 times", and the current data lineage diffusion degree is 0.6; "car owner mobile phone number data" meets the condition of "2 downstream applications + general importance level + number of references 3 times", and the current data lineage diffusion degree is 0.3.
[0035] Next, the sensitivity level of the current data field in the "Battery Safety Warning" scenario is queried. "Battery Voltage Data" is set to "Extremely High" because it is directly related to vehicle safety, while "Owner's Phone Number Data" is set to "Medium" because it is not core data for this scenario. Conversely, in an after-sales service scenario, "Battery Voltage Data" is set to "Low" because it is not related to after-sales service, while "Owner's Phone Number Data" is set to "Extremely High" because it is core data for this scenario.
[0036] Then, the scenario sensitivity level is mapped to the current business scenario sensitivity through the third mapping rule. The rule is set to "extremely high" corresponding to 1.0, "high" corresponding to 0.8, "medium" corresponding to 0.5, and "low" corresponding to 0.2. Therefore, the current business scenario sensitivity of "battery voltage data" is 1.0, and the current business scenario sensitivity of "car owner mobile phone number data" is 0.5.
[0037] After the above three mappings, this embodiment of the invention uses a fourth mapping rule (such as the weighted summation formula "target safety score = business scenario sensitivity × 0.4 + legal and regulatory relevance × 0.3 + data lineage diffusion × 0.3") to fuse the above three dimensions of data, and calculates that the target safety score of "battery voltage data" is 1.0 × 0.4 + 1.0 × 0.3 + 0.6 × 0.3 = 0.4 + 0.3 + 0.18 = 0.88, and the target safety score of "vehicle owner's mobile phone number data" is 0.5 × 0.4 + 0.6 × 0.3 + 0.3 × 0.3 = 0.2 + 0.18 + 0.09 = 0.47.
[0038] Finally, the target safety score is mapped to the target sensitivity level label through the fifth mapping rule. Assuming that the rule sets 0-0.3 to "public", 0.3-0.6 to "internal", 0.6-0.8 to "secret", and 0.8-1.0 to "top secret", then "battery voltage data" corresponds to the "top secret" label and "car owner's mobile phone number data" corresponds to the "internal" label.
[0039] This invention achieves precise and personalized data sensitivity grading. By matching data categories to a regulatory knowledge base, extracting scenario-based data links, and combining multi-dimensional factors for quantitative calculation, it breaks through the subjective limitations of traditional manual qualitative grading. This makes the grading results more closely match the actual attributes of the data and the needs of business scenarios, avoiding the drawbacks of a "one-size-fits-all" approach. This invention enhances the targeting and operability of data compliance management. The precise matching of legality checks with the regulatory knowledge base ensures that compliance judgments for different types of data are based on professionalism and comprehensiveness. It provides enterprises with a clear execution path to meet the requirements of GDPR, the Data Security Law, and other laws and regulations, significantly reducing compliance risks. This invention improves the efficiency and automation level of data grading. The entire process, from data category determination and compliance checks to security score calculation and labeling, requires no manual intervention, greatly reducing labor costs and avoiding errors that may be caused by human operation, ensuring the consistency and reliability of the grading results. The embodiments of this invention provide scientific support for the secure management and efficient utilization of data assets. The precise sensitivity classification label can clearly define the security management level of data, which not only ensures the security and controllability of highly sensitive data, but also provides a basis for the efficient circulation and sharing of low-sensitivity data, avoiding the waste of data resources caused by over-protection.
[0040] In some optional implementations, the standardized processing method for vehicle data assets provided in this embodiment of the invention further includes the following steps: Step c1: Standardize the third data field that fails the quality check based on the data correction standard; Step c2: Use the standardized fourth data field as the first data field, and return the steps for performing data quality verification on the first data field based on data quality standards.
[0041] Step c3: If the fourth data field obtained after the preset number of corrections still fails the data quality check, then the loading of the fourth data field into the target database is blocked, and an alarm message is output. Step c4: When the data correction standard corresponding to the current business scenario is updated, the fourth data field is re-standardized based on the updated data correction standard. Step c5: Use the standardized fourth data field as the first data field, and return the steps for performing data quality verification on the first data field based on data quality standards.
[0042] Specifically, this embodiment of the invention performs compliance checks on data fields labeled with sensitivity levels according to preset quality rules through data quality verification, ensuring that the data meets the requirements of usability and accuracy. The quality standard knowledge base also includes data correction standards corresponding to the current business scenario. The data correction standards are a set of rules used to standardize data fields that fail quality verification. Passing quality verification means that the data field fully complies with the preset quality rules. The target database is a dedicated database that stores data assets after standardization and compliance processing. The alarm information refers to prompts containing information such as the type of data quality problem, the fields involved, and the reason for the violation, used to notify relevant personnel to handle the matter in a timely manner.
[0043] In embodiments of the present invention, such as Figure 3 As shown, taking the "Battery Voltage Data Field" and "Vehicle VIN Field" with pre-labeled sensitivity levels in the "Battery Safety Warning" business scenario for new energy vehicles as examples, we will explain the embedded quality verification process. First, we initiate a data quality verification process for the two data fields with pre-labeled sensitivity levels. The quality verification rules are pre-packaged as standardized quality inspection plugins and embedded in the data processing pipeline. Specifically, for the "Battery Voltage Data Field," the "Voltage Range Check" plugin is called, with a preset effective voltage range of 2.5V-4.2V. For the "Vehicle VIN Field," the "Non-empty Check" plugin is called, with the rule that the VIN field cannot be empty. At the same time, the data correction standards corresponding to the scenario are matched (such as voltage data format unification rules, VIN field completion rules, etc.).
[0044] If the detected value of the "Battery Voltage Data Field" is 3.7V, which is within the preset valid range, and the detected result of the "Vehicle VIN Field" is not empty, both of which pass the quality verification, the system will load these two data fields into the target database according to the preset data storage protocol, ensuring that high-quality data can be called in subsequent business scenarios.
[0045] If the "Battery Voltage Data Field" test value is 1.8V (exceeding the preset value range), or the "Vehicle VIN Field" test result is empty (not in compliance with the non-empty verification rule), then the third data field is determined to be a failure of quality verification.
[0046] At this point, the system will standardize the data based on the data correction standard. For example, if the voltage value is abnormal, the system will correct 1.8V to a reasonable interpolation value that meets the value range requirements (such as 3.5V calculated based on historical data) according to the correction standard. If the VIN field is empty, the system will trigger a query to complete the data based on the associated data source according to the correction standard (such as obtaining the corresponding VIN code from the vehicle registration information database) to obtain the fourth data field after standardization.
[0047] The fourth data field is then used as the first data field again, and the quality verification step is returned to perform the verification again. If the fourth data field still fails the quality verification after the preset number of corrections (e.g., 2 times) (e.g., the VIN code cannot be completed after multiple queries), the system will immediately block the loading of this data field into the target database to prevent unqualified data from polluting data resources. At the same time, it will automatically output alarm information (e.g., "Battery safety warning scenario - the vehicle VIN field (grading label: internal) is still empty after 2 corrections, which does not meet the non-empty verification rule") so that relevant technical personnel can promptly investigate the root cause of the problem in the data collection or transmission process.
[0048] If the data correction standard corresponding to the current business scenario is updated during the correction process (such as adding a new VIN field format validation rule), the system will re-standardize the fourth data field that failed the validation based on the updated data correction standard, and then return the processed field to the quality validation step to ensure that the correction logic is consistent with the latest standard.
[0049] This invention utilizes pipelined embedded quality rule automated verification technology, combined with a dynamic data correction mechanism, to transform quality verification from a "batch processing task" to a "pipeline plug-in and closed-loop correction" model. Through the coordination of standardized "quality inspection plug-ins" and "correction rules," a closed-loop process of verification, correction, and re-inspection is automatically completed within the data synchronization and processing workflow. Data that fails to meet standards after multiple corrections is blocked from entering the process and receives real-time alerts, ensuring that only "compliant" data enters the data warehouse. This achieves proactive data quality assurance and closed-loop governance, abandoning traditional post-audit and single-interception models. It prevents data contamination at the source and improves data usability through automated correction, ensuring the credibility and integrity of data resources in the target database. Furthermore, the verification and correction rules are adapted to different data fields, making the logic more targeted and ensuring that the data meets both business scenario requirements and data asset standardization requirements. Simultaneously, it reduces data management costs and business risks. Through automated verification, correction, and real-time alerts, it reduces manual intervention workload and avoids decision-making errors and business anomalies caused by unqualified data flowing into business processes, providing a solid guarantee for the safe and efficient utilization of data assets in the new energy vehicle industry.
[0050] In some alternative implementations, the step of dynamically updating the quality standard knowledge base includes: Step d1: Monitor the update status of internal data standards and national data standards according to the preset cycle; Step d2 involves monitoring changes in the distribution of initial data assets acquired at different time periods. Step d3: When the data distribution change meets the preset change conditions, query the update status of the internal data standards and national data standards. Step d4: If the update status indicates that any data standard has been updated, then obtain the updated standard content and update the quality standard knowledge base based on the updated standard content.
[0051] Specifically, the quality standard knowledge base is a dedicated database that stores various data quality standards corresponding to different business scenarios. Data quality standards include specific verification rules for data integrity, accuracy, uniqueness, value range, etc. Internal data standards refer to data quality specifications formulated by enterprises according to their business needs. National data standards refer to mandatory or recommended industry data quality standards issued by relevant national departments. The preset cycle is a fixed time interval (e.g., once a day) for monitoring the update status of data standards. Data distribution change refers to the differences in numerical range, frequency of occurrence, format type, etc. of the initial data assets acquired in different time periods. The preset change condition is a threshold standard for judging whether the data distribution has changed significantly (e.g., the data mean fluctuation exceeds 20%, the proportion of newly added data formats exceeds 15%, etc.).
[0052] In this embodiment of the invention, taking the "battery safety warning" business scenario for new energy vehicles as an example, it is assumed that the device monitors the update status of internal and national data standards in real time through interfaces connecting to the enterprise's internal standard management system and the national data standard release platform, according to a preset daily cycle. For example, it monitors whether the internal accuracy standard for battery voltage data has been updated, and whether the state has released new specifications related to data security for new energy vehicles. For the initial data assets of the "battery safety warning" scenario collected at different time periods (such as 9:00, 15:00, and 21:00 daily), statistical analysis tools are used to monitor changes in data distribution, such as comparing the average battery temperature data and the difference in the range distribution of voltage data values at the same time this week and last week.
[0053] Furthermore, if the monitored data distribution changes meet preset change conditions, such as the mean fluctuation of battery voltage data reaching 25%, exceeding the preset threshold of 20%, the system detects that the data fluctuation is large and the data standard may have changed. Therefore, it will immediately query the update status of internal and national data standards and proactively evolve the data standard. If the query finds that the internal data standard has updated the accuracy requirements for battery voltage (e.g., from retaining one decimal place to retaining two decimal places), or that the national data standard has added specifications for the collection range of battery safety data, the system will automatically obtain the updated standard content through the interface. Subsequently, it will synchronously update the battery voltage accuracy verification rules and data collection range verification rules corresponding to the "battery safety warning" scenario in the quality standard knowledge base, ensuring that the data quality standards in the quality standard knowledge base are always consistent with the latest internal requirements and national standards. If the query finds that the data standard has not been updated, it is determined that the user-input data is non-compliant. Based on the above improvements, a complete flowchart of this embodiment can be found. Figure 4 .
[0054] This invention achieves dynamic adaptation and continuous optimization of quality standards. By regularly monitoring the update status of data standards and confirming the update status based on changes in data distribution, it overcomes the shortcomings of traditional static quality standards, which are difficult to adapt to the rapid iteration of business and industry norms. This ensures that data quality verification rules always align with actual needs. It improves the compliance and accuracy of data quality verification by optimizing the quality standard knowledge base in real time based on updated internal and national data standards. This ensures that data quality verification meets both enterprise business management requirements and strictly adheres to industry regulations, effectively reducing compliance risks. It enhances the usability and credibility of data assets by dynamically updating quality standards, ensuring accurate quality verification of data assets with different time periods and distribution characteristics. This prevents unqualified data from flowing into the target database due to outdated standards, providing a reliable guarantee for the standardized and high-quality management of data assets in the new energy vehicle industry. Simultaneously, the automated monitoring and update process significantly reduces manual maintenance costs and improves data governance efficiency.
[0055] In some alternative implementations, the steps for creating a data lineage analysis system include: Step e1: Create a full lineage graph to describe the dependencies between data based on the full lifecycle relationships between data from generation to use. The full lineage graph includes the complete flow path of all data assets in the data asset catalog. Step e2: Query the corresponding set of scenario data assets for each business scenario based on the business keywords corresponding to each business scenario; Step e3: Create a scenario-data mapping model based on the correspondence between business scenarios and scenario data asset sets; Step e4: Create a path focusing engine based on graph traversal technology. The path focusing engine is used to cut out data flow paths that are not related to a certain target business scenario from the full lineage graph according to the scenario-data mapping model, and extract the data links associated with the target business scenario.
[0056] Specifically, the full data lineage graph is a visual representation of the relationships between data throughout its entire lifecycle, from generation and transmission to processing and use. It includes dependencies between all data assets, such as tables, fields, metrics, and dashboards, and represents the complete flow path of all data assets in the data asset catalog. Business keywords are terms that characterize the core needs of a specific business scenario (e.g., "battery," "voltage," and "alarm" in a "battery safety warning" scenario). The scenario data asset set is all data assets (including data tables, metrics, and dashboards) strongly related to a particular business scenario. The scenario-data mapping model is a database model that stores the relationships between business scenarios and their corresponding data asset sets, supporting dynamic updates. The path focusing engine is a core component built on graph traversal technology, used to accurately extract target data links from the complex full data lineage graph.
[0057] Traditional lineage systems only display the technical dependencies between tables, fields, and tasks, making it difficult for business personnel to understand their business implications and quickly answer questions like "Where does the data for this business scenario come from, and which key reports are affected?" This lack of targeted lineage analysis based on business scenarios leads to inefficient impact analysis or attribution efforts, resulting in either overly broad or overly narrow scopes.
[0058] like Figure 5As shown, this embodiment of the invention uses two core business scenarios—"battery safety early warning" and "user driving behavior analysis"—of new energy vehicles as examples for illustration. The system collects information on the entire process of data flow, including data collection by vehicle sensors, transmission by vehicle terminals, processing by the backend system, and invocation by business applications. It then sorts out the upstream and downstream dependencies between various data assets. For example, "battery voltage data" is generated by vehicle sensors, processed by the battery management system into "voltage anomaly indicators," and then referenced by the "battery safety early warning dashboard." The system constructs a full lineage graph of these relationships in the form of nodes (representing data assets) and edges (representing dependencies), fully presenting the flow path of all data.
[0059] refer to Figure 6 For example, in the "battery safety warning" scenario, we extract the business keywords "battery," "voltage," "temperature," and "alarm." Using natural language processing, we retrieve all data assets containing these keywords from the metadata repository, filtering out Table B (stores raw battery voltage and temperature data), Indicator X (battery voltage anomaly indicator), and Dashboard 1 (battery safety warning dashboard), thus forming the scenario data asset set for this scenario. Similarly, in the "user driving behavior analysis" scenario, we retrieve data assets such as Table C (user driving trajectory data) and Indicator J (average driving speed indicator) using the keywords "driving," "speed," and "path," forming its scenario data asset set.
[0060] Based on the correspondence between the above business scenarios and the set of scenario data assets, a scenario-data mapping model is created. The model clarifies the relationship between "battery safety warning" and Table B, indicator X, and dashboard 1, and the relationship between "user driving behavior analysis" and Table C and indicator J. It also supports automatic updating of the relationship through keyword matching when new business scenarios or data assets are added.
[0061] Furthermore, this embodiment of the invention also creates a path focusing engine based on graph traversal technology. This engine has preset correlation calculation rules and can traverse the entire lineage graph and identify nodes and edges directly related to the target business scenario data asset set according to the correlation relationships in the scenario-data mapping model. It automatically prunes irrelevant branches. For example, for the "battery safety warning" scenario, the engine will focus on the nodes corresponding to Table B, indicator X, and dashboard 1 and the dependency edges between them, ignoring branches of irrelevant data assets such as Table A (user registration information table), Table C, Table D (media playback record table), indicator Y, and dashboard 2, accurately extracting the core data link of "battery safety warning → Table B → indicator I → dashboard D". The above process can be referred to Figure 7 .
[0062] The data lineage analysis system created in this invention solves the problem of the disconnect between technology and business in traditional lineage analysis. By matching data assets with business keywords and constructing a scenario-data mapping model, lineage analysis directly aligns with business needs. Business personnel can quickly locate target data links without needing to understand complex technical dependencies. Furthermore, it improves the efficiency and accuracy of lineage analysis. The full lineage graph completely preserves data relationships, and the path focusing engine uses graph traversal technology to prune irrelevant branches, avoiding analysis interference caused by the complexity of the full lineage graph, making data link extraction more efficient. Thirdly, it enhances the system's adaptability and scalability. The scenario-data mapping model supports dynamic updates and can adapt to new business scenarios or data assets. The graph traversal logic of the path focusing engine can flexibly adapt to the link extraction needs of different business scenarios, providing a standardized and efficient solution for data lineage analysis in the diverse business scenarios of the new energy vehicle industry. It also provides precise link support for subsequent data governance stages such as data sensitivity assessment and data quality verification, facilitating standardized management of the entire data assetization process.
[0063] In some alternative implementations, such as Figure 8 As shown, step b4 above includes: Step f1: Input the current business scenario into the scenario-data mapping model, so as to output the corresponding set of data assets for the current scenario through the scenario-data mapping model; Step f2 involves reading the current scene data asset set through the path focusing engine and extracting the paths and nodes where the current scene data asset set is located from the full lineage graph to form the target data link.
[0064] Specifically, the current business scenario refers to the specific application scenarios that are being carried out in the actual operation of new energy vehicles (such as battery safety warning, user driving behavior analysis, etc.), the current scenario data asset set is all data assets (including data tables, indicators, dashboards, etc.) that are strongly related to the current business scenario, and the target data link is the core data flow path that is directly related to the current business scenario.
[0065] When a user initiates a data sensitivity assessment request in the system, this embodiment of the invention first automatically inputs the current business scenario (taking the "battery safety warning" business scenario of new energy vehicles as an example for explanation) into the scenario-data mapping model. This model has established a dynamic association between the business scenario and data assets in advance through natural language processing technology. It will quickly retrieve all data assets associated with "battery safety warning" in the model and output a set of current scenario data assets including Table B (stores raw data of battery voltage and temperature), Indicator I (abnormal battery voltage indicator), and Dashboard D (battery safety warning dashboard).
[0066] After reading the current scene data asset set, the path focusing engine calls the graph traversal algorithm to analyze the full lineage graph. The full lineage graph contains data assets and dependencies unrelated to battery safety warnings, such as Table A (user registration information table) and Table C (media playback record table). The path focusing engine automatically identifies and retains the nodes corresponding to each data asset in the current scene data asset set, as well as the dependency edges between nodes, and prunes the branch paths of irrelevant data assets such as Table A and Table C. Finally, it extracts the core data flow path of "vehicle sensor acquisition → Table B → indicator I → dashboard D", forming a target data link that accurately matches the "battery safety warning" scenario. At the same time, it displays the target data link in the user interface by highlighting it, generating and presenting a lineage graph exclusive to the battery safety warning scenario.
[0067] In some optional implementations, based on lineage analysis, this invention further adds a lineage value assessment method. When a data source table changes, the system not only lists the affected downstream data but also assesses the degree of impact (e.g., high, medium, low) of the changed data on related business scenarios. By establishing an impact factor model, a comprehensive impact score is calculated by considering factors such as the importance of the business scenarios supported by the downstream data assets and the frequency of data updates. Taking the three major scenarios of "battery safety warning," "after-sales rights protection," and "in-vehicle entertainment recommendation" for new energy vehicles as examples, the data source table T (stores battery operation data) has downstream related indicators X (supporting safety warnings), report B (supporting after-sales rights protection), and tag C (supporting entertainment recommendations). When a new "Battery Internal Pressure" field is added to table T, the system first identifies the affected downstream assets using a lineage diagram, then activates the impact factor model (business importance weight 0.6, data update frequency weight 0.4) to score: 10 points for safety alert scenarios (assuming importance 10 points, real-time updates 10 points), 7.2 points for after-sales rights protection scenarios (assuming importance 8 points, daily updates 6 points), and 4.2 points for entertainment recommendation scenarios (assuming importance 5 points, weekly updates 3 points). The final impact level is determined as follows: safety alerts are classified as "high impact," after-sales rights protection as "medium impact," and entertainment recommendations as "low impact," helping to prioritize high-priority adaptation needs.
[0068] This invention achieves precise extraction of target data paths. Through the collaborative work of a scenario-data mapping model and a path focusing engine, it accurately filters out core data paths relevant to the current business scenario from a complex full lineage graph, avoiding interference from irrelevant data paths. Furthermore, it improves the business adaptability of data lineage analysis. The extraction process is entirely centered around the current business scenario, ensuring a high degree of alignment between the target data path and business needs, providing precise technical support for subsequent data sensitivity assessment. Additionally, it reduces the complexity of data processing and labor costs. It eliminates the need for manual screening of data paths from the full lineage graph, significantly improving data governance efficiency through an automated extraction process. Simultaneously, it avoids errors that may arise from manual screening, ensuring the accuracy and consistency of the target data path, and laying a solid foundation for the standardized and compliant processing of data assets in the new energy vehicle industry.
[0069] In some optional embodiments, the standardized processing method for vehicle data assets provided by the present invention further includes: Step g1: Use the data lineage analysis system to detect whether the target data link in the current business scenario has changed its dependent objects at a preset period. Step g2: Use the data lineage analysis system to detect whether the target data attribute information of the current business scenario has changed at a preset period. Step g3: When a change occurs in the dependent object or the attribute information, return to the step of mapping the target data link and target data attribute information to the current data lineage diffusion degree through the second mapping rule.
[0070] Specifically, to ensure that the calculation of data lineage diffusion degree always remains consistent with the actual situation of data flow, the method of this embodiment of the invention also includes a dynamic monitoring and updating mechanism for target data links and target data attribute information.
[0071] Preset periods (e.g., hourly, daily) can be flexibly configured according to the real-time data requirements of business scenarios. Changes in dependent objects refer to changes in the dependent entities in the target data chain, such as the data source, processing nodes, and downstream applications (e.g., adding a downstream referencing system, replacing the data collection source, adjusting the processing flow, etc.); changes in attribute information refer to changes in any item of the target data attribute information (e.g., adjusting the sensitivity level of the cited data, upgrading the importance level of the business scenario, the cumulative number of citations reaching a new threshold, etc.).
[0072] In this embodiment of the invention, taking the "battery safety early warning" business scenario of new energy vehicles as an example, the system automatically executes the monitoring process according to a preset cycle (e.g., every 2 hours) through the data lineage analysis system. The data lineage analysis system traverses the target data link of "battery voltage data" (e.g., "on-board sensor acquisition → battery management system processing → safety early warning dashboard display") and checks whether the dependent objects of each link have changed. For example, if a "vehicle health diagnosis platform" is added as a downstream application referencing the data, or the data processing node is switched from "local battery management system" to "cloud data processing center", it is determined that the dependent objects of the target data link have changed. Simultaneously, it checks whether the target data attribute information corresponding to the data link (e.g., "citation data sensitivity level: high, business scenario importance level: core, number of citations: 8 times") has been updated. For example, if the "number of citations" increases to 15 times, or the "business scenario importance level" is upgraded from "core" to "extremely high" due to business adjustments, it is determined that the attribute information of the target data has changed.
[0073] If any of the above changes (changes in dependent objects or attribute information) are detected, the system will automatically return to the second mapping rule execution step. Based on the updated target data link and target data attribute information, the system will recalculate the current data lineage diffusion degree. For example, after a new downstream application is added to "battery voltage data," the dependent objects of the data link change. The system will remap the data according to the new link characteristics (such as an increase in the number of downstream applications) and the updated attribute information (such as a synchronous increase in the number of references) to obtain a new lineage diffusion degree value. This ensures that the indicator can reflect the latest status of data flow in real time. Then, the system will recalculate the security score, re-label the data, perform the data verification process, and adjust the data's status in the database.
[0074] This invention, through a dynamic monitoring mechanism, achieves real-time calibration of data lineage diffusion, avoiding delays in grading criteria due to changes in data links or attribute information, and further improving the accuracy and timeliness of data sensitivity grading. Simultaneously, automated periodic monitoring and process triggering, requiring no manual intervention to complete updates, reduce the labor costs of data management and ensure that the grading of data assets always accurately adapts to the actual business flow, providing reliable support for the dynamic compliance management of data assets in the new energy vehicle industry.
[0075] Figure 9 A schematic diagram of an embodiment of a vehicle data asset standardization processing apparatus according to the present invention is shown. The apparatus 900 includes: The data asset acquisition module 901 is used to acquire initial data assets that match the current business scenario from the data asset catalog; The first matching module 902 is used to match a quantitative model according to the current business scenario. The quantitative model is built based on the data sensitivity factor. The second matching module 903 is used to match the data quality standards corresponding to the current business scenario from the quality standard knowledge base. The information in the quality standard knowledge base is dynamically updated when the update conditions are met. The scoring module 904 is used to calculate security scores for data fields in the initial data assets using a quantitative model. Sensitivity grading module 905 is used to label each data field with a sensitivity grading label based on the security score to obtain the first data field; The data quality inspection module 906 is used to perform data quality verification on the first data field based on data quality standards. The data import module 907 is used to load the second data field that has passed quality verification into the target database. In an optional embodiment, the sensitivity grading module 905 includes: A classification unit is used to determine the data category of the current data field. The regulatory knowledge base matching unit is used to match the corresponding target regulatory knowledge base according to the data category, and use the target regulatory knowledge base to check the legality of the current data field to obtain the current legality check result. The legal relevance calculation unit is used to map the current legality check result to the current legal relevance through the first mapping rule; The lineage analysis unit is used to extract target data links and target data attribute information related to the current business scenario using the data lineage analysis system. The lineage diffusion calculation unit is used to map the target data link and target data attribute information to the current data lineage diffusion degree through the second mapping rule; The scenario sensitivity analysis unit is used to query the target scenario sensitivity level defined for the current data field in the current business scenario. The scenario sensitivity calculation unit is used to map the target scenario sensitivity level to the current business scenario sensitivity through a third mapping rule. The security score calculation unit is used to integrate the current legal and regulatory relevance, current data lineage diffusion, and current business scenario sensitivity using the fourth mapping rule to obtain the target security score of the current data field; The sensitivity grading unit is used to map the target security score to the target sensitivity grading label of the current data field using the fifth mapping rule.
[0076] In one alternative embodiment, the apparatus further includes: The first update monitoring unit is used to monitor the update status of internal data standards and national data standards according to a preset cycle. The data monitoring unit is used to monitor changes in the distribution of initial data assets acquired at different time periods. The second update monitoring unit is used to query the update status of internal data standards and national data standards when the data distribution changes meet the preset change conditions. The knowledge base dynamic update unit is used to retrieve the updated standard content and update the quality standard knowledge base based on the updated standard content if the update status indicates that any data standard has been updated.
[0077] Figure 10 The diagram shows a structural schematic of an embodiment of the electronic device of the present invention. The specific embodiments of the present invention do not limit the specific implementation of the vehicle.
[0078] like Figure 10 As shown, the electronic device may include: a processor 1002, a communications interface 1004, a memory 1006, and a communications bus 1008.
[0079] The processor 1002, communication interface 1004, and memory 1006 communicate with each other via communication bus 1008. Communication interface 1004 is used to communicate with other network elements such as clients or other servers. The processor 1002 executes program 1010, specifically performing the relevant steps described above in the method embodiment.
[0080] Specifically, program 1010 may include program code, which includes computer-executable instructions.
[0081] The processor 1002 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention. The vehicle may include one or more processors of the same type, such as one or more CPUs; or processors of different types, such as one or more CPUs and one or more ASICs.
[0082] Memory 1006 is used to store program 1010. Memory 1006 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.
[0083] Specifically, program 1010 can be invoked by processor 1002 to execute the relevant steps described above in the method embodiment.
[0084] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.
[0085] The algorithms or displays provided herein are not inherently related to any particular computer, virtual system, or other device. Furthermore, the embodiments of this invention are not directed to any particular programming language.
[0086] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of the invention may be practiced without these specific details. Similarly, for the sake of brevity and to aid in understanding one or more aspects of the invention, in the description of exemplary embodiments of the invention above, various features of the embodiments are sometimes grouped together in a single embodiment, figure, or description thereof. The claims, which follow the detailed description, are hereby expressly incorporated into that detailed description, wherein each claim itself is a separate embodiment of the invention.
[0087] Those skilled in the art will understand that the modules in the device of the embodiment can be adaptively changed and placed in one or more devices different from that embodiment. Modules, units, or components in the embodiment can be combined into a single module, unit, or component, and further, they can be divided into multiple sub-modules, sub-units, or sub-components, except that at least some of such features and / or processes or units are mutually exclusive.
[0088] It should be noted that the above embodiments are illustrative of the invention and not restrictive, and that those skilled in the art can devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The invention can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the unit claims enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names. The steps in the above embodiments, unless otherwise specified, should not be construed as limiting the order of execution.
Claims
1. A standardized processing method for vehicle data assets, characterized in that, The method includes: Retrieve initial data assets from the data asset catalog that match the current business scenario; A quantitative model is matched based on the current business scenario, and the quantitative model is constructed based on a data sensitivity factor; The data quality standards corresponding to the current business scenario are matched from the quality standard knowledge base, and the information in the quality standard knowledge base is dynamically updated when the update conditions are met. The security score is calculated for the data fields in the initial data asset using the quantitative model. Based on the security score, sensitivity level labels are assigned to each data field to obtain the first data field; The data quality of the first data field is verified based on the aforementioned data quality standards. Load the second data field that passed the quality verification into the target database.
2. The method according to claim 1, characterized in that, The data sensitivity factors include business scenario sensitivity, legal and regulatory relevance, and data lineage diffusion. Business scenario sensitivity is an indicator of the sensitivity of a data field under different business scenarios. Legal and regulatory relevance is an indicator of the degree to which a data field complies with laws and regulations. Data lineage diffusion is an indicator of the extent and depth to which data is cited.
3. The method according to claim 2, characterized in that, The steps for constructing the quantization model based on the data sensitivity factor include: Adapt the corresponding legal knowledge base to different categories of data fields; Create a data lineage analysis system, which is used to extract data links associated with business scenarios and data attribute information in the data links. The data links are used to characterize the breadth of data citation, and the data attribute information is used to characterize the depth of data citation. The data attribute information includes at least one of the following: sensitivity level of cited data, importance level of business scenario, and number of citations. Define scenario sensitivity levels for different categories of data fields in various business scenarios; Define a first mapping rule, which is used to map the legality check results of the data fields in the legal knowledge base to the legal relevance degree; A second mapping rule is defined, which is used to map the data link and data attribute information extracted by the data lineage analysis system from the business scenario to the data lineage diffusion degree. Define a third mapping rule, which is used to map the scenario sensitivity level of a data field in any business scenario to the business scenario sensitivity. A fourth mapping rule is defined, which is used to integrate the legal and regulatory relevance, the data lineage diffusion, and the business scenario sensitivity to obtain the security score of the data field; A fifth mapping rule is defined, which is used to map the security score to the corresponding sensitivity rating label.
4. The method according to claim 3, characterized in that, The quantitative model is used to calculate security scores for the data fields in the initial data asset, and based on the security scores, sensitivity grading labels are assigned to each data field, including: Determine the data category of the current data field; Match the corresponding target legal knowledge base according to the data category, and use the target legal knowledge base to check the legality of the current data field to obtain the current legality check result; The current legality check result is mapped to the current legality relevance to laws and regulations using the first mapping rule; The data lineage analysis system is used to extract target data links and target data attribute information associated with the current business scenario; The target data link and the target data attribute information are mapped to the current data lineage diffusion degree using the second mapping rule; Query the target scenario sensitivity level defined for the current data field in the current business scenario; The target scenario sensitivity level is mapped to the current business scenario sensitivity using the third mapping rule. By fusing the current legal and regulatory relevance, the current data lineage diffusion, and the current business scenario sensitivity using the fourth mapping rule, the target security score of the current data field is obtained. The target security score is mapped to the target sensitivity classification label of the current data field using the fifth mapping rule.
5. The method according to claim 4, characterized in that, The method further includes: The data lineage analysis system detects at preset intervals whether the target data link in the current business scenario has changed its dependent objects. The data lineage analysis system detects whether the target data attribute information of the current business scenario has changed at a preset period. When a change occurs in the dependent object or the attribute information, return to the step of mapping the target data link and the target data attribute information to the current data lineage diffusion degree through the second mapping rule.
6. The method according to claim 1, characterized in that, The steps for dynamically updating the quality standard knowledge base include: Monitor the update status of internal data standards and national data standards according to a preset cycle; Monitor changes in the distribution of initial data assets acquired at different time periods; When the data distribution change meets the preset change conditions, query the update status of the internal data standard and the national data standard; If the update status indicates that any data standard has been updated, then the updated standard content is obtained, and the quality standard knowledge base is updated based on the updated standard content.
7. The method according to claim 3, characterized in that, The creation of the data lineage analysis system includes: Based on the full lifecycle relationships between data from generation to use, a full lineage graph is created to describe the dependencies between data. The full lineage graph includes the complete flow path of all data assets in the data asset catalog. Based on the business keywords corresponding to each business scenario, query the corresponding scenario data asset set for each business scenario. Create a scenario-data mapping model based on the correspondence between business scenarios and scenario data asset sets; A path focusing engine is created based on graph traversal technology. The path focusing engine is used to cut out data flow paths that are not related to a certain target business scenario from the full lineage graph according to the scenario-data mapping model, and extract the data links associated with the target business scenario.
8. The method according to claim 7, characterized in that, The data lineage analysis system is used to extract target data links associated with the current business scenario, including: The current business scenario is input into the scenario-data mapping model, so that the scenario-data mapping model outputs the corresponding current scenario data asset set; The path focusing engine reads the current scene data asset set and extracts the path and node where the current scene data asset set is located from the full lineage graph to form the target data link.
9. The method according to claim 1, characterized in that, The quality standard knowledge base also includes data correction standards corresponding to the current business scenario, and the method further includes: Based on the aforementioned data correction standard, the third data field that fails the quality verification is standardized. The standardized fourth data field is used as the first data field, and the step of performing data quality verification on the first data field based on the data quality standard is returned.
10. The method according to claim 9, characterized in that, The method further includes: If the fourth data field obtained after correcting the preset number of times still fails the data quality check, then the loading of the fourth data field into the target database will be blocked, and an alarm message will be output. When the data correction standard corresponding to the current business scenario is updated, the fourth data field is re-standardized based on the updated data correction standard. The standardized fourth data field is used as the first data field, and the step of performing data quality verification on the first data field based on the data quality standard is returned.
11. A standardized processing device for vehicle data assets, characterized in that, The device includes: The data asset acquisition module is used to acquire initial data assets that match the current business scenario from the data asset catalog; The first matching module is used to match a quantization model according to the current business scenario, wherein the quantization model is constructed based on a data sensitivity factor. The second matching module is used to match the data quality standards corresponding to the current business scenario from the quality standard knowledge base. The information in the quality standard knowledge base is dynamically updated when the update conditions are met. The scoring module is used to calculate a security score for the data fields in the initial data asset using the quantitative model. The sensitivity grading module is used to label each data field with a sensitivity grading tag based on the security score to obtain the first data field; The data quality inspection module is used to perform data quality verification on the first data field based on the data quality standard. The data import module is used to load the second data field that has passed quality verification into the target database.
12. An electronic device, characterized in that, include: A memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, the processor executing the computer instructions to perform the method of any one of claims 1 to 10.