Data transaction security assessment method and system

By performing multi-dimensional feature analysis and dynamic risk calculation on field information in data transactions, the adaptability and accuracy issues of data transaction security assessment in existing technologies have been resolved. Fine-grained security assessment and adaptive protection have been achieved, improving the intelligence level and protection efficiency of data transaction security.

CN121502810APending Publication Date: 2026-02-10ZHIYIN (NANJING) DATA CO LTD

Patent Information

Application Number
CN202511700198.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-19
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

In existing technologies, data transaction security assessment methods cannot adapt to the dynamic and ever-changing data ecosystem, and cannot accurately reflect changes in local high-risk content and transaction scenarios, resulting in assessment results that are too general and lack specificity.

Method used

By extracting field information from the dataset to be traded, performing multi-dimensional feature analysis, generating sensitivity vectors and classifying sensitivity levels, and dynamically adjusting them in conjunction with transaction-related information, the risk values ​​of data content, field associations, and transaction environment are calculated. Based on weight coefficients, weighted calculations are performed, and the evaluation strategy and parameters are dynamically adjusted.

Benefits of technology

It enables fine-grained security assessment of different types and sensitivities of data during data transactions, improves the accuracy of assessment and dynamic response capabilities, builds a self-learning and self-evolving security management system, and enhances the security protection capabilities of the entire data transaction process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121502810A_ABST
    Figure CN121502810A_ABST
Patent Text Reader

Abstract

The invention discloses a data transaction security assessment method and system, and relates to the technical field of data security and intelligent assessment. The method comprises the steps of performing feature analysis on a data field, determining a sensitivity weight and generating a sensitivity vector; the fields are graded according to the sensitivity vectors, and dynamic adjustment is carried out in combination with transaction related information to obtain a grading result; calculating a data content risk, a field association risk, and a transaction environment risk based on the sensitivity feature set; determining a weight according to a transaction scene, and weighting various risks to obtain a comprehensive score; and judging the risk level according to the comprehensive score and a threshold value dynamically determined based on historical data. According to the invention, through hierarchical sensitivity modeling and dynamic risk calculation, fine-grained security assessment of different types and sensitivity data is realized, and the risk result can be adaptively adjusted along with changes of data contents, transmission paths and transaction environments, so that the accuracy, real-time performance and scene adaptability of assessment are improved, and the risk assessment efficiency is improved. The defects of a static evaluation method in the prior art are overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data security and intelligent assessment technology, specifically a data transaction security assessment method and system. Background Technology

[0002] With the development of the data element market, data trading has gradually become an important way for data assets to circulate and realize their value. In typical data trading scenarios, the two parties usually upload, share, download, and access data through a third-party platform. However, the types of data involved in the transactions are complex and diverse, and different types of data have significantly different levels of sensitivity. For example, personal identity information, financial records, medical data, and corporate operational data are all highly sensitive data. Once leaked, they can not only cause serious economic losses but also potentially lead to privacy violations and compliance risks.

[0003] In existing technologies, data transaction security assessments primarily rely on static assessment methods, specifically two types: one is rule-based or blacklist / whitelist-based static security review, which determines transaction security by manually pre-setting key fields, risk words, or source rules. This static rule configuration method is poorly adaptable to new data structures and complex transaction paths, and struggles to cope with dynamic and ever-changing data ecosystems. The other type uses fixed algorithms or models to perform risk calculations on the entire dataset. This method uses uniform assessment standards and preset parameters to statically assess all data, often ignoring differences in the sensitivity of fields within the data and failing to adjust assessment strategies in real time according to dynamic factors such as the transaction environment and data usage. This results in assessments that are too general and lack specificity, failing to accurately reflect local high-risk content and the actual risk level that changes with the transaction scenario. The limitations of this static assessment method are particularly prominent in the context of increasingly complex data transaction scenarios and increasingly diversified data flow paths. Summary of the Invention

[0004] The purpose of this application is to provide a data transaction security assessment method and system to solve the problems mentioned in the background art.

[0005] In a first aspect, one embodiment of this application provides a data transaction security assessment method, which includes: extracting field information from a dataset to be traded; performing feature analysis on the field information to determine the sensitivity weight of each field and generating a sensitivity vector; classifying each field into sensitivity levels according to the sensitivity vector, and dynamically adjusting the sensitivity levels based on transaction-related information to obtain a dynamic grading result; integrating the sensitivity vector and the dynamic grading result to generate a sensitivity feature set; calculating the data content risk value, field association risk value, and transaction environment risk value based on the sensitivity feature set; determining weight coefficients according to transaction scenario information, and weighting the data content risk value, field association risk value, and transaction environment risk value according to the weight coefficients to obtain a comprehensive score; and determining the risk level based on the comprehensive score and a preset threshold, wherein the preset threshold is dynamically determined based on historical data.

[0006] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: triggering corresponding protection strategies based on risk levels; collecting transaction execution results and comparing and analyzing the comprehensive score with the transaction execution results; and adjusting sensitivity weights, correlation relationships, weight coefficients, and preset thresholds based on the comparative analysis results.

[0007] In conjunction with the first aspect, in certain implementations of the first aspect, feature analysis is performed on the field information to determine the sensitivity weight of each field and generate a sensitivity vector. This includes: performing multi-dimensional feature analysis on the field information to identify the sensitivity category of each field; determining the basic weight of each field based on the sensitivity category; correcting the basic weight to obtain the sensitivity weight of each field; and generating a sensitivity vector based on the sensitivity weight of each field. The multi-dimensional feature analysis includes format pattern, statistical features, and semantic analysis. The sensitivity category represents the sensitivity of the field content. The basic weight is a preset initial weight, the sensitivity weight is a quantified field sensitivity value, and the sensitivity vector is the set of sensitivity weights for each field.

[0008] In conjunction with the first aspect, in some implementations of the first aspect, transaction-related information includes data usage tags, field access frequencies, and historical risk event records. Each field is classified into sensitivity levels based on a sensitivity vector, and the sensitivity levels are dynamically adjusted based on transaction-related information to obtain a dynamic classification result. This includes: classifying each field into sensitivity levels based on the correspondence between the sensitivity weights of each field in the sensitivity vector and preset weight ranges; and dynamically adjusting the sensitivity levels based on at least one of the following methods: the usage risk coefficient corresponding to the data usage tag, the comparison result of the field access frequency with a preset frequency threshold, and the risk level of historical risk event records, to obtain a dynamic classification result.

[0009] In conjunction with the first aspect, in some implementations of the first aspect, the sensitivity vector and the dynamic grading results are integrated to generate a sensitivity feature set, including: associating and mapping the weights of each field in the sensitivity vector with the corresponding dynamic grading results, and combining auxiliary features to generate a sensitivity feature set; wherein, the auxiliary features include at least one of data usage labels, access frequency statistics, and historical risk event records.

[0010] In conjunction with the first aspect, in some implementations of the first aspect, based on the sensitivity feature set, the data content risk value, field association risk value, and transaction environment risk value are calculated, including: calculating the data content risk value based on the sensitivity feature set and the scale information of the dataset to be traded; constructing the association relationship between different fields in the sensitivity feature set and calculating the field association risk value based on the association relationship; and calculating the transaction environment risk value based on the environmental parameters of the data transaction.

[0011] In conjunction with the first aspect, in some implementations of the first aspect, the data content risk value is calculated based on the size information of the sensitivity feature set and the dataset to be traded; the association relationship between different fields in the sensitivity feature set is constructed, and the field association risk value is calculated based on the association relationship; the transaction environment risk value is calculated based on the environmental parameters of the data transaction, including: calculating the data content risk value according to the sensitivity weight of each field in the sensitivity vector, the number of fields in the dataset to be traded, the number of data records, and the number of transaction participants; calculating the field association risk value according to the association coefficient in the association relationship and the sensitivity weight of each field in the sensitivity vector; and calculating the transaction environment risk value according to the transmission path security level, transmission protocol strength, and access subject reputation in the environmental parameters.

[0012] In conjunction with the first aspect, in certain implementations of the first aspect, weighting coefficients are determined based on transaction scenario information. The data content risk value, field association risk value, and transaction environment risk value are then weighted according to these weighting coefficients to obtain a comprehensive score. This includes: determining a first weighting coefficient corresponding to the data content risk value, a second weighting coefficient corresponding to the field association risk value, and a third weighting coefficient corresponding to the transaction environment risk value based on the transaction scenario type and industry category in the transaction scenario information; multiplying the data content risk value by the first weighting coefficient to obtain a first weighted value, multiplying the field association risk value by the second weighting coefficient to obtain a second weighted value, and multiplying the transaction environment risk value by the third weighting coefficient to obtain a third weighted value; and summing the first, second, and third weighted values ​​to obtain a comprehensive score.

[0013] Secondly, one embodiment of this application provides a data transaction security assessment system, which includes: a field extraction module for extracting field information from a dataset to be traded; a sensitivity analysis module for performing feature analysis on the field information, determining the sensitivity weight of each field, and generating a sensitivity vector; a dynamic grading module for dividing each field into sensitivity levels according to the sensitivity vector, and dynamically adjusting the sensitivity levels based on transaction-related information to obtain a dynamic grading result; a feature integration module for integrating the sensitivity vector and the dynamic grading result to generate a sensitivity feature set; a risk calculation module for calculating the data content risk value, field association risk value, and transaction environment risk value based on the sensitivity feature set; a scoring calculation module for determining weight coefficients according to transaction scenario information, and weighting the data content risk value, field association risk value, and transaction environment risk value according to the weight coefficients to obtain a comprehensive score; and a risk determination module for determining the risk level based on the comprehensive score and a preset threshold, wherein the preset threshold is dynamically determined based on historical data.

[0014] In conjunction with the second aspect, in some implementations of the second aspect, the system further includes: a protection execution module, used to trigger corresponding protection strategies according to the risk level; a result acquisition module, used to collect transaction execution results and compare and analyze the comprehensive score with the transaction execution results; and a parameter adjustment module, used to adjust the sensitivity weight, correlation, weight coefficient and preset threshold based on the comparative analysis results.

[0015] Compared with the prior art, the beneficial effects of this application are: 1. This application introduces a hierarchical sensitivity modeling and dynamic risk calculation method to achieve fine-grained security assessment of different types and sensitivities of data during data transactions. This method can adaptively adjust the risk assessment results based on real-time changes in data content characteristics, transmission paths, and the transaction environment, thereby significantly improving the accuracy, dynamic response capability, and scenario adaptability of the assessment, effectively overcoming the limitations of traditional static assessment methods.

[0016] 2. This application constructs an "assessment-strategy-feedback" security management system, which can automatically trigger corresponding protection strategies (such as encryption, desensitization, access control, etc.) based on the risk assessment results, and continuously optimize model parameters and strategy thresholds through assessment feedback, forming the system's self-learning and self-evolution capabilities. This not only improves the platform's intelligence and automation level, but also effectively enhances the security protection capabilities and risk management efficiency of the entire data transaction process. Attached Figure Description

[0017] Figure 1 A flowchart illustrating a data transaction security assessment method provided in an embodiment of this application; Figure 2A flowchart illustrating a data transaction security assessment method provided in another embodiment of this application; Figure 3 This is a flowchart illustrating a process for performing feature analysis on field information, determining the sensitivity weight of each field, and generating a sensitivity vector, as provided in an embodiment of this application. Figure 4 This is a flowchart illustrating how, according to an embodiment of the present application, fields are divided into sensitivity levels based on a sensitivity vector, and the sensitivity levels are dynamically adjusted based on transaction-related information to obtain a dynamic classification result. Figure 5 A flowchart illustrating the process of calculating data content risk value, field association risk value, and transaction environment risk value based on a sensitivity feature set, as provided in an embodiment of this application; Figure 6 This is a flowchart illustrating a process for determining weighting coefficients based on transaction scenario information, weighting data content risk values, field association risk values, and transaction environment risk values ​​according to weighting coefficients to obtain a comprehensive score, as provided in one embodiment of this application. Figure 7 A schematic diagram of the structure of a data transaction security assessment system provided in an embodiment of this application; Figure 8 A schematic diagram of the structure of a data transaction security assessment system provided in another embodiment of this application. Detailed Implementation

[0018] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0019] Figure 1 This is a flowchart illustrating a data transaction security assessment method provided in one embodiment of this application. Figure 1 As shown in the embodiments of this application, the data transaction security assessment method includes the following steps: Step 100: Extract field information from the dataset to be traded.

[0020] It should be understood that the dataset to be traded refers to the collection of data prepared for trading on a data trading platform. It exists in various forms such as relational database tables, CSV files, JSON, or XML, and contains complete data structures and content. Field information refers to the structured metadata information of each field in the dataset, including field name identifiers, data type definitions, field length limits, whether null values ​​are allowed, primary key and foreign key relationships, index information, and other constraints.

[0021] Step 101: Perform feature analysis on the field information, determine the sensitivity weight of each field, and generate a sensitivity vector.

[0022] It should be understood that a sensitivity vector is a set of vectors composed of the sensitivity weights of all fields in the dataset to be traded, arranged in order.

[0023] Step 102: Divide each field into sensitivity levels according to the sensitivity vector, and dynamically adjust the sensitivity levels based on transaction-related information to obtain dynamic classification results.

[0024] It should be understood that transaction-related information refers to dynamic information related to data transaction scenarios and usage, including data usage tags, field access frequency statistics, and historical risk event records, which are used to make real-time corrections to the basic sensitivity level.

[0025] Step 103: Integrate the sensitivity vector and dynamic grading results to generate a sensitivity feature set.

[0026] Step 104: Based on the sensitivity feature set, calculate the data content risk value, field association risk value, and transaction environment risk value.

[0027] It should be understood that the data content risk value reflects the exposure risk of data being obtained or accessed by unauthorized parties during a transaction. This risk mainly depends on the sensitivity of the data, the size of the data, and the number of participants involved. The field association risk value refers to the aggregated risk quantification value calculated based on the association coefficients and sensitivity weights of each field in the sensitivity association matrix, reflecting the non-linear risk growth effect generated when multiple sensitive fields appear simultaneously. The transaction environment risk value refers to the transmission risk quantification value calculated based on environmental parameters such as transmission path security level, transmission protocol strength, network environment coefficient, and the reputation of the accessing entity, reflecting the degree of threat of data interception or tampering during transmission.

[0028] Step 105: Determine the weighting coefficients based on the transaction scenario information, and calculate the comprehensive score by weighting the data content risk value, field association risk value, and transaction environment risk value according to the weighting coefficients.

[0029] It should be understood that transaction scenario information refers to the descriptive information of the specific application scenario and business environment in which the data exchange takes place, including transaction scenario type such as cross-border transactions, internal enterprise circulation, etc., as well as industry category such as financial, medical, education, etc.

[0030] Step 106: Determine the risk level based on the comprehensive score and the preset threshold. The preset threshold is dynamically determined based on historical data.

[0031] The beneficial effects of this embodiment are as follows: By introducing field-level sensitivity quantification analysis and multi-dimensional feature recognition, it achieves accurate identification and hierarchical management of fields with different levels of sensitivity in data transactions, providing a fine-grained data foundation for dynamic risk assessment and differentiated security protection, and effectively solving the problem of insufficient attention to the differences in the internal structure of data in traditional methods.

[0032] Figure 2 This is a flowchart illustrating a data transaction security assessment method provided in another embodiment of this application. Figure 2 As shown in the embodiments of this application, the data transaction security assessment method includes the following steps: Step 200: Trigger the corresponding protection strategy based on the risk level.

[0033] Specifically, the system will comprehensively assess security scores. The risk level is determined based on the comparison with a preset threshold and the matching result within the threshold range. Greater than or equal to the high-risk threshold When a data point is deemed high-risk, a combination of mandatory protection strategies is automatically triggered, including: mandatory desensitization of the first-level high-sensitivity fields marked by multi-dimensional sensitivity features, with the desensitization method automatically selecting mask desensitization or placeholder replacement based on the field type; detection of the current transport protocol strength, requiring an upgrade to the transport protocol or application-layer encryption if it is lower than the TLS 1.3 standard; imposing rate limiting on the access frequency of high-risk data, ensuring that the number of accesses by a single requester within a unit of time does not exceed the rate limiting threshold determined based on sensitivity and the requester's reputation level; and for data with a risk score significantly exceeding [a certain threshold], [further measures are taken]. For extremely high-risk transactions, the automated processing flow is suspended and the transaction request is pushed to a manual review queue. Less than or equal to Less than When the risk level is determined to be medium, moderate-strength protection measures are implemented, including partial data masking, requiring TLS 1.2 or higher for the transport protocol, relatively lenient access limiting, and enhanced logging. Less than If the risk level is low, the transaction is allowed to proceed normally, but a complete audit log is retained.

[0034] Step 201: Collect transaction execution results and compare and analyze the comprehensive score with the transaction execution results.

[0035] Specifically, the system establishes a transaction lifecycle tracking system, continuously recording data for each transaction from evaluation to execution and final outcome. Recorded data includes input feature groups, intermediate calculation result groups, final score, risk level determination result, a list of triggered protection strategies, strategy execution status, and the final transaction result. The system periodically performs backtesting analysis on historical transactions, integrating the comprehensive security score... The analysis compares the risk level with actual security outcomes. When a transaction is rated as high-risk but no anomalies occur during the process, it is considered an overestimation or misjudgment. The characteristics of the transaction are extracted, and the reasons for the high score are analyzed, including excessively high sensitivity weights, inaccurate correlation coefficients for clustered risks, or overly pessimistic assessments of environmental risk parameters. When a transaction is rated as low-risk but data breaches or misuse occur, it is considered an underestimation or missed detection. The root causes of the low score are analyzed, including unidentified sensitive fields, new risk patterns not covered by existing models, or vulnerabilities in protection strategies. The system statistically compares and contrasts the misjudgment and missed detection rates, using this data as a basis for adjusting model parameters.

[0036] Step 202: Based on the comparative analysis results, adjust the sensitivity weight, correlation, weight coefficient and preset threshold.

[0037] Specifically, the system employs a supervised learning framework, using historical transaction data as training samples to automatically optimize model parameters. At the sensitivity weight adjustment level, the system analyzes the field features involved in cases of underestimation and missed detection, extracts keywords from these field names, adds keywords not in the existing sensitivity dictionary to the dictionary, and simultaneously calculates the hit frequency and accuracy of each keyword in the existing dictionary, removing words with extremely low hit frequency and low accuracy from the dictionary. At the correlation adjustment level, the system calculates the co-occurrence frequency of different field combinations in historical transaction data and the corresponding security event occurrence rate, and re-estimates the correlation coefficients in the sensitivity correlation matrix C based on the statistical results. For field combinations where the co-occurrence rate is significantly higher than when they occur alone, the correlation coefficient is increased; conversely, the coefficient is decreased. Regarding weight coefficient adjustment, the system adjusts the weight configuration for each scenario based on its evaluation performance in different scenarios. , , When the false negative rate for a specific risk in a particular scenario is high, the weighting coefficient of the corresponding risk in that scenario is appropriately increased, with the adjustment amount determined based on the severity of the false negatives. At the preset threshold adjustment level, the system dynamically adjusts the threshold based on the false positive and false negative rates. and For the threshold, iterate through all possible values ​​of the historical scores, calculate the corresponding false positive rate and false negative rate for each value as a candidate threshold, and select the value that minimizes the overall loss function as the new threshold. ; in, The false negative rate, For the false positive rate, The weight for missed detections (take the larger value, such as 5.0). The misjudgment weight is set to 1.0. Parameter adjustment adopts an incremental learning approach. The system sets an optimization period and extracts the newly added transaction data within the optimization period during each update. Incremental training is performed based on the current model parameters, and the adjusted parameters are then applied to the model to ensure that the model always maintains its timeliness based on the latest data.

[0038] The beneficial effects of this embodiment are as follows: By constructing a security management system of "assessment-execution-feedback-optimization", a deep coupling between risk assessment results and actual security protection is achieved, as well as continuous adaptive optimization of model parameters. This not only enables the automatic triggering of differentiated protection strategies based on risk levels to achieve precise prevention and control, but also allows for the identification of assessment deviations and dynamic adjustment of model parameters through systematic retrospective analysis of transaction execution results. This forms a self-learning and self-evolution capability, effectively improving the intelligence level of the data transaction security assessment system and the accuracy and reliability of its long-term operation.

[0039] Figure 3 This is a flowchart illustrating a process for performing feature analysis on field information, determining the sensitivity weight of each field, and generating a sensitivity vector, as provided in one embodiment of this application. Figure 3 As shown in the embodiments of this application, the data transaction security assessment method performs feature analysis on field information, determines the sensitivity weight of each field, and generates a sensitivity vector, including the following steps: Step 300: Perform multi-dimensional feature analysis on the field information to identify the sensitive categories of each field.

[0040] Specifically, the system automatically determines the sampling ratio based on the dataset size. A 10% sampling rate is used for datasets under one million records, while the sampling rate is reduced for larger datasets to balance analytical accuracy and computational efficiency. For the sampled field values, the system extracts features from three dimensions. In terms of format and pattern, the system has a pre-built format recognition rule library that can identify common sensitive information formats such as mobile phone numbers, ID card numbers, email addresses, bank card numbers, IP addresses, URLs, and postal codes. For field values ​​that conform to a specific format pattern, the system records the format type and matching confidence. Format conformity verification and validity checks are performed using a regular expression template library, and the format matching rate—the proportion of samples conforming to that format—is calculated. In terms of statistical features, the system calculates statistical indicators such as the number of unique values, numerical distribution characteristics, missing rate, and duplication rate of the field. These indicators are used to assess the data quality and potential use of the field. In terms of semantic analysis, the system maintains a continuously updated sensitivity dictionary containing sensitive keywords and their synonyms. A hierarchical organizational structure is used to categorize sensitive keywords according to their sensitivity level and category. The system performs deep semantic understanding by combining field names and content, and uses natural language processing techniques for precise matching, fuzzy matching, and semantic similarity matching. Chinese field names are segmented to extract key semantic components for matching. The system then uses a fusion algorithm to comprehensively determine the matching results from these three dimensions. Weighting coefficients are assigned to field name semantic matching, format pattern matching, and statistical feature determination, and a weighted comprehensive score is calculated. This score is combined with a confidence threshold to determine the final sensitivity category identification result. Sensitive categories include strong identity-identifying information, general personal information, and secondary information.

[0041] Step 301: Determine the basic weight of each field based on the sensitivity category.

[0042] Specifically, the system determines the basic weight of the field based on the sensitive categories identified in step 300, with different basic weights corresponding to different sensitive categories. For strong identity information categories such as ID card numbers and bank card numbers, the basic weight is set to 0.9; for general personal information categories such as names and mobile phone numbers, the basic weight is set to 0.7; and for secondary information categories such as addresses and email addresses, the basic weight is set to 0.5.

[0043] Step 302: Correct the basic weights to obtain the sensitivity weights of each field.

[0044] Specifically, the system adjusts the base weights based on matching confidence. Fields with high confidence have their weights increased, while those with low confidence have their weights decreased, with adjustments within ±0.1. The system also considers field type weighting factors to modify the base weights: primary key and unique index fields have their weights multiplied by an enhancement coefficient of 1.2, foreign key fields by a coefficient of 1.1, and ordinary fields retain their original weights. The system incorporates content distribution characteristics into the weight calculation: fields with over 90% uniqueness receive an additional 0.05 weight, and fields with a missing value rate below 5% receive an additional 0.05 weight. Through this calculation process, the system generates a sensitivity weight between 0 and 1 for each field. This sensitivity weight is a quantified value reflecting the field's sensitivity, structural importance, and data integrity characteristics.

[0045] Step 303: Generate a sensitivity vector based on the sensitivity weights of each field.

[0046] Specifically, the system constructs a sensitivity vector from the sensitivity weights of all fields. ,in, Indicates the first Sensitivity weights for each field, This represents the total number of fields. The sensitivity vector is the set of sensitivity weights for each field, expressed in a vectorized form. Multi-dimensional feature analysis includes format patterns, statistical features, and semantic analysis. Sensitive categories characterize the sensitivity of field content. Base weights are preset initial weights, and sensitivity weights are quantified field sensitivity values. The sensitivity vector is the set of sensitivity weights for each field.

[0047] The beneficial effects of this embodiment are as follows: By establishing a multi-dimensional feature analysis framework and hierarchical weight correction, the system achieves automated and accurate identification and quantitative evaluation of sensitive attributes of data fields. The system comprehensively utilizes format pattern recognition, statistical feature analysis, and semantic understanding technologies to accurately capture the sensitive features of fields and transform them into standardized sensitivity weights. Simultaneously, through multi-layered correction based on matching confidence levels, field structure attributes, and content distribution characteristics, the system ensures the scientific rigor and accuracy of weight calculations, laying a solid data foundation for dynamic grading and risk assessment, and significantly improving the automation level of sensitivity identification and the reliability of evaluation results.

[0048] Figure 4 This is a flowchart illustrating an embodiment of the present application, showing how to divide each field into sensitivity levels based on a sensitivity vector, and dynamically adjust the sensitivity levels based on transaction-related information to obtain a dynamic classification result. Figure 4As shown in the embodiments of this application, the data transaction security assessment method includes transaction-related information such as data usage tags, field access frequency, and historical risk event records. Each field is divided into sensitivity levels based on a sensitivity vector, and the sensitivity levels are dynamically adjusted based on transaction-related information to obtain a dynamic classification result. The method includes the following steps: Step 400: Based on the correspondence between the sensitivity weights of each field in the sensitivity vector and the preset weight range, divide each field into sensitivity levels.

[0049] Specifically, the system uses the sensitivity vector generated in step 303. The sensitivity weights of each field are matched against preset weight ranges to classify data fields into three sensitivity levels. When a field's sensitivity weight is greater than or equal to 0.7, it is classified as Level 1 (High Sensitivity). These fields typically contain strongly identifying information or core privacy data, and their leakage will have serious consequences. When a field's sensitivity weight is between 0.3 and 0.7, it is classified as Level 2 (Medium Sensitivity). These fields contain general personal information or business-sensitive data, and their leakage will have a moderate impact. When a field's sensitivity weight is less than 0.3, it is classified as Level 3 (Low Sensitivity). These fields contain publicly available or weakly sensitive information, and their leakage will have a relatively small impact.

[0050] Step 401: Dynamically adjust the data usage risk coefficient corresponding to the data usage label, the comparison result of the field access frequency and the preset frequency threshold, and the risk level of historical risk event records to obtain a dynamic classification result.

[0051] Specifically, the system performs real-time corrections on the basic classification results from three dimensions. In the data usage label dimension, the system obtains the data usage indicated by the data provider and sets usage risk coefficients based on different uses: 1.5 for data used for public display, 1.3 for data used for identity authentication, 0.8 for data used for statistical analysis, and 0.7 for data used for academic research. The system then multiplies the field's sensitivity weight by the corresponding usage risk coefficient to re-determine the classification, achieving dynamic adjustment based on usage. In the access frequency statistics dimension, the system continuously monitors the access frequency of each field on the data trading platform, including query counts, download counts, and API call counts. The access frequency threshold is set at the 90th percentile of the access frequency of all fields. When the number of accesses to a field exceeds this preset threshold in the past month, its sensitivity level is automatically increased by one level, for example, from Level 2 (medium sensitivity) to Level 1 (high sensitivity), or from Level 3 (low sensitivity) to Level 2 (medium sensitivity), ensuring that frequently accessed hot data receives stricter protection. At the historical transaction record level, the system establishes a field-level historical risk profile, recording security event information for each field in past transactions, including data breaches, user complaints, and reports of unauthorized use. For fields that have experienced data breaches within the past year, the sensitivity level is forcibly increased to Level 1 (High Sensitivity). For fields that have received user complaints within the past six months, the sensitivity weight is increased by 0.2. For fields with reports of unauthorized use, the sensitivity weight is increased by 0.15. The risk level is determined based on the type and timing of the security event. Through dynamic correction using at least one of the above three dimensions, the system ultimately generates a dynamic risk rating result.

[0052] The beneficial effects of this embodiment are as follows: By establishing a grading standard based on weight ranges and multi-dimensional dynamic correction, precise hierarchical management and adaptive adjustment of data field sensitivity are achieved. The system not only performs initial grading based on the inherent sensitivity attributes of fields, but also corrects the grading results by integrating real-time dynamic information such as data usage, access frequency, and historical risk records. This ensures that sensitivity assessments fully reflect the true risk status of fields in actual transaction scenarios. This grading approach, combining static characteristics with a dynamic environment, ensures that security protection strategies can be flexibly adjusted according to changes in data usage, significantly improving the timeliness and scenario adaptability of risk assessments, and providing a scientific basis for differentiated and precise protection.

[0053] In another embodiment, the sensitivity vector and dynamic grading results are integrated to generate a sensitivity feature set.

[0054] Specifically, the weights of each field in the sensitivity vector are mapped to the corresponding dynamic grading results, and combined with auxiliary features to generate a sensitivity feature set. The system then uses the sensitivity vector generated in step 303... Compared with the dynamic grading results generated in step 401 Integrate the sensitivity feature set. Organized using a triplet structure, represented as: ,in This is an auxiliary feature set.

[0055] The system establishes a mapping relationship between sensitivity vectors and dynamic grading results, for each field. Its sensitivity weight With dynamic sensitivity level Perform a one-to-one correspondence to form field-level sensitivity feature pairs. This feature pair contains both quantitative weight information and qualitative ranking information. The system further incorporates auxiliary features into the sensitivity feature set, which, together with the sensitivity vector and dynamic ranking results, constitutes a complete feature set.

[0056] The auxiliary features include at least one of the following: data usage tags, access frequency statistics, and historical risk event records. Data usage tags determine the actual exposure scenarios of the field, access frequency statistics reflect the field's usage frequency, and historical risk event records establish a risk history profile for the field. The set of auxiliary features may also include field relationship information, used to construct a relationship matrix in risk calculation.

[0057] The beneficial effects of this embodiment are as follows: By constructing a sensitivity feature set with a triplet structure, the system achieves multi-dimensional information integration and structured expression of field sensitivity attributes. The system organically integrates quantitative sensitivity weights, qualitative dynamic level classifications, and diverse auxiliary feature information, forming a comprehensive feature representation system that includes both the inherent sensitivity features of the field and reflects the actual usage environment. This structured feature set not only provides data input for risk calculation, but more importantly, by integrating dynamic information such as data usage, access frequency, and historical risk records, it enables risk assessment to fully consider multiple risk factors of the field in real transaction scenarios. This significantly improves the comprehensiveness and accuracy of risk quantification, laying a solid data foundation for achieving precise security protection decisions.

[0058] Figure 5 This is a flowchart illustrating a process for calculating data content risk values, field association risk values, and transaction environment risk values ​​based on a sensitivity feature set, as provided in an embodiment of this application. Figure 5 As shown in the embodiments of this application, the data transaction security assessment method calculates data content risk values, field association risk values, and transaction environment risk values ​​based on a sensitivity feature set, including the following steps: Step 500: Calculate the data content risk value based on the sensitivity feature set and the size information of the dataset to be traded.

[0059] Specifically, the system constructs a quantitative model for data content risk values, and the data content risk values... The calculation formula is: ; in, Representation field The sensitivity weight, ranging from 0 to 1, comes from the sensitivity vector in the sensitivity feature set. This weight reflects the inherent sensitivity of the field. The higher the sensitivity weight, the greater the contribution of the field to the exposure risk calculation. Representation field The weight coefficients are determined based on the importance and functional role of the fields in the data structure. The primary key field and the unique identifier field have a weight coefficient of 1.5 because they have the function of individual identification, while the weight coefficient of ordinary fields is set to 1.0. This indicates the number of data records involved in this transaction, i.e. the number of rows or entries contained in the dataset to be traded. The number of records directly reflects the scale of the data. The more records there are, the more individuals are involved, and the wider the scope of the impact once a leakage event occurs. Therefore, the risk of exposure increases proportionally. This indicates the number of participants in the transaction, including data providers, data requesters, and intermediate nodes that may pass through during data transmission. The number of participants reflects the number of potential data exposure points. Each additional participant adds a potential data leakage point, so the more participants there are, the higher the exposure risk. This represents the normalization factor, used to standardize risk values ​​to the range of 0 to 1 for easier risk comparison and comprehensive scoring calculation. This factor is calculated as the sum of the maximum theoretical risk values ​​of all fields. ; in, The maximum value of all field weight coefficients is 1.5. This represents the theoretical maximum number of records. This represents the theoretical maximum number of participants.

[0060] Step 501: Construct the association relationship between different fields in the sensitivity feature set, and calculate the field association risk value based on the association relationship.

[0061] Specifically, the system constructs a sensitivity correlation matrix. This is used to characterize the correlation between different fields, reflecting the non-linear growth trend of overall risk when multiple highly sensitive fields cluster together in the same dataset. Sensitivity Correlation Matrix for Matrix, where Represents the total number of fields in the dataset, matrix elements Representation field and fields The system calculates the correlation strength coefficient between fields. For classic strongly correlated field combinations, the system pre-sets correlation coefficients: 0.9 for name and ID number, 0.8 for name and mobile phone number, 0.85 for ID number and address, and 0.75 for mobile phone number and address. For field combinations not covered in the pre-set database, the system automatically infers the correlation coefficient using data analysis methods based on mutual information theory and calculates the conditional entropy of the two field values. ; in, and Representing fields respectively and fields The set of possible values, Representation field Values and fields Values The joint probability, and These represent marginal probabilities. The larger the mutual information value, the stronger the information dependence between the two fields, and the higher the correlation coefficient.

[0062] Correlation coefficient Mutual Information The mapping relationship between them is normalized: ; in, and Representing fields respectively and fields The information entropy always remains within the standard range of 0 to 1, and at the same time, it can accurately reflect the relative correlation strength between fields.

[0063] Field associated risk value The calculation formula is: ; This formula represents the calculation of the aggregation risk contribution for all possible combinations of field pairs.

[0064] Step 502: Calculate the transaction environment risk value based on the environmental parameters of the data transaction.

[0065] Specifically, environmental parameters reflect the security status of the transmission environment. Transaction environment risk value. The calculation formula is: ; in, This parameter indicates the reliability of the transmission path, ranging from 0 to 1. A higher value indicates a more reliable transmission path. This parameter comprehensively reflects the security and reliability of the network path and intermediate nodes through which data is transmitted. This parameter represents the strength of the transmission protocol, ranging from 0 to 1. A higher value indicates a higher level of encryption and security. This parameter reflects the technical protection capability of the encryption protocol used during data transmission. When using high-strength encryption protocols such as TLS 1.3, this parameter is close to 1. When using older versions of TLS or SSL protocols, the parameter value is moderately reduced. When transmitting plaintext without encryption, this parameter is 0. The assessment of protocol strength is based on cryptographic standards and known security vulnerabilities to ensure that the assessment results reflect the actual security level of the protocol. The network environment coefficient ranges from 0.5 to 2. This coefficient reflects the overall security status of the current Internet environment. When the network threat level is normal, the coefficient is set to 1.0 as the baseline value. When there is a large-scale network attack or a significant increase in threat activities, the coefficient is increased to 1.5 to 2.0 to enhance the vigilance of risk assessment. When the network security situation is good and the threat level is low, the coefficient can be appropriately reduced to 0.8 to 0.9.

[0066] The beneficial effects of this embodiment are as follows: By constructing a multi-dimensional risk quantification model, a comprehensive and systematic assessment of data transaction security risks is achieved. The system establishes risk calculations from three dimensions: data content, field association, and transaction environment. This not only accurately identifies security risks of different sources and natures but also transforms complex risk factors into comparable quantitative indicators through a scientific mathematical model. This provides accurate and reliable data support for comprehensive scoring and risk level determination, significantly improving the scientific rigor and accuracy of data transaction security assessment.

[0067] Figure 6 This is a flowchart illustrating an embodiment of the present application, showing how a comprehensive score is obtained by weighting data content risk values, field association risk values, and transaction environment risk values ​​according to transaction scenario information, based on the weighted coefficients. Figure 6 As shown in the embodiments of this application, the data transaction security assessment method determines weight coefficients based on transaction scenario information, and calculates a comprehensive score by weighting the data content risk value, field association risk value, and transaction environment risk value according to the weight coefficients. The method includes the following steps: Step 600: Based on the transaction scenario type and industry category in the transaction scenario information, determine the first weight coefficient corresponding to the data content risk value, the second weight coefficient corresponding to the field-related risk value, and the third weight coefficient corresponding to the transaction environment risk value.

[0068] Specifically, the system extracts transaction scenario type and industry category information from transaction scenario information. Based on the differences in risk characteristics across different scenarios and industries, it dynamically determines the weight coefficients of the three risk indicators. For data transaction scenarios in the financial industry, the system sets the second weight coefficient to 0.5 to emphasize the importance of field association risk, the first weight coefficient to 0.3, and the third weight coefficient to 0.2. For cross-border data transaction scenarios, the system increases the third weight coefficient to 0.45 to strengthen the focus on transaction environment risks, sets the first weight coefficient to 0.35, and the second weight coefficient to 0.2. For internal enterprise data flow scenarios, the first weight coefficient is set to 0.45, and the second weight coefficient is set to 0.4.

[0069] Step 601: Multiply the data content risk value by the first weight coefficient to obtain the first weighted value, multiply the field association risk value by the second weight coefficient to obtain the second weighted value, and multiply the transaction environment risk value by the third weight coefficient to obtain the third weighted value.

[0070] Specifically, the system performs weighted calculations on the three risk indicators to obtain their respective weighted contribution values. The first weighted value is obtained by multiplying the data content risk value calculated in step 500 by the first weight coefficient determined in step 600. This weighted value reflects the actual contribution of data exposure risk to the overall security score in the current transaction scenario. The larger the first weight coefficient, the more the scenario values ​​the risk of unauthorized access to the data itself, and the higher the proportion of the first weighted value in the overall score. The second weighted value is obtained by multiplying the field association risk value calculated in step 501 by the second weight coefficient determined in step 600. This weighted value reflects the contribution of nonlinear risks generated by the aggregation of multiple sensitive fields in the current scenario. For application scenarios that highly rely on identity recognition and accurate profiling, field association risk is often the core source of security threats. Therefore, the second weight coefficient is larger, and the second weighted value accounts for a higher proportion. The third weighted value is obtained by multiplying the transaction environment risk value calculated in step 502 with the third weight coefficient determined in step 600. This weighted value reflects the impact of factors such as network environment, transmission protocol and reputation of access subject on overall security during transmission. For scenarios involving public network transmission or cross-border data flow, the uncertainty of the transmission environment is a key risk point, and the third weight coefficient is correspondingly larger. The third weighted value occupies an important position in the comprehensive score.

[0071] Step 602: Sum the first weighted value, the second weighted value, and the third weighted value to obtain the comprehensive score.

[0072] Specifically, the system sums the first weighted value, the second weighted value, and the third weighted value calculated in step 601 to obtain a comprehensive security score. The system then calculates the product of the data content risk value and the first weighted coefficient, the product of the field association risk value and the second weighted coefficient, and the product of the transaction environment risk value and the third weighted coefficient. The comprehensive score is the sum of these three products, and its value ranges from 0 to 1. A value closer to 1 indicates a higher overall security risk, while a value closer to 0 indicates a lower risk.

[0073] The beneficial effects of this embodiment are as follows: By dynamically adjusting the weights of the three types of risks based on the transaction scenario and weighting and integrating the risk values, and combining them with the platform security coefficient for correction, an adaptive and refined quantitative assessment of data transaction security risks is achieved, thereby significantly improving the accuracy and scenario adaptability of the comprehensive score.

[0074] Figure 7 This is a schematic diagram of the structure of a data transaction security assessment system provided in one embodiment of this application. Figure 7 As shown in the embodiment of this application, the data transaction security assessment system includes a field extraction module 700, which serves as the system's data input. This module receives the dataset to be traded from the data transaction platform, extracts its field information, and outputs the extraction results to the sensitivity analysis module 701. The sensitivity analysis module 701 receives the field information output by the field extraction module 700 and performs multi-dimensional feature analysis, including format pattern recognition, statistical feature calculation, and semantic analysis. Through multi-level matching, it identifies the sensitivity category of each field, determines the basic weights, and obtains the sensitivity weights of each field after correction. Finally, it generates a sensitivity vector and outputs this vector simultaneously to the dynamic grading module 702 and the feature integration module 703.

[0075] The dynamic grading module 702 receives the sensitivity vector output by the sensitivity analysis module 701. Based on the correspondence between the sensitivity weights of each field and preset weight ranges, it divides the fields into three basic sensitivity levels: Level 1 (high sensitivity), Level 2 (medium sensitivity), and Level 3 (low sensitivity). Simultaneously, this module also receives transaction-related information from external sources, including data usage tags, field access frequency statistics, and historical risk event records. Based on this information, it dynamically corrects the basic sensitivity levels, obtaining a dynamic grading result which is then output to the feature integration module 703. The feature integration module 703 simultaneously receives the sensitivity vector output by the sensitivity analysis module 701 and the dynamic grading result output by the dynamic grading module 702. It correlates and maps these two data points and combines them with auxiliary feature information to generate a sensitivity feature set containing the sensitivity vector, the dynamic grading result, and an auxiliary feature set, according to a triplet structure. This feature set serves as a standardized input and output to the risk calculation module 704.

[0076] The risk calculation module 704 receives the sensitivity feature set output by the feature integration module 703, and simultaneously obtains the scale information of the dataset to be traded, including the number of data records and the number of trading participants, as well as the environmental parameters of the data transaction, including the transmission path security level, transmission protocol strength, and access subject reputation. Based on the sensitivity feature set and the above parameters, it calculates the data content risk value, field association risk value, and transaction environment risk value, and outputs the three risk values ​​to the scoring calculation module 705. The scoring calculation module 705 receives the three risk values ​​output by the risk calculation module 704, and simultaneously obtains the transaction scenario type and industry category from the transaction scenario information. Based on the scenario characteristics, it determines the first weight coefficient, second weight coefficient, and third weight coefficient corresponding to the three risk values, respectively. It then calculates a comprehensive score by weighting and summing the three risk values ​​according to the weight coefficients, and outputs the comprehensive score to the risk judgment module 706.

[0077] The risk assessment module 706 receives the comprehensive score output by the scoring calculation module 705, compares it with preset high-risk and low-risk thresholds, and determines the risk level as high, medium, or low based on the threshold range matching result. The preset thresholds are dynamically determined and periodically updated based on the statistical analysis results of the security event occurrence rate in historical transaction data. The risk assessment module 706 outputs the risk level determination result, completing the basic security assessment process.

[0078] The beneficial effects of this embodiment are: through modular division of labor and multi-source information fusion, the entire process from field sensitivity identification to dynamic risk calculation and then to comprehensive scoring and risk level determination is automated and adaptive, thereby significantly improving the refinement, accuracy and scenario adaptability of data transaction security assessment.

[0079] Figure 8 This is a schematic diagram of the structure of a data transaction security assessment system provided in another embodiment of this application. Figure 8 As shown in the embodiment of this application, the data transaction security assessment system provides a protection execution module 800 that receives the risk level judgment result output by the risk judgment module 706. Based on different risk levels, it triggers a combination of protection strategies of corresponding strength. For high-risk transactions, it implements multiple protection measures such as mandatory desensitization, mandatory encrypted transmission, access rate limiting, and manual review. For medium-risk transactions, it implements medium-strength protection such as partial desensitization and moderate rate limiting. For low-risk transactions, it allows normal passage but retains audit logs. The execution result of the protection strategy is output to the result acquisition module 801.

[0080] The result acquisition module 801 simultaneously receives the protection strategy execution status output by the protection execution module 800 and the comprehensive score output by the scoring calculation module 705. It establishes a transaction lifecycle tracking system to continuously record the entire process data of the transaction, and collects the final execution result of the transaction, including the status of normal completion, abnormal termination, or occurrence of a security event. It compares and analyzes the comprehensive score with the actual security outcome to identify overestimation and underestimation of misjudgment and underestimation and missed judgment. After statistically analyzing the misjudgment rate and missed judgment rate indicators, it outputs the comparative analysis results to the parameter adjustment module 802.

[0081] The parameter adjustment module 802 receives the comparative analysis results output by the result acquisition module 801 and automatically optimizes the model parameters using a supervised learning framework. This module outputs adjustment instructions to multiple preceding modules via feedback connections. Specifically, it feeds back the updated content of the sensitivity dictionary and the adjustment scheme for field weights to the sensitivity analysis module 701, the corrected values ​​of the correlation coefficients in the correlation matrix to the risk calculation module 704, the optimized configuration of weight coefficients under different scenarios to the scoring calculation module 705, and the dynamic adjustment results of preset thresholds to the risk judgment module 706.

[0082] The beneficial effects of this embodiment are as follows: by introducing protection execution, result collection and parameter adaptive adjustment, a security control of "assessment-protection-feedback-optimization" is realized, which enables the system to continuously correct model parameters according to actual transaction performance, thereby significantly improving the accuracy of assessment, the effectiveness of protection and the overall self-evolution capability.

[0083] It will be apparent to those skilled in the art that this application is not limited to the details of the exemplary embodiments described above, and that this application can be implemented in other specific forms without departing from the spirit or essential characteristics of this application. Therefore, the embodiments should be considered illustrative and non-limiting in all respects, and the scope of this application is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within this application. No reference numerals in the claims should be construed as limiting the scope of the claims.

Claims

1. A data transaction security assessment method, characterized in that, include: Extract field information from the dataset to be traded; Feature analysis is performed on the field information to determine the sensitivity weight of each field and generate a sensitivity vector; Based on the sensitivity vector, each field is divided into sensitivity levels, and the sensitivity levels are dynamically adjusted based on transaction-related information to obtain dynamic classification results; Integrate the sensitivity vector and the dynamic grading result to generate a sensitivity feature set; Based on the aforementioned sensitivity feature set, calculate the data content risk value, field association risk value, and transaction environment risk value; The weighting coefficients are determined based on the transaction scenario information. The risk values ​​of the data content, the risk values ​​of the field associations, and the risk values ​​of the transaction environment are then weighted according to the weighting coefficients to obtain a comprehensive score. The risk level is determined based on the comprehensive score and a preset threshold, wherein the preset threshold is dynamically determined based on historical data.

2. The method according to claim 1, characterized in that, Also includes: Trigger the corresponding protection strategy based on the risk level; Collect transaction execution results and compare and analyze the comprehensive score with the transaction execution results; Based on the comparative analysis results, the sensitivity weight, the correlation relationship, the weight coefficient, and the preset threshold are adjusted.

3. The method according to claim 1, characterized in that, The step of performing feature analysis on the field information, determining the sensitivity weight of each field, and generating a sensitivity vector includes: Perform multi-dimensional feature analysis on the field information to identify the sensitive categories of each field; Based on the aforementioned sensitivity categories, determine the basic weights of each field; The basic weights are modified to obtain the sensitivity weights of each field; Based on the sensitivity weights of each field, the sensitivity vector is generated; The multi-dimensional feature analysis includes format pattern, statistical features and semantic analysis. The sensitive category represents the sensitivity of the field content. The basic weight is a preset initial weight. The sensitivity weight is a quantified field sensitivity value. The sensitivity vector is a set of sensitivity weights for each field.

4. The method according to any one of claims 1-3, characterized in that, The transaction-related information includes data usage tags, field access frequency, and historical risk event records. The process involves classifying each field into sensitivity levels based on the sensitivity vector, and dynamically adjusting these sensitivity levels based on the transaction-related information to obtain a dynamic grading result, including: Based on the correspondence between the sensitivity weights of each field in the sensitivity vector and the preset weight range, each field is divided into sensitivity levels; The dynamic classification result is obtained by dynamically adjusting the data based on at least one of the following methods: the usage risk coefficient corresponding to the data usage label, the comparison result of the field access frequency and the preset frequency threshold, and the risk level of the historical risk event records.

5. The method according to any one of claims 1-3, characterized in that, The process of integrating the sensitivity vector and the dynamic grading result to generate a sensitivity feature set includes: The sensitivity feature set is generated by associating and mapping the weights of each field in the sensitivity vector with the corresponding dynamic grading results and combining them with auxiliary features. The auxiliary features include at least one of the following: data usage tags, access frequency statistics, and historical risk event records.

6. The method according to claim 1, characterized in that, The calculation of data content risk value, field association risk value, and transaction environment risk value based on the sensitivity feature set includes: Based on the sensitivity feature set and the size information of the dataset to be traded, the risk value of the data content is calculated; Construct the association relationships between different fields in the sensitivity feature set, and calculate the association risk value of the field based on the association relationships; The risk value of the transaction environment is calculated based on the environmental parameters of the data transaction.

7. The method according to claim 6, characterized in that, The risk value of the data content is calculated based on the size information of the sensitivity feature set and the dataset to be traded; the correlation between different fields in the sensitivity feature set is constructed, and the correlation risk value of the field is calculated based on the correlation. Based on the environmental parameters of the data transaction, the risk value of the transaction environment is calculated, including: The risk value of the data content is calculated based on the sensitivity weight of each field in the sensitivity vector, the number of fields in the dataset to be traded, the number of data records, and the number of trading participants. Calculate the association risk value of the field based on the association coefficient in the association relationship and the sensitivity weight of each field in the sensitivity vector; The transaction environment risk value is calculated based on the transmission path security level, transmission protocol strength, and access subject reputation in the environmental parameters.

8. The method according to any one of claims 1-3, characterized in that, The step involves determining weighting coefficients based on transaction scenario information, and then weighting the data content risk value, the field association risk value, and the transaction environment risk value according to the weighting coefficients to obtain a comprehensive score, including: Based on the transaction scenario type and industry category in the transaction scenario information, determine the first weight coefficient corresponding to the risk value of the data content, the second weight coefficient corresponding to the risk value associated with the field, and the third weight coefficient corresponding to the risk value of the transaction environment; The first weighted value is obtained by multiplying the data content risk value by the first weight coefficient, the second weighted value is obtained by multiplying the field association risk value by the second weight coefficient, and the third weighted value is obtained by multiplying the transaction environment risk value by the third weight coefficient. The first weighted value, the second weighted value, and the third weighted value are summed to obtain the comprehensive score.

9. A data transaction security assessment system, characterized in that, include: The field extraction module is used to extract field information from the dataset to be traded; The sensitivity analysis module is used to perform feature analysis on the field information, determine the sensitivity weight of each field, and generate a sensitivity vector; The dynamic grading module is used to divide each field into sensitivity levels according to the sensitivity vector, and dynamically adjust the sensitivity levels based on transaction-related information to obtain dynamic grading results; The feature integration module is used to integrate the sensitivity vector and the dynamic grading result to generate a sensitivity feature set; The risk calculation module is used to calculate the data content risk value, field association risk value, and transaction environment risk value based on the sensitivity feature set. The scoring calculation module is used to determine the weighting coefficients based on the transaction scenario information, and to calculate the comprehensive score by weighting the risk value of the data content, the risk value of the field association, and the risk value of the transaction environment according to the weighting coefficients. The risk assessment module is used to determine the risk level based on the comprehensive score and a preset threshold, wherein the preset threshold is dynamically determined based on historical data.

10. The system according to claim 9, characterized in that, Also includes: The protection execution module is used to trigger the corresponding protection strategy according to the risk level; The result acquisition module is used to collect transaction execution results and compare and analyze the comprehensive score with the transaction execution results. The parameter adjustment module is used to adjust the sensitivity weight, the correlation relationship, the weight coefficient, and the preset threshold based on the comparative analysis results.

Citation Information

Patent Citations

  • Data security assessment method and system based on data elements

    CN117993024A

  • Scientific and technological financial platform data sharing method

    CN119720281A

  • Dynamic risk control method and device, equipment and storage medium

    CN120163653A

  • Dynamic sensitive data outbound risk assessment method and system based on multi-source risk information

    CN120470590A

  • Multi-field data desensitization method and device, equipment and medium

    CN120850352A

Cited By

  • Equipment safety monitoring and evaluating system and method based on multi-mode cooperation

    CN121765493A

  • Device safety monitoring and evaluation system and method based on multi-modal collaboration

    CN121765493B

  • Data transaction security assessment method based on gated attention feedforward network

    CN122066518A