Self-adaptive data dynamic grading method and device based on intelligent algorithm model

By constructing an adaptive dynamic data classification model and utilizing technologies such as intelligent algorithms and generative adversarial networks, the problems of accuracy and dynamism in multimodal data classification were solved, enabling accurate classification and secure management of data across multiple industries.

CN121744082APending Publication Date: 2026-03-27FUJIAN ZHONGXIN NET SAFETY INFORMATION TECHNOLOGY CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-17
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing data classification methods are ill-suited to the characteristics of multimodal data, cannot accurately quantify sensitivity, and lack dynamism and anti-interference capabilities. They are unable to cope with changes throughout the data lifecycle and new types of attacks, resulting in classification results that are out of touch with industry needs.

Method used

By acquiring multimodal feature data throughout the entire data lifecycle, and combining intelligent algorithm models and noise-robust machine learning algorithms, an adaptive dynamic data classification model is constructed. Adversarial generative networks and causal inference networks are used to adjust the classification strategy in real time, and dynamic classification is achieved by combining blockchain traceability information.

Benefits of technology

It enables precise classification and security management of multimodal and multi-industry data, improves the dynamics and anti-interference capabilities of classification, adapts to changes in the entire data lifecycle and new types of attacks, and ensures that classification results match industry needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121744082A_ABST
    Figure CN121744082A_ABST
Patent Text Reader

Abstract

The invention provides a self-adaptive data dynamic grading method and device based on an intelligent algorithm model, and is applied to the field of data processing. According to the method, data full life cycle management is taken as a core, firstly, multi-modal feature data, algorithm parameters and industry security tags are processed through supervised learning cross-source verification and a noise robust algorithm, and a standardized feature sequence is output; constructing a comprehensive safety assessment framework through an adversarial generative network, extracting vulnerability weight factors to construct a dynamic matrix, and cooperatively calculating to obtain key grading indexes; then, combining scene extraction feature values to divide sensitive levels, and generating a multi-dimensional hierarchical feature matrix; based on the causal reasoning network, abnormal synergy correction factors are identified; and finally, constructing a grading model according to field grouped data, meta-learning screening key factors, fusion constraint and block chain information, and realizing accurate grading and safety management and control of the data by combining a time sequence output dynamic grading result and a full-period management and control strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing, and in particular to an adaptive dynamic data classification method and apparatus based on an intelligent algorithm model. Background Technology

[0002] Currently, basic data classification and grading methods exist in the industry, mostly relying on static rules (such as determining data level based on a pre-set list of sensitive fields) or a single algorithm (such as a simple machine learning model for sensitive data identification) to achieve preliminary data grading. For example, some solutions use manually defined sensitive keywords (such as "ID number" or "confidential") to match keywords in text data to classify sensitivity levels; for numerical data, fixed thresholds (such as classifying transactions exceeding a certain value as highly sensitive) are used for grading. These methods can initially meet the needs in scenarios with simple data types and application scenarios, but they are difficult to adapt to multimodal data and complex industry scenarios.

[0003] The accuracy of data classification is insufficient and its adaptability is poor. Existing methods are unable to cope with the differences in the characteristics of multimodal data. For example, the recognition accuracy of privacy information (facial features, lesion information) in image data (such as medical CT images) is low, and it is impossible to accurately quantify its sensitivity. At the same time, the business logic and compliance requirements of different industries are not fully considered. For example, the sensitivity judgment standards of "equipment control instructions" in the industrial field and "medical data" in the medical field are completely different, but existing solutions mostly adopt a unified classification logic, which leads to the classification results being out of touch with the actual needs of the industry, resulting in the problem of "high-sensitivity data not being strictly managed" or "low-sensitivity data being over-managed".

[0004] The classification process lacks dynamism and has weak anti-interference capabilities. Traditional classification methods mostly adopt static strategies, which cannot adapt to the dynamic changes throughout the entire data lifecycle (such as the addition or modification of sensitive attributes during data flow, and the iterative upgrade of external attack patterns). For example, when new adversarial attacks (such as AI-driven stealth text tampering attacks) emerge, existing classification models, lacking real-time vulnerability detection and strategy adjustment capabilities, struggle to identify shifts in the sensitivity level of tampered data. Simultaneously, random noise during data transmission (such as field loss due to network fluctuations) and redundant information in the storage process can easily lead to misjudgments in the classification algorithm, further reducing classification accuracy.

[0005] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore includes information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0006] According to one aspect of this application, an adaptive dynamic data classification method based on an intelligent algorithm model is provided, comprising: acquiring multimodal feature data of the entire data lifecycle, core parameters of the intelligent algorithm model, and data security management requirement labels for multiple industries; removing outliers and redundant information in data collection and storage by training verification rules through labeled abnormal samples; filtering random interference signals in the data transmission process by combining noise robustness machine learning algorithms; and outputting a standardized full-cycle data feature sequence; converting the standardized full-cycle data feature sequence into a security assessment map; subdividing incomplete data samples based on adversarial generative networks; and constructing a comprehensive security assessment framework; and extracting vulnerability weight factors of different data types based on the comprehensive security assessment framework to construct a dynamic security assessment matrix for intelligent data classification. The algorithm model's core parameters are collaboratively calculated with data security management requirements to obtain a combination of key grading indicators under the target sensitivity level. Based on data usage scenarios, sensitive feature values, grading efficiency benchmark values, and security and grading matching deviation values ​​are extracted to classify data sensitivity levels and generate a multi-dimensional data grading feature matrix. Real-time grading data is processed based on the multi-dimensional data grading feature matrix to identify abnormal signals such as grading standard deviation and imbalance between sensitive features and grading results. A causal inference network is used to generate dynamic correction factors for the grading strategy. Data is grouped by domain, meta-learning is used to filter key factors, and constraints and blockchain traceability limit information are integrated to construct an adaptive dynamic data grading model. Combined with time-series patterns, dynamic grading results and full-cycle management strategies are output.

[0007] Another aspect of this application discloses an adaptive data dynamic grading device based on an intelligent algorithm model, comprising: an acquisition module for acquiring multimodal feature data throughout the entire data lifecycle, core parameters of the intelligent algorithm model, and data security management requirement labels from multiple industries; removing outliers and redundant information from data collection by training verification rules using labeled abnormal samples; filtering random interference signals during data transmission using a noise robustness machine learning algorithm; and outputting a standardized full-cycle data feature sequence; and a processing module for converting the standardized full-cycle data feature sequence into a security assessment map; subdividing incomplete data samples based on an adversarial generative network; and constructing a comprehensive security assessment framework; and extracting vulnerability weight factors for different data types based on the comprehensive security assessment framework to construct a dynamic security assessment. The matrix performs collaborative calculations on the core parameters of the intelligent algorithm model and the data security management requirements to obtain a combination of key grading indicators under the target sensitivity level. Based on the data usage scenario, it extracts data sensitivity feature values, grading efficiency benchmark values, and security and grading matching deviation values ​​to classify data sensitivity levels and generate a multi-dimensional data grading feature matrix. Based on the multi-dimensional data grading feature matrix, it processes real-time grading data to identify abnormal signals such as grading standard deviation and imbalance between sensitive features and grading results, and uses a causal inference network to generate dynamic correction factors for grading strategies. It groups data by domain, uses meta-learning to filter key factors, and integrates constraint conditions and blockchain traceability limit information to construct an adaptive data dynamic grading model. It outputs dynamic grading results and full-cycle management strategies based on time-series patterns.

[0008] According to another aspect of this application, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a second processor, implements the above-described adaptive dynamic data grading method based on an intelligent algorithm model.

[0009] This application presents an adaptive dynamic data grading method and apparatus based on an intelligent algorithm model. Addressing the need for comprehensive data security management across multiple modalities and industries throughout the entire data lifecycle, it proposes an adaptive dynamic data grading scheme. First, it acquires multimodal features, core algorithm parameters, and industry security labels throughout the data lifecycle. Then, it uses supervised learning for cross-source verification and noise robustness algorithms to output standardized feature sequences. Next, it constructs a comprehensive security assessment framework using an adversarial generative network, extracts vulnerability weight factors to build a dynamic matrix, and collaboratively calculates key grading indicators. Then, it extracts feature values ​​based on the scenario to classify sensitivity levels, generating a multi-dimensional grading feature matrix. Subsequently, it identifies grading anomalies based on the matrix and uses a causal inference network to generate correction factors. Finally, it groups data by domain, uses meta-learning to filter key factors, and integrates encryption constraints and blockchain traceability information to construct a grading model. Combining temporal patterns, it outputs dynamic grading results and full-cycle management strategies. This approach solves the problems of low accuracy, poor dynamism, and lack of end-to-end protection in traditional grading methods, achieving accurate grading and security management of data across multiple industries.

[0010] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0011] Figure 1 The flowchart illustrates an adaptive dynamic data classification method based on an intelligent algorithm model provided in an embodiment of this application. Figure 2 The diagram shows a schematic representation of an adaptive data dynamic classification device based on an intelligent algorithm model, according to an embodiment of this application. Detailed Implementation

[0012] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0013] The following is combined with Figure 1 This application describes an adaptive dynamic data classification method based on an intelligent algorithm model, according to an exemplary embodiment of the present application.

[0014] S101 acquires multimodal feature data throughout the entire data lifecycle, core parameters of intelligent algorithm models, and data security management requirement labels for multiple industries. It eliminates abnormal values ​​and redundant information in data collection and storage by training and verifying rules based on labeled abnormal samples. Combined with noise-robust machine learning algorithms, it filters random interference signals during data transmission and outputs a standardized full-cycle data feature sequence.

[0015] In one implementation, multimodal feature data covering the entire lifecycle of data encompasses the characteristics of each stage from data generation to data loss, and the collection needs to cover multiple data types and key nodes throughout the entire process. For example, in the medical field, the data collection stage requires acquiring patient diagnostic images (such as CT images and ultrasound images), electronic medical record texts (including medical history descriptions and medication records), and laboratory indicator values ​​(such as blood routine tests and hormone levels); the storage stage requires recording the data storage location (such as hospital local servers and cloud storage nodes) and the storage encryption status (such as whether AES-256 encryption is used); and the usage stage requires collecting data access logs (such as doctor viewing records and research access records) and data flow trajectories (such as link information from the laboratory to the department workstation), forming a multimodal, full-process feature dataset.

[0016] The core parameters of intelligent algorithm models focus on key algorithm parameters that support data classification, and are precisely collected according to the type of intelligent model used. For example, when using Generative Adversarial Networks (GANs) to handle incomplete data, it is necessary to obtain the number of network layers in the generator and discriminator (e.g., 12 layers for the generator and 8 layers for the discriminator), the type of activation function (e.g., LeakyReLU), the learning rate (e.g., an initial learning rate of 0.001), and the number of iterations (e.g., 5000 rounds). When using federated learning for privacy protection, it is necessary to collect the number of federated nodes (e.g., 10 hospital nodes), the gradient update frequency (e.g., once per hour), and the aggregation algorithm parameters (e.g., the weight allocation coefficients of the FedAvg algorithm) to ensure that the algorithm parameters are traceable and adjustable.

[0017] A structured labeling system should be developed, taking into account the compliance requirements and business scenarios of different industries. For example, in the medical field, "patient privacy protection level" (e.g., data containing ID numbers and medical record numbers are labeled "high privacy level," while statistical data without personal identification is labeled "low privacy level") and "data compliance standards" (e.g., compliance with the Personal Information Protection Law and the Medical Data Security Guidelines) are required; in the industrial field, "data tampering risk level" is required (e.g., equipment control command data is labeled "high risk," while production environment temperature and humidity data is labeled "low risk"), providing an industry-specific basis for subsequent classification.

[0018] A cross-source data joint verification model based on supervised learning is adopted. By training verification rules through labeled abnormal samples, outliers and redundant information in the data acquisition process are accurately identified and eliminated. For labeling abnormal samples and training verification rules, an abnormal sample library is first constructed, labeling common anomaly types in the data acquisition stage. Then, the verification model is trained based on supervised learning. For example, in medical data acquisition, abnormal samples such as "CT image pixel value anomalies caused by sensor failure" (e.g., pixel values ​​in a local area suddenly change to 0 or 255), "electronic medical record age anomalies caused by manual entry errors" (e.g., age labeled as 200 years old), and "redundant test indicators caused by repeated collection" (e.g., blood routine data recorded repeatedly at the same time point) are labeled. The verification model is trained using a random forest algorithm, enabling the model to learn verification rules such as "image pixel value range (normal CT image pixel values ​​are usually between -1000 and 400 HU)," "reasonable age range (0 to 120 years old)," and "uniqueness of data acquisition time."

[0019] The verification and rejection process involves inputting the collected data into the trained verification model. The model outputs anomaly detection results based on preset rules, rejecting outliers and redundant information. For example, when inputting a batch of hospital outpatient data, the model detects three abnormal data points: one CT image with abnormal pixel values ​​(local pixel value -2000 HU) due to sensor malfunction, one electronic medical record with an age label of 150 years, and two duplicate blood glucose test data points (collected one minute apart with identical values). The model automatically rejects these outliers and redundant data, retaining only valid data that meets the verification rules, ensuring the accuracy of the data collection process.

[0020] By combining noise robustness machine learning algorithms, random interference signals generated during data transmission are filtered to ensure the integrity and reliability of the data after transmission. The specific process and examples are as follows: Identify the types of transmission interference by analyzing common interference signals in the data transmission process and clarifying the noise characteristics. For example, when industrial data is transmitted from workshop equipment to the cloud, unstable transmission links may cause "time-series data misalignment" (such as mismatch between equipment temperature data and timestamps).

[0021] Select an appropriate noise robustness algorithm and perform filtering. Choose the corresponding noise robustness algorithm based on the type of interference to process the transmitted data. For example, to address time-series misalignment interference in industrial data, a robust Long Short-Term Memory (LSTM) network algorithm is used. By learning the patterns of normal time-series data (such as equipment temperature increasing by 1°C every 5 minutes), the misaligned timestamps are corrected, allowing the temperature data to re-match with the timestamps. To address numerical jump interference, a robust regression algorithm (such as Huber regression) is used to identify and correct abnormal values ​​that exceed reasonable ranges (such as correcting a tax amount of 10 million yuan to 1 million yuan), ultimately outputting clean data free from transmission interference.

[0022] The cleaned and denoised full-lifecycle data is standardized to unify the data format, feature dimensions, and calculation methods, outputting a standardized full-lifecycle data feature sequence. The specific process and examples are as follows. Unified rules covering data format, feature dimensions, and numerical range are established. For example, the data format is unified to JSON; the feature dimensions are unified to include five core dimensions: "data type (text / image / numerical)," "acquisition time," "storage encryption status," "usage scenario," and "sensitive attribute labels"; the numerical range is standardized to the [0,1] interval (e.g., mapping CT image pixel values ​​from -1000 to 400 HU to 0 to 1, and mapping age from 0 to 120 years to 0 to 1).

[0023] Data is transformed according to preset rules to generate standardized feature sequences. For example, the original information of a medical data item is "Data type: CT image, Acquisition time: 2024-05-20 14:30, Storage encryption status: Encrypted (AES-256), Usage scenario: Clinical diagnosis, Pixel value range: -1000~400HU, Patient age: 55 years old". After standardization, it is converted into a feature sequence in JSON format: {"Data type":"Image","Acquisition time":"2024-05-20 14:30","Storage encryption status":"1 (1 indicates encrypted, 0 indicates unencrypted)","Usage scenario":"Clinical diagnosis","Standardized pixel value range":"0~1","Standardized age":"0.458 (55 / 120≈0.458)"}. This sequence can be directly used for subsequent security assessment atlas construction and hierarchical model training.

[0024] S102 transforms the standardized full-cycle data feature sequence into a security assessment map, and constructs a comprehensive security assessment framework based on the subdivision of incomplete data samples using adversarial generative networks.

[0025] In one implementation, standardized full-lifecycle data feature sequences are processed and converted into multi-dimensional data security assessment maps through multi-dimensional feature encoding, generating basic security assessment visualization data. First, the encoding dimensions are clearly defined, covering the core security elements throughout the data lifecycle, including "data sensitivity dimension," "algorithm vulnerability dimension," "service trustworthiness dimension," and "transfer security dimension." Specifically, the data sensitivity dimension quantifies the degree to which the data contains sensitive information (e.g., the proportion of personal identification information and trade secrets); the algorithm vulnerability dimension assesses the attack resistance of the algorithms supporting data processing (e.g., tolerance to adversarial attacks, backdoor detection rate); the service trustworthiness dimension measures the service compliance of data usage (e.g., the strictness of access control, the integrity of operation logs); and the transfer security dimension assesses the security level of data transmission and storage (e.g., encryption algorithm strength, transmission link stability). Then, a feature mapping encoding method is used to map each numerical feature in the standardized feature sequence into coordinates or color values ​​that can be displayed in the map. Taking a standardized sequence of CT image data in the medical field as an example, its standardized features include: "Data type: image (encoded as 1), storage encryption status: encrypted (encoded as 1, corresponding to high security), sensitive attribute label: high privacy level (encoded as 0.9), algorithm iteration count: 5000 rounds (encoded as 0.85, corresponding to high algorithm stability), access log integrity: 98% (encoded as 0.98)". Through multi-dimensional feature encoding, these features are mapped to the four coordinate axes of the map: X-axis (data sensitivity, 0.9 corresponds to coordinate 0.9), Y-axis (algorithm vulnerability, 0.85 corresponds to coordinate 0.15, the lower the value, the lower the vulnerability), Z-axis (service trustworthiness, 0.98 corresponds to coordinate 0.98), and color dimension (transfer security, 1 corresponds to dark red, representing high security). Finally, a three-dimensional visualized multi-dimensional data security assessment map is generated. In the map, the CT image data is presented as nodes with "X=0.9, Y=0.15, Z=0.98, color=dark red", intuitively reflecting its security status.

[0026] The multi-dimensional data security assessment graph is imported into the generator module of the Generative Adversarial Network (GAN). Through a conditional GAN ​​architecture, scarce feature samples in incomplete data scenarios are synthesized and subdivided to generate an expanded, complete feature sample set. First, the input conditions and generation target of the conditional GAN ​​are determined. The input conditions are the "domain label," "data type label," and "sensitivity level label" from the multi-dimensional data security assessment graph. The generation target is the scarce feature samples missing under these conditions (such as anomalous attack samples in highly sensitive data or niche data type samples from specific industries).

[0027] An adversarial example evolution algorithm is employed to iteratively optimize the expanded and complete feature sample set. Combined with persistent attack response (PARP) technology, it simulates new attack scenarios, generating an enhanced sample set labeled with attack scenarios. The adversarial example evolution algorithm tracks changes in attack patterns in real time, automatically triggering sample iteration when attack features are updated. First, an adversarial example evolution algorithm framework is built, comprising an "attack pattern tracking module," a "sample mutation module," and an "iterative verification module." The attack pattern tracking module obtains features of new attack patterns (such as the recently emerging "AI-driven stealth text tampering attack," which achieves semantic tampering through minute character replacements and is difficult to detect by conventional methods) by crawling network security vulnerability platforms (such as the National Information Security Vulnerability Database) and industry security reports in real time. The sample mutation module optimizes existing samples based on the tracked new attack features (e.g., adding AI-generated stealth tampering characters to existing text tampering samples). The iterative verification module interacts with a discriminator to verify the effectiveness of the mutated samples (i.e., whether they can successfully simulate attack scenarios and conform to data feature distribution). If the verification is successful, the sample is retained; otherwise, it is mutated again. Taking a sample set of user transaction data in the financial sector as an example, the initially expanded sample set includes two types of conventional attack samples: "account information leakage samples" and "transaction amount tampering samples." The adversarial sample evolution algorithm, through the attack pattern tracking module, discovers a new type of "cross-platform transaction link hijacking attack" (characterized by tampering with the transaction redirect address, causing funds to be transferred to an illegal account), and then triggers sample iteration: the sample mutation module adds the attack feature of "minor modification of the transaction redirect address (changing the official address 'bank.xxx' to 'bank.xxx')" to the original transaction data samples, generating mutated samples; the iterative verification module verifies whether the mutated sample can successfully trigger link hijacking by simulating a transaction system environment. If successful, it adds the "attack scenario label: cross-platform transaction link hijacking" to the sample, ultimately generating an enhanced sample set containing 6 types of attack scenario labels (including the newly added "cross-platform transaction link hijacking"), increasing the attack scenario coverage of the samples from 40% to 85%.

[0028] Vulnerability characteristic parameters of different data types in open environments are collected to construct a data type-specific vulnerability constraint library, generating data type constraint data. The constraint library is customized with parameters for the data characteristics of multiple industries. First, the dimensions for vulnerability parameter collection are defined. For four core data types—text, image, audio, and numerical—four categories of parameters are collected: "vulnerability type in open environments," "attack success probability threshold," "security protection threshold," and "compliance requirement parameters." Vulnerability type refers to common attack methods for this data type in open environments; the attack success probability threshold is the maximum allowed probability of a successful attack on this data type; the security protection threshold is the minimum strength of protection measures required to ensure data security (such as encryption algorithm key length and vulnerability detection frequency); and compliance requirement parameters refer to security indicators that comply with industry regulations (such as medical data needing to meet the requirement of a privacy leakage rate of ≤0.001% under the Personal Information Protection Law). Next, parameters are collected according to industry customization, and a constraint library is constructed. Taking the education sector as an example, for "student registration text data" (text type), the collected vulnerability parameters include: vulnerable attack types are "text content tampering, personal information crawling"; the attack success probability threshold is "≤0.005%"; the security protection threshold is "using RSA-2048 encryption algorithm, daily vulnerability scan ≥3 times"; and the compliance requirement parameter is "complying with the 'Education Data Security Guidelines', with a desensitization rate of sensitive information such as student ID numbers ≥99.9%". These parameters are entered into the constraint library and labeled with "Domain: Education, Data Type: Text". For "Equipment Operation Numerical Data" (numerical type) in the industrial field, the collected parameters include: vulnerable attack type as "numerical jump tampering, time series data misalignment attack"; attack success probability threshold as "≤0.001%"; security protection threshold as "using SM4 encryption algorithm, transmission link bit error rate ≤10^-6"; compliance requirement parameter as "complies with the 'Industrial Data Security Management Measures', critical equipment numerical tampering alarm response time ≤1 second". These parameters are also labeled and entered into the constraint library, ultimately forming a vulnerability constraint library covering multiple industries such as medical, education, and industry, distinguishing different data types.

[0029] By integrating enhanced sample sets and data type constraint data, and combining noise robustness testing and privacy inversion testing, a multi-level stability reverse verification system is constructed to generate a comprehensive security assessment framework. First, the enhanced sample set and constraint data are fused and matched. Based on the "domain label" and "data type label" of the enhanced samples, corresponding constraint data is retrieved from the vulnerability constraint library, and constraint thresholds are added to each sample.

[0030] Based on the results of two tests, a multi-level stability inverse verification system is constructed, consisting of a basic stability layer, an advanced anti-interference layer, and an advanced privacy protection layer. The basic stability layer verifies the feature integrity of samples under normal conditions; the advanced anti-interference layer measures the sample's resistance to environmental interference based on noise robustness test results; and the advanced privacy protection layer assesses the sample's privacy protection level based on privacy inversion test results. Integrating all the above modules (multi-dimensional graph, adversarial generation samples, evolutionary optimization samples, vulnerability constraint library, and multi-level verification system), a comprehensive security assessment framework is formed, covering the data layer (sample generation and optimization), the algorithm layer (adversarial generation and evolutionary algorithms), the verification layer (multi-level inverse testing), and the standard layer (vulnerability constraint library). Taking medical applications as an example, this framework can receive standardized feature sequences of medical data, output a visualized security graph, a complete sample set, and a stability test report. It also provides security optimization suggestions based on the constraint library (e.g., "CT image data algorithms are highly vulnerable; it is recommended to upgrade the adversarial training model"), achieving a comprehensive assessment of data security throughout its entire lifecycle.

[0031] S103 extracts vulnerability weight factors of different data types based on the comprehensive security assessment framework to construct a dynamic security assessment matrix, and performs collaborative calculations on the core parameters of the intelligent algorithm model and the data security management and control requirements data to obtain the combination of key graded indicators under the target sensitivity level.

[0032] In one implementation, a feature weight extraction process is used to match multi-dimensional vulnerability features in the comprehensive security assessment framework with preset vulnerability assessment standards for different data types. This extracts vulnerability parameters for each data type at the data layer, algorithm layer, and service layer, generating a data type-vulnerability feature mapping table and initial vulnerability weight values. The multi-dimensional vulnerability features and assessment standards are clearly defined, and the vulnerability features must cover the three-layer architecture of "data layer-algorithm layer-service layer": data layer features include "data leakage risk coefficient" (measuring the probability of unauthorized data access) and "data integrity breach probability" (measuring the possibility of data tampering); algorithm layer features include "adversarial attack tolerance" (the algorithm's ability to resist adversarial sample attacks) and "backdoor detection coverage" (the proportion of backdoor vulnerabilities detected by the algorithm); service layer features include "access permission control strength" (the strictness of service control over data access permissions) and "operation log integrity" (the completeness of service recording data operation behavior). The pre-defined vulnerability assessment standards for different data types need to be customized according to industry and data form. For example, the assessment standards for text data in the medical field (such as electronic medical records) are "data leakage risk coefficient ≤ 0.001%, resistance to attacks ≥ 95%, access control strength ≥ 90%", while the assessment standards for numerical data in the industrial field (such as equipment operating parameters) are "data integrity breach probability ≤ 0.0005%, backdoor detection coverage ≥ 98%, operation log integrity ≥ 99%".

[0033] Next, feature matching and parameter extraction are performed. Taking CT image data (image type) in the medical field as an example, its multi-dimensional vulnerability features are retrieved from the comprehensive security assessment framework: In the data layer, "data leakage risk coefficient = 0.002% (due to the presence of patient privacy information), data integrity damage probability = 0.001%"; In the algorithm layer, "adversarial attack tolerance = 92% (the ability of the existing adversarial training model to resist image tampering attacks), backdoor detection coverage = 94%"; In the service layer, "access control strength = 88% (some unauthorized departments can view thumbnails), operation log integrity = 96%". These features are matched with preset medical image data evaluation standards (data leakage risk coefficient ≤ 0.0015%, data integrity failure probability ≤ 0.001%, adversarial attack tolerance ≥ 94%, backdoor detection coverage ≥ 95%, access control strength ≥ 90%, operation log integrity ≥ 95%), and non-compliant parameters are extracted (data leakage risk coefficient exceeds the standard by 0.0005%, adversarial attack tolerance is 2% lower, access control strength is 2% lower, and backdoor detection coverage is 1% lower). A data type and vulnerability feature mapping table is generated, which clearly defines the six vulnerability features and their specific values ​​corresponding to "data type: medical CT image (image)". Meanwhile, based on the importance of the assessment criteria, initial vulnerability weights are assigned to each feature. For example, the "data leakage risk coefficient" is directly related to patient privacy, so the initial weight is set to 0.3; the "adversarial attack tolerance" is related to algorithm security, so the initial weight is set to 0.25; the "access control strength" is related to service compliance, so the initial weight is set to 0.2; and the initial weights of the remaining features are set to 0.08, 0.07, and 0.1 respectively (the total weight is 1).

[0034] By aligning data security management standards across multiple industries with the core parameters of intelligent algorithm models, a vulnerability weight calibration model is constructed. This model calculates the impact weight of each vulnerability feature through a requirement-parameter matching logic, establishing a dynamic adjustment mechanism for vulnerability weight factors across different data types. The calibration input parameters must be determined, and the data security management standards for multiple industries need to be refined to the feature dimension.

[0035] Next, a calibration model and adjustment mechanism are constructed. Taking contract text data (text type) in the industrial field as an example, the vulnerability weight calibration model is constructed by connecting its security requirement standard "data leakage risk has the highest priority" with the core parameter of the intelligent algorithm model "federated learning gradient threshold = 0.001 (improving data transmission privacy protection)". The model adopts a linear weighted algorithm, and the inputs are "requirement standard weight adjustment coefficient" and "algorithm parameter influence coefficient": For the feature of "data leakage risk coefficient", the requirement standard requires a 20% increase in weight (adjustment coefficient = 1.2). In the algorithm parameters, federated learning improves data privacy protection, so the algorithm influence coefficient of this feature = 0.95 (no need to over-rely on weight). Then the calibrated weight = initial weight (0.3) × 1.2 × 0.95 = 0.342; For the feature of "backdoor detection coverage", the industrial field requirement standard has no special priority (adjustment coefficient = 1), but the accuracy of the backdoor detection model in the intelligent algorithm is improved by 10% (algorithm influence coefficient = 0.85). Then the calibrated weight = initial weight (0.25) × 1 × 0.85 = 0.2125. Meanwhile, a dynamic adjustment mechanism is established: when new compliance requirements are introduced in the industrial sector (such as "operation log integrity must meet 100% standard"), the model automatically sets the requirement adjustment coefficient of the "operation log integrity" feature to 1.3 and recalculates the weight; when the intelligent algorithm model iterates (such as the upgrade of the adversarial generative network), the model updates the algorithm influence coefficient in real time to achieve dynamic adaptation of the weight.

[0036] Using data type as the dividing dimension, a dynamic security assessment matrix is ​​constructed by integrating the data type-vulnerability feature mapping table, the initial vulnerability weight values, and the adjusted weights output by the vulnerability weight calibration model. The matrix's row dimensions represent data types, column dimensions represent vulnerability features, and matrix elements represent calibrated vulnerability weight factors. The matrix dimensions and integration rules are determined as follows: the matrix row dimension represents "data types," covering four core types—text, image, audio, and numerical data—as well as various industry-specific sub-types (such as medical text); the column dimension represents "vulnerability features," namely the six core features of the data layer, algorithm layer, and service layer; and the matrix elements represent "calibrated vulnerability weight factors," calculated as "calibrated weight = initial weight × demand adjustment coefficient × algorithm influence coefficient," with the sum of all feature weight factors corresponding to each data type being 1. Next, fusion and matrix generation are performed. Taking student registration text data (text type) in the education field as an example, the values ​​of its six vulnerability features are obtained from the data type and vulnerability feature mapping table: data leakage risk coefficient = 0.0012%, data integrity breach probability = 0.0008%, adversarial attack tolerance = 93%, backdoor detection coverage = 96%, access control strength = 89%, and operation log integrity = 97%. The initial values ​​of vulnerability weights are 0.3, 0.08, 0.25, 0.07, 0.2, and 0.1, respectively. After calculation by the vulnerability weight calibration model, the demand adjustment coefficient (the education field has high requirements for student privacy protection, so the "data leakage risk coefficient" adjustment coefficient = 1.2, and the rest are 1) and the algorithm impact coefficient (in the student registration data processing algorithm, "the backdoor detection model accuracy is improved by 5%, and the impact coefficient = 0.9; the other algorithm parameters have no significant changes, and the impact coefficient = 0.9) are calculated. 1) Obtain the calibrated weights for each feature: "Data leakage risk coefficient" = 0.3 × 1.2 × 1 = 0.36, "Data integrity breach probability" = 0.08 × 1 × 1 = 0.08, "Adversarial attack tolerance" = 0.25 × 1 × 1 = 0.25, "Backdoor detection coverage" = 0.07 × 1 × 0.9 = 0.063, "Access control strength" = 0.2 × 1 × 1 = 0.2, "Operation log integrity" = 0.1 × 1 × 1 = 0.1 (totaling 1). Fill these weight factors into a matrix, labeling the row dimension as "Data type: Educational student records (text)", and the column dimension as the six vulnerability features, with matrix elements corresponding to 0.36, 0.08, 0.25, 0.063, 0.2, and 0.1. Simultaneously, label the specific vulnerability values ​​for each feature to form a dynamic security assessment matrix, visually presenting the risk weight of this data type on each vulnerability feature.

[0037] This system employs a multi-objective collaborative computation mechanism to fuse dynamic security assessment matrix data, core parameters of intelligent algorithm models, and data security management requirements. It strengthens the weighting of key features by combining preset target sensitivity level assessment standards with the output features of the vulnerability weight calibration model, generating a combination of key grading indicators for the target sensitivity level. The multi-objective computation objectives and input data are clearly defined: the computation objectives are "maximizing grading accuracy (ensuring accurate sensitivity level classification), minimizing security risk (reducing vulnerability risk), and maximizing compliance rate (meeting industry standards)." Input data includes dynamic security assessment matrix data (such as vulnerability weight factors for various data types), core parameters of intelligent algorithm models (such as SM9 encryption algorithm key length and the number of federated learning nodes), and data security management requirements data (such as security indicators corresponding to various industry sensitivity levels; L3 high-sensitivity data requires "data leakage risk coefficient ≤ 0.0008%, encryption algorithm strength ≥ 256 bits," and L2 medium-sensitivity data requires "data leakage risk coefficient ≤ 0.002%, encryption algorithm strength ≥ 128 bits"). Next, fusion calculations and key indicator screening are performed. Taking L3 high-sensitivity contract text data (text type) in the industrial field as an example, the vulnerability weight factors in the dynamic security assessment matrix are "data leakage risk coefficient 0.35, data integrity failure probability 0.09, resistance to attacks 0.24, backdoor detection coverage 0.07, access control strength 0.2, and operation log integrity 0.05"; the core parameters of the intelligent algorithm model are "SM9 encryption algorithm key length = 256 bits, number of federated learning nodes = 5 (cross-departmental collaboration)"; in the security control requirements data, the L3 high-sensitivity data standard is "data leakage risk coefficient ≤ 0.0008%, data integrity failure probability ≤ 0.0005%, resistance to attacks ≥ 98%, access control strength ≥ 95%, encryption algorithm strength ≥ 256 bits, and operation log integrity ≥ 99%".

[0038] A multi-objective collaborative computing mechanism (such as particle swarm optimization) is adopted, with "minimizing security risks" as the core objective, combined with the objectives of "graded accuracy" and "compliance." The matrix data and algorithm parameters are fused and calculated. For example, the "data leakage risk coefficient" has the highest weight (0.35) and needs to be optimized first. Using the SM9 encryption algorithm, it is calculated that it needs to be reduced from the current 0.001% to 0.0007% (meeting the standard). The "access control strength" has a weight of 0.2 and needs to be increased from 89% to 96% by adding approval levels (meeting the standard). The "adversarial attack tolerance" has a weight of 0.24 and needs to be upgraded from 95% to 98% (meeting the standard) by upgrading the adversarial training model. Simultaneously, in conjunction with the L3 high-sensitivity level standard, the weights of key features are strengthened, and the weights of "data leakage risk coefficient," "access control strength," and "encryption algorithm strength" are temporarily increased by 10% in the calculation to ensure they are prioritized for achievement. Ultimately, a combination of key grading indicators was selected for the target sensitivity level (L3 high sensitivity), including "data leakage risk coefficient ≤0.0008%, access control strength ≥95%, resistance to attacks ≥98%, SM9 encryption algorithm (256-bit), and operation log integrity ≥99%". These indicators cover the core risk points of the data layer, algorithm layer, and service layer, and are fully matched with the security requirements of L3 high sensitivity data in the industrial field.

[0039] S104 extracts data-sensitive feature values, hierarchical efficiency benchmark values, and security and hierarchical matching deviation values ​​based on data usage scenarios, classifies data sensitivity levels, and generates a multi-dimensional data hierarchical feature matrix.

[0040] In one implementation, the characteristics of data usage scenarios are analyzed using the principle of scenario-based feature extraction. This clarifies the differences in sensitive attributes, hierarchical efficiency requirements, and security matching standards of data under different scenarios, generating preset data sensitive feature value extraction rules, hierarchical efficiency benchmark thresholds, and allowable ranges for security and hierarchical matching deviations. The dimensions of scenario characteristic analysis are determined, focusing on three core dimensions: "sensitive attribute differences," "hierarchical efficiency requirements," and "security matching standards." Specifically, sensitive attribute differences require distinguishing the types of sensitive information (such as personal identification information, trade secrets, and public data) and their degree of sensitivity contained in the data within the scenario; hierarchical efficiency requirements require clarifying the time limit requirements for data hierarchical speed in the scenario (e.g., real-time transmission scenarios require second-level hierarchical processing, while offline storage scenarios can tolerate minute-level hierarchical processing); and security matching standards correspond to the security compliance requirements that the data hierarchical results must meet in the scenario (e.g., medical scenarios must comply with the Personal Information Protection Law).

[0041] Next, the system analyzes and generates preset rules for specific scenarios. Taking the "clinical diagnosis scenario" (data usage scenario) in the medical field as an example: Regarding the differences in sensitive attributes, the data in this scenario (such as patient CT images and electronic medical records) contains a large amount of personal privacy information (ID number, medical history), and the priority of sensitive attributes is "patient identity information > medical records > device parameters"; Regarding the requirements for classification efficiency, clinical diagnosis requires real-time data retrieval, and the classification time must be ≤1 second / data entry; Regarding the security matching standard, the classification result must meet the requirements of "encrypted storage of highly sensitive data and dual authorization for access". Based on this, preset rules are generated: The data sensitive feature value extraction rule is "the proportion of personal identity information fields in statistical data (such as the number of ID number and mobile phone number fields / the total number of fields), and the proportion of confidential fields in medical records"; The classification efficiency benchmark threshold is "the classification time for a single data entry is ≤1 second, and the batch processing (1000 entries) time is ≤5 minutes"; The allowable range for security and classification matching deviation is "the deviation rate between the actual classification result and the preset sensitivity level is ≤0.5% (such as the misclassification rate of data that should be classified as L3 highly sensitive is ≤0.5%)". Taking the "remote equipment monitoring scenario" in the industrial field as an example: Regarding the differences in sensitive attributes, the priority of sensitive attributes of data (such as equipment operating parameters and fault warning logs) is "equipment control commands > operating parameters > ambient temperature and humidity"; Regarding the requirements for graded efficiency, remote monitoring needs to be graded in near real-time, with a processing time of ≤3 seconds per data point; Regarding the security matching standards, the grading results need to meet the requirements of "encryption of high-sensitivity control command transmission and real-time alarm after grading abnormal parameters". Correspondingly, preset rules are generated: The data sensitive feature value extraction rules are "equipment control command field recognition rate (such as 'start / stop command' and 'parameter adjustment command' recognition accuracy rate ≥99%), and the proportion of abnormal values ​​of operating parameters"; The grading efficiency benchmark threshold is "grading time for a single data point ≤3 seconds, and batch processing (1000 data points) time ≤10 minutes"; The allowable range for security and grading matching deviation is "grading deviation rate ≤1% (because industrial non-privacy data has a slightly higher tolerance for deviation)".

[0042] By integrating the data security management requirement tags and feature value extraction results from multiple industries, a scenario-based hierarchical benchmark is constructed. The calculation rules for data sensitivity feature values, hierarchical efficiency benchmark values, and security and hierarchical matching deviation values ​​are determined through the matching logic between scenario requirements and feature values, establishing a scenario-based feature value quantification mechanism. The detailed dimensions of data security management requirement tags across multiple industries are clarified, requiring the association of "sensitivity level tags," "compliance requirement tags," and "business priority tags" according to the scenario. For example, the requirement tags for the "clinical diagnosis scenario" in the medical field are: "Sensitivity level tag: L3-L4 (high sensitivity), Compliance requirement tag: Complies with the 'Medical Data Security Guidelines,' Business priority tag: Emergency > General outpatient."

[0043] Next, the requirements tags and feature value extraction results are integrated. Taking the "confidential contract review scenario" in the industrial field as an example, the initial feature values ​​of the data in this scenario are first extracted: In a certain confidential contract text data, "the proportion of classified keywords (such as 'confidential' and 'top secret') = 5%, the circulation and approval level = 3 levels (department head → business head → general manager), and the integrity of the operation log = 98%". The requirements tags (sensitivity level tag L3, compliance requirement tag "confidential data must be encrypted, approval level ≥ 2 levels") are integrated to construct a scenario-based hierarchical benchmark: "the proportion of classified keywords", "circulation and approval level", and "operation log integrity" are used as core feature dimensions to determine the calculation rules. The calculation rule for data sensitivity feature value is "percentage of classified keywords × 0.6 + approval level of circulation × 0.3 + integrity of operation log × 0.1" (because classified keywords have the greatest impact on sensitivity level). For example, the sensitivity feature value of this confidential contract = 5% × 0.6 + 3 × 0.3 + 98% × 0.1 = 0.03 + 0.9 + 0.098 = 1.028 (equivalent to 102.8 points on a 100-point scale, out of a maximum of 150 points). The calculation rule for the graded efficiency benchmark value is "time spent on graded evaluation of a single contract = time spent on data preprocessing + time spent on feature extraction + time spent on grade determination". The contract preprocessing took 0.8 seconds, feature extraction took 0.5 seconds, and level determination took 0.3 seconds. The baseline value for leveling efficiency is 1.6 seconds (it needs to be ≤2 seconds to meet the scenario efficiency requirements). The calculation rule for the security and level matching deviation value is "|actual leveling result score - preset sensitivity level standard score| / preset standard score × 100%". For example, if the preset standard score for an L3 level contract is 100 points, the actual score of this contract is 102.8 points, and the deviation value is |102.8-100| / 100×100%=2.8% (it needs to be ≤3% to meet the allowable deviation range).

[0044] Simultaneously, a scenario-based feature value quantification mechanism is established to dynamically adjust the feature dimension weights for different scenarios. For example, in the medical "clinical diagnosis scenario," the weight of "patient identity information percentage" is set to 0.5, the weight of "confidentiality percentage of medical records" is set to 0.4, and the weight of "equipment parameter percentage" is set to 0.1; in the industrial "remote equipment monitoring scenario," the weight of "equipment control command recognition rate" is set to 0.5, the weight of "abnormal operating parameter value percentage" is set to 0.3, and the weight of "environmental temperature and humidity percentage" is set to 0.2, ensuring that the quantification results are highly compatible with the scenario's business logic.

[0045] Using data usage scenarios as the coverage dimension, this approach integrates sensitive feature value data, hierarchical efficiency benchmark data, and security and hierarchical matching deviation data from various scenarios. Combined with pre-defined sensitivity level classification standards, it defines the objects for hierarchical data processing in different scenarios. Determining the scenario coverage and data classification requires encompassing typical scenarios in core fields such as healthcare, industry, and education. Within each scenario, data is categorized and integrated according to data format (text, image, numerical value). For example, the education field covers "student registration management scenario," "online examination scenario," and "scientific research data sharing scenario." The "student registration management scenario" data includes student registration text data (such as student ID number and home address), student registration photo image data, and grade numerical data.

[0046] Next, the scenario data fusion and processing object definition are performed. Taking the "student registration management scenario" in the education field as an example: First, sensitive feature value data in this scenario (such as identity information accounting for 30% in student registration text data, facial recognition success rate in photo image data = 95%, sensitive field accounting for 5% in grade data), graded efficiency benchmark data (graded time for a single student registration data = 0.9 seconds, batch of 1000 data = 4 minutes), and security and graded matching deviation data (graded deviation rate = 0.3%) are fused together with the preset sensitivity level classification standard (sensitivity levels in the education field: L1 (low sensitivity, such as grade ranking), L2 (medium sensitivity, such as home address), L3 (high sensitivity, such as ID number, facial photo)) to classify the data. Specifically, student record text data containing ID card numbers is classified as L3 high-sensitivity processing objects, while data containing home addresses is classified as L2 medium-sensitivity processing objects. Student record photo image data, containing facial information, is classified as L3 high-sensitivity processing objects. Grade data containing only scores is classified as L1 low-sensitivity processing objects, while data containing student names and scores is classified as L2 medium-sensitivity processing objects. Furthermore, the priority of processing objects for each scenario is clearly defined. For example, in the "student record management scenario," L3 high-sensitivity data (ID card numbers, facial photos) has the highest processing priority (requiring priority for encryption and access control), followed by L2 medium-sensitivity data. L1 low-sensitivity data can be processed offline in batches, ensuring that tiered resources are allocated to high-priority data, meeting the business needs of the scenario.

[0047] Multi-type data is loaded through data standardization and anomaly filtering preprocessing mechanisms. Contextualized feature extraction rules are used to enhance the targeting of sensitive feature values, hierarchical efficiency benchmark values, and security and hierarchical matching deviation values. This is then integrated with preset sensitivity level classification standards to generate a multi-dimensional data hierarchical feature matrix. Data standardization and anomaly filtering are performed using a unified data format (such as JSON) and feature dimensions (such as "contextual label," "data type," "sensitive feature value," "hierarchical efficiency value," "matching deviation value," and "sensitivity level") to preprocess the multi-type data. For example, CT image data (image type), electronic medical record data (text type), and laboratory test data (numerical type) from the medical "clinical diagnosis scenario" are uniformly converted into JSON format. The standardized CT image data includes "Scene Label: Clinical Diagnosis, Data Type: Image, Sensitive Feature Value: 95 points (out of 100), Grading Efficiency Value: 0.8 seconds / data entry, Matching Deviation Value: 0.2%, Sensitivity Level: L3". Simultaneously, anomaly filtering algorithms (such as anomaly detection based on isolated forests) are used to remove invalid data, such as "age = 200 years old" data due to entry errors in electronic medical records, and "abnormal pixel values" data due to sensor malfunctions in CT images, ensuring data quality. Then, feature extraction is enhanced by combining scenario-based feature extraction rules, dynamically adjusting the extraction algorithm for different scenarios.

[0048] For example, in the text data of the industrial "contract circulation scenario," a keyword extraction algorithm based on the BERT pre-trained model is used to enhance the recognition accuracy of "confidential keywords" (such as "confidential" and "top secret"), improving the extraction rate to 99%. In the numerical data of the industrial "remote equipment monitoring scenario," an outlier detection algorithm based on a sliding window is used to enhance the feature extraction of "equipment control commands," improving the recognition accuracy to 98%. Finally, a multi-dimensional data grading feature matrix is ​​generated by integrating sensitivity level standards. The matrix's row dimension is "data usage scenario - data type" (such as "clinical diagnosis - CT images," "contract circulation - text," "student registration management - text"), and the column dimension is "sensitive feature value, grading efficiency value, matching deviation value, sensitivity level, scenario priority." The matrix elements are specific quantitative values ​​and labels. Taking the "Clinical Diagnosis - CT Imaging" row as an example, the element values ​​are "Sensitive Feature Value: 95 points, Grading Efficiency Value: 0.8 seconds / item, Matching Deviation Value: 0.2%, Sensitivity Level: L3, Scene Priority: High"; the element values ​​of the "Contract Flow - Text" row are "Sensitive Feature Value: 102.8 points, Grading Efficiency Value: 1.6 seconds / item, Matching Deviation Value: 2.8%, Sensitivity Level: L3, Scene Priority: Medium", forming a structured matrix covering multiple scenarios and types of data, providing accurate feature support for subsequent real-time grading and anomaly identification.

[0049] S105 processes real-time classification data based on a multi-dimensional data classification feature matrix, identifies abnormal signals such as classification standard deviation, imbalance between sensitive features and classification results, and generates dynamic correction factors for classification strategies using a causal inference network.

[0050] In one implementation, benchmark feature factors required for real-time graded data processing are extracted based on a multidimensional data grading feature matrix, constructing a graded data comparison benchmark system. The benchmark feature factor extraction uses data-sensitive feature values, grading efficiency benchmark values, security and grading matching deviation values, and sensitivity level labels from the multidimensional data grading feature matrix as core dimensions. The core dimensions and extraction rules of the benchmark feature factors are clearly defined, with the core dimensions strictly corresponding to the "data-sensitive feature values," "grading efficiency benchmark values," "security and grading matching deviation values," and "sensitivity level labels" in the multidimensional data grading feature matrix to ensure that the factors comprehensively reflect the core attributes of data grading. Specifically, data-sensitive feature values ​​extract the quantitative scores (e.g., percentage scores) of data from each scenario in the matrix; grading efficiency benchmark values ​​extract the time consumed for grading a single data point and the time consumed for batch processing; security and grading matching deviation values ​​extract the deviation rate between the actual grading result and the standard; and sensitivity level labels extract the corresponding L1-L4 level identifiers. The extraction rules must adhere to the "scenario consistency" principle, meaning that the factor extraction logic within the same domain and scenario remains consistent to avoid comparison errors caused by rule differences.

[0051] Next, taking the multidimensional data grading feature matrix of the "clinical diagnosis scenario" in the medical field as an example, the extraction was performed as follows: the baseline feature factors of a batch of CT image data (L3 high sensitivity) were extracted from the matrix, with a data sensitivity feature value of 95 points (out of 100, including a high proportion of patient privacy information), a grading efficiency baseline value of 0.8 seconds / data point per line, a batch of 1000 data points taking 4 minutes, a safety and grading matching deviation value of 0.2% (actual grading result and L3 standard deviation), and a sensitivity level label of L3; at the same time, the baseline factors of the test indicator data (L2 medium sensitivity) in the same scenario were extracted: the data sensitivity feature value of 65 points, the grading efficiency baseline value of 0.5 seconds / data point per line, a batch of 1000 data points taking 2.5 minutes, a safety and grading matching deviation value of 0.3%, and a sensitivity level label of L2. These factors are categorized and integrated according to "scenario-data type-factor dimension" to construct a hierarchical data comparison benchmark system. Within the system, the standard values ​​of the above four factors are clearly defined for "medical clinical diagnosis scenario-CT imaging (L3)" and another set of standard values ​​are defined for "medical clinical diagnosis scenario-testing indicators (L2)". This forms a benchmark library that can be directly used for comparison, providing a basis for the anomaly judgment of real-time data grading results.

[0052] The multidimensional data grading feature matrix and real-time grading data are processed. The feature dimensions and calculation methods of the real-time grading data are standardized to match the matrix's baseline feature factor dimensions, generating standardized real-time grading data. Standardization rules are formulated, specifying processing requirements for the four dimensions of the baseline feature factors: data sensitivity feature values ​​must be uniformly converted to a percentage system (e.g., the original calculation of "sensitive field percentage × 100" is uniformly adjusted to a percentage score of "sensitive field percentage × 0.8 + privacy leakage risk × 0.2"); the grading efficiency baseline value must have a unified time unit (single data processing time is in "seconds / data", batch processing time is in "minutes / 1000 data"); the security and grading matching deviation value must have a unified calculation logic (|real-time score - standard score| / standard score × 100%); sensitivity level labels must be uniformly identified as L1-L4 to avoid ambiguity in wording such as "high sensitivity / medium sensitivity".

[0053] Next, we take the real-time grading data of the "clinical diagnosis scenario" in the medical field as an example to perform standardization: the original grading result of a certain real-time CT image data is "sensitive field ratio 35%, single grading time 800 milliseconds, batch of 1000 data takes 250 seconds, and the actual grading label is 'high sensitivity'". According to standardized rules, the data sensitivity feature value is converted to 35% × 0.8 + privacy leakage risk 0.1 × 0.2 = 28% + 2% = 30%. However, the current result is incorrect and needs to be recalculated using a percentage system. In the original matrix, the CT image data sensitivity feature value of 95 points corresponds to a 30% sensitive field ratio, so the real-time data sensitivity field ratio is 32%. Proportionally, this is converted to (32% / 30%) × 95 points ≈ 98.67 points. In the graded efficiency benchmark value, 800 milliseconds is converted to 0.8 seconds / item, and 250 seconds is converted to 250 / 60 ≈ 4.17 minutes / 1000 items. The security and graded matching deviation value is calculated as |98.67 points - 95 points| / 95 points × 100% ≈ 3.86%. The sensitivity level label "high sensitivity" is uniformly corrected to L3. The final standardized real-time classification data is as follows: data sensitivity feature value of 98.67 points, classification efficiency benchmark value of 0.8 seconds / single (single), 4.17 minutes / 1000 (batch), safety and classification matching deviation value of 3.86%, and sensitivity level label L3. This data is completely matched with the factor dimensions of "medical clinical diagnosis scenario - CT image (L3)" in the benchmark system and can be directly used for subsequent comparison.

[0054] A dual-data comparison and analysis model is constructed based on standardized real-time grading data and multi-dimensional data grading feature matrices. By calculating the deviation between real-time and historical data on core benchmark feature factors, it identifies anomalous signals such as grading standard offsets and imbalances between sensitive features and grading results. The comparison model employs a logic of "dimensional deviation calculation + comprehensive anomaly judgment": For the four benchmark factor dimensions, the deviation between real-time data and benchmark values ​​is calculated (e.g., data sensitive feature value deviation = |real-time score - benchmark score| / benchmark score × 100%, grading efficiency deviation = |real-time time consumption - benchmark time consumption| / benchmark time consumption × 100%), and deviation thresholds are set for each dimension (e.g., data sensitive feature value deviation threshold ≤ 5%, grading efficiency deviation threshold ≤ 10%, safety and grading matching deviation threshold ≤ 1%, sensitivity level labels must be completely consistent). When any dimension deviation exceeds the threshold, or the sensitivity level label does not match, anomaly signal judgment is triggered, and the anomaly type (grading standard offset / feature-result mismatch) is labeled.

[0055] Next, taking the "contract circulation scenario" in the industrial field as an example, comparison and identification were performed: The baseline factors for "Industrial Contract Circulation Scenario - Confidential Contract (L3)" were retrieved from the baseline system: data sensitivity feature value 102.8 points, hierarchical efficiency baseline value 1.6 seconds / item, security and hierarchical matching deviation value 2.8%, and sensitivity level label L3. Standardized hierarchical data for a real-time confidential contract in this scenario was obtained: data sensitivity feature value 90 points, hierarchical efficiency baseline value 2.0 seconds / item, security and hierarchical matching deviation value 8%, and sensitivity level label L2. The deviations for each dimension were calculated: data sensitivity feature value deviation = |90-102.8| / 102.8×100%≈12.45% (exceeding the 5% threshold), hierarchical efficiency deviation = |2.0-1.6| / 1.6×100%=25% (exceeding the 10% threshold), security and hierarchical matching deviation value 8% (exceeding the 1% threshold), and the sensitivity level label L2 did not match the baseline L3. Based on the anomaly detection rules, two types of anomaly signals were identified: first, a shift in the grading standard (the data sensitive feature value, grading efficiency, and matching deviation all exceed the standard, indicating that the current grading standard has shifted from the benchmark); second, an imbalance in the matching between sensitive features and grading results (the data sensitive feature value of 90 points should have corresponded to L3, but was misjudged as L2, indicating a mismatch between the feature and the grade). At the same time, the specific deviation values ​​of each anomaly dimension were recorded to provide a basis for subsequent causal analysis.

[0056] The identified anomalous signals and their correlation with features in the multidimensional data grading feature matrix are input into a causal inference network. This network analyzes the causal links generated by the anomalous signals, calculates the weights of the anomalous signals' influence on the grading results, and generates dynamic correction factors for grading strategies adapted to different anomalous scenarios. The causal inference network employs a deep spatiotemporal causal model architecture. The input layer consists of anomalous signals (e.g., a 12.45% deviation in data sensitivity feature values, or a mismatch in sensitivity levels) and the correlations between features in the multidimensional data grading feature matrix (e.g., the mapping relationship between data sensitivity feature values ​​and sensitivity levels, and the correlation between grading efficiency and the number of algorithm iterations). The hidden layer analyzes the potential causes of the anomalous events through a causal graph (e.g., untimely algorithm model iteration, human annotation errors, or updates to compliance standards). The output layer contains the causal links and their influence weights. Network training must be based on historical anomalous data (e.g., grading annomalous cases from the past year) to ensure inference accuracy.

[0057] Next, taking the abnormal signal of the "remote equipment monitoring scenario" in the industrial field as an example, the inference and correction factor generation are performed: The abnormal signal of a certain real-time equipment control command data (should be L3 high sensitivity) is "data sensitivity feature value 55 points (benchmark 70 points, deviation 21.4%), sensitivity level label L1 (mismatch)". This is then input into the causal inference network along with the correlation between "equipment control command - sensitive feature value - sensitivity level" in the matrix (sensitive feature value ≥ 65 points corresponds to L3, 50-64 points corresponds to L2, < 50 points corresponds to L1). The network analyzes the causal links: the reason for the low data sensitivity feature value is "the equipment control command recognition algorithm has not been updated, and new commands are not recognized as sensitive fields" (impact weight 70%), and the sensitivity level mismatch is "the low feature value causes the level judgment logic to erroneously trigger the L1 standard" (impact weight 30%). Based on the influence weight calculation, a dynamic correction factor for the grading strategy is generated: To address the issue of outdated algorithms, an "iteration coefficient of 1.2 for the device control command recognition algorithm" is generated (to improve the algorithm's recognition rate for new commands, correcting the feature value to 55 × 1.2 ≈ 66 points). To address the level determination logic issue, a "sensitivity level determination threshold correction coefficient of 0.9" is generated (to lower the L3 threshold from 65 points to 58.5 points, ensuring that 66 points can be correctly determined as L3)." These two correction factors can be directly used to adjust subsequent grading strategies, ensuring that the grading results of real-time data return to the correct standard. For example, the corrected data sensitivity feature value of 66 points and sensitivity level label L3 are consistent with the benchmark system, resolving anomaly issues.

[0058] S106 uses domain-grouped data, meta-learning to filter key factors, integrates constraints and blockchain traceability limit information to construct an adaptive data dynamic hierarchical model, and combines time series patterns to output dynamic hierarchical results and full-cycle management strategies.

[0059] In one implementation, correlation analysis is performed on domain-specific data grouped by domain and domain-specific feature data in a multidimensional data grading feature matrix to generate key domain grading factors and dynamic correction factors for grading strategies. Domain grouping and correlation dimensions are clearly defined, with grouping based on core domains such as healthcare, industry, and education. Within each domain, the correlation dimension of "data form (text / image / numerical) - usage scenario - sensitivity level" is focused. Domain-specific data includes business data (e.g., patient treatment data in the healthcare field) and compliance standards (e.g., the Medical Personal Information Protection Law). Domain-specific feature data in the multidimensional data grading feature matrix includes historical features such as sensitivity feature values, grading efficiency values, and matching deviation values ​​for each domain.

[0060] Next, taking the medical field as an example, a correlation analysis was performed: Medical-specific data includes CT image data (L3 high sensitivity, containing patient privacy), electronic medical record text data (L3 high sensitivity), laboratory indicator numerical data (L2 medium sensitivity), and medical data compliance standards (privacy leakage rate ≤0.001%, encryption algorithm must comply with national medical data security standards). Medical-specific characteristic data was extracted from the multi-dimensional data grading feature matrix: CT image data had a sensitivity feature value of 95 points, grading efficiency of 0.8 seconds / item, and matching deviation of 0.2%; electronic medical records had a sensitivity feature value of 92 points, grading efficiency of 1.0 second / item, and matching deviation of 0.3%. Through correlation analysis, key grading factors in the medical field were identified—"proportion of patient privacy information (weight 0.4)," "proportion of confidential fields in medical data (weight 0.3)," "compliance of encryption algorithms (weight 0.2)," and "grading response timeliness (weight 0.1)." These factors are directly related to the core requirements of medical data security.

[0061] Simultaneously, by combining historical cases of abnormal classification in the medical field (such as classification deviations caused by inaccurate privacy recognition of CT images), dynamic correction factors for the classification strategy are generated—"Privacy Field Recognition Algorithm Iteration Coefficient 1.1 (to improve the recognition rate of privacy information)" and "Classification Timeliness Compensation Coefficient 0.95 (to appropriately relax the timeliness threshold due to the priority processing of medical data)"—ensuring that the correction factors can specifically address the pain points of classification in the medical field. Similarly, after correlation analysis in the industrial field, key classification factors are generated: "Contract Confidentiality Level (Weight 0.4)", "Circulation Approval Level (Weight 0.3)", "Operation Log Integrity (Weight 0.2)", and "Cross-Departmental Transmission Security (Weight 0.1)", as well as dynamic correction factors: "Confidential Keyword Recognition Optimization Coefficient 1.2" and "Approval Link Verification Coefficient 1.05", which are highly consistent with the compliance requirements of contract classification in the industrial field.

[0062] We conduct weight fusion and effectiveness verification analysis on key domain grading factors and dynamic correction factors of grading strategies to generate a quantitative matrix of key domain factors. We formulate weight fusion rules, adopting a "domain business priority weighting method," allocating weights according to the degree of impact of each factor on domain data security, with key grading factors and correction factors having weight ratios of 0.8 and 0.2 respectively (key factors determine the core logic of grading, and correction factors assist in optimization). The fusion formula is "fused factor value = key grading factor value × sum of corresponding weights × correction factor coefficient," ensuring that the correction factor can effectively adjust the quantitative results of key factors.

[0063] Next, the industrial sector was used as an example for fusion and verification. The key classification factors in the industrial sector were "equipment control command sensitivity (weight 0.5)," "percentage of outliers in operating parameters (weight 0.3)," and "environmental data involvement density (weight 0.2)." The key factor values ​​for a batch of equipment control command data were 90, 85, and 60 points, respectively. The dynamic correction factors were "control command recognition algorithm iteration coefficient 1.2" and "outlier detection accuracy coefficient 1.1," with an average correction factor of 1.15. The fused factor value was calculated using the formula: (90×0.5 + 85×0.3 + 60×0.2)×1.15 = (45 + 25.5 + 12)×1.15 = 82.5×1.15 = 94.875 points. Subsequently, effectiveness verification was conducted using the "historical data backtesting method." The fused factor values ​​were substituted into 1000 historical data classification cases from the past 6 months in the industrial sector to verify the classification accuracy. If the verification results show that the data classification accuracy after fusion factor increases from 88% to 96% and the misclassification rate decreases from 5% to 1.2%, meeting the requirements of "classification accuracy ≥ 95% and misclassification rate ≤ 2%" in the industrial field, then the fusion result is deemed valid.

[0064] The factor values ​​of various industrial data types (such as equipment control commands, operating parameters, and environmental data) are fused and organized according to "data type - factor dimension" to generate a quantitative matrix of key factors in the industrial field. The matrix clearly shows that "equipment control commands (L3)" corresponds to a fused factor value of 94.875 points, "operating parameters (L2)" corresponds to 82 points, and "environmental data (L1)" corresponds to 65 points, providing standardized quantitative data for model input.

[0065] Based on the fundamental feature factors and key factor quantification matrix of domain-specific data, and combined with the hierarchical architecture of the adaptive dynamic data grading model, the sensitivity level classification standards, encryption constraints, and blockchain traceability limit information of different domains are fused and processed to generate domain-factor-constraint correlation features. The model's hierarchical architecture and constraints are clearly defined. The adaptive dynamic data grading model adopts a four-layer architecture: "data layer - factor layer - constraint layer - decision layer": the data layer inputs the fundamental features of the domain, the factor layer inputs the key factor quantification matrix, the constraint layer inputs domain-specific constraints (sensitivity level standards, encryption conditions, and blockchain traceability limits), and the decision layer outputs the grading results. Constraints in each field need to be refined, such as the sensitivity level standards in the medical field (L3: containing patient privacy data, L2: non-privacy data in diagnosis and treatment, L1: public statistical data), encryption constraints (L3 data requires SM9 encryption, L2 requires AES-256 encryption), and blockchain traceability limits (traceability latency ≤ 1 second, evidence storage capacity ≥ 10TB / year); and the sensitivity level standards in the industrial field (L3: confidential contracts, L2: ordinary confidential contracts, L1: public documents), encryption constraints (L3 requires national cryptographic algorithm encryption, L2 requires symmetric encryption), and blockchain traceability limits (traceability immutability rate ≥ 99.99%). Next, taking the education field as an example, the following related features are generated: the basic features for data classification in the education field are "student registration text data (including ID number), grade numerical data, and student registration photo image data".

[0066] In the key factor quantification matrix, the student record text data has a fusion factor value of 92 (L3), the academic performance data has 80 (L2), and the photo data has 95 (L3). The constraints in the education domain are: sensitivity level standard (L3: includes ID number / facial photo, L2: includes home address / grade, L1: public ranking), encryption constraint (L3 uses SM9 encryption, L2 uses differential privacy protection), and blockchain traceability limit (student data traceability retention ≥ 5 years, immutability rate ≥ 99.98%). Combining the four-layer architecture of the model, the basic features, quantification factors, and constraints are associated according to "data type - factor value - constraint standard" to generate education domain - factor - constraint associated features, such as "student record text data (L3) - fusion factor value 92 - SM9 encryption - traceability retention 5 years" and "academic performance data (L2) - fusion factor value 80 - differential privacy - traceability retention 3 years". Each feature contains data attributes, factor quantification value, and constraint requirements, which can be directly called by the model decision layer.

[0067] An adaptive data dynamic grading model is used to analyze and process the correlation features between domain, factors, and constraints. Combined with the temporal changes in data flow, dynamic grading results and full-cycle management strategies are generated. The model's analytical logic and the way it combines temporal patterns are clearly defined. The model's decision layer adopts the logic of "multi-factor weighted voting + temporal trend prediction": First, based on the factor values ​​and constraints in the correlation features, a comprehensive score for data grading is calculated, and sensitivity levels are divided according to the score; then, combined with the temporal patterns of data flow (e.g., medical data must be grading and stored within 1 hour of collection, and industrial contracts must be transferred across departments within 24 hours), the grading pace and the time nodes of management measures are adjusted. Next, taking the industrial field as an example, model analysis and strategy output are performed: The associated features of "confidential contract text data (L3) - fusion factor value of 96 points - national cryptographic algorithm encryption - traceability and tamper-proof rate of 99.99%" are input into the model. The decision layer calculates the comprehensive score = 96 × 0.8 (factor weight) + 99.99% × 0.2 (constraint weight) = 76.8 + 0.19998 ≈ 76.9998 points, which is judged as L3 high sensitivity level. At the same time, the full-cycle control strategy is output.

[0068] For the data collection phase, contracts containing confidential keywords (such as "confidential" and "top secret") are verified in real time to ensure no sensitive information is missed, and initial classification and labeling are completed within 30 minutes of collection. For the storage phase, the national cryptographic algorithm SM9 is used for encrypted storage, and storage partitions are dedicated to "L3 high-sensitivity" partitions. Only storage nodes with a trust score of ≥95 can host the data, and blockchain traceability records the storage location and access logs. For the usage phase, cross-departmental access requires dual authorization, operation logs are uploaded to the blockchain in real time, and usage traces are cleared within one hour after use. For the archiving phase, the classification results are verified again before archiving to ensure no classification deviation. After archiving, the data is stored for a 10-year period, and a blockchain traceability integrity check is conducted quarterly. This strategy covers the entire lifecycle of industrial contracts, and the timeline nodes are perfectly matched with the rhythm of industrial business processes, making it ready for direct implementation.

[0069] In one implementation, such as Figure 2 As shown, this application also provides an adaptive data dynamic classification device based on an intelligent algorithm model, comprising: The acquisition module 201 is used to acquire multimodal feature data of the entire data lifecycle, core parameters of intelligent algorithm models, and data security management and control requirement labels of multiple industries. It adopts a cross-source data joint verification model based on supervised learning, and removes abnormal values ​​and redundant information in data collection by training verification rules through labeled abnormal samples. It also filters random interference signals in the data transmission process by combining noise robustness machine learning algorithms, and outputs a standardized full-cycle data feature sequence. Processing module 202 is used to convert standardized full-cycle data feature sequences into a security assessment map, and to construct a comprehensive security assessment framework based on adversarial generative networks to subdivide incomplete data samples. Based on this framework, it extracts vulnerability weight factors for different data types to construct a dynamic security assessment matrix, and performs collaborative calculations on the core parameters of the intelligent algorithm model and data security control requirements to obtain a combination of key grading indicators under the target sensitivity level. Based on data usage scenarios, it extracts data sensitivity feature values, grading efficiency benchmark values, and security-grading matching deviation values ​​to classify data sensitivity levels and generate a multi-dimensional data grading feature matrix. Based on this matrix, it processes real-time grading data, identifies abnormal signals such as grading standard deviations and imbalances between sensitive features and grading results, and uses a causal inference network to generate dynamic correction factors for grading strategies. It groups data by domain, uses meta-learning to filter key factors, and integrates constraint conditions and blockchain traceability limit information to construct an adaptive dynamic data grading model. Finally, it outputs dynamic grading results and full-cycle control strategies based on temporal patterns.

[0070] The various embodiments in this application are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments for evaluating the adaptive data dynamic classification method based on intelligent algorithm models, electronic devices, electronic devices, and readable storage media are basically similar to the above-described embodiments of the adaptive data dynamic classification method based on intelligent algorithm models, and therefore the descriptions are relatively simple. Relevant parts can be referred to in the descriptions of the above-described embodiments of the adaptive data dynamic classification method based on intelligent algorithm models.

Claims

1. An adaptive dynamic data grading method based on an intelligent algorithm model, characterized in that, include: Acquire multimodal feature data throughout the entire data lifecycle, core parameters of intelligent algorithm models, and data security management and control requirement tags from multiple industries. Adopt a cross-source data joint verification model based on supervised learning, and eliminate abnormal values ​​and redundant information in data collection by training verification rules through labeled abnormal samples. Combine noise robustness machine learning algorithms to filter random interference signals during data transmission and output a standardized full-cycle data feature sequence. The standardized full-cycle data feature sequence is transformed into a security assessment map, and a comprehensive security assessment framework is constructed based on the incomplete data samples subdivided by adversarial generative networks. Based on the comprehensive security assessment framework, vulnerability weight factors of different data types are extracted to construct a dynamic security assessment matrix. The core parameters of the intelligent algorithm model and the data security management and control requirements are calculated collaboratively to obtain the combination of key graded indicators under the target sensitivity level. Based on the data usage scenarios, extract data-sensitive feature values, hierarchical efficiency benchmark values, and security and hierarchical matching deviation values ​​to classify data sensitivity levels and generate a multi-dimensional data hierarchical feature matrix. Based on the multidimensional data classification feature matrix, real-time classification data is processed to identify abnormal signals such as classification standard deviation, imbalance between sensitive features and classification results, and to generate dynamic correction factors for classification strategies using causal inference networks. Data is grouped by domain, key factors are selected through meta-learning, and adaptive dynamic data hierarchical model is constructed by integrating constraints and blockchain traceability limit information. Dynamic hierarchical results and full-cycle management strategies are output by combining time series patterns.

2. The method as described in claim 1, characterized in that, The standardized full-cycle data feature sequence is transformed into a security assessment map. Based on the adversarial generative network, incomplete data samples are further subdivided to construct a comprehensive security assessment framework, including: The standardized full-cycle data feature sequence is processed and converted into a multi-dimensional data security assessment map through multi-dimensional feature encoding, generating basic security assessment visualization data; The multi-dimensional data security assessment map is imported into the generator module of the adversarial generative network. The conditional adversarial generative network architecture is used to synthesize and subdivide the scarce feature samples in the data incomplete scenario, and generate an expanded complete feature sample set. The expanded and complete feature sample set is iteratively optimized, and new attack scenarios are simulated by combining continuous attack response technology to generate an enhanced sample set with attack scenario labels. Among them, the adversarial sample evolution algorithm tracks the changes in attack patterns in real time and automatically triggers sample iteration when the attack features are updated. Collect vulnerability characteristic parameters of different data types in open environments, construct a data type-based vulnerability constraint library, and generate data type constraint data. The constraint library is customized with parameters for the data characteristics of multiple industries. By integrating enhanced sample sets and data type constraint data, and combining noise robustness testing and privacy inversion testing, a multi-level stability reverse verification system is constructed to generate a comprehensive security assessment framework.

3. The method as described in claim 1, characterized in that, Based on a comprehensive security assessment framework, vulnerability weight factors for different data types are extracted to construct a dynamic security assessment matrix. This matrix is ​​then used to collaboratively calculate the core parameters of the intelligent algorithm model and data security management requirements, yielding a combination of key grading indicators for the target sensitivity level, including: The feature weight extraction process matches the multi-dimensional vulnerability features in the comprehensive security assessment framework with the preset vulnerability assessment standards for different data types, extracts the vulnerability parameters of each data type at the data layer, algorithm layer, and service layer, and generates a data type-vulnerability feature mapping table and initial vulnerability weight values. By aligning with the data security management and control requirements standards of multiple industries and the core parameters of intelligent algorithm models, a vulnerability weight calibration model is constructed. Through the logic of matching requirements and parameters, the influence weight of each vulnerability feature is calculated, and a dynamic adjustment mechanism for vulnerability weight factors of different data types is established. Using data type as the dividing dimension, a dynamic security assessment matrix is ​​constructed by integrating the data type and vulnerability feature mapping table, the initial value of vulnerability weight, and the adjusted weight output by the vulnerability weight calibration model. The matrix row dimension represents data type, the column dimension represents vulnerability feature, and the matrix elements are the calibrated vulnerability weight factors. By using a multi-objective collaborative computing mechanism, dynamic security assessment matrix data, core parameters of intelligent algorithm models, and data security management requirements data are fused and calculated. The key feature weights are strengthened by combining the preset target sensitivity level assessment standards with the output features of the vulnerability weight calibration model to generate a combination of key graded indicators under the target sensitivity level.

4. The method as described in claim 1, characterized in that, Based on data usage scenarios, data sensitivity feature values, hierarchical efficiency benchmark values, and security and hierarchical matching deviation values ​​are extracted to classify data sensitivity levels and generate a multi-dimensional data hierarchical feature matrix, including: By analyzing the characteristics of data usage scenarios through the principle of scenario-based feature extraction, we can clarify the differences in sensitive attributes of data, hierarchical efficiency requirements and security matching standards under different scenarios, and generate preset data sensitive feature value extraction rules, hierarchical efficiency benchmark thresholds and allowable ranges for security and hierarchical matching deviations. By connecting the data security management and control requirements of multiple industries with the extracted labels and feature values, a scenario-based hierarchical benchmark is constructed. The calculation rules for data sensitive feature values, hierarchical efficiency benchmark values, and security and hierarchical matching deviation values ​​are determined through the matching logic of scenario requirements and feature values. A scenario-based feature value quantification mechanism is established. Using data usage scenarios as the coverage dimension, it integrates sensitive feature value data, hierarchical efficiency benchmark data, and security and hierarchical matching deviation data in various scenarios, and combines them with preset sensitivity level classification standards to define the data hierarchical processing objects in different scenarios. By loading multiple types of data through data standardization and anomaly filtering preprocessing mechanisms, and combining scenario-based feature extraction rules to enhance the targeting of sensitive feature values, hierarchical efficiency benchmark values, and security and hierarchical matching deviation values, and integrating them with preset sensitivity level classification standards, a multi-dimensional data hierarchical feature matrix is ​​generated.

5. The method as described in claim 4, characterized in that, Real-time classification data is processed based on a multi-dimensional data classification feature matrix to identify anomalous signals such as classification standard offset, imbalance between sensitive features and classification results, and to generate dynamic correction factors for the classification strategy using a causal inference network. Based on the multidimensional data hierarchical feature matrix, the benchmark feature factors required for real-time hierarchical data processing are extracted, and a hierarchical data comparison benchmark system is constructed. The benchmark feature factor extraction takes the data sensitivity feature value, hierarchical efficiency benchmark value, security and hierarchical matching deviation value and sensitivity level label in the multidimensional data hierarchical feature matrix as the core dimensions. The multidimensional data hierarchical feature matrix and real-time hierarchical data are processed. The feature dimensions and calculation methods of the real-time hierarchical data are standardized to match the matrix baseline feature factor dimensions, thereby generating standardized real-time hierarchical data. A dual-data comparison and analysis model is constructed based on standardized real-time grading data and multi-dimensional data grading feature matrix. By calculating the deviation between real-time and historical data on core benchmark feature factors, abnormal signals such as grading standard deviation and imbalance between sensitive features and grading results are identified. The identified abnormal signals and the feature correlations in the multidimensional data hierarchical feature matrix are input into the causal inference network to analyze the causal links generated by the abnormal signals, calculate the influence weight of the abnormal signals on the hierarchical results, and generate dynamic correction factors for hierarchical strategies that are adapted to different abnormal scenarios.

6. The method as described in claim 1, characterized in that, Data is grouped by domain, key factors are selected through meta-learning, and an adaptive dynamic data grading model is constructed by integrating constraints and blockchain traceability limit information. This model, combined with time-series patterns, outputs dynamic grading results and full-cycle management strategies, including: A correlation analysis is performed on the domain-specific data grouped by domain and the domain feature data in the multidimensional data hierarchical feature matrix to generate key domain hierarchical factors and dynamic correction factors for hierarchical strategies. We perform weight fusion and effectiveness verification analysis on key domain grading factors and dynamic correction factors of grading strategies to generate a quantitative matrix of key domain factors. Based on the domain data classification basic feature factors and domain key factor quantification matrix, combined with the hierarchical architecture of the adaptive data dynamic classification model, the sensitivity level classification standards, encryption constraints and blockchain traceability limit information of different domains are fused and processed to generate domain-factor-constraint correlation features. Based on the adaptive data dynamic classification model, the domain-factor-constraint correlation characteristics are analyzed and processed. Combined with the data flow time sequence change pattern, dynamic classification results and full-cycle management and control strategies are generated.

7. An adaptive data dynamic grading device based on an intelligent algorithm model, characterized in that, The device includes: The acquisition module is used to acquire multimodal feature data throughout the entire data lifecycle, core parameters of intelligent algorithm models, and data security management and control requirement tags from multiple industries. It removes abnormal values ​​and redundant information in data collection by training and verifying rules through labeled abnormal samples, and filters random interference signals during data transmission by combining noise robustness machine learning algorithms, and outputs a standardized full-cycle data feature sequence. The processing module transforms standardized full-cycle data feature sequences into a security assessment map. Based on adversarial generative networks, it subdivides incomplete data samples and constructs a comprehensive security assessment framework. Based on this framework, it extracts vulnerability weight factors for different data types to construct a dynamic security assessment matrix. It then collaboratively calculates core parameters of the intelligent algorithm model with data security management requirements to obtain key grading indicator combinations under the target sensitivity level. Based on data usage scenarios, it extracts data sensitivity feature values, grading efficiency benchmark values, and security-grading matching deviation values ​​to classify data sensitivity levels and generate a multi-dimensional data grading feature matrix. Based on this matrix, it processes real-time grading data, identifies abnormal signals such as grading standard deviations and imbalances between sensitive features and grading results, and uses a causal inference network to generate dynamic correction factors for grading strategies. Finally, it groups data by domain, uses meta-learning to filter key factors, integrates constraints and blockchain traceability limit information to construct an adaptive dynamic data grading model, and outputs dynamic grading results and full-cycle management strategies based on temporal patterns.

8. An electronic device, characterized in that, include: First processor; and a memory for storing executable instructions of the first processor; wherein the first processor is configured to execute the adaptive data dynamic classification method based on any one of claims 1 to 6 by executing the executable instructions.