Dynamic hierarchical data encryption method and system fusing scene and compliance
By employing a dynamic hierarchical data encryption method that combines semantic roles, industry risks, and compliance coefficients, dynamic hierarchical encryption of multiple data types is achieved. This solves the problems of data type adaptability, compliance adaptability, and sensitivity quantification in existing technologies, thereby improving data security and protection efficiency.
Patent Information
- Application Number
- CN202511822123.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-05
- Publication Date
- 2026-01-02
AI Technical Summary
Existing data encryption and hierarchical protection schemes suffer from poor data type adaptability, static hierarchical rules, lagging and inefficient compliance adaptation, coarse sensitivity quantification, and weak linkage between encryption strategies and risks, making it difficult to meet the security protection needs of various data types, dynamic scenarios, and compliance requirements.
A dynamic hierarchical data encryption method that integrates scenario and compliance is adopted. By acquiring multi-source data, extracting keywords and assigning weights according to semantic roles and industry risk coefficients, and combining data type feature values, scenario coefficients and compliance coefficients, the data sensitivity level is dynamically determined, and a differentiated encryption strategy is used for processing.
It achieves unified hierarchical encryption for structured, unstructured, and semi-structured data, supports real-time adaptation to changes in business scenarios and regulations, accurately quantifies data sensitivity, builds a closed-loop dynamic linkage mechanism, ensures data security, avoids resource waste, and reduces labor costs and overall protection costs.
Smart Images

Figure CN121262020A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data security, and in particular to a dynamic hierarchical data encryption method and system integrating scenarios and compliance. BACKGROUND
[0002] With the rapid development of digital economy, data as a core production factor, its types are increasingly diversified, covering structured data (such as database records), unstructured data (such as text documents, image files) and semi-structured data (such as JSON, XML files), and is widely used in government affairs, finance, e-commerce and other industries. Data encryption and hierarchical protection is a key means to protect data security, but there are many problems in the actual application of the existing technology that need to be solved: Poor data type adaptability: existing hierarchical encryption schemes are mostly designed for specific types of data (such as bill data, financial unstructured data), lack of universality, and cannot achieve unified hierarchical encryption processing of structured, unstructured and semi-structured data of all types, making it difficult to meet the diversified data security needs in multiple scenarios.
[0003] Static hierarchical rules, insufficient scene adaptation: traditional data classification is mostly based on pre-set sensitive word libraries to develop fixed rules, which can only achieve static classification and cannot adjust the classification results in real time with changes in business scenarios, resulting in mismatch between classification results and actual scenario risks.
[0004] Compliance adaptation lags behind and is inefficient: existing solutions require manual analysis of safety regulations and manual updates of encryption strategies to respond to changes in regulations, which not only lags behind, but also requires a large amount of manpower, making it difficult to adapt to dynamic revisions and new requirements of regulations in real time.
[0005] Quantitative sensitivity is rough: the classification is based on a one-sided basis, only through simple sensitive word matching or single-dimensional evaluation to determine the sensitivity level, the quantitative precision is insufficient, and the actual sensitivity of the data cannot be accurately reflected.
[0006] Weak linkage between encryption strategy and risk: the encryption strength and data risk level of the existing solution lack precise linkage, which is prone to over-encryption or insufficient encryption, and cannot achieve a balance between security protection and resource efficiency.
[0007] The above problems result in the difficulty of the existing data encryption and hierarchical protection scheme to meet the security protection needs of the current multi-type data, dynamic scenarios and compliance requirements. SUMMARY
[0008] Therefore, it is necessary to provide a dynamic hierarchical data encryption method and system integrating scenarios and compliance to solve at least one of the above problems in the existing technology.
[0009] In a first aspect, the embodiments of the present application are implemented in the following manner. A dynamic hierarchical data encryption method fusing a scene and compliance is provided, comprising: Obtaining multi-source data and preprocessing the same to obtain to-be-encrypted data; Extracting keywords in the to-be-encrypted data and assigning weight coefficients to the keywords according to semantic roles and / or industry risk coefficients; Determining feature values corresponding to each keyword according to data types; Determining a scene coefficient and a compliance coefficient, wherein the scene coefficient is used to dynamically reflect the influence of a business scene on a sensitive level, and the compliance coefficient is used to map mandatory requirements of regulations on the sensitive level; Determining a data sensitive level corresponding to the to-be-encrypted data based on the weight coefficients, the feature values, the scene coefficient, and the compliance coefficient; Encrypting the to-be-encrypted data by using an encryption strategy corresponding to the data sensitive level.
[0010] In a possible implementation, the assigning of the weight coefficients to the keywords according to the semantic roles and / or the industry risk coefficients comprises: Dividing the keywords into different roles according to semantics; Assigning corresponding weight values to the keywords corresponding to different roles; and / or Establishing an industry sensitive word library; Associating a corresponding industry risk coefficient with each keyword based on the industry sensitive word library; Assigning different weight values to the same keyword in different industry scenes based on the industry risk coefficient.
[0011] In a possible implementation, the determining of the feature values corresponding to each keyword according to data types comprises: If the data type is text data, determining the feature value based on the number of occurrences of the keyword and the total number of effective characters in the text; If the data type is image data, determining the feature value based on the number of sensitive region pixels and the total number of image pixels; If the data type is structured data, determining the feature value based on the number of record lines containing the keyword and the total number of record lines.
[0012] In a possible implementation, the scene coefficient is determined in the following manner: Determining risk values of a business scene from multiple business dimensions; Performing corresponding weighted calculation on each risk value according to a preset dimension weight to obtain the scene coefficient.
[0013] In a possible implementation, the compliance coefficient is determined by the following manner: acquiring preset compliance content according to a preset acquisition frequency; analyzing new or revised content in the preset compliance content; transforming the new or revised content into an adjustment rule for the compliance coefficient by using a preset association model; determining the compliance coefficient based on the adjustment rule.
[0014] In a possible implementation, the data sensitivity level corresponding to the data to be encrypted is determined based on the weight coefficient, the feature value, the scene coefficient and the compliance coefficient, including: determining a data sensitivity score corresponding to the data to be encrypted based on the weight coefficient, the feature value, the scene coefficient and the compliance coefficient; determining the data sensitivity level based on the data sensitivity score.
[0015] In a possible implementation, the data to be encrypted is encrypted by using an encryption strategy of a corresponding level based on the data sensitivity level, including: determining a corresponding encryption key, encryption module and protection mechanism based on the data sensitivity level; if the data sensitivity level of the data to be encrypted is a high level, using an encryption key with a first preset time length of a life cycle, using a hardware cryptographic module for encryption, and protecting the data to be encrypted in terms of confidentiality and integrity; if the data sensitivity level of the data to be encrypted is a medium level, using an encryption key with a second preset time length of a life cycle, using a hardware cryptographic module for encryption, and protecting the data to be encrypted in terms of confidentiality; if the data sensitivity level of the data to be encrypted is a low level, using an encryption key with a third preset time length of a life cycle, using a software cryptographic module for encryption, and protecting the data to be encrypted in terms of integrity; wherein the first preset time length is less than the second preset time length, and the second preset time length is less than the third preset time length.
[0016] In a second aspect, a dynamic grading data encryption system fusing a scene and compliance is provided, including: a data to be encrypted acquisition unit configured to acquire multi-source data, and perform preprocessing to obtain data to be encrypted; a weight coefficient distribution unit configured to extract keywords in the data to be encrypted, and distribute weight coefficients to the keywords according to semantic roles and / or industry risk coefficients; a feature value determination unit configured to determine feature values corresponding to each of the keywords according to data types; A scene coefficient and compliance coefficient determination unit is configured to determine a scene coefficient and a compliance coefficient, wherein the scene coefficient is used to dynamically reflect the influence of a business scenario on a sensitivity level, and the compliance coefficient is used to map mandatory requirements of regulations on the sensitivity level. A sensitivity level determination unit is configured to determine a data sensitivity level corresponding to the data to be encrypted based on the weight coefficient, the feature value, the scene coefficient, and the compliance coefficient. An encryption processing unit is configured to perform encryption processing on the data to be encrypted by using an encryption strategy corresponding to the level based on the data sensitivity level.
[0017] In a third aspect, a computer device is provided, which includes a memory, a processor, and computer readable instructions stored in the memory and executable on the processor, and the processor implements the steps of the dynamic hierarchical data encryption method integrating scenarios and compliance as described above when executing the computer readable instructions.
[0018] In a fourth aspect, a readable storage medium is provided, which stores computer readable instructions, and the computer readable instructions implement the steps of the dynamic hierarchical data encryption method integrating scenarios and compliance as described above when executed by a processor.
[0019] The above provides a dynamic hierarchical data encryption method and system fusing a scene and compliance, a method implementation thereof, including: acquiring multi-source data, and pre-processing to obtain to-be-encrypted data; extracting keywords in the to-be-encrypted data, and assigning weight coefficients to the keywords according to semantic roles and / or industry risk coefficients; determining characteristic values corresponding to each keyword according to data types; determining a scene coefficient and a compliance coefficient, wherein the scene coefficient is used for dynamically reflecting an influence of a business scene on a sensitive level, and the compliance coefficient is used for mapping mandatory requirements of regulations on the sensitive level; determining a data sensitive level corresponding to the to-be-encrypted data based on the weight coefficients, the characteristic values, the scene coefficient and the compliance coefficient; and performing encryption processing on the to-be-encrypted data by using an encryption strategy of a corresponding level based on the data sensitive level. In the embodiment of the application, all types of data such as structured data, unstructured data and semi-structured data can be adapted, the scene coefficient is calculated by weighting in four dimensions of time, space, equipment and operation, the compliance coefficient adjustment rule is automatically generated by combining a "clause-sensitive feature" correlation model, the automation and real-time adaptation of compliance requirements are realized, the weight of the keyword is dynamically assigned by combining the semantic role and the industry risk coefficient, the characteristic value is calculated differently according to the data types, the sensitive degree of the data is accurately quantified, the sensitive level is determined based on the sensitive value and the differential encryption strategy is configured, the closed-loop dynamic linkage mechanism of "sensitive value-sensitive level-encryption strategy" is constructed, the safety of high-sensitive data is ensured, the resource waste caused by excessive encryption is avoided, the whole process is automatically realized, the dynamic hierarchical data encryption fusing the scene and the compliance is realized, the human cost and the overall cost of security protection are greatly reduced, and the comprehensive efficiency of data processing and security protection is improved. BRIEF DESCRIPTION OF DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the description of the embodiments of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without any creative labor.
[0021] Figure 1 is a flowchart of a dynamic hierarchical data encryption method fusing a scene and compliance in an embodiment of the present application; Figure 2 is a structural schematic diagram of a dynamic hierarchical data encryption system fusing a scene and compliance in an embodiment of the present application; Figure 3 is a schematic diagram of a computer device in an embodiment of the present application. DETAILED DESCRIPTION
[0022] Clearly, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative labor shall fall within the protection scope of the present application.
[0023] In an embodiment, as shown in Figure 1 A dynamic hierarchical data encryption method combining scene and compliance is provided, comprising the following steps: In step S110, multi-source data is acquired and preprocessed to obtain data to be encrypted. The multi-source data can include structured data (such as database tables, CSV files, Excel files, etc.), unstructured data (such as text files, images, audio, etc.), semi-structured data (such as log files, XML documents, JSON documents, emails, etc.), etc.
[0024] It should be noted that when collecting multi-source data, appropriate collection methods should be adopted according to the different characteristics of structured data, unstructured data and semi-structured data: For example, for structured data, such as database tables, you can use the official drivers provided by the database (such as relational databases like MySQL and Oracle) or third-party database connection tools (such as JDBC and ODBC). By writing code or configuring connection parameters, you can establish a connection to the database and then execute SQL queries (such as SELECT statements) to retrieve the required data. Alternatively, you can use ETL (Extract, Transform, Load) tools like Apache NiFi, Talend, and Informatica, which allow you to configure data source connections and data extraction rules through a graphical interface, quickly extracting data from the database. Simultaneously, you can perform preprocessing operations such as data cleaning and transformation during the extraction process. For unstructured data, such as web text, you can use Python tools like Scrapy (a web scraping framework) and BeautifulSoup (a parsing library) to write programs that analyze the webpage structure and extract text content from HTML pages. For local text files (such as .txt format), you can also use programming languages to read them; for example, in Python, you can use the `open()` function to open the file and read its text content. For text log data generated by applications, log collection tools such as Fluentd and Logstash can be used. These tools can monitor changes in log files in real time, collect newly generated log data, and send it to a designated data storage or processing platform. For semi-structured data, such as XML / JSON documents, Python's ElementTree library (for XML parsing) and json library (for JSON parsing), as well as Java's DOM4J and Jackson libraries, can be used to write code to read the content of XML / JSON documents, parse the tags or key-value pairs, and extract the required data. Alternatively, tools like XSLT (for XML transformation) can be used to convert XML documents into more easily processed formats (such as CSV), or online JSON parsing tools can be used to convert JSON data into other formats for further processing and collection.
[0025] In addition, after obtaining multi-source data, corresponding preprocessing methods can be adopted according to each data type to remove redundant information, unify the data format, and improve the data quality, providing a standardized data foundation for subsequent keyword extraction, eigenvalue calculation, and hierarchical encryption. For example, for structured data, the database table structure / CSV column attributes can be parsed to mark potential sensitive fields such as "user identity information" and "transaction amount", and null values and outliers can be cleaned. For unstructured data, taking text data as an example, meaningless symbols (such as special characters, punctuation marks, and whitespace), stop words (such as common words without sensitive meanings like "of", "is", "and", etc.), are removed, the formats of dates, amounts, etc. are normalized (such as "2025-08-18" unified into the "YYYY-MM-DD" format, and "1 million" converted to "1000000"), words with too short lengths (such as single characters) or no sensitive associations are filtered out, and the core valid words are retained, etc. For image data, sensitive regions (such as identity information) are identified, irrelevant background regions are cropped and removed, and only the regions related to sensitive information are retained; finally, the image data is converted into a standardized pixel matrix format. For semi-structured data, such as log documents, first, the format is parsed, the fields are split according to log delimiters (such as spaces, commas), and core fields such as "timestamp", "device ID", and "operation behavior" are extracted; then, invalid logs (such as entries with incorrect formats or no key information) are filtered; subsequently, the time format is unified (such as converting times in different time zones to UTC time) and the fields are standardized (such as unifying "login" and "登录" in the "operation type" field to "登录操作"); finally, it is converted into a structured data table format.
[0026] In step S120, keywords are extracted from the data to be encrypted, and weight coefficients are assigned to the keywords according to semantic roles and / or industry risk coefficients; Optionally, differentiated keyword extraction techniques are used for different types of data to be encrypted: core words are extracted from structured data, unstructured text data is extracted through word segmentation and keyword ranking algorithms (such as TF-IDF, TextRank), image / audio data is converted into text through recognition (image OCR recognition, audio speech to text) and key information is extracted, semi-structured data is parsed to extract core node words, and finally core keywords related to the sensitivity of the data are selected, such as identity identifiers (name / identity ID), property information (transaction amount / account balance), behavioral data (login IP / operation log), etc. Then, a dual dimension of "semantic role classification" and / or "industry risk adaptation" can be used to assign weight coefficients to keywords. The semantic role dimension requires dividing keywords into different roles such as "operational terminology" and "sensitive events" according to their semantic attributes, and assigning differentiated basic weights to different roles (for example, setting a basic weight value of 1, assigning a basic weight of 1 to keywords of the "operational terminology" category, and assigning a weight of 2 (i.e., twice the basic weight) to keywords of the "sensitive events" category). The industry risk adaptation dimension relies on a pre-set industry sensitive word library, first associating each keyword with the risk coefficient of the corresponding industry (based on the strictness of industry supervision and the quantification of the level of harm of data leakage), and then assigning differentiated weights to the same keyword in different industry scenarios according to the coefficient (for example, the keyword "address" is assigned a basic weight of 1 in the e-commerce industry, and a weight of 3 in other industries due to the association with more privacy compliance requirements). Finally, the final weight coefficient of each keyword is determined by a single dimension or a combination of the two dimensions.
[0027] In step S130, the feature values corresponding to each keyword are determined according to the data type; Optionally, for different types of preprocessed structured, unstructured, and semi-structured data, a differentiated approach adapted to the data characteristics is adopted to determine the feature value corresponding to each keyword. Specifically, for structured data, the feature value can be calculated by combining the preset importance weight of the core field where the keyword is located (e.g., the "user identity information" field has a higher weight than the "delivery address" field) and the proportion of the keyword's occurrence frequency in the corresponding field's valid records (number of keyword occurrences / total number of valid records in the field). The higher the proportion of occurrence frequency and the greater the importance weight of the field, the higher the feature value. For unstructured data: for text data, the feature value is calculated based on the number of times the keyword appears in the valid content, the proportion (number of occurrences / total number of valid words in the text), and the sensitivity of the context. For image data, the feature value is determined by the pixel proportion of the keyword's information in the sensitive area (number of pixels in the sensitive area / total number of pixels in the image), clarity, and completeness of the sensitive information. For audio data, the feature value is determined based on the duration proportion of the audio content corresponding to the keyword in the audio (sensitive audio duration / total audio duration), recognizability, and sensitivity of the semantics. For semi-structured data: logs, XML / JSON documents, etc., feature values are determined by the frequency of keywords in core nodes or key fields and the proportion of related data entries (number of data entries including keywords / total number of data entries); email data combines the frequency of keywords in the subject, body and attachments with the degree of sensitivity association to determine feature values, ensuring that the feature values of keywords in different types of data can objectively reflect their impact weight on the overall sensitivity level of the data.
[0028] Among them, eigenvalues It is used to reflect the intensity and scope of influence of keywords in the data.
[0029] In step S140, a scenario coefficient and a compliance coefficient are determined, wherein the scenario coefficient is used to dynamically reflect the impact of business scenarios on the sensitivity level, and the compliance coefficient is used to map the mandatory requirements of regulations on the sensitivity level. Optionally, the determination of the scenario coefficient needs to comprehensively consider multiple business dimensions such as time, space, equipment, and operation of the business scenario, and combine the risk level corresponding to each business dimension (such as the business peak risk in the time dimension, the cross-border access risk in the space dimension, the risk of untrusted equipment in the equipment dimension, and the high-risk operation risk in the operation dimension). It is obtained by comprehensively quantifying the risk impact weight of each dimension, which can dynamically capture the real-time impact of changes in business scenarios on the data sensitivity level. The determination of the compliance coefficient first collects the latest content of global data security regulations (such as national standards, EU GDPR, US CCPA) and industry standards at a preset frequency, and then uses the "clause-sensitive feature" association model (such as a semantic mapping model based on natural language processing) to transform it into quantifiable coefficient adjustment rules (such as increasing the compliance coefficient by 0.3 for "requires encrypted storage"). Finally, the compliance coefficient of the corresponding sensitive data is determined according to the rule, accurately mapping the mandatory constraints of regulations on the data sensitivity level. This application deeply integrates "scenario factors (time, space, equipment, and operation dimensions)" with "compliance factors (dynamic regulatory requirements)" as the core basis for data classification, enabling real-time dynamic adjustment of classification standards as scenarios change and regulations are updated, fundamentally solving the problems of lagging classification and poor adaptability of traditional solutions.
[0030] In step S150, the data sensitivity level corresponding to the data to be encrypted is determined based on the weight coefficient, feature value, scenario coefficient, and compliance coefficient. Optionally, by first combining the weight coefficient and feature value of each keyword, a basic sensitivity quantification result of the data itself is obtained. Then, this basic result is comprehensively correlated with the scenario coefficient and compliance coefficient to fully superimpose the impact of dynamic risks in the scenario and mandatory requirements of regulations, and finally obtain a data sensitivity score that can comprehensively and accurately reflect the actual sensitivity of the data. Subsequently, based on the preset grading standards (e.g., high level: score ≥ 80, medium level: 40 ≤ score < 80, low level: score < 40, specific thresholds can be adjusted as needed), the data sensitivity level of the data to be encrypted is determined according to the range of the sensitivity score, providing a clear basis for subsequent differentiated encryption processing.
[0031] In step S160, based on the data sensitivity level, the corresponding encryption strategy is used to encrypt the data to be encrypted.
[0032] Optionally, after determining the data sensitivity level of the data to be encrypted, a differential encryption strategy adapted to the level is adopted to perform precise encryption processing on the data. For example, if the data is of a high sensitivity level, an encryption key with a shorter life cycle is selected, and the encryption operation is performed through a hardware password module, while taking into account the protection of data confidentiality and integrity to comprehensively resist security risks; if the data is of a medium sensitivity level, an encryption key with a medium life cycle is adopted, and the encryption is also completed with the help of a hardware password module, focusing on ensuring data confidentiality to meet core security requirements; if the data is of a low sensitivity level, an encryption key with a longer life cycle is used, and the encryption is performed through a software password module, focusing on data integrity protection to achieve a balance between security protection and resource efficiency, ensuring that data of different sensitivity levels can obtain encryption protection adapted to their risk levels, avoiding resource waste caused by over-encryption, and preventing security risks caused by insufficient encryption.
[0033] Exemplarily, taking the example of an e-commerce platform that needs to perform dynamic hierarchical encryption on user order data (including user name, delivery address, payment amount, order log, etc.), first, multi-source data is obtained, such as structured data: order table (including fields such as "user name", "payment amount", "order time", etc.), logistics information table in CSV format; unstructured data: identity information photos uploaded by users (for real-name authentication); semi-structured data: order operation logs in JSON format (including login IP, operation time, etc.). Then, preprocessing is performed on it. For example, for text data (order log): stop words such as "de" and "le" are removed, and "25 / 08 / 2025" is normalized to "2025-08-25"; for image data (identity information photos): sensitive information areas are identified and marked; for structured data: the order table structure is parsed, and "user name" and "payment amount" are marked as potential sensitive fields, and null order records are cleaned (a total of 1000 valid orders). Then, keywords such as "user name", "delivery address", "payment amount", "identity information", and "login IP" are extracted. Among them, "user name" and "identity information": belong to "identity identification", the semantic role is "sensitive event" (associated with personal privacy), the basic weight is multiplied by 2, that is ; "delivery address": in the e-commerce industry scenario, according to the industry adaptation rules, the weight ; "payment amount": belongs to "property information", the semantic role is "sensitive event" (involving financial security), the weight ; "login IP": belongs to "behavior data", the semantic role is "operation term", the weight .
[0034] Subsequently, the eigenvalue corresponding to each keyword is calculated dynamically , for example, "user name" (text data, order log): appears 5 times in a 1000-word log, "Shipping Address" (structured data, order table): Each of the 1000 orders contains a shipping address. "Payment Amount" (structured data, order table): 200 sensitive records with amounts > 5000 yuan. "Identity Information" (Image Data): The area containing identity-sensitive information occupies 15% of the total image pixels. "Login IP" (semi-structured data, operation log): appears 300 times in 500 log entries. .
[0035] Furthermore, scenario coefficients are determined from the dimensions of time, space, device, and operation. For example, the time dimension (30%): weekday 14:00 (low risk), risk value Spatial dimension (20%): Domestic IP (low risk), risk value Device dimension (30%): User's frequently used mobile phone (low risk), risk value Operational Dimension (20%): Normal order placement (low risk), risk value Scene coefficient: , (Involving personal opinions and privacy, such as name and address), therefore compliance level .
[0036] Then the sum =:
[0037] Sensitive score: ; Grading results: This belongs to the medium-level data.
[0038] The encryption strategy is as follows: a key with a lifespan of one quarter is used; encryption is performed by a hardware cryptographic module; and "confidentiality" protection is implemented for order data (to prevent unauthorized access).
[0039] This application provides a dynamic hierarchical data encryption method that integrates scenario and compliance, comprising: acquiring multi-source data and preprocessing it to obtain data to be encrypted; extracting keywords from the data to be encrypted and assigning weight coefficients to the keywords according to semantic roles and / or industry risk coefficients; determining feature values corresponding to each keyword according to data type; determining scenario coefficients and compliance coefficients, wherein the scenario coefficients are used to dynamically reflect the impact of business scenarios on sensitivity levels, and the compliance coefficients are used to map the mandatory requirements of regulations on sensitivity levels; determining the data sensitivity level corresponding to the data to be encrypted based on the weight coefficients, feature values, scenario coefficients, and compliance coefficients; and encrypting the data to be encrypted using a corresponding level of encryption strategy based on the data sensitivity level. In this embodiment, it can adapt to all types of data, including structured, unstructured, and semi-structured data. It calculates scenario coefficients by weighting them across four dimensions: time, space, device, and operation. Combined with a "clause-sensitive feature" association model, it automatically generates compliance coefficient adjustment rules, achieving automated and real-time adaptation to compliance requirements. It can also dynamically assign weights to keywords based on semantic roles and industry risk coefficients, calculate feature values differently according to data types, accurately quantify data sensitivity, determine sensitivity levels based on sensitivity scores, and configure differentiated encryption strategies. It constructs a closed-loop dynamic linkage mechanism of "sensitivity score-sensitivity level-encryption strategy," which not only ensures the security of highly sensitive data but also avoids the waste of resources from excessive encryption. At the same time, the entire process is automated, realizing dynamic hierarchical data encryption that integrates scenarios and compliance, significantly reducing labor costs and overall security protection costs, and improving the overall efficiency of data processing and security protection.
[0040] In one embodiment of this application, assigning weight coefficients to the keywords according to semantic roles and / or industry risk coefficients includes: The keywords are categorized into different roles based on their semantics; Assign corresponding weight values to keywords for different roles; and / or Establish an industry-specific sensitive word database; Each keyword is associated with a corresponding industry risk coefficient based on the aforementioned industry sensitive word library; Based on the industry risk coefficient, the same keywords are assigned different weight values in different industry scenarios.
[0041] Optionally, based on the semantic function and sensitivity relevance of keywords in the data, keywords are semantically categorized into roles such as "operational terminology" and "sensitive events." Then, for each semantic role, a corresponding weight value is assigned to the keywords based on their varying impact on data sensitivity. Different roles can correspond to different weight values. For example, if the role is "operational terminology," the weight... The base value, that is (e.g., "Password" in "The Role of Passwords"), if the role is "Sensitive Event", weight... The base value is multiplied by 2 (e.g., "password" in "password leak"), which is... .
[0042] Furthermore, weight allocation can be based on industry adaptation. First, an industry-sensitive keyword database covering multiple sectors such as finance, e-commerce, and healthcare is constructed. This database pre-collects frequently occurring sensitive keywords from each industry, clearly defining the application scenarios and sensitive association attributes of each keyword within its corresponding industry. Then, based on factors such as the strictness of regulations, the severity of data breaches, and business risk characteristics in each industry, a quantified industry risk coefficient is assigned to each sensitive keyword in the database. The coefficient directly reflects the degree of influence of the keyword on the data sensitivity level within its corresponding industry. Finally, based on this associated industry risk coefficient, differentiated weight values are assigned to the same keyword in different industry scenarios. For example, the weight of "address" in "shipping address". Weight of "address" in "registered address" (Due to the association with personal privacy).
[0043] In one embodiment of this application, determining the feature value corresponding to each keyword according to the data type includes: If the data type is text data, the feature value is determined based on the number of times the keyword appears and the total number of valid characters in the text; If the data type is image data, the feature value is determined based on the number of pixels in the sensitive area and the total number of pixels in the image; If the data type is structured data, the feature value is determined based on the number of record rows containing the keyword and the total number of record rows.
[0044] Optionally, based on the different data types (e.g., unstructured data may include text data and image data, structured data may include database tables, Excel spreadsheets, etc., and semi-structured data may include XML / JSON documents, logs, etc.), a method matching the data characteristics is used to determine the feature values corresponding to each keyword. This data-data-adaptive quantification method enables the feature values to objectively reflect the actual impact of keywords on sensitivity levels in different types of data.
[0045] For example, for text data, the feature value is determined by counting the number of times keywords appear in the text and combining this with the total number of effective words (the actual number of effective words after removing redundant information). The more keywords appear and the higher their proportion in the total number of effective words, the larger the feature value. Specifically, the feature value corresponding to the text data can be calculated using the following formula: = (Number of keyword occurrences / Total number of valid characters in the text) × 100%; For example, if "user name" appears 3 times in a 1000-word text, =31000×100%=0.3%), supports sliding window calculation (window size = 1000 characters).
[0046] For image data, first locate the region containing sensitive information corresponding to the keyword (such as the "identity information" region in an identity information image), count the number of pixels in this sensitive region, and then combine it with the total number of pixels in the image. Determine the feature value based on the proportion of pixels in the sensitive region to the total number of pixels. The higher the proportion, the more prominent the sensitive information, and the larger the feature value. Specifically, the feature value corresponding to the image data can be calculated using the following formula: = (Number of pixels in sensitive areas / Total number of pixels in the image) × 100%; For example, if the identity information area occupies 20% of the total pixels of the image, then the "identity information" =20%, refreshing every 30 frames for dynamic images (such as video frames).
[0047] For structured data (such as database tables, Excel spreadsheets, etc.), count the number of records containing the keyword, compare it with the total number of records, and determine the feature value based on the ratio between the two. The higher the percentage of records containing the keyword, the wider the sensitive association range of the keyword in the dataset, and the larger the feature value. Specifically, the feature value corresponding to the structured data can be calculated using the following formula: = (Number of records containing sensitive keywords / Total number of records) × 100%; For example, in the "Transaction Amount" field, there are 50 sensitive records with amounts greater than 100,000 yuan, out of a total of 1,000 records. It supports incremental calculation (only updates newly added records).
[0048] Furthermore, if the data type is semi-structured data, the feature value (e.g., ...) is determined based on the number of fields containing the keyword and the total number of fields. = (Number of fields including keywords / Total number of fields) × 100%.
[0049] This application dynamically adjusts keyword weights through two dimensions: "semantic role awareness (distinguishing between 'operational terms' and 'sensitive events')" and "industry adaptation (differentiated assignment of sensitive words to different industries)". Meanwhile, based on the characteristics of text, images, and structured data, a differentiated approach is used to calculate feature values. (For example, text can be quantified by frequency of occurrence, and images by the percentage of pixels in sensitive areas), thus achieving precise quantification of data sensitivity and overcoming the shortcomings of traditional methods that produce coarse quantification.
[0050] In one embodiment of this application, the scene coefficient is determined in the following manner: Determine the risk value of each business scenario from multiple business dimensions; According to the preset dimension weights, the risk values are weighted and calculated accordingly to obtain the scenario coefficients.
[0051] Optionally, scene coefficient To dynamically reflect the impact of business scenarios on sensitivity levels, the risk value of each business scenario can be determined from multiple business dimensions (such as time, space, device, and operation). These risk values reflect the degree of danger of the business scenario in different aspects. Next, according to pre-defined weights for each dimension, the risk values determined for each dimension are weighted and calculated accordingly. Through this weighted aggregation method, a scenario coefficient is finally obtained, which dynamically reflects the impact of business scenarios on data sensitivity levels.
[0052] Specifically, taking four business dimensions—time, space, device, and operation—as an example, scenario risks are defined for each dimension (reflecting the degree of danger of the business scenario under that dimension), and the dimension weights are preset as follows: time (30%), space (20%), device (30%), and operation (20%). The business scenario risks are quantified, and the final scenario coefficient is the weighted comprehensive result of the risks of each dimension, as shown below: ; in, This is a time-based risk value (e.g., a quantitative value between 0 and 1, with 0.8 for peak periods and 0.2 for off-peak periods). The risk value is a spatial dimension (e.g., 0.1 for domestic visits and 0.9 for cross-border visits). The risk value is set at the device level (e.g., 0.1 for trusted devices and 0.8 for unfamiliar devices). The risk value is set for the operation dimension (e.g., 0.2 for query operation and 0.9 for transfer operation).
[0053] In one embodiment of this application, the compliance coefficient is determined in the following manner: Collect preset compliant content according to the preset collection frequency; Analyze the newly added or revised content in the preset compliance content; By using a pre-defined association model, the newly added or revised content is transformed into adjustment rules for the compliance coefficient; The compliance coefficient is determined based on the adjustment rules.
[0054] Optionally, compliance coefficient This is used to map the mandatory requirements of regulations regarding sensitivity levels. First, relevant content can be collected from pre-set compliance sources (such as official channels like national standards, EU GDPR, US CCPA, and industry standards) according to a pre-set collection frequency (e.g., once every 24 hours). Then, the collected content is analyzed to filter out newly added or revised clauses (such as GB / T42400-2022's requirements for personal information classification, and GDPR's restrictions on cross-border data transfer). Next, through a pre-set "clause-sensitive feature" association model (such as a semantic matching model based on natural language processing (NLP) to identify core information such as sensitive data types and protection requirements in the clauses), this newly added or revised content is transformed into specific rules that can be directly used to adjust compliance coefficients. Based on these adjustment rules, the compliance coefficient of the corresponding sensitive data is determined. For example, if a clause requires "strengthened protection of user payment data," then it involves sensitive data related to payment information. The value automatically increases to 1.2. When a clause takes effect, it includes text data containing user comments (which may involve personal opinions and privacy). The version has been upgraded from 1.0 to 1.1; new regulations in the financial industry require "separate classification of customer asset data," which will affect the association of keywords such as "account balance." The value is increased to 1.3, thus accurately mapping the mandatory requirements of regulations regarding data sensitivity levels. By periodically capturing compliance texts, parsing clauses, and converting them into adjustment rules, the compliance coefficient is dynamically determined. This enables automatic responses to changes in scenarios and updates in regulations, addressing the issues of weak scenario adaptability and lagging compliance in traditional solutions.
[0055] In one embodiment of this application, determining the data sensitivity level corresponding to the data to be encrypted based on the weight coefficient, feature value, scenario coefficient, and compliance coefficient includes: Based on the weighting coefficient, feature value, scenario coefficient, and compliance coefficient, the data sensitivity score corresponding to the data to be encrypted is determined; The data sensitivity level is determined based on the data sensitivity score.
[0056] Optionally, after obtaining the weight coefficients, feature values, scenario coefficients, and compliance coefficients, they can be substituted into the following formula to calculate the data sensitivity score: ; Where: n is the total number of keywords extracted; The weight of the i-th keyword; Let i be the feature value of the i-th keyword; For scene coefficients; This is the compliance coefficient.
[0057] like This indicates that the data to be encrypted is high-level data and requires high-strength encryption protection; like This indicates that the data to be encrypted is of medium strength and requires medium-level encryption protection. like This indicates that the data to be encrypted is low-level data and requires basic-strength encryption protection.
[0058] In one embodiment of this application, the step of encrypting the data to be encrypted using a corresponding encryption strategy based on the data sensitivity level includes: Based on the data sensitivity level, the corresponding encryption key, encryption module, and protection mechanism are determined; If the data sensitivity level of the data to be encrypted is high, an encryption key with a lifespan of a first preset duration is used to encrypt the data using a hardware cryptographic module, thereby protecting the confidentiality and integrity of the data to be encrypted. If the data sensitivity level of the data to be encrypted is medium, an encryption key with a lifespan of the second preset duration is used to encrypt the data using a hardware cryptographic module, and the confidentiality of the data to be encrypted is protected. If the data sensitivity level of the data to be encrypted is low, an encryption key with a lifespan of a third preset duration is used to encrypt the data using a software cryptographic module, and the integrity of the data to be encrypted is protected. The first preset duration is less than the second preset duration, and the second preset duration is less than the third preset duration.
[0059] Optionally, based on the determined data sensitivity level of the data to be encrypted, a corresponding encryption key, encryption module, and protection mechanism are determined, and then differentiated encryption processing is performed: If the data is highly sensitive, an encryption key with a lifespan of the first preset duration (e.g., 1 month, meaning the key replacement cycle is 1 month) is used, and encryption is completed through a hardware cryptographic module, while simultaneously ensuring the confidentiality and integrity of the data; if the data is moderately sensitive, an encryption key with a lifespan of the second preset duration (e.g., a key with a quarterly lifespan, meaning the key replacement cycle is 1 quarter) is used, also encrypted using a hardware cryptographic module, focusing on ensuring the confidentiality of the data; if the data is low sensitive, an encryption key with a lifespan of the third preset duration (e.g., a key with a 1-year lifespan, meaning the key replacement cycle is 1 year) is used, and encryption is implemented through a software cryptographic module, primarily ensuring the integrity of the data. This application quantifies the data sensitivity level through a sensitivity score calculation formula and directly maps the classification results to differentiated encryption strategies (key lifespan, encryption module, protection mechanism), forming a "classification-encryption" closed loop, ensuring that the encryption strength is accurately matched with the data risk, and avoiding over-encryption or under-encryption.
[0060] In this embodiment, it can adapt to all types of data, including structured, unstructured, and semi-structured data. It calculates scenario coefficients by weighting them across four dimensions: time, space, device, and operation. Combined with a "clause-sensitive feature" association model, it automatically generates compliance coefficient adjustment rules, achieving automated and real-time adaptation to compliance requirements. It can also dynamically assign weights to keywords based on semantic roles and industry risk coefficients, calculate feature values differently according to data types, accurately quantify data sensitivity, determine sensitivity levels based on sensitivity scores, and configure differentiated encryption strategies. It constructs a closed-loop dynamic linkage mechanism of "sensitivity score-sensitivity level-encryption strategy," which not only ensures the security of highly sensitive data but also avoids the waste of resources from excessive encryption. At the same time, the entire process is automated, realizing dynamic hierarchical data encryption that integrates scenarios and compliance, significantly reducing labor costs and overall security protection costs, and improving the overall efficiency of data processing and security protection.
[0061] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0062] In one embodiment, a dynamic hierarchical data encryption system integrating scenario and compliance is provided. This system corresponds one-to-one with the dynamic hierarchical data encryption method integrating scenario and compliance described in the above embodiments. For example... Figure 2 As shown, this dynamic hierarchical data encryption system integrating scenario and compliance includes a data acquisition unit 10, a weight coefficient allocation unit 20, a feature value determination unit 30, a scenario coefficient and compliance coefficient determination unit 40, a sensitivity level determination unit 50, and an encryption processing unit 60. Detailed descriptions of each functional module are as follows: The data acquisition unit 10 is used to acquire multi-source data and perform preprocessing to obtain the data to be encrypted. The weight coefficient allocation unit 20 is used to extract keywords from the data to be encrypted and assign weight coefficients to the keywords according to semantic roles and / or industry risk coefficients. The feature value determination unit 30 is used to determine the feature value corresponding to each keyword according to the data type; The scenario coefficient and compliance coefficient determination unit 40 is used to determine the scenario coefficient and compliance coefficient, wherein the scenario coefficient is used to dynamically reflect the impact of the business scenario on the sensitivity level, and the compliance coefficient is used to map the mandatory requirements of regulations on the sensitivity level. Sensitivity level determination unit 50 is used to determine the data sensitivity level corresponding to the data to be encrypted based on the weight coefficient, feature value, scenario coefficient and compliance coefficient; The encryption processing unit 60 is used to encrypt the data to be encrypted using an encryption strategy corresponding to the data sensitivity level.
[0063] In one embodiment of this application, the weighting coefficient allocation unit 20 is further configured to: The keywords are categorized into different roles based on their semantics; Assign corresponding weight values to keywords for different roles; and / or Establish an industry-specific sensitive word database; Each keyword is associated with a corresponding industry risk coefficient based on the aforementioned industry sensitive word library; Based on the industry risk coefficient, the same keywords are assigned different weight values in different industry scenarios.
[0064] In one embodiment of this application, the feature value determination unit 30 is further configured to: If the data type is text data, the feature value is determined based on the number of times the keyword appears and the total number of valid characters in the text; If the data type is image data, the feature value is determined based on the number of pixels in the sensitive area and the total number of pixels in the image; If the data type is structured data, the feature value is determined based on the number of record rows containing the keyword and the total number of record rows.
[0065] In one embodiment of this application, the scene coefficient is determined in the following manner: Determine the risk value of each business scenario from multiple business dimensions; According to the preset dimension weights, the risk values are weighted and calculated accordingly to obtain the scenario coefficients.
[0066] In one embodiment of this application, the compliance coefficient is determined in the following manner: Collect preset compliant content according to the preset collection frequency; Analyze the newly added or revised content in the preset compliance content; By using a pre-defined association model, the newly added or revised content is transformed into adjustment rules for the compliance coefficient; The compliance coefficient is determined based on the adjustment rules.
[0067] In one embodiment of this application, the sensitivity level determination unit 50 is further configured to: Based on the weighting coefficient, feature value, scenario coefficient, and compliance coefficient, the data sensitivity score corresponding to the data to be encrypted is determined; The data sensitivity level is determined based on the data sensitivity score.
[0068] In one embodiment of this application, the encryption processing unit 60 is further configured to: Based on the data sensitivity level, the corresponding encryption key, encryption module, and protection mechanism are determined; If the data sensitivity level of the data to be encrypted is high, an encryption key with a lifespan of a first preset duration is used to encrypt the data using a hardware cryptographic module, thereby protecting the confidentiality and integrity of the data to be encrypted. If the data sensitivity level of the data to be encrypted is medium, an encryption key with a lifespan of the second preset duration is used to encrypt the data using a hardware cryptographic module, and the confidentiality of the data to be encrypted is protected. If the data sensitivity level of the data to be encrypted is low, an encryption key with a lifespan of a third preset duration is used to encrypt the data using a software cryptographic module, and the integrity of the data to be encrypted is protected. The first preset duration is less than the second preset duration, and the second preset duration is less than the third preset duration.
[0069] In this embodiment, it can adapt to all types of data, including structured, unstructured, and semi-structured data. It calculates scenario coefficients by weighting them across four dimensions: time, space, device, and operation. Combined with a "clause-sensitive feature" association model, it automatically generates compliance coefficient adjustment rules, achieving automated and real-time adaptation to compliance requirements. It can also dynamically assign weights to keywords based on semantic roles and industry risk coefficients, calculate feature values differently according to data types, accurately quantify data sensitivity, determine sensitivity levels based on sensitivity scores, and configure differentiated encryption strategies. It constructs a closed-loop dynamic linkage mechanism of "sensitivity score-sensitivity level-encryption strategy," which not only ensures the security of highly sensitive data but also avoids the waste of resources from excessive encryption. At the same time, the entire process is automated, realizing dynamic hierarchical data encryption that integrates scenarios and compliance, significantly reducing labor costs and overall security protection costs, and improving the overall efficiency of data processing and security protection.
[0070] Specific limitations regarding the dynamic hierarchical data encryption system for integrated scenarios and compliance can be found in the limitations of the dynamic hierarchical data encryption method for integrated scenarios and compliance described above, and will not be repeated here. Each module in the aforementioned dynamic hierarchical data encryption system for integrated scenarios and compliance can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0071] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 3As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes a readable storage medium and internal memory. The readable storage medium stores an operating system, computer-readable instructions, and a database. The internal memory provides an environment for the operation of the operating system and computer-readable instructions in the readable storage medium. The database stores data related to a dynamic hierarchical data encryption method for converged scenarios and compliance. The network interface communicates with external terminals via a network connection. When the computer-readable instructions are executed by the processor, they implement a dynamic hierarchical data encryption method for converged scenarios and compliance. The readable storage medium provided in this embodiment includes both non-volatile readable storage media and volatile readable storage media.
[0072] In one embodiment, a computer device is provided, including a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor. When the processor executes the computer-readable instructions, it implements the aforementioned dynamic hierarchical data encryption method that integrates scenarios and compliance.
[0073] In one embodiment, one or more computer-readable storage media storing computer-readable instructions are provided. The readable storage media provided in this embodiment include non-volatile readable storage media and volatile readable storage media. The readable storage media stores computer-readable instructions, which, when executed by one or more processors, implement the aforementioned dynamic hierarchical data encryption method for integrated scenarios and compliance.
[0074] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware with computer-readable instructions. These computer-readable instructions can be stored in a non-volatile readable storage medium or a volatile readable storage medium. When executed, these computer-readable instructions can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0075] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0076] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A dynamic hierarchical data encryption method that integrates scenario and compliance, characterized in that, The method includes: Acquire multi-source data and preprocess it to obtain the data to be encrypted; Extract keywords from the data to be encrypted, and assign weight coefficients to the keywords according to semantic roles and / or industry risk coefficients; According to the data type, determine the feature value corresponding to each keyword; Determine the scenario coefficient and the compliance coefficient, wherein the scenario coefficient is used to dynamically reflect the impact of business scenarios on the sensitivity level, and the compliance coefficient is used to map the mandatory requirements of regulations on the sensitivity level; Based on the weighting coefficient, feature value, scenario coefficient, and compliance coefficient, the data sensitivity level corresponding to the data to be encrypted is determined; Based on the data sensitivity level, the corresponding encryption strategy is used to encrypt the data to be encrypted.
2. The dynamic hierarchical data encryption method for integrating scenarios and compliance as described in claim 1, characterized in that, The assignment of weight coefficients to the keywords according to semantic roles and / or industry risk coefficients includes: The keywords are categorized into different roles based on their semantics; Assign corresponding weight values to keywords for different roles; and / or Establish an industry-specific sensitive word database; Each keyword is associated with a corresponding industry risk coefficient based on the aforementioned industry sensitive word library; Based on the industry risk coefficient, the same keywords are assigned different weight values in different industry scenarios.
3. The dynamic hierarchical data encryption method for integrating scenarios and compliance as described in claim 1, characterized in that, The step of determining the feature values corresponding to each keyword according to the data type includes: If the data type is text data, the feature value is determined based on the number of times the keyword appears and the total number of valid characters in the text; If the data type is image data, the feature value is determined based on the number of pixels in the sensitive area and the total number of pixels in the image; If the data type is structured data, the feature value is determined based on the number of record rows containing the keyword and the total number of record rows.
4. The dynamic hierarchical data encryption method for integrating scenarios and compliance as described in claim 1, characterized in that, The scenario coefficients are determined in the following manner: Determine the risk value of each business scenario from multiple business dimensions; According to the preset dimension weights, the risk values are weighted and calculated accordingly to obtain the scenario coefficients.
5. The dynamic hierarchical data encryption method for integrating scenarios and compliance as described in claim 1, characterized in that, The compliance coefficient is determined in the following manner: Collect preset compliant content according to the preset collection frequency; Analyze the newly added or revised content in the preset compliance content; By using a pre-defined association model, the newly added or revised content is transformed into adjustment rules for the compliance coefficient; The compliance coefficient is determined based on the adjustment rules.
6. The dynamic hierarchical data encryption method for integrating scenarios and compliance as described in claim 1, characterized in that, The process of determining the data sensitivity level of the data to be encrypted based on the weighting coefficient, feature value, scenario coefficient, and compliance coefficient includes: Based on the weighting coefficient, feature value, scenario coefficient, and compliance coefficient, the data sensitivity score corresponding to the data to be encrypted is determined; The data sensitivity level is determined based on the data sensitivity score.
7. The dynamic hierarchical data encryption method for integrating scenarios and compliance as described in any one of claims 1-6, characterized in that, The step of encrypting the data to be encrypted using a corresponding encryption strategy based on the data sensitivity level includes: Based on the data sensitivity level, the corresponding encryption key, encryption module, and protection mechanism are determined; If the data sensitivity level of the data to be encrypted is high, an encryption key with a lifespan of a first preset duration is used to encrypt the data using a hardware cryptographic module, thereby protecting the confidentiality and integrity of the data to be encrypted. If the data sensitivity level of the data to be encrypted is medium, an encryption key with a lifespan of the second preset duration is used to encrypt the data using a hardware cryptographic module, and the confidentiality of the data to be encrypted is protected. If the data sensitivity level of the data to be encrypted is low, an encryption key with a lifespan of a third preset duration is used to encrypt the data using a software cryptographic module, and the integrity of the data to be encrypted is protected. The first preset duration is less than the second preset duration, and the second preset duration is less than the third preset duration.
8. A dynamic hierarchical data encryption system that integrates scenario and compliance, characterized in that, The system includes: The data to be encrypted acquisition unit is used to acquire multi-source data and perform preprocessing to obtain the data to be encrypted; The weight coefficient allocation unit is used to extract keywords from the data to be encrypted and assign weight coefficients to the keywords according to semantic roles and / or industry risk coefficients. The feature value determination unit is used to determine the feature value corresponding to each keyword according to the data type. The scenario coefficient and compliance coefficient determination unit is used to determine the scenario coefficient and compliance coefficient, wherein the scenario coefficient is used to dynamically reflect the impact of business scenarios on the sensitivity level, and the compliance coefficient is used to map the mandatory requirements of regulations on the sensitivity level; The sensitivity level determination unit is used to determine the data sensitivity level corresponding to the data to be encrypted based on the weight coefficient, feature value, scenario coefficient and compliance coefficient; An encryption processing unit is used to encrypt the data to be encrypted using an encryption strategy corresponding to the data sensitivity level.
9. A computer device comprising a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor, characterized in that, When the processor executes the computer-readable instructions, it implements the steps of the dynamic hierarchical data encryption method for fusion scenarios and compliance as described in any one of claims 1 to 7.
10. A readable storage medium storing computer-readable instructions, characterized in that, When the computer-readable instructions are executed by the processor, they implement the steps of the dynamic hierarchical data encryption method for fusion scenarios and compliance as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Social network privacy data protection method and system
CN120257334A
Method and system for encrypting sensitive data based on large model
CN120455159A
Network security big data processing system and method based on artificial intelligence
CN120528707A
Cited By
Method for generating data security protection requirements based on data classification and grading
CN122046409A
A method for generating data security protection requirements based on data classification grading
CN122046409B
Sensitive data dynamic desensitization strategy and compliance verification method
CN122065344A