Data security sharing and management method and system

By extracting dominant genes and risk values ​​from shared data, calculating recessive and mutation risk values, and formulating hierarchical verification rules, we solve the problems of singleness and implicit identification of risk assessment in data sharing, and achieve dynamic domain management of data and a balance between security and efficiency.

CN120729639AInactive Publication Date: 2025-09-30SICHUAN TECH & BUSINESS UNIV

Patent Information

Application Number
CN202511205317.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2025-09-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies have a single risk assessment dimension in data sharing, fail to effectively integrate the implicit identification risks formed by the combination of illegal sensitive fields, and lack dynamic tracking of risk-enhancing mutations in data version iterations, resulting in the inability to comprehensively and dynamically assess the overall data risk. The data access verification rules have a low match with the sensitivity level, affecting security and efficiency.

Method used

By extracting the dominant genes and risk values ​​of shared data, calculating the recessive risk values ​​and mutation risk values, formulating hierarchical verification rules, and combining the dominant, recessive and mutation fingerprint segments for multi-dimensional risk assessment and verification, dynamic domain management of data can be achieved.

Benefits of technology

It has achieved a comprehensive risk assessment of data from static to dynamic, from single to combined, ensuring that the protection strength is strictly adapted to the data risk, balancing security and sharing efficiency, and achieving full-process security and controllability of data access and traceability of operation tracks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120729639A_ABST
    Figure CN120729639A_ABST
Patent Text Reader

Abstract

The invention discloses a data security sharing and management method and system, and belongs to the field of data security management, and the method comprises the steps: extracting a shared data field name and format characteristics, marking legal sensitive information, forming dominant genes and risk values, calculating recessive and mutation risk values, and matching a belonging sensitive layer domain; analyzing and determining sensitive field weights, generating dominant, implicit and variant fingerprint segments, formulating a hierarchical verification rule, and binding the hierarchical verification rule with a data ID for storage; the user submits a calling application, the encrypted token and the fingerprint segment are obtained through auditing, data are fed back according to a layer domain after verification is passed, and otherwise, abnormity is intercepted and recorded; according to the method, full-dimension dynamic risk assessment of the data is realized, a protection mechanism adaptive to the sensitive hierarchy is constructed, the controllability and traceability of the whole data access process are ensured, and the safety and the sharing efficiency are balanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data security management, and in particular relates to a data security sharing and management method and system. Background Art

[0002] With the growing demand for data sharing, data security and compliance management face challenges. Existing technologies have shortcomings in sensitive information identification, dynamic risk assessment, and access control, which affect the security and efficiency of data sharing. Specific technical issues include: The risk assessment dimension is single-dimensional, focusing only on identifying legally explicit sensitive fields. It fails to effectively integrate the implicit identification risks formed by combinations of non-legally sensitive fields. Furthermore, there is a lack of dynamic tracking of risk-enhancing mutations during data version iterations, resulting in an inability to comprehensively and dynamically assess the overall data risk. The data access verification rules have a poor match with the sensitivity level, which either leads to excessive restrictions affecting sharing efficiency or insufficient protection causing security risks. To this end, we propose a data security sharing and management method and system. Summary of the Invention

[0003] The purpose of the present invention is to provide a data security sharing and management method and system to solve the problems raised in the above background technology.

[0004] To achieve the above objectives, the present invention provides the following technical solution: a data security sharing and management method, comprising the following steps: Step 1: Extract shared data field names, analyze format features, and mark legally sensitive information to form dominant genes and risk values; process non-legally sensitive fields, form field combinations, and calculate data implicit risk values; analyze data variation and calculate mutation risk values, based on which the data domain values ​​are analyzed and matched to the sensitive layer domain to which the data belongs; Step 2: Analyze and sort the weights of the legally sensitive fields in the shared data, extract the core fields, and analyze the explicit fingerprint segments. Combine the implicit risk values ​​and the corresponding field combinations to analyze the implicit fingerprint segments. Generate the mutation fingerprint segments based on the mutation risk value and the timestamp of the most recent risk-enhancing mutation. Develop hierarchical verification rules, bind the fingerprint segments to the data's unique name ID, and store them in an encrypted repository. Step 3: The user submits an application to retrieve shared data. After the permission is reviewed and approved, the encrypted token and the corresponding fingerprint segment are obtained. After submitting the retrieval request, the fingerprint segment and identity permission are verified. If the identity matches, the data is fed back according to the layer domain. Otherwise, it is intercepted and the exception is recorded.

[0005] Preferably, the analysis process of shared data dominant genes and risk values ​​is: Get the shared data uploaded by the user; For each piece of shared data, traverse all fields and extract the text identifiers of each field as field names; analyze the data format of each field and record the format characteristics; Compare the scope of sensitive information specified in the regulations and mark whether each field is legally sensitive information; Bind the extracted field names, data format features, and legally sensitive identifiers to form the dominant gene for each field; The total number of fields marked as legally sensitive is counted and recorded as the explicit risk value XF.

[0006] Preferably, the specific process of processing non-legally sensitive fields is: From all fields of the shared data, the fields with the legally sensitive flag set as “no” in the dominant gene are selected to form a set of non-legally sensitive fields. Based on the industry attributes of shared data, obtain all core business scenarios under the industry attributes; Label each business scenario with a core business tag; extract the business attribute tags of each field in the non-legally sensitive field set and match them with the core business tags of each business scenario to form a field-scenario association table; For each business scenario, filter out all non-legally sensitive fields associated with it from the field-scenario association table and build a dedicated field pool for that scenario. Based on the correlation strength analysis method, select at least three fields with the highest correlation from the scenario-specific field pool as the core correlation fields of the scenario; Verify the combination validity of core related fields and eliminate combinations that have no practical business significance; select the field combination with the highest correlation as the field association combination for this business scenario; All business scenarios and their corresponding three-field association combinations are organized and archived to form an industry business classification library; Extract all associated field combinations from the business classification library to form a set of non-legally sensitive field combinations for shared data; For each field combination in the set of non-legal sensitive field combinations, the values ​​of all fields in the combination are cross-matched to form several theoretical field value combination units; from all field value combination units, the combinations that actually exist in the shared data are screened out and recorded as valid combination units.

[0007] Preferably, the specific process of calculating the implicit risk value of shared data is: Count the number of groups corresponding to each valid combination unit in the shared data, and mark the smallest group number as the minimum group number that can be located; at the same time, obtain the total number of groups corresponding to the field combination in the shared data; and obtain the recognition probability corresponding to each field combination by dividing the minimum group number corresponding to each field combination by the total number of groups corresponding to the field combination; Set multiple recognition probability intervals, preset each recognition probability interval to correspond to a sensitive association value, match the recognition probability corresponding to each field combination with all recognition probability intervals, and output the sensitive association value corresponding to each field combination; The largest sensitive correlation value among all field combinations is screened out and recorded as the implicit risk value YF.

[0008] Preferably, the specific process of calculating the mutation risk value, analyzing the data domain value based on it, and matching the sensitive layer domain to which the data belongs is as follows: Retrieve historical version records of shared data; Compare the current version data with the historical version data at the field level and content level; identify the differences at the data field level and content level and make structured records to generate variant genes; Identify the impact of each data mutation on sensitive risks; If the recognition probability increases or the sensitivity property is enhanced after the mutation, the mutation is marked as a risk-enhancing mutation; otherwise, it is marked as a risk-reducing mutation; The number of mutations of risk-enhancing mutations is counted and recorded as the mutation risk value TF; After normalizing the explicit risk value XF, implicit risk value YF, and mutation risk value TF corresponding to the current shared data, the domain value FYZ is obtained using the formula: FYZ=XF×a1+YF×a2+TF×a3; Among them, a1, a2, and a3 are preset weight coefficients; According to the sensitivity level of shared data, the shared data is divided into public domain, desensitized domain and encrypted domain; and each sensitive layer domain is preset to correspond to a domain value range; By matching the domain value corresponding to the shared data with the domain value interval corresponding to each sensitive layer domain, the corresponding sensitive layer domain is output.

[0009] Preferably, the specific process of analyzing the explicit fingerprint segment is: Establish a dominant gene sensitivity weight evaluation library, in which each field marked as legally sensitive corresponds to a sensitivity weight; For each shared data, obtain the fields marked as legally sensitive in the shared data, substitute them into the dominant gene sensitivity weight evaluation library for matching, and output the corresponding sensitivity weight; Sort the above legally sensitive fields in descending order of sensitivity weight to generate an explicit sensitivity priority sequence; Obtain the explicit risk value corresponding to the shared data, preset the extraction coefficient, and obtain the field extraction number by multiplying the explicit risk value by the corresponding extraction coefficient; According to the value corresponding to the field extraction number, the first corresponding number of fields are extracted from the explicit sensitive priority sequence, and the extracted fields are marked as core explicit sensitive fields; For each core explicit sensitive field, the corresponding data format features are extracted and segmented hashing is performed. Then, the hash results of all fields are concatenated in the order of explicit sensitivity priority to form an explicit fingerprint segment.

[0010] Preferably, the specific process of analyzing latent and variant fingerprint segments is: From all field combinations in the shared data, extract the field combinations corresponding to the implicit risk values ​​and record them as core implicit sensitive combinations; Convert the implicit risk value corresponding to the core implicit sensitive combination into 8-bit binary; Extract the data type identifier of the core implicit sensitive combination, preset numeric type = 1, string type = 0; generate a binary sequence according to the order of the fields in the combination; The two binary sequences are concatenated to obtain the implicit fingerprint segment; Obtain the mutation risk value of the shared data and extract the timestamp of the most recent risk-enhancing mutation in the data; After hashing the timestamp, the first 8 characters are intercepted as the timestamp hash fragment, and the mutation risk value of the shared data is converted into a 4-bit binary number; then it is spliced ​​with the above timestamp hash fragment in sequence to form a mutation fingerprint segment.

[0011] Preferably, the specific process of formulating hierarchical verification rules and binding the fingerprint segment with the data unique name ID and storing it in the encrypted storage library is as follows: Differentiated fingerprint verification rules are formulated for different sensitive domains of shared data, specifically: Public domain: no fingerprint verification is required and can be accessed directly; Desensitized domain: requires double verification of both the explicit fingerprint segment and the variant fingerprint segment; Encrypted domain: requires triple verification of explicit fingerprint segment, implicit fingerprint segment and variant fingerprint segment; For each piece of shared data, first determine the sensitive layer domain to which it belongs; then extract the fingerprint segments required for the corresponding layer domain from the complete sensitive fingerprint of the data, and associate and bind these fingerprint segments with the unique name ID of the data; Finally, the data's unique name ID and the bound fingerprint segment are submitted as a set of associated records to the fingerprint verification center's encrypted repository for storage.

[0012] Preferably, the specific process of step three is: If the user wishes to retrieve shared data, he / she must submit a shared data retrieval application for permission review; Identify the sensitive layer domain where the shared data is located based on the name ID of the shared data to be retrieved; If the user's access permission is approved, a one-time encrypted token corresponding to the shared data will be generated; After the user logs in to the fingerprint distribution interface with the encrypted token, the fingerprint segment corresponding to the shared data is obtained through the end-to-end encrypted channel according to the sensitive layer domain of the data to be retrieved; After obtaining the fingerprint segment corresponding to the shared data, the user submits a retrieval request to the data retrieval interface, including the unique name ID of the shared data to be retrieved, his / her own identity authentication information, and the obtained fingerprint segment; According to the unique name ID of the shared data to be retrieved, the fingerprint segment and sensitive layer domain information bound to the ID are retrieved from the encrypted storage library of the fingerprint verification center; The fingerprint segment submitted by the user is checked for consistency with the fingerprint segment retrieved from the encrypted repository, and the matching degree between the user's identity authentication information and the permissions of the sensitive layer domain of the data to be retrieved is verified; If the fingerprint segments are verified to be consistent and the identity permissions match, the corresponding data content is fed back according to the sensitive layer domain rules. The specific process is as follows: the public domain data returns the original information, the desensitized domain data returns the desensitized processing results, and the encrypted domain data is pushed to the authorized access environment through an encrypted transmission link; If the fingerprint segment verification is inconsistent or the identity authority does not match, the call request will be intercepted immediately, and an exception record containing the request time, user information and failure reason will be generated and stored in the audit log.

[0013] Compared with the prior art, the present invention has the following beneficial effects: (1) This data security sharing and management method and system extracts dominant genes and risk values, calculates recessive risk values, tracks mutation risk values, integrates multi-dimensional risk indicators to form domain values, and achieves comprehensive risk assessment of data from static to dynamic, from single to combined, thus solving the problem of omissions in risk assessment.

[0014] (2) This data security sharing and management method and system formulates differentiated fingerprint verification rules based on the data sensitive layer domain, namely: no verification in the public domain, double verification in the desensitized domain, and triple verification in the encrypted domain. It also ensures that the protection strength is strictly adapted to the data risk and balances security and sharing efficiency through the precise generation and binding of explicit, implicit, and variant fingerprint segments.

[0015] (3) This data security sharing and management method and system, through multi-link control such as permission review, encryption token, fingerprint segment verification, identity matching, combined with exception records and audit logs, realizes the security and control of the entire process from application to access of data, and the operation track is traceable, effectively reducing the risk of unauthorized access. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a flowchart of the present invention. DETAILED DESCRIPTION

[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0018] Embodiment 1; See also Figure 1 ,The present invention provides a data security sharing and management method, including: Step 1: Extract shared data field names, analyze format features, and mark legally sensitive information to form dominant genes and risk values; process non-legally sensitive fields, form field combinations, and calculate data implicit risk values; analyze data variation and calculate mutation risk values. Based on this, analyze data domain values ​​and match the sensitive layer domain to which the data belongs. The specific process is as follows: Get the shared data uploaded by the user; For each shared data, traverse all fields and extract the text identifiers of the fields one by one as field names; field names include: ID number, name, age, etc. Analyze the data format of each field and record the format characteristics; For example: ID card number field: the format feature is an 18-digit mixed string of numbers and letters; name field: the format feature is a Chinese character string; age field: the format feature is a 1-3 digit value; Furthermore, the data format analysis method of the field can adopt the method of sample sampling + rule matching + boundary checking; For example: randomly select 100 entries under the field and count the character type (numbers, Chinese characters, letters, symbols), length range, and structural patterns of the data; Compare sample features with a preset format rule library (including common format templates, such as the 18-digit structure of ID numbers and the 11-digit structure of mobile phone numbers) to preliminarily determine the format type; Check whether there are any abnormal values ​​outside the normal range in the sample (such as "150" or "-5" in the age field). If the proportion of abnormal values ​​is ≤5%, use the majority sample characteristics as the format characteristics and record the type of abnormal values. Based on the above method, determine the format characteristics of each field.

[0019] Compare the sensitive information scope specified in industry regulations (such as ID number, name, biometric information, etc.) and mark whether each field is legally sensitive information; For example: Field 1 ID number: Regulations clearly list it as sensitive personal information → mark "Yes"; Field 2 "Name": can directly identify an individual and is considered sensitive information → mark "Yes"; Field 3 "Age": This field cannot directly identify an individual when it exists alone and is not classified as sensitive by regulations → mark it "No"; Bind the extracted field name, data format characteristics, and legal sensitivity flag to form the dominant gene of each field. For example, the dominant gene of ID card number is: {field name: ID card number, data format: 18-digit mixed string of numbers / alphabets, legal sensitivity flag: yes}; Count the total number of fields marked as legally sensitive and record it as the explicit risk value XF; From all fields of the shared data, the fields with the legally sensitive flag set as “no” in the dominant gene are selected to form a set of non-legally sensitive fields. Based on the industry attributes of the shared data, all core business scenarios under the industry attributes are obtained; (e.g., the medical industry includes "outpatient diagnosis and treatment", "inpatient care", and "chronic disease management"); Label each business scenario with core business tags; for example, the core tags for "outpatient diagnosis and treatment" are "diagnosis basis + prescription issuance + efficacy tracking"; Extract the business attribute labels of each field in the non-legally sensitive field set and match them with the core business labels of each business scenario to form a field-scenario association table; For each business scenario, filter out all non-legally sensitive fields associated with it from the field-scenario association table and build a dedicated field pool for that scenario. Based on the correlation strength analysis method (calculating the correlation between fields through industry expert ratings or historical data interaction frequency), select at least three fields with the highest correlation from the scenario-specific field pool as the core correlation fields for the scenario; Based on the core objectives of the business scenario and the data interaction logic (e.g., medical scenarios center around "disease diagnosis and treatment," and financial scenarios center around "risk assessment"), the validity of the core associated fields is verified, and combinations without practical business significance are eliminated (if the fields in the combination are irrelevant to the core objectives or there is no business logic association between the fields, they are deemed invalid). The three-field combination with the highest degree of association is selected as the field association combination for the business scenario; All business scenarios and their corresponding three-field association combinations are organized and archived to form an industry business classification library; Extract all three-field combinations from the business classification library to form a set of non-legally sensitive field combinations for shared data. (That is, strip away the business scenario information associated with each combination, retain only the field combination itself, and integrate these pure field combinations into an independent set, which is the set of non-legally sensitive field combinations for shared data.) For each field combination in the set of non-legally sensitive field combinations, cross-match the values ​​of all fields in the combination to form several theoretical field value combination units. From all field value combination units, select the combinations that actually exist in the shared data and record them as valid combination units. Count the number of groups corresponding to each valid combination unit in the shared data, and mark the smallest group number as the minimum group number that can be located; at the same time, obtain the total number of groups corresponding to the field combination in the shared data; and obtain the recognition probability corresponding to each field combination by dividing the minimum group number corresponding to each field combination by the total number of groups corresponding to the field combination; Set multiple recognition probability intervals, each preset to correspond to a sensitive association value. Match the recognition probability corresponding to each field combination with all recognition probability intervals, and output the sensitive association value corresponding to each field combination. The larger the recognition probability, the larger the corresponding sensitive association value. The larger the sensitive association value, the higher the risk of the field combination being able to locate a specific individual or a small group through cross-matching, that is, the stronger the implicit recognition ability of the combination; Filter out the largest sensitive correlation value among all field combinations and record it as the implicit risk value YF; Retrieve historical version records of shared data, including original data and various intermediate modified versions; Compare the current version data with the historical version data at the field level (such as field additions, deletions, name changes, and adjustments to legally sensitive identifiers) and content level (such as changes in field value rules, formats, and sensitivity levels). Identify differences at the data field and content levels and create structured records to generate variant genes. Variant genes are standardized records of specific changes that occur to the data during version iterations. The contents of the mutated gene record include: the mutated field (clearly identifying the field that has changed), the mutation type (such as field addition, deletion, content masking, etc.), the characteristics before and after the mutation (recording the specific characteristics before and after the change), and the mutation time (marking the version and corresponding time when the modification occurred); For each data mutation (i.e., data modification), the sensitive attribute mutation impact analysis method is used to identify the direction in which the mutation affects sensitive risks. If the recognition probability increases or the sensitive attribute is enhanced (risk increases) after the mutation, the mutation is marked as a risk-enhancing mutation; otherwise, it is marked as a risk-reducing mutation; The number of mutations of risk-enhancing mutations is counted and recorded as the mutation risk value TF; After normalizing the explicit risk value XF, implicit risk value YF, and mutation risk value TF corresponding to the current shared data, the domain value FYZ is obtained using the formula: FYZ=XF×a1+YF×a2+TF×a3; Among them, a1, a2, and a3 are preset weight coefficients; According to the sensitivity level of shared data, the shared data is divided into public domain, desensitized domain and encrypted domain; and each sensitive layer domain is preset to correspond to a domain value range; By matching the domain value corresponding to the shared data with the domain value interval corresponding to each sensitive layer domain, the corresponding sensitive layer domain is output; Furthermore, the public domain refers to a collection of low-risk data that can be directly shared with all parties; Desensitized domain: refers to a collection of medium-risk data that needs to be desensitized (such as obfuscated or de-identified) before sharing; Encrypted domain: refers to a collection of high-risk data that must be protected by encryption technology and accessed only by authorized parties.

[0020] It should be noted that It not only clearly marks legally sensitive fields such as ID numbers (explicit risks) through regulatory benchmarking, but also mines implicit identification risks (such as cross-locating individuals using multiple fields) through combined analysis of non-legal sensitive fields, achieving "explicit + implicit" dual coverage of sensitive information and avoiding risk omissions caused by single judgments; By recording the field and content changes in the version iteration of the mutant gene record data, the number of risk-enhancing mutations (mutation risk) is counted. By combining explicit and implicit risk values ​​to calculate the domain value, the vague risk perception is converted into a quantifiable indicator, improving the accuracy of risk assessment. Based on the domain value, data is divided into public domain, desensitized domain, and encrypted domain. Different domains correspond to different access rights and processing methods, realizing refined classification management of data. At the same time, it provides a clear basis for differentiated fingerprint verification rules in subsequent steps (such as triple verification required for encrypted domain), ensuring that the protection strength matches the data risk level, taking into account both security and sharing efficiency.

[0021] Step 2: Analyze and sort the weights of the legally sensitive fields in the shared data, extract the core fields, and analyze the explicit fingerprint segments. Combine the implicit risk value and the corresponding field combination to analyze the implicit fingerprint segments. Generate the mutation fingerprint segments based on the mutation risk value and the timestamp of the most recent risk-enhancing mutation. Formulate hierarchical verification rules, bind the fingerprint segments to the data's unique name ID, and store them in an encrypted repository. The specific process is as follows: Establish a dominant gene sensitivity weight evaluation library. Each field marked as legally sensitive in the library corresponds to a sensitivity weight. The sensitivity weight can be defined according to the industry's protection level for the field. For each shared data, obtain the fields marked as legally sensitive in the shared data, substitute them into the dominant gene sensitivity weight evaluation library for matching, and output the corresponding sensitivity weight; Sort the above legally sensitive fields in descending order of sensitivity weight to generate an explicit sensitivity priority sequence; Obtain the explicit risk value corresponding to the shared data, preset the extraction coefficient, and obtain the field extraction number by multiplying the explicit risk value by the corresponding extraction coefficient; if the explicit risk value is an odd number, round it up; According to the value corresponding to the field extraction number, the first corresponding number of fields are extracted from the explicit sensitive priority sequence, and the extracted fields are marked as core explicit sensitive fields; For each core explicit sensitive field, the corresponding data format features are extracted and hashed in sections. The hash results of all fields are then concatenated in order of explicit sensitivity priority to form an explicit fingerprint segment. From all field combinations in the shared data, extract the field combinations corresponding to the implicit risk values ​​and record them as core implicit sensitive combinations; Convert the implicit risk value corresponding to the core implicit sensitive combination into 8-bit binary; Extract the data type identifier of the core implicit sensitive combination, preset numeric type = 1, string type = 0; generate a binary sequence according to the order of the fields in the combination; The two binary sequences are concatenated to obtain the implicit fingerprint segment (i.e., an 11-bit binary sequence). Obtain the mutation risk value of the shared data and extract the timestamp of the most recent risk-enhancing mutation in the data; After hashing the timestamp, extract the first 8 characters as the timestamp hash fragment, and convert the mutation risk value of the shared data into a 4-bit binary number. This is then sequentially concatenated with the timestamp hash fragment to form a mutation fingerprint segment (i.e., a 12-bit sequence). Differentiated fingerprint verification rules are formulated for different sensitive domains of shared data, specifically: Public domain: no fingerprint verification is required and can be accessed directly; Desensitized domain: requires double verification of both the explicit fingerprint segment and the variant fingerprint segment; Encrypted domain: requires triple verification of explicit fingerprint segment, implicit fingerprint segment and variant fingerprint segment; For each piece of shared data, first determine the sensitive layer domain to which it belongs; then extract the fingerprint segments required for the corresponding layer domain from the complete sensitive fingerprint of the data, and associate and bind these fingerprint segments with the unique name ID of the data; Finally, the data's unique name ID and the bound fingerprint segment are submitted as a set of associated records to the fingerprint verification center's encrypted repository for storage.

[0022] It should be noted that the explicit fingerprint segment is generated based on the weight sorting and hashing of legally sensitive fields, the implicit fingerprint segment is constructed by combining the implicit risk value and the field type identifier, and the mutation fingerprint segment associates the mutation risk value with the timestamp of the most recent risk-enhancing mutation. These three types of fingerprints map data sensitive features from different dimensions to form a complete sensitive fingerprint chain. Differentiated rules for "no verification," "double verification," and "triple verification" are formulated for public, desensitized, and encrypted domains, ensuring that verification intensity strictly matches the data risk level, avoiding the impact of excessive verification on efficiency while ensuring access security for highly sensitive data. Binding the fingerprint segment to the data's unique name ID and storing it in encrypted form not only achieves accurate association between fingerprints and data, but also prevents fingerprint information leakage through encrypted storage library protection, laying the foundation for verification accuracy in subsequent retrieval links.

[0023] Step 3: The user submits an application to retrieve shared data. After the permission is approved, the encrypted token and the corresponding fingerprint segment are obtained. After submitting the retrieval request, the fingerprint segment and identity permission are verified. If the identity matches, the data is fed back according to the layer domain. Otherwise, it is intercepted and the exception is recorded. The specific process is as follows: When users access shared data, they must submit a shared data access application for permission review (manual authorization or other methods may be used). The access application includes: the unique name ID of the shared data to be accessed, access qualification certificate, and data usage scenario description; Identify the sensitive layer domain where the shared data is located based on the name ID of the shared data to be retrieved; If the user's access permission is approved, a one-time encrypted token corresponding to the shared data will be generated; After the user logs in to the fingerprint distribution interface with the encrypted token, the fingerprint segment corresponding to the shared data is obtained through the end-to-end encrypted channel according to the sensitive layer domain of the data to be retrieved; After obtaining the fingerprint segment corresponding to the shared data, the user submits a retrieval request to the data retrieval interface, including the unique name ID of the shared data to be retrieved, his / her own identity authentication information, and the obtained fingerprint segment; According to the unique name ID of the shared data to be retrieved, the fingerprint segment and sensitive layer domain information bound to the ID are retrieved from the encrypted storage library of the fingerprint verification center; The fingerprint segment submitted by the user is checked for consistency with the fingerprint segment retrieved from the encrypted repository, and the matching degree between the user's identity authentication information and the permissions of the sensitive layer domain of the data to be retrieved is verified; If the fingerprint segments are verified to be consistent and the identity permissions match, the corresponding data content is fed back according to the sensitive layer domain rules. The specific process is as follows: the public domain data returns the original information, the desensitized domain data returns the desensitized processing results, and the encrypted domain data is pushed to the authorized access environment through an encrypted transmission link; If the fingerprint segment verification is inconsistent or the identity authority does not match, the call request will be intercepted immediately, and an exception record containing the request time, user information and failure reason will be generated and stored in the audit log.

[0024] It should be noted that the access application must submit a unique name ID, qualification certificate and usage scenario description, and complete the permission review through manual authorization and other methods to ensure that the user has legal access qualifications and control access rights from the source; After users obtain the corresponding fingerprint segment with an encrypted token, they must simultaneously pass fingerprint segment consistency verification and identity authority matching verification. This dual verification mechanism significantly reduces the risk of unauthorized access, especially the triple fingerprint verification for encrypted domain data, strengthening the protection of highly sensitive data. Data is fed back according to different rules for public, desensitized, and encrypted domains, balancing security and efficiency. Requests that fail verification are immediately intercepted and exception records containing the time, user, and reason are generated, making the entire operation auditable and traceable, providing a basis for data security incident investigation. Encrypted channels are used from token generation to fingerprint segment acquisition and data push to avoid information leakage during transmission and form an end-to-end secure transmission link.

[0025] A data security sharing and management system, comprising: Data sensitive layer domain division module: Extract shared data field names, analyze format features and mark legally sensitive information to form dominant genes and risk values; process non-legally sensitive fields, form field combinations, and calculate data implicit risk values; analyze data variation, calculate mutation risk values, analyze data domain values ​​based on this, and match the data to the sensitive layer domain; Sensitive fingerprint segment generation and storage module: Analyzes and sorts the weights of legally sensitive fields in shared data, extracts core fields, and analyzes explicit fingerprint segments. Combines implicit risk values ​​with corresponding field combinations to analyze implicit fingerprint segments. Generates mutation fingerprint segments based on mutation risk values ​​and the timestamp of the most recent risk-enhancing mutation. Formulates hierarchical verification rules, binds fingerprint segments to the data's unique name ID, and stores them in an encrypted repository. Shared data retrieval and verification module: The user submits an application for retrieval of shared data, and obtains an encrypted token and corresponding fingerprint segment after the permission is reviewed and approved; submits the retrieval request, verifies the fingerprint segment and identity permission, and if the identity matches, feedback is given according to the layer domain; otherwise, it is intercepted and the exception is recorded.

[0026] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A data security sharing and management method, characterized by: Including the following steps: Step 1: Extract shared data field names, analyze format features, and mark legally sensitive information to form dominant genes and risk values; process non-legally sensitive fields, form field combinations, and calculate data implicit risk values; analyze data variation and calculate mutation risk values, based on which the data domain values ​​are analyzed and matched to the sensitive layer domain to which the data belongs; Step 2: Analyze and sort the weights of the legally sensitive fields in the shared data, extract the core fields, and analyze the explicit fingerprint segments. Combine the implicit risk values ​​and the corresponding field combinations to analyze the implicit fingerprint segments. Generate the mutation fingerprint segments based on the mutation risk value and the timestamp of the most recent risk-enhancing mutation. Develop hierarchical verification rules, bind the fingerprint segments to the data's unique name ID, and store them in an encrypted repository. Step 3: The user submits an application to retrieve shared data. After the permission is reviewed and approved, the encrypted token and the corresponding fingerprint segment are obtained. After submitting the retrieval request, the fingerprint segment and identity permission are verified. If the identity matches, the data is fed back according to the layer domain. Otherwise, it is intercepted and the exception is recorded.

2. A data security sharing and management method according to claim 1, characterized in that: In step 1, the analysis process of shared data dominant genes and risk values ​​is as follows: Get the shared data uploaded by the user; For each piece of shared data, traverse all fields and extract the text identifiers of each field as field names; analyze the data format of each field and record the format characteristics; Compare the scope of sensitive information specified in the regulations and mark whether each field is legally sensitive information; Bind the extracted field names, data format features, and legally sensitive identifiers to form the dominant gene for each field; The total number of fields marked as legally sensitive is counted and recorded as the explicit risk value XF.

3. The data security sharing and management method according to claim 1, characterized in that: In step 1, the specific process of handling non-legally sensitive fields is as follows: From all fields of the shared data, the fields with the legally sensitive flag set as “no” in the dominant gene are selected to form a set of non-legally sensitive fields. Based on the industry attributes of shared data, obtain all core business scenarios under the industry attributes; Label each business scenario with core business tags; Extract the business attribute labels of each field in the non-legally sensitive field set and match them with the core business labels of each business scenario to form a field-scenario association table; For each business scenario, filter out all non-legally sensitive fields associated with it from the field-scenario association table and build a dedicated field pool for that scenario. Based on the correlation strength analysis method, select at least three fields with the highest correlation from the scenario-specific field pool as the core correlation fields of the scenario; Verify the combination validity of core related fields and eliminate combinations that have no practical business significance; select the field combination with the highest correlation as the field association combination for this business scenario; All business scenarios and their corresponding three-field association combinations are organized and archived to form an industry business classification library; Extract all associated field combinations from the business classification library to form a set of non-legally sensitive field combinations for shared data; For each field combination in the set of non-legal sensitive field combinations, the values ​​of all fields in the combination are cross-matched to form several theoretical field value combination units; from all field value combination units, the combinations that actually exist in the shared data are screened out and recorded as valid combination units.

4. The data security sharing and management method according to claim 1, characterized in that: In step 1, the specific process of calculating the implicit risk value of shared data is as follows: Count the number of groups corresponding to each valid combination unit in the shared data, and mark the smallest group number as the minimum group number that can be located; at the same time, obtain the total number of groups corresponding to the field combination in the shared data; and obtain the recognition probability corresponding to each field combination by dividing the minimum group number corresponding to each field combination by the total number of groups corresponding to the field combination; Set multiple recognition probability intervals, preset each recognition probability interval to correspond to a sensitive association value, match the recognition probability corresponding to each field combination with all recognition probability intervals, and output the sensitive association value corresponding to each field combination; The largest sensitive correlation value among all field combinations is screened out and recorded as the implicit risk value YF.

5. The data security sharing and management method according to claim 1, characterized in that: In step 1, the mutation risk value is calculated, the data domain value is analyzed based on it, and the specific process of matching the sensitive layer domain to which the data belongs is as follows: Retrieve historical version records of shared data; Compare the current version data with the historical version data at the field level and content level; identify the differences at the data field level and content level and make structured records to generate variant genes; Identify the impact of each data mutation on sensitive risks; If the recognition probability increases or the sensitivity property is enhanced after the mutation, the mutation is marked as a risk-enhancing mutation; otherwise, it is marked as a risk-reducing mutation; The number of mutations of risk-enhancing mutations is counted and recorded as the mutation risk value TF; After normalizing the explicit risk value XF, implicit risk value YF, and mutation risk value TF corresponding to the current shared data, the domain value FYZ is obtained using the formula: FYZ=XF×a1+YF×a2+TF×a3; Among them, a1, a2, and a3 are preset weight coefficients; According to the sensitivity level of shared data, the shared data is divided into public domain, desensitized domain and encrypted domain; and each sensitive layer domain is preset to correspond to a domain value range; By matching the domain value corresponding to the shared data with the domain value interval corresponding to each sensitive layer domain, the corresponding sensitive layer domain is output.

6. A data security sharing and management method according to claim 1, characterized in that: In step 2, the specific process of analyzing the explicit fingerprint segment is as follows: Establish a dominant gene sensitivity weight evaluation library, in which each field marked as legally sensitive corresponds to a sensitivity weight; For each shared data, obtain the fields marked as legally sensitive in the shared data, substitute them into the dominant gene sensitivity weight evaluation library for matching, and output the corresponding sensitivity weight; Sort the above legally sensitive fields in descending order of sensitivity weight to generate an explicit sensitivity priority sequence; Obtain the explicit risk value corresponding to the shared data, preset the extraction coefficient, and obtain the field extraction number by multiplying the explicit risk value by the corresponding extraction coefficient; According to the value corresponding to the field extraction number, the first corresponding number of fields are extracted from the explicit sensitive priority sequence, and the extracted fields are marked as core explicit sensitive fields; For each core explicit sensitive field, the corresponding data format features are extracted and segmented hashing is performed. Then, the hash results of all fields are concatenated in the order of explicit sensitivity priority to form an explicit fingerprint segment.

7. The data security sharing and management method according to claim 1, characterized in that: In step 2, the specific process of analyzing latent and variant fingerprint segments is as follows: From all field combinations in the shared data, extract the field combinations corresponding to the implicit risk values ​​and record them as core implicit sensitive combinations; Convert the implicit risk value corresponding to the core implicit sensitive combination into 8-bit binary; Extract the data type identifier of the core implicit sensitive combination, preset numeric type = 1, string type = 0; generate a binary sequence according to the order of the fields in the combination; The two binary sequences are concatenated to obtain the implicit fingerprint segment; Obtain the mutation risk value of the shared data and extract the timestamp of the most recent risk-enhancing mutation in the data; After hashing the timestamp, the first 8 characters are intercepted as the timestamp hash fragment, and the mutation risk value of the shared data is converted into a 4-bit binary number; then it is spliced ​​with the above timestamp hash fragment in sequence to form a mutation fingerprint segment.

8. The data security sharing and management method according to claim 1, characterized in that: In step 2, the specific process of formulating hierarchical verification rules, binding the fingerprint segment with the data's unique name ID, and storing it in the encrypted repository is as follows: Differentiated fingerprint verification rules are formulated for different sensitive domains of shared data, specifically: Public domain: no fingerprint verification is required and can be accessed directly; Desensitized domain: requires double verification of both the explicit fingerprint segment and the variant fingerprint segment; Encrypted domain: requires triple verification of explicit fingerprint segment, implicit fingerprint segment and variant fingerprint segment; For each piece of shared data, first determine the sensitive layer domain to which it belongs; then extract the fingerprint segments required for the corresponding layer domain from the complete sensitive fingerprint of the data, and associate and bind these fingerprint segments with the unique name ID of the data; Finally, the data's unique name ID and the bound fingerprint segment are submitted as a set of associated records to the fingerprint verification center's encrypted repository for storage.

9. The data security sharing and management method according to claim 1, characterized in that: The specific process of step three is: If the user wishes to retrieve shared data, he / she must submit a shared data retrieval application for permission review; Identify the sensitive layer domain where the shared data is located based on the name ID of the shared data to be retrieved; If the user's access permission is approved, a one-time encrypted token corresponding to the shared data will be generated; After the user logs in to the fingerprint distribution interface with the encrypted token, the fingerprint segment corresponding to the shared data is obtained through the end-to-end encrypted channel according to the sensitive layer domain of the data to be retrieved; After obtaining the fingerprint segment corresponding to the shared data, the user submits a retrieval request to the data retrieval interface, including the unique name ID of the shared data to be retrieved, his / her own identity authentication information, and the obtained fingerprint segment; According to the unique name ID of the shared data to be retrieved, the fingerprint segment and sensitive layer domain information bound to the ID are retrieved from the encrypted storage library of the fingerprint verification center; The fingerprint segment submitted by the user is checked for consistency with the fingerprint segment retrieved from the encrypted repository, and the matching degree between the user's identity authentication information and the permissions of the sensitive layer domain of the data to be retrieved is verified; If the fingerprint segments are verified to be consistent and the identity permissions match, the corresponding data content is fed back according to the sensitive layer domain rules. The specific process is as follows: the public domain data returns the original information, the desensitized domain data returns the desensitized processing results, and the encrypted domain data is pushed to the authorized access environment through an encrypted transmission link; If the fingerprint segment verification is inconsistent or the identity authority does not match, the call request will be intercepted immediately, and an exception record containing the request time, user information and failure reason will be generated and stored in the audit log.

10. A data security sharing and management system, applied to a data security sharing and management method according to any one of claims 1 to 9, characterized in that: Data sensitive layer domain division module: extracts shared data field names, analyzes format features, and marks legally sensitive information to form dominant genes and risk values; Process non-legally sensitive fields, form field combinations, and calculate data implicit risk values; Analyze data variation, calculate mutation risk value, analyze data domain value based on this, and match the sensitive layer domain to which the data belongs; Sensitive fingerprint segment generation and storage module: Analyzes and sorts the weights of legally sensitive fields in shared data, extracts core fields, and analyzes explicit fingerprint segments. Combines implicit risk values ​​with corresponding field combinations to analyze implicit fingerprint segments. Generates mutation fingerprint segments based on mutation risk values ​​and the timestamp of the most recent risk-enhancing mutation. Formulates hierarchical verification rules, binds fingerprint segments to the data's unique name ID, and stores them in an encrypted repository. Shared data retrieval and verification module: The user submits an application for retrieval of shared data, and obtains an encrypted token and corresponding fingerprint segment after the permission is reviewed and approved; submits the retrieval request, verifies the fingerprint segment and identity permission, and if the identity matches, feedback is given according to the layer domain; otherwise, it is intercepted and the exception is recorded.

Citation Information

Patent Citations

  • Scientific and technological financial platform data sharing method

    CN119720281A

  • Online document processing method and system based on document processing model

    CN120337879A

  • Trusted sharing method and system for multi-party industrial data

    CN120415783A

Cited By

  • File information processing method and system based on digital security

    CN121118113A