Self-adaptive government affair shared data conformal encryption method and system

By conducting structured analysis and sensitivity assessment of government data, combined with adaptive encryption algorithms, we can achieve accurate sensitive identification and differentiated encryption of government data, solve the problems of insufficient encryption strength and resource waste in existing technologies, and improve the security and efficiency of government data sharing.

CN120705910AActive Publication Date: 2025-09-26HUNAN FENGHUI YINJIA SCI & TECH CO LTD

Patent Information

Application Number
CN202510863894.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-09-26
Estimated Expiration
2045-06-26

AI Technical Summary

Technical Problem

Existing government data encryption methods fail to effectively distinguish the sensitivity of data, resulting in a mismatch between encryption strength and resource consumption. They are unable to meet the needs of high-frequency sharing and personalized desensitization, and lack sensitivity analysis and dynamic encryption protection at the granular level of data content.

Method used

By performing structured analysis on government data, using an improved entity recognition model to identify potentially sensitive content, constructing a sensitivity assessment matrix and performing sensitivity level clustering, combining the adaptive neighborhood radius to dynamically adjust the clustering perception scale, and using an adaptive encryption algorithm to perform differentiated encryption processing on data entities of different sensitivity levels.

Benefits of technology

It achieves accurate sensitive identification and encryption of government shared data, improves the flexibility and security of encryption processing, ensures data availability and legal compliance, and supports data sharing in multiple scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705910A_ABST
    Figure CN120705910A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of government affair shared data encryption, and discloses a self-adaptive government affair shared data conformal encryption method and system, and the method comprises the steps: recognizing a data entity containing potential sensitive content in original government affair data after structural analysis through employing an improved entity recognition model, and generating a sensitive tag; carrying out sensitivity evaluation on the data entities, and carrying out sensitive grading processing on the data entities by utilizing an improved clustering algorithm; and carrying out encryption processing and data replacement on the data entity after sensitive grading by adopting a self-adaptive encryption algorithm to obtain government affair shared data. The invention provides a government affair data protection method fusing entity identification, sensitivity evaluation and an adaptive encryption strategy, and the method achieves the conformal encryption protection of a high-sensitivity entity and the desensitization replacement of a low-sensitivity entity through the construction of a multi-dimensional sensitivity matrix, the fine grading of potential sensitive fields, and the matching of an adaptive encryption mode. And the privacy security and the structure availability of government affair data sharing are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of encryption of government data sharing, and in particular to an adaptive conformal encryption method and system for government data sharing. Background Art

[0002] With the continuous advancement of government data resource sharing, cross-departmental and cross-level data interactions are becoming increasingly frequent, and data types and structures are becoming complex, characterized by diversity, unstructuredness, and high sensitivity. Government data contains a large amount of sensitive personal information, including personal identities, addresses, contact information, and account numbers. Without effective protection during data sharing and disclosure, it can easily lead to privacy leaks, illegal tampering, and even data abuse. Therefore, how to accurately identify and differentially encrypt sensitive data while ensuring data availability and semantic integrity has become a research focus in the field of government information security.

[0003] Existing technologies, such as patent CN108924104B, propose an e-government encryption and decryption method. By introducing custom system authentication codes, user encryption codes, and user decryption codes, this method establishes a "double authentication + single sign-on" access control mechanism, achieving preliminary encryption protection for the e-government platform at the user authentication and system access levels. This method has some practical value in terms of access-side security, but it primarily focuses on the identity verification and encryption authorization processes, lacking sensitivity analysis and dynamic encryption protection at the granular level of the data itself.

[0004] At the same time, the current mainstream government data encryption schemes generally have the following problems: they do not distinguish the sensitivity of the data, resulting in a mismatch between encryption strength and resource consumption; sensitive entities such as ID numbers, addresses, and contact information are difficult to automatically identify, and fail to achieve refined identification at the content structure level; they cannot meet the dual needs of high-frequency sharing and personalized desensitization, and cannot adaptively adjust encryption strategies in different sharing scenarios, which can easily lead to information leakage or data redundancy.

[0005] In summary, traditional encryption methods often use a unified algorithm to process all sensitive data, ignoring the diverse needs of entities with different sensitivity levels, which can easily lead to insufficient encryption strength or waste of computing resources. Although some methods introduce sensitivity assessment, their assessment methods are too static and fail to dynamically adjust the sensitivity level by combining contextual semantic features and word structure attributes, making them difficult to adapt to the complex multi-source government text environment. Therefore, there is an urgent need for an adaptive government data encryption method that features multi-dimensional sensitivity modeling, joint semantic structure perception, and support for differentiated conformal encryption. This method can improve the accuracy of sensitive information identification and the flexibility of encryption processing, while ensuring data availability and meeting legal compliance and privacy protection requirements. Summary of the Invention

[0006] In view of this, the present invention provides an adaptive conformal-preserving encryption method for government data sharing. The method performs structured analysis on original government data, combines the semantic association strength between words and word attributes, identifies data entities and sensitive labels containing potentially sensitive content, constructs a sensitivity assessment matrix, and performs sensitivity level clustering. The sensitivity assessment results are introduced into the clustering process to prevent the misjudgment of high- and low-sensitivity data entities from being aggregated. An adaptive neighborhood radius is introduced to dynamically adjust the clustering perception scale based on the sensitivity score, so that highly sensitive data entities have a more fine-grained separation capability, improve the accuracy of cluster boundaries, and effectively divide different sensitivity levels. An adaptive encryption algorithm is then used to select conformal-preserving or homomorphic encryption with different encryption strengths for data entities of different sensitivity levels to avoid insufficient encryption strength or waste of computing resources. During the encryption process, high- and medium-sensitivity data entities are conformally encrypted to avoid using data structure to determine whether they are encrypted. By replacing the original data entity with encrypted ciphertext, a structurally complete, secure, and controllable version of government data sharing is constructed to adapt to the sharing needs of multiple scenarios. This effectively achieves the coexistence of controllable desensitization and structural preservation of government data sharing, providing technical support and practical value for data sharing in high-security scenarios.

[0007] To achieve the above objectives, the present invention provides an adaptive shape-preserving encryption method for government data sharing, comprising the following steps: S1: Perform structured parsing on the original government data, use the improved entity recognition model to identify data entities containing potentially sensitive content in the structured parsed original government data, and generate sensitivity labels for the data entities containing potentially sensitive content; S2: Construct a sensitivity assessment matrix, perform sensitivity assessment on data entities containing potentially sensitive content based on sensitivity labels, and use the improved clustering algorithm to perform sensitivity classification on data entities containing potentially sensitive content, obtaining a set of data entities with three sensitivity classifications; The three types of sensitive data entity sets include highly sensitive data entities, medium sensitive data entities and low sensitive data entities with gradually decreasing sensitivity; S3: Encrypt the data entities after sensitivity classification using an adaptive encryption algorithm to obtain the encrypted ciphertext corresponding to the data entities; the highly sensitive data entities and the medium sensitive data entities are encrypted using a conformal encryption method, and the low sensitive data entities are encrypted using a homomorphic encryption method; S4: Replace the data entities containing potentially sensitive content in the original government data with the corresponding encrypted ciphertext to obtain the government shared data and share it.

[0008] Optionally, the structured parsing includes using regular expressions to match the original government data character by character, cleaning the successfully matched characters, and using a word segmentation algorithm to perform word segmentation to obtain the original government data after structured parsing; the constructed regular expression includes special symbols, HTML tags and garbled characters, and constructs an improved entity recognition model, and the improved entity recognition model includes an input representation layer, a bimodal graph construction layer, a hybrid information propagation layer and a CRF decoding layer.

[0009] Optionally, using the improved entity recognition model to identify data entities containing potentially sensitive content in the original government data after structured parsing, and generating sensitivity labels for the data entities containing potentially sensitive content, includes: The input representation layer is a word vector model, which is used to receive the original government data after structured analysis and generate a word vector for each word; The bimodal graph construction layer is used to calculate the semantic association relationship between any two words based on word vectors, construct a semantic association graph, and extract the word attribute information of each word in the original government data to construct a word structure graph. The word attribute information includes position, TF-IDF value, word type, and the word type. , where word type Corresponding to Chinese, numbers and English respectively; The hybrid information propagation layer is used to convert the word structure graph into an adjacency matrix, where the matrix elements in the adjacency matrix are the connection relationships between any two words. The adjacency matrix is ​​converted into a normalized adjacency matrix for graph convolution based on the degree matrix, and semantic feature extraction and structural feature extraction are performed on the word vectors through multiple rounds of propagation, respectively, to obtain a hybrid information vector for each word in the original government data after structured parsing; The CRF decoding layer is used to receive the mixed information vectors of all words in the original government data after structured parsing, and use the fully connected layer to generate a label mapping matrix of the mixed information vector, use the Viterbi algorithm to perform label transfer and calculate the label transfer score, and use the label transfer scores of all words and the probabilities between words and labels as path scores to obtain a set of label sequences with the largest path scores; the length of the label sequence is consistent with the number of words in the original government data after structured parsing, and the labels in the label sequence are respectively the labels of N words in the original government data after structured parsing, where N represents the number of words in the original government data after structured parsing; the label mapping matrix is ​​a probability matrix between the mixed information vector and different labels, and the labels are divided into two categories, one is a non-sensitive label category that does not contain potential sensitive content, and the other is a sensitive label category that contains potential sensitive content, and the sensitive label category includes several sensitive labels; Extract words with label categories of sensitive label categories as data entities containing potential sensitive content, and extract corresponding sensitivity labels.

[0010] Optionally, semantic feature extraction and structural feature extraction are performed on the word vectors through multiple rounds of propagation based on the normalized adjacency matrix to obtain a mixed information vector for each word in the original government data after structured parsing, including: Construct a word vector matrix corresponding to the word vector in the original government data after structured parsing. The word vector matrix is ​​in the form of a matrix with N rows and Len columns. Len represents the dimension of the word vector. The nth row in the word vector matrix is ​​the word vector of the nth word in the original government data after structured parsing. ; The word vector matrix is ​​subjected to H rounds of semantic feature extraction and structural feature extraction, and the h-th round of propagation formula is: ; ; in, represents the mixed propagation matrix of the word vector matrix after the hth round of propagation, represents the mixed propagation matrix of the word vector matrix after the h-1th round of propagation, represents the feature extraction matrix of the h-th round of propagation, Represents the activation function

[0011] Represents the initial mixing matrix corresponding to the word vector matrix, C represents the word vector matrix, represents the normalized adjacency matrix, represents element-by-element addition, represents a semantic association graph, represents the structural feature extraction matrix, represents the semantic feature extraction matrix; The mixed information vector matrix is ​​in the form of a matrix with N rows, corresponding to the mixed information vectors of N words in the original government data after structured parsing.

[0012] Optionally, a sensitivity assessment matrix is ​​constructed to perform sensitivity assessment on data entities containing potentially sensitive content based on sensitivity labels, including: The sensitivity assessment matrix is ​​the sensitive attributes of different sensitive labels in multi-dimensional sensitivity, the multi-dimensional sensitivity includes legal sensitivity, personal privacy and data uniqueness, and the range of the sensitive attributes is 0-1; The sensitive attributes of the sensitive label in the sensitivity assessment matrix are used as the initial sensitivity assessment result of the data entity, and the word attribute information of the data entity is extracted. The initial sensitivity assessment result is dynamically optimized and adjusted to obtain the sensitivity assessment result of the data entity. The sensitivity assessment result is a three-dimensional vector, which represents the dynamic sensitive attributes of the data entity in terms of legal sensitivity, personal privacy, and data uniqueness, respectively. The word attribute information includes the location of the data entity, the TF-IDF value, and the word type.

[0013] Optionally, based on the sensitivity assessment results of the data entities containing potentially sensitive content, the improved clustering algorithm is used to perform sensitivity classification processing on the data entities containing potentially sensitive content, thereby obtaining a set of data entities with three sensitivity classifications, including: Calculate a sensitivity score for the data entity containing potentially sensitive content based on the sensitivity assessment result, where the sensitivity score is the sum of all vector values ​​in the sensitivity assessment result; Calculate the improved distance between data entities and the adaptive sensitive neighborhood radius of the data entities based on the sensitivity score; Combining the improved distance between data entities and the adaptive sensitive neighborhood radius of data entities, the DBSCAN clustering algorithm is used to perform sensitivity classification on data entities containing potential sensitive content, and a set of data entities with three types of sensitivity classifications is obtained. ,in They are high-sensitivity data entity, medium-sensitivity data entity and low-sensitivity data entity respectively.

[0014] Optionally, an adaptive encryption algorithm is used to encrypt the data entity after sensitivity classification to obtain an encrypted ciphertext corresponding to the data entity, including: Generate an encryption algorithm identifier for the data entity after sensitivity classification, use a pre-trained model to vectorize the data entity, and use an adaptive encryption algorithm to encrypt the vectorized data entity. Concatenate the encryption algorithm identifier and the encryption processing result as the encrypted ciphertext corresponding to the data entity.

[0015] Optionally, data entities containing potentially sensitive content in the original government data are replaced with corresponding encrypted ciphertexts to obtain government shared data, including: Based on the data entity sets with three types of sensitivity levels, the data entities containing potential sensitive content in the original government data are traversed, and the encrypted ciphertext corresponding to the data entities is obtained. The data entities containing potential sensitive content in the original government data are replaced with the corresponding encrypted ciphertext to obtain the government shared data.

[0016] In order to solve the above problems, the present invention provides an adaptive government data sharing and conformal encryption system, which includes a server and a data acquisition device. The data acquisition device is used to collect original government data and perform structured analysis on the original government data. The server includes an entity identification module and an adaptive encryption module: The entity recognition module is used to use the improved entity recognition model to identify data entities containing potentially sensitive content in the original government data after structured parsing, generate sensitivity labels for the data entities containing potentially sensitive content, construct a sensitivity assessment matrix, perform sensitivity assessment on the data entities containing potentially sensitive content based on the sensitivity labels, and use the improved clustering algorithm to perform sensitivity classification processing on the data entities containing potentially sensitive content, thereby obtaining a set of data entities with three sensitivity classifications; The adaptive encryption module is used to encrypt the data entity after sensitivity classification using an adaptive encryption algorithm to obtain the encrypted ciphertext corresponding to the data entity.

[0017] Compared with the prior art, the present invention has the following beneficial effects: First, this application introduces a bimodal graph construction mechanism, integrating two heterogeneous information sources: semantic association graphs and word structure graphs, to fully capture the multi-level relationships between words in the original government data, making up for the shortcomings of traditional sequence models in modeling structural information. In the semantic dimension, the semantic association strength between words is calculated through word vectors to form a semantic graph to reflect contextual dependencies. In the structural dimension, a word structure graph is constructed based on word position information, TF-IDF value, and word type, introducing features in entity representation that are irrelevant to semantics but critical for sensitivity discrimination. Subsequently, in the hybrid information propagation layer, graph convolution propagation of features on the two types of graphs is achieved through the normalized adjacency matrix, effectively fusing semantic and structural signals and enhancing the representational expressiveness of each word. By introducing a CRF decoder to globally annotate the hybrid information vector and using the Viterbi algorithm to obtain the optimal label path, the recognition accuracy of sensitive labels is not only improved, but also the transfer constraints between labels are strengthened, which is significantly better than the recognition method based only on context or a single feature. The overall method can achieve high-accuracy identification of potential sensitive words, providing strong input support for subsequent sensitivity assessment and encryption strategies.

[0018] At the same time, this application innovatively improves the DBSCAN clustering algorithm by introducing an improved distance metric driven by sensitivity scores and an adaptive neighborhood radius mechanism, thus achieving accurate classification processing of potentially sensitive data entities. Compared with traditional clustering methods based only on semantic features or spatial density, this application integrates the multi-dimensional sensitive attribute evaluation results of data entities, so that when the sensitivity differences are significant, the "perception distance" between data entities can be effectively widened to prevent the misjudgment of high and low sensitivity data entities from being aggregated. At the same time, the adaptive neighborhood radius dynamically adjusts the clustering perception scale based on the sensitivity score, so that highly sensitive data entities have a more fine-grained separation capability and improve the accuracy of clustering boundaries. In the encryption stage, differentiated encryption processing is performed based on the sensitivity classification results, and an encryption method selection strategy based on sensitive labels is proposed to improve the security and computational efficiency of the encryption process. Highly sensitive data entities and medium sensitive data entities use conformal encryption to maintain structural consistency, and low sensitive data entities use homomorphic encryption to support computational operations. This effectively achieves the coexistence of controllable desensitization and structural preservation of government shared data, providing technical support and practical value for data sharing in high-security scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 A flowchart of an adaptive conformal encryption method for government data sharing provided by one embodiment of the present invention.

[0020] Figure 2 A schematic diagram of the process of data entity sensitivity classification provided by one embodiment of the present invention. DETAILED DESCRIPTION

[0021] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0022] The present invention provides an adaptive conformal encryption method for shared government data. This method can be performed by at least one of a server, a terminal, or other electronic device capable of executing the method provided by the present invention. In other words, the adaptive conformal encryption method for shared government data can be performed by software or hardware installed on a terminal or server device. The software can be a blockchain platform. The server can include, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster.

[0023] Reference Figure 1 , embodiment 1 of the present invention is: An adaptive shape-preserving encryption method for government data sharing, comprising the following steps: S1: Perform structured analysis on the original government data, use the improved entity recognition model to identify data entities containing potential sensitive content in the original government data after structured analysis, and generate sensitive labels for the data entities containing potential sensitive content.

[0024] Furthermore, the structured parsing includes using regular expressions to match the original government data character by character, cleaning the successfully matched characters, and using a word segmentation algorithm to perform word segmentation to obtain the original government data after structured parsing. The constructed regular expression includes special symbols, HTML tags and garbled characters, and an improved entity recognition model is constructed. The improved entity recognition model includes an input representation layer, a bimodal graph construction layer, a hybrid information propagation layer and a CRF decoding layer.

[0025] It should be noted that special symbols are symbols excluding Chinese characters, English letters, numbers, spaces, and punctuation marks. The garbled characters include Unicode illegal areas. The word segmentation algorithm is the JIEBA word segmentation algorithm.

[0026] Furthermore, the improved entity recognition model is used to identify data entities containing potentially sensitive content in the original government data after structured parsing, and a sensitivity label for the data entity containing potentially sensitive content is generated, including: The input representation layer is a word vector model, which is used to receive the original government data after structured analysis and generate a word vector for each word; Specifically, the word vector model includes the BERT word vector model and the government domain word vector model built based on the government dictionary. The word vector output result of each word in the original government data after structured parsing is: ; in, Represents the word vector of the nth word in the original government data after structured analysis, Represents the BERT word vector of the nth word output by the BERT word vector model, The government affairs domain word vector representing the nth word output by the government affairs domain word vector model; is the vector concatenation processing, N represents the number of segmented words in the original government data after structured parsing; Further explanation: the word vector in the government affairs field is a 5-dimensional one-hot encoding vector. The words in the original government affairs data after structured analysis are matched with the government affairs dictionary to obtain the government affairs category of the word, including unmatched, numbers, place names, institution names and administrative area codes. The one-hot encoding method is used to convert the government affairs category of the word into a one-hot encoding vector; the dimension of the BERT word vector is 128, and the word vector The dimension of is 133.

[0027] The bimodal graph construction layer is used to calculate the semantic association relationship between any two words based on word vectors, construct a semantic association graph, and extract the word attribute information of each word in the original government data to construct a word structure graph. The word attribute information includes position, TF-IDF value, word type, and the word type. , where word type Corresponding to Chinese, numbers, and English, respectively; specifically, TF represents the word frequency, and IDF represents the inverse document frequency of a word, corresponding to the prevalence of the word in all original government data. The lower the IDF, the more common the word appears in multiple original government data. The higher the TF-IDF value, the more important the word is to the current original government data. Assuming that the position of the nth word is n, the calculation formula of the semantic association relationship is: ; in, Indicates the difference between the nth word and the nth word in the original government data after structured analysis. The semantic relationship between words, Both 133 attention mapping matrix, is the scaling factor used to control the stability of the dimension. is 64, They are the nth word and the nth word in the original government data after structured analysis. word vectors of words; They are used to map the projection results of two word vectors in different projection directions. By constructing a loss function that makes semantically related words have a higher semantic association relationship, the loss function is optimized and solved to obtain the attention mapping matrix , is the ReLU activation function; T represents transposition; The representation of the semantic association graph and word structure graph is: ; Among them, e represents the semantic association graph, which is in the form of a matrix with N rows and N columns; Represents a word structure diagram, is the word attribute information of the nth word. In this embodiment, the word attribute information is a 3-dimensional vector, and the word structure diagram is in the form of a matrix with N rows and 3 columns; Specifically, the semantic relationship graph constructs edge weights through the semantic similarity between words, dynamically reflecting the semantic coupling and contextual neighbor relationships between entities, and is suitable for identifying implicit polysemy or weak boundary-sensitive fields in government data; while the word structure graph forms an entity-level graph structure based on the position, importance, and part of speech of the word, which helps to improve the structural perception and inductive generalization capabilities of standardized entities (such as identity cards, address codes, and unit codes).

[0028] The hybrid information propagation layer is used to convert the word structure graph into an adjacency matrix, where the matrix elements in the adjacency matrix are the connection relationships between any two words. The adjacency matrix is ​​converted into a normalized adjacency matrix for graph convolution based on the degree matrix, and semantic feature extraction and structural feature extraction are performed on the word vectors through multiple rounds of propagation, respectively, to obtain a hybrid information vector for each word in the original government data after structured parsing; Specifically, the nth word and the nth word in the original government data after structured analysis The connection between the words is: ; in, Indicates the relationship between the nth word and the nth word in the original government data after structured analysis. The connection between words, represents an exponential function with a natural constant as the base, Indicates the preset position threshold, set is 5; The adjacency matrix is ​​in the form: ; in, Represents the adjacency matrix, which is in the form of a matrix with N rows and N columns; The degree matrix D is a diagonal matrix with N rows and N columns. The matrix element in the nth row and nth column of the degree matrix is ; The transformation formula of the normalized adjacency matrix is: ,in is the normalized adjacency matrix; The CRF decoding layer is used to receive the mixed information vectors of all words in the original government data after structured parsing, and use the fully connected layer to generate a label mapping matrix of the mixed information vector, use the Viterbi algorithm to perform label transfer and calculate the label transfer score, and use the label transfer scores of all words and the probabilities between words and labels as path scores to obtain a set of label sequences with the largest path scores; the length of the label sequence is consistent with the number of words in the original government data after structured parsing, and the labels in the label sequence are respectively the labels of N words in the original government data after structured parsing, where N represents the number of words in the original government data after structured parsing; the label mapping matrix is ​​a probability matrix between the mixed information vector and different labels, and the labels are divided into two categories, one is a non-sensitive label category that does not contain potential sensitive content, and the other is a sensitive label category that contains potential sensitive content, and the sensitive label category includes several sensitive labels; Extract words with label categories of sensitive label categories as data entities containing potential sensitive content, and extract corresponding sensitivity labels.

[0029] Based on the normalized adjacency matrix, semantic feature extraction and structural feature extraction are performed on word vectors through multiple rounds of propagation. The mixed information vector of each word in the original government data after structured analysis is obtained, including: Construct a word vector matrix corresponding to the word vector in the original government data after structured parsing. The word vector matrix is ​​in the form of a matrix with N rows and Len columns. Len represents the dimension of the word vector. The nth row in the word vector matrix is ​​the word vector of the nth word in the original government data after structured parsing; specifically, Len is 133.

[0030] The word vector matrix is ​​subjected to H rounds of semantic feature extraction and structural feature extraction, and the h-th round of propagation formula is: ; ; in, represents the mixed propagation matrix of the word vector matrix after the hth round of propagation, represents the mixed propagation matrix of the word vector matrix after the h-1th round of propagation, represents the feature extraction matrix of the h-th round of propagation, Represents the activation function

[0031] Represents the initial mixing matrix corresponding to the word vector matrix, C represents the word vector matrix, represents the normalized adjacency matrix, represents element-by-element addition, represents a semantic association graph, represents the structural feature extraction matrix, Represents the semantic feature extraction matrix.

[0032] Specifically, the feature extraction matrix, structural feature extraction matrix and semantic feature extraction matrix are all parameters to be trained. A loss function is constructed with the goal of minimizing the error between the corresponding label of the generated mixed information vector and the true label of the word. The loss function is optimized and solved to obtain the parameters to be trained.

[0033] The mixed propagation matrix after the Hth round of propagation is extracted as a mixed information vector matrix. The mixed information vector matrix is ​​in the form of a matrix with N rows, corresponding to the mixed information vectors of N words in the original government data after structured parsing.

[0034] It should be noted that the hybrid information propagation layer in the present invention is based on the normalized adjacency matrix and adopts a multi-round information transmission mechanism, so that the semantic association graph and word structure graph that describe semantic information and structural constraints in the bimodal graph construction layer are efficiently integrated and enhanced in the graph space. This mechanism effectively overcomes the problems of semantic dilution and structural misrecognition of traditional models when processing long texts and cross-domain entity recognition, and improves the sensitive recognition ability of complex nested entities, combined entities and fields with fuzzy semantic boundaries.

[0035] It is further explained that this application introduces a bimodal graph construction mechanism, integrating two types of heterogeneous information sources, semantic association graph and word structure graph, to fully capture the multi-level relationship between words in the original government data, and make up for the defect of the traditional sequence model's insufficient ability to model structural information. In the semantic dimension, the semantic association strength between words is calculated through word vectors to form a semantic graph to reflect contextual dependencies; in the structural dimension, a word structure graph is constructed based on the word's position information, TF-IDF value and word type, and features in the entity representation that are irrelevant to semantics but critical to sensitivity judgment are introduced. Subsequently, in the mixed information propagation layer, the graph convolution propagation of features on the two types of graphs is realized through the normalized adjacency matrix, effectively fusing semantic and structural signals and enhancing the representational expression ability of each word. By introducing a CRF decoder to globally annotate the mixed information vector and using the Viterbi algorithm to obtain the optimal label path, not only the recognition accuracy of sensitive labels is improved, but also the transfer constraints between labels are strengthened, which is significantly better than the recognition method based only on context or a single feature.

[0036] S2: Construct a sensitivity assessment matrix, perform sensitivity assessment on data entities containing potential sensitive content based on sensitivity labels, and use the improved clustering algorithm to perform sensitivity classification on data entities containing potential sensitive content, obtaining a set of data entities with three sensitivity classifications.

[0037] Construct a sensitivity assessment matrix and perform sensitivity assessment on data entities containing potentially sensitive content based on sensitivity labels, including: The sensitivity assessment matrix is ​​the sensitive attributes of different sensitive labels in multi-dimensional sensitivity, the multi-dimensional sensitivity includes legal sensitivity, personal privacy and data uniqueness, and the range of the sensitive attributes is 0-1; Specifically, the sensitive labels include citizen identity identification, residence information, contact information, financial account information, biometric information, political status, ethnicity, certificate information and work information; as an embodiment of the present invention, there are multiple types of data under the sensitive labels, for example, citizen identity identification includes ID card number, passport number, etc., and certificate information includes driver's license, business license, tax registration number, etc.

[0038] Specifically, we distributed privacy sensitivity questionnaires to 30 data managers or security auditors in the government sector, collected multi-dimensional sensitivity scores for various sensitive labels, and normalized them to obtain a sensitivity assessment matrix. The sensitive attributes of the sensitive label in the sensitivity assessment matrix are used as the initial sensitivity assessment result of the data entity, and the word attribute information of the data entity is extracted. The initial sensitivity assessment result is dynamically optimized and adjusted to obtain the sensitivity assessment result of the data entity. The sensitivity assessment result is a three-dimensional vector, which represents the dynamic sensitive attributes of the data entity in terms of legal sensitivity, personal privacy, and data uniqueness, respectively. The word attribute information includes the location of the data entity, the TF-IDF value, and the word type.

[0039] Specifically, the initial sensitivity assessment result of the data entity is ,in These are the legal sensitivity, personal privacy, and data uniqueness of the sensitive labels corresponding to the data entities in the sensitivity assessment matrix; The dynamic optimization adjustment formula is: ; ; ; ; in, Indicates the sensitivity assessment result of the data entity. Represents the dynamic optimization coefficient, Set to 0.7, They are the word frequency TF value, TF-IDF value and position of the data entity.

[0040] It should be noted that the original sensitive attribute vector of the sensitive label (legal sensitivity, personal privacy, and data uniqueness) is used as prior knowledge, but this prior knowledge is a static evaluation result, ignoring the differences in the specific data entities in the actual context. By introducing features such as word frequency, position, and TF-IDF value, the sensitive attributes are dynamically adjusted in a context-driven manner. The higher the word frequency, the more likely the data entity is a common word, which reduces the legal sensitivity. The higher it is, the more likely it is a personalized identifier, thus increasing personal privacy; the further back the position is, the further away the data entity is from the beginning of the text and the title area, reducing data uniqueness.

[0041] Based on the sensitivity assessment results of the data entities containing potentially sensitive content, the improved clustering algorithm is used to perform sensitivity classification processing on the data entities containing potentially sensitive content, and a set of data entities with three sensitivity classifications is obtained, including: Calculate a sensitivity score for the data entity containing potentially sensitive content based on the sensitivity assessment result, where the sensitivity score is the sum of all vector values ​​in the sensitivity assessment result; The improved distance between data entities and the adaptive sensitive neighborhood radius of the data entities are calculated based on the sensitivity scores; specifically, the improved distance between the mth data entity and the xth data entity is: ; in, represents the improved distance between the mth data entity and the xth data entity, Represent the sensitivity assessment results of the mth data entity and the xth data entity respectively, represents the L2 norm, Represents the sensitivity score of the mth data entity and the xth data entity, Represents a sensitive regulatory factor, Respectively represent the mixed information vectors of the mth data entity and the xth data entity in the improved entity recognition model, M represents the number of data entities containing potential sensitive content, and set is 0.2, For selection The minimum value in The smaller it is, the smaller the sensitivity difference between the two data entities is. Represents the semantic distance between two data entities.

[0042] It should be noted that the improved distance effectively reflects the constraint logic that the greater the sensitivity difference, the farther the entity semantic distance is, preventing entities with similar semantics but significant differences in sensitive attributes from being incorrectly aggregated, and ensuring the principle of sensitivity isolation.

[0043] The adaptive sensitive neighborhood radius of the mth data entity is: ; in, represents the adaptive sensitive neighborhood radius of the mth data entity, Indicates the radius control coefficient, which controls the extent of radius contraction of data entities with high sensitivity scores. is 0.6; Indicates the initial neighborhood radius, set is 1.8.

[0044] It should be noted that the adaptive sensitive neighborhood radius can dynamically compress the cluster perception range according to the entity sensitivity score. The higher the sensitivity score of the entity, the more compact its local neighborhood, thereby improving the boundary protection capability of highly sensitive entities and preventing them from being "pulled in" by medium and low-sensitivity entities or having blurred boundaries. At the same time, low-sensitivity entities maintain a larger radius to improve recall efficiency and achieve differential adjustment of cluster density; the dual information flow of semantics and sensitivity evaluation results is integrated, and it has good semantic retention and sensitivity graded perception capabilities.

[0045] Combining the improved distance between data entities and the adaptive sensitive neighborhood radius of data entities, the DBSCAN clustering algorithm is used to perform sensitivity classification on the data entities containing potential sensitive content, and a set of data entities with three types of sensitivity classifications is obtained. ,in The data entities are classified as highly sensitive, medium sensitive, and low sensitive, respectively. Specifically, the higher the sensitivity, the greater the privacy risk and harm caused by leakage.

[0046] like Figure 2 As shown in the figure, the process diagram of data entity sensitivity classification is shown. Word1 to wordN are N consecutive words in the original government data after structured parsing, F1 to FN are mixed information vectors of N consecutive words, G_1 to G_M are M data entities containing potential sensitive content, data is a set of data entities with three types of sensitivity classification, data1, data2, and data3 are data entity sets consisting of high-sensitive data entities, medium-sensitive data entities, and low-sensitive data entities respectively.

[0047] It should be noted that this application innovatively improves the DBSCAN clustering algorithm by introducing an improved distance metric driven by sensitivity scores and an adaptive neighborhood radius mechanism, thus achieving accurate hierarchical processing of potentially sensitive data entities. Compared with traditional clustering methods based only on semantic features or spatial density, this application integrates the multi-dimensional sensitive attribute evaluation results of data entities, so that when the sensitivity differences are significant, the "perception distance" between entities can be effectively widened to prevent the misjudgment of high and low sensitive entities from being aggregated; at the same time, the adaptive neighborhood radius dynamically adjusts the clustering perception scale according to the sensitivity score, so that highly sensitive entities have a more fine-grained separation capability and improve the accuracy of cluster boundaries.

[0048] S3: Adopt an adaptive encryption algorithm to encrypt the data entity after sensitivity classification to obtain the encrypted ciphertext corresponding to the data entity. The highly sensitive data entity and the medium sensitive data entity adopt conformal encryption, and the low sensitive data entity adopts homomorphic encryption.

[0049] Adopting an adaptive encryption algorithm to encrypt the data entity after sensitivity classification, the encrypted ciphertext corresponding to the data entity is obtained, including: Generate an encryption algorithm identifier for the data entity after sensitivity classification, use a pre-trained model to vectorize the data entity, and use an adaptive encryption algorithm to encrypt the vectorized data entity. Concatenate the encryption algorithm identifier and the encryption processing result as the encrypted ciphertext corresponding to the data entity.

[0050] It should be noted that in the encryption stage, differentiated encryption processing is performed in combination with the sensitivity classification results, and an encryption method selection strategy based on sensitive labels is proposed to improve the security and computational efficiency of the encryption process. Conformal encryption is used for high-sensitivity and medium-sensitivity entities to maintain structural consistency, and homomorphic encryption is used for low-sensitivity entities to support computing operations. This effectively achieves the coexistence of controllable desensitization and structural preservation of government shared data, providing technical support and practical value for data sharing in high-security scenarios.

[0051] S4: Replace the data entities containing potentially sensitive content in the original government data with the corresponding encrypted ciphertext to obtain the government shared data and share it.

[0052] The data entities containing potentially sensitive content in the original government data are replaced with the corresponding encrypted ciphertext to obtain the government shared data, including: Based on the data entity sets with three types of sensitivity levels, the data entities containing potential sensitive content in the original government data are traversed, and the encrypted ciphertext corresponding to the data entities is obtained. The data entities containing potential sensitive content in the original government data are replaced with the corresponding encrypted ciphertext to obtain the government shared data.

[0053] As an embodiment of the present invention, the encryption method is identified according to the encryption algorithm identifier in the encrypted ciphertext, and the encrypted ciphertext in the government shared data is decrypted in combination with the encryption key and the inverse operation of the encryption algorithm.

[0054] Example 2: An adaptive government affairs shared data conformal encryption system, the system comprising a server and a data acquisition device; The data collection device is used to collect original government data and perform structured analysis on the original government data; The server includes an entity identification module and an adaptive encryption module: The entity recognition module is used to use the improved entity recognition model to identify data entities containing potentially sensitive content in the original government data after structured parsing, generate sensitivity labels for the data entities containing potentially sensitive content, construct a sensitivity assessment matrix, perform sensitivity assessment on the data entities containing potentially sensitive content based on the sensitivity labels, and use the improved clustering algorithm to perform sensitivity classification processing on the data entities containing potentially sensitive content, thereby obtaining a set of data entities with three sensitivity classifications; The adaptive encryption module is used to encrypt the data entity after sensitivity classification using an adaptive encryption algorithm to obtain the encrypted ciphertext corresponding to the data entity.

[0055] It should be understood that the embodiment is for illustration only and the scope of the patent application is not limited to this structure.

[0056] It should be noted that the serial numbers of the above-mentioned embodiments of the present invention are for descriptive purposes only and do not represent the advantages or disadvantages of the embodiments. In addition, the terms "including", "comprising" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, device, article or method comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, device, article or method. In the absence of further restrictions, an element defined by the sentence "including a ..." does not exclude the presence of other identical elements in the process, device, article or method comprising the element.

[0057] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present invention.

[0058] The above are only preferred embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. An adaptive shape-preserving encryption method for government data sharing, characterized in that: The method comprises: S1: Perform structured parsing on the original government data, use the improved entity recognition model to identify data entities containing potentially sensitive content in the structured parsed original government data, and generate sensitivity labels for the data entities containing potentially sensitive content; S2: Construct a sensitivity assessment matrix, perform sensitivity assessment on data entities containing potentially sensitive content based on sensitivity labels, and use the improved clustering algorithm to perform sensitivity classification on data entities containing potentially sensitive content, obtaining a set of data entities with three sensitivity classifications; The three types of sensitive data entity sets include highly sensitive data entities, medium sensitive data entities and low sensitive data entities with gradually decreasing sensitivity; S3: Using an adaptive encryption algorithm to encrypt the data entities after sensitivity classification to obtain the encrypted ciphertext corresponding to the data entities. The highly sensitive data entities and the medium sensitive data entities are encrypted using a conformal encryption method, and the low sensitive data entities are encrypted using a homomorphic encryption method; S4: Replace the data entities containing potentially sensitive content in the original government data with the corresponding encrypted ciphertext to obtain the government shared data and share it.

2. The adaptive government data sharing and shape-preserving encryption method according to claim 1, characterized in that: The structured parsing includes matching the original government data character by character using regular expressions, cleaning the successfully matched characters, and segmenting the words using a word segmentation algorithm to obtain the original government data after structured parsing; The constructed regular expression includes special symbols, HTML tags and garbled characters, and an improved entity recognition model is constructed. The improved entity recognition model includes an input representation layer, a bimodal graph construction layer, a hybrid information propagation layer and a CRF decoding layer.

3. The adaptive government data sharing and shape-preserving encryption method according to claim 2, characterized in that: The improved entity recognition model is used to identify data entities containing potentially sensitive content in the original government data after structured parsing, and to generate sensitivity labels for the data entities containing potentially sensitive content, including: The input representation layer is a word vector model, which is used to receive the original government data after structured analysis and generate a word vector for each word; The bimodal graph construction layer is used to calculate the semantic association relationship between any two words based on word vectors, construct a semantic association graph, and extract the word attribute information of each word in the original government data to construct a word structure graph. The word attribute information includes position, TF-IDF value, word type, and the word type. , where word type Corresponding to Chinese, numbers and English respectively; The hybrid information propagation layer is used to convert the word structure graph into an adjacency matrix. The matrix elements in the adjacency matrix are the connection relationships between any two words. Based on the degree matrix, the adjacency matrix is ​​converted into a normalized adjacency matrix for graph convolution. The word vectors are subjected to multiple rounds of semantic feature extraction and structural feature extraction, respectively, to obtain a hybrid information vector for each word in the original government data after structured parsing. The CRF decoding layer is used to receive the mixed information vectors of all words in the original government data after structured parsing, and use the fully connected layer to generate a label mapping matrix of the mixed information vector, use the Viterbi algorithm to perform label transfer and calculate the label transfer score, and use the label transfer scores of all words and the probabilities between words and labels as path scores to obtain a set of label sequences with the largest path scores; the length of the label sequence is consistent with the number of words in the original government data after structured parsing, and the labels in the label sequence are respectively the labels of N words in the original government data after structured parsing, where N represents the number of words in the original government data after structured parsing; the label mapping matrix is ​​a probability matrix between the mixed information vector and different labels, and the labels are divided into two categories, one is a non-sensitive label category that does not contain potential sensitive content, and the other is a sensitive label category that contains potential sensitive content, and the sensitive label category includes several sensitive labels; Extract words with label categories of sensitive label categories as data entities containing potential sensitive content, and extract corresponding sensitivity labels.

4. The adaptive government data sharing and shape-preserving encryption method according to claim 3, characterized in that: Based on the normalized adjacency matrix, semantic feature extraction and structural feature extraction are performed on word vectors through multiple rounds of propagation. The mixed information vector of each word in the original government data after structured analysis is obtained, including: Construct a word vector matrix corresponding to the word vector in the original government data after structured parsing. The word vector matrix is ​​in the form of a matrix with N rows and Len columns. Len represents the dimension of the word vector. The nth row in the word vector matrix is ​​the word vector of the nth word in the original government data after structured parsing. ; The word vector matrix is ​​subjected to H rounds of semantic feature extraction and structural feature extraction, and the h-th round of propagation formula is: ; ; in, represents the mixed propagation matrix of the word vector matrix after the hth round of propagation, represents the mixed propagation matrix of the word vector matrix after the h-1th round of propagation, represents the feature extraction matrix of the h-th round of propagation, represents the activation function; Represents the initial mixing matrix corresponding to the word vector matrix, C represents the word vector matrix, represents the normalized adjacency matrix, represents element-by-element addition, represents a semantic association graph, represents the structural feature extraction matrix, represents the semantic feature extraction matrix; The mixed propagation matrix after the Hth round of propagation is extracted as a mixed information vector matrix. The mixed information vector matrix is ​​in the form of a matrix with N rows, corresponding to the mixed information vectors of N words in the original government data after structured parsing.

5. The adaptive government data sharing and shape-preserving encryption method according to claim 1, characterized in that: Construct a sensitivity assessment matrix and perform sensitivity assessment on data entities containing potentially sensitive content based on sensitivity labels, including: The sensitivity assessment matrix is ​​the sensitive attributes of different sensitive labels in multi-dimensional sensitivity, the multi-dimensional sensitivity includes legal sensitivity, personal privacy and data uniqueness, and the range of the sensitive attributes is 0-1; The sensitive attributes of the sensitive label in the sensitivity assessment matrix are used as the initial sensitivity assessment result of the data entity, and the word attribute information of the data entity is extracted. The initial sensitivity assessment result is dynamically optimized and adjusted to obtain the sensitivity assessment result of the data entity. The sensitivity assessment result is a three-dimensional vector, which represents the dynamic sensitive attributes of the data entity in terms of legal sensitivity, personal privacy, and data uniqueness, respectively. The word attribute information includes the location of the data entity, the TF-IDF value, and the word type.

6. The adaptive government data sharing and shape-preserving encryption method according to claim 5, characterized in that: Based on the sensitivity assessment results of the data entities containing potentially sensitive content, the improved clustering algorithm is used to perform sensitivity classification processing on the data entities containing potentially sensitive content, and a set of data entities with three sensitivity classifications is obtained, including: Calculate a sensitivity score for the data entity containing potentially sensitive content based on the sensitivity assessment result, where the sensitivity score is the sum of all vector values ​​in the sensitivity assessment result; Calculate the improved distance between data entities and the adaptive sensitive neighborhood radius of the data entities based on the sensitivity score; Combining the improved distance between data entities and the adaptive sensitive neighborhood radius of data entities, the DBSCAN clustering algorithm is used to perform sensitivity classification on data entities containing potential sensitive content, and a set of data entities with three types of sensitivity classifications is obtained. ,in They are high-sensitivity data entity, medium-sensitivity data entity and low-sensitivity data entity respectively.

7. The adaptive government data sharing and shape-preserving encryption method according to claim 6, characterized in that: Adopting an adaptive encryption algorithm to encrypt the data entity after sensitivity classification, the encrypted ciphertext corresponding to the data entity is obtained, including: Generate an encryption algorithm identifier for the data entity after sensitivity classification, use a pre-trained model to vectorize the data entity, and use an adaptive encryption algorithm to encrypt the vectorized data entity. Concatenate the encryption algorithm identifier and the encryption processing result as the encrypted ciphertext corresponding to the data entity.

8. The adaptive government data sharing and shape-preserving encryption method according to claim 1, characterized in that: The data entities containing potentially sensitive content in the original government data are replaced with the corresponding encrypted ciphertext to obtain the government shared data, including: Based on the data entity sets with three types of sensitivity levels, the data entities containing potential sensitive content in the original government data are traversed, and the encrypted ciphertext corresponding to the data entities is obtained. The data entities containing potential sensitive content in the original government data are replaced with the corresponding encrypted ciphertext to obtain the government shared data.

9. An adaptive government data sharing and conformal encryption system, characterized in that: The adaptive government affairs shared data conformal encryption system includes a server and a data acquisition device; The data collection device is used to collect original government data and perform structured analysis on the original government data; The server includes an entity identification module and an adaptive encryption module: The entity recognition module is used to use the improved entity recognition model to identify data entities containing potentially sensitive content in the original government data after structured parsing, generate sensitivity labels for the data entities containing potentially sensitive content, construct a sensitivity assessment matrix, perform sensitivity assessment on the data entities containing potentially sensitive content based on the sensitivity labels, and use the improved clustering algorithm to perform sensitivity classification processing on the data entities containing potentially sensitive content, thereby obtaining a set of data entities with three sensitivity classifications; The adaptive encryption module is used to encrypt the data entity after sensitivity classification using an adaptive encryption algorithm to obtain the encrypted ciphertext corresponding to the data entity; To implement an adaptive government data sharing conformal encryption method as described in any one of claims 1-8.

Citation Information

Patent Citations

  • A method for encryption and decryption in e-government

    CN108924104B

  • Method for carrying out security protection on sensitive data through natural language analysis

    CN110795751A

  • Dynamic desensitization method for structured data

    CN119066698A

  • Government affair data sharing system based on data security law risk control mode

    CN119989417A

  • Format-preserving encryption method based on stream cipher

    US20210135839A1

Cited By

  • Self-adaptive encryption method and system based on mail content semantic features

    CN121967371A

  • Cloud government affair sensitive data encryption protection method

    CN122153941A