Desensitization information conversion system based on work order circulation

Through a multi-module collaborative desensitized information conversion system, the system achieves accurate identification, layered encryption, and missing information compensation for power grid work orders. This solves the problem of balancing sensitive information protection, processing efficiency, and file integrity in existing technologies, ensuring information security and efficient flow of power grid business.

CN121919903APending Publication Date: 2026-04-24国网福建省电力有限公司营销服务中心
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
国网福建省电力有限公司营销服务中心
Filing Date
2025-12-27
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing power grid work order desensitization technology cannot simultaneously meet the balance requirements of sensitive information protection, efficient processing and document integrity. It suffers from problems such as insufficient accuracy in sensitive information identification, semantic breaks caused by desensitization, insufficient flexibility of encryption strategies and insufficient security of mobile devices.

Method used

The desensitized information conversion system employs a multi-module collaborative approach, including an information identification module, a missing information compensation module, an encryption and encapsulation module, and a burning and decryption module. It identifies sensitive information through a multi-dimensional evaluation algorithm, generates sensitive information annotations, performs layered encryption and missing information compensation, and decrypts the information in conjunction with the unique permission conditions of the mobile device, ensuring information security and integrity.

Benefits of technology

It achieves precise location and layered encryption of sensitive information in power grid work orders, ensuring information security and circulation efficiency, guaranteeing document integrity, adapting to cross-scenario circulation, and supporting the efficient operation of power grid business.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121919903A_ABST
    Figure CN121919903A_ABST
Patent Text Reader

Abstract

The invention relates to a work order circulation-based desensitization information conversion system, which comprises an information identification module, a loss compensation module, an encryption packaging module, a combustion decryption module and a data interaction module, and realizes power grid work order sensitive information full-process management and control through multi-module cooperation. The information identification module accurately locates sensitive information, and the missing compensation module ensures that a work order is complete and readable. The encryption packaging module carries out layered encryption to enhance security, and the combustion decryption module can control decryption and can clear and prevent leakage after being used. And the data interaction module adapts to cross-scene circulation, finally balances information security, processing efficiency and work order integrity, and supports efficient development of power grid services.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power grid data processing technology, and more specifically, to a de-identified information conversion system based on work order flow. Background Technology

[0002] With the advancement of smart grid construction, work orders have become a key data carrier for core businesses such as power grid equipment operation and maintenance, fault handling, and user services. Power grid work orders not only contain core business information such as equipment parameters, operation and maintenance records, and dispatch instructions, but also involve sensitive content such as user identity information, address information, and electricity consumption data. At the same time, they need to meet the needs of cross-departmental and cross-level transfer, and their processing quality directly affects the efficiency of power grid business operations and data security.

[0003] In practical applications, the processing of power grid work orders faces a balancing challenge of three core requirements: First, information sensitivity protection. If sensitive data in power grid work orders is leaked, it may lead to problems such as user privacy leaks and power grid operation safety risks. Data security needs to be ensured through desensitization processing. Second, circulation and processing efficiency. Power grid business has high requirements for work order response speed. Complex encryption or desensitization processes may lead to circulation delays, affecting the operation of time-sensitive businesses such as fault handling and operation and maintenance scheduling. Third, document integrity. The semantic coherence and information integrity of work orders are the foundation for subsequent business processing. Simple removal or replacement of sensitive information can easily cause semantic breaks and missing key information, resulting in the work order being unable to be accurately interpreted and used.

[0004] Existing work order desensitization technologies have several shortcomings: First, the accuracy of sensitive information identification is insufficient, relying heavily on single-rule matching, making it difficult to identify related and structured sensitive data in power grid work orders, easily leading to missed or false positives; Second, desensitization often uses direct deletion or fixed replacement methods, lacking a compensation mechanism for missing information in the semantic context of power grid work orders, resulting in compromised work order integrity and affecting the efficiency of subsequent business integration; Third, encryption and decryption strategies lack flexibility, failing to dynamically adapt to different users' business permissions, either leading to excessive encryption and low processing efficiency, or loose permission control causing the risk of sensitive information leakage; Fourth, some desensitization systems do not consider the security of mobile device processing scenarios, leaving residual decrypted data that can easily cause secondary leakage, and lacking an effective risk management mechanism.

[0005] In summary, existing technologies cannot simultaneously meet the balance requirements of power grid work orders in terms of sensitive information protection, processing efficiency, and document integrity. There is an urgent need for a desensitized information conversion system adapted to power grid business scenarios to solve the above-mentioned technical pain points. Summary of the Invention

[0006] In view of this, the purpose of this invention is to provide a de-identified information conversion system based on work order flow.

[0007] To solve the above-mentioned technical problems, the technical solution of the present invention is: a de-identified information conversion system based on work order flow, comprising:

[0008] The information identification module includes an information identification unit and an information tagging unit. The information identification unit is used to identify sensitive information in the work order file, and the information tagging unit is used to tag the sensitive information to generate sensitive information annotations.

[0009] The missing information compensation module includes a missing information identification unit and a missing information compensation unit. The missing information identification unit is used to identify missing positions in the work order file after removing sensitive information, and the missing information compensation unit is used to write replacement data to the missing positions.

[0010] The encryption and encapsulation module includes a sensitive encryption unit, a layer encapsulation unit, and a layer encryption unit. The sensitive encryption unit is configured with a first encryption strategy to encrypt sensitive information to generate sensitive information ciphertext and a corresponding sensitive information key. The layer encapsulation unit divides the sensitive information ciphertext into layers according to the sensitive information annotation and encapsulates each layer containing the sensitive information ciphertext to generate a layer ciphertext encapsulation dataset. The layer encryption unit is configured with a second encryption strategy to encrypt the layer ciphertext encapsulation dataset to generate layer encrypted ciphertext and a layer encryption key. The layer encryption key includes an index subkey and a matching subkey.

[0011] The burning and decryption module includes an index burning unit, a layer decryption unit, and a sensitive decryption unit. The index burning unit matches each layer of encrypted ciphertext with an index subkey and deletes other layers of encrypted ciphertext based on the unique matching result. The layer decryption unit is used to decrypt the layer of encrypted ciphertext using the matching subkey and a second decryption strategy to generate sensitive information ciphertext. The sensitive decryption unit is used to decrypt the sensitive information ciphertext using a sensitive information key and a first decryption strategy to obtain the corresponding sensitive information.

[0012] The data interaction module includes a data sending unit and a data receiving unit. The data sending unit is used to send the de-identified work order file and the corresponding layered encrypted dataset. The data receiving unit is used to receive the work order file and the layered encrypted dataset, and generate the corresponding layered encryption key and sensitive information key.

[0013] Furthermore: the information identification unit is configured with a sensitive information identification model, which is used to identify sensitive information in the information identification unit. The sensitive information identification model is configured with a multi-dimensional information evaluation algorithm, which is used to calculate the sensitivity value of each field. When the sensitivity value is greater than a preset sensitivity value benchmark, the field is output as sensitive information.

[0014] Furthermore, the multidimensional evaluation algorithm is configured to weight different types of sensitivity sub-values ​​to obtain the sensitivity value. The types of sensitivity sub-values ​​include: association sensitivity sub-values, content sensitivity sub-values, structure sensitivity sub-values, and type sensitivity sub-values. The association sensitivity sub-value reflects the sensitivity corresponding to fields with association relationships. The content sensitivity sub-value reflects the sensitivity corresponding to the field itself. The structure sensitivity sub-value reflects the sensitivity corresponding to the structure of the field in the segment. The type sensitivity sub-value reflects the sensitivity of the description type corresponding to the field.

[0015] Furthermore: the information tag unit is configured with an annotation clue database, which stores a number of annotation clues. Each annotation clue corresponds to an access permission feature. The information tag unit identifies the access permission feature corresponding to the sensitive information based on the annotation clue to generate the sensitive information annotation.

[0016] Furthermore, the missing information identification unit is equipped with a pre-trained semantic recognition model, which is used to identify whether there are semantic missing information in the segment after the sensitive information is extracted and to generate the corresponding missing position.

[0017] Furthermore: the missing information compensation unit determines the corresponding load condition alternative data group based on the sensitive information annotation corresponding to the sensitive information, and calculates the completion evaluation value of each alternative data in the alternative data group through a preset missing information evaluation algorithm. The completion evaluation value is used to evaluate the completeness of the segment corresponding to the alternative data, and the alternative data with the highest completion evaluation value is output.

[0018] Furthermore, it also includes a central server, which is configured with an encryption parameter library. The encryption parameter library is configured with several encryption factors and corresponding decryption factors. The central server updates the corresponding encryption factors and decryption factors according to permissions and sends them to the user terminal.

[0019] Furthermore, each user terminal is configured with corresponding permission information. The user terminal updates the permission information in real time and uploads it to the central server. The central server updates the corresponding encryption factor and decryption factor according to the permission information.

[0020] Furthermore, the combustion decryption module is configured in a mobile device, which is configured with a unique permission condition. When the decryption feedback information of the second decryption strategy meets the unique permission condition, the corresponding sensitive information is allowed to be output.

[0021] Furthermore, the combustion decryption module also includes a decryption combustion unit, which is configured with combustion triggering conditions. The decryption combustion unit determines whether the combustion triggering conditions are met based on the decryption feedback information of the second decryption strategy. If the combustion triggering conditions are met, the current layer encryption ciphertext and the corresponding layer encryption key are deleted.

[0022] The main technical advantages of this invention are reflected in the following aspects: The solution achieves full-process control of sensitive information in power grid work orders through multi-module collaboration. The information identification module accurately locates sensitive information, while the missing information compensation module ensures the integrity and readability of the work order. The encryption and encapsulation module uses layered encryption to enhance security, and the burning and decryption module provides controllable decryption and immediate clearing after use to prevent leakage. The data interaction module adapts to cross-scenario workflows, ultimately balancing information security, processing efficiency, and work order integrity to support the efficient operation of power grid services. Attached Figure Description

[0023] Figure 1 : System architecture schematic diagram of this invention;

[0024] Figure 2 : Flowchart of the encryption logic of this invention;

[0025] Figure 3 : Flowchart of the decryption logic of this invention. Detailed Implementation

[0026] The specific embodiments of the present invention will be further described in detail below with reference to the accompanying drawings, so that the technical solution of the present invention can be more easily understood and mastered.

[0027] Reference Figure 1 As shown, a de-identified information conversion system based on work order flow includes:

[0028] The information recognition module includes an information recognition unit and an information tag unit.

[0029] The information identification unit is used to identify sensitive information in work order files. The information identification unit is configured with a sensitive information identification model, which is used to identify sensitive information within the information identification unit. The sensitive information identification model is also configured with a multi-dimensional information evaluation algorithm, which is used to calculate the sensitivity value of each field. When the sensitivity value exceeds a preset sensitivity value benchmark, the field is output as sensitive information. The training of the sensitive information identification model must follow the characteristics of power grid work order data and business requirements. The specific steps are as follows: First, collect multi-type work order data from the power grid industry over the past three years, covering equipment maintenance work orders, user repair work orders, dispatch instruction work orders, metering and testing work orders, etc., to ensure that the data covers the core business scenarios of the power grid. The collected work order data is manually labeled, including whether each field is sensitive information, the specific type of sensitive information, and auxiliary features such as associated fields, the work order module, and field description type for each field, forming an initial training dataset. Secondly, the initial training dataset was preprocessed: blank work orders, incorrectly formatted work orders, and other invalid data were removed; the field text was segmented using the Jieba word segmentation tool, splitting text like "User ID number: 110101XXXX" into words such as "User ID number" and "110101XXXX"; non-text structured data such as device number and date were converted into a unified text format such as "Device number: D12345" and "Work order date: 2024-10-01"; and the word-segmented text was converted into a 256-dimensional vector form using the Word2Vec tool as the input features of the model. Next, the preprocessed training dataset was divided into training, validation, and test sets in a 7:2:1 ratio. A pre-trained BERT-base model was selected as the base model, and fine-tuning was performed on the training set. During training, the error between the calculated sensitivity value and the manually labeled sensitivity judgment result was used as the loss function. The Adam optimizer was used to iteratively optimize the model parameters, with a learning rate of 2e-5 and a batch size of 32. Every 10 iterations, the model's accuracy, recall, and F1 score were calculated on the validation set. Training was stopped when the F1 score on the validation set showed no improvement for three consecutive iterations to avoid overfitting. Finally, the trained model was applied to the test set to verify its performance. The model was required to achieve a sensitivity recognition accuracy of at least 95% and a recall of at least 92%. If these requirements were not met, the model parameters were adjusted and retrained until the model performance met the standards. The sensitive information identification model takes a text vector of a single field in a work order as input and outputs the associated sensitive sub-values, content sensitive sub-values, structure sensitive sub-values, type sensitive sub-values, and the final sensitive value of that field, providing data support for sensitive information determination.

[0030] The multidimensional evaluation algorithm is configured to weight different types of sensitive sub-values ​​to obtain the sensitivity value. The types of sensitive sub-values ​​include: association sensitive sub-values, content sensitive sub-values, structure sensitive sub-values, and type sensitive sub-values. The association sensitive sub-value reflects the sensitivity of fields with which the field has an association relationship; the content sensitive sub-value reflects the sensitivity of the field itself; the structure sensitive sub-value reflects the sensitivity of the field's structure within a segment; and the type sensitive sub-value reflects the sensitivity of the description type corresponding to the field. The information tagging unit is configured with an annotation clue database, which stores several annotation clues. Each annotation clue corresponds to a permission feature. The information tagging unit identifies the permission features corresponding to sensitive information based on the annotation clues to generate the sensitive information annotation. The information multidimensional evaluation algorithm configured in the sensitive information identification model obtains the final sensitivity value by weighting the four types of sensitive sub-values ​​of a field, thereby determining whether a field is sensitive information. Specifically, it includes:

[0031] The final sensitivity value is obtained by weighted summation of the four types of sensitivity sub-values, as shown in the following formula:

[0032] S=w1×S1+w2×S2+w3×S3+w4×S4

[0033] Where S represents the final sensitivity value of the field, with a value range of [0,1]. The larger the value, the higher the sensitivity of the field. w1, w2, w3, and w4 are the weight coefficients of the association sensitivity sub-value, content sensitivity sub-value, structure sensitivity sub-value, and type sensitivity sub-value, respectively, and satisfy w1+w2+w3+w4=1. This weight is determined by the grid search method: traverse all weight combinations in the interval [0,1] with a step size of 0.1, substitute each combination into the model validation set to calculate the identification F1 value, and select the combination with the highest F1 value as the final weight, for example, w1=0.25, w2=0.35, w3=0.2, w4=0.2. S1 is the association sensitivity sub-value, S2 is the content sensitivity sub-value, S3 is the structure sensitivity sub-value, and S4 is the type sensitivity sub-value, all of which have a value range of [0,1].

[0034] Sensitive sub-value S1: This reflects the degree of association between a field and other sensitive fields. It is calculated as the average sensitivity value of all associated fields of this field, as shown in the following formula:

[0035]

[0036] Where n is the total number of associated fields for the current field. Associated fields are determined through a power grid work order field association graph, which is constructed by domain experts based on power grid business logic and includes semantic and business relationships between fields; S′ i Let S be the final sensitivity value of the i-th associated field. If the associated field is not determined to be sensitive information, then S′ iTake 0.

[0037] Content Sensitivity Sub-value S2: This reflects the sensitivity of the field's content. It is calculated based on the similarity between the field and a preset sensitive word library, using the following formula:

[0038]

[0039] Where m represents the total number of sensitive words in the preset sensitive word library, which is jointly constructed by power grid security experts and business experts, and includes user privacy words and power grid operation sensitive words, and is updated quarterly according to new business scenarios; TF ij IDF is the word frequency (IF) of the j-th sensitive word in the current field text, which is the ratio of the number of times the sensitive word appears to the total number of words in the field. j The inverse document frequency of the j-th sensitive word is calculated as follows:

[0040]

[0041] N is the total number of work orders in the training dataset. j S2 represents the number of work orders containing the j-th sensitive word; the calculated S2 needs to be normalized to the [0,1] interval by Min-Max to ensure consistency with other sub-value dimensions.

[0042] Structural Sensitivity Sub-value S3: This reflects the sensitivity of a field to its structural position within a work order segment. The formula is as follows:

[0043] S3 = W str ×I str

[0044] Among them, W str The structural weights of the work order module are set by experts based on the probability that the module contains sensitive information; str As an indicator variable, if the field is located in a preset high-sensitivity module, then I str =1, otherwise I str =0.

[0045] Type Sensitivity Subvalue S4: Used to reflect the sensitivity of the field description type, its value is the preset field type weight W. type , that is, S4=W type The preset field type weight table is divided by the power grid company according to the information security level, for example, "ID number" W type =0.9, "User's mobile phone number" W type =0.85, "Equipment Model" W type =0.2, "Work Order Number" W type =0.1, to ensure that the sensitivity of different types of fields matches the business security requirements.

[0046] The sensitivity threshold is the threshold for determining whether a field is sensitive information. The determination process is as follows: First, collect cases of sensitive information leakage in the power grid industry over the past five years, statistically analyze the leakage risk probability of fields with different sensitivity values, and clarify the correlation that "the higher the sensitivity value, the greater the harm of leakage." Second, a review group composed of security experts, business experts, and technical experts conducts multiple rounds of review using the Delphi method to initially determine the candidate range of the threshold. Finally, substitute the candidate thresholds into the model test set, calculate the false negative rate and false positive rate under different thresholds, and select candidate values ​​with a false negative rate ≤3% and a false positive rate ≤5% as the final threshold. When the final sensitivity value S of a field > 0.6, the field is determined to be sensitive information and output. The information tagging unit is used to tag sensitive information to generate sensitive information annotations. The function of the information tagging unit is to add sensitive information annotations to the sensitive information output by the information identification unit. These annotations contain the attributes and permission association information of the sensitive information, providing a basis for the layered encapsulation of the subsequent encryption encapsulation module. Its core relies on the configuration of the annotation clue library and the annotation generation process.

[0047] The endorsement clue database is a database that stores the correspondence between "sensitive information characteristics and permission characteristics". Its initial construction process is as follows: First, the permission system of power grid business is sorted out, and user permissions are divided into different levels according to the division of responsibilities, clarifying the scope of responsibilities and the types of sensitive information that can be accessed at each level; Second, the expert team extracts the characteristics of various types of sensitive information to form sensitive information feature items; Finally, the correspondence between feature items and permission characteristics is established to form endorsement clues, and all clues are stored in the endorsement clue database.

[0048] The mechanism for updating the endorsement lead database is as follows: feedback on permission matching in actual applications is collected every six months, the expert team evaluates the feedback data, adjusts leads with inaccurate matching, and adds leads based on new business scenarios to ensure that the lead database is synchronized with the needs of the power grid business.

[0049] The specific steps for the information tagging unit to generate sensitive information annotations are as follows: First, receive the sensitive information output by the information identification unit and extract its features. The extracted features include field type, key content information, associated field features, and the module in which it resides. For example, the extracted features for a certain sensitive information might be "Field type = user ID number, content contains 18 digits, associated field = user mobile phone number, module = user information module". Second, match the extracted features with the clues in the annotation clue database using a feature matching algorithm: traverse all clues and determine whether the sensitive information features contain the sensitive information feature items of the clues. If multiple clues match successfully, select... The first step is to select the clue with the most matching feature terms as the optimal matching clue. The second step is to determine the permission features of the sensitive information based on the optimal matching clue. The third step is to generate a sensitive information annotation, which includes: the sensitive ID generated by the UUID algorithm, the sensitive information type, the corresponding permission features, the identification time, and the identification model version number. For example, the generated annotation is "Sensitive ID: f81d4fae-7dec-11d0-a765-00a0c91e6bf6, Type: User Identity Information, Permission Feature: Operation and Maintenance Administrator Level, Identification Time: 2024-10-01 14:30:25, Model Version: V2.1".

[0050] The generated sensitive information annotations will be bound and stored with the corresponding sensitive information. On the one hand, this provides a basis for the subsequent encryption and encapsulation module to encapsulate sensitive information in layers according to permission characteristics, ensuring that users with different permissions can only access the sensitive information at the corresponding level. On the other hand, it provides an identifier for matching sensitive information with the key in the decryption process, ensuring the accuracy and security of the decryption process.

[0051] The missing information compensation module includes a missing information identification unit and a missing information compensation unit.

[0052] The missing information identification unit is used to identify missing positions in the work order file after sensitive information has been removed, and the missing information compensation unit is used to write replacement data to the missing positions. The missing information identification unit is configured with a pre-trained semantic recognition model, which is used to identify whether there are semantic missing positions in the segments after sensitive information extraction and generate corresponding missing positions. The core function of the missing information identification unit is to accurately locate the semantic missing positions in the work order file after sensitive information has been removed. It accomplishes this task by configuring a pre-trained semantic recognition model, which is a deep learning model optimized based on the text features of power grid work orders. It can determine whether there are missing positions in segments from the perspective of semantic coherence and business logic integrity, rather than relying solely on character gaps for judgment. The training of the semantic recognition model needs to be closely adapted to the text style and business logic of power grid work orders to ensure the accuracy of identifying missing expressions specific to the power grid field. The specific steps are as follows: First, a training dataset is constructed by collecting raw work order data from the power grid industry over the past three years, covering multiple scenarios such as equipment operation and maintenance, user repair requests, dispatch instructions, and metering and testing. Each work order data is manually processed: sensitive information is marked and a removal operation is simulated. Then, power grid business experts judge whether there are semantic missing segments after removing sensitive information. If so, the start and end positions of the missing positions and the type of missing information are marked, forming a labeled dataset containing "work order segments after removing sensitive information - missing label results". At the same time, complete work order segments without removing sensitive information are collected as negative samples to ensure that the ratio of positive to negative samples is 1:1, thereby improving the model's ability to distinguish between "with missing / without missing". Secondly, the labeled dataset is preprocessed: a special word segmentation dictionary for the power grid field is used to segment the work order segments to avoid incorrect segmentation of professional terms by general word segmentation tools; the segmented text is converted into word vectors with a dimension of 300 using the Word2Vec tool, and the syntactic features of the segments are extracted and converted into feature vectors. The word vectors and syntactic feature vectors are concatenated and used as the input features of the model, with the dimension uniformly set to 512; the missing label results are encoded, with "no missing" marked as 0, "information completion missing" marked as 1, and "logical connection missing" marked as 2, which are used as the model output labels. Next, a pre-trained Bi LSTM was selected as the base model. A fully connected layer and a Softmax activation function were added to the model output layer to construct a model structure for input features and missing type determination. The preprocessed dataset was divided into training, validation, and test sets in an 8:1:1 ratio. The cross-entropy loss function was used as the model optimization objective, and the AdamW optimizer was used for training. The initial learning rate was set to 3e-5, the batch size was set to 64, and the number of training epochs was set to 50. Every 5 epochs of training, the missing identification accuracy and missing position localization accuracy of the model were calculated on the validation set. When the accuracy and localization accuracy of the validation set did not improve for 3 consecutive epochs, an early stopping strategy was adopted to stop training to avoid model overfitting.Finally, the trained model is validated. On the test set, the model is required to achieve a missing word recognition accuracy of at least 94% and a missing word location accuracy of at least 90%. If these requirements are not met, the model structure is adjusted or supplementary training data is added and the model is retrained until the performance meets the requirements. The input of this semantic recognition model is the feature vector of the work order segment after removing sensitive information, and the output is the missing word determination result and the corresponding missing word location information, providing accurate location information for subsequent missing word compensation. The missing information identification unit performs missing information identification based on the above semantic recognition model. The process is as follows: 1. Receive the work order file output by the information recognition module after removing sensitive information, and split the work order file into several independent segments according to business logic. The splitting is based on the text structure rules preset by the power grid work order to ensure that each segment corresponds to a single business theme and avoid missing information identification errors caused by cross-theme segments; 2. Preprocess each segment after splitting, and generate segment feature vectors according to the standard process during model training; 3. Input the segment feature vectors into the semantic recognition model to obtain the missing information judgment result and missing position information output by the model; 4. Perform secondary verification on the model output result, and confirm the rationality of the missing information judgment result in combination with the power grid business logic rules. If the model judgment result conflicts with the business rules, the missing information mark is corrected according to the business rules. Finally, a missing information list containing segment number, missing position, and missing information type is formed and sent to the missing information compensation unit.

[0053] The missing information compensation unit determines the corresponding load condition alternative data group based on the sensitive information annotations corresponding to the sensitive information, and calculates the completion evaluation value of each alternative data in the alternative data group using a preset missing information evaluation algorithm. The completion evaluation value is used to evaluate the completeness of the segment corresponding to the alternative data, and the alternative data with the highest completion evaluation value is output. The load condition alternative data group refers to a pre-constructed set of non-sensitive alternative data that conforms to the work order business scenario, based on the permission characteristics and sensitive information types associated with the sensitive information annotations. The load conditions specifically refer to the permission adaptation requirements and business adaptation requirements that the alternative data must meet. The construction and matching process of this data group is as follows:

[0054] First, a classification system for sensitive information is established. Based on the characteristics of power grid work orders, sensitive information is divided into three main categories: user privacy, core equipment parameters, and business secrets. Each main category is further subdivided into specific subtypes. Second, for each subtype of sensitive information, alternative data rules that meet the permission requirements are designed based on its corresponding sensitive information annotation permission characteristics, while ensuring that the alternative data conforms to the power grid business scenario. Third, according to the alternative data rules, several specific alternative data entries are generated. Each alternative data entry needs to be labeled with its corresponding sensitive information subtype, permission characteristics, and business scenario tag, forming an initial alternative data pool. Finally, the initial alternative data pool is filtered to remove alternative data that is semantically unreasonable or does not match the business. The data is then grouped according to sensitive information subtype and permission characteristics. Each group is a load condition alternative data group, which is stored in the system database. The data group content is updated quarterly based on new power grid business scenarios, adding alternative data that conforms to the new scenarios.

[0055] After receiving the list of missing bits, the missing bit compensation unit first extracts the sensitive information annotation associated with the original sensitive information corresponding to the missing bit, obtaining the sensitive information subtype and permission characteristics. Secondly, based on the sensitive information subtype and permission characteristics, it searches the system database for the corresponding load condition replacement data group. If a unique matching data group is found, it is directly called. If multiple matching data groups are found, further filtering is performed based on the business scenario tag of the segment containing the missing bit, selecting the data group that perfectly matches the business scenario tag to ensure that the replacement data is adapted to the current business scenario of the work order. The missing bit evaluation algorithm is used to calculate the completion evaluation value of each replacement data in the load condition replacement data group. This evaluation value comprehensively reflects the semantic coherence, business adaptability, and data consistency of the work order segment after the replacement data fills the missing bit, thereby determining the optimal replacement data. Specifically, it includes the algorithm formula, parameter definitions, calculation methods for each dimension, and the optimal replacement data selection process.

[0056] The completion evaluation value is obtained by weighted summation of three dimensions: semantic coherence, business adaptability, and data consistency, as shown in the following formula:

[0057] C = w a ×C a +w b ×C b +w c ×C c

[0058] Where C represents the completion evaluation value of a certain substitute data, and the value ranges from [0,1]. The larger the value, the better the completeness of the segment after the substitute data fills in the missing position; w a w b w c These are semantic coherence weight, business adaptability weight, and data consistency weight, respectively, and satisfy wa +w b +w c =1, this weight is determined using the analytic hierarchy process (AHP): an evaluation group composed of power grid business experts and language processing experts compares the importance of the three dimensions pairwise, constructs a judgment matrix and calculates the weight vector, and determines the final weight after a consistency check. For example, determining w a =0.4, w b =0.35, w c =0.25; C a For semantic coherence scoring, C b For business adaptability scoring, C c For data consistency scoring, all three values ​​range from [0,1].

[0059] Semantic coherence score C a This is used to evaluate the semantic coherence between the substitute data and the text before and after the missing position. It is achieved by calculating the cosine similarity between the substitute data and the text before and after it, as shown in the following formula:

[0060]

[0061] in, The word vectors used as replacement data (generated using the Word2Vec tool during the training of the semantic recognition model, with a dimension of 300); The concatenated word vectors of the text before and after the missing position are obtained by taking the average of the word vectors of the 10 characters before and after the missing position. (If the preceding and following characters are less than 10, take the average of the word vectors of all characters); cos(·) is the cosine similarity function, and the calculation result is directly used as C. a The value ranges from [0,1], and the larger the value, the better the semantic coherence.

[0062] Business adaptability score: C b This is used to evaluate the matching degree between alternative data and the work order business scenario. A rule-based matching scoring method is employed, with the following steps: First, for the business scenario of the segment containing the missing part, 5-8 business adaptation rules are preset; second, it is checked whether the alternative data meets each rule, awarding 1 point for each rule met and 0 points for each rule not met; finally, the ratio of the score to the total number of rules is calculated, which is C. b The formula is as follows:

[0063]

[0064] Where k is the number of business adaptation rules that the alternative data satisfies, and t is the total number of preset business adaptation rules, with a value range of [0,1]. The larger the value, the better the business adaptability.

[0065] Data consistency score C cThis is used to evaluate the logical consistency between substitute data and other non-sensitive information in the work order, avoiding data contradictions. The calculation steps are as follows: First, extract non-sensitive reference information related to missing bits in the work order; second, construct consistency judgment rules; third, determine whether the substitute data and reference information conform to all consistency rules, scoring 1 point for each rule conformed and 0 points for each rule not conformed; finally, calculate the ratio of the score to the total number of rules, which is C. c The formula is as follows:

[0066]

[0067] Where m is the number of consistency rules that the substitute data satisfies, and n is the total number of preset consistency rules, with a value range of [0,1]. The larger the value, the better the data consistency.

[0068] The missing data compensation unit selects the optimal replacement data according to the following steps: 1. Traverse each replacement data in the load condition replacement data group and calculate the completion evaluation value C for each data according to the above formula; 2. Sort all replacement data in descending order of C value and select the top 3 candidate replacement data with the highest C value. If there are less than 3 replacement data in the data group, all of them are considered as candidates; 3. Manually review the candidate replacement data to check for any hidden problems not covered by the formula; 4. If there are no problems after review, select the replacement data with the highest C value as the final replacement data; if the candidate data with the highest C value has hidden problems, check the next candidate data in turn until the replacement data without problems is determined; 5. Write the final replacement data into the missing position located by the missing identification unit to complete the work order missing compensation, generate a complete work order file with desensitization and completion, and send it synchronously to the encryption and encapsulation module and the data interaction module to prepare for subsequent encryption and transfer.

[0069] The encryption encapsulation module includes a sensitive encryption unit, a layer encapsulation unit, and a layer encryption unit:

[0070] The sensitive encryption unit is configured with a first encryption strategy to encrypt sensitive information to generate sensitive information ciphertext and a corresponding sensitive information key. The first encryption strategy adopts the AES-256-GCM symmetric encryption algorithm, which combines high security and efficiency, supports data integrity verification during the encryption process, and is suitable for the encryption requirements of sensitive information in power grid work orders. The specific implementation steps are as follows: The sensitive encryption unit receives the sensitive information and corresponding sensitive information annotation output by the information identification module, and simultaneously sends an encryption factor acquisition request to the central server, carrying the type identifier of the sensitive information in the request; The central server retrieves the corresponding encryption factor from the encryption parameter library according to the sensitive information type identifier, and transmits the encryption factor encrypted to the sensitive encryption unit; The sensitive encryption unit generates a sensitive information key K1 and an initialization vector IV based on the acquired encryption factor: where K1 is obtained through a seed key S. seedThe unique identifier for sensitive information is derived using the PBKDF2 algorithm, with the following formula:

[0071] K1 = PBKDF2 - HMAC - SHA256(S seed ID sen ,iter,32)

[0072] In the formula, PBKDF2-HMAC-SHA256 is a key derivation function based on SHA256 hash, iter is the number of iterations, 32 is the output key length, and IV is a 12-byte random number generated according to the initial vector generation rules provided by the central server to ensure that the IV is unique for each encryption.

[0073] The AES-256-GCM algorithm is used to encrypt the plaintext P of sensitive information, while the sensitive information signature A is included as additional authentication data in the encryption process to generate the ciphertext C of sensitive information. sen With certification label T sen The formula is:

[0074] (C sen ,T sen =AES-256-GCM K1,IV (P,A)

[0075] In the formula, AES-256-GCM K1,IV This indicates an AES-256-GCM encryption operation using key K1 and initialization vector IV, T sen The authentication tag is used to verify the integrity and authenticity of the ciphertext during subsequent decryption, preventing ciphertext tampering; the sensitive encryption unit ciphers the sensitive information ciphertext C. sen Certification label T sen It is associated with the sensitive information key K1 and stored, with the association identifier being the sensitive ID. sen Simultaneously embedding IV into C sen The header is used to facilitate extraction during subsequent decryption.

[0076] The layer encapsulation unit divides the sensitive information ciphertext into layers based on sensitive information annotations, and encapsulates each layer containing the sensitive information ciphertext to generate a layered ciphertext encapsulation dataset. The core basis of the layering logic is the permission feature in the sensitive information annotation: the layer encapsulation unit receives the "ID" output by the sensitive encryption unit. sen -C sen -T sen "-K1" associates the dataset, via ID. sen Retrieve the corresponding sensitive information endorsements and extract the permission features (Perm) from the endorsements. sen A hierarchical access control system for sensitive information in power grid work orders is pre-defined, with the system defined by power grid security experts based on their business responsibilities; and based on access characteristics (Perm).sen Encrypt sensitive information C sen Classify into the corresponding level: for example, Perm sen When the administrator level is C, sen and the corresponding T sen ID sen Sensitive information is categorized into different levels based on their access permissions. If a sensitive information tag contains multiple access permissions, it is categorized into a higher level based on the highest access permission level. The encrypted data at each level is then aggregated to form initial data blocks for each level. Each initial data block contains the C of all sensitive information at that level. sen T sen ID sen And the corresponding sensitive information annotation summary.

[0077] The layer encryption unit is configured with a second encryption strategy for encrypting the layer ciphertext encapsulated dataset to generate layer encrypted ciphertext and a layer encryption key, wherein the layer encryption key includes an index subkey and a matching subkey.

[0078] The second encryption strategy employs a hybrid encryption scheme combining the RSA-2048 asymmetric encryption algorithm with the AES-128-GCM symmetric encryption algorithm. RSA-2048 is used to encrypt the AES-128-GCM key, while AES-128-GCM is used to encrypt the ciphertext encapsulated dataset. This scheme balances the key distribution security of asymmetric encryption with the data encryption efficiency of symmetric encryption: the layer encryption unit receives the ciphertext encapsulated dataset D. encap First, generate the layer encryption master key K2. The generation method is as follows: generate a 16-byte random number using CSPRNG as the raw data for K2, then perform SHA-256 hash calculation and take the first 16 bytes to obtain the final K2. The formula is:

[0079] K2=trunc 16 (SHA-256(R 16 ))

[0080] In the formula, R 16 trunc is a 16-byte cryptographically secure random number. 16 (·) indicates that the first 16 bytes of the hash result are taken;

[0081] Generate the initial vector IV2 (12 bytes, generated via CSPRNG) required for AES-128-GCM, and apply the AES-128-GCM algorithm to D. encap Encryption is performed to generate layer-encrypted ciphertext C. layer With certification label T layer The formula is:

[0082]

[0083] In the formula, This indicates an AES-128-GCM encryption operation using key K2 and initialization vector IV2, T layer This is a 16-byte authentication tag used to verify C during subsequent decryption. layer To ensure integrity; an RSA public key acquisition request is sent to the central server, which then encapsulates the hierarchical permission features in the dataset using layered ciphertext. layer Retrieve the RSA public key Pub with the corresponding permissions perm and Pub perm Encrypted transmission to the layer encryption unit; RSA-2048 algorithm applied to Pub. perm Encrypt K2 to generate the encrypted layer encryption master key. The formula is:

[0084]

[0085] In the formula, Indicates the use of the public key Pub perm RSA-2048 encryption operation;

[0086] Embed IV2 in C layer The head, will T layer Embedded C layer The tail end forms the final layer of encrypted ciphertext. At the same time and Associative storage, with the association identifier being the layer identifier. ID The layer encryption key consists of the index subkey K. index Matching subkey K match Both are composed of layered encrypted ciphertext. Binding: Generate index subkey K index : Use layer as an identifier ID Encapsulated timestamp time encap and 32-byte random number R 32 The input is used to generate the hash using the SHA-256 hash algorithm. The formula is:

[0087] K index =SHA-256(layer) ID ||time encap ||R 32 )

[0088] In the formula, || represents the string concatenation operation, and R 32 Generated via CSPRNG, ensuring even layer ID With time encap Same, K index Still unique; K indexIts function is to enable the indexed combustion unit to access K in the subsequent combustion decryption module. index Quickly match the corresponding Avoid traversing all layers of encrypted ciphertext;

[0089] The second step is to generate the matching subkey K. match : Encrypt the master key of the encrypted layer Perm with hierarchical permission features layer The hash values ​​are concatenated and then generated using SHA-256 hash calculation. The formula is:

[0090]

[0091] In the formula, K match Its purpose is to be used by the Perm decryption unit when verifying user permissions in subsequent layers. layer Decrypting the matched RSA private key and through K match Verify the validity of the decrypted K2 to prevent decryption with an illegal key;

[0092] Establish the association between the layer encryption key and the layer encryption ciphertext: K index K match , layer ID Perm layer The data is then aggregated to form a "layered encrypted data packet," and the associated information of this data packet is uploaded to the central server for record-keeping, facilitating subsequent user permission verification and key tracing.

[0093] The layer encryption unit associates the "layer encrypted data packet" with the desensitized work order output by the missing compensation module. The association identifier is the unique ID of the work order. Finally, it outputs the combined data of "desensitized work order - layer encrypted data packet" to the data interaction module, which then handles the transmission. At the same time, it encrypts and stores the sensitive information key K1 in the encryption parameter library of the central server, which will be authorized to call upon subsequent requests from the combustion decryption module.

[0094] The combustion decryption module includes an index combustion unit, a layer decryption unit, and a sensitive decryption unit:

[0095] The index burning unit matches each layer of encrypted ciphertext using an index subkey and deletes other layers of encrypted ciphertext based on the unique match result; it also receives layer-encrypted data packets output by the data interaction module, which contain the layer-encrypted ciphertext. Layer encryption key and layer identifier ID Meanwhile, the mobile device automatically collects its own hardware identifier and packages it together with the user's identity information into matching request data;

[0096] The matching request data is encrypted and transmitted to the central server. The central server first verifies the legitimacy of the user's identity information and device hardware identifier: it compares the user account's current permissions with the hierarchical permission features (Perm) in the encrypted data packet. layer The system checks for a match and verifies whether the device hardware identifier is in the "Authorized Device List" registered with the central server. If both checks pass, it retrieves the index subkey registration hash value corresponding to the encrypted data packet at that layer from the encryption parameter library. Kindex备案 And compare the verification result with the hash Kindex备案 Encryption feedback is sent to the index combustion unit;

[0097] After receiving feedback, the index burning unit first verifies whether the verification result is successful. If it fails, the operation is terminated and a permission or device unauthorized prompt is displayed to the user. If it succeeds, the local index subkey K is calculated. index hash value Kindex本地 The integrity of the index subkey is confirmed by hash comparison.

[0098] If the index subkey verification passes, the index burning unit traverses all layers of encrypted ciphertext in local storage, extracts the index subkey corresponding to each group of ciphertext, calculates its hash value, and compares it with the hash value. Kindex本地 A one-to-one comparison is performed to find the uniquely matching layer-encrypted ciphertext; all non-target layer-encrypted ciphertexts and their corresponding layer-encrypted keys are deleted, and only the target layer-encrypted ciphertext and the matching layer-encrypted key are retained. At the same time, a burning operation log is generated and uploaded to the central server for record-keeping to ensure that the operation is traceable.

[0099] The layer decryption unit is used to decrypt the layer encrypted ciphertext by matching the subkey and the second decryption strategy to generate sensitive information ciphertext. The core function of the layer decryption unit is to decrypt the target layer encrypted ciphertext to obtain the layer ciphertext encapsulated dataset by matching the subkey with the decryption credential issued by the central server based on the second decryption strategy. The second decryption strategy is the reverse operation of the second encryption strategy in the encryption encapsulation module. It must strictly follow the "permission-key" matching principle to ensure that only authorized users can decrypt the corresponding layer data.

[0100] The second decryption strategy corresponds to the decryption process of the RSA-2048+AES-128-GCM hybrid encryption scheme in the encryption encapsulation module: the layer decryption unit extracts the matching subkey K from the layer encryption key. match Combined with the hierarchical permission features corresponding to the target layer encrypted ciphertext, Perm layer Generate a "decryption request" containing K match Perm layer User identity information and device hardware identifiers are encrypted and transmitted to the central server; after receiving the "decryption request," the central server first parses the K... match :according to The generation rules, reverse verification of K match With Perm layer The correlation, i.e., the calculation in The encrypted master key uploaded by the encryption encapsulation module, if the result is different from K match If they match, then K is confirmed. match Valid; the central server is based on Perm. layer Retrieve the corresponding RSA private key Priv perm and through the security channel to Priv perm The data is sent to the layer decryption unit; the layer decryption unit applies the RSA-2048 algorithm and uses Priv... perm Decrypted layer encryption master key The formula for obtaining the original layer encryption master key K2 is:

[0101]

[0102] In the formula, This indicates the use of the private key Priv. perm RSA-2048 decryption operation, To extract the encrypted master key from the encrypted data packet, the decrypted K2 needs to be calculated using SHA-256 hash and compared with the K2 filing hash value issued by the central server to confirm that K2 has not been tampered with.

[0103] Extract target layer encrypted ciphertext The initial vector IV2 and the authentication tag T layer The remaining part is the core encrypted data C. layer核心 Applying the AES-128-GCM algorithm, using K2 and IV2 to decrypt C layer核心 At the same time, a decrypted authentication tag is generated. The formula is:

[0104]

[0105] In the formula, This indicates an AES-128-GCM decryption operation using key K2 and initialization vector IV2, D encap The decrypted layered ciphertext is used to encapsulate the dataset; With the extracted T layer Perform a byte-by-byte comparison; if they match, then confirm D. encap If the encrypted text is complete and unaltered, it is output to the sensitive decryption unit; if it is inconsistent, it is determined that the encrypted text has been tampered with, the decryption and burning unit is immediately triggered, and the central server is reported at the same time.

[0106] The sensitive decryption unit is used to decrypt the sensitive information ciphertext using the sensitive information key and the first decryption strategy to obtain the corresponding sensitive information; the first decryption strategy corresponds to the decryption process of the AES-256-GCM symmetric encryption algorithm in the encryption encapsulation module: the sensitive decryption unit decrypts the dataset D from the layered ciphertext encapsulation. encap Extracting associated data containing sensitive target information, including the encrypted sensitive information C. sen Certification label T sen Sensitive ID sen And sensitive information is labeled A, generating a sensitive decryption request, which includes ID. sen User identity information and A are encrypted and transmitted to the central server;

[0107] After receiving a sensitive decryption request, the central server first verifies whether the user's identity information matches the permission characteristics corresponding to the sensitive information annotation A. If they match, it retrieves the corresponding ID from the encryption parameter library. sen The associated sensitive information key K1 is sent to the sensitive decryption unit through a secure channel; the sensitive information ciphertext C is extracted. sen The initialization vector IV is in the middle, and the remaining part is the core sensitive encrypted data C. sen核心 Applying the AES-256-GCM algorithm, using K1 and IV to decrypt C sen核心 At the same time, a decrypted authentication tag is generated. The formula is:

[0108]

[0109] In the formula, This indicates an AES-256-GCM decryption operation using key K1 and initialization vector IV, where A is the sensitive information annotation and P is the original plaintext sensitive information obtained after decryption. With the extracted T sen A comparison is performed. If they match, P is confirmed to be complete and valid, and is ready to be output to the user. If they do not match, it is determined that the sensitive information has been tampered with, and the decryption and burning unit is immediately triggered and reported to the central server. The business validity of the plaintext sensitive information P is verified by logically comparing P with the associated non-sensitive information in the desensitized work order. Only after confirming that there is no logical contradiction can the subsequent output stage be entered.

[0110] The combustion decryption module also includes a decryption combustion unit. The decryption combustion unit is configured with combustion triggering conditions. The decryption combustion unit determines whether the combustion triggering conditions are met based on the decryption feedback information of the second decryption strategy. If the combustion triggering conditions are met, the current layer encryption ciphertext and the corresponding layer encryption key are deleted.

[0111] The core function of the decryption and burning unit is to proactively delete the current layer's encrypted ciphertext and corresponding layer encryption key when preset burning trigger conditions are met, achieving zero data residue after decryption. The burning trigger conditions are not publicly known terms; they are defined as specific events or states that trigger the data destruction operation after decryption, including decryption errors, operation timeouts, and unauthorized access. This aims to prevent sensitive data from being stored in the device for extended periods, potentially leading to leakage. The burning trigger conditions must cover various risk scenarios during the decryption process. The specific definitions and judgment methods are as follows:

[0112] Decryption error trigger: This refers to the occurrence of a preset threshold N of consecutive errors such as authentication tag verification failure or key comparison failure during layer decryption or sensitive decryption. The determination method is: the decryption combustion unit records the number of decryption errors in real time. error When count error When N ≥ N, the triggering condition is met; Operation timeout trigger: refers to the period after the sensitive information is decrypted, during which the user has not performed any legal operation for a preset timeout period T. The determination method is: start counting time from the moment decryption is completed. idle When time idle When T ≥ T, the triggering condition is deemed met. Unauthorized access trigger: This refers to the mobile device detecting unauthorized operation behavior, including inconsistencies between the device hardware identifier and the registration, failed user biometric verification, or the device connecting to an untrusted network. The determination method is as follows: The decryption and burning unit receives real-time security status feedback from the mobile device. If the feedback information contains any type of unauthorized behavior identifier, the triggering condition is deemed met. When the decryption and burning unit determines that any burning triggering condition is met, it immediately performs the following operations: 1. Lock the access permissions of the current layer encrypted ciphertext and the layer encryption key, prohibiting any read and write operations to prevent data from being illegally obtained during the destruction process; 2. Delete the target layer encrypted ciphertext stored locally. Layer encryption key, containing K index With K match The sensitive information key K1 is deleted using multiple overwrites to ensure the data cannot be recovered by data recovery tools; 3. The encrypted dataset D is deleted. encap 4. Generate a burning operation record, including the trigger condition type, burning time, list of deleted data, device security status, etc., and encrypt and upload it to the central server for filing. The central server will synchronously update the decryption status of the work order to "burned"; 5. Output a prompt to the user that the data has been safely destroyed, and lock the decryption function of the mobile device for 30 minutes to prevent malicious cracking attempts.

[0113] It also includes a central server configured with an encryption parameter library containing several encryption factors and corresponding decryption factors. The central server updates and sends the corresponding encryption and decryption factors to the user terminal according to permissions. Mobile devices must pre-complete the following security configurations and register with the central server: bind the device's unique hardware identifier to the user account; only registered devices can receive layered encrypted data packets; the device must integrate at least one biometric module and encrypt and store the user's biometric template in a local security chip, prohibiting its export; install a customized power grid-specific operating system, close unnecessary ports and services, and enable real-time virus protection and intrusion detection functions; configure an independent hardware encryption chip to store the RSA private key Priv issued by the central server. perm The sensitive information key K1 is stored entirely outside the operating system, preventing software-level theft. Each user terminal is configured with corresponding permission information, which is updated in real time and uploaded to the central server. The central server updates the corresponding encryption and decryption factors based on the permission information. The combustion decryption module is configured in the mobile device, which is configured with unique permission conditions. When the decryption feedback information of the second decryption strategy meets the unique permission conditions, the corresponding sensitive information is allowed to be output. After the sensitive decryption unit completes the decryption and verification of the sensitive information plaintext P, it must first pass the unique permission condition verification before outputting P. The process is as follows: The mobile device automatically initiates a permission verification request, requiring the user to complete biometric identification, and simultaneously collects the current device's hardware identifier, network connection status, and operating system security status to generate decryption feedback information; the burning decryption module verifies the decryption feedback information: verifying whether the biometric identification result is consistent with the local template, whether the device hardware identifier is consistent with the registration, whether the network connection is a trusted network, and whether the operating system has no security vulnerabilities; if all verification items pass, it is determined that the decryption feedback information meets the unique permission condition, allowing the output of the sensitive information plaintext P to the mobile device's authorization interface; if any verification item fails, it is determined that the permission condition is not met, the decryption burning unit is immediately triggered, and the central server is notified and the device is locked.

[0114] The data interaction module includes a data sending unit and a data receiving unit.

[0115] The data sending unit is used to send the de-identified work order file and the corresponding layered encrypted dataset. Before data sending, data association, integrity verification, and receiver permission verification must be completed to ensure that the sent data is legal and meets the requirements of the receiving end. The specific steps are as follows: The data sending unit receives the de-identified work order output by the missing data compensation module, denoted as D. desen The layered encrypted data packet output by the encryption encapsulation module is denoted as D. layer包 First, assign a unique data interaction identifier ID to both parties. interThis identifier is a 32-bit string, generated by hashing the work order's unique ID, the sending timestamp, and an 8-bit random number using SHA-256. The formula is as follows:

[0116] ID inter =substr(SHA-256(WO ID ||time send ||R8),0,32)

[0117] In the formula, WO ID The time is the unique ID for the work order. send For sending timestamps, R8 is an 8-bit cryptographically secure random number, and substr(·,0,32) represents taking the first 32 characters of the hash result, ID. inter Its function is to ensure D desen With D layer包 The connection is maintained throughout the process to avoid separation that could lead to mismatches in subsequent decryption.

[0118] Data integrity verification calculates D separately desen With D layer包 The integrity hash value is used by the receiving end to verify whether the data has been tampered with. The formula is:

[0119] hash desen =SHA-256(D desen )

[0120] hash layer包 =SHA-256(D layer包 )

[0121] In the formula, hash desen The hash value represents the integrity of the work order after it has been anonymized. layer包 The integrity hash value of the layered encrypted data packet, and the ID. inter Packaged together as a verification information packet D check The data is then uploaded to the central server for record-keeping, forming the basis for subsequent verification by the receiving end.

[0122] The data sending unit obtains the identification information of the target receiving end, such as the receiving end device hardware ID Dev. ID Receive user account User ID Generate an authorization verification request, which includes the ID. inter Dev ID User ID and D layer包 Perm hierarchical permission features layer The encrypted data is transmitted to the central server; the central server then transmits the data according to the User's... ID Query its current permissions, based on Dev IDCheck if the device is in the authorized device list, and also confirm user permissions and Perm. layer The system checks whether the user is a mid-level administrator and can only receive data packets with intermediate access privileges. If all three verifications pass, the system sends a verification success message and the encrypted communication address of the receiving end to the data sending unit. If any verification fails, the system terminates transmission and prompts the receiving end that it lacks the necessary permissions. Data transmission employs a transmission session encryption + data segmentation transmission method to ensure transmission link security and stability for large data volume transmissions. The specific steps are as follows: Establishing an encrypted transmission session: The data sending unit initiates a session establishment request to the receiving end based on the encrypted communication address of the receiving end fed back by the central server. The request carries the sending end's device certificate. After verifying the certificate's validity, the receiving end generates a transmission session key K. session K is encrypted using the public key in the sending device certificate via the RSA-2048 algorithm. session The data is fed back to the data sending unit; the data sending unit uses its own private key to decrypt and obtain K. session Both parties are based on K session Establish an encrypted transmission channel with the TLS 1.3 protocol;

[0123] If D desen With D layer包 If the total size exceeds a preset threshold, it will be segmented: D will be divided into 1MB units. desen Split into D desen1 D desen2 ,...,D desenn D layer包 Split into D layer1 D layer2 ,...,D layerm Each segment is assigned a segment identifier and segment sequence number, and a hash value for each segment is calculated to form a segment verification table. The data sending unit transmits each segment sequentially according to its segment sequence number, and each segment is verified through K... session Encryption involves transmitting data packets containing segment identifiers, segment sequence numbers, encrypted segment data, and segment hash values; the receiving end, upon receiving each segment, first uses K... session The data transmission unit decrypts the data segment, calculates its hash value, and compares it with the hash value of the segment in the transmitted data packet. If they match, the segment reception is confirmed as complete, and a successful segment reception is reported to the sender. If they do not match, a retransmission of the segment is requested. After all segments have been transmitted, the data transmission unit sends a transmission completion command to the receiver and simultaneously sends a verification packet D. check The receiving end is based on D check hash in desen With hash layer包 Calculate the spliced ​​D respectively desen With D layer包The hash value is used for global integrity verification; if the verification passes, the receiving end sends a successful transmission confirmation message to the sending end, and the data sending unit uploads the transmission result to the central server for record-keeping; if the verification fails, a retransmission process is triggered.

[0124] The data receiving unit is used to receive work order files and layered encrypted datasets, and generate corresponding layered encryption keys and sensitive information keys. The core function of the data receiving unit is to receive the de-identified work orders and layered encrypted data packets transmitted from the sender, complete the legality verification, generate the corresponding layered encryption keys and sensitive information keys, and securely store the data and keys to support the subsequent burning and decryption module. Its implementation needs to cover the entire process of pre-reception verification, in-reception verification, and post-reception processing.

[0125] After data reception is complete, global integrity verification, key generation, and secure storage must be performed. The specific steps are as follows:

[0126] The receiving end splices all the segmented data to recover the complete D. desen With D layer包 Extract the "verification information packet" D transmitted by the sending end. check Get hash desen With hash layer包 ; Calculate D after recovery respectively desen hash value desen接收 With D layer包 hash value layer包接收 Perform global validation:

[0127] The data receiving unit needs to generate data with D. layer包 The matching layer encryption key contains the index subkey K index接收 Matching subkey K match接收 and sensitive information key K 1接收 The generation process requires access to the encryption factor of the central server, as detailed below:

[0128] First, a request to obtain the encryption factor is sent to the central server. The request includes the ID. inter With Perm layer The central server uses the ID as the basis for its operation. inter Retrieve the layer identifier of the filing ID With encapsulated timestamp time encap Send this information and a 32-byte random number R to the receiving end. 32接收 The receiver generates K. index接收 The formula is the same as that of the encryption encapsulation module:

[0129] K index接收 =SHA-256(layer) ID ||time encap ||R32接收 )

[0130] Next, the receiving end from D layer包 Extract the encrypted layer encryption master key. Calculate SHA-256 (Perm) layer Then, K is generated using the following formula. match接收 :

[0131]

[0132] Sensitive Information Key K 1接收 Generation: The receiving end sends a sensitive information key factor request to the central server, carrying the ID. inter With sensitive IDID sen From D layer包 Extracted from; the central server based on ID. sen Retrieve the seed key S from the registration seed接收 The data is sent to the receiving end; the receiving end generates K using the PBKDF2 algorithm. 1接收 The formula is the same as that of the encryption and encapsulation module:

[0133] K 1接收 =PBKDF2-HMAC-SHA256(S seed接收 ID sen ,iter,32)

[0134] In the formula, iter is still taken 10,000 times to ensure K 1接收 It is consistent with the K1 format generated by the encryption and encapsulation module and can be used for subsequent decryption of sensitive information;

[0135] The receiving end will D desen With D layer包 An encrypted partition stored on a mobile device, pre-stored using a K-code. session Encrypt again; put K index接收 K match接收 With K 1接收 A separate key area stored in a hardware encryption chip is prohibited from being exported to the operating system level; simultaneously, a data-key association table is established to record IDs. inter With K index接收 K match接收 K 1接收 The corresponding relationship facilitates the use of the combustion decryption module; after storage, the data receiving unit reports the completion of reception and processing to the central server, and the central server updates the ID. inter The data transfer status is "received".

[0136] Taking a power supply failure repair order in a residential community as an example, this system enables the secure transfer of anonymized information via WeChat. The information identification module uses a sensitive information identification model to calculate the sensitivity values ​​(including weighted sums of four sub-values ​​such as association and content) of the "user's mobile phone number" and "home address" fields in the work order. If the values ​​exceed a preset threshold, the information is marked as sensitive information, and a sensitive information annotation with permission features is generated.

[0137] The missing information compensation module uses a semantic recognition model to locate the missing parts after removing sensitive information. It then selects the highest-scoring completion value from the alternative data group with appropriate permissions, such as "****6789" and "XX neighborhood area," and writes them to ensure the semantic integrity of the work order. The encryption and encapsulation module uses the AES-256-GCM algorithm to encrypt the original sensitive information. After layered encapsulation according to signatures, it generates layered encrypted ciphertext, index subkeys, and matching subkeys using a hybrid strategy of RSA-2048 + AES-128-GCM.

[0138] The data interaction module calls the WeChat API to send the de-identified work order and layered encrypted data packet to the maintenance personnel's WeChat account. The maintenance personnel, through a WeChat mini-program with a built-in combustion decryption module, completes WeChat real-name authentication and enterprise permission verification (the only permitted condition). They then use the index subkey to match the target ciphertext and remove redundancy. By matching the subkey and employing a dual decryption strategy, they decrypt layer by layer to obtain the sensitive information. The decryption combustion unit then automatically deletes the ciphertext and key, achieving secure data transfer within the WeChat platform.

[0139] Of course, the above are just typical examples of the present invention. In addition, the present invention may have many other specific embodiments. All technical solutions formed by equivalent substitution or equivalent transformation fall within the scope of protection claimed by the present invention.

Claims

1. A de-identified information conversion system based on work order flow, characterized in that: include: The information identification module includes an information identification unit and an information tagging unit. The information identification unit is used to identify sensitive information in the work order file, and the information tagging unit is used to tag the sensitive information to generate sensitive information annotations. The missing information compensation module includes a missing information identification unit and a missing information compensation unit. The missing information identification unit is used to identify missing positions in the work order file after removing sensitive information, and the missing information compensation unit is used to write replacement data to the missing positions. The encryption encapsulation module includes a sensitive encryption unit, a layer encapsulation unit, and a layer encryption unit; The sensitive encryption unit is configured with a first encryption strategy to encrypt sensitive information to generate sensitive information ciphertext and a corresponding sensitive information key. The layer encapsulation unit divides the sensitive information ciphertext into layers according to the sensitive information annotation and encapsulates each layer containing the sensitive information ciphertext to generate a layer ciphertext encapsulation dataset. The layer encryption unit is configured with a second encryption strategy to encrypt the layer ciphertext encapsulation dataset to generate layer encrypted ciphertext and a layer encryption key. The layer encryption key includes an index subkey and a matching subkey. The burning and decryption module includes an index burning unit, a layer decryption unit, and a sensitive decryption unit. The index burning unit matches each layer of encrypted ciphertext with an index subkey and deletes other layers of encrypted ciphertext based on the unique matching result. The layer decryption unit is used to decrypt the layer of encrypted ciphertext using the matching subkey and a second decryption strategy to generate sensitive information ciphertext. The sensitive decryption unit is used to decrypt the sensitive information ciphertext using a sensitive information key and a first decryption strategy to obtain the corresponding sensitive information. The data interaction module includes a data sending unit and a data receiving unit. The data sending unit is used to send the de-identified work order file and the corresponding layered encrypted dataset. The data receiving unit is used to receive the work order file and the layered encrypted dataset, and generate the corresponding layered encryption key and sensitive information key.

2. The de-identified information conversion system based on work order flow as described in claim 1, characterized in that: The information identification unit is equipped with a sensitive information identification model, which is used to identify sensitive information in the information identification unit. The sensitive information identification model is also equipped with a multi-dimensional information evaluation algorithm, which is used to calculate the sensitivity value of each field. When the sensitivity value is greater than a preset sensitivity value benchmark, the field is output as sensitive information.

3. The de-identified information conversion system based on work order flow as described in claim 2, characterized in that: The multidimensional evaluation algorithm is configured to weight different types of sensitivity sub-values ​​to obtain the sensitivity value. The types of sensitivity sub-values ​​include: association sensitivity sub-values, content sensitivity sub-values, structure sensitivity sub-values, and type sensitivity sub-values. The association sensitivity sub-value reflects the sensitivity of the field corresponding to the fields with association relationships. The content sensitivity sub-value reflects the sensitivity of the field itself. The structure sensitivity sub-value reflects the sensitivity of the field in the structure of the segment. The type sensitivity sub-value reflects the sensitivity of the description type corresponding to the field.

4. The de-identified information conversion system based on work order flow as described in claim 1, characterized in that: The information tag unit is configured with a signature clue database, which stores a number of signature clues. Each signature clue corresponds to a permission feature. The information tag unit identifies the permission feature corresponding to sensitive information based on the signature clue to generate the sensitive information signature.

5. The de-identified information conversion system based on work order flow as described in claim 1, characterized in that: The missing information identification unit is equipped with a pre-trained semantic recognition model, which is used to identify whether there are semantic missing information in the segment after sensitive information is extracted and to generate the corresponding missing position.

6. The de-identified information conversion system based on work order flow as described in claim 5, characterized in that: The missing information compensation unit determines the corresponding load condition replacement data group based on the sensitive information annotation corresponding to the sensitive information, and calculates the completion evaluation value of each replacement data in the replacement data group through a preset missing information evaluation algorithm. The completion evaluation value is used to evaluate the completeness of the segment corresponding to the replacement data, and the replacement data with the highest completion evaluation value is output.

7. The de-identified information conversion system based on work order flow as described in claim 1, characterized in that: It also includes a central server, which is configured with an encryption parameter library. The encryption parameter library is configured with several encryption factors and corresponding decryption factors. The central server updates the corresponding encryption factors and decryption factors according to permissions and sends them to the user terminal.

8. The de-identified information conversion system based on work order flow as described in claim 7, characterized in that: Each user terminal is configured with permission information. The user terminal updates the permission information in real time and uploads it to the central server. The central server updates the corresponding encryption factor and decryption factor according to the permission information.

9. The de-identified information conversion system based on work order flow as described in claim 1, characterized in that: The combustion decryption module is configured in a mobile device, which is configured with a unique permission condition. When the decryption feedback information of the second decryption strategy meets the unique permission condition, the corresponding sensitive information is allowed to be output.

10. The de-identified information conversion system based on work order flow as described in claim 1, characterized in that: The combustion decryption module also includes a decryption combustion unit. The decryption combustion unit is configured with combustion triggering conditions. The decryption combustion unit determines whether the combustion triggering conditions are met based on the decryption feedback information of the second decryption strategy. If the combustion triggering conditions are met, the current layer encryption ciphertext and the corresponding layer encryption key are deleted.