An archive information processing method and system based on digital security
By combining a dynamically iterative tagging system template with rule matching and manual review, sensitive tags in archival information are identified and processed differently, solving the problem of insufficient tag recognition accuracy in existing technologies and achieving efficient and secure archival information processing.
Patent Information
- Application Number
- CN202511642874.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-11
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2045-11-11
AI Technical Summary
Existing technologies for secure archival information processing rely on statically preset tag systems that lack dynamic iteration capabilities, resulting in insufficient accuracy and an inability to adapt to changes in business scenarios. Tag recognition depends on single rule matching or manual review, which is inefficient. Furthermore, the desensitization methods are homogenized and cannot simultaneously ensure data integrity and security.
It adopts a hierarchical logic that combines rule matching and manual review. Through the learning of historical samples and the dynamic iteration of new case verification, a preset label system template is used to identify sensitive and non-sensitive labels, perform differentiated desensitization processing, and then perform double verification after integration.
It has achieved long-term accuracy and improved processing efficiency in sensitive label identification, avoiding missed and false judgments, ensuring data security and availability, adapting to compliance requirements of different business scenarios, and reducing labor costs.
Smart Images

Figure CN121118113B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of digital data processing, and in particular to an archive information processing method and system based on digital security. BACKGROUND
[0002] In recent years, with the deepening of big data and digital transformation, archive management has accelerated into the digital stage, and massive traditional archives have been converted into digital information, significantly improving storage, transmission and utilization efficiency. However, in the digital environment, the sensitive information including identity identification and technical secrets in the archives faces the risk of leakage, which poses a challenge to data security. Digital data processing technology becomes the key to accurately identifying sensitive information and implementing security protection while ensuring the security of archive information and taking into account the utilization value, which is the core support for the standard development of archive digitization.
[0003] The prior art has obvious deficiencies in archive information security processing. The label system is mostly static preset, lacking dynamic iteration ability based on historical sample learning and new case verification, and being difficult to adapt to changing business scenarios and compliance requirements. Label identification relies on single rule matching or manual review, which has insufficient accuracy or low efficiency, and cannot realize efficient identification under hierarchical logic. The desensitization processing method is homogeneous, and no differentiated strategy is adopted for different sensitive label types, which easily leads to excessive desensitization affecting data usability or insufficient desensitization causing security risks. Moreover, there is a lack of systematic verification mechanism for processed data, making it difficult to balance data integrity and security. SUMMARY
[0004] The technical problem solved by the present application is that the prior art has obvious deficiencies in archive information security processing. The label system is mostly static preset, lacking dynamic iteration ability based on historical sample learning and new case verification, and being difficult to adapt to changing business scenarios and compliance requirements. Label identification relies on single rule matching or manual review, which has insufficient accuracy or low efficiency, and cannot realize efficient identification under hierarchical logic. The desensitization processing method is homogeneous, and no differentiated strategy is adopted for different sensitive label types, which easily leads to excessive desensitization affecting data usability or insufficient desensitization causing security risks. Moreover, there is a lack of systematic verification mechanism for processed data, making it difficult to balance data integrity and security.
[0005] To solve the above technical problems, the present application provides the following technical solution: an archive information processing method based on digital security, comprising the following steps:
[0006] Step S100, acquiring first metadata of archive information;
[0007] Step S200, based on a preset label system template, using rule matching and artificial auditing combined with hierarchical logic, sensitive label and non-sensitive label identification are performed on the first metadata, and the preset label system template is dynamically iterated through historical sample learning and new label case verification;
[0008] Step S300, for the first metadata corresponding to the sensitive label, differential de-sensitization processing is performed according to the sensitive label, and the first metadata corresponding to the non-sensitive label is not de-sensitized.
[0009] Step S400, the first metadata corresponding to the sensitive label after de-sensitization processing and the first metadata corresponding to the non-sensitive label without de-sensitization processing are integrated to obtain second metadata.
[0010] As a preferred scheme of the file information processing method based on digital security, wherein: the first metadata includes field identification data, content description data, management attribute data and associated subject data of the file information;
[0011] The sensitive label includes an identity sensitive label, a technical secret sensitive label, a commercial data sensitive label and a privacy information sensitive label;
[0012] The non-sensitive label includes an open information non-sensitive label, a statistical summary non-sensitive label, an expired secret non-sensitive label and a public matter non-sensitive label;
[0013] The corresponding relationship between the sensitive label and the non-sensitive label and the first metadata is as follows:
[0014] The identity sensitive label corresponds to the associated subject data;
[0015] The technical secret sensitive label corresponds to the content description data;
[0016] The commercial data sensitive label corresponds to the field identification data and the content description data;
[0017] The privacy information sensitive label corresponds to the content description data;
[0018] The open information non-sensitive label corresponds to the content description data and the management attribute data;
[0019] The statistical summary non-sensitive label corresponds to the content description data;
[0020] The expired secret non-sensitive label corresponds to the management attribute data;
[0021] The public matter non-sensitive label corresponds to the management attribute data.
[0022] As a preferred scheme of the digital security-based archive information processing method, the construction mode of the preset label system template comprises:
[0023] The preset label system template comprises classification rules and a corresponding common feature library;
[0024] The classification rules comprise classification rules of sensitive labels and classification rules of non-sensitive labels, wherein the classification rule judgment logic of the sensitive labels comprises field features corresponding to the sensitive labels and content ranges corresponding to the sensitive labels, and the classification rules of the non-sensitive labels comprise exclusion conditions of the non-sensitive labels;
[0025] The field features corresponding to the sensitive labels comprise field names, data types and field levels corresponding to identity identification sensitive labels, technical secret sensitive labels, commercial data sensitive labels and privacy information sensitive labels respectively, wherein the field name corresponding to the identity identification sensitive labels is a unique subject identification type name, the data type is an individual identity code type, and the field level is a core association field;
[0026] The field name corresponding to the technical secret sensitive label is a technical parameter name, the data type is a technical document type, and the field level is a core technology field;
[0027] The field name corresponding to the commercial data sensitive label is a commercial information name, the data type is a commercial transaction or asset type, and the field level is a core business field;
[0028] The field name corresponding to the privacy information sensitive label is a personal privacy name, the data type is a private information record type, and the field level is a privacy association field;
[0029] The content range corresponding to the sensitive label comprises information content boundaries corresponding to the identity identification sensitive label, the technical secret sensitive label, the commercial data sensitive label and the privacy information sensitive label respectively, wherein the information content boundary corresponding to the identity identification sensitive label is an information range that can uniquely identify the identity of the associated subject;
[0030] The information content boundary corresponding to the technical secret sensitive label is an undisclosed core technology information range;
[0031] The information content boundary corresponding to the commercial data sensitive label is a commercial information range with commercial value and not disclosed;
[0032] The information content boundary corresponding to the privacy information sensitive label is a protected personal private information range;
[0033] The exclusion condition is specifically:
[0034] The first metadata excluding the field feature corresponding to the sensitive label, the first metadata falling into the content range corresponding to the sensitive label, and the first metadata having a sensitive management attribute associated feature, wherein the sensitive management attribute associated feature includes an attribute feature of a first metadata secret level being secret or above, a storage period limit being permanent confidentiality, and an access permission being limited to core personnel only;
[0035] The samples of the docking history metadata are manually reviewed to label the sensitive labels and the non-sensitive labels in the samples to obtain labeled samples, the common features corresponding to the sensitive labels and the non-sensitive labels in the labeled samples are extracted, a preset common feature library is constructed, the common features are stored in feature items of corresponding label attributes, the common feature library and the classification rule are cooperated to provide feature matching basis for label identification of the first metadata, and the common features are supplemented to the classification rule;
[0036] According to a preset period, new label cases generated in an archive information processing process and not covered by a current classification rule are collected, the new label cases include first metadata content, associated fields and business scenario information;
[0037] The new label cases are verified by a preset audit subject to determine label attributes of the new label cases, the label attributes include sensitive labels or non-sensitive labels;
[0038] And first metadata features corresponding to the label attributes are extracted, the first metadata features include field identification features, content keyword features and management attribute associated features.
[0039] As a preferred scheme of the archive information processing method based on digital security, wherein: based on the verified label attributes and the first metadata features, the determination logic of the classification rule is updated, and the feature items of the corresponding label attributes in the common feature library are supplemented;
[0040] The updated classification rule and the common features are integrated into a preset label system template, template iteration is realized by replacing an old version template or forming a version superposition, the preset label system template is adapted to a business scenario and compliance requirements corresponding to the new label cases, so that long-term accuracy of sensitive label and non-sensitive label identification in the first metadata is maintained;
[0041] And the iterated preset label system template is applied to the sensitive label and the non-sensitive label identification process of the first metadata in real time.
[0042] As a preferred scheme of the archive information processing method based on digital security, wherein: a layered logic combining rule matching and manual review is used to identify sensitive labels and non-sensitive labels of the first metadata, including:
[0043] obtaining the first metadata, calling the preset label system template, the preset label system template including the classification rule of the sensitive label and the classification rule of the non-sensitive label and the corresponding common feature library;
[0044] structurally splitting the first metadata according to the field structure to obtain a plurality of independent metadata units, the metadata units corresponding to the field identifier data segment, the content description data segment, the management attribute data segment or the associated subject data segment in the first metadata;
[0045] for each of the metadata units, extracting the features corresponding to the metadata units;
[0046] the features include field identification features, content keyword features and management attribute association features,
[0047] wherein the field identification features include field name, data type and field level;
[0048] the content keyword features include core expression vocabulary, data format identifier and semantic direction;
[0049] the management attribute association features include data security level, storage period and access permission identifier.
[0050] As a preferred scheme of the file information processing method based on digital security, wherein: the features of the metadata units are matched with the classification rule of the sensitive label in the preset label system template, specifically including:
[0051] if the field identification features match the field features corresponding to the sensitive label, and the content keyword features fall within the content range corresponding to the sensitive label, the metadata unit is marked as a sensitive label;
[0052] for the metadata units not marked as sensitive labels, first screen according to the exclusion conditions of non-sensitive labels, if the exclusion conditions are not triggered, the features corresponding to the metadata units are compared with the common feature library corresponding to the non-sensitive labels, if the matching degree exceeds the preset threshold, the metadata unit is marked as a non-sensitive label;
[0053] for the metadata units that neither match the sensitive label nor match the non-sensitive label, mark as a to-be-confirmed label, and generate an audit request including the features of the metadata units and the associated business scenarios, and send to a preset audit terminal;
[0054] receive the label attribute confirmation result returned by the audit terminal, update the to-be-confirmed label to a sensitive label or a non-sensitive label, and summarize the label identification results of all metadata units to form a label list of the first metadata;
[0055] The label identification result includes sensitive labels and non-sensitive labels corresponding to the metadata unit and label sources, the label sources include field identification feature matching and content keyword feature matching of sensitive labels, exclusion condition screening and threshold comparison of common feature library of non-sensitive labels, and audit terminal confirmation of labels to be confirmed.
[0056] As a preferred scheme of the file information processing method based on digital security, wherein: the desensitization processing of the sensitive label corresponding to the first metadata includes:
[0057] If the sensitive label is an identity sensitive label, the identity sensitive label corresponding to the first metadata is associated subject data, and an irreversible feature coding mode in a biological recognition technology is used for desensitization and original format preservation, and the steps are as follows:
[0058] Extract the field identification feature, content keyword feature and management attribute association feature of the associated subject data.
[0059] The field identification feature includes the format rule and length parameter of the identity identification field.
[0060] The content keyword feature includes the core coding segment of the identity identification.
[0061] The management attribute association feature includes the privacy protection level corresponding to the identity identification.
[0062] Based on a preset random projection matrix or a nonlinear transformation algorithm, the extracted features are mapped into low-dimensional irreversible feature codes.
[0063] The conditions that the feature code needs to meet include that the length parameter is consistent with the length parameter of the original associated subject data, and the data type and field structure of the original associated subject data match the format rule of the original associated subject data.
[0064] By adjusting the dimension parameter of the random projection matrix or the complexity parameter of the nonlinear transformation, the irreversible mapping strength of the feature code and the original associated subject data is controlled, and the irreversible mapping strength is positively correlated with the privacy protection level of the associated subject data.
[0065] As a preferred scheme of the file information processing method based on digital security, wherein: the irreversible feature code is generated to replace the sensitive content of the original associated subject data, and the conditions that the feature code needs to meet include that the feature code difference degree of different original associated subject data and the original data difference degree are in a preset proportion, and the irreversibility is verified to ensure that the original associated subject data content cannot be inversely deduced through the feature code.
[0066] If the sensitive label is a technical secret sensitive label, a commercial data sensitive label or a privacy information sensitive label, the sensitive label corresponds to first metadata, which is content description data or field identification data, and the content is processed by adopting a desensitization mode adapted to the data type, and the first metadata original format is retained;
[0067] The desensitization mode adapted to the data type refers to a text desensitization mode for content description data and a field desensitization mode for field identification data.
[0068] As a preferred scheme of the file information processing method based on digital security, for the non-sensitive label corresponding first metadata, no processing is performed, including:
[0069] If the non-sensitive label is a public information non-sensitive label, a statistical summary non-sensitive label, an expired secret non-sensitive label or a public matter non-sensitive label, the non-sensitive label corresponds to first metadata, which is content description data or management attribute data, and the content and format of the first metadata are directly retained without any desensitization processing;
[0070] The first metadata corresponding to the sensitive label after processing and the first metadata corresponding to the non-sensitive label are integrated to obtain second metadata, and the second metadata needs to pass through integrity verification and security verification to ensure that it can be directly used for subsequent file information query, statistics and sharing scenarios.
[0071] The file information processing system based on digital security comprises an acquisition module, an identification module, a processing module and an integration module.
[0072] The acquisition module acquires first metadata corresponding to file information.
[0073] The identification module identifies sensitive labels and non-sensitive labels of the first metadata based on a preset label system template by adopting a hierarchical logic combining rule matching and manual auditing, and the preset label system template is dynamically iterated through historical sample learning and new label case verification.
[0074] The processing module performs differential desensitization processing on the first metadata corresponding to the sensitive label according to the sensitive label, and does not perform desensitization processing on the first metadata corresponding to the non-sensitive label.
[0075] The integration module integrates the first metadata corresponding to the sensitive label after desensitization processing and the first metadata corresponding to the non-sensitive label without desensitization processing to obtain second metadata.
[0076] The application has the beneficial effects that: in the three dimensions of information security, processing efficiency and value preservation, a dynamic iteration label system is formed, the long-term accuracy of sensitive and non-sensitive label identification is greatly improved by learning supplementary features from historical samples and updating rules by combining new cases, the problems of missed judgment and misjudgment caused by rule solidification are effectively avoided, the hierarchical identification logic combines automatic feature matching and manual review, which not only significantly improves the processing efficiency and reduces the labor cost, but also complements the blind area of automatic identification, and stably guarantees the label identification quality, the differentiated desensitization strategy is very targeted, and special processing schemes are matched for different types of sensitive data, which can not only build a strong information security line, but also avoid the data value loss caused by excessive desensitization, in addition, the integrated metadata can be directly used in query, statistics and sharing scenes after double verification, seamlessly connecting the subsequent application of digital archives, and truly realizing the safe and efficient circulation of archive resources. BRIEF DESCRIPTION OF DRAWINGS
[0077] Figure 1 A step flow diagram of an archive information processing method based on digital security provided by an embodiment of the application is provided.
[0078] Figure 2 A basic flow diagram of an archive information processing system based on digital security provided by an embodiment of the application is provided. DETAILED DESCRIPTION
[0079] In order to make the above-mentioned purposes, features and advantages of the application more obvious and easy to understand, the specific embodiments of the application will be described in detail below with reference to the drawings of the specification. Obviously, the described embodiments are part of the embodiments of the application, rather than all the embodiments.
[0080] Embodiment 1, refer to Figure 1 An archive information processing method based on digital security is provided by an embodiment of the application, including the following steps:
[0081] Step S100, acquiring first metadata of archive information;
[0082] Step S200, based on a preset label system template, using a hierarchical logic combining rule matching and manual review to identify sensitive labels and non-sensitive labels for the first metadata, and the preset label system template is dynamically iterated through historical sample learning and new label case verification;
[0083] Step S300, for the first metadata corresponding to the sensitive label, differentiating desensitization processing according to the sensitive label, and not desensitizing the first metadata corresponding to the non-sensitive label;
[0084] Step S400, the desensitization processing of the sensitive label corresponding to the first metadata is integrated with the non-sensitive label corresponding to the first metadata without desensitization processing, and the second metadata is obtained.
[0085] In one of the embodiments, first, the step S100 is performed to obtain the first metadata corresponding to the archive information, the first metadata including the field identification data, the content description data, the management attribute data and the associated subject data of the archive, then the step S200 is entered to identify the sensitive label and the non-sensitive label based on the preset label system template, wherein the preset label system template is dynamically iterated through historical sample learning and new label case verification, the identification process adopts the hierarchical logic combining rule matching and manual review, then the step S300 is performed to differentially desensitize the first metadata corresponding to the sensitive label according to the label type, such as using irreversible feature coding method to desensitize the associated subject data corresponding to the sensitive label to prevent original information leakage, and the first metadata corresponding to the non-sensitive label is not desensitized, and the original content and format are directly retained, finally, the step S400 is performed to integrate the desensitized metadata corresponding to the sensitive label and the non-desensitized metadata corresponding to the non-sensitive label, and the second metadata which can be directly used for subsequent archive query, data statistics and information sharing is obtained. Through accurate label identification and differential desensitization, the digital security of sensitive content in the archive is effectively guaranteed, and the usability of non-sensitive information is maximally retained. The dynamically iterated preset label system template can adapt to changes in different business scenarios and compliance requirements, long-term maintain the accuracy and flexibility of archive information processing, and improve the overall efficiency of digital archive management.
[0086] The first metadata includes the field identification data, the content description data, the management attribute data and the associated subject data of the archive information;
[0087] The sensitive label includes an identity identification sensitive label, a technical secret sensitive label, a commercial data sensitive label and a privacy information sensitive label;
[0088] The non-sensitive label includes a public information non-sensitive label, a statistical summary non-sensitive label, an expired secret non-sensitive label and a public notice non-sensitive label;
[0089] The corresponding relationship between the sensitive label and the non-sensitive label and the first metadata is as follows:
[0090] The identity identification sensitive label corresponds to the associated subject data;
[0091] The technical secret sensitive label corresponds to the content description data;
[0092] The commercial data sensitive label corresponds to the field identification data and the content description data;
[0093] The privacy information sensitive label corresponds to the content description data;
[0094] The public information non-sensitive label corresponds to the content description data and the management attribute data;
[0095] The statistical summary non-sensitive label corresponds to the content description data;
[0096] The expired secret non-sensitive label corresponds to the management attribute data;
[0097] The public matter non-sensitive label corresponds to the management attribute data.
[0098] The construction mode of the preset label system template comprises:
[0099] The preset label system template comprises classification rules and corresponding common feature library;
[0100] The classification rules comprise sensitive label classification rules and non-sensitive label classification rules, wherein the sensitive label classification rule judgment logic comprises a field feature corresponding to the sensitive label and a content range corresponding to the sensitive label, and the non-sensitive label classification rule comprises an exclusion condition of the non-sensitive label;
[0101] The field feature corresponding to the sensitive label comprises a field name, a data type and a field level corresponding to an identity sensitive label, a technical secret sensitive label, a commercial data sensitive label and a privacy information sensitive label respectively, wherein the field name corresponding to the identity sensitive label is an associated subject unique identifier type, the data type is an individual identity code type, and the field level is a core associated field;
[0102] The field name corresponding to the technical secret sensitive label is a technical parameter name, the data type is a technical document type, and the field level is a core technical field;
[0103] The field name corresponding to the commercial data sensitive label is a commercial information name, the data type is a commercial transaction or asset type, and the field level is a core business field;
[0104] The field name corresponding to the privacy information sensitive label is a personal privacy name, the data type is a private information record type, and the field level is a privacy associated field;
[0105] The content range corresponding to the sensitive label comprises an information content boundary corresponding to the identity sensitive label, the technical secret sensitive label, the commercial data sensitive label and the privacy information sensitive label respectively, wherein the information content boundary corresponding to the identity sensitive label is an information range that can uniquely identify the identity of the associated subject;
[0106] The information content boundary corresponding to the technical secret sensitive label is an undisclosed core technical information range;
[0107] The information content boundary corresponding to the commercial data sensitive label is the range of commercial information with commercial value and not disclosed;
[0108] The information content boundary corresponding to the privacy information sensitive label is the range of protected personal privacy information;
[0109] The exclusion condition is specifically:
[0110] The first metadata excluding the field feature corresponding to the sensitive label, the first metadata falling into the content range corresponding to the sensitive label, and the first metadata having the sensitive management attribute associated feature, wherein the sensitive management attribute associated feature includes the attribute features of the first metadata classification level being secret or above, the storage period being permanently secret, and the access permission being limited to core personnel only;
[0111] The sample of the historical metadata is connected, the sensitive label and the non-sensitive label in the sample are labeled by manual review to obtain the labeled sample, the common features corresponding to the sensitive label and the non-sensitive label in the labeled sample are extracted, a preset common feature library is constructed, the common features are stored in the feature items of the corresponding label attributes, the common feature library and the classification rule are cooperated to provide the feature matching basis for the label identification of the first metadata, and the common features are supplemented to the classification rule;
[0112] According to a preset period, new label cases generated in the process of processing the archive information and not covered by the current classification rule are collected, the new label cases include the first metadata content, the associated field and the business scenario information;
[0113] The new label cases are verified by a preset audit subject to determine the label attributes of the new label cases, and the label attributes include sensitive labels or non-sensitive labels;
[0114] And the first metadata features corresponding to the label attributes are extracted, the first metadata features include the field identification feature, the content keyword feature and the management attribute associated feature.
[0115] In one of the embodiments, the composition of the first metadata includes field identification data (such as customer transaction serial number, product research and development parameter number) of the archives, content description data (such as product core research and development scheme, company public recruitment information), management attribute data (such as secret level: secret, storage period: 5 years, secret level: public, storage period: 3 years) and associated subject data (such as employee ID number, customer mobile number), and at the same time, the label types are divided, the sensitive labels include identity sensitive label, technical secret sensitive label, commercial data sensitive label and privacy information sensitive label, the non-sensitive labels include public information non-sensitive label, statistical summary non-sensitive label, expired secret non-sensitive label and public matter non-sensitive label, and the correspondence between the sensitive labels and the non-sensitive labels and the first metadata is clear, the identity sensitive label corresponds to the associated subject data (such as employee ID number), the technical secret sensitive label corresponds to the content description data (such as product core research and development scheme), the commercial data sensitive label corresponds to the field identification data (such as customer transaction serial number) and the content description data (such as quarterly revenue details), the privacy information sensitive label corresponds to the content description data (such as customer health record abstract), the public information non-sensitive label corresponds to the content description data (such as company public recruitment information) and the management attribute data (such as secret level: public file attribute), the statistical summary non-sensitive label corresponds to the content description data (such as annual customer quantity statistical total), the expired secret non-sensitive label corresponds to the management attribute data (such as 2010 secret file, current secret period expired storage attribute), and the public matter non-sensitive label corresponds to the management attribute data (such as public period: 2024.1-2024.2 public attribute).
[0116] The construction of the preset label system template is based on classification rules and corresponding common feature library: the classification rules include sensitive label and non-sensitive label rules, wherein the sensitive label classification rule judgment logic covers field features and content range, the field features need to match the field name, data type and field level of each sensitive label, for example, the field name of the identity sensitive label is the employee ID number, the associated subject unique identifier class name of the customer mobile phone number, the data type is individual identity code type (such as 18-digit ID number, 11-digit mobile phone number), and the field level is the core associated field, the field name of the technical secret sensitive label is the product R&D parameter, the device core drawing number, the data type is technical document type, and the field level is the core technology field, and the content range clearly defines the information boundary of each sensitive label, such as the identity sensitive label corresponding to the information that can uniquely identify the associated subject (such as the complete ID number), and the technical secret sensitive label corresponding to the core technical information that has not been disclosed (such as the product chip design scheme that has not been released), and the classification rule of the non-sensitive label is the exclusion condition, which needs to exclude the first metadata including the sensitive label field feature, falling into the sensitive label content range and having the sensitive management attribute associated feature (such as secret and above, permanent secrecy storage period, and core personnel access permission), for example, the product document with secret level: confidential needs to be excluded from the non-sensitive label. In the construction process, the historical metadata samples (such as the archive metadata in the past 3 years) are connected first, the sample labels are marked by manual review (such as marking the employee ID number as an identity sensitive label, and marking the company public annual report data as public information non-sensitive label), and then the common features of the marked samples are extracted (such as the common features of the identity sensitive label including the field name including ID number and mobile phone number, and the data type being 18 / 11 code), and the common feature library is constructed according to the features, and the features are supplemented to the classification rules, and then new label cases (such as customer biometric information, including metadata content customer fingerprint template, associated field biometric field and business scenario customer identity verification) that are not covered by the current classification rules are collected according to the preset period (such as once a month), the label attributes of the cases are verified by the preset audit subject (such as the archive compliance audit group) (such as determining that the customer biometric information is a privacy information sensitive label), and the corresponding first metadata features (such as the field name including biometric, the data type being template file type, and the management attribute not being public) are extracted, and finally the classification rules and the common feature library are integrated and updated, and the construction of the preset label system template is completed.
[0117] Based on the verified and confirmed label attributes and first metadata features, the judgment logic of the classification rules is updated, and the feature items corresponding to the label attributes in the common feature library are supplemented;
[0118] The updated classification rule and common characteristics are integrated into a preset label system template, the template iteration is realized by replacing the old version template or forming version superposition, the preset label system template is adapted to the business scene and compliance requirements corresponding to the new label case, so that the long-term accuracy of identifying sensitive labels and non-sensitive labels in the first metadata is maintained;
[0119] And the preset label system template after iteration is applied to the sensitive label and non-sensitive label identification process of the first metadata in real time.
[0120] In one of the embodiments, after the verification of the new label case is completed, the classification rules and common feature library of the preset label system template are first updated based on the label attributes and corresponding first metadata features confirmed by the verification. For example, for the new label case of customer biometric information, if the preset audit subject has verified that the label attribute is a privacy information sensitive label, and the extracted first metadata features are that the field name includes biometric identification, the data type is a template file type, and the management attribute has no public identifier, then the field feature judgment condition of the field name including biometric identification and the data type being a template file type needs to be supplemented in the classification rule judgment logic of the privacy information sensitive label, and the management attribute without public identifier is added as a common feature in the common feature library of the privacy information sensitive label, so that the classification rules and the feature library can cover the identification requirements of the new case. Subsequently, the updated classification rules and common features are integrated into the preset label system template, and the iteration mode is selected according to the actual application requirements: if the new case only involves partial feature supplementation (such as adding a certain type of privacy information field feature), the old version of the template is replaced with the updated template to directly cover the original template; if the new case involves business scenario expansion (such as adding a medical industry dedicated sensitive label), the version stacking method is used to add a template version suitable for the medical scenario based on the old version of the template, and the version number (such as V2.1-medical scenario version) is marked for tracing and switching. Through the two iteration methods, the template can adapt to the business scenarios (such as financial industry biometric identification scenario and medical industry privacy data scenario) and compliance requirements (such as the protection of biometric information in the Personal Information Protection Law) of the new label case. After the iteration is completed, the updated preset label system template is applied to the sensitive label and non-sensitive label identification process of the first metadata in real time. For example, after the updated template is put into operation, when the system obtains the first metadata of the customer iris template, it can immediately call the biometric identification related field features and common features supplemented in the new template to automatically identify it as a privacy information sensitive label without manual intervention or secondary configuration of the template. Through the process of feature-driven update-flexible integration iteration-real-time application, the misjudgment of the first metadata in the new business scenario (such as preventing the customer biometric information from being mislabeled as a non-sensitive label) is avoided, and the new compliance requirements can be adapted without redeveloping the entire label identification system, which significantly reduces the system maintenance cost. The real-time application mechanism ensures that the new data processing can enjoy the benefits of template iteration, maintaining the long-term accuracy and efficiency of label identification in the processing of archival information.
[0121] The hierarchical logic combining rule matching and manual audit is used to identify sensitive labels and non-sensitive labels from the first metadata, including:
[0122] The first metadata is obtained, and the preset label system template is called. The preset label system template includes the classification rules of sensitive labels and non-sensitive labels and the corresponding common feature library.
[0123] The first metadata is structurally split by field structure to obtain a plurality of independent metadata units, and the metadata units correspond to field identification data segments, content description data segments, management attribute data segments or associated subject data segments in the first metadata;
[0124] For each metadata unit, the features corresponding to the metadata unit are extracted;
[0125] The features include field identification features, content keyword features, and management attribute association features,
[0126] The field identification features include field name, data type and field level;
[0127] The content keyword features include core expression vocabulary, data format identification and semantic direction;
[0128] The management attribute association features include data security level, storage period and access permission identification.
[0129] The features of the metadata unit are matched with the classification rules of the sensitive labels in the preset label system template, specifically including:
[0130] If the field identification features match the field features corresponding to the sensitive labels, and the content keyword features fall within the content range corresponding to the sensitive labels, the metadata unit is marked as a sensitive label;
[0131] For metadata units that are not marked as sensitive labels, first screen according to the exclusion conditions of non-sensitive labels, if the exclusion conditions are not triggered, the features corresponding to the metadata units are compared with the common feature library corresponding to the non-sensitive labels, if the matching degree exceeds the preset threshold, the metadata unit is marked as a non-sensitive label;
[0132] For metadata units that do not match sensitive labels or non-sensitive labels, mark them as pending labels, and generate an audit request including the features of the metadata units and the associated business scenarios, and send it to the preset audit terminal;
[0133] Receive the label attribute confirmation result returned by the audit terminal, update the pending label to a sensitive label or a non-sensitive label, and summarize the label identification results of all metadata units to form a label list of the first metadata;
[0134] The label identification result package includes the sensitive labels and non-sensitive labels corresponding to the metadata units and the label sources, and the label sources include the field identification feature matching and the content keyword feature matching of the sensitive labels, the exclusion condition screening and the threshold comparison of the common feature library of the non-sensitive labels, and the audit terminal confirmation of the pending label.
[0135] In one of the embodiments, first, the first metadata to be processed (such as the employee profile metadata of an enterprise, including the employee ID number: 110101XXXXXX001234, the product non-disclosure chip design parameters: XX model chip power consumption control scheme, and the annual customer total statistics: 1.2 million people accumulated in 2024) is acquired, and a preset label system template is synchronously called, which has included the classification rules of sensitive labels (such as identity and technical secret) and non-sensitive labels (such as public information and statistical summary), and the corresponding common feature library (such as the common features of the identity sensitive label, including the field name including the ID number, the mobile phone number, and the data type of 18 / 11 coding). Then, the acquired first metadata is structured and split according to the field structure, and is disassembled into several independent metadata units, each of which corresponds to a type of data segment in the first metadata, for example, the employee ID number: 110101XXXXXX001234 (associated subject data segment), product non-disclosure chip design parameters: XX model chip power consumption control scheme (content description data segment), and annual customer total statistics: 1.2 million people accumulated in 2024 (content description data segment) three metadata units. The corresponding features are extracted for each metadata unit, including the field identification feature, the content keyword feature, and the management attribute association feature: taking the employee ID number metadata unit as an example, the field identification feature is the field name: employee ID number, the data type: 18 coding type, and the field level: core associated field, the content keyword feature is the core expression vocabulary: ID number, the data format identification: 18 digits + letter combination, and the semantic direction: individual identity identification, and the management attribute association feature is the data security level: secret, the storage period: permanent, and the access permission identification: only for the personnel department. The corresponding features of the metadata units are matched with the classification rules of the sensitive labels in the preset label system template: if the field identification feature of a certain metadata unit matches the field feature of a sensitive label, and the content keyword feature falls within the content range of the sensitive label, it is marked as a sensitive label, for example, the product non-disclosure chip design parameter metadata unit, the field identification feature matches the field name: technical parameter name, and the data type: technical document type of the technical secret sensitive label, and the content keyword feature (core word non-disclosure chip design parameter, semantic direction core technology) falls within the range of non-disclosure core technology information of the technical secret sensitive label, so it is marked as a technical secret sensitive label.For metadata units not marked as sensitive labels, first screen according to the exclusion conditions of non-sensitive labels (exclusion includes sensitive field characteristics, falling into sensitive content range, units with sensitive management attribute association characteristics), if the exclusion conditions are not triggered, compare their characteristics with the common feature library of non-sensitive labels, and if the matching degree exceeds the preset threshold (such as 80%), mark it as a non-sensitive label. For example, the annual customer total statistics metadata unit does not trigger the exclusion condition (no sensitive field, the content is statistical total, and the management attribute secret level is public), the matching degree of its characteristics (content keyword customer total statistics, semantic pointing data summary) with the common feature library of statistical summary non-sensitive labels reaches 90%, so it is marked as a statistical summary non-sensitive label. For metadata units that do not match sensitive labels and non-sensitive labels (such as customer consumption preference analysis report: including customer purchase category distribution but no specific identity information), mark it as a to-be-confirmed label, and generate an audit request including the characteristics of the unit (field name: customer consumption preference analysis, data type: report document type, management attribute secret level: internal public), and the associated business scenario (customer operation analysis), and send it to the preset audit terminal (such as the archive compliance audit group). After receiving the confirmation result returned by the audit terminal (such as determining that it is a non-sensitive label), update the to-be-confirmed label to the corresponding label attribute. Finally, summarize the label identification results of all metadata units to form a label list of the first metadata. The list needs to indicate the label of each metadata unit (such as sensitive label-technical secret, non-sensitive label-statistical summary) and the source of the label (such as sensitive label from field identification feature matching + content keyword feature matching, non-sensitive label from exclusion condition screening + common feature library threshold comparison, to-be-confirmed label from audit terminal confirmation).
[0136] The desensitization processing of the first metadata corresponding to the sensitive label includes:
[0137] If the sensitive label is an identity identification sensitive label, the first metadata corresponding to the identity identification sensitive label is associated subject data, and the irreversible feature coding method in the biometric identification technology is used for desensitization and the original format is retained. The steps are as follows:
[0138] Extract the field identification feature, content keyword feature and management attribute association feature of the associated subject data;
[0139] Among them, the field identification feature includes the format rule and length parameter of the identity identification field;
[0140] The content keyword feature includes the core coding segment of the identity identification;
[0141] The management attribute association feature includes the privacy protection level corresponding to the identity identification;
[0142] Based on the preset random projection matrix or nonlinear transformation algorithm, the extracted features are mapped to low-dimensional irreversible feature codes;
[0143] The conditions that the feature code needs to meet include that the length parameter is consistent with the length parameter of the original associated subject data, and the data type and field structure of the original associated subject data match the format rule of the original associated subject data;
[0144] By adjusting the dimension parameter of the random projection matrix or the complexity parameter of the nonlinear transformation, the strength of the irreversible mapping of the feature code and the original associated subject data is controlled, and the strength of the irreversible mapping is positively correlated with the privacy protection level of the associated subject data.
[0145] The generated irreversible feature code replaces the sensitive content of the original associated subject data, and the feature code meets the conditions including that the difference degree of the feature codes of different original associated subject data and the original data is in a preset proportion, and the original associated subject data content cannot be inversely deduced through the feature code by ensuring the irreversibility.
[0146] If the sensitive label is a technical secret sensitive label, a commercial data sensitive label or a privacy information sensitive label, the sensitive label corresponds to the first metadata being content description data or field identification data, the content is processed by a desensitization method adapted to the data type, and the first metadata original format is retained.
[0147] Among them, the desensitization method adapted to the data type refers to adopting a text desensitization method for content description data and a field desensitization method for field identification data.
[0148] In one of the embodiments, first, for the case of sensitive label for identity identification sensitive label, the label corresponds to the first metadata for the associated subject data (such as employee identity card number 110101XXXXXX001234, customer mobile phone number 138XXXX5678), the irreversible feature coding method in biometric identification technology is used for desensitization, the steps are: first, extract the field identification feature, content keyword feature and management attribute association feature of the associated subject data, take the employee identity card number as an example, the field identification feature includes the format rule (18 digit combination, the first 6 digits are administrative region code, the middle 8 digits are birth date code, and the last 4 digits are check code) and length parameter (fixed 18 digits) of identity identification field, the content keyword feature includes the core coding segment (middle 8 digits of birth date code XXXXXX) of identity identification, and the management attribute association feature includes the privacy protection level (high level, which needs the highest level of desensitization protection) corresponding to the identity identification. Then, based on the preset random projection matrix or nonlinear transformation algorithm (such as using 300-dimensional random projection matrix), the extracted three types of features are mapped into low-dimensional irreversible feature code, and the feature code needs to meet the conditions of consistent length parameter with the original associated subject data (such as identity card number feature code is still 18 digits), data type (numeric type) and field structure (segmented format) with the original format rule, at the same time, the irreversible mapping strength is controlled by adjusting the algorithm parameter, if the privacy protection level of the associated subject data is high (such as identity card number), the dimension parameter of the random projection matrix is increased (such as from 300 dimensions to 500 dimensions) or the complexity of the nonlinear transformation is improved (such as using multilayer perception transformation), so that the mapping strength is positively related to the privacy protection level, finally, the generated irreversible feature code is used to replace the sensitive content of the original associated subject data (such as replacing 110101XXXXXX001234 with 110101582173001234), and the feature code needs to meet the condition that the difference degree of different original data feature codes is a preset proportion of the difference degree of the original data (such as the last 4 digits of the original identity card number are different, and the difference degree of the last 4 digits of the feature code is maintained above 80%), and the irreversibility is verified (such as using SHA-256 hash algorithm to verify that the original identity card number content cannot be inversely deduced from the feature code).For sensitive labels that are technical secret sensitive labels, commercial data sensitive labels, or privacy information sensitive labels, these labels correspond to the first metadata, which is content description data (such as XX model chip power consumption control scheme, customer health record summary) or field identification data (such as customer transaction serial number: 202405010001, technical parameter number: JS2024001), and is processed using a desensitization method adapted to the data type and retains the original format: For content description data, a text desensitization method is used, such as the XX model chip power consumption control scheme (content description data) corresponding to the technical secret sensitive label, which replaces the undisclosed core parameters (such as static power consumption ≤ 5mA) with static power consumption ≤ mA”, retains the original format of the document paragraph structure and title level, and for field identification data, a field desensitization method is used, such as the customer transaction serial number: 202405010001 (field identification data) corresponding to the commercial data sensitive label, which retains the first 8-bit date code 20240501 and replaces the last 4-bit serial code with ****, maintains the total length of the field consistent with the original format, and through the differential design of the desensitization scheme according to the type of sensitive label, it ensures the absolute safety of identity identification sensitive data (irreversible feature coding prevents original information leakage), and realizes precise desensitization for technical secret and commercial data sensitive information (only hides the core sensitive content, retains the availability of non-sensitive parts), while all desensitization processing retains the original format of the metadata, avoiding functional failure due to format abnormalities during subsequent archive query and statistics, and balancing the security of sensitive information and the practicality of archive information, adapting to the security management needs of digital archives in multiple scenarios of enterprises and government affairs.
[0149] For non-sensitive labels, the first metadata is not processed, including:
[0150] If the non-sensitive label is a public information non-sensitive label, a statistical summary non-sensitive label, an expired secret non-sensitive label, or a public matter non-sensitive label, the first metadata corresponding to the non-sensitive label is content description data or management attribute data, and the content and format of the first metadata are directly retained without any desensitization processing.
[0151] The first metadata corresponding to the processed sensitive label and the first metadata corresponding to the non-sensitive label are integrated to obtain the second metadata, and the second metadata needs to pass integrity verification and security verification to ensure that it can be directly used in subsequent archive information query, statistics, and sharing scenarios.
[0152] In one of the embodiments, for the case that the non-sensitive tags are public information non-sensitive tags, statistical summary non-sensitive tags, expired confidential non-sensitive tags or public affairs non-sensitive tags, the first metadata corresponding to these tags are only content description data or management attribute data, without any desensitization processing, directly retaining the original content and format, for example, the company's 2024 public recruitment notice (content description data) corresponding to the public information non-sensitive tag, retaining the text content, paragraph structure and attachment format of the full text of the notice, the 2024 annual customer total statistics: a total of 12,000 people (content description data) corresponding to the statistical summary non-sensitive tag, retaining the original expression of statistical value, unit and statistical period, the 2010 product confidential file corresponding to the expired confidential non-sensitive tag, the storage period is marked as 2010-2020 (the confidentiality period has passed) (management attribute data), retaining the marking information of the secret level, the original storage period and the passed confidentiality period, the 2024 first quarter financial public announcement corresponding to the public affairs non-sensitive tag, the public announcement period: 2024.4.1-2024.4.15 (management attribute data), retaining the original format of the public announcement content, public announcement period and public announcement number. After completing the retention processing of the first metadata corresponding to the non-sensitive tags and the desensitization processing of the first metadata corresponding to the sensitive tags, the two types of metadata are integrated according to the field association relationship of the original archive information to form the second metadata, for example, in the second metadata of the employee archive of a certain enterprise, there are not only the employee ID number (110101582173001234) processed by irreversible feature coding (sensitive tag metadata), but also the employee's department: technology department (public information non-sensitive tag metadata) and 2024 annual department employee attendance statistics: attendance rate 98% (statistical summary non-sensitive tag metadata) retained directly, and the field association logic of each metadata unit is consistent with the original archive. After integration, the second metadata needs to be checked for integrity and security: integrity check checks whether the key metadata unit is missing (such as employee archive needs to include identity + department + attendance statistics, and if one is missing, it is judged to be incomplete), and checks whether the field association relationship is broken (such as attendance statistics needs to be correctly associated with the department), to ensure that the second metadata can support subsequent business use, and security check checks the desensitization effect of sensitive tag metadata by sampling (such as randomly selecting 10% of the identity metadata to verify whether the feature code meets the irreversibility and whether the format is consistent with the original), and confirms that there is no sensitive information mixed into the non-sensitive tag metadata (such as checking whether the employee contact information is included in the public recruitment notice without desensitization), to ensure that the second metadata has no security risk.The non-sensitive metadata directly retains the original content and format, maximally retains the practicability of the archive information, avoids information value loss caused by excessive processing (for example, statistical data can be directly used for report generation without desensitization), and the integration and double verification of the sensitive and non-sensitive metadata ensure the structural integrity of the second metadata (which can be directly used for query and statistics) and the safety of sensitive information (which can be safely used for multi-department sharing), finally realizes the management goal of safe and controllable, ready-to-use and ready-to-take digital archives, and adapts to the archive use requirements of multiple scenes such as daily operation of enterprises and government information disclosure.
[0153] Embodiment 2, refer to Figure 2 For another embodiment of the application, which is different from the first embodiment, a digital security-based archive information processing system is provided, comprising an acquisition module, an identification module, a processing module and an integration module.
[0154] The acquisition module acquires first metadata corresponding to the archive information.
[0155] The identification module identifies the sensitive tags and the non-sensitive tags of the first metadata based on a preset tag system template by using a hierarchical logic combining rule matching and manual auditing, and the preset tag system template is dynamically iterated through historical sample learning and new tag case verification.
[0156] The processing module performs differential desensitization processing on the first metadata corresponding to the identified sensitive tags according to the sensitive tags, and does not perform desensitization processing on the first metadata corresponding to the non-sensitive tags.
[0157] The integration module integrates the first metadata corresponding to the desensitized sensitive tags and the first metadata corresponding to the non-sensitive tags which have not been desensitized, to obtain second metadata.
[0158] In one of the embodiments, first, the first metadata corresponding to the archive information is acquired by the acquisition module, the first metadata including field identification data, content description data, management attribute data and associated subject data of the archive, then the preset label system template is called by the identification module to identify sensitive labels and non-sensitive labels of the first metadata (wherein the preset label system template is dynamically iterated through historical sample learning and new label case verification, and the identification process adopts hierarchical logic combining rule matching and manual review), then the processing module differentiates the desensitization processing of the first metadata corresponding to the sensitive labels according to the label type, such as irreversible feature coding method for desensitization of the associated subject data corresponding to the identity identification sensitive label to prevent original information leakage, and the first metadata corresponding to the non-sensitive labels is not desensitized, and the original content and format are directly retained, finally, the integrated module integrates the desensitized metadata corresponding to the sensitive labels and the non-desensitized metadata corresponding to the non-sensitive labels to obtain the second metadata which can be directly used for subsequent archive query, data statistics and information sharing. Through accurate label identification and differential desensitization, the digital security of sensitive content in the archive is effectively guaranteed, and the usability of non-sensitive information is maximized. The dynamically iterated preset label system template can adapt to changes in different business scenarios and compliance requirements, long-term maintain the accuracy and flexibility of archive information processing, and improve the overall efficiency of digital archive management.
[0159] The present application forms an efficient and dynamic iteration label system in three dimensions of information security, processing efficiency and value retention, supplements features through historical sample learning, and updates rules by combining new case verification, greatly improves the long-term accuracy of sensitive and non-sensitive label identification, effectively avoids the problems of missed and mistaken judgments caused by rule solidification, and the hierarchical identification logic combines automatic feature matching and manual review, which not only significantly improves the processing efficiency and reduces the labor cost, but also complements the blind area of automatic identification, and stably guarantees the label identification quality. The differential desensitization strategy is highly targeted, matches exclusive processing schemes for different types of sensitive data, can build a strong information security line, and can avoid data value loss caused by excessive desensitization. In addition, the integrated metadata can be directly used in query, statistics and sharing scenarios after double verification, seamlessly connects the subsequent application of archive digitization, and truly realizes the safe and efficient circulation of archive resources.
[0160] Those skilled in the art will appreciate that embodiments of the present application can be readily used as a method, a system or a computer program product. Accordingly, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product on one or more computer-usable storage media (or computer- readable storage media) having computer-usable program code embodied in the medium. The medium can be any available medium or combination thereof that is accessible by a general purpose or special purpose computer. By way of example, such computer-usable storage media can include a volatile memory, such as a random access memory (RAM), a non-volatile memory, such as a read-only memory (ROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a floppy diskette, a compact disk, a hard disk, or any other medium that can be used to carry or store computer-usable program code in the form of computer-usable instructions or data structures and that can be accessed by a general purpose or special purpose computer, or a general-purpose or special-purpose processor. Also, the present application can be embodied in a computer program product which can be executed in particu Figure 1 one or more functions specified in the flow or flows and / or blocks Figure 1 one or more functions specified in the flow or flows and / or blocks
[0161] It should be noted that the above-mentioned embodiments are only used to illustrate but not to limit the technical solutions of the present application. Although the present application is described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or equivalent replaced, without departing from the spirit and scope of the technical solutions of the present application, which should be covered in the scope of the claims of the present application.
Claims
1. A digital security-based archive information processing method, characterized by, The method comprises the following steps: Step S100, acquiring first metadata of archive information; Step S200, based on a preset label system template, using a hierarchical logic combining rule matching and manual review to identify sensitive labels and non-sensitive labels from the first metadata, and the preset label system template is dynamically iterated through historical sample learning and new label case verification; Step S300, for the first metadata corresponding to the sensitive labels, performing differential de-sensitization processing according to the sensitive labels, and not performing de-sensitization processing on the first metadata corresponding to the non-sensitive labels; Step S400, integrating the first metadata corresponding to the sensitive labels after de-sensitization processing and the first metadata corresponding to the non-sensitive labels without de-sensitization processing to obtain second metadata; The construction method of the preset label system template comprises: The preset label system template comprises classification rules and corresponding common feature library; The classification rules comprise classification rules of sensitive labels and classification rules of non-sensitive labels, wherein the classification rule of the sensitive label comprises field features corresponding to the sensitive label and content range corresponding to the sensitive label, and the classification rule of the non-sensitive label comprises exclusion conditions of the non-sensitive label; The field features corresponding to the sensitive label comprise field names, data types and field levels corresponding to identity identification sensitive labels, technical secret sensitive labels, commercial data sensitive labels and privacy information sensitive labels respectively, wherein the field name corresponding to the identity identification sensitive label is a correlated subject unique identification type name, the data type is an individual identity code type, and the field level is a core correlation field; The field name corresponding to the technical secret sensitive label is a technical parameter name, the data type is a technical document type, and the field level is a core technology field; The field name corresponding to the commercial data sensitive label is a commercial information name, the data type is a commercial transaction or asset type, and the field level is a core business field; The field name corresponding to the privacy information sensitive label is a personal privacy name, the data type is a private information record type, and the field level is a privacy correlation field; The content range corresponding to the sensitive label comprises information content boundaries corresponding to the identity identification sensitive label, the technical secret sensitive label, the commercial data sensitive label and the privacy information sensitive label respectively, wherein the information content boundary corresponding to the identity identification sensitive label is an information range capable of uniquely identifying the identity of the correlated subject; The information content boundary corresponding to the technical secret sensitive label is an undisclosed core technology information range; The information content boundary corresponding to the commercial data sensitive label is a commercial information range with commercial value and not disclosed; The information content boundary corresponding to the privacy information sensitive label is a protected personal private information range; The exclusion conditions are specifically: Excluding the first metadata comprising the field features corresponding to the sensitive label, excluding the first metadata falling within the content range corresponding to the sensitive label, and excluding the first metadata having sensitive management attribute correlation features, wherein the sensitive management attribute correlation features comprise attribute features of a first metadata classification level being secret or above, a storage period being marked as permanent confidentiality, and access permission being identified as being limited to core personnel; The sample of the docking historical metadata is manually reviewed and labeled with sensitive labels and non-sensitive labels to obtain labeled samples, and the common features corresponding to the sensitive labels and non-sensitive labels in the labeled samples are extracted, and a preset common feature library is constructed, the common features are stored in the feature items of the corresponding label attributes, the common feature library and the classification rule are cooperated to provide a feature matching basis for the label identification of the first metadata, and the common features are supplemented to the classification rule; According to a preset period, new label cases generated in the archive information processing process which are not covered by the current classification rule are collected, the new label cases include first metadata content, associated field and business scenario information; The new label cases are verified by a preset audit subject to determine the label attributes of the new label cases, the label attributes include sensitive labels or non-sensitive labels; And the first metadata features corresponding to the label attributes are extracted, the first metadata features include field identification features, content keyword features and management attribute association features; Based on the verified label attributes and the first metadata features, the determination logic of the classification rule is updated, and the feature items of the corresponding label attributes in the common feature library are supplemented; The updated classification rule and common features are integrated into a preset label system template, the template iteration is realized by replacing the old version template or forming a version superposition, the preset label system template is adapted to the business scenarios and compliance requirements corresponding to the new label cases to maintain the long-term accuracy of the sensitive label and non-sensitive label identification of the first metadata; And the iterated preset label system template is applied to the sensitive label and non-sensitive label identification process of the first metadata in real time; The sensitive label and non-sensitive label identification of the first metadata is performed by using the hierarchical logic combining rule matching and manual review, including: Obtaining the first metadata, calling the preset label system template, the preset label system template includes the classification rule of the sensitive label and the classification rule of the non-sensitive label and the corresponding common feature library; The first metadata is structurally split according to the field structure to obtain a plurality of independent metadata units, the metadata units correspond to the field identification data segment, the content description data segment, the management attribute data segment or the associated subject data segment in the first metadata; For each metadata unit, the features corresponding to the metadata unit are extracted; The features include field identification features, content keyword features and management attribute association features, The field identification features include field name, data type and field level; The content keyword features include core expression vocabulary, data format identifier and semantic direction; The management attribute association features include data security level, storage period and access permission identifier; The features of the metadata unit are matched with the classification rule of the sensitive label in the preset label system template, specifically including: If the field identification features match the field features corresponding to the sensitive label, and the content keyword features fall within the content range corresponding to the sensitive label, the metadata unit is marked as a sensitive label; For metadata units not marked as sensitive labels, first screen according to the exclusion conditions of non-sensitive labels, if the exclusion conditions are not triggered, compare the features of the metadata units with the common feature library corresponding to the non-sensitive labels, if the matching degree exceeds the preset threshold, mark the metadata units as non-sensitive labels; For metadata units that do not match sensitive labels or non-sensitive labels, mark them as pending labels, and generate an audit request including the features of the metadata units and the associated business scenarios, and send it to the preset audit terminal; Receive the label attribute confirmation result returned by the audit terminal, update the pending label to sensitive label or non-sensitive label, and summarize the label identification results of all metadata units to form a label list of the first metadata; The label identification result includes the sensitive label and the non-sensitive label corresponding to the metadata unit and the label source, and the label source includes the field identification feature matching and the content keyword feature matching of the sensitive label, the exclusion condition screening and the threshold comparison of the common feature library of the non-sensitive label, and the audit terminal confirmation of the pending label.
2. The digital security based archival information processing method as claimed in claim 1 wherein: The first metadata includes field identification data, content description data, management attribute data and associated subject data of the archive information; The sensitive label includes an identity sensitive label, a technical secret sensitive label, a commercial data sensitive label and a privacy information sensitive label; The non-sensitive label includes a public information non-sensitive label, a statistical summary non-sensitive label, an expired secret non-sensitive label and a public matter non-sensitive label; The corresponding relationship between the sensitive label and the non-sensitive label and the first metadata is as follows: The identity sensitive label corresponds to the associated subject data; The technical secret sensitive label corresponds to the content description data; The commercial data sensitive label corresponds to the field identification data and the content description data; The privacy information sensitive label corresponds to the content description data; The public information non-sensitive label corresponds to the content description data and the management attribute data; The statistical summary non-sensitive label corresponds to the content description data; The expired secret non-sensitive label corresponds to the management attribute data; The public matter non-sensitive label corresponds to the management attribute data.
3. The digital security-based archive information processing method of claim 2, wherein: The desensitization processing of the first metadata corresponding to the sensitive label includes: If the sensitive label is an identity sensitive label, the first metadata corresponding to the identity sensitive label is associated subject data, and an irreversible feature coding method in biometric identification technology is used for desensitization and original format is retained, and the steps are as follows: Extract the field identification feature, content keyword feature and management attribute association feature of the associated subject data; The field identification feature includes the format rule and length parameter of the identity field; The content keyword feature includes the core coding segment of the identity; The management attribute association feature includes the privacy protection level corresponding to the identity; Map the extracted features to low-dimensional irreversible feature codes based on a preset random projection matrix or a nonlinear transformation algorithm; The conditions that the feature code needs to meet include that the length parameter is consistent with the length parameter of the original associated subject data, and the data type and field structure of the original associated subject data match the format rule of the original associated subject data; By adjusting the dimension parameter of the random projection matrix or the complexity parameter of the nonlinear transformation, the strength of the irreversible mapping of the feature code and the original associated subject data is controlled, and the strength of the irreversible mapping is positively correlated with the privacy protection level of the associated subject data.
4. The digital security based archival information processing method as claimed in claim 3, wherein: The generated irreversible feature code replaces the sensitive content of the original associated subject data, and the conditions that the feature code needs to meet include that the difference degree of the feature codes of different original associated subject data is in a preset proportion with the difference degree of the original data, and the original associated subject data content is ensured to be unable to be inversely deduced through the feature code through the irreversibility verification; If the sensitive label is a technical secret sensitive label, a commercial data sensitive label or a privacy information sensitive label, the first metadata corresponding to the sensitive label is content description data or field identification data, the content is processed in a desensitization manner adapted to the data type, and the original format of the first metadata is retained. The desensitization manner adapted to the data type refers to a text desensitization manner for content description data and a field desensitization manner for field identification data.
5. The digital security based archival information processing method as claimed in claim 4, wherein: The first metadata corresponding to the non-sensitive label is not processed, including: If the non-sensitive label is a public information non-sensitive label, a statistical summary non-sensitive label, an expired secret non-sensitive label or a public notice non-sensitive label, the first metadata corresponding to the non-sensitive label is content description data or management attribute data, and the content and format of the first metadata are directly retained without any desensitization processing; The first metadata corresponding to the sensitive label and the first metadata corresponding to the non-sensitive label after processing are integrated to obtain second metadata, and the second metadata needs to pass integrity verification and security verification to ensure that it can be directly used for subsequent archive information query, statistics and sharing scenarios.
6. A digital security-based archive information processing system applied to the digital security-based archive information processing method according to any one of claims 1-5, characterized in that: It includes an acquisition module, an identification module, a processing module and an integration module; The acquisition module acquires first metadata corresponding to archive information; The identification module identifies sensitive labels and non-sensitive labels of the first metadata based on a preset label system template by using a hierarchical logic combining rule matching and manual auditing, and the preset label system template is dynamically iterated through historical sample learning and new label case verification; The processing module differentially desensitizes the first metadata corresponding to the sensitive labels according to the sensitive labels, and does not desensitize the first metadata corresponding to the non-sensitive labels; The integration module integrates the desensitized first metadata corresponding to the sensitive labels and the first metadata corresponding to the non-sensitive labels without desensitization to obtain second metadata.
Citation Information
Patent Citations
Human resource archive filing management system
CN120743852A
User information storage control method, electronic device, and non-volatile computer-readable storage medium
US20250111085A1