Automatic information secret setting method and system based on intelligent identification
Through intelligent identification technology, the information is automatically identified and the problems of inefficient and insufficient accuracy of traditional information are solved, and the information is efficient and accurate.
Patent Information
- Application Number
- CN202411433676.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-15
- Publication Date
- 2025-06-24
AI Technical Summary
The traditional information encryption process is inefficient and insufficiently accurate, making it difficult to fully reflect the sensitivity and importance of information, especially when dealing with multi-level and multi-dimensional features.
The automatic information encryption method based on intelligent identification is adopted, and the feature data of the classified file is obtained and the data in the pending information database is compared, and the matching feature set is generated, and the intelligent identification module is used to perform the secret level determination process.
It significantly improves the efficiency and accuracy of information determination, ensures the consistency and accuracy of judgment results, makes full use of multi-level feature information, and provides a more comprehensive determination solution.
Smart Images

Figure CN120196751A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical fields of information security and automated classification, and specifically relates to an information automatic classification method and system based on intelligent recognition. Background Art
[0002] With the rapid development of information technology, the amount of information generated, stored, and managed by organizations and enterprises has grown explosively. Against this backdrop, information security management has become particularly important. Information classification, that is, dividing levels according to the sensitivity and importance of information, is one of the key links in information security management. The traditional information classification process mainly relies on manual judgment, which is not only time-consuming and laborious, but also prone to inconsistencies and errors in classification results due to human factors. In addition, traditional methods often only consider features in a single dimension and are difficult to comprehensively reflect the sensitivity and importance of information.
[0003] In actual operation, information usually has multi-level and multi-dimensional features. For example, a file not only has its basic attributes (such as file name, size, type, etc.), but may also contain content descriptions (such as keywords, subject classification, content summary) and associated identifiers (such as unique identifiers, version numbers, information about the creator or modifier, etc.). Traditional methods face challenges in dealing with these complex features and cannot effectively utilize a large amount of potential information, thus affecting the accuracy of classification.
[0004] To improve the efficiency and accuracy of classification, researchers have begun to explore information classification methods based on intelligent algorithms and automated processes. Through advanced data mining techniques and machine learning algorithms, useful features can be extracted from massive data to support automated information classification and classification. However, most current systems still lack a comprehensive and efficient solution that can fully integrate multi-level feature information to achieve intelligent classification of information. Summary of the Invention
[0005] (1) Technical Problems to be Solved
[0006] This invention mainly aims at the above problems and proposes an information automatic classification method and system based on intelligent recognition, aiming to solve the problems of low efficiency and insufficient accuracy in the traditional information classification process, and to achieve efficient and accurate classification of information by adopting intelligent algorithms and automated processes.
[0007] (2) Technical Solutions
[0008] To achieve the above object, this invention provides an information automatic classification method based on intelligent recognition, including the following steps:
[0009] Obtain the feature data of classified files from the information storage system, where the feature data consists of multi-level information description files and their associated identifiers;
[0010] Extract the feature information of the data to be classified from the information repository to be processed, where the feature information includes detailed structure and hierarchical descriptions;
[0011] By comparing the feature data of the classified files with the feature information of the data to be processed, exclude the mismatched features and generate a set of matching features;
[0012] Using the generated set of matching features, retrieve the corresponding target information data from the information repository to be processed;
[0013] Transmit the target information data to the intelligent recognition module for classification level determination processing.
[0014] Furthermore, the classification level determination processing includes:
[0015] Obtain the operation parameters of the intelligent recognition module;
[0016] According to the obtained operation parameters, divide the target information data into several information data groups, and each information data group consists of a series of screened information layers;
[0017] Input each information data group into the intelligent recognition module one by one for further classification level determination.
[0018] Furthermore, obtaining the feature data of the classified files includes: retrieving and obtaining the unique identifier of the classified files, as well as the multi-level feature tags corresponding to each identifier.
[0019] Furthermore, extracting the feature information of the data to be classified includes: extracting the unique identifier of the data file to be processed and the associated hierarchical feature description.
[0020] Furthermore, the process of obtaining feature data includes the following steps:
[0021] Connect to the database interface in the information storage system;
[0022] Execute a predefined query statement to extract the feature data related to the classified files;
[0023] Format the extraction result and convert it into a multi-level information description format for subsequent processing.
[0024] Furthermore, the process of extracting feature information includes the following steps:
[0025] Initialize the access permission of the information repository to be processed;
[0026] Use regular expressions to identify the hierarchical structure and keyword tags of the data to be classified;
[0027] Classify and store the recognized hierarchical structures and keyword tags to form a complete set of feature information.
[0028] Further, excluding the mismatched features to form a set of matching features includes: sequentially comparing the hierarchical features of the data to be classified with the hierarchical features of the classified files, and only retaining the same hierarchical features as part of the set of matching features.
[0029] Further, the process of generating the set of matching features includes the following steps:
[0030] Construct a set of rules for feature comparison, where the rules consider lexical similarity, semantic consistency, and hierarchical correspondence;
[0031] Apply the set of rules to compare the feature information of the data to be classified with the feature data of the classified files one by one;
[0032] Retain the feature pairs that conform to the set of rules and eliminate the feature pairs that do not conform, finally obtaining the set of matching features.
[0033] Further, the classification level determination process also includes:
[0034] Encrypt and compress the information data group into a data packet;
[0035] Use a randomly generated password to encrypt the data packet through a preset encryption algorithm;
[0036] Encrypt the random password using the public key and transmit the finally generated encrypted and compressed packet to the intelligent recognition module for parsing.
[0037] To achieve the above object, the present invention provides an information automatic classification system based on intelligent recognition, including the following modules:
[0038] Feature data acquisition module: used to acquire the feature data of the classified files from the information storage system, where the feature data is composed of a multi-level information description file and its associated identifier;
[0039] Feature information extraction module: used to extract the feature information of the data to be classified from the information library to be processed, and the feature information includes detailed structure and hierarchical descriptions;
[0040] Feature matching module: used to exclude the mismatched features by comparing the feature data of the classified files with the feature information of the data to be processed, and generate a set of matching features;
[0041] Target information retrieval module: use the generated set of matching features to retrieve the corresponding target information data from the information library to be processed;
[0042] Intelligent recognition module: It is used to receive and process the target information data for classified level determination processing.
[0043] (III) Beneficial effects
[0044] Compared with the prior art, the embodiments of the present disclosure propose an information automatic classification method based on intelligent recognition, aiming to overcome the deficiencies of the prior art. Through intelligent algorithms and automated processes, this method can not only significantly improve the efficiency of information classification, but also ensure the consistency and accuracy of the determination results. Specifically, by obtaining the characteristic data of the classified files and comparing them with the data in the information database to be processed, the system can exclude the mismatched characteristics and form a set of matching characteristics. Subsequently, the set is used to retrieve the target information data, and the intelligent recognition module is used to perform classified level determination processing on the target data. This process makes full use of the advantages of multi-level characteristic information and provides a more comprehensive classification solution. Description of the drawings
[0045] Figure 1 It is a flowchart of an information automatic classification method based on intelligent recognition disclosed in this application.
[0046] Figure 2 It is a flowchart of information classified level determination based on an intelligent recognition module disclosed in this application.
[0047] Figure 3 It is a timing diagram of the generation of a matching set disclosed in this application.
[0048] Figure 4 It is an overall architecture diagram of an information automatic classification system based on intelligent recognition disclosed in this application. Specific embodiments
[0049] The following combines the drawings and embodiments to further describe in detail the specific embodiments of the present invention. The following embodiments are used to illustrate the present invention, but not to limit the scope of the present invention.
[0050] The information classification process usually relies on manual judgment, which is not only time-consuming and laborious, but also prone to inconsistencies and errors in the classification results due to human factors. In addition, traditional information classification methods often only consider features in a single dimension and are difficult to comprehensively reflect the sensitivity and importance of information. To solve these problems, the embodiments of the present disclosure provide an information automatic classification method based on intelligent recognition, which realizes efficient and accurate classification of information through intelligent algorithms and automated processes.
[0051] As Figure 1 - Figure 2 shown, the method of the embodiments of the present disclosure mainly includes the following steps:
[0052] Step S100: Obtain the feature data of the classified files from the information storage system, where the feature data consists of a multi-level information description file and its associated identifier;
[0053] Extract the feature data of the already classified files from an existing information storage system (such as a database, a file system, or cloud storage). These feature data are multi-level, including not only the basic attributes of the files (such as file name, size, type), but also content descriptions (such as keywords, subject classifications, content summaries) and the associated identifiers of the files (such as unique identifiers, version numbers, information of the creator or modifier, etc.).
[0054] Step S200: Extract the feature information of the data to be classified from the information repository to be processed, where the feature information contains detailed structural and hierarchical descriptions;
[0055] In Step S200, turn to the information repository to be processed, which refers to a collection of data that has not been classified or needs to be reclassified. Extract the feature information from this data. This information also has detailed structural and hierarchical descriptions, aiming to comprehensively capture all aspects of the data for subsequent comparison and classification work.
[0056] Step S300: By comparing the feature data of the classified files with the feature information of the data to be processed, exclude the mismatched features and generate a set of matching features;
[0057] Compare the feature data of the classified files obtained in Step S100 with the feature information of the data to be processed extracted in Step S200. Identify and exclude the mismatched features, and finally generate a set containing all the matching features. For example: Compare the feature information of "Latest Market Analysis.pdf" with the feature data of classified files such as "2023 Q1 Sales Report.docx". It is found that there are differences between the two in terms of file type (PDF vs. Word), keywords (market analysis vs. sales), etc., but there are matches in some keywords (such as "market") and possible topic similarities (both are about business analysis). Therefore, generate a set containing these matching features, such as "Business Analysis", "Market-related".
[0058] Step S400: Use the generated set of matching features to retrieve the corresponding target information data from the information repository to be processed;
[0059] Based on the set of matching features generated in step S300, perform an exact search in the information repository to be processed, with the aim of finding target information data that is highly similar or matching to the classified files in specific features. For example, using matching features such as "business analysis", "market-related", etc., retrieve all relevant documents from the information repository to be processed, including "Latest Market Analysis.pdf" and other documents that may contain similar keywords or topics.
[0060] Step S500: Transmit the target information data to the intelligent recognition module for classification level determination processing.
[0061] It can be understood that if the target information contains highly sensitive personal data or key technical details, this module will recommend marking it as high-level confidential. Ultimately, the classification level determination result will help the organization effectively manage and protect its information assets.
[0062] Example illustration: Suppose there is an information file to be processed and its classification level needs to be determined. First, obtain the feature data of the classified files from the information storage system. These feature data include information in multiple dimensions such as the unique identifier of the file, content description, source, etc. Then, extract the feature information of the data to be classified from the information repository to be processed, such as the name of the file, keywords, hierarchical structure, etc. Next, generate a set of matching features by comparing the feature data of the classified files with the feature information of the data to be processed. Use this set of matching features to retrieve the file in the information repository to be processed that best matches the target information data. Finally, transmit the target information data to the intelligent recognition module for classification level determination processing. The intelligent recognition module conducts in-depth analysis on the target information data according to the preset operation parameters and algorithm models, and ultimately determines its classification level. Through the above steps, the method of the embodiments of the present disclosure can achieve efficient and accurate classification of information, providing strong support for information security management.
[0063] In the classification determination process of step S500, the system needs to go through several key steps to ensure the accurate classification and secure management of information data. First, in step S501, the system obtains the operation parameters of the intelligent recognition module. These parameters include algorithm settings, risk assessment criteria, sensitivity thresholds, etc., which jointly determine how to process and analyze the target information data. Then, in step S502, based on the obtained operation parameters, the system divides the target information data into several information data groups. Each information data group consists of a series of screened information layers, enabling the system to break down larger chunks of information into more manageable small units according to the nature, structure, and complexity of the content. For example, a technical manual containing multiple chapters will be split into different data groups, with each group focusing on an independent topic or area. This not only helps improve the accuracy of analysis but also optimizes the processing efficiency. Finally, in step S503, the system inputs each information data group into the intelligent recognition module one by one for further classification determination. During this process, the intelligent recognition module uses its built-in analysis capabilities and the previously obtained operation parameters to conduct a detailed review and evaluation of each information data group to determine its appropriate security level. This step-by-step refinement method ensures that even the most minute content details can be accurately determined, and the resulting classification determination results will help the organization effectively formulate security policies and ensure the protection of information assets.
[0064] It should be noted that the intelligent recognition module analyzes and evaluates each information data group through the following steps to determine its appropriate security level: receiving and loading the previously obtained operation parameters, which are used to configure the analysis algorithm of the intelligent recognition module, including risk assessment criteria, sensitivity thresholds, and classification rules; preprocessing the input information data group to extract relevant feature information, such as keywords, topic structures, and background relationships; using machine learning models or expert systems to analyze the extracted feature information to identify potential sensitive content and security risks; comparing the analysis results with the preset risk assessment criteria to conduct a preliminary classification determination for each information data group; combining context information and historical determination records to review and adjust the preliminary determination results to improve the accuracy and reliability of the determination; outputting the final classification determination results and applying them to the information management process.
[0065] In the intelligent recognition module, in order to perform classification level determination more accurately, considering that classification level determination may involve tasks such as text classification, sentiment analysis, or entity recognition, one or a combination of the following models can be adopted: Support Vector Machine (SVM), Naive Bayes classifier, deep learning models (such as Convolutional Neural Network CNN or Recurrent Neural Network RNN). To train the selected machine learning model, a labeled training dataset needs to be prepared. This dataset contains multiple classified file samples, and each sample should be labeled with the corresponding classification level label (such as "public", "confidential", etc.).
[0066] Before training the model, perform feature engineering on the data to extract features useful for the classification level determination task; for example: remove irrelevant characters, stop words, and punctuation marks, perform stemming or lemmatization, etc., and use methods such as TF-IDF (Term Frequency - Inverse Document Frequency), word2vec, or BERT to convert the text into numerical feature vectors, and select the most important feature subset according to statistical tests, model weights, or other criteria.
[0067] Use the selected machine learning model and the prepared training dataset to train the model. During the training process, the model performance can be optimized by adjusting model parameters (such as learning rate, regularization strength, etc.). The goal of training is to minimize the loss function, which measures the difference between the model prediction and the actual label.
[0068] In this embodiment, obtaining the feature data of the classified file includes: retrieving and obtaining the unique identifier of the classified file, and the multi-level feature labels corresponding to each identifier.
[0069] In this embodiment, extracting the feature information of the data to be classified includes: extracting the unique identifier of the data file to be processed and the hierarchical feature description associated with it.
[0070] In step S100, the feature data acquisition process includes the following steps: connect to the database interface in the information storage system; execute a predefined query statement to extract the feature data related to the classified file; format the extraction result and convert it into a multi-level information description format for subsequent processing.
[0071] In step S300, the characteristic data of the classified files is compared with the characteristic information of the data to be processed through the following steps, the mismatched characteristics are excluded, and a set of matching characteristics is generated: Extract the characteristic data from the classified files, which includes but is not limited to keywords, subject tags, format patterns, and statistical attributes; Extract the characteristic information from the data to be processed, ensuring that the characteristic information has the same format and structure as the characteristic data of the classified files; Compare the characteristic data of the classified files with the characteristic information of the data to be processed item by item to identify the characteristic items with high similarity; Exclude the mismatched characteristics below the preset similarity threshold to improve the accuracy of the matching results; Summarize the retained matching characteristics to form a set of matching characteristics, which reflects the characteristic association of the data to be processed in the classified files; Use this set of matching characteristics for subsequent data classification and analysis processes to enhance the effectiveness of the processing.
[0072] In step S200, the specific implementation manner of the characteristic information extraction process is as follows: First, in step S201, the system initializes the access rights to the information library to be processed to ensure the legality and security of subsequent data processing operations. This process includes verifying the user identity, assigning appropriate access levels, and setting necessary encryption measures. Then, in step S202, the data to be classified is analyzed using a preset regular expression to identify the hierarchical structure and keyword tags in the data. The regular expression is a flexible and powerful text processing tool that can efficiently capture the regular patterns in the data, such as the chapter titles of documents, date formats, special terms, etc. Subsequently, in step S203, the hierarchical structure and keyword tags identified by the regular expression are classified and stored. In this process, the system organizes the related data elements into logically connected groups according to the predefined classification rules, thus forming a complete set of characteristic information.
[0073] As Figure 3 shown, in step S300, excluding the mismatched characteristics to form a set of matching characteristics includes: Sequentially comparing the hierarchical characteristics of the data to be classified with the hierarchical characteristics of the classified files, and only retaining the same hierarchical characteristics as part of the set of matching characteristics.
[0074] In step S300, the specific implementation of the generation process of the matching feature set is as follows: First, a set of rules for feature comparison is constructed, which comprehensively considers factors such as lexical similarity, semantic consistency, and hierarchical correspondence. Lexical similarity involves calculating the similarity degree of terms or phrases in the feature information through algorithms, while semantic consistency evaluates the meaning relevance between different feature items through natural language processing techniques. At the same time, the hierarchical correspondence ensures the matching on the data structure. Next, the constructed rule set is applied to compare the feature information of the data to be classified with the feature data of the classified files one by one. In this step, the system will analyze each feature pair in detail according to the criteria of the rule set and identify which feature pairs meet the established conditions. Finally, the feature pairs that conform to these rule sets are retained, while the non-conforming feature pairs are excluded, and an accurate matching feature set is finally formed.
[0075] In step S500, the specific implementation of the classification level determination process includes the following steps: First, the information data group is preprocessed, that is, it is compressed into a smaller data packet through a compression algorithm to improve the efficiency of transmission and storage. Then, a random password is generated, and the compressed data packet is symmetrically encrypted using this random password in combination with a preset encryption algorithm. This encryption process ensures the confidentiality of the data packet during transmission and is not affected by unauthorized access. Subsequently, in order to securely transmit this random password, the system encrypts the random password using the public key in the asymmetric encryption technology. This method ensures that even if the interceptor can obtain the encrypted data packet, they cannot decrypt the data in it because they lack the private key to decrypt the encrypted random password. Finally, the encrypted and compressed packet after the above processing is transmitted to the intelligent recognition module, which has the corresponding private key and can successfully parse and decrypt the data, so as to realize the determination and management of the classification level of the information data group.
[0076] As Figure 4 shown, the present invention also provides an information automatic classification system based on intelligent recognition, including the following modules:
[0077] Feature data acquisition module: used to acquire the feature data of the classified files from the information storage system, where the feature data consists of a multi-level information description file and its associated identifier;
[0078] Feature information extraction module: used to extract the feature information of the data to be classified from the information library to be processed, and the feature information includes detailed structure and hierarchical descriptions;
[0079] Feature matching module: used to generate a matching feature set by comparing the feature data of the classified files with the feature information of the data to be processed and excluding the non-matching features;
[0080] Target information retrieval module: Retrieves corresponding target information data from the information repository to be processed by using the generated matching feature set;
[0081] Intelligent recognition module: Receives and processes the target information data for classification determination processing.
[0082] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. A method for automatically determining confidentiality of information based on intelligent identification, characterized in that: The following steps are involved: Acquire characteristic data of the classified files from the information storage system, wherein the characteristic data is composed of multi-level information describing the files and their associated identifiers; Extracting characteristic information of the data to be classified from the information database to be processed, the characteristic information including detailed structure and hierarchical description; By comparing the feature data of the classified files with the feature information of the data to be processed, unmatched features are eliminated to generate a matching feature set; Using the generated matching feature set, the corresponding target information data is retrieved from the information database to be processed; The target information data is transmitted to the intelligent identification module for confidentiality level determination processing.
2. The method for automatically determining confidentiality of information based on intelligent identification according to claim 1, characterized in that: The classification process includes: Obtaining the operating parameters of the intelligent recognition module; According to the acquired operation parameters, the target information data is divided into a plurality of information data groups, each of which is composed of a series of screened information layers; Each information data group is input into the intelligent recognition module one by one for further confidentiality determination.
3. The method for automatically determining confidentiality of information based on intelligent identification according to claim 1, characterized in that: Acquiring characteristic data of the classified files includes: retrieving and acquiring unique identifiers of the classified files and multi-level characteristic labels corresponding to each identifier.
4. The method for automatically determining confidentiality of information based on intelligent identification according to claim 1, characterized in that: Extracting feature information of the data to be classified includes: extracting a unique mark of the data file to be processed and a hierarchical feature description associated therewith.
5. The method for automatically determining confidentiality of information based on intelligent identification according to claim 1, characterized in that: The feature data acquisition process includes the following steps: A database interface connected to the information storage system; Execute predefined query statements to extract feature data related to the classified files; The extraction results are formatted and converted into a multi-level information description format for subsequent processing.
6. The method for automatically determining confidentiality of information based on intelligent identification according to claim 1, characterized in that: The feature information extraction process includes the following steps: Initialize access permissions to the information repository to be processed; Use regular expressions to identify hierarchical structures and keyword tokens in the data to be classified; The identified hierarchical structures and keyword tags are classified and stored to form a complete feature information set.
7. The method for automatically determining confidentiality of information based on intelligent identification according to claim 1, characterized in that: Excluding unmatched features to form a matching feature set includes: sequentially comparing hierarchical features of the data to be classified with hierarchical features of the classified files, and retaining only the same hierarchical features as part of the matching feature set.
8. The method for automatically determining confidentiality of information based on intelligent identification according to claim 1, characterized in that: The matching feature set generation process includes the following steps: Construct a set of rules for feature comparison, which considers lexical similarity, semantic consistency and hierarchical correspondence; Apply the rule set to compare the feature information of the data to be classified with the feature data of the classified files one by one; The feature pairs that meet the rule set are retained, and the feature pairs that do not meet the rule set are eliminated, and finally a matching feature set is obtained.
9. The method for automatically determining confidentiality of information based on intelligent identification as claimed in claim 2, characterized in that: The confidentiality level determination process also includes: Encrypting and compressing information data groups into data packets; Encrypt the data packet using a randomly generated password using a preset encryption algorithm; The random password is encrypted using the public key, and the resulting encrypted compressed package is sent to the intelligent recognition module for parsing.
10. An automatic information classification system based on intelligent identification, characterized in that: Includes the following modules: Feature data acquisition module: used to acquire feature data of classified files from the information storage system, wherein the feature data is composed of multi-level information description files and their associated identifiers; Feature information extraction module: used to extract feature information of the data to be classified from the information database to be processed, and the feature information contains detailed structure and hierarchical description; Feature matching module: used to compare the feature data of the classified files with the feature information of the data to be processed, eliminate unmatched features, and generate a matching feature set; Target information retrieval module: uses the generated matching feature set to retrieve the corresponding target information data from the information database to be processed; Intelligent identification module: used to receive and process the target information data to perform confidentiality level determination.
Citation Information
Patent Citations
Method and device for identifying entity in statement and electronic equipment
CN111144102A
File classification method, computing device and computer readable storage medium
CN113946548A
Text encryption method and apparatus, non-volatile storage medium, processor
CN114936376A
Data classification method and device, medium and equipment
CN118350040A
Intelligent and automatic data classification and grading method and device
CN118734126A
Cited By
Company information disclosure data processing system based on set operation and confidentiality constraint
CN122046384A