Cross-platform data security circulation supervision system and method
By obtaining initial review records, identifying reviewers, and utilizing neural network models and knowledge graph analysis to establish a benchmark set, the problems of low efficiency and insufficient data security in cross-platform archival document classification were solved, achieving more efficient and reliable archival document management.
Patent Information
- Application Number
- CN202510521497.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2045-04-24
AI Technical Summary
Existing technologies are inefficient and highly subjective in cross-platform file classification, which can easily lead to data misclassification, increase the risk of data leakage, and result in incomplete data security mechanisms.
By obtaining the initial review records, identifying the reviewers, reclassifying the documents, and using neural network model training and knowledge graph analysis to establish a benchmark set, the document category can be determined, and early warning prompts can be provided.
It improves the accuracy of archival document classification and data security, reduces the possibility of misclassification, and enhances the reliability and security of data.
Smart Images

Figure CN120611409B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data security technology, specifically to a cross-platform data security circulation supervision system and method. Background Technology
[0002] With the advent of the information age, big data has become an important tool for social research. For enterprises, the construction of big data platforms can not only safeguard their development but also improve their overall management level and capabilities. Currently, with the growth of data volume and the increasing demand for data value mining, enterprises generate a large number of archives during the production process. In order to better manage archives, they usually choose to migrate and classify them across platforms. Reasonable classification and management of archives helps to ensure data security in a cross-platform environment.
[0003] Currently, when migrating and classifying archival documents across platforms, the common methods are manual classification or semi-automated classification based on simple rules. However, given the massive amount of archival documents across different platforms, this approach is not only inefficient and highly subjective, but also prone to data misclassification, increasing the risk of data leakage. For example, improper classification can lead to unauthorized access to archival documents, resulting in data security issues and incomplete data security mechanisms. Summary of the Invention
[0004] The purpose of this invention is to provide a cross-platform data security circulation supervision system and method to solve the problems raised in the prior art.
[0005] To achieve the above objectives, the present invention provides the following technical solution:
[0006] A cross-platform data security circulation supervision method includes the following steps:
[0007] Step S100: Obtain the historical operation records of classifying archival documents into the corresponding archival categories as the initial review records, extract the initial reviewers corresponding to the initial review records, and determine the reviewers who will reclassify the archival documents from the classification personnel other than the initial reviewers, and use the operation records of the reviewers who reclassify the archival documents as the review records.
[0008] Step S200: Extract the content summary of the archival documents, as well as the archival descriptions written by the initial reviewers and the secondary reviewers after classifying the archival documents, and extract the relationships to obtain the entities and relationships; and obtain the benchmark set corresponding to each archival category based on the entities and relationships in the operation records corresponding to each archival document.
[0009] Step S300: Establish a corresponding neural network model for each archive category. Train each neural network model based on the archive description and content summary corresponding to each archive category to obtain the trained neural network model.
[0010] Step S400: Extract the archive category, archive description and content summary corresponding to the archive file to be detected, and obtain the entity and relationship corresponding to the archive file to be detected. Based on the benchmark set and neural network model corresponding to the archive category, determine whether to issue an early warning prompt for the archive category to which the archive file to be detected belongs.
[0011] Furthermore, step S100 includes:
[0012] Step S110: Obtain several historical preliminary review records. The preliminary review records are records of cross-platform transmission of archive files and initial archive classification. Extract the preliminary reviewer number and archive category automatically recorded by the computer program in each preliminary review record, as well as the archive description manually recorded by the preliminary reviewer. The archive description is a record of the preliminary reviewer's summary of the contents of the archive files and the explanation of the reasons for the classification after completing the classification of the archive files.
[0013] Step S120: If a file that has already been classified is found to have an incorrect classification during subsequent inspection, the file will be classified as an abnormal file; let A be the initial reviewer corresponding to a file F, and let N be the total number of initial review records corresponding to another classification person B. B And the number of abnormal files N0, the classification personnel B as the reviewer of file F, the first degree coefficient T1=(N B -N0) / N B This allows us to obtain the first degree coefficient for each category of personnel, excluding the initial reviewer A.
[0014] Extract all keywords from the content summary of file F to generate a keyword set; obtain the initial review records of several files that are not abnormal files corresponding to classifier B, extract keywords from the file description in each initial review record, and obtain the occurrence frequency of each keyword to generate a vocabulary list; add the occurrence frequencies of words in the vocabulary list that are the same as keywords in the keyword set to obtain the total number of words for classifier B; obtain the total number of words for each other classifier, and normalize them, using the normalized result as the second degree coefficient T2 for classifier B as a reviewer;
[0015] Step S130: The overall conformity level of classifier B as reviewer is then obtained as T. B=W1×T1+W2×T2, where W1 and W2 are the first weight and the second weight, respectively; obtain the total compliance degree of each category of personnel, and select the category of personnel with the highest total compliance degree as the reviewers; and have the reviewers reclassify the file F.
[0016] Furthermore, step S200 includes:
[0017] Step S210: Obtain the file description X corresponding to the preliminary review record of a certain file file F. A The file description corresponding to the review record X B And the content summary X, if the archival document F is consistent in the archival category in the initial review and the second review, then the archival description X is obtained. A File Description X B And all entities and relationships between entities in the content summary X, and create a separate profile description X. A File Description X B The knowledge graphs K1, K2, and K3 corresponding to the content summary X are used to extract all element combinations of entity-relationship-entity.
[0018] Step S220: Obtain a combination M1(E1) of elements in K1. M 1,R M 1,E2 M 1) and a combination of elements in K2, M2(E1) M 2,R M 2,E2 M 2), E1 M 1. E2 M 1. E1 M 2 and E2 M Both 2 are entities, R M 1 and R M Both are related, and E1 is obtained respectively. M 1 and E1 M 2. R M 1 and R M 2. E2 M 1 and E2 M Word similarity S between 2 E1 S R S E2 Calculate the combinatorial similarity X between combinatorial combinations M1 and M2. M =S R (S E1 +S E2 If X MIf the value exceeds a preset first threshold, then combination M1 is marked with a first label to obtain all first label combinations in knowledge graph K1; then, the element combinations in knowledge graph K1 and K3 are compared to obtain all second label combinations in knowledge graph K1; and the element combination that is both a first and second label combination is taken as the target combination of file F.
[0019] Relation extraction and knowledge graph construction are existing techniques in the field of large language models, and will not be elaborated here. "Entity-relationship-entity" is also a common basic unit in knowledge graphs, and it is combined as an element in this scheme. The specific extraction method can be obtained through existing technologies. The calculation of combination similarity is determined based on all entities and relationships, which helps to ensure the reliability of the combination similarity calculation. The target combination obtained from the combination of the first and second tags is determined by the archive description and content summary. The more target combinations there are, the greater the similarity and correlation between the archive description and the content summary. That is, the more reliable the content record in the manually recorded archive description is, the more convincing the classification reasons for the archive category in the archive description is, and the more reliable the target combination is as the classification basis for the corresponding archive category. This also shows the rationality of obtaining the following benchmark set.
[0020] Step S230: Establish a baseline set with empty elements for each archive category; based on the target combination and archive category corresponding to the archive documents that are consistent between the initial review and the final review, add the target combination to the baseline set of the corresponding archive category, and thus obtain the baseline set corresponding to each archive category.
[0021] Further, step S300 includes: taking all archive descriptions and content summaries corresponding to a certain archive category as relevant information, and all archive descriptions and content summaries corresponding to other archive categories as irrelevant information, setting the correlation between all relevant information and a certain archive category to 1, and the correlation between irrelevant information and a certain archive category to 0, and then taking several relevant and irrelevant information, as well as their respective correlation levels, as a dataset, and substituting them into the neural network model corresponding to a certain archive category for training, to obtain the trained neural network model.
[0022] Furthermore, step S400 includes:
[0023] Step S410: The archive file to be tested is a file that has completed cross-platform transmission and initial archive category classification. The archive category of the archive file to be tested is taken as C0. The archive description and content summary of the initial review record of the archive file to be tested are extracted, and all element combinations are obtained. The number of element combinations is taken as Y. The combination similarity between the y-th element combination and each combination in the baseline set of archive category C0 is calculated. The maximum combination similarity is taken as the target value of the y-th element combination. The target value of each element combination is obtained, thus determining the category fit of the archive file to be tested. Where e is the natural logarithm, q is the adjustment factor coefficient, and Sy is the target value of the y-th element combination;
[0024] It should be noted that the function h = 1 - e -x When x ≥ 0, h takes the value [0, 1), and is a function that increases as x increases. Furthermore, when x is small, the increase in h is greater than when x is large. In this scheme, the target value is determined by the combination of the file to be detected and the benchmark set. A small target value is sufficient to indicate a high degree of category fit between the two. Therefore, the function h = 1 - e is used here. -x The design should be carried out, and the adjustment factor coefficient q is used as an adjustment factor for the degree of category fit. The specific value should be determined according to the actual situation.
[0025] Step S420: Substitute the file description and content summary of the preliminary review record of the file to be tested into each neural network model to obtain the degree of correlation between the file to be tested and each file category. If the degree of category fit Z is less than the preset degree threshold or the degree of correlation of file category C0 is not the maximum, then issue an early warning for the file category to which the file to be tested belongs and prompt relevant personnel to handle it.
[0026] In this solution, when classifying an archival document across platforms, the criteria for determining whether the classification is reasonable are whether it conforms to the category and whether it is the most appropriate category compared to others.
[0027] The cross-platform data security circulation supervision system includes a module for identifying reviewers, a module for establishing a benchmark set, a module for training a neural network model, and a module for early warning and alerts.
[0028] The module for determining reviewers is used to obtain historical operation records of classifying archival documents into corresponding archival categories as preliminary review records, extract the preliminary reviewers corresponding to the preliminary review records, and determine the reviewers who will reclassify the archival documents from the classification personnel other than the preliminary reviewers. The operation records of the reviewers reclassifying the archival documents are then used as review records.
[0029] The benchmark set establishment module is used to extract the content summary of the archival documents, as well as the archival descriptions written by the initial reviewers and the secondary reviewers after classifying the archival documents. Relationships are extracted to obtain the entities and relationships. Based on the entities and relationships in the operation records corresponding to each archival document, the benchmark set corresponding to each archival category is obtained.
[0030] Neural network model training module: used to build a corresponding neural network model for each archive category. Based on the archive description and content summary corresponding to each archive category, each neural network model is trained to obtain the trained neural network model.
[0031] Early warning module: It is used to extract the archive category, archive description and content summary of the archive file to be detected, and obtain the entity and relationship of the archive file to be detected. Based on the benchmark set and neural network model corresponding to the archive category, it determines whether to issue an early warning for the archive category to which the archive file to be detected belongs.
[0032] Furthermore, the reviewer determination module includes a preliminary review record analysis unit, a degree coefficient calculation unit, and a reviewer determination unit;
[0033] Preliminary review record analysis unit: used to obtain several historical preliminary review records, extract the preliminary reviewer number and file category automatically recorded by the computer program in each preliminary review record, as well as the file description manually recorded by the preliminary reviewer;
[0034] Degree coefficient calculation unit: used to set abnormal files; obtain the total number of preliminary review records corresponding to the classifiers, and the number of file documents that are abnormal files, to obtain the first degree coefficient of the classifiers as file reviewers; extract all keywords from the content summary of the file documents to generate a keyword set; obtain the preliminary review records of several file documents that are not abnormal files corresponding to the classifiers, and combine them with analysis to obtain the second degree coefficient of the classifiers as reviewers;
[0035] Reviewer determination unit: used to obtain the total compliance degree of the classifiers as reviewers, and based on the total compliance degree of each classifier, select the classifier with the highest total compliance degree as the reviewer.
[0036] Furthermore, the early warning module includes a category matching degree calculation unit and an early warning unit;
[0037] Category fit calculation unit: used to extract the archival description and content summary of the preliminary review record of the archival document to be tested, and to obtain the category fit of the archival document to be tested;
[0038] Early warning unit: Used to determine whether to issue an early warning based on the neural network model corresponding to the file category and the degree of category fit.
[0039] Compared with existing technologies, the beneficial effects of this invention are as follows: This invention provides a cross-platform data security circulation supervision system and method, including: acquiring historical preliminary review records of classifying archival documents into corresponding archival categories; extracting preliminary reviewers; identifying reviewers who will reclassify the archival documents; and obtaining review records for reclassification; extracting relationships from both the content summary and description of the archival documents to obtain a baseline set corresponding to each archival category; establishing a neural network model and training each neural network model based on the archival description and content summary; extracting the archival category, description, and content summary corresponding to the archival document to be detected, and determining whether to issue a warning for the archival category to which the archival document to be detected belongs. This invention, by summarizing patterns from historical operation records, determines whether the classification of archival documents is reasonable, which helps to reduce the occurrence of misclassification of archives and improve data security and reliability. Attached Figure Description
[0040] Figure 1 This is a flowchart illustrating the cross-platform data security circulation supervision method of the present invention;
[0041] Figure 2 This is a structural diagram of the cross-platform data security circulation supervision system of the present invention. Detailed Implementation
[0042] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0043] Example: Figure 1 As shown, this invention provides a technical solution for a cross-platform data security circulation supervision method, including the following steps:
[0044] Step S100: Obtain the historical operation records of classifying archival documents into the corresponding archival categories as the initial review records, extract the initial reviewers corresponding to the initial review records, and determine the reviewers who will reclassify the archival documents from the classification personnel other than the initial reviewers, and use the operation records of the reviewers who reclassify the archival documents as the review records.
[0045] Step S110: Obtain several historical preliminary review records. The preliminary review records are records of cross-platform transmission of archive files and initial archive classification. Extract the preliminary reviewer number and archive category automatically recorded by the computer program in each preliminary review record, as well as the archive description manually recorded by the preliminary reviewer. The archive description is a record of the preliminary reviewer's summary of the contents of the archive files and the explanation of the reasons for the classification after completing the classification of the archive files.
[0046] Step S120: If a file that has already been classified is found to have an incorrect classification during subsequent inspection, the file will be classified as an abnormal file; let A be the initial reviewer corresponding to a file F, and let N be the total number of initial review records corresponding to another classification person B. B And the number of abnormal files N0, the classification personnel B as the reviewer of file F, the first degree coefficient T1=(N B -N0) / N B This allows us to obtain the first degree coefficient for each category of personnel, excluding the initial reviewer A.
[0047] Extract all keywords from the content summary of file F to generate a keyword set; obtain the initial review records of several files that are not abnormal files corresponding to classifier B, extract keywords from the file description in each initial review record, and obtain the occurrence frequency of each keyword to generate a vocabulary list; add the occurrence frequencies of words in the vocabulary list that are the same as keywords in the keyword set to obtain the total number of words for classifier B; obtain the total number of words for each other classifier, and normalize them, using the normalized result as the second degree coefficient T2 for classifier B as a reviewer;
[0048] The first degree coefficient is determined based on the classifier's historical classification experience, which is equivalent to the classifier's reliability in classification. The larger the first degree coefficient, the better the classifier is as a reviewer. However, it should also be determined by the second degree coefficient, which is determined by the current archival document and the archival documents classified by the classifier in the past. Generally speaking, the greater the similarity between the current archival document and the archival documents classified by the classifier in the past, the more experience the classifier has in classifying such archival documents. The more experience the classifier has, the more authoritative their classification of archival documents is. Through comprehensive comparison, a reviewer can be selected.
[0049] Step S130: The overall conformity level of classifier B as reviewer is then obtained as T. B=W1×T1+W2×T2, where W1 and W2 are the first weight and the second weight, respectively; obtain the total compliance degree of each category of personnel, and select the category of personnel with the highest total compliance degree as the reviewers; and have the reviewers reclassify the file F.
[0050] Step S200: Extract the content summary of the archival documents, as well as the archival descriptions written by the initial reviewers and the secondary reviewers after classifying the archival documents, and extract the relationships to obtain the entities and relationships; and obtain the benchmark set corresponding to each archival category based on the entities and relationships in the operation records corresponding to each archival document.
[0051] Step S210: Obtain the file description X corresponding to the preliminary review record of a certain file file F. A The file description corresponding to the review record X B And the content summary X, if the archival document F is consistent in the archival category in the initial review and the second review, then the archival description X is obtained. A File Description X B And all entities and relationships between entities in the content summary X, and create a separate profile description X. A File Description X B The knowledge graphs K1, K2, and K3 corresponding to the content summary X are used to extract all element combinations of entity-relationship-entity.
[0052] Step S220: Obtain a combination M1(E1) of elements in K1. M 1,R M 1,E2 M 1) and a combination of elements in K2, M2(E1) M 2,R M 2,E2 M 2), E1 M 1. E2 M 1. E1 M 2 and E2 M Both 2 are entities, R M 1 and R M Both are related, and E1 is obtained respectively. M 1 and E1 M 2. R M 1 and R M 2. E2 M 1 and E2 M Word similarity S between 2 E1 S R S E2 Calculate the combinatorial similarity X between combinatorial combinations M1 and M2. M =S R (S E1 +S E2If X M If the value exceeds a preset first threshold, then combination M1 is marked with a first label to obtain all first label combinations in knowledge graph K1; then, the element combinations in knowledge graph K1 and K3 are compared to obtain all second label combinations in knowledge graph K1; and the element combination that is both a first and second label combination is taken as the target combination of file F.
[0053] By analogy with the element combinations in knowledge graphs K1 and K3, we obtain the following combinations of all second-label combinations in knowledge graph K1:
[0054] Get a combination M1(E1) of elements in K1 M 1,R M 1,E2 M 1) and a combination of elements in K3, M3(E1) M 3,R M 3,E2 M 3), E1 M 1. E2 M 1. E1 M 3 and E2 M 3 are all entities, R M 1 and R M All three are related relationships; obtain E1 respectively. M 1 and E1 M 3. R M 1 and R M 3. E2 M 1 and E2 M Word similarity S between 3 E3 S R1 S E4 Calculate the combinatorial similarity X between combinatorial combinations M1 and M3. M1 =S R1 (S E3 +S E4 If X M1 If the value is greater than the preset second threshold, then combination M1 will be marked with a second label to obtain all combinations of second labels in the knowledge graph K1.
[0055] Step S230: Create a baseline set with empty elements for each archive category; based on the target combination and archive category corresponding to the archive documents that are consistent between the initial review and the final review, add the target combination to the baseline set of the corresponding archive category, and thus obtain the baseline set corresponding to each archive category.
[0056] Step S300: Establish a corresponding neural network model for each archive category. Train each neural network model based on the archive description and content summary corresponding to each archive category to obtain the trained neural network model.
[0057] All archive descriptions and content summaries corresponding to a certain archive category are considered as relevant information, while all archive descriptions and content summaries corresponding to other archive categories are considered as irrelevant information. The correlation between all relevant information and a certain archive category is set to 1, and the correlation between irrelevant information and a certain archive category is set to 0. Then, several relevant and irrelevant information, as well as their respective correlation levels, are used as a dataset and substituted into the neural network model corresponding to a certain archive category for training to obtain the trained neural network model.
[0058] The input to the trained neural network model is the archive description and content summary, and the output is the degree of correlation between the archive file and the archive category corresponding to the neural network model. The degree of correlation ranges from [0,1].
[0059] Step S400: Extract the archive category, archive description and content summary corresponding to the archive file to be detected, and obtain the entity and relationship corresponding to the archive file to be detected. Based on the benchmark set and neural network model corresponding to the archive category, determine whether to issue an early warning prompt for the archive category to which the archive file to be detected belongs.
[0060] Step S410: The archive file to be tested is a file that has completed cross-platform transmission and initial archive category classification. The archive category of the archive file to be tested is taken as C0. The archive description and content summary of the initial review record of the archive file to be tested are extracted, and all element combinations are obtained. The number of element combinations is taken as Y. The combination similarity between the y-th element combination and each combination in the baseline set of archive category C0 is calculated. The maximum combination similarity is taken as the target value of the y-th element combination. The target value of each element combination is obtained, thus determining the category fit of the archive file to be tested. Where e is the natural logarithm, q is the adjustment factor coefficient, and Sy is the target value of the y-th element combination;
[0061] Step S420: Substitute the file description and content summary of the preliminary review record of the file to be tested into each neural network model to obtain the degree of correlation between the file to be tested and each file category. If the degree of category fit Z is less than the preset degree threshold or the degree of correlation of file category C0 is not the maximum, then issue an early warning for the file category to which the file to be tested belongs and prompt relevant personnel to handle it.
[0062] This solution also provides a cross-platform data security circulation supervision system, as shown in the attached document. Figure 2 As shown, it includes:
[0063] The module for determining reviewers is used to obtain historical operation records of classifying archival documents into corresponding archival categories as preliminary review records, extract the preliminary reviewers corresponding to the preliminary review records, and determine the reviewers who will reclassify the archival documents from the classification personnel other than the preliminary reviewers. The operation records of the reviewers reclassifying the archival documents are then used as review records.
[0064] The benchmark set establishment module is used to extract the content summary of the archival documents, as well as the archival descriptions written by the initial reviewers and the secondary reviewers after classifying the archival documents. Relationships are extracted to obtain the entities and relationships. Based on the entities and relationships in the operation records corresponding to each archival document, the benchmark set corresponding to each archival category is obtained.
[0065] Neural network model training module: used to build a corresponding neural network model for each archive category. Based on the archive description and content summary corresponding to each archive category, each neural network model is trained to obtain the trained neural network model.
[0066] Early warning module: It is used to extract the archive category, archive description and content summary of the archive file to be detected, and obtain the entity and relationship of the archive file to be detected. Based on the benchmark set and neural network model corresponding to the archive category, it determines whether to issue an early warning for the archive category to which the archive file to be detected belongs.
[0067] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.
Claims
1. A cross-platform-based data security circulation supervision method, characterized in that, The method comprises the following steps: Step S100: obtaining historical operation records of filing files into corresponding archive categories as preliminary examination records, extracting preliminary examination personnel corresponding to the preliminary examination records, determining review personnel for re-classification of archive categories of the archive files from classification personnel other than the preliminary examination personnel, and obtaining operation records of the review personnel for re-classification of the archive files as review records; Step S200: extracting a content summary of the archive files, and archive descriptions written by the preliminary examination personnel and the review personnel after classification of the archive files, and performing relation extraction to obtain entities and associated relations therein; and obtaining a reference set corresponding to each archive category according to the entities and associated relations in the operation records corresponding to each archive file; Step S300: establishing a corresponding neural network model for each archive category, training each neural network model according to the archive descriptions and the content summaries corresponding to each archive category, and obtaining a trained neural network model; Step S400: extracting an archive category, an archive description and a content summary corresponding to a to-be-detected archive file, and obtaining entities and associated relations corresponding to the to-be-detected archive file, and judging whether to give a pre-warning prompt for an archive category to which the to-be-detected archive file belongs according to the reference set and the neural network model corresponding to the archive category; Step S410: The file to be detected is the file that has completed cross-platform transmission and completed the first archive category classification, taking the archive category of the file to be detected as C0, extracting the archive description and content summary of the preliminary examination record of the file to be detected, obtaining all element combinations therein, taking the number of element combinations as Y, calculating the combination similarity between the yth element combination and each combination in the reference set of the archive category C0, taking the maximum value of the combination similarity as the target value of the yth element combination, obtaining the target value of each element combination, and further obtaining the category fitting degree of the file to be detected wherein e is a natural logarithm, q is an adjustment factor coefficient, and Sy is the target value of the yth element combination. Step S420: substituting the archive description and the content summary of the preliminary examination record of the to-be-detected archive file into each neural network model to obtain a correlation degree between the to-be-detected archive file and each archive category, and giving a pre-warning for the archive category to which the to-be-detected archive file belongs and prompting relevant personnel to handle if a category fitting degree Z is less than a preset degree threshold or a correlation degree of the archive category C0 is not the maximum.
2. The cross-platform-based data security circulation supervision method according to claim 1, characterized in that, Step S100 comprises: Step S110: obtaining historical preliminary examination records of cross-platform transmission of archive files and first archive category classification, extracting a preliminary examination personnel number and an archive category recorded automatically by a computer program and an archive description recorded manually by a preliminary examination personnel in each preliminary examination record, and the archive description is a content record of summarizing and elaborating on the classification reason and the content in the archive file after the preliminary examination personnel completes classification of the archive file; Step S120: If a file classified by a certain category is found to have a classification error in the subsequent detection process, the file is regarded as an abnormal file; set the initial reviewer corresponding to a certain file F as A, and according to the total number N of initial review records corresponding to another reviewer B B , and the number N0 of abnormal files, obtain the first degree coefficient T1=(N B -N0) / N B of reviewer B as the review personnel of file F, and then obtain the first degree coefficient of each reviewer other than the initial reviewer A. extracting all keywords in the content summary of the archive file F to generate a keyword set, obtaining preliminary examination records of archive files corresponding to the classification personnel B and not being abnormal archives, extracting keywords in the archive description in each preliminary examination record, and comprehensively obtaining the number of occurrences of each keyword to generate a vocabulary table, adding the number of occurrences of the same words in the vocabulary table and the keyword set to obtain the total word quantity of the classification personnel B, and obtaining the total word quantity of each classification personnel and normalizing to take the normalized result as a second degree coefficient T2 of the classification personnel B as a review personnel; Step S130: further obtaining the total coincidence degree T of the classified personnel B as the review personnel B =W1×T1+W2×T2, wherein W1 and W2 are the first weight and the second weight respectively; obtaining the total coincidence degree of each classified personnel, taking the classified personnel with the largest total coincidence degree as the review personnel; and reclassifying the archive file F by the review personnel. 3.The cross-platform based data security circulation supervision method according to claim 1, characterized in that, Step S200 comprises: Step S210: Obtain the archive description X corresponding to the record of the certain archive file F in the first instance A , the record of the reexamination, the archive description X B , and the content summary X. If the archive file F is consistent in the archive category in the first instance and the reexamination, obtain all entities and the association relationship between the entities in the archive description X A , the archive description X B , and the content summary X, and respectively establish the knowledge graph K1, K2, and K3 corresponding to the archive description X A , the archive description X B , and the content summary X, and extract all entity-association relationship-entity element combinations therein. Step S220: obtaining a certain element combination M1 (E1 M 1,R M 1,E2 M 1) and a certain element combination M2 (E1 M 2,R M 2,E2 M 2), E1 M 1, E2 M 1, E1 M 2 and E2 M 2 are entities, R M 1 and R M 2 are association relationships, respectively obtaining the word similarity S M , S M , S M between E1 M 1 and E1 M 2, R M 1 and R E1 2, E2 R 1 and E2 E2 2, calculating the combination similarity X M between the combination M1 and the combination M2 = S R (S E1 +S E2 ); if X M is greater than a preset first threshold, marking the combination M1 as the first mark, obtaining all the first marked combinations in the knowledge graph K1; further comparing the element combinations in the knowledge graphs K1 and K3 to obtain all the second marked combinations in the knowledge graph K1; and taking the element combination which is simultaneously the first and second marked combination as the target combination of the archive file F; Step S230: establishing a reference set with no elements for each file category; adding the target combination to the reference set of the corresponding file category according to the target combination corresponding to the file of the same file category of each preliminary review and review, and then obtaining the reference set corresponding to each file category.
4. The cross-platform-based data security circulation supervision method according to claim 3, characterized in that, Step S300 includes: taking all file descriptions and content summaries corresponding to a certain file category as relevant information, and taking all file descriptions and content summaries corresponding to other file categories as irrelevant information; setting the relevance of all relevant information to the certain file category as 1, and setting the relevance of all irrelevant information to the certain file category as 0; then taking the relevant information and irrelevant information and the respective relevance as a data set, and inputting the data set into the neural network model corresponding to the certain file category for training to obtain the trained neural network model.
5. A data security circulation supervision system for performing the cross-platform based data security circulation supervision method according to any one of claims 1-4, characterized in that, The system includes a reviewer determining module, a reference set establishing module, a neural network model training module, and a pre-warning prompting module; The reviewer determining module is configured to obtain historical operation records of classifying file into corresponding file categories as preliminary review records, extract preliminary reviewers corresponding to the preliminary review records, determine review reviewers for re-classifying file categories from classification personnel other than the preliminary reviewers, and obtain operation records of the review reviewers re-classifying file as review records; The reference set establishing module is configured to extract content summaries of file, and file descriptions written after classification of file by preliminary reviewers and review reviewers, and perform relationship extraction on the content summaries and file descriptions to obtain entities and association relationships therein; and obtain reference sets corresponding to each file category according to the entities and association relationships in the operation records corresponding to each file; The neural network model training module is configured to establish a corresponding neural network model for each file category, train each neural network model according to file descriptions and content summaries corresponding to each file category, and obtain a trained neural network model; The pre-warning prompting module is configured to extract file categories, file descriptions, and content summaries corresponding to a to-be-detected file, obtain entities and association relationships corresponding to the to-be-detected file, and determine whether to pre-warn the file category to which the to-be-detected file belongs according to the reference set and the neural network model corresponding to the file category.
6. The data security in-transit monitoring system of claim 5, wherein, The reviewer determining module includes a preliminary review record analysis unit, a degree coefficient calculation unit, and a review reviewer determining unit; The preliminary review record analysis unit is configured to obtain historical preliminary review records, extract preliminary reviewer numbers and file categories recorded automatically by computer programs, and file descriptions recorded manually by preliminary reviewers in each preliminary review record. The degree coefficient calculation unit is configured to set the abnormal archives, obtain the total number of the initial review records corresponding to the classification personnel and the number of the archive files being the abnormal archives, and obtain the first degree coefficient of the classification personnel as the archive file review personnel; extract all keywords in the content summary of the archive file to generate a keyword set; obtain the initial review records of a plurality of archive files corresponding to the classification personnel which are not abnormal archives, and obtain the second degree coefficient of the classification personnel as the review personnel in combination with analysis; The review personnel determination unit is configured to obtain the total coincidence degree of the classification personnel as the review personnel, and determine the classification personnel with the maximum total coincidence degree as the review personnel according to the total coincidence degree of each classification personnel.
7. The data security in-transit monitoring system of claim 5, wherein, The early warning prompt module comprises a category coincidence degree calculation unit and an early warning prompt unit. The category coincidence degree calculation unit is configured to extract the archive description and the content summary of the initial review record of the to-be-detected archive file, and obtain the category coincidence degree of the to-be-detected archive file. The early warning prompt unit is configured to determine whether to perform early warning prompt on the archive category to which the to-be-detected archive file belongs according to the neural network model corresponding to the archive category and the category coincidence degree.
Citation Information
Patent Citations
Full-process archive data security management system and method
CN119004543A
Archive management system, method and equipment for intelligent miniature archive room
CN119226226A