Information management device and method for intelligent archive
By cutting and classification of the FF1 algorithm and Yamamoto method of the information management device of the intelligent archives, combined with the Laplace mechanism algorithm, the balance problem of archive information between confidentiality protection and information public confidence is solved, and the accuracy of mobile confidentiality handling of confidentiality of archive information is improved.
Patent Information
- Application Number
- CN202510533851.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-08-01
AI Technical Summary
In the prior art, there is a lack of balanced handling between confidential information protection and information public confidence in the information management device of intelligent archives, resulting in low accuracy in mobile confidential handling of confidential information of archives.
By performing upload feasibility checks on uploaded files, the FF1 algorithm and Yamamoto method are used to cut and classify archive information, and combined with the Laplace mechanism algorithm, the information of archive objects in the group is accurately and confidential, forming an uploaded file to be uploaded.
It realizes accurate confidentiality disposal of archive information, improves the accuracy of maneuver confidentiality disposal of archive information, reduces the impact of global noise on confidence, and improves the security of information management.
Smart Images

Figure CN120408713A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of information management, and specifically relates to an information management device and method for intelligent archives. Background Art
[0002] Intelligent archives are archive management systems constructed using modern information technologies. Through digital and network means, centralized management and efficient utilization of archives are achieved. It not only includes the storage and retrieval of archives, but also covers multiple aspects such as the security protection, sharing, and services of archives.
[0003] In practical applications, current information management devices for intelligent archives often use the prior art solutions mentioned in the patent publication number "CN118862036B" to perform upload feasibility checks on files to be uploaded, and then perform intelligent classification on the files to be uploaded that are permitted to be uploaded.
[0004] Performing an upload feasibility check on the file to be uploaded means that because the file information of different types of files to be uploaded has different confidentiality levels and importance, just as the disciplinary material information in the file information of the file to be uploaded often involves the experience and privacy of the file object, and the party and league information in the party and league material information of the file information of the file to be uploaded involves the organizational mobility of the file object, detailed confidentiality information identification and protection rules are required.
[0005] Currently, during the upload feasibility check of the file to be uploaded, it is necessary to identify and protect confidential information. However, the addition of random noise in the protection does not fully consider the differences in the confidentiality changes of the information of different types of file objects in the specific file information, resulting in frequent deficiencies in the balance between the confidentiality protection of the file information of the currently to-be-uploaded file and the confidence level of information sharing, and the accuracy of the dynamic confidentiality processing of the confidential information of the file information of the to-be-uploaded file is not high. Summary of the Invention
[0006] To solve the defects existing in the prior art, the present invention proposes an information management device and method for intelligent archives, effectively avoiding the frequent deficiencies in the balance between the confidentiality protection of the file information of the to-be-uploaded file and the confidence level of information sharing in the prior art, and the low accuracy of the dynamic confidentiality processing of the confidential information of the file information of the to-be-uploaded file.
[0007] The present invention uses the following technical solutions.
[0008] An information management method for intelligent archives, comprising: Performing an upload feasibility check on the file to be uploaded, and then performing intelligent classification on the file to be uploaded that is permitted to be uploaded; The method for performing an upload feasibility check on the file to be uploaded includes: Step 1. Obtain several types of file information of the files to be uploaded corresponding to each file object, and perform modulation processing on the confidentiality characteristic information; Furthermore, in Step 1, initially obtain the information of the files to be uploaded in the file object that has not performed the upload feasibility check, that is, obtain the file information of the files to be uploaded corresponding to [[[NUMBER]]] file objects. For the non-numeric information in the file information of the files to be uploaded here, take its GBK code to represent it. For the numeric information in the file information of the files to be uploaded here, the number of samples of each type of file information is .
[0009] Furthermore, in Step 1, the number of file objects is one thousand.
[0010] Furthermore, in Step 1, The value of [[[VARIABLE]]] can be five hundred.
[0011] Furthermore, in Step 1, the file information includes eleven types of file information such as the name, gender, age, height, vital capacity, address, salary, information on materials for joining political parties and organizations, information on disciplinary materials, and information on reward materials of the file object, which are sequentially represented by character codes respectively.
[0012] Furthermore, in Step 1, for the confidentiality characteristic information in the information of the files to be uploaded perform initial processing. Therefore, apply the FF1 algorithm to the file information of each file object to perform initial processing on the confidentiality characteristic information to obtain the modulation values of the file information of each file object, and maintain the original information font specifications and sizes.
[0013] Furthermore, in Step 1, take the queue formed by arranging the file information of each type of file object of the file object in the order of their generation time points as the information queue of each type of file information.
[0014] Step 2. According to the types of file information, divide the other file information of each file object except the confidentiality characteristic information into Information One and Information Two. Use the Yamamoto method to perform cutting on the file information of each type in Information One at the file information of their entire generation time points to obtain the sub-queues of the file information of each type in Information One; Determine the discrimination factors of each sub-queue based on the differences between each sub-queue and other sub-queues; Furthermore, in Step 2, regard all the numeric information in the remaining file information except the confidentiality characteristic information as Information One, and regard all the non-numeric information in the remaining file information as Information Two.
[0015] Further, in Step2, for Information 1 and Information 2 of each file object, for the information queues of each type of file information in Information 1, use the Yamamoto method for each information queue to obtain the segmentation information of each information queue. For each sub-queue in the segmentation information, calculate the L2 norm between each sub-queue and each other sub-queue.
[0016] Further, in Step2, the discrimination factor has the following operation equation: ; where represents the discrimination factor of the th sub-queue of the th type of file information in Information 1; represents the L2 norm between the th sub-queue and the th sub-queue of the th type of file information in Information 1; represents the quantity obtained by subtracting the number of file information of the th sub-queue from the number of file information of the th sub-queue of the th type of file information in Information 1, represents the Euler number; represents the number of sub-queues of the th type of file information in Information 1.
[0017] Step3, construct the information vectors of each file information in Information 1 according to the discrimination factors of each information sub-queue in Information 1, analyze the approximation of the information vectors of each file information in Information 1 and other information, and determine the information confidentiality factors of each file information in Information 1; according to the difference range between each file information in Information 2 and other file information, determine the status comparison factors of each file information in Information 2; combine the information confidentiality factors of all file information at all generation time points in Information 1 of all file objects and the status comparison factors of all generation time point information in Information 2, and perform segmentation on the file objects; Further, in Step3, for each type of file information in Information 1 of each file object, take the average of the approximations corresponding to the information vectors between each type of file information and all other generation time point information as the information confidentiality factor of each type of file information, and the approximation corresponding to the information vector is the Pearson coefficient.
[0018] Further, in Step3, for the information two of each file object, for the information queues of each type of file information in the information two, the average of the differences between each file information in the information two and all other generation time point information in the information two corresponding to the information queue is used as the condition comparison factor of each file information in the information two.
[0019] Further, in Step3, the average of the information confidentiality factors of all generation time point information in the information one of each file object is used as the characteristic quantity one of each file object; the average of the condition comparison factors of the file information at all generation time points in the information two of each file object is used as the characteristic quantity two of each file object. The characteristic quantity one and the characteristic quantity two are respectively used as the values on the X-axis and the Y-axis in the corresponding Cartesian coordinate system of each file object, so as to form coordinate points.
[0020] Through the above analysis, in this application, the characteristic quantity one and the characteristic quantity two are respectively used as the values on the X-axis and the Y-axis in the corresponding Cartesian coordinate system of each file object, so as to form the coordinate points corresponding to the file object. The DBSCAN algorithm is used to perform grouping and segmentation on all file objects.
[0021] Step4, use the Laplace mechanism algorithm to perform confidentiality processing on the information of the file objects in each group to form the files to be uploaded that are allowed to be uploaded.
[0022] An information management device for intelligent files includes: A modulation module, which is used to obtain several types of file information of the files to be uploaded corresponding to each file object, and perform modulation processing on the confidentiality characteristic information; A cutting module, which is used to divide the other file information of each file object except the confidentiality characteristic information into information one and information two according to the type of file information, and use the Yamamoto method to cut each type of file information in the information one at all its generation time points of the file information to obtain the sub-queues of each type of file information in the information one; determine the difference factors of each sub-queue according to the differences between each sub-queue and other sub-queues; A comparison module, which is used to construct the information vectors of each file information in the information one according to the difference factors of each information sub-queue in the information one, analyze the approximation of the information vectors of each file information in the information one and other information, and determine the information confidentiality factors of each file information in the information one; determine the condition comparison factors of each file information in the information two according to the difference range between each file information in the information two and other file information; combine the information confidentiality factors of all generation time point information in the information one of all file objects and the condition comparison factors of all generation time point information in the information two to perform segmentation on the file objects; A confidentiality module, which is used to perform confidentiality processing on the information of file objects in each group by using the Laplace mechanism algorithm to form the files to be uploaded that are allowed to be uploaded.
[0023] The beneficial effects of the present invention are that, compared with the prior art, the technical effects of the present invention include: Obtain several types of file information of the corresponding files to be uploaded of each file object, and perform modulation processing on the confidentiality characteristic information; divide the other file information of each file object except the confidentiality characteristic information into Information One and Information Two according to the type of file information, and use the Yamamoto method to perform cutting on the file information of each type in Information One at the file information at the time of its overall generation to obtain sub-queues of the file information of each type in Information One; determine the discrimination factors of each sub-queue according to the differences between each sub-queue and other sub-queues; construct information vectors of each file information in Information One according to the discrimination factors of each information sub-queue in Information One, analyze the approximation of the information vectors of each file information in Information One and other information, and determine the information confidentiality factors of each file information in Information One; determine the situation comparison factors of each file information in Information Two according to the difference range between each file information in Information Two and other file information; combine the information confidentiality factors of the file information at the overall generation time of all of Information One of all file objects and the situation comparison factors of the information at the overall generation time of Information Two to perform segmentation on the file objects; perform confidentiality processing on the information of file objects in each group by using the Laplace mechanism algorithm to form the files to be uploaded that are allowed to be uploaded; accordingly, combine the confidentiality information differences reflected by the file objects to perform precise segmentation on the information with similar confidentiality characteristics at different times, so as to perform precise Laplace mechanism processing on the information of file objects, weaken the effect of the discrete Laplace mechanism processing with too high globalization effect during the privacy processing on the confidence difference of the information after confidentiality processing, and thus improve the precision of the flexible confidentiality processing of the file information of the files to be uploaded. Description of the Drawings
[0024] Figure 1 is the flowchart of the information management method for intelligent files described in the present invention; Figure 2 is the partial structure diagram of the information management device for intelligent files described in the present invention. Detailed Embodiments
[0025] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will, in conjunction with the accompanying drawings in the embodiments of the present invention, clearly and completely describe the technical solutions of the present invention. The embodiments described in this application are only partial embodiments of the present invention, not all embodiments. According to the spirit of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0026] As Figure 1 shown, an information management method for intelligent archives according to the present invention includes: Performing an upload feasibility check on the files to be uploaded, and then performing intelligent classification on the files to be uploaded that are permitted to be uploaded; The method for performing an upload feasibility check on the files to be uploaded includes: Step1, obtaining several types of file information of the files to be uploaded corresponding to each file object, and modulating and processing the confidentiality characteristic information; In a preferred but non-limiting embodiment of the present invention, in Step1, initially obtain the information of the files to be uploaded that have not undergone an upload feasibility check within the file object, that is, obtain the file information of the files to be uploaded corresponding to each file object. For non-numeric information such as the name, gender, and award material information in the file information of the files to be uploaded, take its GBK code to represent. For numeric information such as age, height, and vital capacity in the file information of the files to be uploaded here, the number of samples for each type of file information is .
[0027] In a preferred but non-limiting embodiment of the present invention, in Step1, the number of file objects is one thousand.
[0028] In a preferred but non-limiting embodiment of the present invention, in Step1, the value of can be five hundred.
[0029] In a preferred but non-limiting embodiment of the present invention, in Step1, the file information includes eleven types of file information such as the name, gender, age, height, vital capacity, address, salary, information on joining the Party or Youth League materials, disciplinary action materials information, and award material information of the file object, and are sequentially represented by the character codes .
[0030] In a preferred but non-limiting embodiment of the present invention, in Step1, for the confidentiality characteristic information in the information of the files to be uploaded, perform initial processing. Just like the name, gender, and address of the file object, so apply the FF1 algorithm to the file information of each file object for the confidentiality characteristic information Execute the starting disposal, obtain the modulation values of the file information of each file object, and maintain the original information font specifications and sizes.
[0031] In a preferred but non-limiting embodiment of the present invention, in Step1, it is necessary to further analyze the differences in the changes of the confidential information of the file information, and use the queue formed by arranging the file information of each type of file object in the order of their generation time points (that is, the creation time points of the file information) as the information queue of each type of file information.
[0032] Thus, the information queues of all file objects at the generation time points are obtained. [[ID=,7]]
[0033] Step2, according to the types of file information, divide the other file information of each file object except for the confidential characteristic information into Information One and Information Two, use the Yamamoto method to perform cutting on the file information of each type in Information One at the file information of all its generation time points, and obtain the sub-queues of the file information of each type in Information One; determine the difference factors of each sub-queue according to the differences between each sub-queue and other sub-queues; Generally, when using the Laplace mechanism to perform confidentiality processing on the coded file information, there will be a globalization characteristic, that is, adding random noise to the file information at all generation time points, resulting in an enhanced globalization situation of the noise processing of the file information at all generation time points, and often resulting in poor confidence of the file information at all generation time points after executing the Laplace mechanism; and the de-globalization method often uses the information of each file object in the file information of the file to be uploaded to perform Laplace mechanism processing, and the addition of its random noise is affected by the different time-varying differences of different file object information, resulting in a large error in the noise processing of the file information, not only increasing the complexity of the file information processing, but also not significantly improving the confidence of the processed file information compared to the globalization processing.
[0034] In the face of the above defects, this application obtains the file information of the file to be uploaded, and performs the FF1 algorithm on the information with confidentiality characteristics in the information, and then takes into account the different confidentiality change differences in the information manifestations of the file information of different file objects at different specific generation times, can perform a comparative analysis on the information change characteristics of the file objects, and perform characteristic segmentation on the file information of the file objects through the comparative analysis values.
[0035] In a preferred but non-limiting embodiment of the present invention, in Step 2, initially perform feature extraction via the file information corresponding to each file object, and perform precise analysis on the confidentiality feature of the file information of the file object according to the value of the feature extraction, that is, perform category segmentation on the file information of the file object. Here, the numerical information (age, height, vital capacity, salary, numerical information in the reward material information) in the remaining file information outside the confidentiality feature information is regarded as Information One, and the non-numerical information (text information in the Party / League membership application materials, disciplinary action materials, and reward material information) in the remaining file information is regarded as Information Two. What segments Information One and Information Two is to analyze according to the relevant differences in the power information of the file object at different times, highlighting the confidentiality feature of the file information of the file object.
[0036] In a preferred but non-limiting embodiment of the present invention, in Step 2, for Information One and Information Two of each file object, for the information queue of each type of file information in Information One, use the Yamamoto method for each information queue to obtain the segmentation information of each information queue, which is suitable for analyzing the differences in short-term information changes of each Information One of the file object. For each sub-queue in the segmentation information, calculate the L2 norm between each sub-queue and other sub-queues. The higher the L2 norm, the higher the difference in the file information changes at different times.
[0037] In a preferred but non-limiting embodiment of the present invention, in Step 2, for each type of file information in Information One, calculate the difference between each sub-queue corresponding to each type of file information and the number of file information corresponding to other sub-queues in the queue. The higher the number difference, the higher the difference in the time period size of the action interval of the file information changes at different times, that is, the higher the difference in the characteristics of the file object information in different time periods (the time period formed between the generation time points) presented by the information changes, and the smaller the probability that the current file information reflects the confidentiality information characteristics of the file object; through the comparison differences of each type of file information of each file object at different times, calculate the difference factor for comparing the confidentiality characteristics of the file object information The calculation equation is: ; where represents the difference factor of the th sub-queue of the th type of file information in Information One; represents the L2 norm between the th sub-queue and the th sub-queue of the th type of file information in Information One; represents the th type of file information in Information One The quantity obtained by subtracting the number of archive information of the th sub-queue from the number of archive information of the nth sub-queue represents the Euler number; Represents the number of sub-queues of the archive information of the th category within Information 1; the higher the obtained discrimination factor, the higher the difference in partial characteristics of the current archive object's information with the change of time point during the analysis of the information confidentiality characteristics, indicating that the information confidentiality characteristic of the archive object's information is relatively low.
[0038] Step3. Construct information vectors of each archive information within Information 1 according to the discrimination factors of each information sub-queue within Information 1, analyze the approximation of the information vectors of each archive information within Information 1 and other information, and determine the information confidentiality factors of each archive information within Information 1; according to the difference range between each archive information within Information 2 and other archive information, determine the status comparison factors of each archive information within Information 2; combine the information confidentiality factors of the archive information at all generation time points within Information 1 and the status comparison factors of the information at all generation time points within Information 2 of all archive objects to perform segmentation on the archive objects; According to the above method, the difference in information changes at different time periods of the information of each archive object can be analyzed. Then, for each category of archive information within the information of each archive object, the vector formed by the discrimination factors of all the sub-queues at the corresponding generation time points of each category of archive information (the elements of the vector are the discrimination factors of all the sub-queues at the generation time points) is used as the information vector of each category of archive information for performing confidentiality characteristic analysis, that is, according to the difference in information changes at different time periods of different categories of archive information of the archive object during the upload feasibility check, the confidentiality characteristics reflected by each category of archive information are accurately analyzed.
[0039] In a preferred but non-limiting embodiment of the present invention, in Step3, for each category of archive information within Information 1 of each archive object, the average of the approximation corresponding to the information vector between each category of archive information and other information at all generation time points is used as the information confidentiality factor of each category of archive information. In this application, the approximation corresponding to the information vector is the Pearson coefficient. The higher the information confidentiality factor, the closer the difference characteristics of the information between the current archive information and other archive information with the change of time at different time periods during the analysis of the information confidentiality characteristics of each archive object's information, that is, the more obvious the corresponding information confidentiality characteristic.
[0040] In a preferred but non-limiting embodiment of the present invention, in Step 3, to determine the information characteristics of each file object, for the information two of each file object, for the information queues of each type of file information in the information two, the average of the differences between each file information in the information two and all other generation time point information in the information two corresponding to the information queue is used as the condition comparison factor of each file information in the information two. Just as the average of the L2 norms corresponding to the information queue between each type of file information in the information two and all other generation time point information in the information two is used as the condition comparison factor of each type of file information in the information two, the higher the condition comparison factor, the higher the difference in the change of the information state of the file object, and the more obvious the confidentiality characteristic of the file object information.
[0041] In a preferred but non-limiting embodiment of the present invention, in Step 3, to accurately segment the file object information, the average of the information confidentiality factors of all generation time point information in the information one of each file object is used as the characteristic quantity one of each file object; the average of the condition comparison factors of all generation time point file information in the information two of each file object is used as the characteristic quantity two of each file object. The characteristic quantity one and the characteristic quantity two are respectively used as the values on the X-axis and the Y-axis in the corresponding Cartesian coordinate system of each file object, so as to form coordinate points, which is to combine the confidentiality change difference characteristics of digital information and non-digital information in each file object to identify the information difference between different file objects, and is used to segment the file object information, so as to accurately identify the information characteristics between file objects.
[0042] Through the above analysis, in this application, the characteristic quantity one and the characteristic quantity two are respectively used as the values on the X-axis and the Y-axis in the corresponding Cartesian coordinate system of each file object, so as to form the coordinate points corresponding to the file object. The DBSCAN algorithm is used to perform grouped segmentation on all file objects. Here, to accurately compare the confidentiality characteristics of different file object information, during the grouping process of this application, the information confidentiality factor of all generation time point information in the information one of each file object and the condition comparison factor of all generation time point information in the information two are used to form the identification vector of each file object (the elements of an identification vector are the binary group formed by an information confidentiality factor and a condition comparison factor at the corresponding time point), and the quantity obtained by multiplying the approximation quantity (Pearson coefficient) of the identification vectors between different file objects by the L2 norm between the coordinate points corresponding to different file objects is used as the radius between the DBSCAN algorithms. Finally, the segmentation information of all generation time point file objects can be obtained.
[0043] Then, according to the comparative analysis of the information characteristics of the file objects, the file objects are segmented.
[0044] Step 4: Apply the Laplace mechanism algorithm to the information of the file objects in each group to perform confidentiality processing to form the files to be uploaded that are allowed to be uploaded.
[0045] The file objects are segmented in the above manner, and then the information of the file objects is segmented accordingly. The information queues of all the generated-time file objects in each group after the group segmentation form the information groups of each group. The Laplace mechanism algorithm is used to process the information in the information groups of each group.
[0046] Therefore, according to the above manner of the present application, the Laplace mechanism can be used to process the information groups in each group respectively, which can improve the accuracy of the information processing of each file object. Through the confidentiality processing of the information of the file objects in each group, the confidentiality processing performance of the file object information is improved.
[0047] As Figure 2 shown, an information management device for intelligent files according to the present invention includes: A modulation module, which is used to obtain several types of file information of the files to be uploaded corresponding to each file object, and perform modulation processing on the confidentiality characteristic information; A cutting module, which is used to divide the other file information of each file object except the confidentiality characteristic information into Information One and Information Two according to the types of file information, and use the Yamamoto method to cut the file information of each type in Information One at the file information of its entire generation time to obtain sub-queues of the file information of each type in Information One; determine the difference factors of each sub-queue according to the difference between each sub-queue and other sub-queues; A comparison module, which is used to construct information vectors of each file information in Information One according to the difference factors of each information sub-queue in Information One, analyze the approximation of the information vectors of each file information in Information One and other information, and determine the information confidentiality factors of each file information in Information One; determine the situation comparison factors of each file information in Information Two according to the difference range between each file information in Information Two and other file information; combine the information confidentiality factors of the file information of all the generated-time in Information One of all the file objects and the situation comparison factors of the information of all the generated-time in Information Two to segment the file objects; A confidentiality module, which is used to apply the Laplace mechanism algorithm to the information of the file objects in each group to perform confidentiality processing to form the files to be uploaded that are allowed to be uploaded.
[0048] The beneficial effects of the present invention are that, compared with the prior art, the technical effects of the present invention include: Obtain several types of file information of the files to be uploaded corresponding to each file object, and modulate and process the confidentiality characteristic information; divide the other file information of each file object except the confidentiality characteristic information into Information One and Information Two according to the types of file information, and use the Yamamoto method to perform cutting on the file information of each type in Information One at the file information at the time of its overall generation to obtain sub-queues of the file information of each type in Information One; determine the difference factors of each sub-queue according to the differences between each sub-queue and other sub-queues; construct information vectors of each file information in Information One according to the difference factors of each information sub-queue in Information One, analyze the approximation of the information vectors of each file information in Information One and other information, and determine the information confidentiality factors of each file information in Information One; determine the situation comparison factors of each file information in Information Two according to the difference ranges between each file information in Information Two and other file information; combine the information confidentiality factors of the file information at the overall generation time in Information One of all file objects and the situation comparison factors of the information at the overall generation time in Information Two to perform segmentation on the file objects; perform confidentiality processing on the information of the file objects in each group using the Laplace mechanism algorithm to form the files to be uploaded that are allowed to be uploaded; accordingly, combine the confidentiality information differences reflected by the file objects to perform precise segmentation on the information with similar confidentiality characteristics at different times, so as to perform precise Laplace mechanism processing on the information of the file objects, weaken the influence of the discrete Laplace mechanism processing with too high globalization effect during the privacy processing on the confidence difference of the information after confidentiality processing, so as to improve the precision of the mobile confidentiality processing of the file information of the files to be uploaded.
[0049] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: still can modify the specific implementation manners of the present invention or make equivalent replacements, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered within the protection scope of the claims of the present invention.
Claims
1. An information management method for intelligent archives, characterized in that, Including: Perform an upload feasibility check on the files to be uploaded, and then perform intelligent classification on the files to be uploaded that are permitted to be uploaded; A method for performing an upload feasibility check on the files to be uploaded, including: Step1, Obtain several types of file information of the files to be uploaded corresponding to each file object, and perform modulation processing on the confidentiality characteristic information; Step2, According to the types of file information, divide the other file information of each file object except the confidentiality characteristic information into Information One and Information Two. Use the Yamamoto method to perform cutting on the file information of each type in Information One at the time point of its overall generation to obtain sub-queues of the file information of each type in Information One. Determine the discrimination factors of each sub-queue according to the differences between each sub-queue and other sub-queues; Step3, Construct information vectors of the file information of each type in Information One according to the discrimination factors of each information sub-queue in Information One, analyze the approximation amounts of the information vectors of the file information of each type in Information One and other information, and determine the information confidentiality factors of the file information of each type in Information One. According to the difference ranges between the file information of each type in Information Two and other file information, determine the situation comparison factors of the file information of each type in Information Two. Combine the information confidentiality factors of the file information of all file objects in Information One at the overall generation time point and the situation comparison factors of the information of all file objects in Information Two at the overall generation time point to perform segmentation on the file objects; Step4, Perform confidentiality processing on the information of the file objects in each group using the Laplace mechanism algorithm to form the files to be uploaded that are permitted to be uploaded.
2. The information management method for intelligent archives according to claim 1, wherein, In Step1, initially obtain the information of the files to be uploaded within the file object that have not undergone the upload feasibility check, that is, obtain the file information of the files to be uploaded corresponding to the file objects. For the non-numeric information in the file information of the files to be uploaded here, take its GBK code for representation. For the numeric information in the file information of the files to be uploaded here, the number of samples for each type of file information is .
3. The information management method for intelligent archives according to claim 2, wherein, In Step 1, the number of file objects is one thousand; In Step1, The value can be five hundred; In Step1, the file information includes eleven types of file information, namely, the name, gender, age, height, vital capacity, address, salary, information on materials for joining political parties or youth leagues, information on disciplinary materials, and information on award materials of the file object, which are sequentially represented by the character code respectively; In Step1, in the information of the file to be uploaded, face the confidentiality characteristic information Execute the initial processing, so apply the FF1 algorithm to the file information of each file object for the confidentiality characteristic information Execute the initial processing, obtain the modulation value of the file information of each file object, and keep the original information font specifications and sizes.
4. The information management method for intelligent archives according to claim 3, wherein, In Step1, the queue formed by arranging the file information of each type of file object in the order of their generation time points is used as the information queue of the file information of each type.
5. The information management method for intelligent archives according to claim 4, characterized in that In Step2, all the numerical information in the remaining file information except the confidentiality characteristic information is regarded as Information One, and all the non-numerical information in the remaining file information is regarded as Information Two.
6. The information management method for intelligent archives according to claim 5, characterized in that, In Step2, for Information One and Information Two of each file object, for the information queue of the file information of each type in Information One, use the Yamamoto method for each information queue to obtain the segmentation information of each information queue. For each sub-queue in the segmentation information, calculate the L2 norm between each sub-queue and other sub-queues.
7. The information management method for intelligent archives according to claim 6, wherein, In Step 2, the operation equation of the discrimination factor is as follows: ; Here, The discrimination factor of the nth sub-queue of the file information of the first type of characterization information; The L2 norm between the nth sub-queue and the mth sub-queue of the file information of the first type of characterization information; The quantity obtained by subtracting the number of file information of the nth sub-queue from the number of file information of the mth sub-queue of the file information of the first type of characterization information, representing the Euler number; The number of sub-queues of the file information of the first type of characterization information.
8. The information management method for intelligent archives according to claim 7, wherein, In Step3, for the file information of each type in Information One of each file object, regard the average of the approximation amounts corresponding to the information vectors between the file information of each type and all other information at the generation time point as the information confidentiality factor of the file information of each type. The approximation amount corresponding to the information vector is the Pearson coefficient.
9. The information management method for smart archives according to claim 8, characterized in that, In Step3, for Information Two of each file object, for the information queue of the file information of each type in Information Two, regard the average of the differences corresponding to the information queue between the file information of each type in Information Two and all other information at the generation time point in Information Two as the situation comparison factor of the file information of each type in Information Two; In Step 3, the average of the information confidentiality factors of all the generation time point information in Information 1 of each file object is used as the characteristic quantity 1 of each file object; the average of the situation comparison factors of the file information at the generation time point of all in Information 2 of each file object is used as the characteristic quantity 2 of each file object. The characteristic quantity 1 and the characteristic quantity 2 are respectively used as the values on the X-axis and the Y-axis in the corresponding Cartesian coordinate system of each file object, thereby forming coordinate points.
10. An information management device for intelligent archives, characterized in that: Including: A modulation module, which is used to obtain several types of file information of the files to be uploaded corresponding to each file object, and perform modulation processing on the confidentiality characteristic information; A cutting module, which is used to divide the other file information of each file object except the confidentiality characteristic information into Information 1 and Information 2 according to the types of file information, and use the Yamamoto method to perform cutting on the file information of each type in Information 1 at the generation time point of all of them, to obtain sub-queues of the file information of each type in Information 1; determine the difference factors of each sub-queue according to the differences between each sub-queue and other sub-queues; A comparison module, which is used to construct information vectors of each file information in Information 1 according to the difference factors of each information sub-queue in Information 1, analyze the approximation of the information vectors of each file information in Information 1 and other information, and determine the information confidentiality factors of each file information in Information 1; determine the situation comparison factors of each file information in Information 2 according to the difference range between each file information in Information 2 and other file information; combine the information confidentiality factors of the file information at the generation time point of all in Information 1 of all file objects and the situation comparison factors of the information at the generation time point of all in Information 2, and perform segmentation on the file objects; A confidentiality module, which is used to perform confidentiality processing on the information of the file objects in each group by using the Laplace mechanism algorithm to form the files to be uploaded that are allowed to be uploaded.
Citation Information
Patent Citations
An intelligent archive management system and method based on big data
CN118862036B