Archive management system based on cloud archive library
By calculating the vector cosine similarity between users in the cloud archive library, dynamically merging users with high similarity to the same data partition, the problem of large amounts of partitions occupying resources in the cloud archive library is solved, and more efficient resource utilization and retrieval efficiency is achieved.
Patent Information
- Application Number
- CN202510681343.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-26
AI Technical Summary
In the cloud archive, when there are many visitors, a large number of data partitions need to be opened, resulting in waste of storage space and system resources.
By calculating the vector cosine similarity between the enter user and the online user, dynamically merge users with high similarity to the same data partition, and only supplement differentiated content to avoid repeated opening of independent partitions.
It significantly reduces the number of partitions, reduces the use of storage space and system resources, maintains the refinement and controllability of the archive library structure, and improves retrieval efficiency and response flexibility.
Smart Images

Figure CN120217446A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of file management, and particularly to a file management system based on a cloud file library. Background Art
[0002] A cloud file library is a digital information management mode that migrates traditional paper or local electronic files to a cloud computing architecture. With the help of distributed storage, elastic computing, and multiple backup technologies, it centrally hosts a large amount of file data in the cloud, enabling access and collaboration anytime and anywhere.
[0003] Data partitioning refers to a management strategy in a cloud file library that logically isolates and divides the cloud storage and computing resources into logical units according to dimensions such as departments, business domains, or user identities. By creating corresponding data partitions for each user, physical-logical dual isolation of file data, refined permission control, and independent backup of multiple copies can be achieved on the same set of distributed storage and elastic computing architectures. However, when there are a large number of personnel accessing the file library, more data partitions will be created, resulting in more storage space and system resources being occupied. Summary of the Invention
[0004] The purpose of the present invention is to provide a file management system based on a cloud file library to solve the above technical problems.
[0005] The purpose of the present invention can be achieved through the following technical solutions: A file management system based on a cloud file library, comprising: A collection and processing module: when a user enters the cloud file library, mark the user as an entering user, and generate a to-be-oriented vector based on the access keyword sequence of the entering user; mark the entering user who has an opened data partition and is online as a target user, and obtain the comparison vector of the target user; Obtain the cosine similarity A between the to-be-oriented vector and the comparison vector, and set the cosine similarity threshold Ays; A first partitioning module: when the cosine similarity A > Ays, use the corresponding target user as a merging user, add supplementary content B to the data partition of the merging user to obtain data partition X, and use the data partition X as the common data partition of the merging user and the entering user, where the supplementary content B ∈ B1 and B ∉ B2, and B1 and B2 respectively represent the content that the entering user and the merging user are allowed to access; A second partitioning module: when the cosine similarity A ≤ Ays, obtain target keywords based on the browsing records of the entering user in the cloud file library, and create a separate data partition for the entering user based on the target keywords.
[0006] As a further solution of the present invention: obtaining the to-be-oriented vector and the comparison vector includes: For the incoming user, convert a single keyword in the corresponding access keyword sequence into a word vector, and perform TF-IDF weighted averaging on all the word vectors to obtain a target vector, denoted as the to-be-oriented vector; Obtain the target vector of the target user, denoted as the comparison vector.
[0007] As a further solution of the present invention: creating a separate data partition for the incoming user includes: Obtain the browsing records of the incoming user within a preset monitoring period, mark the files in a single browsing record as target files, and use the browsed part in a single target file as the target part; Determine target keywords based on the target part, mark the files in the cloud archive library whose types and quantities of the target keywords included exceed the preset values as partition files, create a data partition Y, store the partition files in the data partition Y, and use the data partition Y as the data partition of the incoming user.
[0008] As a further solution of the present invention: determining target keywords includes: Extract the keywords in the target part, denoted as the to-be-determined words, group the to-be-determined words, and the quantities and / or types of the to-be-determined words in different groups are different; Determine the quantity n of the to-be-determined words in the target part i included in the group j, and calculate the fitness k between the group j and the target part i ij =n / N j where N j represents the quantity of the to-be-determined words in the group j, and calculate the total fitness K of the group j j ; Generate a set of total fitnesses K jh =(K1, K2,..., K m ), m represents the total number of groups, and use the to-be-determined words in the group corresponding to the maximum total fitness Kmax = max(K jh ) as the target keywords.
[0009] As a further solution of the present invention: when the data partition corresponding to the merged user corresponds to the remaining target users, perform the following steps: Mark the remaining target users corresponding to the data partition of the merged user as to-be-determined users. When the cosine similarity between the comparison vector of any one of the to-be-determined users and the to-be-oriented vector is less than the cosine similarity threshold, no longer use this merged user as the merged user.
[0010] As a further solution of the present invention: when the cosine similarity of two or more of the above is greater than the cosine similarity threshold, a prompt message is sent for reporting.
[0011] As a further solution of the present invention: when the target user does not exist, the steps in the second partition module are executed to open a data partition for the entering user.
[0012] As a further solution of the present invention: when the registration time of the entering user is less than a preset duration, the subsequent steps are not executed, and a prompt message is sent to a preset administrator, and the administrator manually opens the data partition of the entering user.
[0013] Advantages of the present invention: Compared with the prior art: 1) By calculating the vector cosine similarity between the entering user and the online user, the present invention merges the two into the same data partition when the threshold condition is met, and only supplements the differential content, avoiding the need to create independent partitions for each new user repeatedly; this dynamic merging mechanism enables the number of partitions to adaptively converge with similar requirements, significantly reducing the disk partitioning and metadata overhead, fundamentally reducing the risk of the storage space and system resources being occupied unnecessarily, and keeping the archive structure refined and controllable. 2) When the similarity is not sufficient to trigger the merging, the target keywords are automatically extracted based on the actual browsing segments of the user during the monitoring period, and then the relevant archives are screened from the entire library according to the keyword density to immediately create an exclusive partition for the user. This content-driven partition generation method can accurately fit the user's real retrieval interest; at the same time, the corresponding relationship between the keywords and the archives can be continuously updated with the user's behavior, and the system automatically adjusts the partition content without the need for frequent manual migration or reconstruction, realizing the dynamic alignment of the archive resources and the business requirements, maintaining an efficient, refined and sustainable management granularity, and significantly shortening the search path and improving the response flexibility. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The present invention will be further described below with reference to the accompanying drawings.
[0015] Figure 1 It is a schematic flowchart of an archive management system based on a cloud archive library of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0017] Please refer to Figure 1As shown in the figure, the present invention is an archive management system based on a cloud archive library. The present invention is an archive management system based on a cloud archive library, including: Collection and processing module: When a user enters the cloud archive library, mark the user as an entering user, and generate a to-be-oriented vector based on the access keyword sequence of the entering user; mark the entering user who has opened a data partition and is online as a target user, and obtain the comparison vector of the target user; In a preferred embodiment of the present invention, obtaining the to-be-oriented vector and the comparison vector includes: For the entering user, convert a single keyword in the corresponding access keyword sequence into a word vector, and perform TF-IDF weighted average on all the word vectors to obtain a target vector, denoted as the to-be-oriented vector; Obtain the target vector of the target user, denoted as the comparison vector; It should be noted that when it is detected that a user logs in to the cloud archive library, first mark the user as an "entering user" in the session context, and extract the access keyword sequence from the user's most recent retrieval, click or filtering operation, such as "contract", "lease", "2023". Subsequently, call the pre-trained Chinese word vector model (such as Word2Vec or BERT Token Embedding) to map each keyword into a word vector of the same dimension; the system synchronously queries the word frequency and document frequency of the whole library corpus. After obtaining the corresponding TF values for the keywords "contract", "lease", "2023" respectively, multiply them by their IDF values to obtain weights, and then weight and sum the three groups of vectors one by one in the way of "word vector × weight", and finally divide by the sum of the weights to obtain the weighted average vector representing the user's interest intention. Store this vector in the cache and mark it as the "to-be-oriented vector"; at the same time, the system traverses other users who have currently created data partitions and are in an online state, directly load the target vectors generated and persisted when they entered into the memory, and mark them as "comparison vectors" respectively; after preparing the to-be-oriented vector and multiple comparison vectors, subsequent cosine similarity calculation and partition merging or independent decision-making can be carried out; By first mapping the access keyword sequence word by word into word vectors and then performing weighted average with TF-IDF weights, it is possible to highlight the key information in each user's retrieval intention while keeping the vector dimension fixed, weaken the noise brought by stop words or generalization words, so that the obtained to-be-oriented vector and comparison vector not only have the overall semantic expression ability but also take into account the importance of words, so as to more accurately measure the true proximity between the access requirements of different users when calculating the cosine similarity subsequently; this method has a small amount of calculation, is easy to update online in real time, and is convenient for the system to quickly complete similarity determination and partition merging decision in a multi-person concurrent scenario; Obtain the cosine similarity A between the to-be-oriented vector and the comparison vector, and set the cosine similarity threshold Ays; The first partition module: when the cosine similarity A > Ays, the corresponding target user is regarded as the merged user, the supplementary content B is added to the data partition of the merged user to obtain the data partition X, and the data partition X is used as the common data partition of the merged user and the incoming user. The supplementary content B ∈ B1 and B ∉ B2, where B1 and B2 respectively represent the content that the incoming user and the merged user are allowed to access; Exemplarily, when the system detects that the cosine similarity between the incoming user and a certain target user exceeds the threshold Ays, the target user is identified as the merged user. For example, if the incoming user U1 has a high similarity with the online user U2, the system first retrieves the access content set B1 of U1 and the existing content set B2 of U2, and calculates the difference set B, which is the part that exists in B1 but not in B2. Assuming that B1 contains files a, b, c and B2 only contains a, b, then B is c. Subsequently, the system creates or attaches the file c in the physical partition storing U2's data, and can use soft links, hard links or block-level references instead of full replication to quickly generate the extended partition X; Such processing can avoid repeatedly partitioning the physical space for each new user with highly similar needs. At the same time, by only appending the differential content instead of overall replication, it reduces data redundancy and disk fragmentation, keeping the total number of partitions and the scale of a single partition within a controllable range. The centralized storage of files brought about by partition sharing simplifies operation and maintenance operations such as directory indexing, backup, and recovery. The system does not need to frequently create, mount, and scan a large number of partitions in a high-concurrency scenario, thus helping the overall solution to achieve comprehensive optimization of storage space, I / O scheduling, and management processes; The second partition module: when the cosine similarity A ≤ Ays, the target keyword is obtained based on the browsing record of the incoming user in the cloud archive, and a separate data partition is opened for the incoming user based on the target keyword; Another preferred embodiment of the present invention, opening a separate data partition for the incoming user includes: Obtain the browsing record of the incoming user within a preset monitoring period, mark the files in a single browsing record as target files, and regard the browsed part of a single target file as the target part; Determine the target keyword based on the target part, mark the files in the cloud archive whose types and quantities containing the target keyword both exceed the preset value as partition files, open the data partition Y, store the partition files in the data partition Y, and use the data partition Y as the data partition of the incoming user; Exemplarily, first, continuously capture each page access record of the user within a set monitoring period. For example, the user opens three types of documents, namely "Project Contract", "Budget Approval Form", and "Progress Weekly Report", multiple times within three days. The complete file entries where they are located are marked as target files, and the actual text, attachments, pictures, etc. that the user scrolls or retrieves and locates are recorded as target parts. Subsequently, the platform performs Chinese word segmentation and named entity recognition on these target parts, calculates the word frequency and position weights after filtering stop words, and extracts keywords that reflect the user's focus of attention, such as "contract amount", "payment node", and "progress milestone". Then, retrieve all files in the entire library that contain these keywords simultaneously and have a frequency of occurrence exceeding the system threshold, and classify the eligible entries as partitioned files. For example, retrieve contracts in the same series, supplementary agreements, payment vouchers, and corresponding weekly reports. Next, create a new data partition Y, store the retrieved partitioned files in this partition in the form of symbolic links or physical copies, and write an independent partition mapping pointing to Y for the accessing user in the access control list, so that when the user logs in later, they will directly land in a dedicated file set that highly matches their interests; It should be noted that using the user's real browsing behavior to drive keyword extraction and file aggregation can make the partition content accurately reflect the user's current business tasks, avoiding access redundancy or gaps caused by static allocation according to organizational roles or initial application lists. At the same time, screening files with keyword coverage and occurrence types as thresholds can ensure that the partition size is moderate and key information is not overly missing, shortening the subsequent retrieval path, improving the cache and index hit rates, and the partition is automatically updated as the user's interests change, eliminating the need for administrators to manually migrate frequently, which not only reduces the operation and maintenance burden but also reduces the risk of storage fragmentation, providing efficient and sustainable support for the entire cloud file library to achieve on-demand partitioning and dynamic scaling; In a preferred case of this embodiment, determining the target keywords includes: Extract the keywords in the target part, denoted as undetermined words, group the undetermined words, and the number and / or type of undetermined words in different groups are different; Determine the number n of undetermined words in the target part i included in the group j, and calculate the fit degree k between the group j and the target part i ij =n / N j where N j represents the number of undetermined words in the group j, and calculate the total fit degree K of the group j j ; Generate a set of total fit degrees K jh =(K1, K2,..., K m ), where m represents the total number of groups, and take the undetermined words in the group corresponding to the maximum total fit degree Kmax = max(K jh ) as the target keywords; It should be noted that, firstly, several pending words are extracted from the target part, such as the four keywords of contract amount, payment node, progress milestone, and acceptance criteria, and then groups are constructed one by one according to mathematical permutations and combinations: first, the four words are regarded as a complete combination as a whole, and then all triple combinations are enumerated, such as contract amount + payment node + progress milestone, and the remaining groups are not repeated; the fit is defined as the ratio of the number of words that have appeared in the combination to the size of the combination, which is essentially a coverage measurement: it not only measures the completeness of the occurrence of a keyword combination in a single browsing fragment, but also normalizes the combinations of different sizes through the denominator, so that the combinations containing two words, three words, and even all words can be compared under the same dimension, thereby avoiding the bias of naturally obtaining a higher absolute number of matches due to the large number of words; the coverage is accumulated on all target parts to further examine the overall adaptability of the combination to the user's continuous browsing behavior, reflecting its stability and lasting relevance; finally, the combination with the highest total fit is taken, which is equivalent to selecting the keyword set that is most commonly mentioned and relatively complete, and maximizing the restoration of user attention topics in a statistical sense.
[0018] When the data partition corresponding to the merged user corresponds to the remaining target users, the following steps are performed: Mark the remaining target users corresponding to the data partition corresponding to the merged user as pending users, and when the cosine similarity between the comparison vector of any of the pending users and the pending vector is less than the cosine similarity threshold, the merged user is no longer considered as a merged user; It is understandable that when the merged user has shared the same data partition with other target users, the system will mark all of these sharers as pending users, and then calculate the semantic vector similarity between each pending user and the newly entered user; as long as it is found that the similarity between any pending user and the newly entered user is lower than the set threshold, the merge will be canceled immediately and the existing partition structure will be maintained unchanged. This ensures that all users in the same data partition maintain a high degree of consistency in their access needs, and prevents irrelevant users from being forcibly aggregated; It should be noted that when two or more cosine similarities are greater than the cosine similarity threshold, a prompt message is sent for reporting; When the semantic similarity between two or more users and a new user exceeds the preset threshold at the same time, the report is immediately sent, allowing administrators to immediately pay attention to possible group high-overlapping access situations, and promptly troubleshoot risk scenarios such as account fraud, data leakage, or abnormal collaboration, to avoid the system automatically merging to form an excessively large sharing range and weaken the isolation strength; It should be noted that, when the target user does not exist, the steps in the second partitioning module are executed to open a data partition for the incoming user.
[0019] When there are no online target users in the library for comparison, the system instead creates an independent partition for the incoming user, which can ensure that their legitimate business operations are not delayed, and at the same time avoid binding users with low relevance or unknown requirements rigidly to the previous partitions, resulting in redundant data. Thus, in two extreme cases, the security closed-loop is completed through manual intervention and automatic implementation respectively, taking into account both real-time performance and fineness, and ensuring the integrity and sustainability of the overall isolation strategy of the cloud archive library; It can be understood that when the registration time of the incoming user is less than a preset duration, the subsequent steps are not executed, and a prompt message is sent to the preset management personnel, who manually create a data partition for the incoming user; When the user registration time is too short, there is not enough data to support their identity credibility, access intention, and subsequent activity. If they immediately enter the automatic partition process, it may result in over-authorization due to misjudging similarity due to the lack of historical keywords. Therefore, the system first suspends the automatic processing and pushes the information to the management personnel for manual review and manual creation of the partition. This can verify the user's identity, business background, and the scope of required archives, prevent malicious registration and the rapid acquisition of sensitive files by temporary accounts, and also lay an accurate starting point for subsequent automatic merging or independent partitioning, helping the cloud archive library to maintain a reasonable number of partitions and efficient utilization of storage resources while ensuring security and compliance.
[0020] The above has described in detail an embodiment of the present invention, but the content described is only the preferred embodiment of the present invention and cannot be considered as limiting the scope of implementation of the present invention. All equivalent changes and improvements made according to the scope of the present invention application should still fall within the scope covered by the patent of the present invention.
Claims
1. An archive management system based on a cloud archive repository, characterized in that, Including: Collection and processing module: When a user enters the cloud archive, mark the user as an entering user, and generate a to-be-oriented vector based on the access keyword sequence of the entering user; mark the entering user who has an opened data partition and is online as a target user, and obtain the comparison vector of the target user. Obtain the cosine similarity A between the to-be-oriented vector and the comparison vector, and set the cosine similarity threshold Ays. First partition module: When the cosine similarity A > Ays, regard the corresponding target user as a merged user, add supplementary content B to the data partition of the merged user to obtain data partition X, and regard data partition X as the common data partition of the merged user and the entering user, where the supplementary content B ∈ B1 and B ∉ B2, and B1 and B2 respectively represent the content that the entering user and the merged user are allowed to access. Second partition module: When the cosine similarity A ≤ Ays, obtain the target keyword based on the browsing record of the entering user in the cloud archive, and open a separate data partition for the entering user based on the target keyword.
2. The archival management system based on a cloud archive repository according to claim 1, characterized in that, Obtaining the to-be-oriented vector and the comparison vector includes: For the entering user, convert a single keyword in the corresponding access keyword sequence into a word vector, and perform TF-IDF weighted average on all the word vectors to obtain a target vector, denoted as the to-be-oriented vector. Obtain the target vector of the target user, denoted as the comparison vector.
3. An archive management system based on a cloud archive repository according to claim 1, characterized in that, Opening a separate data partition for the entering user includes: Obtain the browsing record of the entering user within a preset monitoring period, mark the files in a single browsing record as target files, and regard the browsed part in a single target file as the target part. Determine the target keyword based on the target part, mark the files in the cloud archive whose types and quantities containing the target keyword both exceed the preset value as partition files, open data partition Y, store the partition files in data partition Y, and regard data partition Y as the data partition of the entering user.
4. An archive management system based on a cloud archive repository according to claim 3, characterized in that, Determining the target keyword includes: Extract the keywords in the target part, denoted as to-be-determined words, group the to-be-determined words, and the quantities and / or types of to-be-determined words in different groups are different. Determine the number \(n\) of pending words in the target part \(i\) included in the group \(j\), and calculate the fit degree \(k\) between the group \(j\) and the target part \(i\). ij =n / N j , N j represents the number of pending words in the group \(j\), and calculate the total fit degree \(K\) of the group \(j\). j ; Generate the total matching degree set K jh = (K1, K2,..., K m ), where m represents the total number of groups. Take the undetermined words in the group corresponding to the maximum total matching degree Kmax = max(K jh ) as the target keywords.
5. A file management system based on a cloud archive according to claim 1, characterized in that, When the data partition corresponding to the merged user corresponds to the remaining target users, perform the following steps: Mark the remaining target users corresponding to the data partition of the merged user as to-be-determined users. When the cosine similarity between the comparison vector of any one of the to-be-determined users and the to-be-oriented vector is less than the cosine similarity threshold, no longer regard this merged user as a merged user.
6. The archival management system based on a cloud archive repository according to claim 1, wherein When two or more of the cosine similarities are greater than the cosine similarity threshold, send a prompt message for reporting.
7. A file management system based on a cloud archive according to claim 1, characterized in that, When there is no target user, perform the steps in the second partition module to open a data partition for the entering user.
8. An archive management system based on a cloud archive repository according to claim 1, characterized in that, When the registration time of the entering user is less than the preset duration, do not perform the subsequent steps, and send a prompt message to the preset management personnel, and the management personnel manually open the data partition of this entering user.
Citation Information
Patent Citations
Strong privacy protection method and system for spatial keyword query
CN117171802A
Sensitive partition cloud data deduplication method and system supporting similarity detection
CN117420961A
Archive information security risk management system
CN117972174A
Dynamic web portal management system based on big data
CN118094049A
Cloud-based data backup system and method
CN119537100A