Digital archive management method and system
Through vectorized representation and feature extraction of archive instances, and combining supportive archive tags for classification and modeling, the problems of inefficient archive management and high error rate in the existing technology are solved, and more efficient and accurate archive management is achieved.
Patent Information
- Application Number
- CN202411199848.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-29
- Publication Date
- 2025-06-24
AI Technical Summary
The existing digital archive management methods are inefficient and have high error rates when processing large-scale archive data, lack flexibility and accuracy, difficult to effectively cluster and label archives, and consume a lot of resources.
A digital archive management method is proposed, by obtaining and vectorizing archival instances, extracting time period information and supportive archival tags, performing feature extraction and classification, generating preliminary and secondary archival sets, using algorithms to model and calculate label weights, and finally selecting the label with the largest weight as archive tags.
It significantly improves the efficiency and accuracy of archive management, enables rapid processing and analysis of large-scale data sets, provides more accurate archival tag decisions, reduces errors and improves automation.
Smart Images

Figure CN120196800A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data management, and specifically relates to a digital file management method and system. Background Art
[0002] In the current information age, the efficiency and accuracy of data management and processing are crucial for the development of all industries. With the deepening of digital transformation, how to efficiently manage and utilize a large number of digital files has become an urgent problem to be solved. Traditional file management methods rely mostly on manual operations, which are not only inefficient but also prone to errors. In addition, due to the increasing quantity and complexity of digital files, traditional methods are no longer able to meet the needs of modern management.
[0003] The development of digital file management technology aims to improve the automation and intelligence level of file management by using modern information technology, thereby effectively enhancing the efficiency and accuracy of file management. For example, by applying advanced algorithms such as machine learning and deep learning, rapid processing of large-scale file data and intelligent decision support can be achieved, significantly improving work efficiency and decision-making quality.
[0004] However, existing digital file management methods still have some deficiencies. First, these methods often lack sufficient flexibility and accuracy in the vectorization representation, feature extraction, and classification processing of files. Second, how to effectively cluster and label files according to the content and relevance of files, and how to process and optimize file data sets to improve the effect of model training, remains a challenge. In addition, existing methods may encounter problems of low efficiency and high resource consumption when dealing with a large number of files. Summary of the Invention
[0005] (I) Technical Problems to be Solved
[0006] The present invention mainly aims at the above problems and proposes a digital file management method and system, the purpose of which is to solve the problems of low efficiency and error rate in dealing with large-scale file data by traditional methods.
[0007] (II) Technical Solutions
[0008] To achieve the above object, the first aspect of the present invention provides a digital file management method, which includes the following steps:
[0009] Obtain a set of digital file instances, and extract the time period information and multiple supporting file tags in each file instance;
[0010] Perform vectorization representation on each file instance, and extract features from these vectorized data to obtain a featureized data set;
[0011] Classify the corresponding time periods in the characterized dataset based on the supporting archive labels, group the time periods of the same category together, and form a preliminary archive set;
[0012] Perform preprocessing operations on the preliminary archive set to generate a secondary archive set;
[0013] Use an algorithm to model the secondary archive set, train multiple sub-models for each archive instance, and each sub-model is associated with a supporting archive label;
[0014] Calculate the weight value of each supporting archive label, and adjust the output results of each sub-model based on the weight value;
[0015] Select the label with the largest weight value from the candidate labels output by the sub-model as the finally determined archive label.
[0016] Furthermore, the method for vectorizing each archive instance includes the following steps:
[0017] Obtain the text content of each archive instance;
[0018] Perform word segmentation on the obtained text content to generate a list of words;
[0019] Use a pre-trained word vector model to convert each word in the list of words into a corresponding word vector;
[0020] Perform weighted averaging on the converted word vectors to obtain a comprehensive vector;
[0021] Combine the comprehensive vector with the time period information to form a vector representation containing time features;
[0022] Perform normalization processing on the vector representation containing time features to generate the final vectorized representation.
[0023] Furthermore, the method for classifying the corresponding time periods in the characterized dataset based on the supporting archive labels includes the following steps:
[0024] Determine the time period information of each archive instance in the characterized dataset;
[0025] Extract the archive instances containing specific supporting archive labels from the characterized dataset;
[0026] Sort the archive instances containing specific supporting archive labels according to the time period information;
[0027] Divide the sorted time period information into multiple time intervals;
[0028] For each time interval, calculate the similarity measure of the archive instances within it;
[0029] Based on similarity measurement, time intervals with high similarity are grouped into the same category;
[0030] Summarize all category information to complete the classification of time periods.
[0031] Furthermore, the method for preprocessing the preliminary archive set to generate a secondary archive set includes the following steps:
[0032] Check the integrity of each archive instance in the preliminary archive set;
[0033] Perform data filling or deletion on archive instances with missing information;
[0034] Standardize the formats of archive instances from different sources to obtain a unified data structure;
[0035] Apply a denoising algorithm to filter out the noise data existing in the preliminary archive set;
[0036] Identify and merge redundant or duplicate archive instances;
[0037] Reorganize the preprocessed archive instances and sort them according to time periods and supporting archive labels;
[0038] Sort out and save the preprocessed archive instances to form a secondary archive set.
[0039] Furthermore, the specific steps for modeling the secondary archive set using an algorithm include the following steps:
[0040] Extract the feature vector X of each archive instance from the secondary archive set i ;
[0041] Standardize the extracted feature vectors using the formula:
[0042]
[0043] where μ is the mean of the feature vector and σ is the standard deviation of the feature vector;
[0044] Select a machine learning algorithm as the modeling tool;
[0045] Divide the secondary archive set into a training set and a validation set, and use the cross-validation method for data division;
[0046] Use the training set to train the selected machine learning algorithm, and calculate the loss function during the training process:
[0047]
[0048] where f(Z i ,θ) is the model prediction result, yi is the true label, l is the loss function; N is the number of training samples;
[0049] During the training process, adjust the model parameters θ to optimize the loss function L(θ), and use the gradient descent method to update the parameters:
[0050]
[0051] where η is the learning rate, represents the gradient symbol, which is used to represent the gradient of the loss function L(θ) with respect to the parameter θ;
[0052] Use the validation set to evaluate the performance of the trained model;
[0053] Save the finally obtained model, and train multiple sub-models for each file instance, and each sub-model is associated with a supporting file label.
[0054] Furthermore, the calculation and adjustment steps of the weight values specifically include:
[0055] Calculate the weight value w for each supporting file label j , according to the output results y of each sub-model j , adjust the model output:
[0056]
[0057] where M is the number of sub-models, w j represents the weight value of the j-th sub-model; y j represents the output result of the j-th sub-model; P is the final prediction result after comprehensively considering the outputs of all sub-models.
[0058] Furthermore, the method of selecting the label with the largest weight value from the candidate labels output by the sub-models as the finally determined file label specifically includes the following steps:
[0059] Obtain the prediction results of each sub-model for the file instance and their corresponding weight values;
[0060] Perform weighted summation on the prediction results and their weight values of all sub-models to obtain a comprehensive score;
[0061] Select the label with the highest comprehensive score as the final label of the file instance.
[0062] Furthermore, when modeling the secondary file set, the machine learning algorithms used include but are not limited to one or more of the following: linear regression, logistic regression, support vector machine, random forest, gradient boosting decision tree, neural network and its variants.
[0063] Furthermore, the preprocessing operation on the preliminary archive set also includes desensitizing sensitive information in the archive instances.
[0064] To achieve the above object, the second aspect of the present invention provides a digital archive management system, including the following modules:
[0065] A data acquisition module, configured to acquire a set of digital archive instances, and extract the time period information and multiple supporting archive tags in each archive instance;
[0066] A vectorization representation module, configured to vectorize each archive instance, and perform feature extraction on these vectorized data to obtain a featureized data set;
[0067] A classification module, based on the supporting archive tags, classifies the corresponding time periods in the featureized data set, aggregates the time periods of the same category together to form a preliminary archive set;
[0068] A preprocessing module, configured to perform preprocessing operations on the preliminary archive set to generate a secondary archive set;
[0069] A modeling module, using an algorithm to model the secondary archive set, training multiple sub-models for each archive instance, and each sub-model is associated with a supporting archive tag;
[0070] A weight calculation module, configured to calculate the weight value of each supporting archive tag, and adjust the output results of each sub-model based on the weight value;
[0071] A tag selection module, selects the tag with the largest weight value from the candidate tags output by the sub-model as the finally determined archive tag.
[0072] (III) Beneficial effects
[0073] Compared with the prior art, the digital archive management method and system provided by the present invention, firstly, by vectorizing the archive instances and extracting key features, this method can quickly and effectively process and analyze large data sets. Then, intelligent classification is carried out using supporting archive tags, enhancing the organizational structure of the data, so that the data of relevant time periods can be accurately aggregated. In addition, through preprocessing and modeling of the data, and using sub-models and weight adjustment strategies, this system can provide more accurate archive tag decisions, reduce errors and improve the automation level. These steps work together to significantly improve the overall efficiency and accuracy of archive management. Especially when facing a large amount of complex archive data, the method and system of the present invention show excellent processing capabilities and efficient management performance. Description of the drawings
[0074] Figure 1Flowchart of a digital file management method disclosed in this application.
[0075] Figure 2 Schematic diagram of a data acquisition process disclosed in this application.
[0076] Figure 3 Process diagram of a vectorization representation disclosed in this application.
[0077] Figure 4 Process diagram of classifying feature data based on supportive file tags disclosed in this application.
[0078] Figure 5 Process diagram of preprocessing for displaying a preliminary file set disclosed in this application.
[0079] Figure 6 Process diagram of model training and weight calculation disclosed in this application.
[0080] Figure 7 Process diagram of final label selection disclosed in this application.
[0081] Figure 8 Structure diagram of a digital file management system disclosed in this application. Detailed implementation manners
[0082] The present invention will be described in detail below with reference to the accompanying drawings. The technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0083] As Figure 1 shown, the present invention provides a digital file management method, and the method includes the following steps:
[0084] Step S100, obtain a set of digital file instances, and extract the time period information and multiple supportive file tags in each file instance;
[0085] In this step, first, a set of digital archive instances are obtained. These archive instances can be documents, images, videos, or any other form of digital information. Then, continue to extract the time period information from each archive instance, such as the creation or modification date of a document, or the shooting time range of a video. In addition, multiple supporting archive tags are also extracted from the archives. The supporting archive tags can be automatically or manually labeled according to attributes such as content, theme, author, location, etc., and are used for subsequent classification and analysis work. For example, if dealing with a historical document library, step S100 also involves identifying the creation time of each document and related keywords such as "historical events", "persons", etc., which are all instances of supporting archive tags.
[0086] It can be understood that "supporting archive tags" refer to a set of tags assigned to digital archive instances, and these tags are defined based on specific attributes or characteristics of the archive content. The purpose is to support subsequent data processing and management processes, such as classification, retrieval, and analysis. The tags cover a wide range of elements, including but not limited to the theme, author, location, time period of the archive, and related events or activities.
[0087] Step S200: Represent each archive instance in vector form and extract features from these vectorized data to obtain a featureized data set;
[0088] In this step, the text content, image features, or other data formats in the archive instance are first converted into vector form. For example, as Figure 2 、 Figure 3 shown, text data can be converted into numerical vectors through word segmentation and using a pre-trained word vector model, and images can extract key features through image processing algorithms. These vectorized data will then undergo further feature extraction, aiming to identify and retain the information that is most critical for subsequent tasks. Finally, these processed data are organized into a featureized data set, which will be used for further analysis and model training to support the automatic classification, retrieval, and other management tasks of the archives.
[0089] Step S300: Based on the supporting archive tags, classify the corresponding time periods in the featureized data set, gather the time periods of the same category together to form a preliminary archive set;
[0090] In this step, as Figure 4As shown, supportive tags based on archival instances—such as event types, related topics, or activities—group similar or related time periods in the dataset together. This classification enables archives with similar characteristics to be organized into a preliminary set of archives, facilitating further analysis and management. For example, if the characterized dataset contains historical documents, step S300 clusters the archives according to document tags and relevant time information, thus forming a preliminary set of archives regarding a specific historical event.
[0091] Step S400: Perform preprocessing operations on the preliminary set of archives to generate a secondary set of archives;
[0092] As Figure 5 shown, when processing a file archive containing multiple formats and sources, this step ensures that all documents are converted to a unified format and irrelevant or duplicate content is removed, thereby making the set of archives more standardized and easier to manage.
[0093] Step S500: Use an algorithm to model the secondary set of archives, training multiple sub-models for each archival instance, with each sub-model associated with a supportive archival tag;
[0094] As Figure 6 shown, first, the feature vectors of each archival instance are extracted from the secondary set of archives. Then, for each supportive archival tag, a sub-model associated with it is trained. Each sub-model focuses on predicting the archival characteristics related to its associated tag, such as a specific event type, document topic, or activity category. In this way, the multi-dimensional characteristics of the archives can be processed and predicted in a fine-grained manner, improving the accuracy and relevance of the prediction. For example, when managing the archives of a large digital library, different sub-models are trained for different literary genres (such as science fiction, history, romance, etc.), and each model learns to identify and predict documents belonging to a specific genre.
[0095] Step S600: Calculate the weight values of each supportive archival tag and adjust the output results of each sub-model based on the weight values;
[0096] In this step, first, the corresponding weight values are determined according to the performance and importance of each sub-model. The weight values reflect the relative importance of each tag in the overall archival classification and prediction. Then, the weight values are used to adjust the output results of each sub-model to ensure that the comprehensive prediction result of the output is more accurate and reliable. For example, if a tag such as "government documents" is extremely crucial in a specific set of archives, the output of its corresponding sub-model will be given a higher weight. In this way, when the system makes the final tag decision, it pays more attention to the prediction results of those key tags, thereby enhancing the efficiency and accuracy of the entire archival management system.
[0097] Step S700: Select the tag with the largest weight value from the candidate tags output by the sub-models as the finally determined file tag.
[0098] As Figure 7 shown, the output results of each sub-model will be weighted and evaluated according to the weight values calculated in the previous step. Each sub-model outputs a candidate tag. By comparing the weighted scores of these tags, the tag with the highest score is selected as the finally determined tag for this file instance. For example, if a file instance is classified as possibly belonging to three categories: "Education", "History", and "Law", and among the sub-models for these three categories, the weight value of the "History" tag is the highest, then "History" will be selected as the final tag for this file.
[0099] In step S200, the purpose is to convert each file instance into a mathematical representation form that can be processed by a computer, that is, a vectorized representation, which involves the following detailed steps (see Figure 3 ):
[0100] Step S201: First, the system obtains the text content of each file instance. For example, extract all the text from a PDF document.
[0101] Step S202: Next, perform word segmentation on the extracted text content, breaking the continuous text into a list of independent words, applicable to various language environments, such as space segmentation in English or word segmentation in Chinese.
[0102] Step S203: Then, use a pre-trained word vector model (such as Word2Vec or GloVe) to convert each word in the word list into a corresponding word vector. This step converts the text data into a numerical form for easy algorithm processing.
[0103] Step S204: Perform weighted averaging on these word vectors to generate a comprehensive vector, so that the information of all the words in the text can be comprehensively considered to form a vector representing the entire text.
[0104] Step S205: Combine the obtained comprehensive vector with the time period information to add time features. For example, add the creation date of the document or relevant time markers to the vector representation to provide more context information.
[0105] Step S206: Finally, perform normalization processing on the vector representation containing time features, such as Z-score normalization, to ensure that the data has the same scale in different dimensions for easy subsequent algorithm processing, and generate the final vectorized representation.
[0106] The method of classifying the characterized dataset according to the supportive file tags and time period information is described in detail in step S300, ensuring the effective organization and management of the file data (see Figure 4 ):
[0107] Step S301: Identify and record the time period information in each file instance, such as the creation and modification dates of the document or the time range of the event.
[0108] Step S302: Filter out the file instances containing specific supportive file tags from the dataset, for example, filter all the documents labeled with "financial news".
[0109] Step S303: Sort these filtered file instances according to their time period information, and organize them in chronological order or other relevant time criteria.
[0110] Step S304: Split the sorted files into different time intervals according to time continuity or logical relationship, and each interval may contain one or more time points or ranges.
[0111] Step S305: Calculate the similarity measure for the file instances within each time interval, such as by comparing the similarity of the text content or other relevant features.
[0112] Step S306: Based on the above similarity measure, classify the time intervals with high similarity into the same category to form a more concentrated data cluster.
[0113] Step S307: Summarize all the classified information to complete the classification process for the entire time period, so as to achieve the effective management and easily accessible storage structure of the file instances.
[0114] In step S400, the method of preprocessing the preliminary file set to generate the secondary file set includes the following steps (see Figure 5 ):
[0115] First, verify the integrity of each file instance (step S401) to ensure that all files are complete, or fill in the missing information. Incomplete files may be deleted (step S402). Subsequently, unify the formatting of the files from different sources to ensure the consistency of the data structure (step S403). In addition, apply a denoising algorithm to remove the noise in the data and improve the data quality (step S404). Step S405 involves identifying and merging duplicate file instances to reduce redundancy. Finally, the preprocessed files will be reorganized, sorted systematically according to the time period and supportive file tags (step S406), and the sorted files will be saved as the secondary file set for convenient subsequent access and management (step S407).
[0116] In step S500, the specific steps of modeling the secondary archive set using an algorithm include the following steps (see Figure 6 ): Extract feature vectors from the secondary archive set and perform normalization processing, select an appropriate machine learning algorithm (such as support vector machine, random forest, or deep neural network) for model training, divide the training set and validation set through cross-validation, and optimize the loss function during the training process to adjust the model parameters. Next, evaluate the model performance using the validation set, and train multiple sub-models for each archive instance, with each sub-model corresponding to a supporting archive label. Finally, calculate the weight values of each supporting archive label, perform a weighted sum based on the output results of the sub-models, and select the label with the largest weight value as the finally determined archive label.
[0117] The specific steps are as follows:
[0118] Extract the feature vector X of each archive instance from the secondary archive set i ;
[0119] Normalize the extracted feature vectors using the formula:
[0120]
[0121] where μ is the mean of the feature vector and σ is the standard deviation of the feature vector;
[0122] Select a machine learning algorithm as the modeling tool;
[0123] Divide the secondary archive set into a training set and a validation set, and use the cross-validation method for data division;
[0124] Use the training set to train the selected machine learning algorithm model, and calculate the loss function during the training process:
[0125]
[0126] where f(Z i ,θ) is the model prediction result, y i is the true label, l is the loss function; N is the number of training samples;
[0127] Adjust the model parameter θ during the training process to optimize the loss function L(θ), and use the gradient descent method to update the parameter:
[0128]
[0129] where η is the learning rate, represents the gradient symbol, which is used to represent the gradient of the loss function L(θ) with respect to the parameter θ;
[0130] Evaluate the performance of the trained model using the validation set;
[0131] Save the finally obtained model, and train multiple sub-models for each file instance, with each sub-model associated with a supporting file label.
[0132] Calculate the weight value w for each supporting file label j , and adjust the model output according to the output result y j of each sub-model:
[0133]
[0134] where M is the number of sub-models, and w j represents the weight value of the j-th sub-model; y j represents the output result of the j-th sub-model; P is the final prediction result after comprehensively considering the outputs of all sub-models.
[0135] The above formula mainly reflects two characteristics: one is to standardize the feature vector through the standardization formula to reduce the influence brought by the difference in different feature scales, so as to improve the stability and effect of model training; the other is to measure the difference between the model prediction result and the true label through the loss function L(θ) in the machine learning algorithm, and use the gradient descent method to update the model parameters to minimize the loss function. The weighted summation formula further integrates the output results of each sub-model, and realizes the effective fusion of the outputs of multiple sub-models by calculating the weight values of each supporting file label, and finally selects the label with the largest weight as the finally determined file label.
[0136] Preferably, in step S700, the method of selecting the label with the largest weight value from the candidate labels output by the sub-model as the finally determined file label specifically includes the following steps (see Figure 7 ): Obtain the prediction result of each sub-model for the file instance and its corresponding weight value; perform weighted summation on the prediction results and weight values of all sub-models to obtain a comprehensive score; select the label with the highest comprehensive score as the final label of the file instance.
[0137] Preferably, when modeling the secondary file set, the machine learning algorithms used include but are not limited to one or more of the following: linear regression, logistic regression, support vector machine, random forest, gradient boosting decision tree, neural network and its variants.
[0138] Preferably, the preprocessing operation on the preliminary file set further includes desensitizing the sensitive information in the file instance.
[0139] As Figure 8 shown, the present invention also provides a digital file management system, including the following modules:
[0140] The data acquisition module 100 is used to acquire a set of digital archive instances and extract time period information and a plurality of supporting archive tags in each archive instance;
[0141] The vectorization representation module 200 is used to vectorize each archive instance and extract features from the vectorized data to obtain a feature data set;
[0142] A classification module 300 , based on supporting archive tags, classifies the corresponding time periods in the characterization data set, and groups time periods of the same category together to form a preliminary archive set;
[0143] The preprocessing module 400 is used to perform preprocessing operations on the preliminary archive set to generate a secondary archive set;
[0144] A modeling module 500 uses an algorithm to model the secondary archive set, trains multiple sub-models for each archive instance, and each sub-model is associated with a supporting archive label;
[0145] A weight calculation module 600 is used to calculate the weight value of each supporting archive tag and adjust the output results of each sub-model based on the weight value;
[0146] The tag selection module 700 selects the tag with the largest weight value from the candidate tags output by the sub-model as the final file tag.
[0147] Therefore, no matter from which point of view, the embodiments should be regarded as illustrative and non-restrictive, and the scope of the present application is limited by the appended claims rather than the above description, and it is intended that all changes falling within the meaning and scope of the equivalent elements of the claims are included in the present application. Any figure mark in the claims should not be regarded as limiting the claims involved. In addition, it is obvious that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices stated in the device claim can also be implemented by the same unit or device through software or hardware. The words first, second, etc. are used to indicate names, and do not indicate any particular order.
[0148] The above embodiments are only used to illustrate the technical solution of the present application and are not intended to limit it. Although the present application has been described in detail with reference to the preferred embodiments, a person skilled in the art should understand that the technical solution of the present application may be modified or replaced by equivalents without departing from the spirit and scope of the technical solution of the present application.
Claims
1. A digital archive management method, characterized in that: The steps include: Obtain a set of digital archive instances, and extract time period information and a plurality of supporting archive tags in each archive instance; Each archive instance is represented by vectorization, and features are extracted from the vectorized data to obtain a feature data set; Based on the supporting archive tags, the corresponding time periods in the characterization dataset are classified, and time periods of the same category are clustered together to form a preliminary archive set; Performing preprocessing operations on the preliminary archive set to generate a secondary archive set; Using an algorithm to model the secondary archive set, multiple sub-models are trained for each archive instance, and each sub-model is associated with a supporting archive label; Calculate the weight value of each supporting archive label and adjust the output results of each sub-model based on the weight value; The label with the largest weight value is selected from the candidate labels output by the sub-model as the final archive label.
2. A digital archive management method according to claim 1, characterized in that: The method of vectorizing each archive instance includes the following steps: Get the text content of each archive instance; Perform word segmentation on the acquired text content to generate a word list; Use the pre-trained word vector model to convert each word in the word list into a corresponding word vector; Perform weighted averaging on the converted word vectors to obtain a comprehensive vector; Combine the comprehensive vector with the time period information to form a vector representation containing time features; The vector representation containing time features is standardized to generate the final vectorized representation.
3. A digital archive management method according to claim 2, characterized in that: The method for classifying the corresponding time periods in the characterization data set based on the supporting archive tags comprises the following steps: Determining time period information for each archive instance in the characterization data set; extracting archival instances containing specific supporting archival labels from the featurized dataset; sorting archive instances containing specific supporting archive tags by time period information; Divide the sorted time period information into multiple time intervals; For each time interval, calculate the similarity measure of its internal archive instances; Based on the similarity measurement, the time intervals with high similarity are classified into the same category; Summarize all category information and complete the classification of time periods.
4. A digital archive management method according to claim 3, characterized in that: The method of performing preprocessing operations on the preliminary archive set to generate a secondary archive set comprises the following steps: Check the integrity of each archive instance in the preliminary archive set; Fill in or delete data for archive instances with missing information; Standardize the formats of archive instances from different sources to obtain a unified data structure; Apply denoising algorithms to filter out noisy data present in the preliminary archive set; Identify and merge redundant or duplicate archive instances; Reorganize the preprocessed archive instances and sort them according to time periods and supporting archive tags; Arrange and save the preprocessed archive instances to form a secondary archive set.
5. A digital archive management method according to claim 4, characterized in that: The specific steps of using the algorithm to model the secondary archive set include the following steps: Extract the feature vector X of each archive instance from the secondary archive set i ; Normalize the extracted feature vectors using the formula: Among them, μ is the mean of the eigenvector, σ is the standard deviation of the eigenvector; Choose a machine learning algorithm as the modeling tool; Divide the secondary archive set into training set and validation set, and use cross-validation method to divide the data; Use the training set to train the model of the selected machine learning algorithm and calculate the loss function during the training process: Among them, f(Z i ,θ) is the model prediction result, y i is the true label, l is the loss function; N is the number of training samples; During the training process, the model parameters θ are adjusted to optimize the loss function L(θ), and the parameters are updated using the gradient descent method: Where η is the learning rate, Represents the gradient symbol, which is used to represent the gradient of the loss function L(θ) with respect to the parameter θ; Use the validation set to evaluate the performance of the trained model; The resulting model is saved, and multiple sub-models are trained for each archival instance, with each sub-model associated with a supporting archival label.
6. A digital archive management method according to claim 5, characterized in that: The steps of calculating and adjusting the weight value specifically include: Calculate the weight value w for each supporting archive label j , according to the output results y of each sub-model j , adjust the model output: Where M is the number of sub-models, w j represents the weight value of the jth sub-model; y j represents the output result of the jth sub-model; P is the final prediction result after comprehensively considering the outputs of all sub-models.
7. A digital archive management method according to claim 1, characterized in that: The method of selecting the label with the largest weight value from the candidate labels output by the sub-model as the finalized archive label specifically includes the following steps: Obtain the prediction results of each sub-model for the archive instance and its corresponding weight value; The prediction results of all sub-models and their weight values are weighted and summed to obtain a comprehensive score; The label with the highest comprehensive score is selected as the final label of the archive instance.
8. A digital archive management method according to claim 1, characterized in that: When modeling the secondary archive set, the machine learning algorithms used include, but are not limited to, one or more of the following: linear regression, logistic regression, support vector machine, random forest, gradient boosted decision tree, neural network and their variants.
9. A digital archive management method according to claim 1, characterized in that: The preprocessing operation on the preliminary archive set also includes desensitizing sensitive information in the archive instances.
10. A digital archive management system, characterized in that: Includes the following modules: A data acquisition module, used to acquire a set of digital archive instances and extract time period information and a plurality of supporting archive tags in each archive instance; A vectorization representation module is used to vectorize each archive instance and extract features from the vectorized data to obtain a feature data set; The classification module, based on the supporting archival labels, classifies the corresponding time periods in the characterization dataset and clusters time periods of the same category together to form a preliminary archival set; A preprocessing module, used for performing preprocessing operations on the preliminary archive set to generate a secondary archive set; A modeling module, which uses an algorithm to model the secondary archive set, trains multiple sub-models for each archive instance, and each sub-model is associated with a supporting archive label; The weight calculation module is used to calculate the weight value of each supporting archive label and adjust the output results of each sub-model based on the weight value; The label selection module selects the label with the largest weight value from the candidate labels output by the sub-model as the final archive label.