An archive management method and system based on a cloud archive repository
Through automated classification and construction of knowledge graphs, the problem of manual labeling and search efficiency of archive management platforms in the existing technology is solved, and efficient and accurate archive management and personalized recommendations are achieved.
Patent Information
- Application Number
- CN202311396364.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-26
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2043-10-26
AI Technical Summary
The existing digital archive management platform requires manual labeling and archive archive data, which is time-consuming and labor-intensive and error-prone. It only provides basic text search functions, and is inefficient in searching, making it difficult to quickly find files that users are interested in.
By obtaining archival data, extracting data characteristics, automatically classifying and managing, calculating the correlation between archival data, building a knowledge graph, and providing personalized archival recommendations and query responses based on user query input and the reviewed archival data.
It realizes automated file classification and management, improves the efficiency and accuracy of file management, reveals the relationships and connections between files, improves the efficiency of review, and helps users quickly find files of interest.
Smart Images

Figure CN117648433B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing, and particularly relates to an archive management method and system based on a cloud archive repository. Background Art
[0002] Archives, as the precipitation and accumulation of history, provide the most fundamental basis for the development of humanity. With the continuous update and development of modern technologies, archive management work has gradually tended towards digitization and informatization. The current forms of archive management are mainly divided into two types: electronic archives and paper archives, and the storage mode of most archive management still remains in the form of traditional paper archive management. With the continuous development of archive management business, the conventional storage, classification, and management of paper archives have made the demand for physical storage space in archives more and more tense. At the same time, the heavy work of sorting out archives not only brings inconvenience to the daily work of management personnel, but also invisibly increases the workload of managers, and does not essentially improve the utilization rate of archive materials. Therefore, a digital archive management platform has emerged as the times require.
[0003] However, after the current digital archive management platform collects archive data, it is necessary to manually annotate and file the archive data, which is time-consuming, laborious and error-prone. And the current digital archive management platform often only provides a basic text search function, and displays the archives whose names contain the search term in the form of a list. When the number of archives whose names contain the search term is too large, users need to click on each archive in turn to confirm when consulting archives, and the consulting efficiency is low, and it is difficult to quickly find the archives that users are interested in. Summary of the Invention
[0004] In order to solve the technical problems that the existing digital archive management platform needs to manually annotate and file archive data, which is time-consuming, laborious and error-prone, and only provides a basic text search function, and displays the archives whose names contain the search term in the form of a list. When the number of archives whose names contain the search term is too large, users need to click on each archive in turn to confirm when consulting archives, and the consulting efficiency is low, and it is difficult to quickly find the archives that users are interested in, the present invention provides an archive management method and system based on a cloud archive repository.
[0005] First Aspect
[0006] The present invention provides an archive management method based on a cloud archive repository, including:
[0007] S101: Obtain archive data;
[0008] S102: Extract the data features of the archive data;
[0009] S103: Classify the archival data according to the data characteristics of the archival data, and store the archival data in the corresponding category folders;
[0010] S104: Calculate the correlation degree between each piece of archival data;
[0011] S105: Construct a knowledge graph based on the correlation degree between each piece of archival data;
[0012] S106: Obtain the query input of the user;
[0013] S107: In response to the query input, display the archival data that matches the query input;
[0014] S108: Recommend archival data based on the knowledge graph according to the viewed archival data.
[0015] Second aspect
[0016] The present invention provides an archival management system based on a cloud archive library for executing the archival management method based on a cloud archive library in the first aspect.
[0017] Compared with the prior art, the present invention has at least the following beneficial technical effects:
[0018] (1) In the present invention, by extracting the data characteristics of the archival data and classifying according to these characteristics, automated archival classification and management can be achieved. It reduces the burden of manual annotation and classification, and improves the efficiency and accuracy of archival management.
[0019] (2) In the present invention, calculating the correlation degree between archival data can reveal the relationships and connections between different archives. By constructing a knowledge graph, the relevance between archives can be better understood and displayed, which helps to discover valuable information hidden in the data.
[0020] (3) In the present invention, based on the user's query input and the viewed archival data, the system can provide personalized archival recommendations and query responses, improving the viewing efficiency and helping users quickly find the archives they are interested in. Brief description of the drawings
[0021] The following will further illustrate the above characteristics, technical features, advantages and their implementation manners of the present invention in a clear and understandable manner in combination with the drawings in the preferred embodiments.
[0022] Figure 1 is a schematic flowchart of an archival management method based on a cloud archive library provided by the present invention;
[0023] Figure 2It is a schematic structural diagram of an archive management method based on a cloud archive library provided by the present invention. Detailed implementation manners
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will describe the specific implementation manners of the present invention with reference to the accompanying drawings. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings, and other implementation manners can also be obtained.
[0025] To make the drawings concise, only the parts related to the invention are schematically shown in each drawing, and they do not represent the actual structure of the product. In addition, to make the drawings concise and easy to understand, in some drawings, components with the same structure or function are only schematically shown for one of them, or only one of them is marked. In this document, "one" not only means "only this one", but also means the case of "more than one".
[0026] It should also be further understood that the term "and / or" used in the specification and appended claims of the present invention refers to any combination and all possible combinations of one or more of the related listed items, and includes these combinations.
[0027] In this document, it should be noted that unless otherwise clearly specified and limited, the terms "installation", "connection", and "connection" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection. It can be a mechanical connection or an electrical connection. It can be directly connected, or indirectly connected through an intermediate medium, and can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0028] In addition, in the description of the present invention, the terms "first", "second", etc. are only used for distinguishing descriptions, and cannot be understood as indicating or implying relative importance.
[0029] Embodiment 1
[0030] In one embodiment, referring to the attached drawings of the specification Figure 1 , a flowchart of the archive management method based on a cloud archive library provided by the present invention is shown. Referring to the attached drawings of the specification Figure 2 , a schematic structural diagram of an archive management method based on a cloud archive library provided by the present invention is shown.
[0031] An archive management method based on a cloud archive library provided by the present invention includes:
[0032] S101: Obtain archive data.
[0033] Optionally, the paper archives can be scanned to obtain archive data.
[0034] Optionally, various archive data, including historical documents, photos, files, etc., can be obtained from institutions such as public archives, libraries, and museums.
[0035] Optionally, there is a large amount of historical literature, pictures, and files on the Internet, and a search engine can be used to obtain specific archive data.
[0036] S102: Extract the data features of the archive data.
[0037] Among them, the data features can be word embedding features, word frequency features, subject model features, structured features, etc.
[0038] In a possible implementation manner, S102 specifically includes sub-steps S1021 to S1025:
[0039] S1021: Segment the archive data.
[0040] Specifically, a space-based word segmentation method, maximum matching method, dictionary-based word segmentation method, deep learning word segmentation method, etc. can be adopted to achieve word segmentation of the archive data.
[0041] S1022: Extract the word frequency of each word in the archive data:
[0042]
[0043] Among them, tf i represents the word frequency of the i-th word, c i represents the number of times the i-th word appears, and n represents the total number of words.
[0044] Among them, the word frequency can represent the importance or frequency of a certain word in the archive data.
[0045] S1023: Calculate the inverse document frequency of each word in the archive data:
[0046]
[0047] Among them, idf i represents the inverse document frequency of the i-th word, D represents the total number of archive data in the cloud archive, and D i represents the total number of archive data containing the i-th word.
[0048] Among them, the inverse document frequency can measure the rarity or particularity of a word in the archive data.
[0049] S1024: Calculate the text feature value of each word:
[0050] w i =tf ij ·idf i
[0051] Among them, w i Represents the text feature value of the i-th word.
[0052] It should be noted that extracting TF-IDF features can accurately reflect the importance and rarity of words in archival data, and has a good effect on distinguishing features between different archival data.
[0053] S1025: Sort the words in descending order of text feature values, and select the first preset number of text feature values that are ranked first to form a vector as the feature vector of the archive data:
[0054]
[0055] Where W represents the data feature, It represents the first ith feature value in the order of text feature values from large to small, s represents the first preset number, and s can also be called the dimension of the feature vector.
[0056] Among them, those skilled in the art can set the size of the first preset number according to actual conditions, and the present invention does not limit it.
[0057] It should be noted that in the archival data, some words may be very common words, such as "的", "是", etc. They appear in most documents, but their information content is relatively low. After sorting the text feature values from large to small, these common words that do not have high information content can be excluded from the feature vector, thereby reducing the dimension of the feature vector and improving the relevance and importance of the features. Furthermore, words ranked at a larger feature value are often keywords and subject terms with higher importance in the current archival data. By selecting these keywords as part of the feature vector, the theme and content of the document can be better expressed.
[0058] S103: Classify the archival data according to the data features of the archival data, and store the archival data in corresponding category folders.
[0059] In the present invention, by extracting data features of archive data and classifying according to these features, automatic archive classification and management can be achieved, which reduces the burden of manual labeling and classification and improves the efficiency and accuracy of archive management.
[0060] In a possible implementation, S103 specifically includes sub-steps S1031 to S1034:
[0061] S1031: Initialize K classification centers.
[0062] Among them, each classification center corresponds to an archive category.
[0063] Specifically, the archive categories can specifically be: historical archives, geographical information archives, scientific research archives, cultural heritage archives, legal archives, and so on.
[0064] S1032: According to the feature vectors of each archive data, calculate the distances from the current archive data to the center points of each classification center:
[0065]
[0066] Among them, D j represents the distance from the current archive data to the j-th classification center, represents the i-th eigenvalue in the feature vector, c ij represents the i-th eigenvalue in the feature vector of the center point of the j-th classification center, and s represents the dimension of the feature vector.
[0067] S1033: Divide the current archive data into the classification with the smallest D j and update the classification center.
[0068] S1034: Continue to select the next archive data until the classification of all archive data is completed.
[0069] In the present invention, by initializing K classification centers, the archive data can then be automatically clustered according to the feature vectors of the archive data without the need to know the data categories in advance. Each classification center corresponds to an archive category, so that the archive data is automatically grouped according to similarity, and automated archive classification and management can be achieved. It reduces the burden of manual annotation and classification, and improves the efficiency and accuracy of archive management.
[0070] S104: Calculate the correlation degrees between each archive data.
[0071] Specifically, the correlation degrees between each archive data can be calculated by calculating cosine similarity, Euclidean distance, Pearson correlation coefficient, etc.
[0072] In a possible implementation manner, S104 specifically includes:
[0073] Calculate the correlation degrees between each archive data according to the following formula:
[0074]
[0075] Among them, σ ijrepresents the degree of association between the i-th file data and the j-th file data, W i represents the feature vector of the i-th file data, W j represents the feature vector of the j-th file data, and ||·|| represents the modulus operation of the vector.
[0076] S105: Construct a knowledge graph according to the degree of association between each file data.
[0077] Among them, the knowledge graph is a structured data representation method used to store and represent the relationships, attributes, and other knowledge between entities. The knowledge graph can be regarded as a semantic network for presenting the entities, concepts, and their connections in the real world.
[0078] In a possible implementation manner, S105 is specifically:
[0079] Use each file data as a node and use the degree of association between each file data as the path connecting the nodes to construct a knowledge graph.
[0080] Among them, the calculation method of the path length is:
[0081]
[0082] Among them, l ij represents the path length between the i-th node and the j-th node, l ij can also be called the path length between the i-th file data and the j-th file data, σ ij represents the degree of association between the i-th node and the j-th node, σ ij can also be called the degree of association between the i-th file data and the j-th file data, and e represents the natural logarithm.
[0083] It should be noted that through the above formula, it can be seen that the higher the degree of association between two file data, the shorter the distance between the nodes corresponding to the two file data. The distance between nodes briefly reflects the degree of association between file data, which can make the visual presentation of the degree of association in the graph more intuitive. Data pairs with high degrees of association are closer in the graph, which helps to more accurately express the close association between data.
[0084] In the present invention, calculating the degree of association between file data can reveal the relationships and connections between different files. By constructing a knowledge graph, the relevance between files can be better understood and displayed, which helps to discover valuable information hidden in the data.
[0085] In a possible implementation manner, in the knowledge graph, each node is displayed in a circular shape, and the calculation method of the circular radius of each node is:
[0086]
[0087] Among them, r represents the circular radius of the node, λ represents the conversion factor, It represents the first ith feature value in the order of text feature values from large to small, s represents the preset number, and s can also be called the dimension of the feature vector.
[0088] In the present invention, the circular radius of a node can be used to indicate the importance of the node. A larger circular radius can attract the user's attention, thereby highlighting important nodes. This can make the knowledge graph more visually attractive and informative, providing users with a more interesting and effective way to analyze and explore data.
[0089] S106: Obtaining the user's query input.
[0090] The query input may be a search input after the user types text in the search box, or may be a voice input initiated by the user to query archive data. The present invention does not limit the specific form of the query input.
[0091] S107: In response to the query input, display the archive data matching the query input.
[0092] In actual application, current digital archive management platforms often only provide basic text search functions, and display archives containing search words in the archive names in the form of a list. This search efficiency is low. In the present invention, a faster and more accurate archive search method is provided.
[0093] In a possible implementation, S107 specifically includes:
[0094] S1071: Parse the query input, combine the query history records, and obtain a query feature vector of the query input through a recurrent neural network.
[0095] It should be noted that the recurrent neural network, combined with the user's query history, can identify the user's query intention over a period of time, thereby providing the user with personalized query results. Compared with traditional text search, this method can locate the archival data that the user is interested in more quickly and improve the user's information acquisition efficiency.
[0096] In a possible implementation, S10711: construct a recurrent neural network, the recurrent neural network comprising: an input layer, a state layer, an attention layer and an output layer.
[0097] S10712: At the input layer, each query text in the query history is input to form a query text sequence [x 1 ,…,x t,…x n , where x t represents the query text at the t-th query, and n represents the number of queries.
[0098] S10713: In the state layer, calculate the hidden state of the query text at the t-th query:
[0099]
[0100]
[0101]
[0102] where h t represents the hidden state of the query text at the t-th query, represents the previous state of the forward loop, represents the previous state of the backward loop, GRU() represents the non-linear calculation through the recurrent neural network, and u t represents 's weight coefficient, and v t represents 's weight coefficient, and p t represents the bias term of the hidden state at the t-th query.
[0103] S10714: In the attention layer, assign weights to each query text and accumulate to obtain the hidden state of the current attention layer:
[0104]
[0105] where s represents the hidden state of the current attention layer, and γ t represents the weight of the query text at the t-th query, and h t represents the hidden state of the query text at the t-th query.
[0106] S10715: In the output layer, output the query feature vector A of the current query input:
[0107] A = tan(u n s + p n )
[0108] where u n represents the weight coefficient of the previous state of the forward loop of the current query input, i.e., the n-th query, and p n represents the bias term of the hidden state of the current query input, i.e., the n-th query.
[0109] In the present invention, through the recurrent neural network, combined with the user's query history, the query intention of the user over a period of time is recognized, so as to provide personalized query results for the user.
[0110] S1072: Calculate the matching degree between the query feature vector and the archive data in the cloud archive according to the following formula:
[0111]
[0112] where, τ i represents the matching degree between the query feature vector and the i-th archive data in the cloud archive, W i represents the feature vector of the i-th archive data, A represents the query feature vector, and ||·|| represents the modulus operation of the vector.
[0113] It should be noted that by calculating the matching degree between the query feature vector and the feature vector of the archive data, the archive data with a relatively high relevance to the user's query intention can be quickly screened out.
[0114] S1073: Sort each archive data in descending order of the matching degree, and select the second preset number of archive data with the top ranking for display.
[0115] Among them, those skilled in the art can set the size of the second preset number according to the actual situation, and the present invention does not make a limitation.
[0116] It should be noted that by sorting the archive data according to the matching degree and displaying the archive data with the top ranking, a recommendation function can also be provided for the user to help the user discover potentially interesting data.
[0117] In the present invention, through a recurrent neural network, combined with the user's query history, the query intention of the user in a period is recognized, so as to provide personalized query results for the user. Compared with traditional text search, this method can locate the archive data of interest to the user more quickly and improve the information acquisition efficiency of the user.
[0118] S108: Recommend archive data based on the knowledge graph according to the viewed archive data.
[0119] In a possible implementation manner, S108 specifically includes:
[0120] S1081: Obtain the viewed archive data.
[0121] S1082: In the knowledge graph, calculate the comprehensive path length from the viewed archive data to each unviewed archive data:
[0122]
[0123] where, L i represents the comprehensive path length from the viewed archive data to the i-th unviewed archive data, l ijDenote the path length between the i-th unexamined archive data and the j-th examined archive data, and m represents the total number of examined archive data.
[0124] Among them, the comprehensive path can be understood as the sum of the path lengths from each examined archive data to a certain unexamined archive data.
[0125] It should be noted that the shorter the comprehensive path length, the more similar this unexamined archive data is to the previously examined archive data. By calculating the comprehensive path length from the examined archive data to the unexamined archive data, unexamined archive data with a relatively high degree of association with the examined archive data can be found. This kind of relevance recommendation helps users discover data related to their interests.
[0126] S1083: Sort each unexamined archive data in ascending order of the comprehensive path length, and select the third preset number of unexamined archive data with the top ranking for recommendation.
[0127] Among them, those skilled in the art can set the size of the third preset number according to the actual situation, and the present invention does not make a limitation.
[0128] In the present invention, by analyzing the archive data examined by the user, the interests and preferences of the user can be understood, so as to realize personalized archive recommendation. This helps users quickly find the content they are interested in. Further, by calculating the comprehensive path length from the examined archive data to the unexamined archive data in the knowledge graph, unexamined archive data with a relatively high degree of association with the examined archive data can be found. This kind of recommendation based on relevance can more accurately meet the information needs of users.
[0129] Compared with the prior art, the present invention has at least the following beneficial technical effects:
[0130] (1) In the present invention, by extracting the data features of the archive data and classifying them according to these features, automated archive classification and management can be realized. The burden of manual annotation and classification is reduced, and the efficiency and accuracy of archive management are improved.
[0131] (2) In the present invention, by calculating the degree of association between archive data, the relationships and connections between different archives can be revealed. By constructing a knowledge graph, the relevance between archives can be better understood and displayed, which helps to discover valuable information hidden in the data.
[0132] (3) In the present invention, based on the user's query input and the examined archive data, the system can provide personalized archive recommendation and query response, improve the access efficiency, and help users quickly find the archives they are interested in.
[0133] Embodiment 2
[0134] In one embodiment, an archive management system based on a cloud archive repository provided by the present invention is used to execute the archive management method based on the cloud archive repository in Embodiment 1.
[0135] The archive management system based on the cloud archive repository provided by the present invention can achieve the steps and effects of the archive management method based on the cloud archive repository in Embodiment 1 above. To avoid repetition, the present invention will not elaborate further.
[0136] Compared with the prior art, the present invention has at least the following beneficial technical effects:
[0137] (1) In the present invention, by extracting the data features of archive data and classifying them according to these features, automated archive classification and management can be achieved. This reduces the burden of manual annotation and classification, and improves the efficiency and accuracy of archive management.
[0138] (2) In the present invention, by calculating the correlation degree between archive data, the relationships and connections between different archives can be revealed. By constructing a knowledge graph, the relevance between archives can be better understood and displayed, which helps to discover valuable information hidden in the data.
[0139] (3) In the present invention, based on the user's query input and the viewed archive data, the system can provide personalized archive recommendations and query responses, improving the viewing efficiency and helping users quickly find the archives they are interested in.
[0140] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combinations of these technical features do not conflict, they should be considered as within the scope described in this specification.
[0141] The above embodiments only represent several implementation manners of the present invention, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the appended claims.
Claims
1. An archive management method based on a cloud archive repository, characterized in that, it includes: S101: Obtain archive data; S102: Extract the data features of the archive data; S103: Classify the archive data according to the data features of the archive data, and store the archive data in the corresponding category folders; S104: Calculate the correlation degree between each archive data; S105: Construct a knowledge graph according to the correlation degree between each archive data; S106: Obtain the query input of the user; S107: In response to the query input, display the archive data that matches the query input; S108: Recommend archive data based on the knowledge graph according to the viewed archive data; wherein, the S102 specifically includes: S1021: Segment the archive data; S1022: Extract the word frequency of each word in the archive data: Among them, tf i represents the term frequency of the i-th word, c i represents the number of times the i-th word appears, and n represents the total number of words; S1023: Calculate the inverse document frequency of each word in the archive data; Among them, idf i represents the inverse document frequency of the i-th word, D represents the total number of document data in the cloud database, D i represents the total number of document data containing the i-th word; S1024: Calculate the text feature value of each word; w i = tf i ·idf i Among them, w i represents the text feature value of the i-th word; S1025: Sort each word in descending order of the text feature value, and select the first preset number of text feature values sorted in the front to form a vector as the feature vector of the archive data: Where W represents the data feature, represents the i-th feature value in the order of text feature values from large to small, s represents the first preset number, and s can also be called the dimension of the feature vector; wherein, the S104 specifically includes: Calculate the correlation degree between each archive data according to the following formula: Among them, σ ij represents the correlation degree between the i-th file data and the j-th file data, W i represents the feature vector of the i-th file data, W j represents the feature vector of the j-th file data, ||·|| represents the modulus operation of the vector; wherein, the S105 is specifically: Take each archive data as a node, and take the correlation degree between each archive data as the path connecting the nodes to construct a knowledge graph; wherein, the calculation method of the path length is: where, l ij represents the path length between the i-th node and the j-th node, and l ij can also be referred to as the path length between the i-th file data and the j-th file data. σ ij represents the correlation degree between the i-th node and the j-th node, and σ ij can also be referred to as the correlation degree between the i-th file data and the j-th file data. e represents the natural logarithm; wherein, the S108 specifically includes: S1081: Obtain the viewed archive data; S1082: In the knowledge graph, calculate the comprehensive path length from the viewed archive data to each unviewed archive data; Among them, L i represents the comprehensive path length from the archived data that has been accessed to the i-th unaccessed archived data, and l ij represents the path length between the i-th unaccessed archived data and the j-th archived data that has been accessed, and m represents the total number of archived data that has been accessed; S1083: Sort each unviewed archive data in ascending order of the comprehensive path length, and select the third preset number of unviewed archive data sorted in the front for recommendation.
2. The archive management method based on a cloud archive repository according to claim 1, characterized in that, the S103 specifically includes: S1031: Initialize K classification centers, where each classification center corresponds to an archive category; S1032: Calculate the distance from the current archive data to the center points of each classification center according to the feature vectors of each archive data: Among them, D j represents the distance from the current file data to the j-th classification center, represents the i-th eigenvalue in the feature vector, c ij represents the i-th eigenvalue in the feature vector of the center point of the j-th classification center, and s represents the dimension of the feature vector; S1033: Divide the current file data into the smallest classification of D j and update the classification center; S1034: Continue to select the next archive data until all archive data are classified.
3. The archive management method based on a cloud archive repository according to claim 1, characterized in that, in the knowledge graph, each node is displayed in a circular manner, and the calculation method of the circular radius of each node is: where r represents the circular radius of the node, and λ represents the conversion coefficient. represents the i-th eigenvalue among the top i eigenvalues sorted in descending order according to the text eigenvalue, s represents the preset quantity, and s can also be referred to as the dimension of the eigenvector.
4. The archive management method based on a cloud archive repository according to claim 1, characterized in that, the S107 specifically includes: S1071: Parse the query input, combine the query history record, and obtain the query feature vector of the query input through a recurrent neural network; S1072: Calculate the matching degree between the query feature vector and the archive data in the cloud archive according to the following formula: Among them, τ i represents the matching degree between the query feature vector and the i-th archival data in the cloud archive, W i represents the feature vector of the i-th archival data, A represents the query feature vector, and ||·|| represents the modulus operation of the vector; S1073: Sort each piece of archive data in descending order of the matching degree, and select the second preset number of archive data with the top rankings for display.
5. The archive management method based on a cloud archive according to claim 4, wherein, the specific steps of S1071 include: S10711: Construct a recurrent neural network, which includes an input layer, a state layer, an attention layer, and an output layer; S10712: At the input layer, each query text in the query history is input to form a query text sequence [x 1 , …, x t , … x n , where x t represents the query text at the t-th query, and n represents the number of queries; S10713: In the state layer, calculate the hidden state of the query text at the t-th query; Among them, h t represents the hidden state of the query text at the t-th query, represents the previous state of the forward loop, represents the previous state of the backward loop, GRU() represents the non-linear calculation through the recurrent neural network, u t represents the weight coefficient of, v t represents the weight coefficient of, p t represents the bias term of the hidden state at the t-th query; S10714: In the attention layer, assign weights to each query text and accumulate them to obtain the hidden state of the current attention layer; where s represents the hidden state of the current attention layer, γ t represents the weight of the query text at the t-th query, h t represents the hidden state of the query text at the t-th query; S10715: In the output layer, output the query feature vector A of the current query input; A = tan(u n s + p n ) Among them, u n represents the weight coefficient of the previous state of the forward loop at the current query input, that is, the nth query, and p n represents the bias term of the hidden state at the current query input, that is, the nth query.
6. An archive management system based on a cloud archive, wherein, it is used to execute the archive management method based on a cloud archive according to any one of claims 1 to 5.
Citation Information
Patent Citations
Text clustering method
CN106951498A
NLP-based scientific research archive management method and system
CN116909991A