Intelligent archive editing and research method and system based on intelligent agent

Through intelligent archive editing and research methods based on intelligent objects, the problem of low efficiency of traditional archive editing and research has been solved, efficient and intelligent archive editing and research have been achieved, and the efficiency of utilization of archive resources has been improved.

CN120218065AActive Publication Date: 2025-06-27WUHAN UNIV

Patent Information

Application Number
CN202510349210.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-06-27
Estimated Expiration
2045-03-24

AI Technical Summary

Technical Problem

The archive editing and research methods in the existing technology are inefficient and time-consuming, and it is difficult to meet the needs of the information explosion era.

Method used

The intelligent archival editing and research method based on the agent is adopted. By searching archival data from designated data sources, preprocessing and streamlining, embedding vectors are extracted and keyword databases are constructed, target keywords are determined, and abstracts are generated using the agent and intelligently compiled and researched.

Benefits of technology

It has greatly improved the efficiency of editing and research, reduced the cost of editing and research, promoted the development and utilization of archive resources, realized the intelligence and personalization of intelligent archive editing and research, and brought more possibilities to archive work.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218065A_ABST
    Figure CN120218065A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent archive editing and research method and system based on an intelligent agent, and belongs to the technical field of archive data processing. Keywords of archive data are extracted, then abstracts corresponding to the archive data are generated on the basis of the extracted keywords, and finally the archive data are edited and researched according to the keywords corresponding to the archive data and the abstracts. The intelligent agent is adopted to achieve intelligent file editing and research, the editing and research efficiency can be greatly improved, the editing and research cost is reduced, development and utilization of file resources are promoted, intelligent file editing and research are more intelligent and personalized, and more possibilities are brought to file work.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of archival data processing, and particularly relates to an intelligent archival compilation method and system based on an intelligent agent. Background Art

[0002] Archival data is an important carrier for recording information such as history, culture, society, and economy, including various types such as text, images, audio, video, and electronic archives. Its characteristics include authenticity, integrity, diversity, long-term nature, and confidentiality. The sources are extensive, covering government agencies, enterprises and institutions, social organizations, and individuals. Archival data has multiple values such as historical, legal, scientific research, educational, and decision-making, and is an important basis for research, proof, reference, inheritance, and decision-making. With the development of informatization, the digital and intelligent management of archival data has become a trend, and the efficient processing and utilization of archival data have become the key to archival management. Archival compilation is an important task for excavating the value of archives and serving society. The traditional archival compilation method is inefficient and time-consuming, and it is difficult to meet the needs of the information explosion era. Summary of the Invention

[0003] The present invention provides an intelligent archival compilation method and system based on an intelligent agent to solve the problems in the prior art that the archival compilation method is inefficient and time-consuming and difficult to meet the needs of the information explosion era.

[0004] On the one hand, the present invention provides an intelligent archival compilation method based on an intelligent agent, including: Retrieving archival data from a specified data source or a specified database, and preprocessing the archival data to obtain the preprocessed archival data; Performing a trimming process on the preprocessed archival data to obtain the trimmed archival data, and extracting the embedding vectors in the trimmed archival data; After performing a keyword extraction operation on the embedding vectors in all the archival data, constructing the extracted keywords into an archival keyword library; For any piece of archival data, based on the archival keyword library, determining at least one target keyword corresponding to the archival data; Based on at least one target keyword corresponding to the archival data, using an intelligent agent to extract the abstract corresponding to the archival data to obtain the archival directory corresponding to the archival data; According to at least one target keyword corresponding to the archival data and the archival directory, performing intelligent compilation on the archival data to complete the intelligent archival compilation based on the intelligent agent.

[0005] In a possible implementation manner, retrieving archival data from a specified data source or a specified database, and preprocessing the archival data to obtain the preprocessed archival data, includes: Retrieve archival data from a specified data source or a specified database to obtain multiple pieces of archival data; among them, the archival data includes text data; Perform deduplication processing on the archival data to obtain the archival data after deduplication processing; Perform format standardization processing on the archival data after deduplication processing to obtain the archival data after format standardization, and use the archival data after format standardization as the archival data after preprocessing.

[0006] In a possible implementation manner, perform refinement processing on the archival data after preprocessing to obtain the archival data after refinement processing, including: Remove the images and tabular data in the archival data after preprocessing to obtain the archival data after initial processing; Perform text cleaning processing on the archival data after initial processing to obtain the archival data after text cleaning processing; Perform word segmentation and stop word removal processing on the archival data after text cleaning processing, and use the archival data after word segmentation and stop word removal processing as the archival data after refinement processing.

[0007] In a possible implementation manner, extract the embedding vectors in the archival data after refinement processing, including: Input the archival data after refinement processing into the BERT language model to process the archival data after refinement processing through the word embedding function of the BERT language model, and obtain the embedding vectors in the archival data after refinement processing.

[0008] In a possible implementation manner, after performing keyword extraction operations on the embedding vectors in all archival data, construct the extracted keywords into an archival keyword library, including: Use the UMAP algorithm to perform dimensionality reduction processing on the embedding vectors in the archival data after refinement processing to obtain the embedding vectors after dimensionality reduction processing; Based on the embedding vectors after dimensionality reduction processing, use the HDBSCAN algorithm to perform clustering processing to obtain multiple clustering topics; For any one clustering topic, use the c-TD-IDF algorithm to obtain the c-TD-IDF values corresponding to the vocabulary in each clustering topic; For any one clustering topic, use the vocabulary with the largest c-TD-IDF value as the keyword corresponding to the clustering topic; For any one clustering topic, calculate the MMR values corresponding to the other vocabulary except the keyword in turn, and select the vocabulary with the largest MMR value as the new keyword, and repeat this step until the keyword extraction end condition is met, to obtain all the keywords corresponding to each clustering topic; Construct an archive keyword library based on all the keywords corresponding to all clustering topics.

[0009] In a possible implementation, for any archive data, based on the archive keyword library, determining at least one target keyword corresponding to the archive data includes: For any archive data to be processed, obtain the weighted vocabulary graph corresponding to the archive data; Based on the weighted vocabulary graph corresponding to the archive data and the archive keyword library, use the TextRank algorithm to obtain the score corresponding to each vocabulary in the archive data to be processed; According to the scores corresponding to each vocabulary in the archive data to be processed, determine at least one target keyword corresponding to the archive data to be processed.

[0010] In a possible implementation, based on at least one target keyword corresponding to the archive data, use an agent to extract the abstract corresponding to the archive data to obtain the archive directory corresponding to the archive data, including: Take at least one target keyword corresponding to the archive data and the archive data as the input of the first agent to obtain the target keyword vector and the archive data text vector; wherein, the first agent includes an XLNet model; Use a second agent to process the archive data text vector to obtain the forward semantic feature text vector and the reverse semantic feature text vector; wherein, the second agent includes a BiGRU model; Use a third agent to process the target keyword vector, the archive data text vector, the forward semantic feature text vector, and the reverse semantic feature text vector to obtain the archive directory corresponding to the archive data; wherein, the third agent includes a decoder based on an attention mechanism.

[0011] In a possible implementation, before using the agent, it further includes: training the hyperparameters of the agent.

[0012] In a possible implementation, according to at least one target keyword corresponding to the archive data and the archive directory, perform intelligent compilation and research on the archive data, including: According to at least one target keyword corresponding to the archive data, take the N keywords with the highest scores as the title corresponding to the archive data; Take the title corresponding to the archive data and the archive directory corresponding to the archive data together as the personalized compilation and research outline corresponding to the archive data to complete the intelligent archive compilation based on the agent.

[0013] On the other hand, the present invention provides an intelligent archival compilation and research system based on an agent, including: a data preprocessing module, a vector extraction module, a keyword library construction module, a keyword extraction module, an abstract extraction module, and an intelligent compilation and research module; The data preprocessing module is used to retrieve archival data from a specified data source or a specified database, and preprocess the archival data to obtain the preprocessed archival data; The vector extraction module is used to perform streamlining processing on the preprocessed archival data to obtain the streamlined archival data, and extract the embedding vectors in the streamlined archival data; The keyword library construction module is used to extract keywords from the embedding vectors in all archival data, and construct the extracted keywords into an archival keyword library; The keyword extraction module is used to determine at least one target keyword corresponding to the archival data based on the archival keyword library for any piece of archival data; The abstract extraction module is used to extract the abstract corresponding to the archival data by using an agent based on at least one target keyword corresponding to the archival data, and obtain the archival directory corresponding to the archival data; The intelligent compilation and research module is used to perform intelligent compilation and research on the archival data according to at least one target keyword corresponding to the archival data and the archival directory, and complete the intelligent archival compilation and research based on the agent.

[0014] An intelligent archival compilation and research method and system based on an agent provided by the present invention can greatly improve the compilation and research efficiency, reduce the compilation and research cost, promote the development and utilization of archival resources, make the intelligent archival compilation and research more intelligent and personalized, and bring more possibilities to archival work by extracting keywords from archival data, then generating the abstract corresponding to the archival data based on the extracted keywords, and finally realizing intelligent archival compilation and research by using an agent according to the keywords and abstract corresponding to the archival data. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention.

[0016] Figure 1 It is a flowchart of an intelligent archival compilation and research method based on an agent provided by an embodiment of the present invention.

[0017] Figure 2 It is a schematic structural diagram of an intelligent archival compilation and research system based on an agent provided by an embodiment of the present invention.

[0018] Through the above-mentioned accompanying drawings, specific embodiments of the present invention have been shown, and will be described in more detail hereinafter. These drawings and the written description are not intended to limit the scope of the inventive concept in any way, but to illustrate the concept of the present invention to those skilled in the art by reference to specific embodiments. Detailed Description of the Embodiments

[0019] Here, exemplary embodiments will be described in detail, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numerals in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. On the contrary, they are merely examples of apparatuses and methods consistent with some aspects of the present invention as detailed in the appended claims.

[0020] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0021] As Figure 1 shown, an agent-based intelligent compilation and research method for archives provided by an embodiment of the present invention includes: S101. Retrieve archive data from a specified data source or a specified database, and preprocess the archive data to obtain the preprocessed archive data; Before the agent processes, preprocessing the archive data is a crucial step. As a carrier of historical records and information resources, archive data often contains rich knowledge and value, but its original form is not always suitable for direct input into the agent for analysis and mining. Therefore, preprocessing has become an indispensable link to ensure data quality and improve model performance. The primary task of preprocessing is data cleaning. Since archive data may come from different historical periods and different recording systems, its formats, standards, and quality are often uneven. First, data cleaning is performed, and its purpose is to remove impurities, redundant information, and errors in the data, such as deleting duplicate entries, correcting spelling mistakes, and standardizing formats, to ensure the correctness and uniformity of the data.

[0022] S102. Perform a reduction process on the preprocessed archive data to obtain the reduced archive data, and extract the embedding vectors from the reduced archive data; Archive data often has multiple data types, and some irrelevant data not only has no effect on archive compilation and research but also affects the implementation of the archive compilation and research plan. Therefore, some irrelevant data in the archive data can be reduced to improve the accuracy of archive compilation and research.

[0023] To achieve archive compilation and research, it is also necessary to extract the embedding vectors from the reduced archive data to achieve data identification and extraction.

[0024] S103. After performing the operation of extracting keywords from the embedding vectors in all archival data, construct an archival keyword library with the extracted keywords. One of the focuses of archival compilation and research in the embodiments of the present invention is to extract archival catalogs. Although the research on automatic text summarization has gradually increased in recent years, the archival compilation and research datasets in some specific fields are still relatively scarce. Therefore, the embodiments of the present invention can construct an archival keyword library based on a batch of archival data. When conducting archival compilation and research subsequently, the archival keyword library is used as the basis for archival compilation and research, thereby improving the accuracy of archival compilation and research.

[0025] S104. For any piece of archival data, based on the archival keyword library, determine at least one target keyword corresponding to the archival data. The target keywords can be selected according to a preset quantity, so as to select a preset number of target keywords; or they can be selected according to preset conditions, so as to select target keywords that meet the conditions.

[0026] S105. Based on at least one target keyword corresponding to the archival data, use an agent to extract the abstract corresponding to the archival data, and obtain the archival catalog corresponding to the archival data. Traditional extractive summarization methods are widely used in the field of automatic text summarization. It forms a summary by extracting some key sentences or phrases from the original text. Although this method shows certain advantages in maintaining the meaning of the original text, how to build a powerful model that can understand and capture the deep meaning of the text is still a challenging problem. In particular, how to effectively extract key information from long texts and generate a compact, accurate, and fluent summary is crucial for the quality of text summarization. With the rapid development of the natural language processing (NLP) field, generative summarization methods, especially agent-based summary generation methods, have gradually become a research hotspot. Therefore, the embodiments of the present invention use an agent to extract the abstract corresponding to the archival data to improve the accuracy of abstract generation.

[0027] S106. According to at least one target keyword corresponding to the archival data and the archival catalog, conduct intelligent compilation and research on the archival data, and complete the archival intelligent compilation and research based on the agent.

[0028] For example, the compilation and research conditions can be preset as follows: use one or more target keywords as the title corresponding to the archival data (such as setting a title and a subtitle, that is, using the target keyword with the highest score as the title and the target keyword with the second highest score as the subtitle), then use the storage order of the archival data as the serial number corresponding to the archival data, and at the same time supplement with the archival catalog to achieve the compilation and research of the archival data.

[0029] An agent-based intelligent compilation method for archives provided by the present invention extracts keywords from archive data, then generates an abstract corresponding to the archive data based on the extracted keywords, and finally uses an agent to achieve intelligent compilation of archives according to the keywords and abstracts corresponding to the archive data, which can greatly improve the compilation efficiency, reduce the compilation cost, promote the development and utilization of archive resources, make the intelligent compilation of archives more intelligent and personalized, and bring more possibilities to archive work.

[0030] In a possible implementation, retrieve archive data from a specified data source or a specified database, and preprocess the archive data to obtain the preprocessed archive data, including: Retrieve archive data from a specified data source or a specified database to obtain multiple archive data; among them, the archive data includes text data; Perform duplicate removal processing on the archive data to obtain the archive data after duplicate removal processing; for example, a duplicate check threshold can be set in advance. When the text similarity between two archive data exceeds this duplicate check threshold, it indicates that these two archive data are duplicates, and one of them can be randomly removed or the one with an older time can be removed.

[0031] Perform format standardization processing on the archive data after duplicate removal processing to obtain the archive data after format standardization, and use the archive data after format standardization as the preprocessed archive data.

[0032] For archive data from different sources, there may be different formats. For unified processing, the archive data can be converted into a unified format, which is more convenient for processing.

[0033] In a possible implementation, perform refinement processing on the preprocessed archive data to obtain the refined archive data, including: Remove the images and table data in the preprocessed archive data to obtain the initially processed archive data; the images and table data in the archive data do not play a major role in archive compilation. Therefore, the embodiments of the present invention can remove these useless data to improve the accuracy of archive compilation.

[0034] Perform text cleaning processing on the initially processed archive data to obtain the archive data after text cleaning processing; for example, special symbols, numerical values, and punctuation can be removed from the initially processed archive data. These elements often interfere with archive compilation, so they can be removed to ensure the accuracy of the final result.

[0035] Perform word segmentation and stop word removal processing on the archive data after text cleaning processing, and use the archive data after word segmentation and stop word removal processing as the refined archive data.

[0036] In a possible implementation, extracting the embedding vectors from the file data after the extraction and refinement process includes: Inputting the file data after the extraction and refinement process into the BERT language model to process the file data after the extraction and refinement process through the word embedding function of the BERT language model, and obtaining the embedding vectors in the file data after the extraction and refinement process.

[0037] The BERT model has been pre-trained on a large amount of text and can understand the complex models and context relationships in language. The sentences in the document are converted into a fixed-length vector. This vector represents the semantic content of the document in a multi-dimensional space. When the file data set is input into the model, the output layer of the model will generate the embedding vector of this text.

[0038] In a possible implementation, after performing the keyword extraction operation on the embedding vectors in all file data, constructing the extracted keywords into a file keyword library includes: Using the UMAP algorithm to perform dimensionality reduction processing on the embedding vectors in the file data after the extraction and refinement process to obtain the embedding vectors after the dimensionality reduction processing; UMAP (Uniform Manifold Approximation and Projection) is a non-linear dimensionality reduction algorithm for high-dimensional data visualization. It aims to map high-dimensional data into a low-dimensional space while preserving the local and global structure of the data. The UMAP algorithm has been widely used in many fields, such as bioinformatics, image processing, and natural language processing.

[0039] The main features of the UMAP algorithm include: (1) Non-linear dimensionality reduction: UMAP can capture the non-linear structure in high-dimensional data and map it into a low-dimensional space. (2) Preservation of local and global structure: UMAP considers both the local and global structure of the data during the dimensionality reduction process, enabling the low-dimensional representation to better reflect the characteristics of the high-dimensional data. (3) High efficiency: The UMAP algorithm is relatively efficient in terms of computation and can handle large-scale data sets. (4) Scalability: UMAP can adapt to data sets of different scales and complexities. (5) Few parameters: The UMAP algorithm has relatively few parameters and is easy to adjust and optimize.

[0040] Based on the embedding vectors after the dimensionality reduction processing, using the HDBSCAN algorithm to perform clustering processing to obtain multiple clustering themes; HDBSCAN (Hierarchical Density-Based Spatial Clustering of Applications with Noise) **is a density-based spatial clustering algorithm that can identify clusters of arbitrary shapes and handle noisy data. HDBSCAN does not require the number of clusters to be specified in advance, which makes it very useful in many practical applications.

[0041] Using UMAP text embedding dimensions for dimensionality reduction and the HDBSCAN algorithm for clustering based on document embeddings, similar documents will be assigned to the same category, which can make data processing more efficient.

[0042] For any clustering topic, the c-TD-IDF algorithm is used to obtain the c-TD-IDF values corresponding to the vocabulary in each clustering topic; c-TD-IDF (Cluster-based Term Frequency-Inverse Document Frequency) is an improved TF-IDF algorithm for text feature representation. The traditional TF-IDF algorithm evaluates the importance of a word in a document set or corpus through term frequency (TF) and inverse document frequency (IDF). On this basis, c-TD-IDF introduces clustering information to better capture the semantic features of the text.

[0043] For any clustering topic, the vocabulary with the largest c-TD-IDF value is used as the keyword corresponding to that clustering topic; For any clustering topic, the MMR values corresponding to the other vocabulary except the keyword are calculated in turn, and the vocabulary with the largest MMR value is selected as the new keyword. Repeat this step until the keyword extraction end condition is met (such as when the number of keywords reaches the requirement or the MMR value does not meet the threshold requirement), to obtain all the keywords corresponding to each clustering topic; The MMR value usually refers to the value of Maximum Marginal Relevance. MMR is a technique used in information retrieval and document summarization, aiming to balance relevance and diversity. It attempts to avoid redundancy of results while returning the most relevant results. Therefore, the archive keyword library constructed in the embodiments of the present invention can make the extraction of keywords more accurate. While considering the relevance of keywords to the topic, so that the keywords can accurately describe the topic, it is also necessary to consider the difference between the newly selected keywords and the already selected keywords, making the keyword list diverse and avoiding repetition. According to all the keywords corresponding to all clustering topics, an archive keyword library is constructed.

[0044] In a possible implementation, for any piece of archival data, based on the archival keyword library, determining at least one target keyword corresponding to the archival data includes: For any piece of archival data to be processed, obtain the weighted vocabulary graph corresponding to the archival data; Based on the weighted vocabulary graph corresponding to the archival data and the archival keyword library, use the TextRank algorithm to obtain the score corresponding to each vocabulary in the archival data to be processed; Match the keywords to be extracted in the archival data with the archival keyword library. For the vocabulary that appears in both, assign a greater initial weight to these vocabulary, and at the same time adjust the weight of the edges in the graph, and thereby reflect that the co-occurrence relationship between keywords is more important than other vocabulary in agricultural literature. If it is an edge where two keywords co-occur, a weight two levels higher can be assigned, which can improve the accuracy of the importance evaluation of keywords.

[0045] The TextRank algorithm is a graph model-based ranking algorithm for text processing. It uses the mutual relationship (co-occurrence relationship) between local vocabulary to rank the vocabulary or sentences in the text. The TextRank algorithm is inspired by the PageRank algorithm and can be used for tasks such as keyword extraction and text summarization.

[0046] According to the score corresponding to each vocabulary in the archival data to be processed, determine at least one target keyword corresponding to the archival data to be processed.

[0047] For example, the N vocabulary with the largest scores can be selected as at least one target keyword corresponding to the archival data to be processed. However, it is worth noting that other conditions can also be set to select target keywords.

[0048] Optionally, in the embodiments of the present invention, by extracting target keywords, an editing and research outline with distinct levels and clear logic can be automatically generated according to the relevance of the archival data content, such as including chapter titles and / or sub-titles.

[0049] Optionally, conditions can also be set by the user to extract both person and time information at the same time to improve the personalization degree of archival editing and research, and ultimately achieve a better archival editing and research effect.

[0050] In a possible implementation, based on at least one target keyword corresponding to the archival data, use an intelligent agent to extract the abstract corresponding to the archival data to obtain the archival directory corresponding to the archival data, including: Take at least one target keyword corresponding to the archival data and the archival data as the input of the first intelligent agent, and obtain the target keyword vector and the archival data text vector; wherein, the first intelligent agent includes the XLNet model; The XLNet model is a general pre-trained language model based on Transformer, jointly developed by researchers from Google and Carnegie Mellon University (CMU). The XLNet model adopts a hybrid approach of autoregressive and autoencoding during the pre-training stage, aiming to combine the advantages of autoencoding models such as BERT and traditional autoregressive models. In the embodiments of the present invention, the XLNet model is used to extract text vectors, providing a basis for subsequent abstract extraction.

[0051] The second agent is used to process the text vectors of the archive data to obtain the forward semantic feature text vectors and the reverse semantic feature text vectors; wherein, the second agent includes a BiGRU model. The text vectors of the archive data are passed to the BiGRU model. The BiGRU model is divided into two GRU units. One is used to extract the semantic features of the text vectors of the archive data in the forward direction, and the other is used to extract the semantic features of the text vectors of the archive data in the reverse direction. Then, training is carried out through multiple GRU hidden units to obtain the context features of the two text vectors of the archive data. The forward semantic feature text vectors and the reverse semantic feature text vectors are respectively the context features of the text vectors of the archive data after forward and reverse extraction and training by the BiGRU.

[0052] Optionally, the second agent may further include a multi-layer attention mechanism model. A multi-layer attention mechanism is introduced after the BiGRU to strengthen the model's attention to important parts of the text. Based on the previous keyword extraction and the semantic features of the text vectors of the archive data, an archive directory is generated by combining the attention mechanism and the coverage mechanism.

[0053] The third agent is used to process the target keyword vectors, the text vectors of the archive data, the forward semantic feature text vectors, and the reverse semantic feature text vectors to obtain the archive directory corresponding to the archive data; wherein, the third agent includes a decoder based on the attention mechanism.

[0054] Optionally, the decoder based on the attention mechanism may include a self-attention mechanism (Self-Attention), a first normalization layer (Layer Norm), an attention layer (Attention Model), a second normalization layer (Layer Norm), a feed-forward layer (Feed-forward), a third normalization layer (Layer Norm), a linear layer (Linear), and a softmax layer connected in sequence. Among them, the input and output of the self-attention mechanism are simultaneously used as the input of the first normalization layer, the input and output of the attention layer are simultaneously used as the input of the second normalization layer, and the input and output of the feed-forward layer are simultaneously used as the input of the third normalization layer.

[0055] In the embodiments of the present invention, by introducing a decoder based on the attention mechanism, an additional attention mechanism is introduced into the decoder to further optimize the processed text vectors. The attention mechanism can enhance the model's ability to focus on key information, capture the key content of the text more accurately, and avoid generating redundant or repetitive summary information. To integrate and optimize all the above information, residual connections and layer normalization are used to fuse and standardize various feature vectors. Residual connections help solve the problem of vanishing gradients in deep networks, while layer normalization ensures smoother data flow in the network, promoting the stability and efficiency of model training. Finally, the decoder converts the fused feature vectors into a predicted vocabulary probability distribution through a linear layer, and then converts these probability distributions into the final text summary output through a softmax layer. This process ensures that the generated summary not only compactly and accurately reflects the core content of the original text, but also is fluent and natural in language expression, meeting the objective requirements of summary generation.

[0056] In a possible implementation manner, before using the agent, it further includes: training the hyperparameters of the agent.

[0057] The training process may include: A1. Construct a main task node and multiple sub-task nodes for the agent, and allocate private memory and a common memory for each task node; wherein, each node is deployed on a different core or server; A2. Initialize the iteration counter t = 1, and initialize the scale matrix = I; where I represents the identity scale matrix; A3. Obtain the error function value corresponding to the agent, and put this error function value into the common memory for each task node to call; wherein, the error function value is obtained through the agent on any task node, represents the hyperparameters of the agent during training (such as the connection weights between network layers); t A4. Obtain the gradient matrix of the agent, and put the scale matrix and the gradient matrix into the common memory for each task node to call; wherein, each column in the gradient matrix is the gradient of a hyperparameter in the agent; wherein, represents the gradient operator, is an order square matrix, M represents the total number of dimensions of the hyperparameters of the agent; M ​A5. Each task node runs in parallel, and calls the scale matrix from the common memory for each task node and the gradient matrix , and determines the search direction of the agent on the m th task node as the product-accumulation data of the th row data in the scale matrix m and the th column data in the gradient matrix m ; where m = 1, 2, …, M, and the search direction is stored in the private memory corresponding to the m th task node; A6. Retrieve the search direction from the private memory for each task node , and according to the search direction and use the golden section method to determine the update step size corresponding to the m th hyperparameter in the agent model on this task node , and store the update step size in the private memory corresponding to the m th task node; Optionally, according to the search direction and use the golden section method to determine the update step size corresponding to the m th hyperparameter in the agent model on this task node , including: Construct a search matrix with only one row using the search direction , where the mth element in this search matrix is the search direction D , and the other elements are 0. The total number of elements in the search matrix is the same as the total number of hyperparameters of the agent model; Construct a step size matrix with only one column , where the mth element in this step size matrix is the update step size to be determined , and the other elements are 0. The total number of elements in the step size matrix is the same as the total number of hyperparameters of the agent model; Construct a solution condition D according to the search matrix D , and solve this solution condition to determine the update step size ; A7. Retrieve the update step size from the private memory for each task node , and update the th hyperparameter in the agent model on the task node according to the update step size ; According to the m th hyperparameter of the agent model on the m th task nodem An updated hyperparameter to obtain the m th gradient component; Optionally, according to the update step update the m th hyperparameter in the agent model on the task node to:

[0058] where, represents the m th hyperparameter in the agent model, represents the updated ; A8. Each task node runs in parallel. According to the m th gradient component, re-determine the gradient matrix m corresponding to the th task node, and update the scale matrix corresponding to the m th task node according to this gradient matrix ; Optionally, according to the m th gradient component, re-determine the gradient matrix m corresponding to the th task node, and update the scale matrix corresponding to the m th task node according to this gradient matrix , including: Based on the gradient matrix , use the th gradient component to update the m th element in the gradient matrix m corresponding to the th task node to obtain the gradient matrix corresponding to the m th task node; According to the gradient matrix m corresponding to the th task node, update the scale matrix

[0059] where, represents the updated , W represents the weight difference matrix, W = , represents the weight vector of the agent model updated according to the update step , represents the weight vector of the agent model before update,T denotes transpose, Z = , where Z represents the gradient difference matrix; A9. Output the scale matrix through each task node in the m column elements and the m th hyperparameter corresponding to the agent model to the main task node; A10. Through the main task node, m hyperparameters form an updated hyperparameter vector, and determine whether the error function value corresponding to the updated hyperparameter vector is less than a preset threshold. If so, the training end requirement is met and the training is completed. Otherwise, go to step A11; A11. Determine whether the count value of the iteration counter t is greater than the preset maximum number of training times. If so, the training end requirement is met and the training is completed; otherwise, increment the count value of the iteration counter t by one, and use the column elements in the scale matrix output by each task node to form a scale matrix m , use the updated hyperparameter vector as the hyperparameters of the agent model on all task nodes, and return to step A3.

[0060] The training algorithm provided by the embodiments of the present invention enables multiple parties to collaborate and update, can effectively improve the algorithm execution speed, and at the same time enables devices or processors with relatively weak computing power to also perform collaborative training and improve the hyperparameter optimization ability.

[0061] In a possible implementation manner, intelligent compilation and research of archival data is performed according to at least one target keyword corresponding to the archival data and the archival directory, including: According to at least one target keyword corresponding to the archival data, use the N keywords with the highest scores as the title corresponding to the archival data; Use the title corresponding to the archival data and the archival directory corresponding to the archival data together as the personalized compilation and research outline corresponding to the archival data, and complete the archival intelligent compilation and research based on the agent.

[0062] Figure 2 As shown, the present invention provides an archival intelligent compilation and research system based on an agent, including: a data preprocessing module 201, a vector extraction module 202, a keyword library construction module 203, a keyword extraction module 204, an abstract extraction module 205, and an intelligent compilation and research module 206; The data preprocessing module 201 is used to retrieve archival data from a specified data source or a specified database, and preprocess the archival data to obtain the preprocessed archival data; The vector extraction module 202 is used to perform streamlining processing on the preprocessed archive data to obtain the streamlined archive data, and extract the embedding vectors in the streamlined archive data; The keyword library construction module 203 is used to perform keyword extraction operations on the embedding vectors in all archive data, and then construct the extracted keywords into an archive keyword library; The keyword extraction module 204 is used to determine at least one target keyword corresponding to the archive data based on the archive keyword library for any piece of archive data; The abstract extraction module 205 is used to extract the abstract corresponding to the archive data by using an intelligent agent based on at least one target keyword corresponding to the archive data, and obtain the archive directory corresponding to the archive data; The intelligent compilation and research module 206 is used to perform intelligent compilation and research on the archive data according to at least one target keyword corresponding to the archive data and the archive directory, and complete the intelligent archive compilation and research based on the intelligent agent.

[0063] The archive intelligent compilation and research system based on an intelligent agent provided by an embodiment of the present invention can execute the above method technical solution, and its principle and beneficial effects are similar, so details are not described here again.

[0064] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0065] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the specified functions in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0066] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more processes and / or blocks Figure 1 in one or more processes and / or blocks Figure 1 specified in the block or blocks.

[0067] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes and / or blocks Figure 1 in one or more processes and / or blocks Figure 1 specified in the block or blocks.

[0068] Those of ordinary skill in the art can understand that all or part of the steps in implementing the above facts and methods can be completed by instructing relevant hardware through a program. The program involved or the said program can be stored in a computer-readable storage medium. When the program is executed, it includes the following steps: At this time, the corresponding method steps are introduced. The storage medium can be ROM / RAM, magnetic disk, optical disc, etc.

[0069] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. An agent-based intelligent archives compilation and research method, characterized in that: include: Retrieving archival data from a designated data source or a designated database, and preprocessing the archival data to obtain the preprocessed archival data; Simplifying the preprocessed archive data to obtain simplified archive data, and extracting embedded vectors from the simplified archive data; After extracting keywords from the embedding vectors in all archive data, the extracted keywords are constructed into an archive keyword library; For any piece of archive data, based on the archive keyword library, determine at least one target keyword corresponding to the archive data; Based on at least one target keyword corresponding to the archive data, using an intelligent agent to extract a summary corresponding to the archive data to obtain an archive directory corresponding to the archive data; According to at least one target keyword and archive directory corresponding to the archive data, the archive data is intelligently compiled and researched to complete the intelligent compilation and research of archives based on intelligent agents.

2. The agent-based intelligent archive compilation and research method according to claim 1 is characterized in that: The pre-processed archive data is streamlined to obtain streamlined archive data, including: Removing images and table data from the pre-processed archive data to obtain the archive data after initial processing; Performing text cleaning on the archive data after the initial processing to obtain archive data after the text cleaning; The file data after the text cleaning process is segmented and stop words are removed, and the file data after the word segmentation and stop words are used as the file data after the streamlining process.

3. The agent-based intelligent archive compilation and research method according to claim 1 is characterized in that: Extract embedding vectors from the condensed archive data, including: The streamlined archival data is input into the BERT language model, so as to process the streamlined archival data through the word embedding function of the BERT language model and obtain the embedding vector in the streamlined archival data.

4. The agent-based intelligent archives compilation and research method according to claim 3 is characterized in that: After extracting keywords from the embedding vectors in all archive data, the extracted keywords are constructed into an archive keyword library, including: The UMAP algorithm is used to reduce the dimension of the embedded vector in the simplified archive data to obtain the embedded vector after dimensionality reduction. Based on the embedding vector after dimensionality reduction, the HDBSCAN algorithm is used for clustering to obtain multiple clustering topics; For any cluster topic, the c-TD-IDF algorithm is used to obtain the c-TD-IDF value corresponding to the vocabulary in each cluster topic; For any cluster topic, the word with the largest c-TD-IDF value is used as the keyword corresponding to the cluster topic; For any cluster topic, calculate the MMR values ​​of other words except the keywords in turn, and select the word with the largest MMR value as the new keyword. Repeat this step until the keyword extraction end condition is met, and obtain all the keywords corresponding to each cluster topic; Construct an archive keyword library based on all keywords corresponding to all clustered topics.

5. The agent-based intelligent archive compilation and research method according to claim 4 is characterized in that: For any piece of archive data, based on the archive keyword library, at least one target keyword corresponding to the archive data is determined, including: For any archive data to be processed, obtain a weighted vocabulary graph corresponding to the archive data; Based on the weighted vocabulary graph corresponding to the archival data and the archival keyword library, the TextRank algorithm is used to obtain the score corresponding to each word in the archival data to be processed; At least one target keyword corresponding to the archive data to be processed is determined according to the score corresponding to each word in the archive data to be processed.

6. The agent-based intelligent archive compilation and research method according to claim 5 is characterized in that: Based on at least one target keyword corresponding to the archive data, an agent is used to extract a summary corresponding to the archive data to obtain an archive directory corresponding to the archive data, including: At least one target keyword corresponding to the archive data and the archive data are used as inputs of a first agent to obtain a target keyword vector and an archive data text vector; wherein the first agent includes an XLNet model; The second agent is used to process the text vector of the archive data to obtain a positive semantic feature text vector and a reverse semantic feature text vector; wherein the second agent includes a BiGRU model; A third intelligent agent is used to process the target keyword vector, the archive data text vector, the positive semantic feature text vector and the reverse semantic feature text vector to obtain the archive directory corresponding to the archive data; wherein the third intelligent agent includes a decoder based on the attention mechanism.

7. The agent-based intelligent archives compilation and research method according to claim 6 is characterized in that: Before the intelligent agent is used, the method further includes: training the hyper parameters of the intelligent agent.

8. The agent-based intelligent archive compilation and research method according to claim 6 is characterized in that: According to at least one target keyword and archive directory corresponding to the archive data, the archive data is intelligently compiled and researched, including: According to at least one target keyword corresponding to the archive data, N keywords with the highest scores are used as titles corresponding to the archive data; The title corresponding to the archival data and the archival directory corresponding to the archival data are used together as the personalized editing and research outline corresponding to the archival data to complete the intelligent editing and research of the archives based on the intelligent agent.

9. An intelligent archives editing and research system based on an agent, characterized in that: include: Data preprocessing module, vector extraction module, keyword library construction module, keyword extraction module, abstract extraction module and intelligent editing and research module; The data preprocessing module is used to retrieve the archive data from a specified data source or a specified database, and preprocess the archive data to obtain the preprocessed archive data; The vector extraction module is used to simplify the preprocessed archive data to obtain simplified archive data and extract embedded vectors from the simplified archive data; The keyword library construction module is used to extract keywords from the embedded vectors in all archive data and then construct the extracted keywords into an archive keyword library; The keyword extraction module is used to determine at least one target keyword corresponding to any archive data based on the archive keyword library; The summary extraction module is used to extract the summary corresponding to the archive data using an agent based on at least one target keyword corresponding to the archive data to obtain an archive directory corresponding to the archive data; The intelligent editing and research module is used to perform intelligent editing and research on the archive data according to at least one target keyword and the archive directory corresponding to the archive data, and complete intelligent editing and research of the archive based on the intelligent agent.

Citation Information

Patent Citations

  • Automatic archive editing method

    CN104361111A

  • FPGA system and implementation method based on on-line training neural network of quasi-newton method

    CN106528357A

  • Establishing method and system of search engine

    CN107818130A

  • GENERATIVE AUTOMATIC abstracting METHOD BASED ON BERT AND EXTERNAL KNOWLES

    CN114398478A

  • Electronic archive retrieval method and system based on large language model

    CN118643148A

Cited By

  • Archive data processing auxiliary editing and research system

    CN122412368A