A method for digital storage management of files
Through technical means such as optical character recognition, natural language processing, clustering analysis and knowledge graph construction, the archive content is transformed into a multimodal knowledge product with clear structure and easy to understand, solving the problem that the archive content is difficult to efficiently transform into new knowledge products and a single form of dissemination, and realizing the intelligent management and efficient dissemination of archive resources.
Patent Information
- Application Number
- CN202510618931.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2045-05-14
AI Technical Summary
The digital storage method of archive content lacks a systematic processing and transformation mechanism, resulting in insufficient social communication, low user understanding and participation, difficult to efficiently convert into new knowledge products, and difficult to break through the traditional framework of communication.
Key information units are extracted through optical character recognition technology combined with natural language processing algorithms, cluster analysis algorithms are used for topic aggregation, and structured knowledge networks are generated using knowledge graph construction technology, and knowledge nodes are retelled through text generation algorithms. Graphs, audio or video forms are generated by multimodal transformation technology, user behavior analysis is used to optimize the dissemination effect, and finally push the knowledge product to the adaptive scenario through the content distribution network, and continuously optimize the processing technology through machine learning algorithms.
It realizes the intelligent processing of archive content and the effective dissemination of knowledge, improves the utilization efficiency of archive resources and the effectiveness of knowledge dissemination, and enhances the social influence of archives.
Smart Images

Figure CN120124594B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of archive information processing, and in particular to a method for digital storage and management of archives. Background Art
[0002] As a vital carrier of human history and culture, archives not only carry information from the past but also hold immense potential for driving social progress. Their research and application are particularly crucial in the knowledge economy. Through the in-depth development and innovative transformation of archival content, new knowledge products and services can be generated, significantly enhancing their social value and influence. However, research and practice in this field still face numerous bottlenecks that urgently need to be overcome.
[0003] Currently, the application of archival content remains largely at the stage of simple digital storage and display, with relatively simplistic approaches and a lack of systematic processing and transformation mechanisms. Many solutions overly focus on archival preservation while neglecting the revitalization of its content, resulting in insufficient social dissemination of archives and low user understanding and engagement. This situation limits the potential for archives to transform from static resources into dynamic knowledge, preventing their full potential from being realized. Against this backdrop, the core challenges facing archival content transformation are becoming increasingly prominent. First, insufficient processing technology hinders the effective distillation of archival content into structured, reproducible knowledge units. Second, the design of transformation formats lacks universality and comprehensibility, making it difficult to meet the needs of diverse audiences, resulting in poor dissemination effectiveness. Finally, the connection between technical means and social impact objectives is insufficient, making it difficult to precisely integrate processed content into practical application scenarios. Unaddressed these technical factors have led to unique challenges, such as the difficulty in efficiently transforming archival content into new knowledge products and the difficulty in breaking through traditional dissemination frameworks. Summary of the Invention
[0004] The purpose of this invention is to provide a method for digital storage and management of archives. Through innovative processing technology, the archive contents are transformed into new knowledge products with clear structure and easy to understand, and a form suitable for wide dissemination is designed to effectively expand the social influence of the archives.
[0005] To achieve the above-mentioned purpose, the present invention provides the following technical solution: a method for digital storage and management of archives, comprising:
[0006] S1. Obtain the original data of the archive content, and extract key information units from the text using optical character recognition technology combined with natural language processing algorithms to obtain preliminarily classified semantic fragments;
[0007] S2. Use a clustering analysis algorithm to perform thematic aggregation on the content, determine the thematic category to which each segment belongs, and obtain a thematic content collection;
[0008] S3. Use knowledge graph construction technology to associate and map the logical relationships between semantic fragments, and generate a knowledge network with clear structure;
[0009] S4. Apply text generation algorithms to restate knowledge nodes in natural language to obtain more understandable expression texts; by analyzing the node conversion process, determine the semantic consistency during the conversion process. If the evaluation result reaches the preset threshold, output the final expression text;
[0010] S5. Combine multimodal conversion technology to convert text content into forms such as charts, audio, or video, and generate a diverse set of conversion forms; if the data is complete after the conversion technology processing, integrate the chart form, audio form, and video form, and sort them according to the thinking chain using logical organization methods based on the diverse set to obtain an ordered conversion result. Through information comparison technology, judge whether the ordered conversion result is consistent with the concise expression to obtain the final conversion set;
[0011] S6. Adopt user behavior analysis algorithms to judge the dissemination effects of different conversion forms in the target group and obtain an optimized combination of dissemination forms;
[0012] S7. Use content delivery network technology to push the processed knowledge products to the adapted application scenarios and generate the application output of dynamic knowledge;
[0013] S8. Collect user feedback data, iteratively adjust the processing technology through machine learning algorithms to obtain updated technical parameters; use statistical tools to calculate the distribution characteristics of the feedback data, use the preset threshold to judge the parameter change trend, and determine the technology iteration direction; if the parameter change trend exceeds the preset threshold, use the gradient descent algorithm to adjust the technical parameters to obtain updated technical parameters;
[0014] S9. For the updated technical parameters, reprocess the original data of the file content, and repeatedly execute clustering analysis and knowledge graph construction to generate a continuously optimized sequence of knowledge products.
[0015] Preferably, the S1 includes:
[0016] Obtain the original data, process the content of the archives through optical character recognition technology, extract the text information, and obtain the first text set. For the first text set, use natural language processing algorithms to analyze the text information, extract the key information, and obtain the second information set. Extract the information units from the second information set, and conduct a preliminary classification through the preset classification rules to obtain the third classification set. According to the preliminary classification in the third classification set, use clustering algorithms to analyze the semantic fragments and obtain the fourth semantic set. Through the fourth semantic set, judge the relevance between the semantic fragments. If the relevance is higher than the preset threshold, merge the relevant semantic fragments to obtain the fifth merged set. Obtain the fifth merged set, and for the merged semantic fragments, use statistical methods to determine the distribution characteristics of the key information and obtain the sixth feature set. According to the sixth feature set, judge the integrity of the semantic fragments by comparing with the preset semantic templates to obtain the seventh verification set.
[0017] Preferably, S2 includes:
[0018] Through the acquisition of semantic fragments, apply word segmentation technology to conduct a preliminary classification of the text to obtain the initial grouped data. Use clustering analysis algorithms to perform topic aggregation on the initial grouped data to determine the topic categories of each group. According to the results of the topic aggregation, generate a content set containing the topic categories. If the number of topic categories in the content set exceeds the preset threshold, screen the categories through the judgment basis to obtain the refined set. For the refined set, use statistical analysis methods to calculate the distribution characteristics of each topic category to obtain the distribution data. Through the distribution data, determine the core characteristics of the topic categories and generate the analysis results. According to the analysis results, use visualization tools to present the distribution of the topic categories to obtain the final output.
[0019] Preferably, S3 includes:
[0020] Obtain the topic content from the content set, use word segmentation technology to extract semantic fragments to obtain the preliminary semantic units. For the semantic fragments, use dependency syntactic analysis to determine the logical relationships and obtain the dependency structure between the fragments. Obtain the association mapping from the dependency structure, and use knowledge graph construction technology to generate the initial network structure and determine the nodes and edges. If the logical relationships between the nodes meet the preset threshold, adjust the network structure through clustering algorithms to obtain the optimized knowledge network. According to the optimized knowledge network, perform iterative updates for the fragment associations, obtain the semantic expansion after mapping generation, judge the clarity of the structure through the semantic expansion, use topological sorting to generate the final knowledge network, and extract the complete representation of the topic content from the final knowledge network to obtain the structured semantic expression.
[0021] Preferably, S4 includes:
[0022] Through the pre-established knowledge network, node data with clear structure is obtained, and the text generation algorithm is used to perform preliminary conversion on the node content to obtain the initial text. Natural language features are extracted from the initial text, and the content is restated based on the features to generate an intermediate expression text. If the intermediate expression text meets the requirements of the thinking chain, the sentence structure is adjusted through the application of the algorithm to obtain a text with enhanced logical expression. The node conversion process is analyzed according to the enhanced text, the semantic consistency in the conversion process is determined, and the optimized text is generated. The optimized text is evaluated. If the evaluation result reaches the preset threshold, the final expression text is output. By comparing the final expression text with the knowledge network, it is determined whether there are unconverted nodes. If so, the conversion process is repeated to obtain the final expression text of all nodes, which are integrated by splicing to obtain a complete and clearly structured output text.
[0023] Preferably, the S5 includes:
[0024] Obtain the original expression text, extract concise text through text parsing technology, determine the core expression content, use multimodal conversion technology to generate a chart form for the concise text, obtain a chart data set, extract features from the chart data set, use audio synthesis technology to convert it into audio form, obtain audio data, use video rendering technology to combine the chart data set with the audio data to generate a video form, obtain video data, if the data is complete after processing by the conversion technology, then integrate the chart form, audio form and video form to generate a diversified set, based on the diversified set, use a logical organization method to sort according to the thinking chain, obtain an ordered conversion result, use information comparison technology to determine that the ordered conversion result is consistent with the concise expression, and obtain the final conversion set.
[0025] Preferably, the S6 includes:
[0026] Through the behavior log, the behavioral data corresponding to the user behavior in the target group is obtained to obtain the initial data set. The clustering algorithm is used to extract the correlation characteristics between the conversion form and the user behavior from the initial data set to obtain the feature matrix. Based on the feature matrix, the communication effect of each conversion form in the target group is judged to obtain the effect score. If the effect score is lower than the preset threshold, the corresponding conversion form is eliminated from the form set to obtain the filtered form set. The optimization algorithm is used to generate the optimized combination based on the filtered form set to obtain the candidate combination set. Based on the candidate combination set, the behavioral data of each combination in the target group is obtained to judge the communication effect to obtain the final optimized combination. The key conversion form is extracted from the final optimized combination to generate a communication execution plan.
[0027] Preferably, the S7 includes:
[0028] Obtain the original knowledge data through a content distribution network, use processing technologies to perform structured conversion on the data to obtain processed knowledge products, extract key features from the processed knowledge products, combine network technologies to determine the push requirements for the adapted scenarios, judge the push priority. If the push priority is higher than the preset threshold, transmit the processed knowledge products to the adapted scenarios through push technologies to generate initial output data. According to the initial output data, use optimization propagation technologies to adjust the propagation form to obtain a real-time updated version of the dynamic knowledge. Analyze the real-time updated version of the dynamic knowledge through technology combination, obtain feedback data from the adapted scenarios, determine the adjustment direction of the application output. For the feedback data, use machine learning algorithms to iteratively optimize the dynamic knowledge to obtain the final application output, and push the final application output to the target scenarios through the content distribution network to complete the application deployment of the dynamic knowledge.
[0029] Preferably, S8 includes:
[0030] Obtain user feedback data through an external interface, use statistical tools to calculate the distribution characteristics of the feedback data to obtain a preliminary analysis result, extract the dynamic knowledge content according to the preliminary analysis result, match the knowledge application scenarios with preset rules to determine the applicable knowledge scope, screen the feedback data subset through the applicable knowledge scope, use machine learning algorithms to train the model to obtain a feedback data classifier, adjust the processing technology parameters according to the output result of the classifier, use a preset threshold to judge the parameter change trend to determine the technology iteration direction, update the dynamic knowledge base through the technology iteration direction, analyze the consistency of the knowledge base using clustering algorithms to obtain an optimized knowledge structure, reprocess the feedback data according to the optimized knowledge structure. If the parameter change trend exceeds the preset threshold, use the gradient descent algorithm to adjust the technology parameters to obtain updated technology parameters, generate a processing technology configuration through the updated technology parameters, and use a consistency verification tool to detect the configuration integrity to obtain the final technology output.
[0031] Preferably, S9 includes:
[0032] Obtain the original data and perform preprocessing according to the technology parameters to obtain a standardized data set. Process the standardized data set through clustering analysis to determine the preliminary clustering result. Construct a knowledge graph for the preliminary clustering result to obtain a structured knowledge network. Extract features from the structured knowledge network to generate an initial product sequence. Use optimization iteration to adjust the initial product sequence to obtain an optimized product sequence. If the optimized product sequence meets the preset threshold, output the final knowledge product sequence; if not, loop through clustering analysis and knowledge graph construction, continuously optimize and update the technology parameters to obtain a knowledge product sequence adapted to the changes.
[0033] As can be seen from the above technical solutions, the present invention has the following beneficial effects:
[0034] The method for digital storage management of archives extracts key information from the original archive data through optical character recognition and natural language processing technologies, and uses clustering analysis algorithms for topic aggregation. Subsequently, the present invention applies knowledge graph construction technologies to generate a structured knowledge network, and restates knowledge nodes through text generation algorithms to improve comprehensibility. The present invention also adopts multimodal conversion technologies to convert content into forms such as charts, audio, or video, and optimizes the dissemination effect according to user behavior analysis. Finally, the present invention uses a content delivery network to push knowledge products to adapted scenarios, and continuously optimizes processing technologies through machine learning algorithms. This method realizes the intelligent processing of archive content, the effective dissemination of knowledge, and continuous optimization, greatly improving the utilization efficiency of archive resources and the knowledge dissemination effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0037] As Figure 1 shown, the present invention provides a technical solution: a method for digital storage management of archives, including:
[0038] S1. Obtain the original data of the archive content, and extract key information units from the text through optical character recognition technology combined with natural language processing algorithms to obtain preliminarily classified semantic segments;
[0039] S2. Use clustering analysis algorithms to perform topic aggregation on the content, judge the topic category to which each segment belongs, and obtain a themed content set;
[0040] S3. Use knowledge graph construction technologies to perform associative mapping on the logical relationships between semantic segments to generate a knowledge network with clear structure;
[0041] S4. Apply text generation algorithms to perform natural language restatement on knowledge nodes to obtain a more easily understandable expression text; by analyzing the node conversion process, determine the semantic consistency during the conversion process. If the evaluation result reaches a preset threshold, output the final expression text;
[0042] S5. Combine multi-modal conversion technologies to convert text content into chart, audio or video forms, generating a diverse set of conversion forms. If the data is complete after the conversion technology processing, integrate the chart form, audio form and video form, and sort them according to the thinking chain using a logical organization method based on the diverse set to obtain an ordered conversion result. Then, through information comparison technology, judge whether the ordered conversion result is consistent with the concise expression to obtain the final conversion set.
[0043] S6. Use user behavior analysis algorithms to judge the dissemination effects of different conversion forms among the target group and obtain an optimized combination of dissemination forms.
[0044] S7. Utilize content delivery network technology to push the processed knowledge products to the adapted application scenarios, generating the application output of dynamic knowledge.
[0045] S8. Collect user feedback data, iteratively adjust the processing technology through machine learning algorithms to obtain updated technical parameters. Use statistical tools to calculate the distribution characteristics of the feedback data, use preset thresholds to judge the parameter change trend, and determine the technical iteration direction. If the parameter change trend exceeds the preset threshold, use the gradient descent algorithm to adjust the technical parameters to obtain updated technical parameters.
[0046] S9. For the updated technical parameters, reprocess the original data of the file content, and repeatedly execute clustering analysis and knowledge graph construction to generate a continuously optimized sequence of knowledge products.
[0047] The core of this implementation lies in transforming the original archival content from static text documents into dynamic, structured, and disseminable knowledge products. First, the text information in scanned images is recognized through a high-precision OCR module, and the recognition results are further analyzed for semantic structure by a natural language processing module, including entity recognition (such as names of people, organizations, and places), keyword extraction (such as event tags, time nodes), and dependency parsing. The extracted semantic fragments are subject to topic recognition and classification through topic modeling algorithms to achieve content aggregation. Then, based on these semantic fragments, a knowledge graph is constructed, and a graph database such as Neo4j is used to organize entity nodes and relationship edges, making the logical connections between information clear and having reasoning capabilities. Based on the knowledge graph, a pre-trained language model (such as T5 or ChatGLM) is used to automatically generate natural language restatements for each knowledge node, enhancing the comprehensibility and affinity of its language expression. After that, through graphic and text visualization components (such as ECharts, Plotly), the data is converted into graphical displays, and a TTS engine generates voice content, and videos are automatically edited in combination with scripts and visual materials. In the dissemination stage, the interaction behaviors of users with different forms of content (such as click frequency, reading depth, forwarding rate) are analyzed and input into the recommendation system to form a user portrait and content adaptation matrix for precise distribution. Finally, through user feedback data (including behavioral data and subjective evaluations), the algorithm parameters of core modules such as text generation, graph construction, and push strategies are continuously optimized to ensure the continuous evolution of content in subsequent iterations. This method realizes the transformation from "archival information" to "knowledge products" by performing semantic understanding, structural modeling, and content reconstruction on archival data, significantly improving the utilization efficiency of archives. The multi-modal expression and dissemination mechanism not only enriches the content form but also enhances the accessibility and influence of information. The iterative optimization ability ensures the continuous evolution of the system, having significant advantages in aspects such as service accuracy, content timeliness, and dissemination effect. It is applicable to information-intensive scenarios, such as historical archive management, enterprise knowledge base construction, government affairs disclosure, and educational resource collation, etc.
[0048] Taking a local archives as an example, its collection includes a large number of paper historical documents from the early 20th century to the present. After introducing the digital management method of this invention, first, all paper archives are transformed into structured text data through high-resolution scanning and OCR technology; then, the NLP model combined with the historical term library extracts core information such as people, time, and events to construct a local history knowledge graph. Based on the knowledge graph, graphic and text explanation content is generated and accompanied by AI voice broadcasts to form a "collection of local history story videos". These videos are released through multiple channels such as government affairs official accounts and community science popularization exhibition board screens. Finally, combined with feedback such as citizens' likes and viewing times, the system continuously optimizes the content topics and presentation forms to achieve the vivid dissemination of archival culture, greatly enhancing citizens' understanding and interest in local history and culture.
[0049] S1 includes obtaining the original data, processing the content of the archives through optical character recognition technology, extracting text information to obtain the first text set, analyzing the text information for the first text set using natural language processing algorithms, extracting key information to obtain the second information set, extracting information units from the second information set, performing preliminary classification through preset classification rules to obtain the third classification set, analyzing semantic fragments using a clustering algorithm based on the preliminary classification in the third classification set to obtain the fourth semantic set, judging the relevance between semantic fragments through the fourth semantic set. If the relevance is higher than the preset threshold, merging relevant semantic fragments to obtain the fifth merged set, obtaining the fifth merged set, determining the distribution characteristics of key information for the merged semantic fragments using statistical methods to obtain the sixth characteristic set, and judging the integrity of semantic fragments by comparing with the preset semantic template based on the sixth characteristic set to obtain the seventh verification set.
[0050] The overall workflow of this embodiment is achieved sequentially through nine steps, from the extraction of text from the original files to the verification of the semantic structure. Each step relies on specific algorithm modules for data processing and information recognition. In the first step, the system scans and recognizes the file images through an optical character recognition engine (such as Tesseract 5.0), extracts the text content therein, and forms a text data set, which is called the first text set. The text information in each record will serve as the basic data for subsequent processing. In the second step, a natural language processing model (such as BERT or LTP) is used to analyze these texts, including named entity recognition (such as person names, place names, organization names), event extraction, and time expression recognition, etc., so as to extract fragments with key information significance and form the second information set. In the third step, the smallest granularity semantic units are further extracted from the above extracted information. For example, in the sentence "Zhang San moved to Shanghai in 1947", "Zhang San", "1947", "moved to", and "Shanghai" are each regarded as semantic units. These units are input into a classification function and mapped to specific categories according to the set rules, such as "person name category", "time category", "action category", etc., to obtain a preliminary classification result, that is, the third classification set. In the fourth step, based on the preliminary classification results, similarity clustering analysis is performed on the semantic fragments. At this time, the system will convert each fragment into a semantic vector, and the similarity between the vectors is measured by the cosine value of their included angle. If the similarity between two fragments, that is, the cosine value of the included angle between their semantic vectors, is greater than the set threshold (such as 0.85), then these two fragments are determined to be semantically similar. In the fifth step, the system combines the above similar semantic fragments into a whole fragment to form the fifth combined set. This operation aims to integrate semantically repeated or highly relevant information units, reduce redundancy, and improve the conciseness of the content. In the sixth step, the system performs word frequency statistics and co-occurrence analysis on the combined semantic fragments to identify which keywords or key expressions appear with a high frequency, and then forms the sixth feature set. This set reflects the central elements of the information content. In the seventh step, the above feature set is compared with a preset semantic template. This template is generally composed of three parts: "theme - action - result", such as "someone - did something - resulting in something". If the combined fragment can match this template structure, it is regarded as semantically complete and enters the seventh verification set; otherwise, it is regarded as an incomplete fragment expression and needs to be completed or removed in subsequent processing.
[0051] Through this multi-level processing flow, this embodiment significantly improves the accuracy and depth of the structured processing of file data. By adopting the method of quantifying similarity and semantic template comparison, the stability and robustness of information extraction are effectively improved. Especially through the combination of semantic fragments and feature focusing, it can avoid the dilution or omission of key information in subsequent processing, ensure the integrity and clarity of the construction basis of the knowledge graph, and improve the quality of subsequent content generation and dissemination.
[0052] In the project of promoting the digitization of local archives of a certain period by the archives bureau of a certain province, the original archives are a large number of scanned images, and the text content includes historical documents such as civil affairs official documents, land deeds, and family genealogies. After OCR extraction by the system, a rule library is used to classify information such as "household head", "year of birth", and "address" into "identity category", "time category", and "geographical category", and a classification set is constructed. After clustering and merging multiple segments related to the same household registration by similarity, the system uses a template to identify the complete migration record of a certain family. This information is automatically labeled, mapped, and a "Household Registration Migration Map of a Certain Period" is generated, which is widely used in public education and cultural relic exhibition scenarios, greatly improving the dissemination efficiency of historical archives and social service capabilities.
[0053] S2 includes obtaining semantic segments, applying word segmentation technology to preliminarily classify the text to obtain initial grouped data, using a clustering analysis algorithm to perform topic aggregation on the initial grouped data to determine the topic categories of each group, generating a content set containing topic categories according to the results of topic aggregation, if the number of topic categories in the content set exceeds a preset threshold, screening the categories through judgment criteria to obtain a refined set, for the refined set, using statistical analysis methods to calculate the distribution characteristics of each topic category to obtain distribution data, determining the core characteristics of the topic categories through the distribution data to generate an analysis result, and presenting the topic category distribution using a visualization tool according to the analysis result to obtain the final output.
[0054] This implementation mode is mainly used for theme aggregation and structural analysis of semantic fragments in archival texts, and finally presents the distribution characteristics of each theme category in a graphical way. First, the system applies word segmentation technology to the acquired semantic fragments. For example, it uses Jieba word segmentation or LTP Chinese word segmenter to divide the text content into multiple keywords or word units. After word segmentation, the system makes a preliminary classification of these words based on the context relationship and semantic connection of the words, thus forming an initial grouped data set. Subsequently, the system uses clustering analysis method to conduct theme aggregation on the above initial grouped data. Taking the K-means clustering algorithm as an example, the system first converts each semantic fragment into a vector form, and these vectors can be generated by TF-IDF or Word2Vec model to represent the semantic features of the text. Then, the system calculates the distance between each fragment and multiple preset theme center points. This distance is obtained by taking the difference between the vector coordinates of the semantic fragment and the coordinates of the theme center in each dimension, squaring them and summing them up, and finally taking the square root. All semantic fragments will be assigned to the clustering center with the closest distance to them. The goal of the clustering process is to minimize the sum of the distances from all semantic fragments to their respective center points. When the clustering is completed, the system will generate a theme content set containing all classification information according to the theme categories represented by each cluster. To avoid classification chaos caused by too many theme categories, the system will set a threshold for the maximum allowable number of themes, such as 20 categories. If the number of theme categories generated by the system exceeds this threshold, it will eliminate the categories with less influence or weak representativeness according to certain screening conditions, such as the number of text fragments contained in a certain category, the importance of keywords, or the semantic coverage of the theme, to obtain a more compact refined content set. Next, the system will conduct statistical analysis on the retained theme category set. Specific methods include calculating the frequency proportion of a certain keyword appearing in a specific theme. For example, if a keyword appears 50 times in the theme set and the total number of words in the theme set is 5000, then the frequency of this keyword is 1%. After statistically obtaining the keywords and their frequencies of all theme categories, the system can judge the core features of each theme based on this. Finally, the system inputs these statistical data into a visualization tool (such as ECharts or Tableau) to display the theme categories and their distribution characteristics in the form of bar charts, word clouds, heat maps, etc., so that users can intuitively understand the weights, importance and trend distributions of each theme in the entire batch of archival data.
[0055] This method can automatically identify and extract the thematic structure in the archival content, and generate structured classification results and visual diagrams. By introducing the classification screening and theme focusing mechanisms, the information redundancy is greatly reduced, making the final output results more representative and interpretable. Its statistical feature extraction and distribution display provide the basic conditions for further analysis of the digital content of archives (such as trend recognition and key point extraction), and lay a solid foundation for subsequent knowledge modeling and multimodal transformation.
[0056] Taking the digital archives of a university as an example, it stores a large number of previous years' teacher evaluation archival materials. To analyze the key research directions of teachers, the system introduces this implementation method. After text segmentation, it clusters according to research keywords, and themes such as "artificial intelligence", "educational technology", "computational linguistics", etc. quickly emerge. Since the initial number of themes reaches more than thirty categories, the system automatically filters out niche themes with fewer than a certain number of documents and only retains fifteen core research directions. Subsequently, through statistical analysis, the representative words and frequency data of each theme are obtained and displayed in the form of a heat map, enabling the department managers to intuitively grasp the research density distribution of each discipline and optimize the allocation of scientific research resources for the next year accordingly.
[0057] S3 includes obtaining the thematic content through the content set, using word segmentation technology to extract semantic fragments to obtain preliminary semantic units. For the semantic fragments, dependency syntactic analysis is used to determine the logical relationship and obtain the dependency structure between the fragments. An association mapping is obtained from the dependency structure, and knowledge graph construction technology is used to generate an initial network structure, determining nodes and edges. If the logical relationship between nodes meets the preset threshold, the network structure is adjusted through a clustering algorithm to obtain an optimized knowledge network. According to the optimized knowledge network, iterative updates are performed for fragment associations to obtain the semantic expansion after mapping generation. The clarity of the structure is judged through the semantic expansion, and topological sorting is used to generate the final knowledge network. The complete representation of the thematic content is extracted from the final knowledge network to obtain a structured semantic expression.
[0058] This implementation method realizes the construction of a knowledge graph for semantic fragments by combining natural language understanding, graph theory modeling, and optimization algorithms, and outputs a structured expression. First, the system extracts the topic content from the content set obtained in the previous stage. The text is segmented into semantic fragments using word segmentation technology. For example, the sentence "Teacher Wang teaches an artificial intelligence course" is split into several word elements such as "Teacher Wang", "teaches", "artificial intelligence", "course", etc., forming a preliminary set of semantic units. Then, the system applies dependency parsing technology (such as LTP, spaCy, etc.) to these semantic fragments to identify grammatical structures such as subject-predicate relationships, verb-object relationships, and modification relationships between words, and generates a dependency tree. Each dependency relationship in the dependency tree is modeled as an edge. For example, "Teacher Wang" is the subject of "teaches", and "artificial intelligence course" is the object of "teaches". These pieces of information together form the dependency structure of the semantic fragment. The dependency structure is transformed into an initial knowledge graph structure, where each node in the graph represents an entity or concept (such as "Teacher Wang", "artificial intelligence course"), and the edges between nodes represent semantic relationships (such as "teaches"). To determine whether a connection should be established between two nodes, the system calculates based on a logical relationship scoring function. For example, a comprehensive scoring function that combines syntactic path distance and semantic similarity can be used. If the score of two nodes is higher than a set threshold (such as 0.75), a connection is established. The scoring calculation process is as follows: the path distance score is 1 / path length. For example, the distance between "Teacher Wang" and "course" is two steps, and the path score is 0.5; the word vector similarity is calculated by dividing the dot product of the two word vectors by the product of their respective norms; the final score is the weighted average of the path score and the similarity score. Once the graph is initially constructed, the system optimizes the graph structure. If there are redundant connections or weak connection edges in the graph, the graph clustering algorithm (such as the Louvain community discovery algorithm) is called to re-partition the node structure and simplify the graph. Clustering can be merged according to the edge weights between nodes, making the relationships within the subgraphs tight and the relationships between subgraphs loose. Next, the system iteratively updates the graph. That is, when a new semantic fragment overlaps with a part of the original graph, it is extended and connected to the graph, and the relationship strength between it and the existing structure is calculated according to the rules. The extended nodes enhance the information density of the graph by merging similar semantics and introduce new topic paths. After completing the semantic extension, the system generates the final knowledge network based on the topological sorting method of the graph. Topological sorting will generate a structured semantic expression path layer by layer starting from nodes without predecessors according to the dependency relationships between nodes, ensuring that the information logic order is clear. The final output structured semantic expression consists of the knowledge expression chains represented by the paths in the graph structure and can be used as the knowledge input for subsequent natural language generation, multimodal conversion, and other processes.
[0059] This implementation mode realizes the logical modeling of the relationships between complex semantic segments by integrating dependency syntactic analysis and knowledge graph construction, ensuring the integrity, relevance, and scalability of the archival knowledge structure. Through the graph optimization and topological sorting mechanism, the generated semantic network is both clear and extensible, greatly enhancing the reusability and dissemination ability of the implicit knowledge in the archives, and is applicable to scenarios such as intelligent retrieval, automatic summarization, and question-answering systems.
[0060] Taking a local research center as an example, it processes a large number of red archival materials, which contain complex relationships among many historical figures, events, and organizations. After being processed by this implementation mode, the system conducts semantic modeling on sentences such as "the anti-tyrant struggle of the peasant association of a certain organization", and constructs knowledge graph paths such as "a certain organization" - "organization" - "peasant association", "peasant association" - "execute" - "anti-tyrant struggle". This graph can clearly display the historical event process through its topological structure, and is finally transformed into structured knowledge data for researchers to call in the special research system, and can also be exported for digital products such as historical education micro-videos and electronic exhibitions, greatly enhancing the value mining ability of archival materials and the public communication effect.
[0061] S4 includes obtaining well-structured node data through a pre-established knowledge network, initially transforming the node content using a text generation algorithm to obtain an initial text, extracting natural language features from the initial text, restating the content for the features to generate an intermediate expression text. If the intermediate expression text meets the requirements of the thought chain, the sentence pattern is adjusted through algorithm application to obtain a text with enhanced logical expression. Analyze the node conversion process based on the enhanced text to determine the semantic consistency during the conversion process, generate an optimized text, evaluate the optimized text. If the evaluation result reaches the preset threshold, output the final expression text. Compare the final expression text with the knowledge network to determine whether there are untransformed nodes. If so, repeat the conversion process to obtain the final expression texts of all nodes, and integrate them using the splicing method to obtain a complete and well-structured output text.
[0062] The core of this implementation method is to convert the node content in the knowledge graph into a logically coherent and semantically clear natural language expression through natural language generation algorithms (NLG). In the first step, structured node data is extracted from the already constructed knowledge network. Each node generally contains basic elements such as entity names, attributes, and relationship descriptions, serving as the semantic input for subsequent text generation. In the second step, the system calls text generation algorithms, such as the T5, GPT, BART models based on the Transformer architecture, to perform a preliminary conversion on the content of each node. The input at this stage is structured information, and the output is the initial text. For example, for the node "So-and-So - Organization - Peasant Association" activity, it may initially generate "So-and-So organized the Peasant Association". In the third step, the system extracts natural language features from this initial text, such as subject-predicate-object structure, semantic role labeling (SRL), keyword weights, etc., and analyzes the integrity and conciseness of the language expression. In the fourth step, for the extracted language features, the system uses a text rephrasing module to optimize the language, such as forming an intermediate representation text through semantic near-synonym replacement, syntactic structure transformation, etc. For example, "So-and-So organized the Peasant Association" can be rephrased as "The organizer of the Peasant Association is So-and-So". In the fifth step, the system determines whether the intermediate representation meets the requirements of the "chain of thought", that is, whether the information expression follows the chronological order, causal relationship, or hierarchical structure. If it meets the condition of logical fluency, the system further applies a sentence enhancement algorithm, such as using a deep syntactic reconstruction module, to generate a text with enhanced logical expression. In the sixth step, the system analyzes whether the semantics are consistent according to the conversion process of the previous and subsequent nodes. This step can use cosine similarity to judge the proximity of the previous and subsequent texts in the semantic space. If the similarity is higher than the preset threshold (such as 0.90), it is determined that the semantics are consistent. The literal calculation process of semantic similarity is: after converting the two texts into vectors, calculate their vector dot products respectively, and then divide by the product of the lengths of the two vectors. The result is the similarity value. In the seventh step, if the optimized text passes the semantic consistency verification, the system will evaluate its quality according to content evaluation metrics (such as BLEU score, ROUGE metric, language fluency score, etc.). If the score exceeds the set threshold value (such as BLEU score ≥ 0.65), the final representation text is output. In the eighth step, the system compares all the output texts with the original knowledge network to check whether there are unconverted nodes. If any are found missing, it automatically calls the steps described above to perform the conversion again until all nodes are converted into natural language texts. In the ninth step, all the final texts are spliced and synthesized into a complete output text according to the node dependency relationship or logical order. This text is semantically complete and structurally clear, and can be used as input material for application scenarios such as natural language summaries of archive content, document introductions, and voice broadcast texts.
[0063] This implementation method realizes the automatic conversion of structured knowledge into natural language content. Through technical means such as multi-stage language enhancement, logical chain judgment, and semantic consistency verification, it ensures that the finally generated text has strong logic, natural expression, and accurate content, effectively improving the readability, dissemination, and service adaptability of archival content, and is applicable to intelligent service scenarios such as knowledge extraction, abstract generation, and automatic narration.
[0064] Taking the digital guided tour system of a certain museum as an example, a large amount of structured description information of collections is stored in its knowledge base, such as "Collection name: A certain oil lamp, Era: 1934, Use: Night lighting". Through node conversion, this implementation method automatically generates the natural language text "A certain oil lamp was born in 1934 and is an important tool used by someone for night lighting during a certain period". The system can generate similar texts for all collection information and splice them into a complete voice-guided tour copy according to the exhibition theme, realizing the AI explanation function in the digital exhibition and significantly improving the visiting experience.
[0065] S5 includes obtaining the original expression text, extracting the concise text through text parsing technology, determining the core expression content, for the concise text, generating a chart form using multi-modal conversion technology to obtain a chart data set, extracting features from the chart data set, converting it into an audio form using audio synthesis technology to obtain audio data, through video rendering technology, combining the chart data set and the audio data to generate a video form to obtain video data. If the data is complete after the conversion technology processing, integrating the chart form, audio form, and video form to generate a diversified set, according to the diversified set, using a logical organization method to sort by the thinking chain to obtain an ordered conversion result, and through information comparison technology, judging that the ordered conversion result is consistent with the concise expression to obtain the final conversion set.
[0066] This embodiment aims to convert natural language text content into multimodal forms such as charts, audio, and video, enhancing the dissemination form and expression effect of archival information. In the first step, the system obtains the original natural language text and invokes the text parsing module to perform structured processing on it, including sentence splitting, redundancy removal, keyword extraction, etc., to obtain a concise text with refined and clear semantics. In this step, the TF-IDF algorithm is used to extract keywords, that is, the inverse product of the frequency of a word in the current text and its frequency in the entire corpus, which is used to measure the importance of the word to the current text. In the second step, the concise text is input into the multimodal conversion module, and first, a chart form is generated. At this time, the system identifies the data type, time type, and relationship type content contained in it, and determines the chart type accordingly (such as bar chart, pie chart, flowchart). The available structure of the chart data set is: Chart data set = { chart type, axis information, label items, data points}. In the third step, visual features such as color, proportion, trend slope, etc. are extracted from the chart data for subsequent audio-visual coordination generation. For example, in time series data, the system can calculate the slope (change speed), and its literal expression is "slope = difference in data between two time points / time difference". In the fourth step, audio synthesis technologies (such as Tacotron 2, FastSpeech) are invoked to convert the concise text into natural language speech. This module converts the text into a speech spectrum, and then the vocoder (such as WaveGlow) synthesizes natural audio data. In the fifth step, a video rendering engine (such as FFmpeg or WebGL) is used to synchronously integrate the generated charts and audio to form an animated display. The system aligns the audio with the chart changes through the time axis control module to ensure that the explanatory content is consistent with the visual content. In the sixth step, if the charts, audio, and video are all successfully generated and meet the requirements of preset formats, lengths, completeness, etc. (such as video length > 10 seconds, speech coverage rate > 95%), the system combines these three types of content into a diversified set. In the seventh step, to ensure the logicality of the expression, the system organizes the multimodal content based on the "chain of thought" organization strategy (such as the order of presenting the theme first, explaining the data, and sublimating the conclusion), sorts the multimodal content, and splices it according to the logical structure to form a complete and ordered conversion result. In the eighth step, the system compares whether the final conversion set and the original concise text are expressed consistently through natural language comparison algorithms, such as BERT semantic embedding similarity, edit distance calculation, etc. For example, if the similarity of the semantic vectors of the two texts is greater than 0.85, it is considered that the expressions are equivalent. Finally, the conversion set is output, including structural charts, explanatory audio, and synchronized video, forming a complete multimodal expression product. This embodiment takes multimodal conversion as the core, combines chart generation, audio synthesis, video rendering and other links to achieve a complete expression from concise natural language text to images, audio, and video. In the text parsing stage, the system uses the TF-IDF algorithm to identify keywords. The calculation process is as follows: TF (term frequency): the number of times a word appears in a document divided by the total number of words in the document.For example, if the word "reduction and exemption" appears 5 times in the text and the total number of words is 200, then TF = 5÷200 = 0.025. IDF (Inverse Document Frequency): log(total number of documents÷number of documents containing the word). If there are 10,000 documents in the corpus and 200 of them contain "reduction and exemption", then IDF = log(10,000÷200) ≈ log(50) ≈ 1.70. TF-IDF value: 0.025×1.70 ≈ 0.0425. The higher this value, the more important the word is to the text.
[0067] If the text contains continuous numerical data, such as "The sales in 2020 was 100 million yuan, in 2021 it was 120 million yuan, and in 2022 it was 150 million yuan", the system will identify two dimensions, "year" and "sales", and automatically select a line chart according to the following conditions: Condition 1: There are ≥2 time points with an order relationship; Condition 2: Each time point has corresponding continuous numerical values; System determination: "Year" constitutes the X-axis, and "Sales" is the Y-axis → The output chart type is "line chart".
[0068] For a line chart, the system calculates the trend slope (k) to judge the growth or decline rate: The sales in 2020 was 100, in 2021 it was 120, and in 2022 it was 150 (unit: 10,000 yuan). Slope from 2020 - 2021 = (120 - 100)÷(2021 - 2020) = 20 / 1 = 20. Slope from 2021 - 2022 = (150 - 120)÷(2022 - 2021) = 30 / 1 = 30. System analysis: The slope is increasing, determined as "accelerated growth trend".
[0069] Divide the concise text content into N keyword sentence groups (for example, 10), synthesize the voice and mark whether each group is effectively read. If 9 segments of voice cover the original information segments, then the voice coverage rate is: 9÷10 = 90%. If the set threshold is 85%, then the output condition is met.
[0070] The final output text is compared with the original concise text. The system calculates the cosine similarity through embedded vectors: Embed the two texts into vectors A and B respectively; Calculate the dot product A·B, that is, multiply the corresponding dimensions of the two vectors and then sum; Calculate the norms of vectors A and B respectively; Similarity = dot product÷(norm of A×norm of B); For example, A·B = 35, ‖A‖ = 7, ‖B‖ = 6, then similarity = 35÷(7×6) = 35÷42 ≈ 0.833; If the set semantic consistency threshold is 0.80, the current result passes.
[0071] Video length and integrity judgment, verify whether the generated video meets the integrity requirements: Video duration > 10 seconds; The coincidence degree of audio content and chart description ≥ 90%; Data missing rate < 5%. Meeting these three items, the conversion process is qualified and can be integrated into the final multi-modal set.
[0072] This method realizes the conversion of natural language archival content into multi-modal forms such as graphics, audio, and video, which not only enhances the communication performance of the content but also improves the understanding and acceptance effect of the audience. Through integrity verification, logical sorting, and semantic consistency evaluation, it ensures that the generated content is highly reliable and has a rigorous structure, and is particularly suitable for scenarios such as education, cultural dissemination, and smart government affairs that require information to be conveyed in multiple forms.
[0073] In the application of a government affairs disclosure platform, the original policy provisions are often long paragraphs of legal language. After the system extracts key points such as "the value-added tax exemption amount for small and micro enterprises increased by 20% in 2024", it generates video content that includes a line chart showing the annual increase and decrease trends and a voiceover explaining the policy background and impact. All the content is automatically synthesized and published on WeChat official accounts and mini-program platforms, facilitating the public to quickly understand the essence of the policy and widely improving the communication efficiency and social response rate.
[0074] S6 includes obtaining the behavior data corresponding to the user behavior in the target group through the behavior log to obtain the initial data set, using a clustering algorithm to extract the association features between the conversion forms and the user behavior from the initial data set to obtain the feature matrix, judging the communication effect of each conversion form in the target group for the feature matrix to obtain the effect score, if the effect score is lower than the preset threshold, then removing the corresponding conversion form from the form set to obtain the filtered form set, generating an optimized combination using an optimization algorithm through the filtered form set to obtain the candidate combination set, obtaining the behavior data of each combination in the target group for the candidate combination set, judging the communication effect to obtain the final optimized combination, and extracting the key conversion forms from the final optimized combination to generate a communication execution plan.
[0075] This implementation method evaluates the communication effectiveness of different conversion formats within a user group and, through an optimization algorithm, selects the most effective combination, enabling dynamic optimization and execution of communication strategies. First, the system extracts behavioral data about the target group from user behavior logs, including each user's click count, dwell time, like frequency, forwarding behavior, and bounce rate for content in different conversion formats. For example, a user browsed "chart" content for 45 seconds, received no likes or forwarding, while browsing "video" content for 125 seconds, received three likes and one forwarding, and did not exit the page. Next, the system clusters all user behavioral data for each content format and extracts correlation features between conversion format and user behavior. Each conversion format, such as "chart," "audio," or "video," is represented as a sample, and its corresponding behavioral features, such as average dwell time, like rate, forwarding rate, and bounce rate, serve as attribute dimensions for the sample, forming a set of feature vectors. The system then calculates a communication effectiveness score for each conversion format based on these feature values. The scoring formula is: Communication Effectiveness Score = Weight 1 × Average Dwell Time + Weight 2 × Like Rate + Weight 3 × Retweet Rate − Weight 4 × Bounce Rate. For example, for the conversion format "Video," if its average dwell time is 105 seconds, its Like Rate is 31%, its Retweet Rate is 18%, and its Bounce Rate is 10%, and the weights are set to 0.4, 0.2, 0.3, and 0.1, respectively, the Communication Effectiveness Score is calculated as follows: Video Score = 0.4 times 105, plus 0.2 times 0.31, plus 0.3 times 0.18, minus 0.1 times 0.10. This equals: 42 + 0.062 + 0.054 − 0.01 = 42.106. If a conversion format's score falls below the set communication threshold (e.g., 30 points), it is considered to have poor communication effectiveness and will be eliminated from the candidate list. Subsequently, the system uses genetic algorithms or other optimization strategies to arrange and combine the remaining conversion forms, generating multiple communication form combination schemes (such as "chart + video", "audio + video", etc.). Each combination is tested on the user group, and the behavioral performance data under the combination is recollected. Each combination is again comprehensively evaluated according to the above-mentioned scoring formula, and the combination with the highest communication score is selected as the final optimized combination. Finally, the 1-2 conversion forms with the highest key contribution are identified from the optimized combination, such as "video" with the highest score, which is determined as the core conversion method. Combined with the characteristics of the user terminal and the communication channel, such as WeChat official account, learning platform homepage or social media, a communication execution plan containing elements such as content type, release time, and release platform is generated for actual deployment.
[0076] By quantitatively analyzing user behavior and combining optimization algorithms to automatically screen conversion methods, this method can dynamically adjust the content release form, significantly improving the acceptance and dissemination efficiency of archival knowledge products among the target group. The dissemination effect is data-driven rather than subjective selection, enhancing the scientificity and adaptive ability of the strategy.
[0077] In the "Digital Archival Culture Enters Campus" project of a provincial archives, the system is deployed in the background of the official account and educational APP. Behavioral data analysis shows that the student group has the longest stay time and the most likes for the "video + bullet screen explanation" combination. The system automatically eliminates the graphic form and adopts the form of video + voice animation, which is centrally released on the home pages of WeChat and the learning platform from 20:00 to 21:00 in the evening. After implementation, the average dissemination effect score has increased by about 60%, and a large number of student comments and feedback have been collected, forming a good dissemination closed-loop.
[0078] S7 includes obtaining original knowledge data through a content delivery network, using processing technologies to perform structured conversion on the data to obtain processed knowledge products, extracting key features from the processed knowledge products, combining network technologies to determine the push requirements for the adapted scenarios, judging the push priority. If the push priority is higher than the preset threshold, the processed knowledge products are transmitted to the adapted scenarios through push technologies to generate initial output data. According to the initial output data, optimization dissemination technologies are used to adjust the dissemination form to obtain a real-time updated version of dynamic knowledge. The real-time updated version of dynamic knowledge is analyzed through technology combination, feedback data is obtained from the adapted scenarios, the adjustment direction of the application output is determined. For the feedback data, machine learning algorithms are used to iteratively optimize the dynamic knowledge to obtain the final application output, and the final application output is pushed to the target scenario through the content delivery network to complete the application deployment of dynamic knowledge.
[0079] In this embodiment, the knowledge content is structurally processed, adaptively pushed, dynamically adjusted, and feedback optimized through a content delivery network, ensuring that archival knowledge is deployed and disseminated in the most appropriate form in different scenarios. In the first step, the system obtains the original knowledge data from the content delivery network, such as text entries, graph nodes, relationship structures, etc. These data are processed through a preset data processing module for format conversion, semantic chunking, label standardization, etc., to obtain a processed knowledge product with clear structure. In the second step, the key features of the above knowledge product are extracted, including the subject of the field, information freshness, target user matching degree, content popularity, etc. The system calculates a "push priority score" for each piece of knowledge content to determine whether it is worthy of immediate push. The calculation method of this score is: push priority score = user activity × weight 1 + content timeliness × weight 2 + adaptation score × weight 3 + popularity score × weight 4. For example, the features of a certain knowledge content are as follows: the user activity (i.e., the number of visitors in the current target group) is 1000; the content is no more than 2 hours old since generation, corresponding to a timeliness score of 90; the semantic adaptation score is 85 (the matching degree is calculated on a percentage basis); the popularity score (such as the weighted score composed of clicks, forwards, etc.) is 110. The weights are set as follows: weight 1 (user activity) is 0.2, weight 2 (timeliness) is 0.3, weight 3 (adaptation) is 0.25, and weight 4 (popularity) is 0.25. Then the calculation of the push priority is: 1000×0.2 + 90×0.3 + 85×0.25 + 110×0.25 = 200 + 27 + 21.25 + 27.5 = 275.75. If this value exceeds the system-set priority threshold (e.g., 200), it is considered that this knowledge product has the value of immediate push, and the push process is started. In the third step, the system selects the most suitable transmission path (such as the nearest CDN node) according to parameters such as the current user's location, terminal type, network environment, etc., and pushes this knowledge product to specific application scenarios such as government affairs terminals, mobile devices, web interfaces, etc., to form an initial output version. In the fourth step, after the initial output version is put into use in the target scenario, the system continuously collects user behavior feedback data, including access duration, click frequency, browsing depth, bounce rate, user rating, etc. If it is found that the current dissemination form has an unsatisfactory effect (such as the average browsing duration of users is significantly lower than other versions), the system will switch the dissemination form through a content conversion strategy, such as converting the original graphic and text content into a voice explanation or a short video explanation. In the fifth step, for the above "dynamic knowledge update version" after adjustment, the system continues to collect new user feedback and uses machine learning algorithms (such as gradient boosting decision tree models or neural networks) to train a prediction model for identifying which feature combinations can maximize user response. In the sixth step, the system fine-tunes and optimizes the content based on the prediction results, such as adjusting the word order, using more understandable words, simplifying the logical structure, and generating the final application output content.In the seventh step, the final version is pushed to the target platform by the content delivery network again to achieve the deployment of dynamic knowledge and continuous adaptation and update.
[0080] This method provides a full-process automated framework from knowledge data modeling, real-time push, feedback analysis to optimization and iteration, and has the dynamic ability to adapt to the needs of different terminals and business scenarios. By perceiving the user behavior data in real time, the system can continuously adjust the content output strategy, so that the knowledge products always maintain the latest and optimal expression forms, greatly improving the intelligence and practical value of knowledge services.
[0081] In the intelligent archive service cloud platform of a certain province, for different user roles (such as government affairs personnel, community residents, research scholars), the system automatically analyzes their historical access trajectories and behavior characteristics, optimizes the structure of the "major historical event graph", and converts the original text-based display into a voice-guided video. Before pushing, the system identifies that the community users have high activity and the content has high popularity, and the priority score is 92 points. The system immediately schedules the edge nodes to push the updated version to the social APP side, and the access volume increases by 134% after pushing. The system collects the return visit behavior data and continuously optimizes the video structure with a machine learning model to achieve the personalized and intelligent upgrade of government affairs knowledge services.
[0082] S8 includes obtaining user feedback data through an external interface, calculating the distribution characteristics of the feedback data using statistical tools to obtain preliminary analysis results, extracting dynamic knowledge content according to the preliminary analysis results, matching the knowledge application scenarios with preset rules to determine the applicable knowledge scope, screening the feedback data subset through the applicable knowledge scope, training a model using a machine learning algorithm to obtain a feedback data classifier, adjusting the processing technical parameters according to the output result of the classifier, using a preset threshold to judge the parameter change trend to determine the technical iteration direction, updating the dynamic knowledge base through the technical iteration direction, analyzing the consistency of the knowledge base using a clustering algorithm to obtain an optimized knowledge structure, reprocessing the feedback data according to the optimized knowledge structure, if the parameter change trend exceeds the preset threshold, then using the gradient descent algorithm to adjust the technical parameters to obtain updated technical parameters, generating a processing technology configuration through the updated technical parameters, and using a consistency verification tool to detect the configuration integrity to obtain the final technical output.
[0083] The goal of this implementation method is to drive the dynamic adjustment of archive processing technology through user feedback data, and achieve the intelligent iteration and optimization of the knowledge product generation strategy.
[0084] User feedback collection and distribution analysis
[0085] First, collect user feedback data from various usage terminals or platforms through the API interface, such as information like ratings, comment tags, click popularity, and dwell time. Use descriptive statistical methods to calculate the distribution characteristics of these data, including mean, variance, maximum value, minimum value, etc. For example: If a set of user ratings is: [3, 4, 5, 3, 2, 5]; the mean = the sum of all ratings divided by the number of ratings, that is, (3 + 4 + 5 + 3 + 2 + 5) / 6 = 22 / 6 ≈ 3.67; the variance = the average of the squares of the differences between each rating and the mean, that is: ((3 - 3.67)² + (4 - 3.67)² +... + (5 - 3.67)²) / 6 ≈ 1.06. This distribution characteristic is used to reflect the acceptance degree and fluctuation of users towards the current dynamic knowledge content.
[0086] Dynamic Knowledge Extraction and Scenario Matching
[0087] Based on the analysis results, identify high-frequency concerns. For example, if users frequently comment "The picture is too small" or "The concept is difficult to understand", the system extracts the content containing relevant questions from the dynamic knowledge base and identifies the applicable scenarios through the semantic rule matching mechanism, such as "Mobile government affairs browsing" or "Beginner learning path", etc., to determine the applicable scope of this knowledge content.
[0088] Feedback Data Classifier Training and Technical Parameter Adjustment
[0089] Filter the feedback data according to the knowledge adaptation scope, and only retain the strongly relevant data subset for training the classification model. The training model can select Support Vector Machine (SVM), Random Forest, XGBoost, etc., and the output result is the classification label, such as "The content is verbose", "The structure is complex", "The interface is unfriendly", etc. Map the classification results to the technical parameter adjustment dimension. For example, "The content is verbose" corresponds to the "Text paragraph length" parameter. The change trend of each parameter is calculated by the first-order difference mean of the time series data: If the values of a certain parameter in the past 5 iterations are: [100, 95, 90, 85, 80] in sequence, the first-order difference is: [-5, -5, -5, -5], and the average change rate = (-5 + -5 + -5 + -5) / 4 = -5. If this change rate exceeds the set threshold (such as -3), it is considered that this parameter enters the rapid decay interval, triggering key monitoring or reconstruction.
[0090] Knowledge Base Consistency Analysis and Gradient Update
[0091] To ensure the logical consistency of the updated knowledge structure, the system applies a clustering algorithm, such as K-means, to the knowledge base to semantically recluster the knowledge content, ensuring independence between classes and homogeneity within classes. If the parameter change range exceeds the set threshold in multiple dimensions simultaneously, the system uses the gradient descent algorithm to optimize all parameters, that is: new parameter = original parameter - learning rate × current gradient direction. For example, the original parameter is 80, the current gradient is 4, and the learning rate is 0.1. Then the new parameter = 80 - 0.1×4 = 79.6. All updated parameters will be combined to form a new processing technology configuration, and the integrity of logic, interfaces, semantics, etc. will be checked through a consistency verification tool. After confirmation, the final technology output version will be generated for guiding subsequent dynamic knowledge generation and application.
[0092] This method deeply integrates user behavior and feedback data into the technical parameter optimization process, realizing the automated update and iteration of archival knowledge products from content to technical configuration. It no longer relies on manual experience to formulate rules, but continuously evolves through a feedback loop and machine learning models, enabling the system to have a high level of adaptability, stability, and intelligence.
[0093] In the intelligent archival service platform of a national library, users browse the collection archival atlas through a mini-program and submit feedback. The system discovers through feedback collection that the label frequency of "some terms are difficult to understand" has increased significantly, and trains the model to identify its multi-source pointing to the knowledge nodes of the "professional terms" category. Based on this, the system automatically adjusts the default opening parameter of the "term explanation module" from off to on, and dynamically inserts term entry links into the interface. Finally, the average page stay time is increased by 22%, and the user score is increased by 0.8 points, verifying the effectiveness of feedback-driven technical adjustment.
[0094] S9 includes obtaining raw data and preprocessing it according to technical parameters to obtain a standardized data set, processing the standardized data set through clustering analysis to determine the preliminary clustering result, constructing a knowledge graph for the preliminary clustering result to obtain a structured knowledge network, extracting features from the structured knowledge network to generate an initial product sequence, using optimization iteration to adjust the initial product sequence to obtain an optimized product sequence. If the optimized product sequence meets the preset threshold, the final knowledge product sequence is output; if not, clustering analysis and knowledge graph construction are looped, and by continuously optimizing and updating technical parameters, a knowledge product sequence adapted to changes is obtained.
[0095] This implementation method aims to continuously optimize and output a high-quality archival knowledge product sequence through the dual mechanisms of clustering and knowledge graph construction.
[0096] Preprocessing and standardization of raw data
[0097] First, receive the original data input of the file content (which may be text, tables, image tags, etc.), and perform a preprocessing process according to the current technical parameter settings, including operations such as missing value imputation, redundant information removal, and feature scaling. For example: perform stemming and stop word removal on text fields; apply normalization to numerical fields: if the original value of a certain field is 80, the maximum value is 100, and the minimum value is 20, then its normalized value is: (80 - 20) ÷ (100 - 20) = 60 ÷ 80 = 0.75. The processing results form a standardized data set under a unified scale.
[0098] Cluster analysis and generation of preliminary clustering results
[0099] Apply a clustering analysis algorithm (such as K-means) to the standardized data set. The clustering goal is to classify data with similar semantics into the same category. Each data point is represented in vector form, and the Euclidean distance is used to measure its proximity to each cluster center: Suppose a text data vector is A = [0.6, 0.8], and a center point is C = [0.5, 0.7], then the Euclidean distance is: Calculate the value of [(0.6 - 0.5)² + (0.8 - 0.7)²], and then take the square root, which is approximately equal to 0.141. Each data is assigned to the nearest cluster center to form a preliminary clustering result.
[0100] Knowledge graph construction and generation of a structured knowledge network
[0101] Construct a knowledge graph based on the clustering results. Each clustering category represents a theme entity node, and the semantic relationships between nodes (such as "association", "reference", "contrast") are connected as edges. For example, connect the "agricultural policy" clustering category with the "rural land use records" category through the "overlapping attribution time period" relationship. The graph construction is managed using a graph data model (such as Neo4j, RDF) to generate a knowledge network with a clear structure and inferability.
[0102] Initial product sequence generation and optimization
[0103] Extract keywords, knowledge links, and important node paths from the above knowledge network as the basic content of knowledge product entries, and combine them to generate an initial product sequence. Introduce an optimization algorithm (such as a greedy algorithm or a genetic algorithm) to iteratively evaluate and adjust this sequence. The evaluation metrics include information coverage rate, redundancy, node centrality, etc.; if the knowledge nodes included in the product sequence cover more than 80% of the original knowledge network and the redundant content does not exceed 15%, then the evaluation value is "qualified". Set the evaluation function as follows: Evaluation score = coverage rate × weight 1 - redundancy × weight 2 + average path depth × weight 3. If the evaluation score is greater than the set threshold (such as 85 points), then the output is the final knowledge product sequence; otherwise, return to perform a new clustering and graph construction process, and synchronously update technical parameters, such as the number of clusters, the threshold of graph structure complexity, etc.
[0104] Technical parameter continuous optimization mechanism
[0105] If the evaluation score fails to reach the preset target consistently in multiple iteration cycles, the system will trigger the parameter optimization mechanism and call the gradient descent algorithm to adjust the technical parameters. For example, the update of the clustering number K is: new K = current K - learning rate × current gradient. For instance, if the current K = 10, the gradient is 2, and the learning rate is 0.1, then new K = 10 - 0.2 = 9.8 → rounding down to get K = 9.
[0106] This method realizes the dynamic optimization of knowledge product construction through a nested iteration mechanism, combines the double-layer structure of clustering and knowledge graph, making the content organization more accurate and the semantics clearer. By using multi-dimensional index evaluation and parameter adaptive adjustment, it greatly improves the consistency, structure, and business suitability of knowledge product generation, and has significant practical value and scalability.
[0107] In a digital archive system of a certain university, the system collects data such as students' teaching evaluation opinions, teaching arrangement documents, and course score sheets. The original data is clustered after standardization to form topic graphs such as "Teaching effect feedback", "Course structure suggestions", and "Teacher behavior impressions". The system generates a sequence of teaching quality analysis reports accordingly. After optimization and iteration, the final version presents important problem nodes upfront, and the redundant description is compressed by nearly 40%. This sequence is used in the school leadership decision support system, promoting the adjustment of the course structure and the reform of the teacher evaluation mechanism, and realizing the intelligent transformation of archive data into management strategies.
[0108] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A digital archive storage and management method, characterized in that: include: S1. Obtain the original data of the archive content, and extract key information units from the text using optical character recognition technology combined with natural language processing algorithms to obtain preliminarily classified semantic fragments; S2. Use a clustering analysis algorithm to thematically aggregate the content, determine the thematic category to which each segment belongs, and obtain a thematic content collection; S3. Using knowledge graph construction technology, the logical relationships between semantic segments are mapped to generate a clearly structured knowledge network. S4. Apply a text generation algorithm to restate the knowledge nodes in natural language to obtain a more understandable text. By analyzing the node conversion process, the semantic consistency of the conversion process is determined. If the evaluation result reaches a preset threshold, the final text is output. S5. Incorporate multimodal conversion technology to convert text content into charts, audio, or video formats, generating a diverse set of conversion formats. If the data is complete after conversion technology processing, integrate the charts, audio, and video formats. Based on the diverse set, use logical organization methods to sort them according to the chain of thought to obtain an ordered conversion result. Use information comparison technology to determine whether the ordered conversion result is consistent with the concise expression, thereby obtaining the final conversion set. S6. Use user behavior analysis algorithms to determine the communication effects of different conversion methods among target groups and obtain an optimized combination of communication methods; S7. Use content distribution network technology to push processed knowledge products to adapted application scenarios and generate dynamic knowledge application outputs; S8. Collect user feedback data and iteratively adjust the processing technology through machine learning algorithms to obtain updated technical parameters; use statistical tools to calculate the distribution characteristics of feedback data, use preset thresholds to determine parameter change trends, and determine the direction of technical iteration; If the parameter change trend exceeds the preset threshold, the gradient descent algorithm is used to adjust the technical parameters to obtain updated technical parameters; S9. Reprocess the original data of the archive content based on the updated technical parameters, perform cluster analysis and knowledge graph construction in a loop, and generate a continuously optimized sequence of knowledge products.
2. The method for digital storage and management of archives according to claim 1, characterized in that: Said S1 comprises: Obtain original data, process the archive content through optical character recognition technology, extract text information, and obtain a first text set; for the first text set, use a natural language processing algorithm to analyze the text information, extract key information, and obtain a second information set; extract information units from the second information set, perform preliminary classification according to preset classification rules, and obtain a third classification set; based on the preliminary classification in the third classification set, use a clustering algorithm to analyze the semantic segments to obtain a fourth semantic set; through the fourth semantic set, judge the correlation between the semantic segments; if the correlation is higher than a preset threshold, merge the related semantic segments to obtain a fifth merged set; obtain the fifth merged set; for the merged semantic segments, use a statistical method to determine the distribution characteristics of the key information to obtain a sixth feature set; based on the sixth feature set, judge the integrity of the semantic segments by comparing with the preset semantic template to obtain a seventh verification set.
3. The method for digital storage and management of archives according to claim 1, characterized in that: The S2 includes: By acquiring semantic fragments, the word segmentation technology is applied to preliminarily classify the text to obtain initial group data. The initial group data is then theorized using a clustering analysis algorithm to determine the subject categories of each group. Based on the results of the subject aggregation, a content set containing subject categories is generated. If the number of subject categories in the content set exceeds the preset threshold, the categories are filtered based on the judgment criteria to obtain a streamlined set. For the streamlined set, the distribution characteristics of each subject category are calculated using a statistical analysis method to obtain distribution data. The core characteristics of the subject category are determined through the distribution data to generate analysis results. Based on the analysis results, visualization tools are used to present the subject category distribution to obtain the final output.
4. The method for digital storage and management of archives according to claim 1, characterized in that: The S3 includes: The subject content is obtained through the content collection, and the word segmentation technology is used to extract semantic fragments to obtain preliminary semantic units. For the semantic fragments, the dependency syntax analysis is used to determine the logical relationship and obtain the dependency structure between the fragments. The association mapping is obtained from the dependency structure, and the knowledge graph construction technology is used to generate the initial network structure to determine the nodes and edges. If the logical relationship between the nodes meets the preset threshold, the network structure is adjusted through the clustering algorithm to obtain the optimized knowledge network. According to the optimized knowledge network, the fragment association is iteratively updated to obtain the semantic extension after the mapping is generated. The clarity of the structure is judged by the semantic extension, and the final knowledge network is generated by topological sorting. The complete representation of the subject content is extracted from the final knowledge network to obtain a structured semantic expression.
5. The method for digital storage and management of archives according to claim 1, characterized in that: The S4 includes: Through the pre-established knowledge network, node data with clear structure is obtained, and the text generation algorithm is used to perform preliminary conversion on the node content to obtain the initial text. Natural language features are extracted from the initial text, and the content is restated based on the features to generate an intermediate expression text. If the intermediate expression text meets the requirements of the thinking chain, the sentence structure is adjusted through the application of the algorithm to obtain a text with enhanced logical expression. The node conversion process is analyzed according to the enhanced text, the semantic consistency in the conversion process is determined, and the optimized text is generated. The optimized text is evaluated. If the evaluation result reaches the preset threshold, the final expression text is output. By comparing the final expression text with the knowledge network, it is determined whether there are unconverted nodes. If so, the conversion process is repeated to obtain the final expression text of all nodes, which are integrated by splicing to obtain a complete and clearly structured output text.
6. The method for digital storage and management of archives according to claim 1, characterized in that: The S5 includes: Obtain the original expression text, extract concise text through text parsing technology, determine the core expression content, use multimodal conversion technology to generate a chart form for the concise text, obtain a chart data set, extract features from the chart data set, use audio synthesis technology to convert it into audio form, obtain audio data, use video rendering technology to combine the chart data set with the audio data to generate a video form, obtain video data, if the data is complete after processing by the conversion technology, then integrate the chart form, audio form and video form to generate a diversified set, based on the diversified set, use a logical organization method to sort according to the thinking chain, obtain an ordered conversion result, use information comparison technology to determine that the ordered conversion result is consistent with the concise expression, and obtain the final conversion set.
7. The method for digital storage and management of archives according to claim 1, characterized in that: The S6 includes: Through the behavior log, the behavioral data corresponding to the user behavior in the target group is obtained to obtain the initial data set. The clustering algorithm is used to extract the correlation characteristics between the conversion form and the user behavior from the initial data set to obtain the feature matrix. Based on the feature matrix, the communication effect of each conversion form in the target group is judged to obtain the effect score. If the effect score is lower than the preset threshold, the corresponding conversion form is eliminated from the form set to obtain the filtered form set. The optimization algorithm is used to generate the optimized combination based on the filtered form set to obtain the candidate combination set. Based on the candidate combination set, the behavioral data of each combination in the target group is obtained to judge the communication effect to obtain the final optimized combination. The key conversion form is extracted from the final optimized combination to generate a communication execution plan.
8. The method for digital storage and management of archives according to claim 1, characterized in that: The S7 includes: Obtain original knowledge data through the content distribution network, use processing technology to perform structured transformation on the data to obtain processed knowledge products, extract key features from the processed knowledge products, combine network technology to determine the push requirements of the adaptation scenario, and judge the push priority. If the push priority is higher than the preset threshold, the processed knowledge product is transmitted to the adaptation scenario through push technology to generate initial output data. Based on the initial output data, the optimized communication technology is used to adjust the communication form to obtain a real-time updated version of dynamic knowledge. The real-time updated version of dynamic knowledge is analyzed through a combination of technologies, feedback data is obtained from the adaptation scenario, and the adjustment direction of the application output is determined. Based on the feedback data, the dynamic knowledge is iteratively optimized using a machine learning algorithm to obtain the final application output. The final application output is pushed to the target scenario through the content distribution network to complete the application deployment of dynamic knowledge.
9. The method for digital storage and management of archives according to claim 1, characterized in that: The S8 includes: Obtain user feedback data through an external interface, use statistical tools to calculate the distribution characteristics of the feedback data, and obtain preliminary analysis results. Extract dynamic knowledge content based on the preliminary analysis results, use preset rules to match knowledge application scenarios, determine the applicable knowledge scope, filter the feedback data subset based on the applicable knowledge scope, use machine learning algorithms to train the model, and obtain a feedback data classifier. Adjust the processing technology parameters based on the classifier output results, use preset thresholds to judge the parameter change trend, determine the direction of technical iteration, update the dynamic knowledge base based on the direction of technical iteration, use clustering algorithms to analyze the consistency of the knowledge base, and obtain an optimized knowledge structure. Reprocess the feedback data based on the optimized knowledge structure. If the parameter change trend exceeds the preset threshold, use the gradient descent algorithm to adjust the technical parameters to obtain updated technical parameters. Generate the processing technology configuration based on the updated technical parameters, use consistency verification tools to detect the configuration integrity, and obtain the final technical output.
10. The method for digital storage and management of archives according to claim 1, characterized in that: The S9 includes: Obtain the original data and pre-process it according to the technical parameters to obtain a standardized data set. Process the standardized data set through cluster analysis to determine the preliminary clustering results. Construct a knowledge graph based on the preliminary clustering results to obtain a structured knowledge network. Extract features from the structured knowledge network to generate an initial product sequence. Use optimization iteration to adjust the initial product sequence to obtain an optimized product sequence. If the optimized product sequence meets the preset threshold, the final knowledge product sequence is output; if not, the cluster analysis and knowledge graph construction are executed cyclically. By continuously optimizing and updating the technical parameters, a knowledge product sequence that adapts to changes is obtained.
Citation Information
Patent Citations
Video processing method and device, electronic equipment and storage medium
CN114666655A
File multi-mode intelligent compilation method and system based on knowledge graph
CN117216008A