Electronic archives management system based on knowledge graph
By using a knowledge graph-based electronic records management system, the problems of low query efficiency and unstructured data processing in large-scale electronic records management systems have been solved, enabling efficient and accurate record querying and display, and improving the query and utilization efficiency for users.
Patent Information
- Application Number
- CN202311195734.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-15
- Publication Date
- 2026-01-13
- Estimated Expiration
- 2043-09-15
AI Technical Summary
Existing electronic record management systems suffer from low query and retrieval efficiency when faced with a large number of electronic records. They are unable to effectively handle non-standard and unstructured data, making it difficult to meet the requirements for accuracy and recall, and users find it difficult to quickly locate the target record.
An electronic records management system based on knowledge graphs is adopted, including modules for data storage, management, relationship analysis, relationship storage, and relationship display. Through data preprocessing, cleaning, mining, and knowledge graph display, the system constructs the relationships between records to achieve efficient querying and display.
It improves the speed and accuracy of query and retrieval, can handle non-standard and unstructured data, provides more accurate search results, allows users to intuitively understand the relationships between files, quickly locate target files, and improve work efficiency and data utilization.
Smart Images

Figure CN117171105B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of computer technology, and in particular to an electronic record management system based on knowledge graphs. Background Technology
[0002] With the continuous development of digital technology, archival management has gradually shifted from paper-based document management to a digital process. Compared with traditional paper-based archival management, electronic archival management, through online data operation, eliminates the need for physical searching, making it more efficient and convenient. Furthermore, as the number of archives increases over time, traditional paper archives occupy more and more space, and the search and retrieval of archives becomes increasingly complex, impacting the work efficiency of archival management personnel. Compared to traditional paper archives, electronic archive management and retrieval are simpler and more efficient. Archival management personnel can easily locate specific archives online and quickly view their contents. Moreover, electronic archives offer more standardized data access control, with more standardized and easier-to-operate isolation of different data permissions and confidentiality levels, ensuring better data security and greatly reducing the risk of important information leaks due to archival management personnel negligence during access. In addition to managing massive amounts of information, electronic archival management systems can also combine pattern recognition, natural language processing, and other artificial intelligence technologies to mine information from massive archives, revealing deep-seated relationships between archives, fully exploring the value of archival data, and providing more support and assistance for archival management.
[0003] Electronic records management systems are comprehensive solutions for modernizing the management of records in enterprises and institutions. They are computerized management information systems that enable the receipt, management, storage, and utilization of electronic records. These systems possess features such as openness, functional scalability, flexible configuration, and security and reliability, supporting management of multiple categories and formats, and also providing auxiliary management functions for physical records. Electronic records are an important component of national information resources, and the effective utilization of electronic record information is of great significance for improving work efficiency and maximizing data value.
[0004] Currently, many electronic record management systems only store and retrieve archival data in a simple way. As the number of electronic records stored in business systems increases, traditional query and retrieval schemes become inefficient and cannot meet the accuracy and recall requirements when processing non-standard and unstructured data. Analyzing and displaying existing archival data relationships is becoming increasingly difficult. Summary of the Invention
[0005] This specification provides one or more embodiments of a knowledge graph-based electronic records management system to address the technical problems raised in the background section.
[0006] One or more embodiments of this specification employ the following technical solutions:
[0007] This specification provides one or more embodiments of a knowledge graph-based electronic records management system, comprising: a data storage module, a data management module, a relationship analysis module, a relationship storage module, and a relationship display module; wherein,
[0008] The data storage module is used to collect archive data and to store and query various types of business archives through a database.
[0009] The data management module is used to manage the preset business processes of electronic archives;
[0010] The relationship analysis module is used to summarize different archive data stored in the data storage module to obtain the archive relationships of the different archive data;
[0011] The relationship storage module is used to distinguish and store file relationships based on different business operations and different historical periods;
[0012] The relationship display module is used to obtain the archive relationships stored in the relationship storage module, and to display the association between the current archive data and other archive data through a knowledge graph.
[0013] The above-described at least one technical solution adopted in the embodiments of this specification can achieve the following beneficial effects:
[0014] Traditional query and retrieval methods become inefficient as the number of electronic records increases. However, the knowledge graph-based electronic record management system described in this specification can effectively improve the speed and accuracy of query and retrieval through relationship analysis and relationship display modules. Through the display of the knowledge graph, users can more intuitively understand the relationships between records and quickly find the required record information.
[0015] Traditional query and retrieval schemes have limited capabilities in processing non-standard and unstructured data. However, the knowledge graph-based electronic records management system described in this specification can aggregate different types of archival data and differentiate and store archival relationships from different business periods and historical periods through a relational storage module, thereby better handling non-standard and unstructured data. In this way, the system can provide more accurate and comprehensive search results, meeting users' requirements for precision and recall.
[0016] This specification describes an electronic records management system based on a knowledge graph. Through a relationship display module, it visualizes the relationships between records stored in the relationship storage module. This display method intuitively presents the connections between current record data and other record data, helping users better understand the relationships and value between records. In this way, users can quickly locate and utilize relevant record data, improving work efficiency and maximizing data value. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings:
[0018] Figure 1 A schematic diagram of the structure of a knowledge graph-based electronic records management system provided for one or more embodiments of this specification;
[0019] Figure 2 This specification provides a relationship visualization process for one or more embodiments.
[0020] Figure 3 This is a schematic diagram of file relationships provided for one or more embodiments of this specification. Detailed Implementation
[0021] This specification provides an electronic records management system based on a knowledge graph.
[0022] Currently, many electronic records management systems only provide simple storage and retrieval of archival data. As the volume of electronic records stored in business systems grows larger, traditional query and retrieval methods become inefficient and fail to meet accuracy and recall requirements when processing non-standard and unstructured data. Analyzing and displaying existing archival data relationships is becoming increasingly difficult. By analyzing and displaying the relationships between records from different business sources, the utilization efficiency of archival data can be greatly improved. Furthermore, leveraging the advantage of electronic records—the ability to extend the preservation time of materials indefinitely—it is possible to better utilize business data relationships from different periods, greatly facilitating the access and use of archival documents and maximizing the value of archival data.
[0023] In recent years, knowledge graphs have become a popular technology in the field of artificial intelligence, achieving great success in areas such as semantic search, intelligence analysis, and intelligent question answering. Knowledge graphs represent the relationships between different things through "points" and "edges," structuring heterogeneous knowledge within a domain and building knowledge connections. This addresses application scenarios where data is scattered across multiple systems, diverse, complex, and isolated, with low value for individual data points. A unified structured representation combined with rich semantic information constructs rich relationships that can be directly provided to downstream applications. For enterprises and organizations facing challenges in integrating and analyzing heterogeneous data sources with inconsistent formats and standards, knowledge graphs can standardize these data sources and transform them into knowledge graphs. Leveraging the visualization and reasoning capabilities of knowledge graphs facilitates unified querying and analysis, helping users discover hidden relationships and patterns, and supporting more efficient, accurate, and intelligent data analysis.
[0024] Introducing knowledge graphs into electronic record management systems helps users quickly build management relationships between archival data from different business systems. Through the reasoning and visualization capabilities of the graph, users can quickly find knowledge, build knowledge connections, discover hidden data relationships, and effectively link different business data. This achieves efficient and intelligent business processes, empowering enterprises and organizations to utilize knowledge data in multiple dimensions. In traditional electronic record management systems, archives from different business sources are stored and displayed according to type. Users need to filter queries in the database based on specific field information of the archives, resulting in a single query method and slow query speeds as the volume of archive data increases. Furthermore, when users are unclear about the specific information of the archive they are looking for, they cannot directly locate their target archive, further increasing query costs and disrupting normal work processes. By leveraging the archival relationship network built with knowledge graphs, electronic record management systems can link archival data scattered across multiple business sources. This relationship network constructs potential connections between archival information from different business sources, solving the problems of siloed data and low application value of single data. Users can leverage this information, combined with customized diagrams tailored to their business needs, to explicitly solidify and connect domain knowledge, expanding the scenarios for data utilization: by linking archival information from different business sources, the scope of archival queries and retrieval can be expanded, improving the efficiency of users' queries and retrieval of relevant archival information, and providing a better ability to organize, manage, and understand the massive amounts of information on the Internet.
[0025] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.
[0026] Figure 1 This is a schematic diagram of the structure of a knowledge graph-based electronic records management system provided for one or more embodiments of this specification.
[0027] The electronic records management system may include: a data storage module 102, a data management module 104, a relationship analysis module 106, a relationship storage module 108, and a relationship display module 110; among which,
[0028] The data storage module in the embodiments of this specification can be used to collect archive data and to store and query various types of business archives through a database.
[0029] In the embodiments of this specification, it can be first determined which archival data needs to be collected and the source of the data; then, a suitable data collection method can be designed, which may include manual input, importing files, data interfaces with other systems, etc., to ensure that the collected data is accurate and complete, and to consider data security and the standardization of data format; then, a suitable database system can be selected, such as a relational database (e.g., MySQL, Oracle) or a document database (e.g., MongoDB), to store the archival data, considering the structure and relationships of the data, and designing the database table structure, including fields and indexes, to support subsequent query and analysis needs.
[0030] Furthermore, the embodiments in this specification can create corresponding data tables and store data according to the database design. Database management tools or programming languages can be used to create data tables and perform data insertion operations; then, based on the database query language (such as SQL) or programming interface, the function of querying archive data can be implemented. Query interfaces can be designed according to business needs, supporting queries based on keywords, time ranges, business types, and other conditions, ensuring the accuracy of query results and response speed.
[0031] Meanwhile, the embodiments in this specification can establish a data quality management mechanism, including data cleaning, data verification, and data backup, to ensure the accuracy and integrity of archival data; they can also design and implement relevant data access control strategies based on the sensitivity of the archives and business needs to ensure that only authorized personnel can access and modify the corresponding archival data.
[0032] Furthermore, the embodiments in this specification can also test data storage and retrieval functions to ensure system stability and performance. Optimization can be performed based on the test results to improve system response speed and user experience.
[0033] The data management module in the embodiments of this specification can be used to manage the preset business processes of electronic archives.
[0034] In the embodiments of this specification, specific pre-defined business process requirements can be identified first. These pre-defined business processes may include receiving, organizing, utilizing, appraising, statistically analyzing, and auditing archival data. Communication and negotiation can be conducted with archival management personnel and relevant departments to understand the business processes and data requirements. Then, based on the pre-defined business processes, appropriate business processes can be designed, clarifying the specific operational steps, roles, and permissions for each process, as well as the data flow and data processing requirements.
[0035] In the embodiments of this specification, the process and method for receiving archival data can be designed and implemented, which may include file uploading, data import, interface docking, etc., to ensure the accuracy and integrity of archival data.
[0036] In the embodiments of this specification, the received archival data can be organized and evaluated in accordance with archival management standards and requirements. This may include data classification, archiving, integrity checks, quality control, and data verification to ensure the accuracy, integrity, and consistency of the archival data.
[0037] In the embodiments described in this specification, functions and interfaces can be provided to support the viewing, retrieval, analysis and utilization of archival data. Query interfaces and data analysis functions are designed according to business needs to support users' effective utilization of archival data.
[0038] In the embodiments described in this specification, statistical and audit tracking functions can be designed and implemented to statistically analyze and audit the usage of archival data. This includes recording user operation logs, tracking data usage, and auditing operations such as data modification and deletion.
[0039] The embodiments in this specification can provide system management functions, including user permission management, role management, system configuration, log management, etc., to ensure the security, stability and manageability of the system.
[0040] In the embodiments described in this specification, the functionality of the data management module can be tested to ensure the stability and performance of the system. Based on the test results, optimization can be performed to improve the system's response speed and user experience.
[0041] The relationship analysis module in this embodiment can be used to summarize different archive data stored in the data storage module to obtain the archive relationships of the different archive data.
[0042] In the embodiments of this specification, when obtaining the archival relationship of the different archival data, the archival relationship of the different archival data can be obtained by preprocessing, data reduction, data cleaning and data mining analysis of the different archival data. The preprocessing includes missing value processing.
[0043] When handling missing values in different archival data according to the embodiments of this specification, missing values of archival data for the same business in the same period can be filled by averaging, which means that the missing values are filled by the average value of other data for the same business in that period; missing values of archival data for different businesses in different periods can be filled by preset rules, which can be logical rules predefined according to business needs to determine and fill missing values; missing values of archival data for the same business in different periods can be filled by the Last Observation Carried Forward (LOCF) method, which means that the last observed value can be used as the filling value for the missing value, assuming that the value is relatively stable; missing values of archival data for different businesses in the same period can be filled by statistical methods, which means that statistical methods (such as average, median, etc.) are used to calculate a reasonable value to replace the missing value.
[0044] When performing data cleaning on the different archive data in this embodiment of the specification, the content of the different archive data can be detected according to the preset archive data detection standards to see if it conforms to the archive storage standards, and the archive data that does not conform to the archive storage standards can be filtered out.
[0045] Specifically, the embodiments in this specification can first define the archiving and storage standards for the archives, that is, determine which content and rules are considered to conform to the standards. These standards may include data format, data structure, field rules, data types, etc. According to the preset detection standards, content detection is performed on each piece of archive data to determine whether it conforms to the standards. This can be verified by comparing the archive data with the requirements defined in the standards.
[0046] In the content inspection process described in this specification, non-compliant archive data is identified. This data may include formatting errors, missing fields, and field values that do not meet specifications. The detected non-compliant archive data can be processed or marked. Processing may include correcting formatting errors, filling in missing fields, and adjusting non-compliant field values. Marking can be used for subsequent processing or further analysis.
[0047] Furthermore, the embodiments in this specification may perform the data cleaning step multiple times to ensure that the archive data complies with the specifications.
[0048] When performing data mining analysis on the different archival data in the embodiments of this specification, semantic analysis can be performed on the field information of the different archival data based on deep learning networks, NLP frameworks and the data characteristics of different businesses, so as to realize multi-level and multi-dimensional data analysis of characters, words and chapters in the different archival data and obtain the semantic analysis content of the different archival data.
[0049] The embodiments in this specification can use deep learning networks, such as autoencoders, recurrent neural networks (RNNs), or convolutional neural networks (CNNs), to build suitable models based on the data characteristics of different businesses. Natural language processing (NLP) frameworks and techniques, such as bag-of-words models, word embeddings, TF-IDF, and topic models, can be used to perform semantic analysis on the field information of different archive data and extract relevant features.
[0050] The embodiments in this specification can apply NLP technology and deep learning models to perform text parsing and semantic analysis on characters, words, and passages in the different archival data. By analyzing the semantics, sentiment, theme, keywords, etc. of the text, the meaning and relationships of the data can be extracted and understood.
[0051] The embodiments in this specification can perform multi-level, multi-dimensional data analysis based on the results of semantic analysis. Data visualization tools and techniques can be used to display the analysis results in the form of charts, images, word clouds, etc., to better understand and discover relevant patterns and trends in the data.
[0052] Through the above implementation steps, multi-level and multi-dimensional data analysis, including semantic analysis, can be performed on the different archival data. This helps to deeply explore the inherent meaning and relationships of the data, providing richer information and insights to support business decisions and business optimization.
[0053] The embodiments in this specification can reduce the data of the different archives, and the following implementation steps can be taken:
[0054] Data cleaning and organization: This involves cleaning and organizing archival data, including removing redundant data, correcting errors, and standardizing formats and naming conventions. This helps improve data quality and accuracy and prepares the data for subsequent data reduction.
[0055] Data classification and labeling: Classify and label data according to its type, content, and value. Data can be classified using tags, metadata, or other methods to facilitate the development and implementation of subsequent mitigation strategies.
[0056] Data evaluation and filtering: Evaluate and filter each piece of data based on actual needs and objectives. Consider factors such as the timeliness, importance, and availability of the data to determine which data should be retained, deleted, or archived.
[0057] Develop reduction strategies: Based on the results of data evaluation and screening, develop specific reduction strategies. The following strategies can be considered:
[0058] Remove redundant data: Delete duplicate, redundant, or no longer useful data.
[0059] Remove outdated data: Identify and delete expired, obsolete, or no longer needed data.
[0060] Archive storage: Archive data that is not frequently accessed in the long term to free up storage space.
[0061] Compressing data: Using compression algorithms to compress data and reduce storage space usage.
[0062] Implement the data reduction strategy: Based on the established strategy, begin implementing data reduction. Delete, archive, or compress data, ensuring the data reduction process is correct and effective.
[0063] Monitoring and Evaluation: During implementation, monitor the effectiveness and impact of data reduction. Evaluate whether the data reduction results achieved the expected effects and ensure that the reduction process did not lead to data loss or errors.
[0064] Update documents and records: Record all data pruning operations performed, including data deletion, archiving, or compression. Ensure the accuracy and timeliness of documents and records.
[0065] Regular review and updates: Regularly review and update data reduction strategies to adapt to changing needs and requirements. Reassess data value and retention periods, and adjust reduction strategies and standards accordingly.
[0066] By implementing the above steps, different types of archival data can be effectively reduced, eliminating redundant and useless data and optimizing data management and utilization efficiency. At the same time, it also helps reduce storage costs, improve data access speed, and ensure data compliance and reliability.
[0067] The relational storage module in the embodiments of this specification can be used to distinguish and store file relations for different businesses and different historical periods.
[0068] In the embodiments of this specification, when storing archive relationships separately, the current archive data can be stored as the first source archive; the same business archive relationship archived in the same period can be stored as the second source archive; different business archive relationships archived in the same period can be stored as the current business source archive; the same business archive relationship archived in different periods can be stored as the second source archive, and the time interval between the two archives can be stored; different business archive relationships archived in different periods can be stored as the current business source archive, and the time interval between the two archives can be stored.
[0069] It should be noted that, in the embodiments of this specification, storing current archive data as the first source indicates that these archives are up-to-date and directly related to current business. Storing archives related to the same business archived within the same period as the second source indicates that these archives were archived within the same period but have a weaker connection to current business. Storing archives related to different businesses archived within the same period as the source archives of the current business indicates that these archives were not only archived within the same period but also directly related to current business. Storing archives related to the same business archived at different times as the second source archives and recording the time interval between the two archives indicates that these archives were archived within the same business at different historical periods, and the time interval can be used to understand the development of the business. Storing archives related to different businesses archived at different times as the source archives of the current business and recording the time interval between the two archives indicates that these archives were archived by different businesses and at different historical periods, and the time interval can be used to understand the evolution between different businesses.
[0070] By implementing the above steps, archives from different business operations and historical periods can be stored separately, facilitating subsequent management and utilization.
[0071] Furthermore, in conjunction with the above-described method of distinguishing and storing archive relationships between different business operations and different historical periods, when summarizing different archive data stored in the data storage module to obtain the archive relationships of the different archive data, the archive data of the same business obtained at the current time can be summarized, the first association relationship of the archive data of the same business can be analyzed, and the first source archive and the second source archive can be distinguished and recorded when establishing the first association relationship; the archive data of the first designated business obtained at the current time can be summarized with the archive data of the same business previously archived, the second association relationship between the archive data of the same business in different historical periods can be analyzed, and the first source archive and the second source archive can be distinguished and recorded when establishing the second association relationship. The second source archive is distinguished and recorded, and the generation time interval of the archive data is also distinguished and recorded. The archive data of the second designated business acquired at the current moment is summarized and analyzed with the archive data of other businesses acquired at the same time to obtain the third association relationship of all archive data archived at the same time. When establishing the third association relationship, the first source archive and other business source archives are distinguished and recorded. The archive data of the third business acquired at the current moment is summarized and analyzed with the archive data of other previously archived businesses to obtain the fourth association relationship of all archive data archived at the same time. When establishing the fourth association relationship, the current business source archive and other business source archives are distinguished and recorded, and the generation time interval of the archive data is also distinguished and recorded.
[0072] It should be noted that, in the embodiments of this specification, when establishing the first association relationship, the first source file and the second source file are recorded separately. This means that for file data of the same service acquired at the current time, the first association relationship is established by analyzing the relationship between them, and the difference between the first source file and the second source file is recorded. When establishing the second association relationship, the first source file and the second source file are recorded separately, and the distinction is made based on the time interval between the generation of the file data. This means that the file data of the first specified service acquired at the current time is compared and analyzed with the file data of the same service previously archived, the relationship between them is established, and the first source file and the second source file are recorded and distinguished based on the time interval. When establishing the third association relationship, the first source file and other service source files are recorded separately. This means that the relationship between the file data of the second specified service acquired at the current time and other service source data is obtained by analyzing the file data of other services, and the first source file and other service source files are recorded separately. When establishing the fourth relationship, the current business source file and other business source files are distinguished and recorded, and the records are distinguished and recorded according to the generation time interval of the file data. This means that the file data of the third business obtained at the current moment is compared and analyzed with the file data of other businesses to obtain the relationship between them, and the current business source file and other business source files are recorded and distinguished according to the time interval.
[0073] The relationship display module in this embodiment can be used to obtain the file relationships stored in the relationship storage module and display the association between the current file data and other file data through a knowledge graph.
[0074] In the embodiments of this specification, the archive relationships stored in the relation storage module can be obtained first, other archive data associated with the current archive data can be determined, and the dimensions of the other archive data can be determined; then, the nodes and associated edges of the current archive data and the other archive data in the association network can be determined through the knowledge graph, and a related association network can be generated to show the association relationship between the current archive data and the other archive data.
[0075] It should be noted that when the embodiments of this specification obtain the file relationships stored in the relationship storage module, the stored file relationship data can be obtained from the relationship storage module. This data describes the association between files, including the association between the current file data and other file data.
[0076] It should be noted that when determining other archive data associated with the current archive data and their respective dimensions in the embodiments of this specification, other archive data associated with the current archive data can be determined based on the obtained archive relationship data, and the dimensions of these archive data can be further determined. The dimensions can be business dimensions, time dimensions, or other related dimensions.
[0077] It should be noted that when the embodiments of this specification determine the nodes and edges in the association network through knowledge graphs, the technology and methods of knowledge graphs can be used to determine the nodes and related edges of the current archive data and other archive data in the association network based on the determined association relationships and dimension information. Nodes can represent archive data, and related edges can represent the association relationships between archives.
[0078] It should be noted that when generating the association network to display the association relationship in the embodiments of this specification, a related association network graph can be generated based on the nodes and association edges in the determined association network. The association network graph can be displayed in a graphical or other form to show the association relationship between the current archive data and other archive data. Through this graph, the association between archives can be understood intuitively.
[0079] It should be noted that, through the above implementation steps, the embodiments in this specification can obtain and display the relationship between current archival data and other archival data, which helps to understand and analyze the relationship between archival data and further apply it to related management and decision-making.
[0080] Furthermore, in the embodiments of this specification, when determining the nodes and associated edges of the current archive data and the other archive data in the association network through a knowledge graph, the first source archive can be determined as the root node in the association network, the other archive data can be determined as the slave nodes in the association network, and the second source archive of the same period, the current business source archive of the same period, the second source archive of different periods, and the current business source archive of different periods can be determined as the associated edges in the association network.
[0081] It should be noted that the embodiments in this specification first construct a knowledge graph, which can include nodes and related edges of archival data. Nodes represent different archival data, and related edges represent the relationships between different archival data. Next, the root node and slave nodes are determined. Based on the knowledge graph, the first source archive is determined as the root node of the association network, and other archival data are determined as slave nodes of the association network. Then, the related edges are determined. Based on the knowledge graph, the second source archive of the same period, the current business source archive of the same period, the second source archive of different periods, and the current business source archive of different periods are determined as related edges in the association network. Finally, the relationship between nodes and edges is determined. Based on the nodes and related edges in the knowledge graph, the relationship between the current archival data and other archival data is determined, that is, the node and related edge in which it is located in the association network are determined.
[0082] It should be noted that the specific implementation steps of the embodiments in this specification can be carried out in accordance with the above summary, including constructing a knowledge graph, determining the root node and slave nodes, determining associated edges, and determining the relationship between nodes and edges.
[0083] It should be noted that, in response to the problems of low query and retrieval efficiency and limited effective information in existing electronic record management systems when the amount of record data increases, this specification provides a knowledge graph-based efficient electronic record management system. By collecting basic data information of records to establish an electronic record information database, queries can be performed not only using the field information of the records but also through the constructed relationship between the records. This not only results in fast query speed but also high accuracy. When users are unclear about the specific content of the target record, they can also quickly locate it through the relationship network constructed between other records. This facilitates record management and greatly improves the utilization rate of record data.
[0084] This electronic records management system includes a data storage module, a data management module, a relationship analysis module, a relationship storage module, and a relationship display module.
[0085] 1. The data storage module uses a database to store and retrieve various business files;
[0086] 2. The data management module implements the general functional requirements of routine electronic record business processes and electronic record system management, such as receiving, sorting, utilizing, appraising, statistically analyzing, auditing, tracking, and system management.
[0087] 3. The relationship analysis module summarizes the different archive data stored in the data storage module after archiving, and obtains the relationships between the archive data through methods such as data preprocessing, data reduction, data cleaning and evaluation, and data mining analysis;
[0088] 4. The relationship storage module stores the relationships between archives from different business periods and historical periods. This electronic archive management system collects archive data from different business sources and stores it in the data storage module. It also performs correlation analysis on data archived from different time ranges and different sources and stores the obtained archive data in the relationship storage module.
[0089] 5. The relationship display module is used to retrieve the file relationships stored in the relationship storage module and display the association relationships between different files.
[0090] Figure 2 The relationship display process provided in one or more embodiments of this specification sequentially performs file data collection, file data archiving, file relationship analysis, file relationship storage, and file relationship display.
[0091] The data storage module can collect archival data and use a database to store and retrieve various types of business archives.
[0092] The data management module can fulfill the general functional requirements of routine electronic record business processes and electronic record system management, such as receiving, organizing, utilizing, appraising, statistically analyzing, auditing, tracking, and system management.
[0093] The relationship analysis module can aggregate different archive data stored in the data storage module after archiving, and obtain the relationships between the archive data through methods such as data preprocessing, data reduction, data cleaning and evaluation, and data mining analysis.
[0094] The relational storage module can store the relationships between archives from different business periods and historical periods. This electronic archive management system collects archive data from different business sources and stores it in the data storage module. It also performs relational analysis on data archived from different time ranges and different sources through the relational analysis module and stores the obtained archive data in the relational storage module.
[0095] The relationship display module can be used to retrieve file relationships stored in the relationship storage module and display the association between different files.
[0096] It should be noted that the specific technical solutions of the embodiments in this specification are as follows:
[0097] The relationship analysis module can aggregate different archived data stored in the data storage module after archiving, and obtain the relationships between the archived data through methods such as data preprocessing, data reduction, data cleaning and evaluation, and data mining analysis.
[0098] (1) Data preprocessing includes missing value handling, data cleaning, data selection, data transformation, data integration, data reduction and data cleaning;
[0099] (2) Missing values in the same period of archival data are filled by averaging, archival data from different business sources are supplemented by pre-made rules, missing values in archival data from different periods are filled by LOCF method, and archival data from different business sources in the same period are filled by statistical method.
[0100] (3) During the data cleaning process, the data content is checked and verified in accordance with the electronic archive data inspection specifications to see if it conforms to the archive storage specifications. Data that does not conform to the electronic archive archiving specifications is screened.
[0101] (4) Perform semantic analysis on the field information of the archive data, use deep learning networks and NLP frameworks to perform semantic understanding to achieve multi-level and multi-dimensional data analysis of characters, words, and chapters, obtain the semantic analysis content of the archive data, and at the same time combine the data characteristics of different business data sources to perform fine-grained semantic analysis extraction to expand the feature information of the archive data.
[0102] (5) Re-examine and verify the characteristics of the acquired archival data information, correct identifiable errors in the data files, clean up erroneous or conflicting data according to certain rules, and obtain the data feature vector for analysis in the desired format.
[0103] The relational storage module stores archival relationships from different business operations and historical periods. This electronic records management system collects archival data from various business sources, stores it in the data storage module, and performs correlation analysis on data archived from different time ranges and sources, storing the resulting archival data in the relational storage module.
[0104] (1) The data from the same business source obtained at the current time are summarized through the relationship analysis module, the inherent relationship of these data is analyzed, the generated relationship is stored in the relationship storage module, and the first source file and the second source file are distinguished and recorded when the relationship is established;
[0105] (2) The current business file is summarized with the previously archived business file through the relationship analysis module. The relationship between the data of the same business file in different historical periods is analyzed. The obtained relationship is stored in the relationship storage module. When establishing the relationship, the first source file and the second source file are distinguished and recorded. At the same time, the generation time interval of the file data is distinguished and recorded.
[0106] (3) The business file obtained at the current time is summarized and analyzed with other business files obtained at the same time through the relationship analysis module to obtain the association relationship of all archived file data at the same time. The obtained association relationship is stored in the relationship storage module. When establishing the relationship, the first source file is distinguished from other business source files and recorded.
[0107] (4) The current business file is summarized and analyzed with other previously archived business files through the relationship analysis module to obtain the association relationship of all archived file data at the same time. The obtained association relationship is stored in the relationship storage module. When establishing the relationship, the current business source file is distinguished and recorded from other business source files. At the same time, the generation time interval of the file data is distinguished and recorded.
[0108] Specifically, the relationship storage module distinguishes and records the file relationships for each file based on different business sources and business periods. It records the current file as the primary source file and then distinguishes between different business sources and business periods:
[0109] (1) The same business archives filed in the same period are regarded as second source archives;
[0110] (2) The relationship between different business archives filed in the same period shall be regarded as the current business source archive;
[0111] (3) Treat the same business archives filed at different times as second source archives and record the time interval between the two archives;
[0112] (4) The relationship between different business archives filed in different periods is used as the current business source archive, and the time interval between the two archives is recorded.
[0113] The relationship display module is used to retrieve file relationships stored in the relationship storage module and display the associations between different files. By viewing a specific file, the user can choose whether to view the corresponding file's relationships and whether to jump to the next view on the file's display interface. Clicking the "View File Relationships" option will jump to the next view.
[0114] During navigation, the system uses the relational storage module to locate the associated files across various dimensions (this navigation method avoids the time-consuming problem of directly querying all file information). The associated file data is then passed to the system front-end, displaying the relationships between different files in a "node"-"edge" manner. Different dimensions of file relationships and file nodes are differentiated and displayed, intuitively showing the relationships between files at different levels and business files from different periods. The currently displayed file serves as the root node in the relationship network, acting as the primary source file. Then, secondary source files from the relational storage module, the current business source file, secondary source files from different periods, and current business source files from different periods are simultaneously displayed on the relationship display interface. The relationship display interface can be found in [link to relevant documentation]. Figure 3 The diagram shows the relationship between the archives.
[0115] When users are unsure of the specific file information they are looking for, they can perform a fuzzy search using key information from related files associated with that file. Then, by leveraging the association between that file and the target file, they can quickly locate the target file, thus improving the efficiency of their file search and retrieval.
[0116] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.
[0117] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0118] The above description is merely one or more embodiments of this specification and is not intended to limit this specification. Various modifications and variations can be made to the one or more embodiments of this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of one or more embodiments of this specification should be included within the scope of the claims of this specification.
Claims
1. A knowledge graph-based electronic records management system, characterized in that, The system includes: a data storage module, a data management module, a relationship analysis module, a relationship storage module, and a relationship display module; wherein, The data storage module is used to collect archive data and to store and query various types of business archives through a database. The data management module is used to manage the preset business processes of electronic archives; The relationship analysis module is used to summarize different archive data stored in the data storage module to obtain the archive relationships of the different archive data; The relationship storage module is used to distinguish and store file relationships based on different business operations and different historical periods; The relationship display module is used to obtain the archive relationships stored in the relationship storage module, and to display the association between the current archive data and other archive data through a knowledge graph. The method of distinguishing and storing archives based on different business operations and historical periods includes: The current archival data is stored as the primary source archive; Records of the same business that were filed during the same period are stored as second-source records; Different business archives filed in the same period are stored as the current business source archives; The relationship between the same business archives filed at different times is stored as a second source archive, and the time interval between the two archives is also stored. The relationship between different business archives filed at different times is stored as the current business source archive, and the time interval between the two archives is stored; The step of summarizing different archive data stored in the data storage module to obtain the archive relationships of the different archive data includes: The archive data of the same business acquired at the current moment are aggregated, the first association relationship of the archive data of the same business is analyzed, and the first source archive and the second source archive are distinguished and recorded when the first association relationship is established; The archive data of the first designated business obtained at the current moment is summarized with the archive data of the same business previously archived. The second relationship between the archive data of the same business in different historical periods is analyzed. When establishing the second relationship, the first source archive and the second source archive are distinguished and recorded, and the generation time interval of the archive data is distinguished and recorded. The archive data of the second designated business acquired at the current moment is summarized and analyzed with the archive data of other businesses acquired at the same time to obtain the third association relationship of all archived archive data at the same time. When establishing the third association relationship, the first source archive and other business source archives are distinguished and recorded. The archive data of the third business acquired at the current moment is summarized and analyzed with the archive data of other previously archived businesses to obtain the fourth association relationship of all archived archive data at the same time. When establishing the fourth association relationship, the archive source of the current business is distinguished and recorded from the archive source of other businesses, and the generation time interval of the archive data is also distinguished and recorded. The step of acquiring the archive relationships stored in the relation storage module and displaying the association between the current archive data and other archive data through a knowledge graph includes: Retrieve the archive relationships stored in the relation storage module, determine other archive data associated with the current archive data, and the dimension in which the other archive data is located; The knowledge graph is used to determine the nodes and edges of the current archive data and other archive data in the association network, and a related association network is generated to show the relationship between the current archive data and other archive data; The step of determining the nodes and associated edges of the current archive data and other archive data in the association network through a knowledge graph includes: The knowledge graph determines the first source file as the root node in the association network, the other file data as slave nodes in the association network, and the second source file from the same period, the current business source file from the same period, the second source file from different periods, and the current business source file from different periods as the association edges in the association network. The preset business processes include: receiving, organizing, utilizing, identifying, statistically analyzing and auditing archival data, as well as one or more of the processes in system management.
2. The system according to claim 1, characterized in that, The process of obtaining the file relationships of the different file data includes: By preprocessing, reducing, cleaning and mining the different archival data, the archival relationships between the different archival data are obtained. The preprocessing includes handling missing values.
3. The system according to claim 2, characterized in that, When handling missing values in the different archive data, the following steps are included: Missing values in the archive data of the same business during the same period are filled in by averaging. Missing values in archival data for different periods and business operations are filled in using preset rules. The Last Observation-Driven Completion (LOCF) method fills in missing values in the archival data of the same business at different times. Missing values in the archival data of different businesses during the same period were filled in using statistical methods.
4. The system according to claim 2, characterized in that, Data cleaning of the different archive data includes: According to the preset archival data detection standards, the content of different archival data is detected to see if it conforms to the archival storage standards, and archival data that does not conform to the archival storage standards is filtered out.
5. The system according to claim 2, characterized in that, Data mining analysis was performed on the different archive data, including: Based on the characteristics of deep learning networks, NLP frameworks, and different business data, semantic analysis is performed on the field information of the different archive data to achieve multi-level and multi-dimensional data analysis of characters, words, and chapters in the different archive data, and to obtain the semantic analysis content of the different archive data.
Citation Information
Patent Citations
Electronic archiving method and system for power grid operation and maintenance project archives
CN114443923A
Archive data management method and system
CN115033528A