A digital-humanities-based local archives search management system and method
By employing digital humanities-based archive search and management systems, and utilizing information area division, collection, identification, and encryption technologies, the inefficiency and security issues of local archive management have been resolved, achieving efficient and secure archive information management and sharing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUANGXI ACAD OF SCI
- Filing Date
- 2023-05-23
- Publication Date
- 2026-05-01
AI Technical Summary
Local archives management departments lack advanced infrastructure and standardized management practices, resulting in heavy workloads, low efficiency, and high error rates. This makes it impossible to achieve intelligent and precise management, and digital archives are easily damaged or lost, with insufficient security and confidentiality.
The system employs a digital humanities-based archive search and management system. Through the division, collection, identification, database establishment, text information acquisition and matching analysis of archive information areas, it calculates the priority matching coefficient of archives, realizes efficient information retrieval and presentation, and encrypts important archive information through the RSA algorithm.
It enables precise and rapid management and secure storage of local archives, reduces management costs, improves information sharing and security, and prevents human-caused damage and computer virus attacks.
Smart Images

Figure CN116594957B_ABST
Abstract
Description
A Local Archives Search and Management System and Method Based on Digital Humanities Technical Field
[0001] This invention relates to the field of search management technology, and more specifically, to a local archives search management system and method based on digital humanities. Background Technology
[0002] In the era of big data, the way archival information is preserved has undergone tremendous changes, and the digital transformation of document management has become an inevitable trend. In particular, digital archives built on the basis of digital technologies such as computers and networks will make the sharing and use of archival information more convenient and efficient while ensuring the integrity and security of archival data.
[0003] Digital humanities connects diverse information resources, including archives, digital technologies, humanities disciplines, and other fields related to specific issues. It demonstrates how information resources, represented by archives, can be integrated across disciplines to meet humanistic and even broader social pursuits. It also builds management systems based on digital technologies, expanding the understanding and methodology of archive management. On the other hand, it strengthens the problem-oriented awareness of archive resource integration and services, making archive management objects more specific and services more refined, thus broadening new paths for the development of the archive field to meet the needs of humanities and society.
[0004] However, in actual use, it still has some shortcomings. For example, due to lack of funds, some local archives management departments lack advanced infrastructure, archives management work lacks rationality, and related operations lack standardization, resulting in problems such as large workload, low efficiency, and high error rate. It is impossible to achieve intelligent and accurate management of local archives information and improve the quality of work.
[0005] Because digital archives are easily deleted and damaged, data corruption or loss can occur during the management process due to damage or aging of storage equipment or other unforeseen circumstances, threatening the security and confidentiality of archival information. Summary of the Invention
[0006] In order to overcome the above-mentioned defects of the prior art, embodiments of the present invention provide a local archives search and management system and method based on digital humanities, which is used to solve the problems mentioned in the background art.
[0007] To achieve the above objectives, the present invention provides the following technical solution:
[0008] Archive Information Area Division Module: This module is used to divide the archives of the area to be managed into monitoring sub-areas according to the year, and to number each monitoring sub-area of the archives of the area to be managed sequentially as 1, 2, ..., n.
[0009] Archive Information Collection Module: This module is used to collect archive information from various monitoring sub-areas within the region to be managed and synchronize the information to the Archive Information Recognition Module.
[0010] Archival Information Recognition Module: This module is used to acquire archival information collected by the archival information acquisition module. It identifies archival file names and keywords through text information, statistically analyzes user search volume indicators through archival information, converts physical data into digital file format through image scanning, and transmits the recognized information to the archival information database creation module.
[0011] Archival Information Database Establishment Module: This module receives information transmitted from the archival information identification module and establishes an archival information database to store the attributes, characteristics, and location of archives based on their file names.
[0012] Text information acquisition module: used to acquire text information input by users in each monitoring sub-region of the area to be managed through the server, and transmit the text information to the text information extraction module.
[0013] Text information extraction module: This module receives text information transmitted from the text information acquisition module, filters out keyword information, obtains the search volume index of the document user based on the keyword information, and transmits the data information to the information matching and analysis module.
[0014] Information Matching and Analysis Module: Based on the keyword information extracted from the text information module of each monitoring sub-region of the regional archives to be managed, the module calculates the keyword relevance weight index, calculates the search volume weight index based on the user's search volume index, and calculates the archive priority matching coefficient based on the keyword relevance weight index and the search volume weight index, and transmits the results to the archive search and evaluation module.
[0015] The document search and evaluation module receives data from the information matching and analysis module, compares it with the preset document priority matching coefficient threshold, matches the corresponding information in the document information database based on the results, and outputs the document information results, including document name, pictures, audio and video materials.
[0016] Search Engine Optimization Module: Based on the keyword density input by the user in the text information extraction module and the output archive information results, the module dynamically sorts the results by level, prioritizing related searches for the user's next search.
[0017] Archival resource security supervision module: Based on the archival information that needs to be encrypted in each monitoring sub-region of the region to be managed, it is classified into one category and the data is encrypted using the RSA algorithm.
[0018] The year span in the archive information area division module, which is divided by year, is no less than one year. The monitoring sub-areas of the archives to be managed are numbered sequentially as 1, 2, ..., n.
[0019] The specific meaning of the archive information collection module is:
[0020] The system scans and inputs information from local archives, including archive names, images, audio, and video materials. Administrators can add, modify, and delete archive information.
[0021] The text information extraction module is specifically as follows:
[0022] Keywords extracted from the user-input text are numbered as q1, q2, ..., q i , ..., q n .
[0023] The formula for calculating the keyword relevance weight index is as follows:
[0024] Where α represents the keyword relevance weight index, λ i Let q be the impact factor of the i-th keyword. i Let n be the number of the i-th keyword. i This represents the number of characters in the i-th keyword.
[0025] The k i The calculation formula is:
[0026] Where q i This is represented as the number of the i-th keyword.
[0027] The formula for calculating the search volume weight index is as follows:
[0028] Where β represents the search volume weight index, h represents the user's search volume metric, ΔH represents the preset user search volume metric threshold, and λ represents other influencing factors of the search volume weight index.
[0029] The formula for the priority matching coefficient of the archives is:
[0030] Where θ represents the file priority matching coefficient, α represents the keyword relevance weight index, β represents the search volume weight index, μ represents the compensation factor for the file priority matching coefficient, e represents the natural constant, and λ represents other influencing factors of the file priority matching coefficient.
[0031] The specific evaluation method for the document search and evaluation module is as follows:
[0032] The priority matching coefficient θ of each monitored sub-region of the regional archives to be managed is compared with the preset priority matching coefficient threshold Δθ. If θ > Δθ, the matching archive file name in the archive information database is matched and the stored archive information is output. Otherwise, the result is sent to the management personnel for processing.
[0033] A method for searching and managing local archives based on digital humanities, the method being used to implement the local archives search and management system based on digital humanities as described in any one of claims 1-8, characterized by comprising the following steps:
[0034] Step S01: Division of Archive Information Areas: Specifically, the archives of the area to be managed are divided into monitoring sub-areas according to the year, and the monitoring sub-areas of the archives of the area to be managed are numbered sequentially as 1, 2, ..., n.
[0035] Step S02: Archive Information Collection: Specifically, this involves collecting archive information from each monitored sub-region of the region to be managed and synchronizing the information to the archive information identification module.
[0036] Step S03: Archival Information Recognition: Specifically, this involves identifying the archival file name and keyword information through archival text information, statistically analyzing user search volume indicators through archival information, and converting physical data into digital file format through image scanning.
[0037] Step S04: Establishment of the archival information database: Specifically, this involves establishing an archival information database, storing the attributes, characteristics, and location of archives based on their file names.
[0038] Step S05: Text Information Acquisition: Specifically, this involves acquiring the text information input by users in each monitoring sub-region of the region to be managed through the server.
[0039] Step S06: Text information extraction: Specifically, based on the text information, keyword information is filtered out, and the search volume index of the document by users is obtained based on the keyword information.
[0040] Step S07: Information Matching Analysis: Specifically, based on the keyword information extracted from the text information modules of each monitored sub-region of the regional archives to be managed, the keyword relevance weight index is calculated, the search volume weight index is calculated based on the user's search volume index, and the archive priority matching coefficient is calculated based on the keyword relevance weight index and the search volume weight index.
[0041] Step S08: Archive Search Evaluation: Specifically, the archive priority matching coefficient is compared with the preset archive priority matching coefficient threshold. Based on the result, the corresponding information is matched in the archive information database, and the archive information results are output, including archive name, pictures, audio and video materials.
[0042] Step S09: Search Engine Optimization: Specifically, based on the keyword density input by the user in the text information extraction module and the output archive information results, dynamic ranking is performed, and related searches are given priority recommendation when the user searches again.
[0043] Step S10: Security supervision of archival resources: Specifically, the archival information that needs to be encrypted in each monitoring sub-region of the region to be managed is classified into one category, and the data is encrypted using the RSA algorithm.
[0044] The technical effects and advantages of this invention are as follows:
[0045] 1. This invention provides a local archives search and management system and method based on digital humanities. By collecting archives information from each monitored sub-region of the region to be managed, the system extracts the archives information and further establishes an archives information database. Based on the text information input by the user, the system extracts keyword information and search volume, and then analyzes and obtains the archives priority matching coefficient. The coefficient is compared with a preset archives priority matching coefficient threshold, and the system automatically retrieves the archives information searched by the user in the database. The system accurately and quickly outputs the archives information searched by the user. At the same time, through the optimization function of the search engine, the system achieves efficient information retrieval and presentation, enabling the scientific and rational application of digital technology to play a greater role in modern archives management.
[0046] 2. Traditional record management relies heavily on human, material, and financial resources to manage paper-based records. In contrast, digital record management, based on modern information technology, reduces costs and expands the scope, time, and service range of record usage. It also makes the transmission and sharing of record information resources simpler and more efficient. Encrypting and archiving record information that requires encryption is beneficial to the security of records, preventing human damage, computer virus attacks, or malicious deletion, and ensuring the security of digitized records. Attached Figure Description
[0047] Figure 1 is a schematic diagram of the system module connection of the present invention.
[0048] Figure 2 is a schematic diagram of the file search and evaluation module of the present invention. Detailed Implementation
[0049] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0050] Please refer to Figure 1. This invention provides a local archives search and management system and method based on digital humanities, including an archives information area division module, an archives information collection module, an archives information identification module, an archives information database establishment module, a text information acquisition module, a text information extraction module, an information matching and analysis module, an archives search evaluation module, a search engine optimization module, and an archives resource security supervision module.
[0051] The archival information area division module is connected to the archival information acquisition module, the archival information acquisition module is connected to the archival information identification module, the archival information identification module is connected to the archival information database establishment module, the text information acquisition module is connected to the text information extraction module, the text information extraction module is connected to the information matching and analysis module, the information matching and analysis module is connected to the archival information database establishment module, the archival information database establishment module is connected to the archival search and evaluation module, the archival search and evaluation module is connected to the search engine optimization module, and the search engine optimization module is connected to the archival resource security supervision module.
[0052] A local archives search and management system and method based on digital humanities, characterized by comprising:
[0053] Archive Information Area Division Module: This module is used to divide the archives of the area to be managed into monitoring sub-areas according to the year, and to number each monitoring sub-area of the archives of the area to be managed sequentially as 1, 2, ..., n.
[0054] Archive Information Collection Module: This module is used to collect archive information from various monitoring sub-areas within the region to be managed and synchronize the information to the Archive Information Recognition Module.
[0055] Archival Information Recognition Module: This module is used to acquire archival information collected by the archival information acquisition module. It identifies archival file names and keywords through text information, statistically analyzes user search volume indicators through archival information, converts physical data into digital file format through image scanning, and transmits the recognized information to the archival information database creation module.
[0056] Archival Information Database Establishment Module: This module receives information transmitted from the archival information identification module and establishes an archival information database to store the attributes, characteristics, and location of archives based on their file names.
[0057] Text information acquisition module: used to acquire text information input by users in each monitoring sub-region of the area to be managed through the server, and transmit the text information to the text information extraction module.
[0058] Text information extraction module: This module receives text information transmitted from the text information acquisition module, filters out keyword information, obtains the search volume index of the document user based on the keyword information, and transmits the data information to the information matching and analysis module.
[0059] Information Matching and Analysis Module: Based on the keyword information extracted from the text information module of each monitoring sub-region of the regional archives to be managed, the module calculates the keyword relevance weight index, calculates the search volume weight index based on the user's search volume index, and calculates the archive priority matching coefficient based on the keyword relevance weight index and the search volume weight index, and transmits the results to the archive search and evaluation module.
[0060] The document search and evaluation module receives data from the information matching and analysis module, compares it with the preset document priority matching coefficient threshold, matches the corresponding information in the document information database based on the results, and outputs the document information results, including document name, pictures, audio and video materials.
[0061] Search Engine Optimization Module: Based on the keyword density input by the user in the text information extraction module and the output archive information results, the module dynamically sorts the results by level, prioritizing related searches for the user's next search.
[0062] Archival resource security supervision module: Based on the archival information that needs to be encrypted in each monitoring sub-region of the region to be managed, it is classified into one category and the data is encrypted using the RSA algorithm.
[0063] In one possible design, the year span of the archive information area division module is no less than one year, and the monitoring sub-areas of the area archives to be managed are numbered sequentially as 1, 2, ..., n.
[0064] In one possible design, the archive information acquisition module specifically refers to:
[0065] The system scans and inputs information from local archives, including archive names, images, audio, and video materials. Administrators can add, modify, and delete archive information.
[0066] In one possible design, the text information extraction module specifically comprises:
[0067] Keywords extracted from the user-input text are numbered as q1, q2, ..., q i ,...,q n .
[0068] Please refer to Figure 2. The document search evaluation module includes keyword relevance weight index, search volume weight index, and document priority matching coefficient.
[0069] In one possible design, the formula for calculating the keyword relevance weight index is:
[0070] Where α represents the keyword relevance weight index, λ i Let q be the impact factor of the i-th keyword. i Let n be the number of the i-th keyword. i This represents the number of characters in the i-th keyword.
[0071] The k i The calculation formula is:
[0072] Where q i This is represented as the number of the i-th keyword.
[0073] In one possible design, the formula for calculating the search volume weight index is:
[0074] Where β represents the search volume weight index, h represents the user's search volume metric, ΔH represents the preset user search volume metric threshold, and λ represents other influencing factors of the search volume weight index.
[0075] In one possible design, the formula for the file priority matching coefficient is:
[0076] Where θ represents the file priority matching coefficient, α represents the keyword relevance weight index, β represents the search volume weight index, μ represents the compensation factor for the file priority matching coefficient, e represents the natural constant, and λ represents other influencing factors of the file priority matching coefficient.
[0077] In one possible design, the specific evaluation method of the archive search evaluation module is as follows:
[0078] The priority matching coefficient θ of each monitored sub-region of the regional archives to be managed is compared with the preset priority matching coefficient threshold Δθ. If θ > Δθ, the matching archive file name in the archive information database is matched and the stored archive information is output. Otherwise, the result is sent to the management personnel for processing.
[0079] A method for searching and managing local archives based on digital humanities, the method being used to implement the local archives search and management system based on digital humanities as described in any one of claims 1-8, characterized by comprising the following steps:
[0080] Step S01: Division of Archive Information Areas: Specifically, the archives of the area to be managed are divided into monitoring sub-areas according to the year, and the monitoring sub-areas of the archives of the area to be managed are numbered sequentially as 1, 2, ..., n.
[0081] Step S02: Archive Information Collection: Specifically, this involves collecting archive information from each monitored sub-region of the region to be managed and synchronizing the information to the archive information identification module.
[0082] Step S03: Archival Information Recognition: Specifically, this involves identifying the archival file name and keyword information through archival text information, statistically analyzing user search volume indicators through archival information, and converting physical data into digital file format through image scanning.
[0083] Step S04: Establishment of the archival information database: Specifically, this involves establishing an archival information database, storing the attributes, characteristics, and location of archives based on their file names.
[0084] Step S05: Text Information Acquisition: Specifically, this involves acquiring the text information input by users in each monitoring sub-region of the region to be managed through the server.
[0085] Step S06: Text information extraction: Specifically, based on the text information, keyword information is filtered out, and the search volume index of the document by users is obtained based on the keyword information.
[0086] Step S07: Information Matching Analysis: Specifically, based on the keyword information extracted from the text information modules of each monitored sub-region of the regional archives to be managed, the keyword relevance weight index is calculated, the search volume weight index is calculated based on the user's search volume index, and the archive priority matching coefficient is calculated based on the keyword relevance weight index and the search volume weight index.
[0087] Step S08: Archive Search Evaluation: Specifically, the archive priority matching coefficient is compared with the preset archive priority matching coefficient threshold. Based on the result, the corresponding information is matched in the archive information database, and the archive information results are output, including archive name, pictures, audio and video materials.
[0088] Step S09: Search Engine Optimization: Specifically, based on the keyword density input by the user in the text information extraction module and the output archive information results, dynamic ranking is performed, and related searches are given priority recommendation when the user searches again.
[0089] Step S10: Security supervision of archival resources: Specifically, the archival information that needs to be encrypted in each monitoring sub-region of the region to be managed is classified into one category, and the data is encrypted using the RSA algorithm.
[0090] In this embodiment, it should be specifically explained that the present invention provides a local archives search and management system and method based on digital humanities. This system collects archives information from various monitored sub-regions of the area to be managed, identifies the archives information, and further establishes an archives information database. Based on the text information input by the user, a keyword relevance weight index is calculated using keyword information, and a search volume weight index is calculated using the user's search volume index. This analysis yields an archives priority matching coefficient, which is compared with a preset archives priority matching coefficient threshold. The system then automatically retrieves the archives information searched by the user from the database, accurately and quickly outputting the information. Simultaneously, through the optimization function of the search engine, efficient information retrieval and presentation are achieved, enabling the scientific and rational application of digital technology to play a greater role in modern archives management.
[0091] In this embodiment, it should be specifically noted that traditional record management relies on a large amount of human, material, and financial resources to manage paper-based records. In contrast, digital record management, based on the application of modern information technology, simplifies costs, expands the space, time, and service scope of record usage, and makes the transmission and sharing of record information resources simpler and more efficient. Encrypting and archiving record information that requires encryption is beneficial to the security of records, preventing human damage, computer virus attacks, or malicious deletion, and ensuring the security of digitized records.
[0092] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A local archives search and management system based on digital humanities, characterized in that, include: The document consists of several modules: **Archival Information Area Division Module:** This module divides the archives of the area to be managed into monitoring sub-regions according to year, numbering each sub-region sequentially as 1, 2, ..., n. **Archival Information Acquisition Module:** This module collects archival information from each monitoring sub-region of the area to be managed and synchronizes the information to the Archival Information Recognition Module. **Archival Information Recognition Module:** This module acquires the archival information collected by the Archival Information Acquisition Module, identifies file names and keywords from text information, statistically analyzes user search volume, converts physical data into digital file format through image scanning, and transmits the recognized information to the Archival Information Database Establishment Module. **Archival Information Database Establishment Module:** This module receives information from the Archival Information Recognition Module and establishes an archival information database, storing the attributes, characteristics, and location of archives based on their file names. Text information acquisition module: used to acquire text information input by users in each monitoring sub-region of the area to be managed through the server, and transmit the text information to the text information extraction module; Text information extraction module: used to receive the text information transmitted by the text information acquisition module, filter out keyword information, obtain the search volume index of document users based on the keyword information, and transmit the data information to the information matching and analysis module; Information Matching and Analysis Module: Based on the keyword information extracted from the text information module of each monitored sub-region of the archives in the region to be managed, the module calculates the keyword relevance weight index, calculates the search volume weight index based on the user's search volume index, and calculates the archive priority matching coefficient based on the keyword relevance weight index and the search volume weight index. The results are then transmitted to the archive search evaluation module. The formula for calculating the keyword relevance weight index is: ,in This is expressed as a keyword relevance weight index. Let the i-th keyword be the influence factor. This is represented by the number of the i-th keyword. The number of characters in the i-th keyword; The calculation formula is: ,in This is represented by the i-th keyword number; The archive search and evaluation module receives data transmitted from the information matching and analysis module, compares it with a preset archive priority matching coefficient threshold, matches corresponding information in the archive information database based on the result, and outputs the archive information results, including archive name, images, audio, and video materials; the formula for the archive priority matching coefficient is: ,in This is represented as the file priority matching coefficient. This is expressed as a keyword relevance weight index. This is expressed as a search volume weighting index. The compensation factor is represented by the file priority matching coefficient, and e is represented by the natural constant. Other influencing factors are represented as the archive priority matching coefficient; the specific evaluation method of the archive search evaluation module is as follows: the archive priority matching coefficient of each monitored sub-region of the region to be managed is used. Preset file priority matching coefficient threshold To make a comparison, if If the matching file name matches the archival information database, the stored archival information is output; otherwise, the result is sent to the administrator for processing. The search engine optimization module dynamically sorts the results based on the keyword density input by the user in the text information extraction module and the output archival information, prioritizing related searches for the user's next search. The archival resource security supervision module categorizes the archival information requiring encryption within each monitoring sub-region of the managed area into one category and encrypts the data using the RSA algorithm.
2. The local archives search and management system based on digital humanities as described in claim 1, characterized in that: The year span in the archive information area division module, which is divided by year, is no less than one year. The monitoring sub-areas of the archives to be managed are numbered sequentially as 1, 2, ..., n.
3. The local archives search and management system based on digital humanities as described in claim 1, characterized in that: The aforementioned archive information collection module specifically refers to scanning and inputting information from local archives, including archive names, images, audio, and video materials. Administrators can add, modify, and delete archive information.
4. The local archives search and management system based on digital humanities as described in claim 1, characterized in that: The text information extraction module specifically performs the following steps: It extracts keywords from the user-input text information and then assigns them numbers. , ,..., ,..., 。 5. A local archives search and management system based on digital humanities as described in claim 1, characterized in that: The formula for calculating the search volume weight index is as follows: ,in This is expressed as a search volume weighting index. This is represented as a metric based on user search volume. This represents the preset threshold for user search volume metrics. Other influencing factors are represented as search volume weighting indices.
6. A method for searching and managing local archives based on digital humanities, the method being used to implement the local archives search and management system based on digital humanities as described in any one of claims 1-5 above, characterized in that, Includes the following steps: Step S01: Archive Information Area Division: Specifically, the archives in the area to be managed are divided into monitoring sub-areas according to the year, and each monitoring sub-area is numbered sequentially as 1, 2, ..., n; Step S02: Archive Information Collection: Specifically, the archive information data of each monitoring sub-area of the archives in the area to be managed is collected and synchronized to the archive information recognition module; Step S03: Archive Information Recognition: Specifically, the archive file name and keyword information are identified through the archive text information, the user search volume index is statistically analyzed through the archive information, and the physical data is converted into digital file format through image scanning; Step S04: Archive Information Database Establishment: Specifically, an archive information database is established, storing the attributes, characteristics, and location of archives according to the archive file name; Step S05: Text Information Acquisition: Specifically, the text information input by users in each monitoring sub-area of the archives in the area to be managed is obtained through the server; Step S06: Text Information Extraction: Specifically, keyword information is filtered out based on the text information, and the data is extracted based on the keyword information. The document's user search volume metrics; Step S07: Information Matching Analysis: Specifically, based on the keyword information extracted from the text information module of each monitoring sub-region of the archives in the region to be managed, a keyword relevance weight index is calculated, a search volume weight index is calculated based on the user search volume metrics, and an archive priority matching coefficient is calculated based on the keyword relevance weight index and the search volume weight index; Step S08: Archive Search Evaluation: Specifically, the archive priority matching coefficient is compared with the preset archive priority matching coefficient threshold, and corresponding information is matched in the archive information database based on the results, outputting archive information results, including archive name, images, audio, and video materials; Step S09: Search Engine Optimization: Specifically, based on the keyword density input by the user in the text information extraction module and the output archive information results, a dynamic ranking is performed, and related searches are prioritized for the user's next search; Step S10: Archive Resource Security Supervision: Specifically, the archive information that needs to be encrypted in each monitoring sub-region of the archives in the region to be managed is classified into one category, and the data is encrypted using the RSA algorithm.
Citation Information
Patent Citations
Digital archive management system based on big data
CN115994745A