Full-life-cycle controlled hydrogeological digital intelligent archive system

By constructing a hydrogeological digital archive system with full lifecycle management, the problems of data fragmentation, low retrieval efficiency, and cumbersome borrowing process in hydrogeological archive management have been solved. It has realized automatic traceability of archive status, multi-dimensional retrieval, and intelligent recommendation, which has improved resource utilization and security and met the rapid response needs of hydrogeological projects.

CN122019469APending Publication Date: 2026-05-12HYDROGEOLOGY BUREAU OF CHINA COAL GEOLOGY ADMINISTRATION
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HYDROGEOLOGY BUREAU OF CHINA COAL GEOLOGY ADMINISTRATION
Filing Date
2025-12-18
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

The existing hydrogeological archive management system suffers from problems such as data fragmentation and version confusion, low retrieval efficiency, cumbersome borrowing process, and insufficient resource reuse. It lacks full life cycle management, multi-dimensional retrieval, and flexible approval mechanisms, making it difficult to meet the rapid response and safety requirements of hydrological projects.

Method used

A hydrogeological digital archive system with full life-cycle management is constructed, adopting a three-tier B/S architecture, including a full life-cycle management module, a multi-modal retrieval module, a flexible borrowing approval module, and an intelligent recommendation module, to realize automatic traceability of archive status, multi-dimensional retrieval, differentiated approval, and intelligent recommendation.

Benefits of technology

It has enabled the full lifecycle connectivity and status traceability of archival data, improved the accuracy and efficiency of retrieval, shortened the borrowing approval cycle, improved resource utilization and security, and met the rapid response needs of hydrological projects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122019469A_ABST
    Figure CN122019469A_ABST
Patent Text Reader

Abstract

The invention provides a full-life-cycle management and control hydrogeology digital intelligent archive system, relates to the technical field of archive management and hydrogeology application, and aims to solve the problems of management fragmentation, low retrieval accuracy, rigid borrowing process, low resource utilization rate and the like in hydrogeology archive management. A digital intelligent management system covering the whole life cycle is constructed, the system adopts a three-layer B / S architecture, automatic tracing of archive states is realized through a whole life cycle management and control module, retrieval accuracy is improved by using a multi-modal retrieval algorithm, a flexible borrowing approval process is designed based on classified classification, and an intelligent recommendation module is developed in combination with user behaviors. According to the method, multi-mode retrieval, process approval, intelligent pushing and full-life-cycle state tracing of the hydrogeological archive resources are achieved, archive retrieval efficiency is improved, the borrowing approval cycle is shortened, the archive resource utilization rate is improved, and powerful support is provided for efficient decision making and data sharing of hydrogeological work.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of archives management and hydrogeological application technology, and relates to a hydrogeological digital archives system with full life cycle management. Background Technology Hydrogeological archives serve as the data carrier for core tasks such as hydrological exploration, coal mine water control projects, and ecological water conservation projects. Their management efficiency directly determines the scientific validity and safety of hydrological engineering decisions. These archives contain crucial information such as borehole parameters, mine hydrogeological type reports, basic water hazard monitoring data, and water control project design schemes, providing fundamental support for water resource management, disaster prevention, and ecological protection. As hydrogeological work transitions towards digitalization and precision, traditional archive management models are no longer adequate to meet the industry's development needs, highlighting key pain points and necessitating a modern upgrade of archive management through technological innovation.

[0002] Currently, hydrogeological archives management is still mainly based on paper files or scattered electronic documents, lacking a unified full life cycle management standard, which leads to the following key problems: (1) Data fragmentation and version confusion: The archives are independent of each other from formation, filing to destruction, which easily leads to problems such as duplicate storage and inconsistent versions. For example, in coal mine water control projects, the exploration plan report and supporting drawings are often stored in different locations (such as the project team's computer and the archive server), which requires manual verification and correlation, which is inefficient and prone to errors. (2) Low retrieval efficiency: The existing retrieval method relies on shallow matching of "title-keywords", which cannot accurately retrieve the content of the archive (such as borehole depth, water chemical composition, etc.). The average search time is more than 30 minutes, which is difficult to meet the rapid response needs of emergency scenarios such as mine water inrush. (3) Cumbersome borrowing process: Borrowing approval is mainly offline, with a cycle of 3-5 working days, and it is impossible to track the status of the archives in real time (such as the borrower and the remaining time). At the same time, there is a lack of a graded confidentiality mechanism for classified archives (such as water inrush emergency data), which poses a risk of data leakage. (4) Insufficient resource reuse: The ability to share archives is weak, and a large amount of high-value data (such as drilling records in old mining areas and historical water damage cases) is idle. There is a lack of intelligent push mechanism based on user behavior or project association, which restricts the reuse of experience across regions.

[0003] Domestic research has gradually deepened since 2000, resulting in the following representative achievements, but all of them have the problem of insufficient industry adaptation: (1) Limitations of general systems: The "Digital Archives Management System Basic Version" promoted by the State Archives Administration has realized basic functions such as electronic archiving, but it has not been optimized for the characteristics of hydrogeology (such as unstructured data processing and project association needs), and cannot effectively support the dynamic management of data such as borehole curves and monitoring maps. (2) Single function of industry-specific systems: Some hydrological bureaus have developed dedicated systems (such as the "Hydrological Archives Digital Management System" of the Yellow River Hydrological Bureau) that focus on data storage and simple query, but do not cover the entire life cycle of archives, and the retrieval dimension is single (only supporting "site-time" query), making it difficult to associate with complex scenarios such as flood prevention and control. (3) Disconnect between theoretical research and practice: The workflow engine scheme proposed by scholars in journals such as "Archival Science Communications" has optimized the approval efficiency, but it has not combined the confidentiality classification characteristics of hydrological archives (such as the differentiated needs of top secret, secret and public data), resulting in limited security and applicability.

[0004] Foreign research is guided by "long-term preservation" and "open sharing" and is relatively mature, but it is also difficult to directly apply it to the field of hydrogeology: (1) The framework is too general: The OAIS standard of the International Council on Archives (ICA) provides a full life cycle management framework, but does not consider the dynamic updating characteristics of hydrological archives (such as phased supplementation of borehole data and real-time updates of monitoring data), and lacks a flexible management mechanism. (2) The system functions focus on data sharing: The NWIS system of the United States Geological Survey (USGS) realizes centralized storage of national hydrological data, but its functions are limited to simple retrieval and open sharing. It does not have the control functions such as borrowing approval, and cannot meet the needs of classified archive management. (3) The accuracy of technology application is insufficient: The EU's "Smart Archives" project introduces natural language processing (NLP) technology to improve retrieval capabilities, but does not conduct corpus training for hydrological professional terms (such as "transient electromagnetic geophysical data"), and the retrieval accuracy is less than 65%, which is not practical.

[0005] Based on the current situation at home and abroad, there are significant gaps in the management of digital archives: (1) Lack of full life cycle control: The existing system is independent of each link (collection, storage, retrieval, borrowing, archiving, and destruction), and cannot achieve status traceability and data linkage. (2) Insufficient industry-specific intelligent functions: The retrieval mechanism lacks multi-dimensional correlation capabilities (such as project-drilling-water disaster linkage) and has poor compatibility with unstructured data (engineering drawings, geophysical maps). (3) Insufficient balance between security and efficiency: The borrowing process does not combine classified classification design with differentiated approval, and intelligent push technology is lacking, resulting in low resource utilization.

[0006] Based on the fact that industry needs have shifted from "simple digitization" to "full lifecycle digitalization," there is an urgent need to develop a digital system for hydrogeological archives that integrates full-process management of archives, industry-adaptive retrieval, and flexible approval mechanisms to support the digital transformation and high-quality development of hydrological work. Summary of the Invention

[0007] To address the aforementioned issues, this invention provides a hydrogeological digital archive system with full lifecycle management. It addresses problems in hydrogeological archive management such as fragmented management, low retrieval accuracy, rigid borrowing processes, and low resource utilization by constructing a digital management system covering the entire lifecycle of "collection-storage-retrieval-borrowing-archiving-destruction." This improves archive retrieval efficiency, shortens borrowing approval cycles, enhances archive resource utilization, and enables efficient decision-making and data sharing in hydrogeological work.

[0008] To achieve the above objectives, the present invention provides a hydrogeological digital archive system with full life cycle management, including: a system architecture module, a full life cycle management module, a multimodal retrieval module, a flexible borrowing approval module, and an intelligent recommendation module; The system architecture module adopts a three-tier B / S architecture, including a user interface presentation layer, a business logic layer, and a data access layer, for: The interface presentation layer provides a responsive interface that supports access from both PC and mobile devices. It integrates an archive retrieval center, borrowing management, personal center, and system management portal, and includes a dedicated project archive page for the hydrological industry. The business logic layer includes a full lifecycle management engine, a multimodal retrieval engine, a flexible approval engine, and an intelligent recommendation engine, which are used to handle the core logic of document management; The data access layer adopts a hybrid storage mode of relational database and non-relational database. The relational database stores archive metadata, and the non-relational database stores unstructured data. Data operations are implemented through a unified data access interface. The full lifecycle management module is used for: The system manages the entire lifecycle of hydrogeological archives, from collection, storage, retrieval, borrowing, archiving to destruction. It achieves automatic tracking and data linkage of archive status through state transition functions. The archive lifecycle status includes unarchived, in storage, borrowed, overdue, destruction pending review, and destroyed. State transitions are triggered by events. The multimodal retrieval module is used for: Based on metadata, semantic text, and relevance dimensions, the document retrieval module employs a multi-layered retrieval model, including a metadata retrieval layer, a deep text retrieval layer, and a relevance dimension retrieval layer, and sorts the retrieval results using a comprehensive scoring algorithm. The flexible borrowing approval module is used for: The approval process is designed based on the classification of classified documents, which includes open, secret and confidential levels. The approval time is dynamically calculated based on the classification of the documents and the approval level. The intelligent recommendation module is used for: Based on users' historical borrowing behavior and project relevance, a collaborative filtering algorithm is used to recommend archival resources.

[0009] As a further improvement of the present invention, the state transition function in the full lifecycle management module is: in, S i The current status of the archive. E j This is an event triggered by the transfer of file status. S k The target state; The events that trigger the transfer of the archive status include: archive filing, initiating borrowing, returning archive, applying for destruction, approval, and rejection. The system automatically records state transition logs, including trigger time, operator, and related events, enabling full lifecycle traceability.

[0010] As a further improvement of the present invention, the metadata retrieval layer of the multimodal retrieval module performs a combined retrieval based on the fields of file type, project name, borehole number, and formation time, and the retrieval score formula is as follows: in, For the number of fields to be searched, For field weights, For fields With query terms The degree of matching; The deep text retrieval layer uses a BERT-based hydrological corpus model for semantic matching, and the retrieval score formula is as follows: in, and These are the semantic vectors for the query term and the text segment, respectively. A valid match is determined when Score2 ≥ 0.6. The correlation dimension retrieval layer establishes a correlation map of project-borehole-water hazard data, and the correlation degree formula is: The final search results are ranked by overall score as follows: The scores are sorted in descending order.

[0011] As a further improvement of the present invention, the formula for predicting the approval time of the flexible borrowing approval module is as follows: in, The basic approval time is 1 hour for public files, 4 hours for secret files, and 24 hours for confidential files. α is the approval level coefficient, and L is the approval level. Public files are subject to single-level approval, classified files to two-level approval, and confidential files to three-level approval. The system displays the approval progress in real time and automatically sends reminders before the borrowing period expires.

[0012] As a further improvement of the present invention, the recommendation similarity formula of the intelligent recommendation module is as follows: in, For the current user, For recommended files, For users Collection of historical borrowing archives For users For the archives borrowing frequency For archives and The degree of correlation; The intelligent recommendation module calculates the correlation degree based on the user's borrowing records, combined with the project, region, and data type.

[0013] As a further improvement of the present invention, it also includes an archive collection module, which supports online uploading, format verification and automatic archiving of archives. Uploaded files are automatically verified for format, including PDF, JPEG, TIFF and MP4. After verification, a unique number is automatically assigned to the file based on the full lifecycle management module, and the file enters the in-stock status.

[0014] As a further improvement of the present invention, the format of the unique number is: XM-Year-Item Number-Archive Type, where XM represents the item prefix, the year is the four-digit year in which the archive was created, the item number is the unique identifier of the item, and the archive type is a predefined archive classification code used to distinguish archive categories; The unique number is automatically generated when the file is uploaded and is used for full lifecycle status tracking and retrieval association.

[0015] As a further improvement of the present invention, the system hardware environment includes a server and a client; The server is configured with at least 16 CPU cores and 64GB of memory. The storage uses a hybrid architecture of SSD and HDD, with SSD used for database files and HDD used for archive storage. The client supports both PC and mobile devices. The PC version requires an Intel Core i5 or higher CPU and 8GB of RAM, while the mobile version supports Android 10.0 or higher or iOS 14.0 or higher. The network environment requires a server bandwidth of no less than 100Mbps and a network latency of less than 50ms.

[0016] As a further improvement of the present invention, the system software environment includes: a Linux operating system on the server side, and Windows, macOS, Android, and iOS systems on the client side; The database uses MySQL relational database and MongoDB non-relational database, and implements a unified data access interface through the Hibernate framework; The development framework includes a front-end responsive framework and a back-end lightweight framework, while the middleware includes a web server and a caching middleware.

[0017] As a further improvement of the present invention, the system interface presentation layer provides a temporary folder function, which allows users to add files to the temporary folder in batches and initiate borrowing, with the interaction achieved through JavaScript code; The data access interface of the business logic layer includes methods for querying archives by project and archiving storage.

[0018] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention addresses the technical problems of existing hydrogeological archive management systems, namely fragmented management, low retrieval accuracy, rigid borrowing processes, and low resource utilization. By constructing a full lifecycle management model for hydrogeological archives, it achieves data connectivity and status traceability across all stages, from archive formation, archiving, retrieval, borrowing, return, to destruction. Through the design of a multimodal retrieval algorithm specifically for the hydrogeological industry, it supports in-depth retrieval based on metadata, text, and related dimensions, improving retrieval accuracy and efficiency. By establishing a flexible borrowing approval process based on classification levels, it achieves differentiated review and real-time status tracking for archives of different classification levels (public / secret / confidential). Finally, by developing an intelligent recommendation module based on user behavior, it enhances the utilization rate of hydrogeological archive resources and their cross-project sharing capabilities. Attached Figure Description

[0019] Figure 1 This is an overall architecture diagram of a hydrogeological digital archive system for full life-cycle management, as disclosed in one embodiment of the present invention. Figure 2This is an example diagram of the functional interface of a hydrogeological digital archive system with full - life - cycle management and control disclosed in an embodiment of the present invention. Detailed implementation manners

[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0021] The following further describes the present invention in detail with reference to the accompanying drawings: As Figure 1 、 2 shown, a hydrogeological digital archive system with full - life - cycle management and control provided by the present invention includes: a system architecture module, a full - life - cycle management and control module, a multi - modal retrieval module, a flexible borrowing approval module, an intelligent recommendation module, an archive collection module, an intelligent retrieval module, and a borrowing management module; As Figure 1 shown, the system architecture module adopts a three - layer B / S architecture, including a presentation layer, a business logic layer, and a data access layer, and is used for: The presentation layer provides a responsive interface, supports access from both the PC side and the mobile side, integrates an archive retrieval center, borrowing management, a personal center, and a system management entry, and includes a project archive special page dedicated to the hydrogeological industry; specifically, the core interaction code of the presentation layer includes: <!--Entry of the project archive special page and trigger function of the temporary storage folder--> <button class="btn btn-primary" data-bs-toggle="dropdown"> Project Topic< / button> Jinjie Coal Mine Water Control Project <script>function addToTemp(id){$.post(" / temp / add",{fileId:id},res=>alert("已加入暂存夹"));}< / script> The business logic layer contains a full - life - cycle management and control engine, a multi - modal retrieval engine, a flexible approval engine, and an intelligent recommendation engine, and is used to process the core logic of archive management; The data access layer adopts a mixed storage mode of a relational database and a non - relational database. As Figure 2 shown, the relational database (MySQL) stores archive metadata, such as project name, transferor, and confidentiality level, and the non - relational database stores unstructured data, such as attached drawings and text PDFs. A unified data access interface is implemented through the Hibernate framework to achieve data operations. Specifically, the code of the unified data access interface includes: / / Unified interface of the data access layer (implemented by Hibernate) Public interface ArchiveDao{ @Query("from Archive where projectName=?1") List <archive>getByProject(String projectName); / / Search for archives by project Void save(Archive archive); / / Archive storage } Full lifecycle management module, such as Figure 2 As shown, it is used for: The system manages the entire lifecycle of hydrogeological archives, from collection, storage, retrieval, borrowing, archiving to destruction. It achieves automatic tracking and data linkage of archive status through state transition functions. The archive lifecycle status includes unarchived, in storage, borrowed, overdue, destruction pending review, and destroyed. State transitions are triggered by events. in, Define the "state-event" mapping relationship for the entire lifecycle of an archive, and implement automatic triggering and data traceability at each stage through state transition functions: Let the set of archive lifecycle states be: Archives formed (not archived); Archive storage (in the database); Currently borrowed; Overdue payment; Archived and destroyed pending review; Destroyed; The set of file status transition events is as follows: (2) in: Archives; Initiate borrowing; Return the archives; Request for destruction; Approved (Destroyed); : Review rejected (destroyed); Furthermore, The state transition function in the full lifecycle management module is: in, S i The current status of the archive. E j This is an event triggered by the transfer of file status. S k The target state; The events that trigger a change in the status of an archive include: archiving, initiating a borrowing process, returning an archive, applying for destruction, approval, and rejection. The system automatically records state transition logs, including trigger time, operator, and related events, enabling full lifecycle traceability.

[0022] Specifically, for example: (After an inquiry for a file in the archive is made, it will be changed to "in borrowing"); (The borrowed archives will be transferred to "in storage" upon return). (The status will change to "destroyed" after the destruction of pending files is approved).

[0023] This function allows the system to automatically record archive status transition logs (including trigger time, operator, and related events), enabling full lifecycle traceability.

[0024] The multimodal retrieval module is used for: Based on metadata, semantic text, and relevance dimensions, a three-layer retrieval model is designed for file retrieval to improve retrieval accuracy. The multimodal retrieval module adopts a multi-layer retrieval model, including a metadata retrieval layer, a deep text retrieval layer, and a relevance dimension retrieval layer, and sorts the retrieval results through a comprehensive scoring algorithm. The metadata retrieval layer of the multimodal retrieval module performs combined searches based on the fields of archive type, project name, borehole number, and formation time. It supports exact matching (e.g., "archive type = project, project name = Jinjie Coal Mine") and fuzzy matching (e.g., "borehole number contains ZK-2023"). The retrieval scoring formula is as follows: in, For the number of fields to retrieve (e.g.) =4, corresponding to file type, project name, borehole number, and creation time). For field weights (e.g., the weight of "borehole number" in the hydrological industry) =0.4, "Project Name" weight =0.3), For fields F i With query terms Q i Matching degree (0≤ ≤1); The deep retrieval layer of the main text uses a BERT-based hydrological corpus model for semantic matching, and the retrieval score formula is as follows: in, and These are the semantic vectors for the query term and the text segment, respectively. Cosine similarity (0≤ ≤1), a valid match is determined only if Score2≥0.6; The correlation dimension retrieval layer establishes a data correlation map between projects, boreholes, and water hazards. During retrieval, it automatically associates relevant resources from the query archives (e.g., when querying "Jinjie Coal Mine Exploration Plan," it automatically associates the project's borehole data and basic water hazard data). The correlation degree formula is: The final search results are ranked by overall score as follows: The scores are sorted in descending order to ensure optimal accuracy and relevance.

[0025] The flexible borrowing approval module is used for: The approval process is designed based on the classification of classified documents, which includes open, secret and confidential levels. The approval time is dynamically calculated based on the classification of the documents and the approval level. in, The formula for predicting the approval time of the flexible borrowing approval module is as follows: in, Based on the basic approval time, public files take 1 hour, confidential files take 4 hours, and classified files take 24 hours. α is the approval level coefficient (single level). =0, two levels =0.5, three levels =1.0), L is the approval level; Public archives adopt a single-level approval process. =1, classified files are subject to a two-tier approval process. =2, Confidential files are subject to a three-tier approval process. =3; The system displays the approval progress in real time and automatically sends reminders before the borrowing period expires.

[0026] Specifically, For example: publicly available archives (such as hydrological industry standards and specifications): Hourly (single-level approval, department administrator review); confidential files (such as mine flooding emergency plans): (Three-level approval process: department administrator → hydrology bureau archives administrator → chief engineer's office review)

[0027] The intelligent recommendation module is used for: Based on users' historical borrowing behavior (such as borrowing "North China Coal Mine Water Hazard Data" 5 times in the past 3 months) and the relevance of the project, a collaborative filtering algorithm is used to recommend archival resources.

[0028] The recommendation similarity formula for the intelligent recommendation module is as follows: in, For the current user, For recommended files, For users Collection of historical borrowing archives For users For the archives borrowing frequency For archives and The degree of correlation; The intelligent recommendation module calculates the relevance based on the user's borrowing history, combined with the project, region, and data type.

[0029] Archive collection module, such as Figure 2 As shown, it is used for: Supports online file uploading, format verification, and automatic archiving. Uploaded files are automatically verified for format, including PDF, JPEG, TIFF, and MP4. After verification, the file automatically enters the full lifecycle management module. The file is in the "in-stock" status and a unique number is automatically assigned to it.

[0030] in, The format of the unique number is: XM-Year-Project Number-Archive Type. XM represents the project prefix, the year is the four-digit year in which the archive was created, the project number is the unique identifier of the project, and the archive type is a predefined archive classification code used to distinguish archive categories, such as: XM-2023-JJ-001, where JJ represents "Mine Hydrogeological Report"; A unique number is automatically generated when the file is uploaded and is used for full lifecycle status tracking and retrieval association.

[0031] The intelligent search module is used for: The system integrates "metadata retrieval - text retrieval - related retrieval" functions. After the user enters a query term (such as "Jinjie Coal Mine transient electromagnetic method"), the system automatically calculates the comprehensive score and highlights the matching fragments (metadata matching fields are highlighted in yellow, and text matching content is highlighted in red). The borrowing management module is used for: After selecting files, users can add them to a "temporary folder" to initiate batch borrowing. The system automatically identifies the file's security level and assigns an approval process, displaying the approval progress in real time (e.g., "Department administrator is reviewing → 1 hour remaining"). A reminder is automatically sent 24 hours before the borrowing period expires. This invention provides a hydrogeological digital archive system with full lifecycle management, the hardware environment of which includes a server and a client. The server is configured with at least 16 CPU cores and 64GB of memory. The storage uses a hybrid architecture of SSD and HDD, with SSD used for database files and HDD used for archive storage. The client supports both PC and mobile devices. The PC version requires an Intel Core i5 or higher CPU and 8GB of RAM, while the mobile version supports Android 10.0 or higher or iOS 14.0 or higher. The network environment requires a server bandwidth of no less than 100Mbps and a network latency of less than 50ms.

[0032] Specifically, To meet the storage, access, and response requirements throughout the entire lifecycle management of archives, the server, as the core supporting device, must be equipped with an Intel Xeon Gold 6330 or higher CPU with no fewer than 16 cores to ensure efficient processing when multiple users concurrently initiate retrieval and borrowing requests; the memory configuration must be no less than 64GB DDR4 to ensure smooth loading and processing of large-scale hydrogeological archive data and avoid system lag due to insufficient memory; the storage adopts an "SSD+HDD" combined architecture, where SSDs with a capacity of 2TB or more are used to store database files to meet the fast read and write requirements of metadata and frequently accessed archives, while HDDs with a capacity of 10TB or more are used for archival storage of various hydrogeological archive resources, balancing storage capacity and cost control.

[0033] The client hardware needs to be adapted to various usage scenarios. PC clients must be equipped with an Intel Core i5 or higher CPU, 8GB or more of RAM, and a screen resolution of at least 1920×1080 to ensure complete interface display and smooth operation. Mobile clients must support Android 10.0 or higher or iOS 14.0 or higher, with at least 4GB of RAM to facilitate user access to the system in mobile scenarios such as field surveys. Regarding the network environment, the server bandwidth must reach 100Mbps or higher to ensure stability when multiple users simultaneously download large files (such as engineering drawings and survey reports). The network latency between the client and server must be controlled within 50ms to ensure fast loading of search results and timely submission of borrowing requests, meeting the timeliness requirements in emergency scenarios.

[0034] The system software environment includes: a Linux operating system on the server side, and Windows, macOS, Android, and iOS systems on the client side; The database uses MySQL relational database and MongoDB non-relational database, and implements a unified data access interface through the Hibernate framework; The development framework includes a front-end responsive framework and a back-end lightweight framework, while the middleware includes a web server and a caching middleware.

[0035] Specifically, At the operating system level, the server uses a stable and compatible Linux system version to support the long-term operation of backend services and middleware; the client covers mainstream systems such as Windows 10 / 11, macOS 12.0 and above, Android 10.0 and above, and iOS 14.0 and above, adapting to the usage habits of different users. The database adopts a hybrid storage mode: a relational database is used to store archive metadata (such as project information, security classification, and handover records), supporting complex multi-field queries; a non-relational database is used to store unstructured data such as attached images and PDF text, adapting to diverse archive formats and dynamic update requirements.

[0036] In terms of development framework, the front-end uses a responsive framework and interaction library to achieve multi-terminal interface adaptation, ensuring intuitive operation of functions such as retrieval and borrowing. The back-end uses a lightweight framework to encapsulate business logic, and a data access framework to decouple from the database, simplifying the CRUD operations of archive data, while supporting flexible switching between multiple data sources. Regarding middleware, the web server is responsible for receiving and processing client requests, ensuring stable service operation; the caching middleware is used to store frequently accessed data (such as popular archive lists and recent borrowing records), reducing the number of database queries and improving retrieval response speed. The algorithm dependency library provides support for intelligent functions, training a text retrieval model adapted to hydrological terminology using a professional language model, and using Chinese word segmentation tools to achieve accurate text processing, ensuring the efficient operation of multimodal retrieval algorithms.

[0037] The system interface layer of this invention provides a temporary folder function, which allows users to add files to the temporary folder in batches and initiate borrowing, with the interaction implemented through JavaScript code; The data access interface of the business logic layer includes methods for querying archives by project and archiving storage.

[0038] Example: The system of this invention has been deployed and verified in the "Digital Archives of Hydrogeology" of a hydrogeological survey unit. The actual application effect of the technical solution is intuitively demonstrated through the system interface and data statistics module, including: Archive Collection and Digitization Achievements: The system visualizes the quantity and digitization progress of all types of hydrogeological archives through the "Collection Statistics" module. The bar chart shows that project-related archives contain 14,895 items, with 10,970 digitized; borehole-related archives contain 3,204 items, with 3,200 digitized, achieving a digitization rate exceeding 99%; and hydrochemical archives contain 2,161 items, all digitized, achieving a 100% digitization rate. Combined with the digitization statistics table, the total number of archives in the entire database is 25,049, with 20,210 digitized, achieving an overall digitization rate of 80%. Notably, archives such as standards and specifications, laws and regulations, and geological manuals all have a 100% digitization rate, fully validating the system's efficient management capabilities across the entire "archive collection-digitization-storage" process.

[0039] File Types and Resource Distribution: The statistics on file types across the entire database and the distribution of archive types visually present the diversity of formats and the correlation between types in hydrogeological archives. The system supports the storage and retrieval of mainstream file formats such as PDF, JPG, WORD, and EXCEL, including 6,371 PDF files and 26,288 JPG files, covering all types of hydrogeological archives such as project reports, borehole maps, and water hazard data tables. From the perspective of archive type, project-related archives have the most associated files (22,876 files), while borehole and hydrochemical archives also have a significant number of files, demonstrating the system's ability to integrate multi-dimensional related resources of "project-borehole-hydrochemical data," providing a data foundation for subsequent multi-modal retrieval and intelligent recommendation.

[0040] System utilization efficiency and business value: Statistical data shows that since its launch, the system has been used by 181 people, with 1311 usage visits, 568 queries, 35 borrowing requests, and 43 borrowed documents. This data indicates that the system, through its "intelligent retrieval-flexible borrowing" process, effectively improves the utilization efficiency of hydrogeological archives and solves the pain points of "difficult retrieval and slow approval" in the traditional model. This invention has technical advantages and practical business value in the whole life cycle management, multimodal retrieval, and flexible borrowing of hydrogeological archives, and can effectively support the archive management needs of scenarios such as hydrological exploration and coal mine water control projects.

[0041] Advantages of this invention: This invention covers the entire process of archives from collection, storage, retrieval, borrowing to destruction through a "state-event" mapping model, achieving full lifecycle management. The status of archives is traceable throughout the process, and data at each stage is automatically linked, avoiding version confusion and duplicate storage. The status log records the operation time, personnel and events, supports reverse traceability, and improves management transparency.

[0042] This invention features a three-tier B / S architecture that supports multi-terminal adaptation, improving system accessibility. The responsive interface allows users to access the system conveniently in mobile scenarios such as field surveys. The business logic layer encapsulates core algorithms to ensure functional stability. The data layer uses hybrid storage (relational + non-relational databases) to optimize data processing efficiency.

[0043] This invention optimizes the hardware environment to ensure smooth operation in high-concurrency scenarios, with no lag during concurrent retrieval and borrowing by multiple users; network latency is controlled within 50ms to meet the real-time response requirements of emergency scenarios (such as mine water inrush).

[0044] This invention's multimodal retrieval algorithm significantly improves retrieval accuracy and efficiency, the flexible borrowing approval process shortens the approval cycle while balancing security and efficiency, the intelligent recommendation module improves the utilization rate of archival resources, and the standardization of archival numbering enhances manageability.

[0045] The deep retrieval layer of this invention uses the BERT model, which greatly improves the accuracy of matching professional terms. Users can perform natural language queries, and the results are highlighted, enabling automatic recommendation of related resources through related dimension retrieval.

[0046] This invention digitizes the borrowing process, reducing operational complexity, and has strong system compatibility, supporting the management of multiple file formats.

[0047] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.< / archive>

Claims

1. A hydrogeological digital archive system for full life-cycle management, characterized in that, include: The system includes modules for system architecture, full lifecycle management, multimodal retrieval, flexible borrowing approval, and intelligent recommendation. The system architecture module adopts a three-tier B / S architecture, including a user interface presentation layer, a business logic layer, and a data access layer, for: The interface presentation layer provides a responsive interface that supports access from both PC and mobile devices. It integrates an archive retrieval center, borrowing management, personal center, and system management portal, and includes a dedicated project archive page for the hydrological industry. The business logic layer includes a full lifecycle management engine, a multimodal retrieval engine, a flexible approval engine, and an intelligent recommendation engine, which are used to handle the core logic of document management; The data access layer adopts a hybrid storage mode of relational database and non-relational database. The relational database stores archive metadata, and the non-relational database stores unstructured data. Data operations are implemented through a unified data access interface. The full lifecycle management module is used for: The system manages the entire lifecycle of hydrogeological archives, from collection, storage, retrieval, borrowing, archiving to destruction. It achieves automatic tracking and data linkage of archive status through state transition functions. The archive lifecycle status includes unarchived, in storage, borrowed, overdue, destruction pending review, and destroyed. State transitions are triggered by events. The multimodal retrieval module is used for: Based on metadata, semantic text, and relevance dimensions, the document retrieval module employs a multi-layered retrieval model, including a metadata retrieval layer, a deep text retrieval layer, and a relevance dimension retrieval layer, and sorts the retrieval results using a comprehensive scoring algorithm. The flexible borrowing approval module is used for: The approval process is designed based on the classification of classified documents, which includes open, secret and confidential levels. The approval time is dynamically calculated based on the classification of the documents and the approval level. The intelligent recommendation module is used for: Based on users' historical borrowing behavior and project relevance, a collaborative filtering algorithm is used to recommend archival resources.

2. The hydrogeological digital archive system for full life-cycle management as described in claim 1, characterized in that, The state transition function in the full lifecycle management module is: in, S i The current status of the archive. E j This is an event triggered by the transfer of file status. S k The target state; The events that trigger the transfer of the archive status include: archive filing, initiating borrowing, returning archive, applying for destruction, approval, and rejection. The system automatically records state transition logs, including trigger time, operator, and related events, enabling full lifecycle traceability.

3. The hydrogeological digital archive system for full life-cycle management as described in claim 1, characterized in that: The metadata retrieval layer of the multimodal retrieval module performs a combined retrieval based on the fields of file type, project name, borehole number, and formation time. The retrieval score formula is as follows: in, For the number of fields to be searched, For field weights, For fields With query terms The degree of matching; The deep text retrieval layer uses a BERT-based hydrological corpus model for semantic matching, and the retrieval score formula is as follows: in, and These are the semantic vectors for the query term and the text segment, respectively. A valid match is determined when Score2 ≥ 0.

6. The correlation dimension retrieval layer establishes a correlation map of project-borehole-water hazard data, and the correlation degree formula is: The final search results are ranked by overall score as follows: The scores are sorted in descending order.

4. The hydrogeological digital archive system for full life-cycle management as described in claim 1, characterized in that: The formula for predicting the approval time of the flexible borrowing approval module is as follows: in, The basic approval time is 1 hour for public files, 4 hours for secret files, and 24 hours for confidential files. α is the approval level coefficient, and L is the approval level. Public files are subject to single-level approval, classified files to two-level approval, and confidential files to three-level approval. The system displays the approval progress in real time and automatically sends reminders before the borrowing period expires.

5. The hydrogeological digital archive system for full life-cycle management as described in claim 1, characterized in that: The recommendation similarity formula of the intelligent recommendation module is: in, For the current user, For recommended files, For users Collection of historical borrowing archives For users For the archives borrowing frequency For archives and The degree of correlation; The intelligent recommendation module calculates the correlation degree based on the user's borrowing records, combined with the project, region, and data type.

6. The hydrogeological digital archive system for full life-cycle management as described in claim 1, characterized in that: It also includes an archive collection module, which supports online uploading, format verification, and automatic archiving of archives. Uploaded files are automatically verified for format, including PDF, JPEG, TIFF, and MP4. After verification, a unique number is automatically assigned to the file based on the full lifecycle management module, and the file enters the in-stock status.

7. The hydrogeological digital archive system for full life-cycle management as described in claim 6, characterized in that, The format of the unique number is: XM-Year-Item Number-Archive Type, where XM represents the item prefix, the year is the four-digit year in which the archive was created, the item number is the unique identifier of the item, and the archive type is a predefined archive classification code used to distinguish archive categories; The unique number is automatically generated when the file is uploaded and is used for full lifecycle status tracking and retrieval association.

8. The hydrogeological digital archive system for full life-cycle management as described in claim 1, characterized in that: The system hardware environment includes servers and clients; The server is configured with at least 16 CPU cores and 64GB of memory. The storage uses a hybrid architecture of SSD and HDD, with SSD used for database files and HDD used for archive storage. The client supports both PC and mobile devices. The PC version requires an Intel Core i5 or higher CPU and 8GB of RAM, while the mobile version supports Android 10.0 or higher or iOS 14.0 or higher. The network environment requires a server bandwidth of no less than 100Mbps and a network latency of less than 50ms.

9. The hydrogeological digital archive system for full life-cycle management as described in claim 1, characterized in that, The system software environment includes: a Linux operating system on the server side, and Windows, macOS, Android, and iOS systems on the client side; The database uses MySQL relational database and MongoDB non-relational database, and implements a unified data access interface through the Hibernate framework; The development framework includes a front-end responsive framework and a back-end lightweight framework, while the middleware includes a web server and a caching middleware.

10. The hydrogeological digital archive system for full life-cycle management as described in claim 1, characterized in that: The system interface presentation layer provides a temporary folder function, which allows users to add files to the temporary folder in batches and initiate borrowing, with the interaction implemented through JavaScript code; The data access interface of the business logic layer includes methods for querying archives by project and archiving storage.