Archive management system based on cloud archive library
By designing an archive management system based on cloud archives, using deep learning, reinforcement learning and blockchain technology, the shortcomings of existing systems in heterogeneous data retrieval, rule management and multi-terminal access security are solved, and efficient unified retrieval, dynamic rule generation and secure and reliable multi-terminal access are achieved.
Patent Information
- Application Number
- CN202510093419.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
Smart Images

Figure CN120011617A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cloud archive management, and in particular to an archive management system based on a cloud archive library. Background Art
[0002] With the rapid development of cloud computing and artificial intelligence technologies, archive management systems have been widely used in judicial, medical, and corporate scenarios. However, existing archive management systems have shortcomings in the following aspects:
[0003] 1. The problem of unified retrieval of heterogeneous data: Archival data is usually stored in multiple formats such as text, images, audio and video. Existing systems usually use separate format-specific retrieval algorithms, which leads to low efficiency of cross-format retrieval, inconsistent results, and inability to meet users' precise query needs.
[0004] 2. Lifecycle management based on static rules: Traditional archive management relies on predefined static rules, such as manually set policies based on fixed storage time limits or access frequencies. It is difficult to dynamically adapt to changes in user behavior and fluctuations in resource load, resulting in low resource utilization.
[0005] 3. Security risks of multi-terminal access: With the increasing demand for distributed storage and multi-terminal collaboration, existing technologies mostly control file access through centralized permission management. However, this method is susceptible to single point failures or permission conflicts, making it difficult to ensure the security and consistency of data access.
[0006] In order to solve the above problems, the industry has proposed some optimization measures. For example, using deep learning technology to improve heterogeneous data retrieval, using basic rule engines to enhance lifecycle management, or improving permission control through blockchain technology. However, these measures still have insufficient semantic understanding of heterogeneous data retrieval in actual applications, and cannot achieve unified semantic representation of data in different formats; lifecycle management lacks dynamic adaptation capabilities, and it is difficult to balance the resource allocation of high-frequency access archives and low-frequency archives; blockchain technology generally lacks effective integration with distributed consistency protocols in permission control, resulting in the inability to guarantee data consistency in multi-terminal scenarios.
[0007] Therefore, how to efficiently realize unified retrieval of heterogeneous data, dynamically generate archive management rules, and enhance the security and consistency of multi-terminal access has become a technical problem to be solved by the present invention. Summary of the invention
[0008] The present invention provides an archive management system, system and computer-readable storage medium based on a cloud archive library, the main purpose of which is to solve the problems of difficult retrieval of heterogeneous data, inflexible rule management and poor security of multi-terminal access.
[0009] To achieve the above-mentioned purpose, the present invention provides an archive management system based on a cloud archive library, the system comprising a data acquisition module, a data preprocessing module, a cross-format semantic retrieval module, an intelligent rule generation module, a distributed authority control module and a dynamic resource optimization module;
[0010] Data collection module: used to collect archival data in heterogeneous formats from multiple data sources in real time, including text data, image data, and audio / video data; the collected data is subjected to preliminary feature extraction through edge computing devices, and the extracted features include text semantic features, image visual features, and audio waveform features; the preliminary features are synchronously uploaded to the cloud storage unit with the original archival data;
[0011] Data preprocessing module: used to perform normalization processing on the uploaded heterogeneous data, including semantic segmentation and vectorization of text, compression and embedding coding of image features, and spectrum analysis and embedding generation of audio data; calculate multimodal semantic embedding vectors based on deep learning models, and build a cross-format unified semantic vector space, where the generation formula of the semantic vector is:
[0012] V f =W t ·V t +W i ·V i +W a ·V a
[0013] Among them, V f is the multimodal semantic vector, V t 、V i 、V a are text, image and audio feature vectors respectively, and W t , W i , W a The weight coefficient is generated based on historical retrieval data and usage scenario fitting;
[0014] Cross-format semantic retrieval module: used to receive the user's natural language retrieval request, parse the request into a unified semantic vector, and realize unified retrieval of cross-format archives by calculating the similarity between the user request semantic vector and the archive data semantic vector; the similarity calculation is based on the cosine similarity formula:
[0015]
[0016] Where Q is the semantic vector of the user's request, D is the semantic vector of the archive data, and Sim(Q,D) is the similarity score. The search results are sorted from high to low according to the similarity score and returned.
[0017] Intelligent rule generation module: used to dynamically generate archive lifecycle management rules based on the user's operation history and the frequency of archive access; rule generation uses a reinforcement learning algorithm, and the objective function is the weighted sum of maximizing the archive access frequency and minimizing the storage resource consumption. The objective function formula is:
[0018] R=α·F-β·C
[0019] Among them, R is the rule adaptation score, F is the file access frequency, C is the storage resource consumption, α and β are dynamically adjusted weight parameters, which are adaptively adjusted by the algorithm according to the user behavior pattern;
[0020] Distributed permission control module: Based on blockchain technology, a distributed permission management mechanism is built to perform fine-grained control over the user's access to archives. The blockchain network records the permission request and access record for each archive access, and verifies the user's permission through a zero-knowledge proof algorithm. Distributed consistency protocols such as PBFT are used to ensure data consistency and prevent permission conflicts in multi-terminal access scenarios.
[0021] Dynamic resource optimization module: used to dynamically adjust the allocation of storage and computing resources according to system load and access frequency; predict the access load in the next time period through the resource scheduling optimization algorithm and calculate the best resource allocation plan; the scheduling optimization formula is:
[0022]
[0023] Among them, L pred (t+1) is the predicted value of the access load in the next time period, L(t) is the current load, A i is the average load in the past n time periods, and γ is the smoothing coefficient.
[0024] Optionally, the data acquisition module includes an edge computing unit and a data upload unit, wherein: the edge computing unit performs real-time feature extraction on the collected heterogeneous data and generates preliminary feature data; the data upload unit synchronously transmits the preliminary feature data and the corresponding original data to a cloud storage unit, and marks the timestamp and data source identifier to support subsequent retrieval and permission management.
[0025] Optionally, the data preprocessing module further includes: a data cleaning module for removing noise, redundant data and invalid fields in the uploaded data; a feature fusion module for calculating weight coefficients of different modal data based on the self-attention mechanism to adjust the generation of multimodal semantic vectors, wherein the weight coefficient is calculated by the following formula:
[0026]
[0027] Among them, W xis the weight coefficient of mode x, S x Score the importance of modality x in a specific context, where n is the total number of all modalities.
[0028] Optionally, the cross-format semantic retrieval module also includes a user intent recognition unit, which is specifically used to: construct a semantic vector based on the user's search request statement, and analyze the user's intent through a deep learning model; generate secondary retrieval conditions based on the intent analysis results, and further refine the cross-format semantic retrieval results to improve the accuracy of the retrieval results.
[0029] Optionally, the intelligent rule generation module further includes a user behavior analysis unit, which is used to: extract behavioral features based on the user's access history and operation behavior to the archives; and dynamically adjust the weight parameters of the objective function using a reinforcement learning algorithm, wherein the behavioral features include access frequency, operation duration, and interaction depth.
[0030] Optionally, the blockchain network in the distributed permission control module adopts a hierarchical structure, specifically including: a primary node for storing archive access records and permission change information; a secondary node for distributed permission verification and consistency maintenance; each node ensures the consistency and integrity of archive access records in multi-terminal scenarios through a distributed consistency protocol PBFT.
[0031] Optionally, the distributed permission control module adopts zero-knowledge proof technology to achieve the following functions: verifying user permissions without disclosing permission details; preventing permission conflicts caused by simultaneous access by multiple users, and ensuring the privacy of data access.
[0032] Optionally, the dynamic resource optimization module further includes a load prediction unit and a resource scheduling unit, wherein: the load prediction unit predicts the future access load based on the time series model; the resource scheduling unit dynamically adjusts the allocation of storage resources and computing resources according to the predicted access load, and the optimization formula is:
[0033]
[0034] Among them, R(t) is the resource scheduling result at time t, P i is the resource priority weight, A i is the demand for the i-th resource, and n is the total number of resource categories.
[0035] Optionally, the cross-format semantic retrieval module further includes a result ranking unit, which generates a comprehensive score Score by fusing the user preference data and the semantic similarity score. The calculation formula of the comprehensive score is:
[0036] Score=λSim(Q,D)+(1-λ)·Pref
[0037] Among them, Sim(Q,D) is the semantic similarity score between the user request and the archive data, Pref is the score based on the user's historical preference, and λ is the weight factor, which is dynamically adjusted according to the user's usage scenario.
[0038] Optionally, the data preprocessing module and the dynamic resource optimization module work together to give priority to processing archive data with high access frequency and high priority after data upload, wherein the priority is determined by the following formula:
[0039] Priority=ω 1 ·F+ω 2 ·U
[0040] Among them, Priority is the priority of archive data, F is the historical access frequency of the archive, U is the emergency use demand score of the archive, ω 1 and ω 2 is the priority weight parameter.
[0041] Compared with the problems described in the background technology, the beneficial effects of the present invention are:
[0042] 1. The cross-format semantic retrieval module overcomes the format difference barrier of heterogeneous archival data in retrieval by constructing a unified multimodal semantic vector space. The semantic embedding generation technology based on the deep learning model enables multimodal data such as text, images, and audio to be efficiently retrieved in a unified semantic vector space, thereby significantly improving the accuracy and efficiency of retrieval. In addition, the user intent recognition unit generates dynamic retrieval conditions by parsing natural language queries, effectively adapting to complex user needs and ensuring the flexibility and versatility of the system. This feature not only solves the core problem of low efficiency of heterogeneous data retrieval in the existing technology, but also further improves the level of intelligent retrieval.
[0043] 2. The intelligent rule generation module uses reinforcement learning algorithms, combined with user behavior analysis and archive access frequency, to dynamically adjust the archive lifecycle management rules. This module adaptively adjusts the weight parameters in the objective function so that the rule generation process can respond to user behavior changes and resource constraints in real time. This dynamic adaptation capability effectively improves the utilization efficiency of archive storage resources while ensuring priority access to high-frequency archives. Compared with static rule management methods, the intelligent rule generation module achieves accurate mapping and management of resources and behaviors, thus showing significant progress in the technical implementation path.
[0044] 3. The distributed permission control module is based on blockchain technology and realizes the decentralized storage and verification of archive access records and permission management. Through the zero-knowledge proof algorithm, the module can verify the validity of user access rights in multi-terminal scenarios while ensuring the privacy of access data. The distributed consistency protocol further ensures the consistency of archive access records between different nodes and prevents permission conflicts caused by concurrent access by multiple terminals. This design significantly enhances the security and data consistency of the archive system and solves the problem of low efficiency of permission control in a multi-terminal access environment in the prior art.
[0045] 4. The dynamic resource optimization module realizes real-time adjustment of resource allocation through the collaborative work of the load prediction unit and the resource scheduling unit. The load prediction unit based on the time series model can make high-precision predictions on access load changes and provide optimization basis for the resource scheduling unit. The resource scheduling unit dynamically allocates storage and computing resources through the resource scheduling optimization algorithm, giving priority to high-access frequency and high-priority archive data, effectively reducing resource waste and improving the overall operation efficiency of the system. The introduction of this module enables the system to have strong dynamic adaptability and energy consumption optimization capabilities, meeting the needs of large-scale concurrent access scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 Schematic diagram of the technical architecture of the cloud archive management system of the present invention.
[0047] Figure 2 It is a schematic diagram of the retrieval and optimization process of the cloud archive management system of the present invention.
[0048] Figure 3 It is a modular functional block diagram of the cloud archive management system of the present invention.
[0049] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings in conjunction with the embodiments. DETAILED DESCRIPTION
[0050] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0051] The embodiment of the present application provides an archive management system based on a cloud archive library. The system includes a data acquisition module, a data preprocessing module, a cross-format semantic retrieval module, an intelligent rule generation module, a distributed authority control module and a dynamic resource optimization module;
[0052] Data collection module: used to collect archival data in heterogeneous formats from multiple data sources in real time, including text data, image data, and audio / video data; the collected data is subjected to preliminary feature extraction through edge computing devices, and the extracted features include text semantic features, image visual features, and audio waveform features; the preliminary features are synchronously uploaded to the cloud storage unit with the original archival data;
[0053] Data preprocessing module: used to perform normalization processing on the uploaded heterogeneous data, including semantic segmentation and vectorization of text, compression and embedding coding of image features, and spectrum analysis and embedding generation of audio data; calculate multimodal semantic embedding vectors based on deep learning models, and build a cross-format unified semantic vector space, where the generation formula of the semantic vector is:
[0054] V f =W t ·V t +W i ·V i +W a ·V a
[0055] Among them, V f is the multimodal semantic vector, V t 、V i 、V a are text, image and audio feature vectors respectively, and W t , W i , W a The weight coefficient is generated based on historical retrieval data and usage scenario fitting;
[0056] Cross-format semantic retrieval module: used to receive the user's natural language retrieval request, parse the request into a unified semantic vector, and realize unified retrieval of cross-format archives by calculating the similarity between the user request semantic vector and the archive data semantic vector; the similarity calculation is based on the cosine similarity formula:
[0057]
[0058] Where Q is the semantic vector of the user's request, D is the semantic vector of the archive data, and Sim(Q,D) is the similarity score. The search results are sorted from high to low according to the similarity score and returned.
[0059] Intelligent rule generation module: used to dynamically generate archive lifecycle management rules based on the user's operation history and the frequency of archive access; rule generation uses a reinforcement learning algorithm, and the objective function is the weighted sum of maximizing the archive access frequency and minimizing the storage resource consumption. The objective function formula is:
[0060] R=α·F-β·C
[0061] Among them, R is the rule adaptation score, F is the file access frequency, C is the storage resource consumption, α and β are dynamically adjusted weight parameters, which are adaptively adjusted by the algorithm according to the user behavior pattern;
[0062] Distributed permission control module: Based on blockchain technology, a distributed permission management mechanism is built to perform fine-grained control over the user's access to archives. The blockchain network records the permission request and access record for each archive access, and verifies the user's permission through a zero-knowledge proof algorithm. Distributed consistency protocols such as PBFT are used to ensure data consistency and prevent permission conflicts in multi-terminal access scenarios.
[0063] Dynamic resource optimization module: used to dynamically adjust the allocation of storage and computing resources according to system load and access frequency; predict the access load in the next time period through the resource scheduling optimization algorithm and calculate the best resource allocation plan; the scheduling optimization formula is:
[0064]
[0065] Among them, L pred (t+1) is the predicted value of the access load in the next time period, L(t) is the current load, A i is the average load in the past n time periods, and γ is the smoothing coefficient.
[0066] Optionally, the data acquisition module includes an edge computing unit and a data upload unit, wherein: the edge computing unit performs real-time feature extraction on the collected heterogeneous data and generates preliminary feature data; the data upload unit synchronously transmits the preliminary feature data and the corresponding original data to a cloud storage unit, and marks the timestamp and data source identifier to support subsequent retrieval and permission management.
[0067] Optionally, the data preprocessing module further includes: a data cleaning module for removing noise, redundant data and invalid fields in the uploaded data; a feature fusion module for calculating weight coefficients of different modal data based on the self-attention mechanism to adjust the generation of multimodal semantic vectors, wherein the weight coefficient is calculated by the following formula:
[0068]
[0069] Among them, W x is the weight coefficient of mode x, S x Score the importance of modality x in a specific context, where n is the total number of all modalities.
[0070] Optionally, the cross-format semantic retrieval module also includes a user intent recognition unit, which is specifically used to: construct a semantic vector based on the user's search request statement, and analyze the user's intent through a deep learning model; generate secondary retrieval conditions based on the intent analysis results, and further refine the cross-format semantic retrieval results to improve the accuracy of the retrieval results.
[0071] Optionally, the intelligent rule generation module further includes a user behavior analysis unit, which is used to: extract behavioral features based on the user's access history and operation behavior to the archives; and dynamically adjust the weight parameters of the objective function using a reinforcement learning algorithm, wherein the behavioral features include access frequency, operation duration, and interaction depth.
[0072] Optionally, the blockchain network in the distributed permission control module adopts a hierarchical structure, specifically including: a primary node for storing archive access records and permission change information; a secondary node for distributed permission verification and consistency maintenance; each node ensures the consistency and integrity of archive access records in multi-terminal scenarios through a distributed consistency protocol PBFT.
[0073] Optionally, the distributed permission control module adopts zero-knowledge proof technology to achieve the following functions: verifying user permissions without disclosing permission details; preventing permission conflicts caused by simultaneous access by multiple users, and ensuring the privacy of data access.
[0074] Optionally, the dynamic resource optimization module further includes a load prediction unit and a resource scheduling unit, wherein: the load prediction unit predicts the future access load based on the time series model; the resource scheduling unit dynamically adjusts the allocation of storage resources and computing resources according to the predicted access load, and the optimization formula is:
[0075]
[0076] Among them, R(t) is the resource scheduling result at time t, P i is the resource priority weight, A i is the demand for the i-th resource, and n is the total number of resource categories.
[0077] Optionally, the cross-format semantic retrieval module further includes a result ranking unit, which generates a comprehensive score Score by fusing the user preference data and the semantic similarity score. The calculation formula of the comprehensive score is:
[0078] Score=λ·Sim(Q,D)+(1-λ)·Pref
[0079] Among them, Sim(Q,D) is the semantic similarity score between the user request and the archive data, Pref is the score based on the user's historical preference, and λ is the weight factor, which is dynamically adjusted according to the user's usage scenario.
[0080] Optionally, the data preprocessing module and the dynamic resource optimization module work together to give priority to processing archive data with high access frequency and high priority after data upload, wherein the priority is determined by the following formula:
[0081] Priority=ω 1 ·F+ω 2 ·U
[0082] Among them, Priority is the priority of archive data, F is the historical access frequency of the archive, U is the emergency use demand score of the archive, ω 1 and ω 2 is the priority weight parameter.
[0083] Embodiment 1:
[0084] In judicial scenarios, the timeliness, accuracy, and security of archive management are crucial. For example, when dealing with complex cases, the judicial department needs to retrieve relevant evidence materials from different sources, which may include text records (such as court transcripts), image files (such as on-site photos), audio files (such as audio records), and video files (such as surveillance videos). The existing technology is difficult to meet the judicial department's requirements for efficiency and security of archive retrieval due to problems such as heterogeneous data formats, low retrieval efficiency, and insufficient security. To this end, this paper provides an innovative solution based on cross-format deep retrieval and dynamic intelligent management of cloud archive management system.
[0085] 1. Collect and upload data.
[0086] This system is deployed in the judicial archive management cloud platform. The front end includes terminal devices connected to courts, public security bureaus, law firms and other institutions. These devices collect archive data in different formats in real time through edge computing nodes.
[0087] a. Collection process: At the beginning of the case, the witness testimony is recorded through the connected audio recording equipment, and the case-related text files (such as trial records) and image evidence (such as crime scene photos) are simultaneously obtained.
[0088] b. Edge computing: Each edge computing node extracts features from the collected multimodal data. For example, audio files use fast Fourier transform (FFT) to extract spectral features, image files use convolutional neural network (CNN) to extract visual features, and text files use natural language processing (NLP) methods to perform word segmentation and semantic vectorization.
[0089] c. Upload and label: The edge computing node packages the extracted features and original files and uploads them to the cloud, attaching the case number, timestamp, and data source label to ensure a one-to-one correspondence between the data and the case.
[0090] 2. Perform data preprocessing and semantic fusion.
[0091] After the data is uploaded to the cloud, the preprocessing module further normalizes and fuses the data to construct a unified semantic vector space across formats.
[0092] a. Normalization: The system normalizes the feature vectors of multimodal data such as text, images, and audio to eliminate the differences in feature scales between different modalities. For example, audio spectrum features are normalized to the range of 0-1 to match the scale of other modal features.
[0093] b. Semantic fusion: Through a deep learning model based on multimodal semantic embedding technology (such as a multimodal fusion model based on the Transformer architecture), the system maps feature vectors of different modalities to a unified semantic vector space. In this way, whether the user enters a natural language search term or a cross-modal query, the relevance can be calculated in the same semantic space.
[0094] 3. Perform cross-format semantic retrieval.
[0095] Judicial personnel enter query requests through the system interface, such as "suspect testimony recorded in the on-site video on May 5, 2023". The system generates a query semantic vector based on the natural language content entered by the user, and calculates the similarity between the vector and the semantic vector of the archive data stored in the cloud.
[0096] a. Retrieval process: Using a calculation method based on cosine similarity, the retrieval results are arranged from high to low according to semantic similarity. For example, the system can retrieve the surveillance video that is most relevant to the user's query semantics, and simultaneously return the audio recordings and on-site photos at the corresponding time point, realizing the automatic association of cross-format retrieval results.
[0097] b. Dynamic optimization: After a user clicks on a search result, the system analyzes the user's behavior (such as the selected file type and content) in real time and optimizes the weight distribution of the search model to make subsequent search results more accurate.
[0098] 4. Generate intelligent rules and perform dynamic management.
[0099] This system dynamically generates archive management rules through reinforcement learning algorithms to cope with the complexity and urgency of archive retrieval in judicial cases.
[0100] a. Rule generation: For example, when evidence related to a case is frequently reviewed, the system will automatically increase the priority of these files and reduce the resource allocation of low-priority files to balance the utilization of storage and computing resources.
[0101] b. Adaptive rule adjustment: When a case is closed and the access frequency of related files decreases, the system automatically lowers its priority and migrates the files to a low-cost storage area to achieve efficient management of resources.
[0102] 5. Carry out distributed permission control and security assurance.
[0103] In judicial scenarios, the security requirements for archive management are extremely high, especially in the context of multi-terminal and multi-department collaboration.
[0104] a. Permission verification: The system uses a distributed permission management module built on blockchain technology to verify each user's access request. For example, when a law firm needs to obtain access rights to a piece of evidence material, the system verifies the validity of the permission through zero-knowledge proof to ensure that the user's identity is not leaked.
[0105] b. Data consistency: In a multi-terminal collaboration scenario, the system uses a distributed consistency protocol (PBFT) to ensure the consistency of archive data among nodes, preventing data tampering or loss due to authority conflicts.
[0106] 6. Perform dynamic resource optimization.
[0107] The system achieves dynamic optimization of storage and computing resources through load prediction unit and resource scheduling unit.
[0108] a. Load prediction: Based on the time series model, the system predicts the access load in the judicial scenario. For example, during the peak period of cases, it is predicted that the access volume to certain specific files will increase significantly.
[0109] b. Resource scheduling: Based on the prediction results, the system prioritizes high-performance storage and computing resources to high-priority archives, while optimizing resource utilization of low-priority archives through data tiered storage. This dynamic resource allocation mechanism significantly reduces system energy consumption while ensuring responsiveness in high-concurrency scenarios.
[0110] This embodiment, through unified and efficient cross-format retrieval: realizes unified retrieval and relevance ranking of text, image, audio and video archives, significantly improving the efficiency and accuracy of retrieval; through intelligent rule generation and dynamic optimization, it achieves a significant improvement in resource utilization and ensures priority calling of key archives in high-load scenarios; uses blockchain technology and distributed consistency protocols to solve permission conflicts and data consistency problems in multi-terminal access, providing strong security protection for judicial archive management.
[0111] Embodiment 2:
[0112] In medical scenarios, patients seeking medical treatment across hospitals need to access their complete medical records, which are usually scattered across independent systems in different hospitals and stored in a variety of formats, including electronic medical records (text format), medical images (such as X-rays, CT scan images), laboratory test reports (tabular format), and audio and video files of surgical records. Due to the complexity of different data sources and data formats, it is difficult for existing technologies to achieve fast, accurate, and secure medical record access and management. Based on this, this embodiment demonstrates the specific application of a cloud archive management system based on cross-format deep retrieval and dynamic intelligent management in medical scenarios.
[0113] 1. Data collection and archiving.
[0114] Medical institutions access medical equipment and information management systems through the system to collect patients' multimodal medical record data in real time.
[0115] a. Data collection: When a patient is admitted to the hospital, the hospital's information system will automatically collect their medical records (text files) and upload them in real time. Medical images generated by imaging devices (such as CT and MRI) are connected to edge computing nodes through the imaging system for preliminary image feature extraction. Laboratory reports and surgical records are also uploaded to the system through corresponding devices and interfaces.
[0116] b. Data labeling: When archiving, the system binds the patient's medical record information to the unique patient ID, and attaches metadata such as timestamp and data source hospital for subsequent retrieval and permission management.
[0117] 2. Data preprocessing and semantic modeling.
[0118] The uploaded data is cleaned and normalized by the preprocessing module to ensure that data from different hospitals can be seamlessly integrated into the system.
[0119] a. Normalization processing: For medical images, the system standardizes the resolution and pixel range; for text data, natural language processing (NLP) technology is used to standardize the medical terminology in medical records, for example, synonyms used in different hospitals are uniformly converted into standardized terms.
[0120] b. Semantic modeling: Through the multimodal fusion model, the system maps the semantic vectors of text medical records, the visual feature vectors of medical images, and the table structure vectors of laboratory reports into a unified semantic space to build cross-modal medical record associations. This semantic modeling enables users to call data in different formats through simple natural language queries.
[0121] 3. Retrieval and use of medical records.
[0122] Doctors enter query requests into the system, such as "the patient's latest CT images and related laboratory reports". The system quickly retrieves relevant files from the medical record library based on the unified representation of semantic vectors.
[0123] a. Retrieval process: The system parses the keywords "CT images" and "laboratory reports" in the query, and filters out the medical records of specific patients in combination with the patient ID. Based on the unified semantic space, the data relevance is calculated, and the CT images and laboratory reports are sorted by relevance and returned. Doctors can directly view the relevant files with the highest priority.
[0124] b. Real-time dynamic retrieval: When doctors screen results, the system records the screening behavior and adjusts the weight parameters of the retrieval model to optimize the relevance of subsequent query results.
[0125] 4. Dynamic resource optimization.
[0126] During medical peak periods, such as the morning hours when outpatient volume is high or during epidemics, the system ensures efficient access to medical records through a dynamic resource optimization module.
[0127] a. Load prediction: By analyzing the hospital's historical access records and current load data, the system predicts the access peak in the next time period. For example, if a patient's medical records are continuously accessed, the system will automatically predict that the access frequency may increase further.
[0128] b. Priority allocation: Based on the prediction results, the system migrates medical record data with high access frequency from low-cost storage areas to cache, while limiting access to low-priority files to ensure efficient use of system resources.
[0129] 5. Authority control and data security.
[0130] Since medical records involve highly private data, the system ensures the security and compliance of data access through a distributed permission management module.
[0131] a. Distributed authority management: The system uses blockchain technology to record every permission change and access behavior of medical records, ensuring that every operation can be traced. For example, doctors need to be authorized by the hospital's authority management department before they can access all the patient's medical records.
[0132] b. Access verification: When accessing data, the system uses a zero-knowledge proof algorithm to verify the doctor's access rights, ensuring that the permission verification is completed without leaking other private information of the patient.
[0133] c. Multi-terminal consistency: For remote consultation scenarios, the system uses a distributed consistency protocol (PBFT) to ensure the consistency of data when accessed by multiple hospital terminals, avoiding misdiagnosis or diagnosis delays caused by data synchronization delays.
[0134] 6. Technical effects and comparative experiments.
[0135] In order to verify the technical effect of this system in medical scenarios, the following comparative experiments are designed:
[0136] 1. Experimental scenario: The system was deployed in three different hospitals, and the cross-hospital call scenario of patient medical records was selected for experiment, and the cross-format retrieval efficiency, data call latency and resource utilization were recorded respectively. The comparison system adopts the traditional distributed retrieval method (existing technology).
[0137] 2. Experimental data:
[0138] The complete medical records of 50 patients were used as experimental samples. The total volume of sample medical records was about 5GB, including about 5,000 text records, 1,000 medical images and 500 laboratory reports.
[0139] 3. Experimental results:
[0140] a. Retrieval efficiency: With the support of multimodal semantic modeling, this system reduces the average retrieval time of cross-format medical records from 1.5 seconds in the existing technology to 0.3 seconds, improving efficiency by 80%.
[0141] b. Call delay: Under high-load scenarios, the system's average call delay is only 120 milliseconds, which is 60% shorter than traditional systems.
[0142] c. Resource utilization: Through the dynamic resource optimization module, the system's cache hit rate increased by 30%, saving 20% of storage energy consumption.
[0143] d. Security: In the permission control test, the system achieved 100% data consistency in multi-terminal access scenarios, significantly better than the 97% of the existing technology.
[0144] Through semantic fusion technology and dynamic resource optimization, this embodiment not only realizes the rapid call of cross-hospital data, but also improves resource utilization efficiency and reduces system energy consumption. It is superior to existing technologies in multiple dimensions such as retrieval efficiency, call latency and resource utilization under high-load medical scenarios, especially in terms of security and consistency, solving key problems in traditional decentralized medical record management.
[0145] Embodiment 3:
[0146] In cross-format data processing and semantic vector generation, this embodiment adopts a dynamic optimization weight calculation method for data features of different modalities, so that the fusion of multimodal data in a unified semantic vector space is more accurate and adaptable.
[0147] For example, feature normalization, for multimodal data such as text, images, and audio, the following normalization steps are used to eliminate the differences in feature scales between modalities. For example, text data: use a word segmentation algorithm to segment the text into semantic units, and generate semantic vectors through a pre-trained Transformer model (such as BERT), and then use regularization to limit the vector value to the range of [-1,1]; image data: use a convolutional neural network (such as ResNet) to extract visual features, and normalize the mean and standard deviation of each feature vector; audio data: extract spectral features through fast Fourier transform (FFT), map them to a fixed-dimensional vector, and normalize them to the range of 0 to 1.
[0148] Next, the dynamic weight is calculated, and its dynamic weight is generated by the contextual importance score (Contextual Importance Score, Sx) of the data context, and the calculation formula is as follows:
[0149]
[0150] Among them, W x represents the weight of mode x, S x S is the context importance score of mode x, and n is the total number of modes. Its variable definition: S x : Based on data sources, historical search frequency and current usage scenarios, it is calculated through weighted calculation; Exp(S x ): It is used to amplify the difference between modes and ensure that more important modes have higher weights.
[0151] This method significantly improves the retrieval accuracy of cross-modal data in a unified semantic vector space. At the same time, the dynamic weights adapt to the characteristics of user retrieval requests in different scenarios, overcoming the limitations of traditional static weight allocation methods. In the generation of archive lifecycle management rules, a reinforcement learning algorithm is used to dynamically adapt the balance between user behavior and system resources. The specific steps are as follows:
[0152] First, in terms of objective function design, the goal of rule generation is to maximize the weighted sum of archive access frequency (F) and minimize storage resource consumption (C). The objective function is as follows:
[0153] R=α·F-β·C
[0154] Its variables are defined as: R: rule adaptation score, used to evaluate the pros and cons of the current generated rules; F: the access frequency of the archive, which is counted by the system in real time; C: storage resource consumption, calculated based on the archive size and storage cost; α, β: dynamically adjusted weight parameters, which are adaptively adjusted based on the reinforcement learning algorithm.
[0155] And in the reinforcement learning model, the state space is such as the access frequency of archives, user behavior patterns, and system resource usage. The action space is such as rule adjustment, such as increasing the priority of a certain archive, reducing the storage cost of low-frequency archives, etc.; the reward mechanism is such as feedback to the model in real time based on the optimization results of the objective function R to adjust the strategy. Its implementation path, for example, the initial rule setting is that the system generates initial life cycle management rules based on historical data; its dynamic adjustment is that the reinforcement learning model regularly analyzes user behavior and system resource usage, and optimizes the rules in real time.
[0156] In order to solve the problem of multi-terminal permission control and data consistency, this embodiment proposes a distributed permission management mechanism based on blockchain technology, which specifically includes the following implementation paths. For example, in the blockchain node design, the blockchain network adopts a layered architecture:
[0157] Level 1 nodes: store archive access records and permission change information; Level 2 nodes: responsible for distributed permission verification and consistency maintenance. The PBFT (Practical Byzantine Fault Tolerance) protocol is used between nodes to ensure data consistency.
[0158] The authorization verification process is that the system uses a zero-knowledge proof algorithm to complete authorization verification. The specific steps are as follows: the user submits an authorization verification request, the system generates an authorization certificate, and broadcasts it through the blockchain network; each node verifies the authorization certificate without disclosing the user's authorization details to ensure privacy; in terms of consistency assurance, each file access operation is recorded as a blockchain transaction and synchronized in all nodes; the PBFT protocol ensures the consistency and integrity of file authorization records in multi-terminal access scenarios. Through blockchain technology, the present invention effectively solves the single point failure and authorization conflict problems existing in traditional centralized authorization management, and significantly improves the security and consistency of file management.
[0159] In this embodiment, a load optimization algorithm based on time series prediction is designed in dynamic resource management, which specifically includes the following steps: First, access load prediction, that is, the system predicts the access load of the next period according to historical access records through the following formula:
[0160]
[0161] Variable definition: L pred (t+1): the predicted access load in the next period; L(t): the current access load; A i : The average load of the past n periods; γ: Smoothing coefficient, used to control the weight ratio of historical data to current data.
[0162] Next, in terms of resource scheduling optimization, the system calculates the resource allocation plan based on the prediction results:
[0163]
[0164] Variable definition: R(t): resource allocation result of the current period; P i : Resource priority weight; A i : The demand for the i-th type of resources.
[0165] Example 4
[0166] This embodiment introduces a multi-level dynamic weight adjustment mechanism and a distributed lightweight permission management module to further improve the accuracy of cross-format semantic retrieval, the adaptability of resource scheduling, and the system efficiency and security in multi-terminal scenarios.
[0167] To ensure the operability of this embodiment, the specific design is as follows: in the construction of the multimodal semantic vector space, the feature vectors of different modal data need to be assigned weights to adapt to specific scenario requirements. However, the traditional weight allocation method is relatively static and fails to dynamically respond to changes in real-time usage scenarios. In addition, the context importance score of the weight source lacks specific calculation logic, which may lead to limited optimization effects.
[0168] This embodiment introduces a two-step optimization process for dynamic weight calculation, namely, feature score calculation: generating a modal importance score based on the contextual relevance and historical retrieval frequency of modal features; dynamic weight allocation: adjusting the modal weight to match current needs in combination with real-time user behavior and scenario characteristics.
[0169] Feature scoring formula: Improving context importance score S x The calculation formula is as follows:
[0170]
[0171] Freq(x): The historical search frequency of modality x. Relevance(x): The relevance of modality x in the current usage scenario, calculated based on the deep learning model. Cost(x): The computational cost of modality x, such as the resources or time required for computation. Weight generation formula: Substitute the score S x Mapped to weight W x , calculated using the normalized exponential function:
[0172]
[0173] Where n is the total number of all modes, W x Used to control the fusion ratio of feature vectors in the semantic space. Through the above formula, weight allocation no longer relies on a single static rule, but dynamically adapts to specific scenarios. Experiments show that this mechanism improves the accuracy of cross-modal retrieval by more than 10%, while reducing the computational cost of low-correlation modalities.
[0174] As for distributed lightweight permission verification, the high energy consumption and response delay of blockchain technology in distributed permission control are the difficulties of current technology application. In addition, the ability to protect user privacy in permission verification also needs to be further enhanced. This embodiment designs a lightweight distributed permission verification mechanism, combining zero-knowledge proof with layered blockchain architecture to reduce system energy consumption and improve verification efficiency. Its implementation path is:
[0175] Lightweight blockchain architecture: The blockchain network is divided into a core layer and an edge layer. Core layer: stores key file access records, with a small number of nodes to reduce energy consumption. Edge layer: performs distributed permission verification, and nodes synchronize permission change information through a fast consensus algorithm.
[0176] Permission verification process: The user submits an access request, and the system generates a zero-knowledge proof ZK. ZK is quickly verified by edge layer nodes without disclosing specific permission details. After verification, the core layer records the access operation and updates the permission change information. Fast consensus algorithm: The improved PBFT algorithm is used to significantly shorten the communication time between nodes:
[0177] T PBFT =O(n·log(n))
[0178] Among them, T PBFT is the consensus time complexity, and n is the number of nodes. In this embodiment, the permission verification time is shortened to 60% of the traditional blockchain solution, and the verification accuracy remains 100%. The system energy consumption is reduced by about 20%, meeting the needs of large-scale multi-terminal scenarios. Zero-knowledge proof ensures that user privacy is not leaked, further improving the security of file access.
[0179] In terms of reinforcement learning model optimization, the dynamic generation of archive lifecycle management rules relies on reinforcement learning, but the existing design fails to fully consider the impact of objective function parameters on resource scheduling results. This embodiment introduces a dynamic weight adjustment mechanism to the objective function of reinforcement learning to improve the adaptability of resource allocation. Objective function:
[0180] R=α(t)·F-β(t)·C
[0181] α(t) and β(t) are time-related dynamic weights, which are adaptively adjusted based on historical data trends. F: archive access frequency. C: resource consumption. Dynamic weight calculation:
[0182]
[0183] Among them, HighPriorityCount is the number of high-priority files, and Total Count is the total number of files. In this way, the dynamic adjustment of weights effectively balances the priority of high-frequency access files and the fairness of resource allocation; during peak load periods, the priority adjustment mechanism improves system response efficiency by 15%.
[0184] It is obvious to those skilled in the art that the present invention is not limited to the details of the above exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0185] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the spirit and scope of the technical solution of the present invention.
Claims
1. A cloud archive management system, comprising a data acquisition module, a data preprocessing module, a cross-format semantic retrieval module, an intelligent rule generation module, a distributed authority control module and a dynamic resource optimization module, characterized in that: Data collection module: used to collect archival data in heterogeneous formats from multiple data sources in real time, including text data, image data, and audio / video data; the collected data is subjected to preliminary feature extraction through edge computing devices, and the extracted features include text semantic features, image visual features, and audio waveform features; the preliminary features are synchronously uploaded to the cloud storage unit with the original archival data; Data preprocessing module: used to perform normalization processing on the uploaded heterogeneous data, including semantic segmentation and vectorization of text, compression and embedding coding of image features, and spectrum analysis and embedding generation of audio data; calculate multimodal semantic embedding vectors based on deep learning models, and build a cross-format unified semantic vector space, where the generation formula of the semantic vector is: V f =W t ·V t +W i ·V i +W a ·V a Among them, V f is the multimodal semantic vector, V t 、V i 、V a are text, image and audio feature vectors respectively, and W t , W i , W a The weight coefficient is generated based on historical retrieval data and usage scenario fitting; Cross-format semantic retrieval module: used to receive the user's natural language retrieval request, parse the request into a unified semantic vector, and realize unified retrieval of cross-format archives by calculating the similarity between the user request semantic vector and the archive data semantic vector; the similarity calculation is based on the cosine similarity formula: Where Q is the semantic vector of the user's request, D is the semantic vector of the archive data, and Sim(Q,D) is the similarity score. The search results are sorted from high to low according to the similarity score and returned. Intelligent rule generation module: used to dynamically generate archive lifecycle management rules based on the user's operation history and the frequency of archive access; rule generation uses a reinforcement learning algorithm, and the objective function is the weighted sum of maximizing the archive access frequency and minimizing the storage resource consumption. The objective function formula is: R=α·F-β·C Among them, R is the rule adaptation score, F is the file access frequency, C is the storage resource consumption, α and β are dynamically adjusted weight parameters, which are adaptively adjusted by the algorithm according to the user behavior pattern; Distributed permission control module: Based on blockchain technology, a distributed permission management mechanism is built to perform fine-grained control over the user's access to archives. The blockchain network records the permission request and access record for each archive access, and verifies the user's permission through a zero-knowledge proof algorithm. Distributed consistency protocols such as PBFT are used to ensure data consistency and prevent permission conflicts in multi-terminal access scenarios. Dynamic resource optimization module: used to dynamically adjust the allocation of storage and computing resources according to system load and access frequency; predict the access load in the next time period through the resource scheduling optimization algorithm and calculate the best resource allocation plan; the scheduling optimization formula is: Among them, L pred (t+1) is the predicted value of the access load in the next time period, L(t) is the current load, A i is the average load in the past n time periods, and γ is the smoothing coefficient.
2. The cloud archive management system according to claim 1, characterized in that: The data acquisition module includes an edge computing unit and a data upload unit, wherein: the edge computing unit performs real-time feature extraction on the collected heterogeneous data and generates preliminary feature data; the data upload unit synchronously transmits the preliminary feature data and the corresponding original data to the cloud storage unit, and marks the timestamp and data source identifier to support subsequent retrieval and permission management.
3. The cloud archive management system according to claim 1, characterized in that: The data preprocessing module further includes: a data cleaning module for removing noise, redundant data and invalid fields in the uploaded data; a feature fusion module for calculating the weight coefficients of different modal data based on the self-attention mechanism to adjust the generation of multimodal semantic vectors, wherein the weight coefficients are calculated by the following formula: Among them, W x is the weight coefficient of mode x, S x Score the importance of modality x in a specific context, where n is the total number of all modalities.
4. The cloud archive management system according to claim 1, characterized in that: The cross-format semantic retrieval module also includes a user intent recognition unit, which is specifically used to: construct a semantic vector based on the user's retrieval request statement, and parse the user's intent through a deep learning model; generate secondary retrieval conditions based on the intent analysis results, and further refine the cross-format semantic retrieval results to improve the accuracy of the retrieval results.
5. The cloud archive management system according to claim 1, characterized in that: The intelligent rule generation module further includes a user behavior analysis unit, which is used to: extract behavior features based on the user's access history and operation behavior to the archive; A reinforcement learning algorithm is used to dynamically adjust the weight parameters of the objective function, where the behavioral characteristics include access frequency, operation duration, and interaction depth.
6. The cloud archive management system according to claim 1, characterized in that: The blockchain network in the distributed permission control module adopts a hierarchical structure, specifically including: a primary node for storing archive access records and permission change information; a secondary node for distributed permission verification and consistency maintenance; each node ensures the consistency and integrity of archive access records in multi-terminal scenarios through a distributed consistency protocol PBFT.
7. The cloud archive management system according to claim 1, characterized in that: The distributed permission control module adopts zero-knowledge proof technology to achieve the following functions: verifying user permissions without disclosing permission details; preventing permission conflicts caused by simultaneous access by multiple users, and ensuring the privacy of data access.
8. The cloud archive management system according to claim 1, characterized in that: The dynamic resource optimization module further includes a load prediction unit and a resource scheduling unit, wherein: the load prediction unit predicts the future access load based on the time series model; the resource scheduling unit dynamically adjusts the allocation of storage resources and computing resources according to the predicted access load, and the optimization formula is: Among them, R(t) is the resource scheduling result at time t, P i is the resource priority weight, A i is the demand for the i-th resource, and n is the total number of resource categories.
9. The cloud archive management system according to claim 1, characterized in that: The cross-format semantic retrieval module further includes a result ranking unit, which generates a comprehensive score Score by fusing user preference data and semantic similarity scores. The calculation formula of the comprehensive score is: Score=λ·Sim(Q,D)+(1-λ)·Pref Among them, Sim(Q,D) is the semantic similarity score between the user request and the archive data, Pref is the score based on the user's historical preference, and λ is the weight factor, which is dynamically adjusted according to the user's usage scenario.
10. The cloud archive management system according to claim 1, characterized in that: The data preprocessing module and the dynamic resource optimization module work together to give priority to the archival data with high access frequency and high priority after the data is uploaded, wherein the priority is determined by the following formula: Priority=ω1·F+ω2·U Among them, Priority is the priority of archive data, F is the historical access frequency of the archive, U is the emergency use demand score of the archive, and ω1 and ω2 are priority weight parameters.
Citation Information
Cited By
Intelligent knowledge base system for hierarchical authority management
CN120181807A
Method for carrying out intelligent abstracting on unhealthy asset association cases
CN120234411A
Multi-terminal-oriented distributed audio and video adaptive production system
CN120301991A
Big data resource processing method based on cloud database service
CN120469843A
Automatic classified storage system for multi-modal archive files
CN120469975A