Unmanned archive remote management method and system based on 5G network
By establishing a dual-channel system for image transmission and control auditing in a 5G unmanned archive, generating traceable session records, performing image standardization processing, and extracting layout and semantic features, the problem of reproducible processing and anomaly assessment of low-quality scanned images was solved. This enabled deterministic nearest neighbor retrieval and auditable remote archiving management, improving the management efficiency and security of the archive.
Patent Information
- Application Number
- CN202610050623.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-15
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2046-01-15
AI Technical Summary
Existing technologies are insufficient to achieve remote closed-loop archiving management for reproducible processing, deterministic anomalies, prevalence assessment, and auditability of low-quality scanned images in 5G unattended archive scenarios.
A dual-channel system for image transmission and control auditing based on a 5G network is established to generate traceable session records. Standardization processing is driven by image quality indicators, and format and semantic features are extracted to form a multi-dimensional index. Deterministic nearest neighbor retrieval is used to output anomaly and prevalence. Combined with key management, fingerprint watermarking, and genealogy tracing, reliable remote management is achieved.
It enables reproducible image processing in 5G unattended archives, stably triggers rescanning and strategy switching, and supports remote management that is space-saving, easy to retrieve, highly auditable, and accountable.
Smart Images

Figure CN121542222A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of digital management technology, and in particular to a remote management method and system for unmanned archives based on a 5G network. Background Technology
[0002] In district / county or park-level digital archives, nighttime batch scanning and daytime remote review have become common practices. Because the archives are unattended, scanning equipment, access control systems, and edge nodes need to transmit large amounts of images and logs back to the remote center via 5G networks, and receive remotely assigned tasks, rescans, and parameter adjustments. Existing solutions often suffer from three main problems: First, image quality is affected by shadows, tilt, reflections, red stamps, and paper creases, leading to fluctuations in layout positioning and text recognition, making it difficult to promptly detect and rescan abnormal pages. Second, archiving and retrieval often rely on a single dimension, resulting in low retrieval efficiency and a high risk of incorrect archiving when using the same template for different business purposes or the same business with different layouts. Third, storage and security are often disconnected from business operations, lacking reproducible processing parameters, auditable session links, and complete lineage tracing, making it difficult to meet the requirements of long-term preservation, immutability, and accountability for leaks.
[0003] Currently, Chinese invention patent application number CN202510258100.2 discloses a multi-dimensional spatial storage management method, system, and storage medium for digital archives. The method includes receiving archives to be stored, scanning the archives to obtain digital archives; the digital archives are a set of scanned images; for any given digital archive, images are read sequentially, images are recognized, text boxes are located, and feature curves for each image are extracted based on the located text boxes; the anomaly degree of any image is determined based on the feature curves, the anomaly degrees of all images in the same digital archive are statistically analyzed to determine the prevalence of the digital archive; and the digital archives are archived and stored based on the prevalence. This invention scans archives to obtain digital archives, extracts features from the digital archives, and then classifies and archives them, greatly improving data retention time and reducing data storage space.
[0004] The aforementioned technologies are insufficient for achieving reproducible processing, deterministic anomaly assessment, prevalence evaluation, and auditable remote closed-loop archiving management of low-quality scanned images in a 5G unattended archive scenario. Summary of the Invention
[0005] The technical problem solved by this invention is that existing technologies are difficult to achieve remote closed-loop archiving management for reproducible processing, deterministic anomalies, prevalence assessment, and auditability of low-quality scanned images in 5G unattended archive scenarios.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: A method for remote management of unmanned archives based on 5G networks includes the following steps: Step S1: Establish an image transmission channel and a control audit channel based on the 5G network, and generate session record data; Step S2: Based on the session record data, perform session tagging and time alignment on the collected original archive image data, calculate image quality index data, generate standardized archive image data based on the image quality index data, obtain archive metadata corresponding to the original archive image data, and write the archive metadata into the session record data; Step S3: Extract layout structure point data from standardized archival image data and fit layout feature curve data, perform text recognition to obtain recognized text data and generate semantic feature data, and combine archival metadata to generate multidimensional index entry data. Step S4: Perform deterministic nearest neighbor retrieval on the multidimensional index entry data and output anomaly data and universality data; Step S5: Generate storage strategy data and genealogy tracing data based on prevalence data, archive metadata, and image quality index data; Step S6: Based on key management, generate encryption key data to encrypt the archive data output according to the storage policy data, generate fingerprint digest data and watermark data, and after associating anomaly data, prevalence data, storage policy data, genealogy tracing data and session record data, output remote management result data through the control audit channel.
[0007] Preferably, step S1 includes the following sub-steps: Step S101: Collect device identity data, which includes the identity of the scanning device, the identity of the access control device, and the identity of the edge node; Step S102: Establish an image transmission channel and a control audit channel based on device identity data; Step S103: Write the device identity data, timestamp, and channel identifier into the session record data.
[0008] Preferably, step S2 includes the following sub-steps: Step S201: Collect raw archive image data and time-align it with session record data to form raw archive image data with session tags; Step S202: Calculate image quality index data for the original archive image data and write the image quality index data into the session record data; Step S203: Obtain the archive metadata corresponding to the original archive image data and write the archive metadata into the session record data.
[0009] Preferably, the standardized archival image data is obtained by standardizing the original archival image data based on image quality index data: Denoising parameters, skew correction parameters, and cropping parameters are generated based on image quality index data. The original archival image data is sequentially subjected to denoising, skew correction, text area cropping, and resolution normalization to output standardized archival image data. Write the denoising parameters, correction parameters, and clipping parameters into the session log data.
[0010] Preferably, step S3 includes the following sub-steps: Step S301: Locate the text region in the standardized archival image data and output the layout structure point data; Step S302: Generate layout feature sequence based on layout structure point data and fit to obtain layout feature curve data; Step S303: Perform character recognition on the standardized archival image data to obtain recognized text data, and generate semantic feature data based on the recognized text data; Step S304: Collect archive metadata and combine layout feature curve data, semantic feature data, archive metadata and image quality index data to generate multidimensional index entry data.
[0011] Preferably, step S4 includes the following sub-steps: Step S401: Write the layout feature curve data into the layout nearest neighbor index library and write the semantic feature data into the semantic nearest neighbor index library. Step S402: Perform deterministic nearest neighbor retrieval on the layout feature curve data based on the layout nearest neighbor index library to obtain the layout nearest neighbor set; perform deterministic nearest neighbor retrieval on the semantic feature data based on the semantic nearest neighbor index library to obtain the semantic nearest neighbor set. Step S403: Calculate the anomaly score data based on the layout nearest neighbor set, the semantic nearest neighbor set, and the image quality index data; Step S404: Calculate the universality data based on the layout nearest neighbor set, the semantic nearest neighbor set, and the archive metadata.
[0012] Preferably, step S5 includes the following sub-steps: Step S501: Generate storage strategy data based on the prevalence data and file metadata. The storage strategy data includes one or a combination of template differential storage strategy, hot and cold tiered storage strategy and tamper-proof storage strategy. Step S502: When the universality data meets the universality condition, execute the template differential storage strategy and record the differential relationship data; Step S503: When the archive metadata meets the requirements for long-term preservation or auditing, execute the tamper-proof storage policy and record the storage location data; Step S504: Associate the multidimensional index entry data, storage strategy data and their execution results to generate genealogy tracing data, and bind and store it with session record data.
[0013] Preferably, step S6 includes the following sub-steps: Step S601: Generate encryption key data for each file based on key management, and use the encryption key data to encrypt the file data output according to the storage strategy. Step S602: Generate fingerprint summary data based on layout feature curve data and semantic feature data, and write the fingerprint summary data into the genealogy tracing data; Step S603: Generate watermark data based on fingerprint digest data and embed it into the corresponding archive data; Step S604: Output remote management result data through the control audit channel. The remote management result data includes at least anomaly data, prevalence data, storage strategy data, and index identifiers for genealogy tracing data.
[0014] Preferably, the remote management result data output through the control audit channel includes remote closed-loop control logic: When the anomaly data meets the anomaly conditions, a rescan instruction is sent to the scanning device corresponding to the device identity data based on the session record data, and new original archive image data is sent back. When the image quality index data does not meet the quality conditions, the standardization processing parameters are adjusted based on the session record data and the standardized archive image data is regenerated. When the prevalence data meets the prevalence condition and the image quality index data meets the quality condition, the template differential storage strategy is triggered and the lineage tracing data is updated.
[0015] A remote management system for unmanned archives based on a 5G network includes a session recording acquisition module, an image data processing module, a fitting curve construction module, a data nearest neighbor retrieval module, a storage strategy generation module, and a remote management and control module. The session recording acquisition module is used to establish an image transmission channel and a control audit channel based on the 5G network, and to generate session recording data; The image data processing module is used to perform session tagging and time alignment on the collected original archive image data based on session record data, calculate image quality index data, generate standardized archive image data based on the image quality index data, obtain archive metadata corresponding to the original archive image data, and write the archive metadata into the session record data. The fitting curve construction module is used to extract layout structure point data from standardized archival image data and fit layout feature curve data, perform text recognition to obtain recognized text data and generate semantic feature data, and combine archival metadata to generate multidimensional index entry data. The data nearest neighbor retrieval module is used to perform deterministic nearest neighbor retrieval on multidimensional index entry data and output anomaly data and generality data; The storage strategy generation module is used to generate storage strategy data and genealogy tracing data based on universality data, archive metadata and image quality index data. The remote management control module is used to generate encryption key data based on key management to encrypt the archive data output according to the storage strategy data, generate fingerprint digest data and watermark data, and associate anomaly data, prevalence data, storage strategy data, genealogy tracing data and session record data, and output remote management result data through the control audit channel.
[0016] The beneficial effects of this invention are as follows: This invention is designed for 5G unattended archives, establishing a dual-channel system for image transmission and control auditing and generating traceable session records. It drives image standardization with quality indicators such as sharpness, tilt, and shadows, ensuring parameter reproducibility. It extracts layout feature curves, semantic features, and archival metadata to form a multi-dimensional index. It uses deterministic nearest neighbor retrieval to output anomaly and prevalence, stably triggering rescanning and strategy switching. Prevalence drives template differentiation and tamper-proof storage. Combined with key management encryption, fingerprint watermarking, and genealogical tracing, it achieves space saving, easy retrieval, strong auditing, and accountability. Attached Figure Description
[0017] Figure 1 A flowchart illustrating the steps of a remote management method for an unmanned archive based on a 5G network, as provided in one embodiment of the present invention; Figure 2 This is a basic flowchart of a remote management system for unmanned archives based on a 5G network, provided as an embodiment of the present invention. Detailed Implementation
[0018] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.
[0019] Example 1, referring to Figure 1 This paper presents a remote management method for unmanned archives based on a 5G network. This embodiment targets a county-level unmanned archive scenario: the archive is unattended at night, and paper archives are automatically scanned in batches by scanning equipment. Edge nodes are installed within the archive (for local processing and network outage buffering). The remote management terminal is located in a city-level data center, and the archive accesses the network via a dedicated 5G network. Two types of communication are distinguished: image transmission and control auditing, to ensure that large-scale image uploads and control auditing do not crowd out bandwidth. The method includes the following steps: Step S1: Establish an image transmission channel and a control audit channel based on the 5G network, and generate session record data.
[0020] Step S2: Based on the session record data, perform session tagging and time alignment on the collected original archival image data, calculate image quality index data, generate standardized archival image data based on the image quality index data, obtain archival metadata corresponding to the original archival image data, and write the archival metadata into the session record data.
[0021] Step S3 involves extracting layout structure point data from standardized archival image data and fitting layout feature curve data, performing text recognition to obtain recognized text data and generating semantic feature data, and combining this with archival metadata to generate multidimensional index entry data.
[0022] Step S4: Perform deterministic nearest neighbor retrieval on the multidimensional index entry data and output anomaly data and universality data.
[0023] Step S5: Generate storage strategy data and genealogy tracing data based on prevalence data, archive metadata, and image quality index data.
[0024] Step S6: Based on key management, generate encryption key data to encrypt the archive data output according to the storage policy data, generate fingerprint digest data and watermark data, and after associating anomaly data, prevalence data, storage policy data, genealogy tracing data and session record data, output remote management result data through the control audit channel.
[0025] This invention targets 5G unmanned archives, establishing a dual-channel system for image transmission and control auditing, and generating traceable session records. It drives image standardization with quality indicators such as sharpness, tilt, and shadows, ensuring parameter reproducibility. It extracts layout feature curves, semantic features, and archival metadata to form a multi-dimensional index. It uses deterministic nearest neighbor retrieval to output anomaly and prevalence, stably triggering rescanning and strategy switching. Prevalence drives template differentiation and tamper-proof storage. Combined with key management encryption, fingerprint watermarking, and genealogical tracing, it achieves space-saving, easy retrieval, strong auditing, and accountability.
[0026] Step S1 includes the following sub-steps: Step S101: Collect device identity data, which includes the identity of the scanning device, the identity of the access control device, and the identity of the edge node.
[0027] After powering on, warehouse-side scanning devices, access control devices, and edge nodes register their device identity data with the remote management terminal. This device identity data includes at least the device ID, device type, warehouse ID, firmware version, and certificate serial number. After verifying the device certificate, the remote management terminal issues a session token and establishes session record data. This session record data includes at least the device identity data, session token, session start time, channel identifier, scanning task number, and batch number. This session record data is subsequently used for: Tag each scanned image with a session tag to ensure that each image corresponds to a device, task, and time. As the basis for instruction signatures and receipts to control the audit channel; As a basis for audit traceability and parameter reproduction, it ensures that the processing parameters at that time can be restored when the same batch is re-scanned or when disputes are reviewed.
[0028] Step S102: Establish an image transmission channel and a control audit channel based on device identity data.
[0029] The image transmission channel is used to upload original and standardized archive images, while the control and audit channel is used to issue scanning tasks, rescan instructions, parameter adjustment instructions, and upload anomaly alarms, audit logs, and pedigree tracing information. Both types of channels use session tokens for authentication, and the session log data records the issuance time, execution result, and receipt summary of each instruction to avoid the problem of having images but not knowing who collected them, when they were collected, or whether they have been modified when there is no one to monitor them.
[0030] Step S103: Write the device identity data, timestamp, and channel identifier into the session record data.
[0031] Step S2 includes the following sub-steps: Step S201: Collect raw archive image data and time-align it with session record data to form raw archive image data with session tags.
[0032] The scanning device performs scanning according to the scanning task number issued remotely, and obtains the original archive image data. After receiving the data, the edge node immediately writes the session tag, page number and acquisition time from the session record data into the header index information of the original archive image data, and sends the original archive image data index back to the remote management terminal, so that the remote terminal can see the progress and page number missing status even if it has not received all the images.
[0033] Step S202: Calculate image quality index data for the original archive image data and write the image quality index data into the session record data.
[0034] Edge nodes calculate image quality metrics for each original archive image and write this data into the session record data and the image index of that page. This data drives the normalization processing parameters and participates in anomaly determination (when poor quality leads to unstable recognition, anomaly calculation is penalized to avoid misjudging poor quality as content anomaly). The image quality metrics in this embodiment include at least sharpness, tilt angle, shadow ratio, and text ratio, calculated as follows: Sharpness index Using the Laplace variance method: First, assume that the single-page image before standardization is converted to grayscale to obtain a grayscale image. The edge response map is obtained by applying the Laplacian operator to the grayscale image. Then calculate its variance; the larger the variance, the clearer the representation. This can be denoted as: ; in," This represents a second-order edge operator. The sharpness index is normalized to the range of 0 to 1. The normalization method is to subtract the minimum value from the current value and then divide by the range. The minimum value and range are derived from the statistical benchmark of the historical batches of this warehouse and written into the session record data.
[0035] The tilt angle index uses text line direction detection: long straight lines (table lines or text line baselines) are extracted from the binarized image, and their main direction angle is calculated as the tilt angle. The larger the absolute value of the tilt angle, the more severe the skewness.
[0036] The shadow ratio index is decomposed using the background brightness map: the background brightness map is estimated using large-scale filtering. Then, the proportion of the entire page where the brightness is below the shadow threshold is calculated. The shadow threshold is automatically determined by the background brightness distribution of this batch and written into the session record data. The specific calculation method is as follows: Define the set of shaded areas : ; Shading ratio index Defined as: ; in, The shadow threshold, It is the set of all pixels in the image. For the number of pixels, The text proportion metric uses connected component filtering: Connectivity analysis is performed on the binarized graph to filter out excessively large connected components (such as entire black borders) and excessively small noise points. The total number of pixels in the remaining connected components is divided by the total number of pixels on the entire page to obtain the text proportion. Text proportion is used to prevent blank pages, cover images, and photo pages from affecting the layout curves and to participate in anomaly penalties.
[0037] Step S203: Obtain the archive metadata corresponding to the original archive image data and write the archive metadata into the session record data.
[0038] Standardized archival image data is obtained by standardizing the original archival image data based on image quality index data: Denoising parameters, skew correction parameters, and cropping parameters are generated based on image quality index data.
[0039] The original archival image data is sequentially subjected to denoising, skew correction, text area cropping, and resolution normalization to output standardized archival image data.
[0040] Write the denoising parameters, correction parameters, and clipping parameters into the session log data.
[0041] Edge nodes generate standardized processing parameters based on the aforementioned quality indicators and perform standardized processing, outputting standardized archival image data. Standardized processing includes denoising, tilt correction, image core cropping, and resolution normalization. This embodiment requires that the denoising intensity, tilt correction angle, cropping boundary, and normalized resolution be written into the session record data. If a rescan or manual review is triggered subsequently, the same processing procedure can be replayed based on the session record data.
[0042] Step S3 includes the following sub-steps: Step S301: Locate the text region in the standardized archival image data and output the layout structure point data.
[0043] Edge nodes locate text regions in the standardized archival image, obtaining layout structure point data. This embodiment employs a robust scheme of connected component merging and projection segmentation: First, candidate text connected components are extracted from the binary graph and merged into text blocks according to distance. Then, the text blocks are divided by horizontal and vertical projection to correct the boundaries, thus obtaining the bounding rectangle of each text block. The four corner points of each rectangle, together with the center point, constitute the layout structure point data. The located points are not intended to resemble text boxes, but rather to be fitted later to form a layout feature curve that is more robust to scanning noise.
[0044] Step S302: Generate layout feature sequence based on layout structure point data and fit to obtain layout feature curve data.
[0045] Edge nodes are fitted with layout feature curve data based on layout structure point data. This embodiment uses a diagonal projection histogram curve to reduce sensitivity to minor clipping and local noise. Map each point to the projected coordinates in the diagonal direction from the top left to the bottom right (add the horizontal and vertical coordinates of the points and then normalize them according to the page width and height), and then count the number of all points according to the number of intervals in which the projected coordinates fall, forming an interval counting curve. The curve is then smoothed and normalized to the 0-1 range to obtain the layout feature curve data. The layout feature curve data is essentially a stable signature of the layout structure, which is more noise-resistant and easier to retrieve nearest neighbors than directly storing text box coordinates.
[0046] Step S303: Perform character recognition on the standardized archival image data to obtain recognized text data, and generate semantic feature data based on the recognized text data.
[0047] Edge nodes perform text recognition on standardized archival images to obtain recognized text data and generate semantic feature data. This embodiment, for practicality and interpretability, employs two layers of semantic features: The first layer is the keyword vector (which is obtained by segmenting the text, removing stop words, counting high-frequency words, and encoding them according to their weights). The second layer is the topic vector (the topic distribution is obtained by reducing the dimensionality of the sentence segment vector).
[0048] If a deep semantic model is deployed, the identified text can also be fed into a sentence vector model to obtain semantic vectors. However, regardless of which model is used, it is required to write the model version number, vocabulary version number, or topic model version number into the session record data.
[0049] Step S304: Collect archive metadata and combine layout feature curve data, semantic feature data, archive metadata and image quality index data to generate multidimensional index entry data.
[0050] The archive metadata is obtained and bound to the aforementioned data to generate multidimensional index entries. The archive metadata is not arbitrary; in this embodiment, it can be used alone or in combination for verification. Firstly, when the remote management terminal issues a scanning task, it carries task ledger fields (such as file number, archival group number, catalog number, security classification, retention period, creation date, batch number), and the edge node directly writes the session record data and binds it to each page; Secondly, edge nodes retrieve the collection ledger from the archive management system based on the archive number and then backfill it; Third, the fixed fields on the homepage are extracted in a structured manner to generate parsed ledger fields, and consistency checks are performed with the first two sources (if there is a discrepancy, an anomaly level is increased and a review work order is generated).
[0051] Ultimately, the multidimensional index entry data includes at least: session record data references, page numbers, image quality index data, layout feature curve data, semantic feature data, and archive metadata.
[0052] Step S4 includes the following sub-steps: Step S401: Write the layout feature curve data into the layout nearest neighbor index library and write the semantic feature data into the semantic nearest neighbor index library.
[0053] Edge nodes write the layout feature curve data of historical batches into the layout nearest neighbor index library and the semantic feature data of historical batches into the semantic nearest neighbor index library. The index library adopts a deterministic nearest neighbor structure to avoid different scores for the same page in multiple runs due to random sampling comparison. In engineering, this embodiment can use two types of indexes: The layout nearest neighbor index uses curve fingerprint binning and precise comparison within each bin; The semantic nearest neighbor index library uses vector nearest neighbor indexes (e.g., graph structure nearest neighbors).
[0054] Regardless of the structure used, the index building process must output the index version number and the batch range for data entry, and write this information into the session log data to ensure traceability.
[0055] Step S402: Perform deterministic nearest neighbor retrieval on the layout feature curve data based on the layout nearest neighbor index library to obtain the layout nearest neighbor set, and perform deterministic nearest neighbor retrieval on the semantic feature data based on the semantic nearest neighbor index library to obtain the semantic nearest neighbor set.
[0056] For the layout feature curve data of the current page, the candidate set is first located using curve fingerprinting, and then the precise similarity is calculated on the candidate set to obtain the layout nearest neighbor set. For the semantic feature data of the current page, vector nearest neighbor retrieval is directly performed to obtain the semantic nearest neighbor set. The similarity is based on cosine similarity, and the similarity is limited to between 0 and 1, with the closer to 1 indicating greater similarity.
[0057] Step S403: Calculate the anomaly data based on the layout nearest neighbor set, the semantic nearest neighbor set, and the image quality index data.
[0058] Anomaly data is used to determine whether a rescan or review is needed. Its prediction logic is based on the assumption that the nearest similar neighbor is not similar enough, and a quality penalty is applied. Specifically: The maximum similarity is obtained by taking the maximum similarity in the layout nearest neighbor set and the semantic nearest neighbor set respectively, and then merging them according to the weight to obtain the comprehensive maximum similarity. The lower the overall maximum similarity, the higher the anomaly.
[0059] Meanwhile, if image quality metrics indicate poor clarity, large shadows, or severe tilt, a penalty term is added to the anomaly score to prevent misinterpreting poor quality as abnormal file content. This section is the core of this method's prediction: it doesn't predict future time series, but rather predicts whether the current page is in an abnormal state, and uses this abnormal state for closed-loop control.
[0060] Step S404: Calculate the universality data based on the layout nearest neighbor set, the semantic nearest neighbor set, and the archive metadata.
[0061] Prevalence data is used to drive storage strategies (template differentiation, hot / cold tiering), and its prediction logic is based on how common something is in the structural and semantic space. Specifically: The average similarity of several most similar nearest neighbors is taken as the basic value of universality; The similarity is corrected by incorporating category constraints from the archive's metadata. For example, similarity within the same archival group or with the same catalog number is given higher weight, while similarity across different archival groups is given lower weight. This ensures that the generality reflects not only the similarity of the layout but also the consistency of the business affiliation. Higher generality indicates that the archive is more likely to belong to a common template or common business type, making it more suitable for template differential storage to save space.
[0062] Step S5 includes the following sub-steps: Step S501: Generate storage strategy data based on generality data and file metadata. The storage strategy data includes one or a combination of template differential storage strategy, hot and cold tiered storage strategy and tamper-proof storage strategy.
[0063] Edge nodes generate storage strategy data based on prevalence data, archive metadata, and image quality metrics. Storage strategies include at least one or a combination of three types: template differential storage strategy, hot / cold tiered storage strategy, and tamper-proof storage strategy. Strategy generation employs an auditable threshold lookup method: for example, images with high prevalence and low security level are processed via template differential and placed in the hot tier for frequent retrieval; images with low prevalence but long retention periods or high security levels are stored using tamper-proof storage and placed in the cold archive tier; images with poor quality but still requiring storage are forcibly retained in their original form with a rescan recommendation recorded. The threshold table and strategy version number are written into the session log data.
[0064] Step S502: When the universality data meets the universality condition, execute the template differential storage strategy and record the differential relationship data.
[0065] When the template differential storage strategy is triggered, the edge node first selects a representative template page based on the nearest neighbor set (generally, the nearest neighbor with the highest similarity is chosen as the template). Then, it aligns the current page with the template page at the text block level and outputs differential relationship data. The differential relationship data includes: the position information of each text block, whether it has been changed, and a summary of the image or text fragment of the changed area. During storage, only the template page reference and differential relationship data are stored, instead of storing the entire page repeatedly. This significantly saves space for template files with high prevalence. During retrieval or browsing, the data is reconstructed and displayed based on the template page and differential relationship data. The reconstruction process is also recorded in the genealogy tracing.
[0066] Step S503: When the archive metadata meets the requirements for long-term preservation or auditing, implement an immutable storage policy and record the storage location data.
[0067] Regardless of the storage strategy adopted, edge nodes will write multidimensional index entry data, storage strategy data, storage location data, template reference relationships, differential relationship data, and session record data references into the genealogy tracing data in chronological order. The genealogy tracing data is synchronized to the remote management terminal through the control audit channel, meeting the rigid requirements of unmanned archives for auditing and accountability.
[0068] Step S504: Associate the multidimensional index entry data, storage strategy data and their execution results to generate genealogy tracing data, and bind and store it with session record data.
[0069] Step S6 includes the following sub-steps: Step S601: Generate encryption key data for each file based on key management, and use the encryption key data to encrypt the file data output according to the storage strategy.
[0070] Edge nodes invoke the key management service to generate encryption key data for each file, and then encrypt the file data using an envelope encryption method: the file content is stored after being encrypted with the encryption key data, and the encryption key data is simultaneously encapsulated and saved using the master key of the key management service. This enables key rotation, revocation, and access auditing, avoiding instability or unauditability caused by directly using layout feature curves as the key source.
[0071] Step S602: Generate fingerprint summary data based on layout feature curve data and semantic feature data, and write the fingerprint summary data into the genealogy tracing data.
[0072] Fingerprint digest data is used for leakage tracing and consistency verification. In this embodiment, the fingerprint digest data is obtained by concatenating the normalized sequence digest of the layout feature curve data, the digest of the semantic feature data, and the digest of the file number and batch number in the file metadata and then performing a hash operation. : ; in, This is a normalized sequence summary of the layout characteristic curve data. A summary of semantic feature data. The file number in the file metadata. This is a summary of the batch number in the archive metadata. For secure hash functions, This is for splicing operations.
[0073] Semantic summarization involves extracting the first few weighted words from semantic feature data and encoding them with their respective weights into a short string. The advantage of this approach is that the fingerprint contains both format stability information and semantic distinguishing information, and is linked to a specific file number, facilitating accountability and comparison.
[0074] Step S603: Generate watermark data based on fingerprint digest data and embed it into the corresponding archive data.
[0075] The watermark data is derived from the fingerprint digest data as a bit sequence and then embedded into the frequency domain coefficients of the archival image to resist common compression and scaling. In engineering, a block-based discrete cosine transform embedding method can be used: the standardized archival image is divided into several blocks, a discrete cosine transform is performed on each block, and small increments or decrements are made to the low-frequency coefficients according to the watermark bits. The embedding strength is then correlated with image quality indicators (the clearer the image, the weaker the embedding can be to reduce visual impact; the worse the image, the stronger the embedding can be to ensure detectability). The embedding position and embedding strength parameters are also written into the session log data for subsequent verification and legal evidence reproduction.
[0076] Step S604: Output remote management result data through the control audit channel. The remote management result data includes at least the index identifiers of anomaly data, prevalence data, storage strategy data, and genealogy tracing data.
[0077] The remote management result data, including remote closed-loop control logic, is output through the control audit channel. When the anomaly data meets the anomaly conditions, a rescan instruction is sent to the scanning device corresponding to the device identity data based on the session record data, and new original archive image data is sent back.
[0078] When the image quality index data does not meet the quality conditions, the standardization processing parameters are adjusted based on the session record data and the standardized archive image data is regenerated.
[0079] When the prevalence data meets the prevalence condition and the image quality index data meets the quality condition, the template differential storage strategy is triggered and the lineage tracing data is updated.
[0080] Edge nodes output anomaly data, prevalence data, storage strategy data, pedigree tracing data index, encrypted result summary, fingerprint summary data, and watermark embedding result receipts as remote management result data through the control audit channel. The remote management terminal performs closed-loop control based on the remote management result data: if the anomaly exceeds the threshold, a rescan instruction is automatically issued and the standardized processing parameters that need to be improved are specified; if the image quality indicators continue to fail to meet the standards, an equipment maintenance work order is issued; if the prevalence is high, template differential is prioritized and batch storage of the same template batch is accelerated.
[0081] Example 2, refer to Figure 2 This paper presents a remote management system for unmanned archives based on a 5G network, which includes a session recording acquisition module, an image data processing module, a fitting curve construction module, a data nearest neighbor retrieval module, a storage strategy generation module, and a remote management and control module.
[0082] The session recording acquisition module is used to establish image transmission channels and control audit channels based on the 5G network and generate session recording data.
[0083] The image data processing module is used to perform session tagging and time alignment on the acquired raw archival image data based on session record data, calculate image quality index data, generate standardized archival image data based on the image quality index data, obtain archival metadata corresponding to the raw archival image data, and write the archival metadata into the session record data.
[0084] The curve fitting module is used to extract layout structure point data from standardized archival image data and fit layout feature curve data, perform text recognition to obtain recognized text data and generate semantic feature data, and combine archival metadata to generate multidimensional index entry data.
[0085] The data nearest neighbor retrieval module is used to perform deterministic nearest neighbor retrieval on multidimensional index entries and output anomaly and generality data.
[0086] The storage strategy generation module is used to generate storage strategy data and genealogy tracing data based on prevalence data, archive metadata, and image quality index data.
[0087] The remote management and control module is used to generate encryption key data based on key management to encrypt the archive data output according to the storage policy data, generate fingerprint digest data and watermark data, and after associating anomaly data, prevalence data, storage policy data, genealogy tracing data and session record data, output remote management result data through the control audit channel.
[0088] By separating the image transmission channel from the control audit channel, and coordinating the data acquisition and alignment, instruction issuance, receipt recording, and audit traceability of session records, the problem of mutual interference between image transmission and control audit and the lack of process traceability during unattended operation is solved. This allows the remote center to complete rescanning, parameter adjustment, and work order closure without interrupting image uploads.
[0089] Image quality metrics such as sharpness, tilt angle, shadow ratio, and text ratio are constructed, and these metrics drive denoising, skew correction, cropping, and resolution normalization. This enables standardized processing to have reproducible parameters, significantly reducing layout drift and recognition fluctuations across devices and batches, and minimizing misjudgments and missed judgments.
[0090] Simultaneously, the layout structure points are extracted and the layout feature curve is fitted. Combined with text recognition to generate semantic features, and then the archival metadata is introduced to form multi-dimensional index entries. This avoids confusion caused by relying solely on template similarity or full-text retrieval, which results in different business types for the same template or different layout types for the same business. This improves positioning efficiency and archival consistency.
[0091] Deterministic nearest neighbor retrieval is used to calculate the similarity between the layout and semantics, thereby obtaining the anomaly and prevalence. An image quality penalty term is introduced to distinguish between poor quality and content anomalies, avoiding the instability of results caused by random sampling comparison, thus triggering rescanning, verification and strategy switching more reliably.
[0092] Based on prevalence, metadata, and quality indicators, storage strategies are generated, supporting the selection of strategies such as template differentiation, hot / cold tiering, and tamper-proof storage. The logic of using differentiation for common templates, tamper-proof for long-term and high-density templates, and archive layer for low-frequency templates is put into an executable mechanism, reducing redundant storage and improving long-term preservation and compliance auditing capabilities.
[0093] Key management is used to generate encryption keys to encrypt archive data. At the same time, fingerprint digests and watermarks are generated and written into the genealogy tracing data to achieve controllable permissions, key rotation, and accountability for leakage. The abnormality, prevalence, policy results and session records are correlated and output to form an end-to-end, verifiable remote management evidence chain.
[0094] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0095] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the protection scope of the present invention.
Claims
1. A method for remote management of unmanned archives based on 5G networks, characterized in that, Includes the following steps: Step S1: Establish an image transmission channel and a control audit channel based on the 5G network, and generate session record data; Step S2: Based on the session record data, perform session tagging and time alignment on the collected original archive image data, calculate image quality index data, generate standardized archive image data based on the image quality index data, obtain archive metadata corresponding to the original archive image data, and write the archive metadata into the session record data; Step S3: Extract layout structure point data from standardized archival image data and fit layout feature curve data, perform text recognition to obtain recognized text data and generate semantic feature data, and combine archival metadata to generate multidimensional index entry data. Step S4: Perform deterministic nearest neighbor retrieval on the multidimensional index entry data and output anomaly data and universality data; Step S5: Generate storage strategy data and genealogy tracing data based on prevalence data, archive metadata, and image quality index data; Step S6: Based on key management, generate encryption key data to encrypt the archive data output according to the storage policy data, generate fingerprint digest data and watermark data, and after associating anomaly data, prevalence data, storage policy data, genealogy tracing data and session record data, output remote management result data through the control audit channel.
2. The method for remote management of unmanned archives based on a 5G network as described in claim 1, characterized in that, Step S1 includes the following sub-steps: Step S101: Collect device identity data, which includes the identity of the scanning device, the identity of the access control device, and the identity of the edge node; Step S102: Establish an image transmission channel and a control audit channel based on device identity data; Step S103: Write the device identity data, timestamp, and channel identifier into the session record data.
3. The method for remote management of unmanned archives based on a 5G network as described in claim 2, characterized in that, Step S2 includes the following sub-steps: Step S201: Collect raw archive image data and time-align it with session record data to form raw archive image data with session tags; Step S202: Calculate image quality index data for the original archive image data and write the image quality index data into the session record data; Step S203: Obtain the archive metadata corresponding to the original archive image data and write the archive metadata into the session record data.
4. The method for remote management of unmanned archives based on a 5G network as described in claim 3, characterized in that, Standardized archival image data is obtained by standardizing the original archival image data based on image quality index data: Denoising parameters, skew correction parameters, and cropping parameters are generated based on image quality index data. The original archival image data is sequentially subjected to denoising, skew correction, text area cropping, and resolution normalization to output standardized archival image data. Write the denoising parameters, correction parameters, and clipping parameters into the session log data.
5. The method for remote management of unmanned archives based on a 5G network as described in claim 4, characterized in that, Step S3 includes the following sub-steps: Step S301: Locate the text region in the standardized archival image data and output the layout structure point data; Step S302: Generate layout feature sequence based on layout structure point data and fit to obtain layout feature curve data; Step S303: Perform character recognition on the standardized archival image data to obtain recognized text data, and generate semantic feature data based on the recognized text data; Step S304: Collect archive metadata and combine layout feature curve data, semantic feature data, archive metadata and image quality index data to generate multidimensional index entry data.
6. The method for remote management of unmanned archives based on a 5G network as described in claim 5, characterized in that, Step S4 includes the following sub-steps: Step S401: Write the layout feature curve data into the layout nearest neighbor index library and write the semantic feature data into the semantic nearest neighbor index library. Step S402: Perform deterministic nearest neighbor retrieval on the layout feature curve data based on the layout nearest neighbor index library to obtain the layout nearest neighbor set; perform deterministic nearest neighbor retrieval on the semantic feature data based on the semantic nearest neighbor index library to obtain the semantic nearest neighbor set. Step S403: Calculate the anomaly score data based on the layout nearest neighbor set, the semantic nearest neighbor set, and the image quality index data; Step S404: Calculate the universality data based on the layout nearest neighbor set, the semantic nearest neighbor set, and the archive metadata.
7. The method for remote management of unmanned archives based on a 5G network as described in claim 6, characterized in that, Step S5 includes the following sub-steps: Step S501: Generate storage strategy data based on the prevalence data and file metadata. The storage strategy data includes one or a combination of template differential storage strategy, hot and cold tiered storage strategy and tamper-proof storage strategy. Step S502: When the universality data meets the universality condition, execute the template differential storage strategy and record the differential relationship data; Step S503: When the archive metadata meets the requirements for long-term preservation or auditing, execute the tamper-proof storage policy and record the storage location data; Step S504: Associate the multidimensional index entry data, storage strategy data and their execution results to generate genealogy tracing data, and bind and store it with session record data.
8. The method for remote management of unmanned archives based on a 5G network as described in claim 7, characterized in that, Step S6 includes the following sub-steps: Step S601: Generate encryption key data for each file based on key management, and use the encryption key data to encrypt the file data output according to the storage strategy. Step S602: Generate fingerprint summary data based on layout feature curve data and semantic feature data, and write the fingerprint summary data into the genealogy tracing data; Step S603: Generate watermark data based on fingerprint digest data and embed it into the corresponding archive data; Step S604: Output remote management result data through the control audit channel. The remote management result data includes at least anomaly data, prevalence data, storage strategy data, and index identifiers for genealogy tracing data.
9. A remote management method for unmanned archives based on a 5G network as described in claim 8, characterized in that, The remote management result data, including remote closed-loop control logic, is output through the control audit channel. When the anomaly data meets the anomaly conditions, a rescan instruction is sent to the scanning device corresponding to the device identity data based on the session record data, and new original archive image data is sent back. When the image quality index data does not meet the quality conditions, the standardization processing parameters are adjusted based on the session record data and the standardized archive image data is regenerated. When the prevalence data meets the prevalence condition and the image quality index data meets the quality condition, the template differential storage strategy is triggered and the lineage tracing data is updated.
10. A remote management system for unmanned archives based on a 5G network, which is applied in the remote management method for unmanned archives based on a 5G network as described in any one of claims 1-9, characterized in that, It includes a session recording acquisition module, an image data processing module, a fitting curve construction module, a data nearest neighbor retrieval module, a storage strategy generation module, and a remote management and control module; The session recording acquisition module is used to establish an image transmission channel and a control audit channel based on the 5G network, and to generate session recording data; The image data processing module is used to perform session tagging and time alignment on the collected original archive image data based on session record data, calculate image quality index data, generate standardized archive image data based on the image quality index data, obtain archive metadata corresponding to the original archive image data, and write the archive metadata into the session record data. The fitting curve construction module is used to extract layout structure point data from standardized archival image data and fit layout feature curve data, perform text recognition to obtain recognized text data and generate semantic feature data, and combine archival metadata to generate multidimensional index entry data. The data nearest neighbor retrieval module is used to perform deterministic nearest neighbor retrieval on multidimensional index entry data and output anomaly data and generality data; The storage strategy generation module is used to generate storage strategy data and genealogy tracing data based on universality data, archive metadata and image quality index data. The remote management control module is used to generate encryption key data based on key management to encrypt the archive data output according to the storage strategy data, generate fingerprint digest data and watermark data, and associate anomaly data, prevalence data, storage strategy data, genealogy tracing data and session record data, and output remote management result data through the control audit channel.
Citation Information
Patent Citations
Multi-dimensional space storage management method and system for digital archives and storage medium
CN119759845A
Near real-time detection and classification of machine anomalies using machine learning and artificial intelligence
CA3128957A1
Hydraulic machinery feasibility research report grading review method based on large language model
CN121257516A
Intelligent AI-driven file digital full-process processing system
CN121330702A