An unmanned archive remote management method and system based on a 5G network
By establishing a dual-channel system for image transmission and control auditing in a 5G unattended archive, generating traceable session records, and performing standardized processing and multi-dimensional indexing driven by image quality indicators, the problems of reproducible processing and auditable archiving of low-quality scanned images are solved, achieving stable remote management and efficient archive management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUNAN SPIDER ROBOT TECH CO LTD
- Filing Date
- 2026-01-15
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies are insufficient to achieve remote closed-loop archiving management for reproducible processing, deterministic anomalies, prevalence assessment, and auditability of low-quality scanned images in 5G unattended archive scenarios.
A dual-channel audit system for image transmission and control based on 5G network is established to generate traceable session records. Standardization processing is driven by image quality indicators, and format and semantic features are extracted to form a multi-dimensional index. Deterministic nearest neighbor retrieval is used to output anomaly and prevalence. Combined with key management, fingerprint watermarking and genealogy tracing, reproducible image processing and auditable archiving are achieved.
It enables reproducible image processing in 5G unattended archives, stably triggers rescanning and strategy switching, and supports space-saving, easy-to-retrieve, strong auditing and accountability-based remote management.
Smart Images

Figure CN121542222B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of digital management, and in particular to a remote management method and system for an unmanned archives based on a 5G network. BACKGROUND
[0002] In a district or park level digital archives, night batch scanning and daytime remote auditing have become a common mode. Since the warehouse is unattended, the scanning equipment, access control and edge nodes need to return a large number of images and logs to the remote center through the 5G network, and accept the tasks, rescan and parameter adjustment issued remotely. The existing scheme often has three types of pain points: first, the image quality is affected by shadow, tilt, reflection, red stamp and paper crease, resulting in fluctuation of layout positioning and text recognition, and it is difficult to find and close-loop rescan the abnormal page; second, archiving and retrieval rely on a single dimension, and the retrieval efficiency is low and the risk of misfiling is high in the case of different businesses with the same template and different formats with the same business; third, storage and security are often out of touch with business, lacking reproducible processing parameters, auditable session links and complete lineage traceability, and it is difficult to meet the requirements of long-term preservation, tamper-proof and leakage accountability.
[0003] At present, the Chinese invention patent with the application number CN202510258100.2 discloses a multi-dimensional space storage management method and system for a digital archives, and a storage medium. The method comprises receiving an archive to be stored, scanning the archive to obtain a digital archive; the digital archive is a set of scanned images; for any digital archive, the images are read in sequence, the images are identified, the text box is positioned, and the feature curve of each image is extracted based on the positioned text box; the abnormality of any image is determined according to the feature curve, the abnormality of all images in the same digital archive is counted, and the universality of the digital archive is determined; and the digital archive is stored according to the universality. The present application scans the archives to obtain digital archives, extracts features from the digital archives, and then classifies and stores the archives, greatly improving the data retention time and reducing the data storage space.
[0004] The above-mentioned technology cannot realize reproducible processing, deterministic anomaly, universality evaluation and auditable remote closed-loop archiving management for low-quality scanned images in the 5G unmanned archives scenario. SUMMARY
[0005] The technical problem solved by the present application is that the existing technology cannot realize reproducible processing, deterministic anomaly, universality evaluation and auditable remote closed-loop archiving management for low-quality scanned images in the 5G unmanned archives scenario.
[0006] To solve the above technical problems, the present application provides the following technical solutions:
[0007] A remote management method of unmanned archives based on a 5G network, comprising the following steps:
[0008] Step S1, establishing an image transmission channel and a control audit channel based on a 5G network, and generating session record data;
[0009] Step S2, based on the session record data, marking the collected original archive image data with a session and aligning the time, calculating image quality index data, generating standardized archive image data according to the image quality index data, obtaining archive metadata corresponding to the original archive image data, and writing the archive metadata into the session record data;
[0010] Step S3, extracting layout structure point data and fitting layout feature curve data from the standardized archive image data, performing text recognition to obtain recognized text data and generating semantic feature data, and combining the archive metadata to generate multi-dimensional index item data;
[0011] Step S4, performing deterministic nearest neighbor search on the multi-dimensional index item data, and outputting abnormality data and universality data;
[0012] Step S5, generating storage strategy data and pedigree tracing data according to the universality data, the archive metadata and the image quality index data;
[0013] Step S6, generating encryption key data based on key management to encrypt the archive data output according to the storage strategy data, generating fingerprint summary data and watermark data, and associating the abnormality data, the universality data, the storage strategy data, the pedigree tracing data and the session record data, and outputting remote management result data through the control audit channel.
[0014] Preferably, the step S1 comprises the following sub-steps:
[0015] Step S101, collecting device identity data, the device identity data comprising scanning device identity, access control device identity and edge node identity;
[0016] Step S102, establishing an image transmission channel and a control audit channel based on the device identity data;
[0017] Step S103, writing the device identity data, the time stamp and the channel identifier into the session record data.
[0018] Preferably, the step S2 comprises the following sub-steps:
[0019] Step S201, collecting original archive image data and aligning the time with the session record data to form original archive image data with session marks;
[0020] Step S202, calculating image quality index data for the original archive image data, and writing the image quality index data into the session record data;
[0021] Step S203, obtaining archive metadata corresponding to the original archive image data, and writing the archive metadata into the session record data.
[0022] Preferably, the original archive image data is standardized to obtain standardized archive image data based on the image quality index data, and the standardized archive image data is:
[0023] The denoising parameter, the rectification parameter, and the cutting parameter are respectively generated according to the image quality index data;
[0024] The original archive image data is sequentially executed with denoising, tilt correction, center cutting, and resolution normalization, and the standardized archive image data is outputted;
[0025] The denoising parameter, the rectification parameter, and the cutting parameter are written into the session record data.
[0026] Preferably, the step S3 comprises the following sub-steps:
[0027] Step S301, locating a text region in the standardized archive image data and outputting layout structure point position data;
[0028] Step S302, generating a layout feature sequence based on the layout structure point position data and fitting to obtain layout feature curve data;
[0029] Step S303, performing character recognition on the standardized archive image data to obtain recognized text data, and generating semantic feature data based on the recognized text data;
[0030] Step S304, collecting archive metadata, and combining the layout feature curve data, the semantic feature data, the archive metadata, and the image quality index data to generate multi-dimensional index entry data.
[0031] Preferably, the step S4 comprises the following sub-steps:
[0032] Step S401, writing the layout feature curve data into a layout neighbor index library, and writing the semantic feature data into a semantic neighbor index library;
[0033] Step S402, performing deterministic neighbor retrieval on the layout feature curve data based on the layout neighbor index library to obtain a layout neighbor set, and performing deterministic neighbor retrieval on the semantic feature data based on the semantic neighbor index library to obtain a semantic neighbor set;
[0034] Step S403, calculating abnormality data according to the layout neighbor set, the semantic neighbor set, and the image quality index data;
[0035] Step S404, calculating the universality data according to the layout neighborhood set, the semantic neighborhood set, and in combination with the archive metadata.
[0036] Preferably, the step S5 comprises the following sub-steps:
[0037] Step S501, generating storage strategy data according to the universality data and the archive metadata, the storage strategy data comprising one or a combination of template differential storage strategy, cold-hot tiered storage strategy, and tamper-proof storage strategy;
[0038] Step S502, when the universality data meets the universality condition, executing the template differential storage strategy and recording the differential relationship data;
[0039] Step S503, when the archive metadata meets the long-term preservation or audit requirement, executing the tamper-proof storage strategy and recording the storage location data;
[0040] Step S504, associating the multi-dimensional index entry data, the storage strategy data, and the execution results thereof to generate the pedigree tracing data, and storing the pedigree tracing data in combination with the session record data.
[0041] Preferably, the step S6 comprises the following sub-steps:
[0042] Step S601, generating encryption key data for each archive based on key management, and encrypting the archive data output according to the storage strategy data by using the encryption key data;
[0043] Step S602, generating fingerprint summary data based on the layout feature curve data and the semantic feature data, and writing the fingerprint summary data into the pedigree tracing data;
[0044] Step S603, generating watermark data based on the fingerprint summary data and embedding the watermark data into the corresponding archive data;
[0045] Step S604, outputting remote management result data through the control audit channel, the remote management result data at least comprising the index identifier of the anomaly degree data, the universality data, the storage strategy data, and the pedigree tracing data.
[0046] Preferably, outputting the remote management result data through the control audit channel comprises remote closed-loop control logic:
[0047] When the anomaly degree data meets the anomaly condition, issuing a rescan instruction to the scanning device corresponding to the device identity data based on the session record data and returning new original archive image data;
[0048] When the image quality index data does not meet the quality condition, adjusting the standardized processing parameters based on the session record data and re-generating standardized archive image data;
[0049] When the prevalence data meets the prevalence condition and the image quality index data meets the quality condition, a template difference storage strategy is triggered and pedigree tracing data is updated.
[0050] The remote management system of the unmanned archive based on the 5G network comprises a session record acquisition module, an image data processing module, a fitting curve construction module, a data near neighbor search module, a storage strategy generation module and a remote management control module.
[0051] The session record acquisition module is used for establishing an image transmission channel and a control audit channel based on the 5G network and generating session record data.
[0052] The image data processing module is used for performing session marking and time alignment on the collected original archive image data based on the session record data, calculating image quality index data, generating standardized archive image data according to the image quality index data, obtaining archive metadata corresponding to the original archive image data and writing the archive metadata into the session record data.
[0053] The fitting curve construction module is used for extracting layout structure point data from the standardized archive image data and fitting layout feature curve data, obtaining recognized text data through character recognition and generating semantic feature data, and generating multi-dimensional index item data in combination with the archive metadata.
[0054] The data near neighbor search module is used for performing deterministic near neighbor search on the multi-dimensional index item data and outputting abnormality degree data and prevalence data.
[0055] The storage strategy generation module is used for generating storage strategy data and pedigree tracing data according to the prevalence data, the archive metadata and the image quality index data.
[0056] The remote management control module is used for generating encryption key data based on key management to encrypt the archive data output according to the storage strategy data, generating fingerprint summary data and watermark data, and associating the abnormality degree data, the prevalence data, the storage strategy data, the pedigree tracing data and the session record data to output remote management result data through the control audit channel.
[0057] The present application has the following advantages: the present application is directed to the 5G unmanned archive, establishes an image transmission channel and a control audit channel and generates traceable session records, drives image standardization with quality indexes such as definition, inclination and shadow and ensures that parameters are reproducible, extracts layout feature curves, semantic features and archive metadata to form multi-dimensional indexes, outputs abnormality degree and prevalence through deterministic near neighbor search, stably triggers rescan and strategy switching, drives template difference and tamper-proof storage with prevalence, and realizes space saving, easy search, strong audit and traceability in combination with key management encryption, fingerprint watermarking and pedigree tracing. BRIEF DESCRIPTION OF DRAWINGS
[0058] Figure 1 A step flow chart of a remote management method of an unmanned archive based on a 5G network is provided for an embodiment of the present application.
[0059] Figure 2 A basic flow schematic diagram of a remote management system of an unmanned archive based on a 5G network is provided for an embodiment of the present application. DETAILED DESCRIPTION
[0060] In order to make the above objectives, features and advantages of the present application more obvious and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments.
[0061] Embodiment 1, with reference to Figure 1 provides a remote management method of an unmanned archive based on a 5G network. The embodiment is aimed at a county-level unmanned archive scene: the warehouse is unattended at night, the paper archives are automatically scanned in batches by the scanning device, the edge node (used for processing and network buffering nearby) is arranged in the warehouse, the remote management end is arranged in the municipal data center, the warehouse side is connected through the 5G private network, and the image transmission and control audit two types of communication are distinguished to ensure that the large amount of image uploading and the control audit do not occupy the bandwidth, and the method comprises the following steps:
[0062] Step S1, establishing an image transmission channel and a control audit channel based on a 5G network, and generating session record data.
[0063] Step S2, based on the session record data, marking the collected original archive image data with a session and aligning the time, calculating image quality index data, generating standardized archive image data according to the image quality index data, obtaining archive metadata corresponding to the original archive image data, and writing the archive metadata into the session record data.
[0064] Step S3, extracting layout structure point data and fitting layout feature curve data from the standardized archive image data, performing text recognition to obtain recognized text data and generating semantic feature data, and combining the archive metadata to generate multi-dimensional index item data.
[0065] Step S4, performing deterministic nearest neighbor search on the multi-dimensional index item data, and outputting abnormality data and universality data.
[0066] Step S5, generating storage strategy data and pedigree tracing data according to the universality data, the archive metadata and the image quality index data.
[0067] Step S6, based on the key management, the encrypted key data is generated to encrypt the archive data output according to the storage strategy data, the fingerprint summary data and the watermark data are generated, and the abnormality data, the universality data, the storage strategy data, the pedigree tracing data and the session record data are associated and output through the control audit channel to output the remote management result data.
[0068] The application faces the 5G unmanned archive, establishes the image transmission and control audit double channels, generates the traceable session record, drives the image standardization by the quality indexes such as the definition, the inclination and the shadow, guarantees the parameter reproducibility, extracts the layout feature curve, the semantic feature and the archive metadata to form the multi-dimensional index, adopts the deterministic neighbor search to output the abnormality and the universality, stably triggers the rescan and the strategy switching, drives the template difference and the non-tamperable storage by the universality, and realizes the space saving, the easy search, the strong audit and the traceability by combining the key management encryption, the fingerprint watermark and the pedigree tracing.
[0069] Step S1 includes the following sub-steps:
[0070] Step S101, device identity data is collected, and the device identity data includes scanning device identity, access control device identity and edge node identity.
[0071] The warehouse side scanning device, access control device and edge node register the device identity data to the remote management end after power on, and the device identity data at least includes device number, device type, belonging warehouse number, firmware version and certificate serial number, the remote management end checks the device certificate, issues a session token for the device and establishes session record data, and the session record data at least records the device identity data, the session token, the session start time, the channel identifier, the scanning task number and the batch number, and the session record data is used for:
[0072] Each scanning image is marked with a session mark to ensure that the image corresponds to the device, task and time one by one;
[0073] As the basis for the instruction signature and the receipt of the control audit channel;
[0074] As the basis for audit tracing and parameter reproduction, the same batch can restore the processing parameters at the time of rescan or dispute review.
[0075] Step S102, based on the device identity data, the image transmission channel and the control audit channel are established.
[0076] The image transmission channel is used to upload the original archive image and the standardized archive image, and the control audit channel is used to issue scanning tasks, rescan instructions, parameter adjustment instructions, and upload abnormal alarm, audit log and pedigree tracing information. Both types of channels are authenticated by session tokens, and the time of issuing each instruction, the execution result and the receipt summary are recorded in the session record data to avoid the problem that the image is taken but the person who takes it, the time when it is taken and whether it is changed are unknown when unattended.
[0077] Step S103, write the device identity data, timestamp and channel identifier into the session record data.
[0078] Step S2 includes the following sub-steps:
[0079] Step S201, collect the original archive image data and time align with the session record data to form the original archive image data with session marks.
[0080] The scanning device performs scanning according to the scanning task number issued remotely to obtain the original archive image data. The edge node immediately writes the session mark, page number and collection time in the session record data into the header index information of the original archive image data after receiving, and indexes the original archive image data back to the remote management end, so that the remote end can see the progress and page number missing even if it does not receive all the images.
[0081] Step S202, calculate the image quality index data of the original archive image data, and write the image quality index data into the session record data.
[0082] The edge node calculates the image quality index data of each original archive image, and writes the quality index data into the session record data and the page image index, which is used to drive the standardization processing parameter and participate in the abnormality degree judgment (when the quality is poor and the recognition is unstable, the abnormality degree calculation will be punished to avoid misjudging the poor quality as content abnormality). The image quality index data of the embodiment at least includes the definition index, the inclination angle index, the shadow ratio index and the text proportion index, and the calculation method is as follows:
[0083] Definition index Laplacian variance method is adopted: first, the grayscale image of the single page image before standardization is obtained by grayscale , the edge response graph is obtained by performing Laplacian operator on the grayscale image , and then the variance is calculated. The larger the variance, the clearer it is. It can be recorded as:
[0084] ;
[0085] , wherein, represents the second-order edge operator. The clarity index will be normalized to the interval 0~1, and the normalization method is to subtract the minimum value from the current value and then divide by the range. The minimum value and the range come from the statistical benchmark of the historical batches of the warehouse and are written in the session record data.
[0086] The skew angle index adopts text line direction detection: extract long straight lines (table lines or text line baselines) from the binary image, and count the main direction angle as the skew angle. The larger the absolute value of the skew angle, the more serious the skew.
[0087] The shadow ratio index adopts background brightness graph decomposition: estimate the background brightness graph with a large-scale filter , and then count the proportion of the area with brightness lower than the shadow threshold in the whole page. The shadow threshold is automatically determined by the background brightness distribution of the batch, and is written in the session record data. The specific calculation method is:
[0088] Define the shadow region set :
[0089] ;
[0090] The shadow ratio index is defined as:
[0091] ;
[0092] wherein, is the shadow threshold, is the whole image pixel set, is the number of pixels,
[0093] The text proportion index adopts connected domain screening: do connected domain analysis on the binary image, screen out too large connected domains (such as whole black edges) and too small noise points, and divide the total number of remaining connected domain pixels by the total number of pixels in the whole page to obtain the text proportion. The text proportion is used to avoid the influence of blank pages, cover pictures, and photo pages on the layout curve and participate in the abnormality punishment.
[0094] In step S203, the archive metadata corresponding to the original archive image data is obtained, and the archive metadata is written into the session record data.
[0095] The standardized archive image data obtained by standardizing the original archive image data based on the image quality index data is:
[0096] The denoising parameter, the skew correction parameter, and the cutting parameter are generated according to the image quality index data, respectively.
[0097] The original archive image data is sequentially executed denoising, skew correction, layout cutting, and resolution normalization, and the standardized archive image data is output.
[0098] Write the denoising parameter, the rectification parameter and the cutting parameter into the session record data.
[0099] The edge node generates the standardization processing parameter based on the quality index and performs the standardization processing, outputs the standardization archive image data, and the standardization processing includes denoising, tilt correction, cutting and resolution normalization. The embodiment requires writing the denoising intensity, the rectification angle, the cutting boundary and the normalized resolution into the session record data, and if the re-scanning or manual review is triggered subsequently, the same set of processing procedures can be replayed according to the session record data.
[0100] Step S3 includes the following sub-steps:
[0101] Step S301, locate the text area in the standardization archive image data and output the layout structure point position data.
[0102] The edge node locates the text area in the standardization archive image and obtains the layout structure point position data. The robust scheme of connected domain merging and projection segmentation is adopted in the embodiment:
[0103] First, extract the candidate text connected domain on the binary image and merge it into a text block according to the distance, and then do horizontal and vertical projection segmentation on the text block to correct the boundary, so as to obtain the circumscribed rectangle of each text block;
[0104] The four corner points and the center point of each rectangle together constitute the layout structure point position data. The located point position is not to look like a text box, but to fit the layout feature curve which is more robust to scanning noise in the subsequent.
[0105] Step S302, generate the layout feature sequence based on the layout structure point position data and fit to obtain the layout feature curve data.
[0106] The edge node fits the layout feature curve data based on the layout structure point position data. The diagonal projection histogram curve is adopted in the embodiment to reduce the sensitivity to slight cutting and local noise:
[0107] Map each point position to the projection coordinate in the left-up to right-down diagonal direction (add the horizontal and vertical coordinates of the point position and then normalize according to the page width and page height), and then count the number of all point positions falling into a fixed number of intervals according to the projection coordinate, to form an interval counting curve;
[0108] Smooth the curve and normalize it to the interval of 0~1 to obtain the layout feature curve data. The layout feature curve data is essentially a stable signature of the layout structure, which is more resistant to noise and more convenient for neighbor search than directly storing the text box coordinates.
[0109] Step S303, perform text recognition on the standardization archive image data to obtain recognized text data, and generate semantic feature data based on the recognized text data.
[0110] The edge node performs text recognition on the standardized archive image to obtain recognized text data and generate semantic feature data. This embodiment is landable and interpretable, and adopts two layers of semantic features:
[0111] The first layer is a keyword vector (high-frequency words are counted after word segmentation and stop words removal from the recognized text and are encoded according to weights);
[0112] The second layer is a theme vector (a theme distribution is obtained by dimension reduction of a sentence segment vector).
[0113] If a deep semantic model is deployed, the recognized text can also be sent to a sentence vector model to obtain a semantic vector, but regardless of which model is adopted, the model version number, word table version number or theme model version number is required to be written into session record data.
[0114] In step S304, archive metadata is collected, and layout feature curve data, semantic feature data, archive metadata and image quality indicator data are combined to generate multi-dimensional index entry data.
[0115] Archive metadata is obtained and combined with the aforementioned data to generate multi-dimensional index entry data. The archive metadata is not obtained in a vacuum, and this embodiment can be used alone or in combination with verification:
[0116] First, the remote management end carries a task account field (such as archive number, fond number, directory number, secret level, storage period, formation date, batch number) when issuing a scanning task, and the edge node directly writes the session record data and binds each page;
[0117] Second, the edge node queries the library account from the archive management system according to the archive number and backfills;
[0118] Third, the fixed fields on the first page are structurally extracted to generate parsed account fields, and consistency verification is performed with the previous two sources (if inconsistent, an abnormality degree is triggered and a review work order is generated).
[0119] Finally, the multi-dimensional index entry data at least includes: session record data reference, page number, image quality indicator data, layout feature curve data, semantic feature data and archive metadata.
[0120] Step S4 includes the following sub-steps:
[0121] In step S401, layout feature curve data is written into a layout neighbor index library, and semantic feature data is written into a semantic neighbor index library.
[0122] The edge node writes the layout feature curve data of the historical batch into a layout neighbor index library, and writes the semantic feature data of the historical batch into a semantic neighbor index library. The index library adopts a deterministic neighbor structure, avoiding different scores caused by random sampling comparison of the same page multiple times. The embodiment can adopt two types of indexes in engineering:
[0123] The layout neighbor index library adopts curve fingerprint bucketing and accurate comparison within the bucket;
[0124] The semantic neighbor index library adopts a vector neighbor index (for example, a graph structure neighbor).
[0125] Regardless of the structure adopted, the index construction process is required to output an index version number and a batch range written into the library, and write session record data to ensure traceability.
[0126] In step S402, a deterministic neighbor search is performed on the layout feature curve data based on the layout neighbor index library to obtain a layout neighbor set, and a deterministic neighbor search is performed on the semantic feature data based on the semantic neighbor index library to obtain a semantic neighbor set.
[0127] For the layout feature curve data of the current page, the curve fingerprint is used to locate a candidate set, and then accurate similarity calculation is performed on the candidate set to obtain the layout neighbor set. For the semantic feature data of the current page, vector neighbor search is directly performed to obtain the semantic neighbor set. The similarity adopts cosine similarity, which is limited between 0 and 1. The closer to 1, the more similar.
[0128] In step S403, the abnormality degree data is calculated according to the layout neighbor set, the semantic neighbor set, and the image quality index data.
[0129] The abnormality degree data is used to judge whether re-scanning or re-checking is needed. The prediction logic is that the most similar neighbor is still not similar enough and the quality is punished. The specific method is as follows:
[0130] The maximum similarity is taken from the layout neighbor set and the semantic neighbor set, respectively, and the comprehensive maximum similarity is obtained by weight fusion;
[0131] The lower the comprehensive maximum similarity, the higher the abnormality degree.
[0132] At the same time, if the image quality index data indicates that the definition is poor, the shadow is large, or the inclination is serious, a penalty term is added to the abnormality degree to avoid mistaking the recognition deviation caused by poor quality as an abnormality of the archive content. This section is the prediction core of the method: it does not predict future time series, but predicts whether the current page is in an abnormal state, and uses the abnormal state for closed-loop control.
[0133] In step S404, the universality degree data is calculated according to the layout neighbor set, the semantic neighbor set, and the archive metadata.
[0134] Prevalence data is used to drive storage strategies (template differentiation, hot / cold tiering), and its prediction logic is based on how common something is in the structural and semantic space. Specifically:
[0135] The average similarity of several most similar nearest neighbors is taken as the basic value of universality;
[0136] The similarity is corrected by incorporating category constraints from the archive's metadata. For example, similarity within the same archival group or with the same catalog number is given higher weight, while similarity across different archival groups is given lower weight. This ensures that the generality reflects not only the similarity of the layout but also the consistency of the business affiliation. Higher generality indicates that the archive is more likely to belong to a common template or common business type, making it more suitable for template differential storage to save space.
[0137] Step S5 includes the following sub-steps:
[0138] Step S501: Generate storage strategy data based on generality data and file metadata. The storage strategy data includes one or a combination of template differential storage strategy, hot and cold tiered storage strategy, and tamper-proof storage strategy.
[0139] Edge nodes generate storage strategy data based on prevalence data, archive metadata, and image quality metrics. Storage strategies include at least one or a combination of three types: template differential storage strategy, hot / cold tiered storage strategy, and tamper-proof storage strategy. Strategy generation employs an auditable threshold lookup method: for example, images with high prevalence and low security level are processed via template differential and placed in the hot tier for frequent retrieval; images with low prevalence but long retention periods or high security levels are stored using tamper-proof storage and placed in the cold archive tier; images with poor quality but still requiring storage are forcibly retained in their original form with a rescan recommendation recorded. The threshold table and strategy version number are written into the session log data.
[0140] Step S502: When the universality data meets the universality condition, execute the template differential storage strategy and record the differential relationship data.
[0141] When the template differential storage strategy is triggered, the edge node first selects a representative template page based on the nearest neighbor set (generally, the nearest neighbor with the highest similarity is chosen as the template). Then, it aligns the current page with the template page at the text block level and outputs differential relationship data. The differential relationship data includes: the position information of each text block, whether it has been changed, and a summary of the image or text fragment of the changed area. During storage, only the template page reference and differential relationship data are stored, instead of storing the entire page repeatedly. This significantly saves space for template files with high prevalence. During retrieval or browsing, the data is reconstructed and displayed based on the template page and differential relationship data. The reconstruction process is also recorded in the genealogy tracing.
[0142] Step S503: When the archive metadata meets the requirements for long-term preservation or auditing, implement an immutable storage policy and record the storage location data.
[0143] Regardless of the storage strategy adopted, the edge node writes the multi-dimensional index entry data, storage strategy data, storage location data, template reference relationship, differential relationship data, session record data reference in chronological order into the lineage tracing data. The lineage tracing data is synchronized to the remote management terminal through the control audit channel, meeting the rigid requirements of the unmanned archives for audit and accountability.
[0144] Step S504, the multi-dimensional index entry data, storage strategy data and its execution result are associated to generate lineage tracing data, and are stored in association with the session record data.
[0145] Step S6 includes the following sub-steps:
[0146] Step S601, based on key management, an encryption key data is generated for each archive, and the archive data output according to the storage strategy data is encrypted using the encryption key data.
[0147] The edge node calls the key management service to generate encryption key data for each archive, and uses envelope encryption to encrypt the archive data: the archive content is encrypted using the encryption key data and stored, and the encryption key data is encapsulated and saved using the master key of the key management service. In this way, key rotation, revocation and permission audit can be achieved, and the instability or un-auditability caused by directly using the layout feature curve as the key source can be avoided.
[0148] Step S602, based on the layout feature curve data and the semantic feature data, a fingerprint summary data is generated, and the fingerprint summary data is written into the lineage tracing data.
[0149] The fingerprint summary data is used for leakage tracing and consistency checking. In the embodiment, the normalized sequence summary of the layout feature curve data, the summary of the semantic feature data, the archive number and batch number summary in the archive metadata are spliced to obtain the fingerprint summary data :
[0150] ;
[0151] wherein, is the normalized sequence summary of the layout feature curve data, is the summary of the semantic feature data, is the archive number in the archive metadata, is the batch number summary in the archive metadata, is a secure hash function, is a splicing operation.
[0152] The semantic digest is a short string of the first several weighted words and their weight codes of the semantic feature data. This has the advantages that the fingerprint contains both layout stability information and semantic differentiation information, and is bound to a specific archive number, facilitating accountability and comparison.
[0153] In step S603, watermark data is generated based on the fingerprint digest data and embedded in the corresponding archive data.
[0154] The watermark data is derived from the fingerprint digest data as a bit sequence, and then embedded in the frequency domain coefficients of the archive image to resist common compression and scaling. In engineering, block discrete cosine transform embedding can be used: the normalized archive image is divided into several blocks, and for each block, a discrete cosine transform is performed, a small increase or decrease is made on the watermark bit in the low-frequency coefficient, and the embedding strength is associated with the image quality index data (the clearer the image, the weaker the embedding to reduce visual impact, and the worse the image, the embedding can be appropriately enhanced to ensure detectability). The embedding position and embedding strength parameters are also written into the session record data, facilitating subsequent verification and legal evidence reproduction.
[0155] In step S604, the remote management result data is output through the control audit channel, and the remote management result data at least includes the index identification of the abnormality data, the universality data, the storage strategy data and the pedigree tracing data.
[0156] The output of the remote management result data through the control audit channel includes remote closed-loop control logic:
[0157] When the abnormality data meets the abnormality condition, a rescan instruction is issued to the scanning device corresponding to the device identity data based on the session record data, and the new original archive image data is returned.
[0158] When the image quality index data does not meet the quality condition, the standardized processing parameters are adjusted based on the session record data, and the standardized archive image data is regenerated.
[0159] When the universality data meets the universality condition and the image quality index data meets the quality condition, the template differential storage strategy is triggered and the pedigree tracing data is updated.
[0160] The edge node returns the abnormality data, the universality data, the storage strategy data, the pedigree tracing data index, the encrypted result digest, the fingerprint digest data and the watermark embedding result receipt as the remote management result data through the control audit channel. The remote management end performs closed-loop control according to the remote management result data: if the abnormality is above the threshold, a rescan instruction is automatically issued and the standardized processing parameters that need to be improved are specified; if the image quality index does not meet the standard, a device maintenance work order is issued; if the universality is high, the template differential is preferentially enabled and batch storage is accelerated for the same template batch.
[0161] Embodiment 2, refer to Figure 2The application provides a 5G network-based unmanned archive remote management system, which comprises a session record acquisition module, an image data processing module, a fitting curve construction module, a data near neighbor search module, a storage strategy generation module and a remote management control module.
[0162] The session record acquisition module is used for establishing an image transmission channel and a control audit channel based on the 5G network and generating session record data.
[0163] The image data processing module is used for performing session marking and time alignment on the collected original archive image data based on the session record data, calculating image quality index data, generating standardized archive image data according to the image quality index data, obtaining archive metadata corresponding to the original archive image data and writing the archive metadata into the session record data.
[0164] The fitting curve construction module is used for extracting layout structure point data from the standardized archive image data and fitting layout feature curve data, performing text recognition to obtain recognized text data and generating semantic feature data, and combining the archive metadata to generate multi-dimensional index item data.
[0165] The data near neighbor search module is used for performing deterministic near neighbor search on the multi-dimensional index item data and outputting abnormality degree data and universality degree data.
[0166] The storage strategy generation module is used for generating storage strategy data and pedigree tracing data according to the universality degree data, the archive metadata and the image quality index data.
[0167] The remote management control module is used for generating encryption key data based on key management to encrypt the archive data output according to the storage strategy data, generating fingerprint summary data and watermark data, and associating the abnormality degree data, the universality degree data, the storage strategy data, the pedigree tracing data and the session record data to output remote management result data through the control audit channel.
[0168] Through the separation of the image transmission channel and the control audit channel, the session record data is used for the whole process of collection and alignment, instruction issuing, receipt recording and audit tracing, the problems of mutual occupation of image return and control audit and untraceable process during unattended operation are solved, the remote center can complete rescan, parameter adjustment and work order closed loop without interrupting image uploading.
[0169] The image quality indicators such as sharpness, inclination angle, shadow ratio and text proportion are constructed, and the quality indicators are used to drive denoising, deviation correction, cutting and resolution normalization, so that the standardized processing has reproducible parameters, the layout drift and recognition fluctuation across devices and batches are significantly reduced, and misjudgment and omission are reduced.
[0170] At the same time, the layout structure point is extracted and the layout feature curve is fitted, the semantic feature is generated by combining the text recognition text, and the multi-dimensional index entry is formed by introducing the archive metadata, which avoids the confusion of the same template different business and the same business different format caused by relying on template similarity or full-text retrieval only, improves the positioning efficiency and the consistency of archiving.
[0171] The layout and semantic similarity is calculated by using deterministic neighbor search, the abnormality and universality are obtained, and the image quality penalty term is introduced to distinguish the quality difference and content anomaly, which avoids the unstable results caused by random sampling comparison, so that the re-scanning, re-checking and strategy switching are triggered more reliably.
[0172] The storage strategy is generated according to the universality, metadata and quality index, which supports the strategy selection of template difference, hot and cold layering and non-tamperable storage. The logic of template common difference, long-term, high-density level non-tamperable, low heat archiving layer is put into executable mechanism, which reduces the repeated storage, improves the long-term preservation and compliance audit ability.
[0173] The encryption key is generated by using key management to encrypt the archive data, and the fingerprint summary and watermark are generated and written into the pedigree tracing data, which realizes the controllable permission, the replaceable key, the traceable leakage, and the association output of the abnormality, the universality, the strategy result and the session record, forming the end-to-end, verifiable remote management evidence chain.
[0174] Those skilled in the art will appreciate that embodiments of the present application can be readily used as a method, a system or a computer program product. Accordingly, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product on one or more computer-usable storage media (or computer- readable storage media) having computer-usable program code embodied in the medium. The medium can be any available medium or combination thereof that is accessible by a general purpose or special purpose computer. By way of example, such computer-usable storage media can include a volatile memory, such as a random access memory (RAM), a non-volatile memory, such as a read-only memory (ROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic disk, a flash memory, a compact disk (CD) or a digital versatile disk (DVD). The computer-usable program code can include any suitable set of instructions, statements or Figure 1 one or more flows and / or blocks Figure 1 one or more flows and / or blocks
[0175] It should be noted that the above-mentioned embodiments are only used to illustrate but not to limit the technical solutions of the present application. Although the present application is described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or replaced equivalently without departing from the spirit and scope of the technical solutions of the present application, and they should be covered in the protection scope of the present application.
Claims
1. A 5G network-based unmanned archive remote management method, characterized by, The method comprises the following steps: Step S1, establishing an image transmission channel and a control audit channel based on a 5G network, and generating session record data; Step S2, marking the collected original archive image data with a session based on the session record data, and time aligning, calculating image quality index data, generating standardized archive image data according to the image quality index data, obtaining archive metadata corresponding to the original archive image data, and writing the archive metadata into the session record data; Step S3, extracting layout structure point data and fitting layout feature curve data from the standardized archive image data, performing text recognition to obtain recognized text data and generating semantic feature data, and combining the archive metadata to generate multi-dimensional index item data; Step S4, performing deterministic nearest neighbor search on the multi-dimensional index item data, and outputting abnormality data and universality data; Step S5, generating storage strategy data and pedigree tracing data according to the universality data, the archive metadata and the image quality index data; Step S6, generating encryption key data based on key management to encrypt the archive data output according to the storage strategy data, generating fingerprint summary data and watermark data, and associating the abnormality data, the universality data, the storage strategy data, the pedigree tracing data and the session record data to output remote management result data through the control audit channel.
2. The unmanned archive remote management method based on a 5G network according to claim 1, wherein, The step S1 comprises the following sub-steps: Step S101, collecting device identity data, the device identity data comprising scanning device identity, access control device identity and edge node identity; Step S102, establishing an image transmission channel and a control audit channel based on the device identity data; Step S103, writing the device identity data, a timestamp and a channel identifier into the session record data.
3. The unmanned archive remote management method based on a 5G network according to claim 2, characterized in that, The step S2 comprises the following sub-steps: Step S201, collecting original archive image data and time aligning with the session record data to form original archive image data with session marks; Step S202, calculating image quality index data for the original archive image data, and writing the image quality index data into the session record data; Step S203, obtaining archive metadata corresponding to the original archive image data, and writing the archive metadata into the session record data.
4. The unmanned archive remote management method based on a 5G network according to claim 3, characterized in that, The original archive image data is standardized to obtain standardized archive image data based on the image quality index data: Denoising parameters, rectification parameters and cutting parameters are respectively generated according to the image quality index data; The original archive image data is sequentially executed with denoising, tilt correction, layout center cutting and resolution normalization, and the standardized archive image data is outputted; The denoising parameters, the rectification parameters and the cutting parameters are written into the session record data.
5. The unmanned archive remote management method based on a 5G network according to claim 4, characterized in that, The step S3 comprises the following sub-steps: Step S301, locating a text area in the standardized archive image data and outputting layout structure point data; Step S302, generating a layout feature sequence based on the layout structure point data and fitting to obtain layout feature curve data; Step S303, performing text recognition on the standardized archive image data to obtain recognized text data, and generating semantic feature data based on the recognized text data; Step S304, collect the archive metadata, and combine the layout feature curve data, semantic feature data, archive metadata, and image quality index data to generate multi-dimensional index entry data.
6. The unmanned archive remote management method based on a 5G network according to claim 5, characterized in that, The step S4 includes the following sub-steps: Step S401, write the layout feature curve data into the layout neighbor index library, and write the semantic feature data into the semantic neighbor index library; Step S402, perform deterministic neighbor search on the layout feature curve data based on the layout neighbor index library to obtain a layout neighbor set, and perform deterministic neighbor search on the semantic feature data based on the semantic neighbor index library to obtain a semantic neighbor set; Step S403, calculate abnormality data according to the layout neighbor set, the semantic neighbor set, and the image quality index data; Step S404, calculate universality data according to the layout neighbor set, the semantic neighbor set, and the archive metadata.
7. The unmanned archive remote management method based on a 5G network according to claim 6, characterized in that, The step S5 includes the following sub-steps: Step S501, generate storage strategy data according to the universality data and the archive metadata, the storage strategy data including one or a combination of template differential storage strategy, hot and cold layered storage strategy, and tamper-proof storage strategy; Step S502, when the universality data meets the universality condition, execute the template differential storage strategy and record the differential relationship data; Step S503, when the archive metadata meets the long-term preservation or audit requirement, execute the tamper-proof storage strategy and record the storage location data; Step S504, associate the multi-dimensional index entry data, the storage strategy data, and the execution results thereof to generate pedigree tracing data, and store the pedigree tracing data in association with the session record data.
8. The unmanned archive remote management method based on a 5G network according to claim 7, characterized in that, The step S6 includes the following sub-steps: Step S601, generate encryption key data for each archive based on key management, and encrypt the archive data output according to the storage strategy data using the encryption key data; Step S602, generate fingerprint summary data based on the layout feature curve data and the semantic feature data, and write the fingerprint summary data into the pedigree tracing data; Step S603, generate watermark data based on the fingerprint summary data and embed the watermark data into the corresponding archive data; Step S604, output remote management result data through the control audit channel, the remote management result data including at least the index identifier of the abnormality data, the universality data, the storage strategy data, and the pedigree tracing data. 9.The 5G network-based remote management method of unmanned archives room according to claim 8, wherein, Outputting the remote management result data through the control audit channel includes remote closed-loop control logic: When the abnormality data meets the abnormality condition, issue a rescan instruction to the scanning device corresponding to the device identity data based on the session record data and return new original archive image data; When the image quality index data does not meet the quality condition, adjust the standardized processing parameters based on the session record data and regenerate the standardized archive image data; When the universality data meets the universality condition and the image quality index data meets the quality condition, trigger the template differential storage strategy and update the pedigree tracing data.
10. The unmanned archive remote management system based on a 5G network is applied to the unmanned archive remote management method based on a 5G network in any one of claims 1-9, characterized in that, The system includes a session record collection module, an image data processing module, a fitting curve construction module, a data neighbor search module, a storage strategy generation module, and a remote management control module; The session record collection module is configured to establish an image transmission channel and a control audit channel based on a 5G network, and generate session record data; The image data processing module is configured to mark and time-align the collected original archive image data based on the session record data, calculate image quality index data, generate standardized archive image data according to the image quality index data, obtain archive metadata corresponding to the original archive image data, and write the archive metadata into the session record data; The fitting curve construction module is configured to extract layout structure point data from the standardized archive image data and fit layout feature curve data, perform text recognition to obtain recognized text data and generate semantic feature data, and generate multi-dimensional index item data in combination with the archive metadata; The data near neighbor search module is configured to perform deterministic near neighbor search on the multi-dimensional index item data, and output abnormality data and universality data; The storage strategy generation module is configured to generate storage strategy data and pedigree tracing data according to the universality data, the archive metadata, and the image quality index data; The remote management control module is configured to generate encrypted key data based on key management to encrypt archive data output according to the storage strategy data, generate fingerprint summary data and watermark data, and output remote management result data through the control audit channel after associating the abnormality data, the universality data, the storage strategy data, the pedigree tracing data, and the session record data.
Citation Information
Patent Citations
Multi-dimensional space storage management method and system for digital archives and storage medium
CN119759845A
Hydraulic machinery feasibility research report grading review method based on large language model
CN121257516A
Intelligent AI-driven file digital full-process processing system
CN121330702A