Intelligent archive digitization processing method and system
By acquiring and processing archival image data in the visible and near-infrared bands using a multispectral linear array scanner, the imaging quality problem of complex-format archives was solved, enabling the high-fidelity generation of fully structured digital archives and improving the quality of digital processing and information analysis capabilities of archives.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUANGZHOU XIEZHENG INFORMATION TECH CO LTD
- Filing Date
- 2025-12-25
- Publication Date
- 2026-07-10
AI Technical Summary
Existing technologies suffer from poor image quality when dealing with complex layout archives, making it difficult to generate machine-readable, semantically related structured data. Furthermore, they lack a deep understanding and analysis of the layout's logical structure, resulting in poor structure of digitized outputs that fail to meet the needs of knowledge-based archiving and intelligent retrieval of high-value archives.
A multispectral linear array scanner is used to simultaneously acquire archival image data in the visible and near-infrared bands. Through non-uniform illumination correction and perspective distortion correction, a dual-channel fused archival image with geometric and color standardization is generated. A multimodal feature fusion network is used for page region segmentation and semantic classification. Combined with a joint entity recognition and relation extraction model, the text content is parsed to generate a fully structured digital archival object.
It achieves high-fidelity and automated generation of fully structured digital archives, improving the accuracy, completeness, and knowledge content of digital processing, and effectively analyzing the layout logic structure and key information of complex archives.
Smart Images

Figure CN121502056B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of digital archives technology, specifically a method and system for intelligent digital archive processing. Background Technology
[0002] With the deepening of archival informatization, the digitization of traditional paper archives has become a crucial link in information preservation, management, and utilization. Current mainstream methods primarily employ high-speed scanners to acquire images of archives and rely on optical character recognition (OCR) technology to extract text content. However, these methods have significant limitations when dealing with archives with complex layouts: firstly, for paper archives with uneven lighting, binding distortion, stains, or fading, poor image quality severely affects the accuracy of subsequent processing; secondly, existing technologies typically treat the OCR-generated text as a single linear sequence, lacking a deep understanding and analysis of the logical structure of the layout (such as tables, signatures, and mixed text / image areas), resulting in poor structure of the digitized output and difficulty in directly forming machine-readable, semantically related structured data. Furthermore, existing processes primarily focus on extracting text content, lacking sufficient automated identification and indexing capabilities for key entities and their relationships within the archives, making it difficult to meet the deeper needs of knowledge-based archiving and intelligent retrieval of high-value archives. Summary of the Invention
[0003] The purpose of this invention is to provide an intelligent method and system for digitizing archives, which addresses the shortcomings of existing technologies and can automatically and faithfully generate fully structured digital archive objects, thereby improving the accuracy, completeness, and knowledge content of digitization.
[0004] One embodiment of this application provides an intelligent archive digitization processing method, the method comprising:
[0005] Paper archives are scanned simultaneously using a multispectral linear array scanner, and archival image data in the visible light band and near-infrared band are collected respectively.
[0006] Non-uniform lighting correction and perspective distortion correction are performed on the archival image data to generate a geometrically and color-normalized dual-channel fused archival image;
[0007] Based on the dual-channel fused archive image, a multimodal feature fusion network is used to perform page region segmentation and semantic classification, outputting structured page description information including text area, table area, signature area and illustration area;
[0008] Based on the structured layout description information, a joint entity recognition and relation extraction model is used to parse the text content, and key field entities and their semantic relationships are extracted to generate a preliminary archival information graph.
[0009] Based on the aforementioned archival information map and the preset archival metadata specifications, a content importance assessment module driven by an attention mechanism is used to index key information, ultimately outputting a fully structured digital archival object that conforms to archiving standards.
[0010] Another embodiment of this application provides an intelligent archive digitization processing system, the system comprising:
[0011] The acquisition module is used to simultaneously scan paper archives using a multispectral linear array scanner, acquiring archive image data in the visible light band and near-infrared band respectively;
[0012] The correction module is used to perform non-uniform lighting correction and perspective distortion correction on the archival image data, and generate a geometrically and color-normalized dual-channel fused archival image.
[0013] The classification module is used to perform page region segmentation and semantic classification based on the dual-channel fused archive image using a multimodal feature fusion network, and output structured page description information including text area, table area, signature area and illustration area;
[0014] The parsing module is used to parse the text content based on the structured layout description information using a joint entity recognition and relation extraction model, while extracting key field entities and their semantic relationships to generate a preliminary archival information graph.
[0015] The output module is used to index key information based on the archival information map and the preset archival metadata specifications, through an attention mechanism-driven content importance assessment module, and finally output a fully structured digital archival object that conforms to the archiving standards.
[0016] Another embodiment of this application provides a storage medium storing a computer program, wherein the computer program is configured to execute the method described in any of the preceding claims when running.
[0017] Another embodiment of this application provides an electronic device including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the method described in any of the preceding claims.
[0018] Compared with existing technologies, the intelligent archive digitization processing method provided by this invention can automatically and with high fidelity generate fully structured digital archive objects, thereby improving the accuracy, completeness and knowledge content of digitization processing. Attached Figure Description
[0019] Figure 1 A hardware structure block diagram of a computer terminal for an intelligent archive digitization processing method provided in an embodiment of the present invention;
[0020] Figure 2 A flowchart illustrating an intelligent archive digitization method provided in an embodiment of the present invention;
[0021] Figure 3 This is a schematic diagram of the structure of an intelligent archive digitization processing system provided in an embodiment of the present invention. Detailed Implementation
[0022] The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0023] This invention first provides an intelligent archive digitization processing method, which can be applied to electronic devices, such as computer terminals, specifically ordinary computers.
[0024] The following detailed explanation uses a computer terminal as an example. Figure 1 This is a hardware structure block diagram of a computer terminal for an intelligent archive digitization processing method provided in an embodiment of the present invention. (See diagram below.) Figure 1 As shown, the computer device includes a processor, memory, and network interface connected via a system bus, wherein the memory may include non-volatile storage media and internal memory.
[0025] See Figure 2 The present invention provides an intelligent archive digitization processing method, which may include the following steps:
[0026] S201 uses a multispectral linear array scanner to simultaneously scan paper archives, acquiring archive image data in the visible light band and near-infrared band respectively;
[0027] Specifically, the scanning parameters of the multispectral linear array scanner can be initialized, including the sensitivity settings and scanning resolution for the visible and near-infrared bands, and a scanning parameter configuration file can be generated.
[0028] This step is the basic parameter calibration process for document scanning. Its core is to customize the scanning parameters according to the document's material and preservation condition to ensure the quality of dual-band image data acquisition, providing high-fidelity raw data for subsequent processing. The specific implementation method is as follows:
[0029] The core of a multispectral linear array scanner is a dual-band linear array image sensor, corresponding to the visible light band (wavelength range 400-760nm) and the near-infrared band (wavelength range 760-1100nm), respectively. Initialization requires parameter configuration via the scanner control software, following the "archive type adaptation principle"—adjusting parameters based on different preservation conditions (e.g., new archives, faded archives, moldy archives) and materials (e.g., Xuan paper, kraft paper, coated paper). Sensitivity settings aim for a "signal strength-to-noise ratio balance," with the visible light band sensitivity unit being V / lux. s (Vol / lux) (seconds), near-infrared band is V / W m - ² (volts / watt) (Square meter), by pre-scanning a 1cm×1cm archival sample to obtain the basic signal value. If the sample signal value is below 200mV, the sensitivity is increased; if it is above 800mV, the sensitivity is decreased to avoid signal saturation or excessive weakness. For example, when processing faded Xuan paper archives, the visible light sensitivity is set to 0.5V / lux. s, near-infrared sensitivity set to 0.8V / W m - ², enhanced weak signal capture capability; when processing modern printed paper archives, visible light sensitivity drops to 0.2V / lux. s, near-infrared drops to 0.3V / W m - ², reduce noise interference.
[0030] The choice of scanning resolution should balance "detail preservation" and "storage efficiency," measured in dpi (dots per inch). A base resolution of 300 dpi is recommended (to meet archival digitization standards). For documents containing small stamps or handwritten annotations, increase the resolution to 600 dpi. For large-format drawings, reduce the resolution to 200 dpi to control file size. Resolution and scanning size are linked. For example, if the document is A4 size (210mm × 297mm), the pixel size of the scanned image at 300 dpi is 2480 × 3508, while at 600 dpi it is 4960 × 7016. These parameters should be automatically calculated and matched during configuration.
[0031] The scanning parameter configuration file is stored in XML format and contains five main modules: "Device Identifier, File Number, Band Parameters, Resolution Parameters, and Pre-scan Verification Value." Each module has specific fields. After the configuration file is generated, it must be verified using MD5 to ensure its integrity, and the verification value is appended to the end of the file.
[0032] According to the scanning parameter configuration file, the scanner is controlled to perform line-by-line synchronous scanning of paper documents, simultaneously acquiring visible light image data and near-infrared image data, and generating the original dual-band image sequence.
[0033] This step is the core acquisition stage for dual-band data. A line-by-line synchronous scanning mechanism ensures the spatial correspondence of the two band images, avoiding subsequent fusion deviations caused by asynchronous acquisition. The specific implementation method is as follows:
[0034] The scanning control adopts a "master-slave signal triggering" mode. The visible light and near-infrared sensors of the multispectral linear array scanner share the same scanning drive axis. The master controller calculates the scanning step size based on the resolution parameters in the configuration file. For example, at 300 dpi, the movement distance per step is 0.0847 mm (1 inch / 300). The drive axis moves the file through the scanning area at a uniform speed, with the transport speed set to 0.5 m / s, ensuring that the acquisition time of a single scan line is stable within 20 microseconds. The master controller sends a synchronization trigger signal to the dual-band sensor, with the signal time difference controlled within ≤10 microseconds, ensuring that the visible light and near-infrared signals at the same scanning position are captured simultaneously, guaranteeing synchronization at the hardware level.
[0035] During progressive scanning, the linear array sensor generates image data in a "line-frame" manner. A single scan line corresponds to one row of pixels in the image, and consecutive scan lines are stitched together to form a complete image frame. The visible light sensor outputs 3-channel (RGB) data, while the near-infrared sensor outputs single-channel data, both stored with 16-bit depth (value range 0-65535). Compared to 8-bit depth, this retains more detail in both light and dark areas, making it particularly suitable for capturing faded text in archives. Image data is transmitted in real-time to a local cache with a capacity of 10GB. A circular overwrite mechanism is used to prevent data overflow, and a CRC32 check is performed every 10 frames transmitted to ensure error-free transmission.
[0036] The original dual-band image sequence consists of a visible light sequence and a near-infrared sequence. The number of frames and the frame size of the two types of sequences are exactly the same. The sequence naming follows the rule of "file number_band type_frame number.format", such as "DA_2025_0036_VIS_001.TIFF" and "DA_2025_0036_NIR_001.TIFF". The TIFF (lossless compression) format is selected to preserve the original data quality. The metadata of a single frame image includes information such as the scan timestamp (accurate to milliseconds), scan line number, and sensor temperature, which are appended to the tag field of the TIFF file. For example, the timestamp "2025-10-1609:23:45.123" and the sensor temperature "28.5℃" provide a time reference for subsequent time alignment.
[0037] During scanning, the document's positional offset needs to be monitored in real time. The scanner's built-in CCD positioning sensor acquires the document's edge coordinates. If the edge coordinate offset exceeds 2 pixels for three consecutive frames, mechanical correction is immediately triggered, driving the document transport platform to make fine adjustments. The correction accuracy is ≤0.5 pixels to prevent image sequence distortion caused by document skew. For example, if the CCD detects a 3-pixel rightward offset during scanning, the platform immediately adjusts to the left by 0.025mm (3×0.0847mm / 10) to ensure accurate positioning of subsequent scan lines.
[0038] The original dual-band image sequence is time-aligned to ensure that the visible light image and the near-infrared image are completely synchronized in space and time, generating aligned dual-band image pairs.
[0039] This step is crucial for eliminating scanning synchronization errors. By achieving both temporal and spatial alignment, it addresses image offset issues caused by sensor trigger delays and mechanical vibrations, providing a precisely matched image foundation for subsequent fusion processing. The specific implementation method is as follows:
[0040] Timing alignment is based on "timestamp association + frame synchronization verification". First, the scanning timestamp of each frame in the original dual-band image sequence is extracted to establish a time mapping relationship between visible light frames and near-infrared frames. Theoretically, the difference between the timestamps of two synchronously acquired frames should be ≤10 microseconds. If the difference exceeds 50 microseconds, it is determined to be a timing out-of-synchronization, which needs to be corrected by interpolation or frame discarding. For example, if the timestamp of visible light frame 001 is "09:23:45.123000" and the timestamp of near-infrared frame 001 is "09:23:45.123060", the difference is 60 microseconds. Near-infrared frame 001 needs to be discarded, and the timestamp of near-infrared frame 002 needs to be adjusted to match that of visible light frame 001. At the same time, the frame adjustment log is recorded to ensure timing continuity.
[0041] Spatial alignment employs a dual verification method combining cross-correlation algorithm and feature point matching. First, the overlapping region of each pair of visible and near-infrared images (typically the central 80% area of the image) is selected, and the cross-correlation coefficient between the two images is calculated. The cross-correlation coefficient ranges from -1 to 1, with values closer to 1 indicating a higher spatial matching degree. If the cross-correlation coefficient is ≥0.95, the spatial alignment is considered successful. If it is <0.95, feature points (such as document edges and text inflection points) are extracted from both images using the SIFT feature point detection algorithm. At least 50 feature points are extracted from each image, and feature point pairs are matched using the K-nearest neighbor algorithm. The average offset (Δx, Δy) is calculated, and then the near-infrared image is translated to correct the spatial offset. For example, if the cross-correlation coefficient of a pair of images is 0.88, after feature point matching, the offset is calculated to be (2,1) pixels, that is, the near-infrared image is offset 2 pixels to the right and 1 pixel down relative to the visible light image. By using an image translation algorithm, the near-infrared image is translated 2 pixels to the left and 1 pixel up. After correction, the cross-correlation coefficient is recalculated and increased to 0.98, which meets the alignment requirements.
[0042] Spatial alignment also needs to handle image scaling and rotation deviations. If there is a slight rotation (≤1°) during the document scanning process, the horizontal baseline (such as the document edge) of the two images is detected by Hough transform, the rotation angle θ is calculated, and the near-infrared image is rotated. The rotation center is set as the image center point. After rotation, cropping is performed to remove black borders. For example, if the near-infrared image is detected to be rotated 0.5° counterclockwise relative to the visible light image, it is rotated 0.5° clockwise around the image center (1240, 1754). After rotation, a 10-pixel black border area is cropped to ensure that the two images are the same size.
[0043] Aligned dual-band image pairs are stored using a one-to-one binding method, named in the format "file number_image pair number_alignment identifier.TIFF", for example, "DA_2025_0036_Pair_001_Aligned.TIFF". The file contains image data for both bands, distinguished by the band identifier field. An alignment verification report is also generated, recording parameters such as the temporal difference, cross-correlation coefficient, and offset correction for each image pair. For example, a report entry might be "Pair_001: temporal difference 8 microseconds, cross-correlation coefficient 0.98, spatial offset correction (2,1) pixels, rotation correction 0°, alignment qualified". Image pairs that fail the verification (e.g., cross-correlation coefficient < 0.9) need to be re-scanned to avoid affecting subsequent processing.
[0044] Image quality metrics are extracted from aligned dual-band image pairs, adaptive image enhancement is performed based on the image quality metrics, and finally archival image data in the visible light and near-infrared bands are generated.
[0045] This step is a key optimization process for improving image quality. By quantifying quality indicators and performing differential enhancement processing, it addresses issues such as noise, blurring, and low contrast in scanned images, providing high-quality data for subsequent correction and fusion. The specific implementation method is as follows:
[0046] Image quality metrics extraction focuses on three core dimensions: sharpness, contrast, and signal-to-noise ratio (SNR). Each dimension has a defined quantification method and evaluation threshold. Sharpness is measured using the Laplacian gradient value. The Laplacian value of each pixel in the image is calculated, and the average value of all pixels is taken as the sharpness metric. The threshold is set to ≥50 (the higher the value, the sharper the image). For example, a visible light image with a Laplacian average of 42 is considered to have insufficient sharpness. Contrast is calculated using the standard deviation of grayscale values. A standard deviation ≥80 is considered acceptable, while a standard deviation <50 indicates low contrast. Signal-to-noise ratio (SNR) is calculated as the ratio of signal power to noise power, measured in dB. An SNR ≥30dB is acceptable, while <25dB indicates low SNR. During metric extraction, the image is divided into 16×16 pixel blocks, and the local metrics for each block are calculated. If the local metric of a block is below 70% of the global threshold, it is marked as a "weak quality area" and requires focused processing in subsequent enhancement.
[0047] Adaptive image enhancement employs a differentiated strategy of "regional and band-based" enhancement, selecting enhancement algorithms specifically based on the shortcomings of quality indicators. For low-resolution images, an unsharpened mask algorithm is used to enhance edge details. The gain coefficient k in the algorithm is dynamically adjusted according to the sharpness index; the lower the sharpness, the larger the k value (ranging from 1.2 to 2.0). For example, when the sharpness index is 42, k=1.8, and when the sharpness index is 55, k=1.3, to avoid excessive enhancement that could amplify noise. For low-contrast images, the CLAHE (contrast-limited adaptive histogram equalization) algorithm is used in the visible light band, dividing the image into 8×8 grid blocks. The histogram clipLimit of each block is set to 2.0 (to control the contrast enhancement). Gamma correction is used in the near-infrared band (γ value ranging from 0.8 to 1.2). For example, when the contrast standard deviation of a near-infrared image is 45, γ=0.8 enhances dark details. For low signal-to-noise ratio images, a bilateral filtering algorithm is used for noise reduction. The filter window size is set to 5×5, the spatial standard deviation σ_d=5, and the grayscale standard deviation σ_r=20, suppressing noise while preserving details.
[0048] Enhancement processing requires localized strengthening of "weak quality areas." For example, if the contrast standard deviation of a 16×16 block in the lower left corner of an image is 38 (global threshold 50), it is marked as a weak area. The CLAHE algorithm is then applied separately to this area, increasing the clipLimit to 2.5, while simultaneously overlaying edge enhancement operators to ensure local quality meets the global standard. During enhancement, the image's grayscale distribution is monitored in real time to avoid pixel saturation (0 or 65535). If the percentage of saturated pixels exceeds 1% after enhancement, the enhancement intensity is reduced and the image is reprocessed. For example, the gain coefficient of the unsharpened mask is reduced from 1.8 to 1.5 to ensure no loss of image detail.
[0049] The final generated visible and near-infrared archival image data must meet the quality standards (resolution ≥ 50, contrast standard deviation ≥ 50, SNR ≥ 30dB). The format remains TIFF lossless compression, 16-bit deep. The metadata includes enhancement processing records, including the enhancement algorithm type, parameter values, and a comparison of quality indicators before and after enhancement. For example, the enhancement record for a visible light image is: "Enhancement algorithm: CLAHE + unsharpened mask; CLAHEclipLimit=2.0, unsharpened k=1.5; Before enhancement: resolution 42, contrast 45, SNR=28dB; After enhancement: resolution 58, contrast 62, SNR=32dB." After image data generation, a secondary verification is performed by the quality detection module. Once verified, the data is stored in a distributed storage system, awaiting subsequent correction processing.
[0050] S202, perform non-uniform lighting correction and perspective distortion correction on the archive image data to generate a geometrically and color-standardized dual-channel fused archive image;
[0051] Specifically, illumination distribution analysis can be performed on archival image data in the visible light band and near-infrared band respectively, and a non-uniform illumination field can be estimated using a polynomial fitting algorithm to generate an illumination distribution model.
[0052] This step is a core prerequisite for resolving uneven brightness in scanned archival images. By accurately analyzing the illumination distribution characteristics of dual-band images, a mathematical model that closely reflects actual illumination changes is constructed, providing a quantitative basis for subsequent illumination compensation. The specific implementation method is as follows:
[0053] Non-uniform illumination in scanned archival images mainly stems from scanner light source attenuation, differences in paper reflectivity, and scanning angle deviations. This manifests as typical characteristics such as "bright center and dark edges," "local shadows," and "striped brightness fluctuations." Furthermore, the illumination distribution patterns differ between the visible and near-infrared bands—visible light is more significantly affected by paper color and texture, while near-infrared is more susceptible to interference from uneven material transmittance. Therefore, separate illumination analyses are necessary. Before analysis, the dual-band images need preprocessing: the visible light image is converted from RGB space to grayscale space (using a weighted average method, grayscale value G = 0.299R + 0.587G + 0.114B). The near-infrared image, being single-channel data, directly uses its original grayscale value for analysis. Simultaneously, median filtering (3×3 window size) is used to remove isolated noise points in the image, avoiding interference from noise in the illumination distribution estimation. For example, a 2×2 pixel bright spot caused by dust in a visible light image can be effectively eliminated after median filtering, ensuring that the illumination analysis focuses on true brightness changes.
[0054] The illumination distribution analysis employs a "grid sampling + statistical modeling" approach. First, the image is divided into non-overlapping grid blocks of a fixed size. The grid size is adaptively adjusted based on the image resolution; a 300 dpi A4 image (2480×3508 pixels) is typically divided into a 40×50 grid (approximately 62×70 pixels per block), ensuring sampling density while avoiding excessive computation. The average grayscale value is calculated for each grid block using the following formula: Where M and N are the width and height of the grid block, and G(i,j) is the gray value of pixel (i,j) within the block. For example, the average gray value of the central grid block in a visible light image is 220, while that of the edge grid blocks is 85, which intuitively reflects the distribution characteristic of "bright center and dark edge". In a near-infrared image, a certain local grid block has an average gray value of only 50 due to the shadow of paper wrinkles, which is significantly lower than the 150 of the surrounding area.
[0055] Polynomial fitting algorithms are the core tool for estimating non-uniform illumination fields. The core idea is to approximate the illumination intensity distribution on the image plane using low-order polynomial functions. Choosing a third-order polynomial as the fitting basis function can accurately capture common illumination gradient features while avoiding overfitting problems caused by higher-order polynomials. For any pixel (x,y) on the image plane, the fitting function for its illumination intensity L(x,y) is: L(x,y)=a0+a1x+a2y+a3x²+a4xy+a5y²+a6x³+a7x²y+a8xy²+a9y³, where x and y are the normalized coordinates of the pixel (range [0,1], eliminating the influence of the pixel's absolute position on the fitting), and a0 to a9 are the polynomial coefficients to be solved.
[0056] The coefficients are solved using the least squares method. The center coordinates (x_c, y_c) of each grid block are used as sampling points, and the average gray value G_block of the grid block is used as the actual illumination observation value at that point. The objective function is constructed as: minΣ[L(x_c, y_c) - G_block]², and the coefficient vector is solved by matrix inversion. For example, after calculating the fitting coefficients for a visible light image, we get a0=100, a1=50, a2=-30, a3=-120, and the absolute values of the remaining coefficients are all less than 10. Substituting these values into the function, we can see that increasing x (horizontal direction) increases illumination, while increasing y (vertical direction) decreases illumination, consistent with the actual distribution of "brighter on the right, darker on the left, brighter at the top, darker at the bottom." After fitting, the coordinates of each pixel are substituted into the polynomial function to obtain the estimated illumination intensity value for that pixel. The estimated values of all pixels form a complete illumination distribution model, which is stored in a two-dimensional array, perfectly matching the size of the original image.
[0057] To verify the model's accuracy, the root mean square error (RMSE) between the fitted illumination estimate and the actual average gray value of the grid blocks is calculated. A model is considered acceptable if RMSE ≤ 15; otherwise, the grid size or polynomial order needs to be adjusted for refitting. For example, the initial RMSE of a near-infrared image was 22. After increasing the number of grid blocks to 50×60, the RMSE decreased to 13, and the model accuracy met the requirements.
[0058] Based on the illumination distribution model, pixel-by-pixel illumination compensation is performed on visible light images and near-infrared images to eliminate brightness unevenness and generate illumination-corrected dual-band images.
[0059] This step is the lighting correction process. Through precise pixel-by-pixel compensation, it restores images under non-uniform lighting to a uniform lighting effect, while preserving the original texture and text details of the file. The specific implementation method is as follows:
[0060] The core principle of illumination compensation is "to strip away the influence of illumination and restore the intrinsic grayscale of the target". That is, it is assumed that the pixel grayscale value G_original(x,y) of the original image is determined by the product of the intrinsic grayscale value of the target G_object(x,y) and the illumination intensity L(x,y) (simplified model). Therefore, the compensated target grayscale value G_corrected(x,y) = G_original(x,y) × (G_avg / L(x,y)), where G_avg is the global average grayscale value of the image, which serves as the standard illumination reference after compensation. The essence of this formula is to uniformly correct the light intensity of each pixel to the global average level. For example, if the original grayscale value of a pixel is 80, the light estimation value is 100, and the global average grayscale value is 150, then the grayscale value after compensation is 80×(150 / 100)=120, which effectively improves the brightness of dark areas; while for pixels with excessive light (such as L(x,y)=200, G_original=240), the compensation value is 240×(150 / 200)=180, which avoids overexposure in bright areas.
[0061] To address the differences in characteristics between visible light and near-infrared images, the compensation strategy needs to be adjusted accordingly: visible light images need to take color consistency into account, so illumination compensation is first performed on the three RGB channels separately (the illumination distribution model and global average gray level are calculated separately for each channel), and then the image is converted back to the RGB space to avoid color distortion. For example, if the global average grayscale of the R channel of an RGB image is 160, the original value of the R channel of a certain pixel is 90, and the estimated illumination value is 80, the compensated R value is 90 × (160 / 80) = 180. The G and B channels are processed according to the same logic, and the color ratio remains unchanged. Near-infrared images are single-channel data, and the compensation is performed directly according to the above formula. However, special attention should be paid to the faded text area. The near-infrared band has a stronger ability to recognize faded ink. When compensating, the compensation coefficient of this area can be appropriately reduced (multiplied by an attenuation factor of 0.8-0.9) to avoid the background brightness from obscuring the details of the text. For example, for pixels in the faded signature area of a near-infrared image, the grayscale value after compensation should be controlled between 100-120, which is higher than the original background of 80 and lower than the text of 60, highlighting the contrast between the text and the background.
[0062] A grayscale constraint mechanism must be incorporated during the compensation process to prevent the compensated pixel grayscale values from exceeding the range of 16-bit images (0-65535). When the calculated G_corrected(x,y) > 65535, it is forcibly set to 65535; when < 0, it is forcibly set to 0. Simultaneously, for "dead black areas" in the image (original grayscale value < 10, estimated illumination value < 20), neighborhood interpolation is used instead of direct compensation. That is, the average grayscale value within a 3×3 neighborhood of the pixel after compensation is taken as its final value, avoiding noise amplification caused by excessively large compensation coefficients in such areas. For example, in the wrinkled shadow area of a file edge, some pixels have an original grayscale value of 5 and an estimated illumination value of 15. Direct compensation would yield 5×(150 / 15) = 50, but since there is no actual file content in this area, taking the neighborhood average of 35 as the final value is more consistent with the overall visual effect.
[0063] The illuminated dual-band image needs to undergo quality verification, with the core indicators being "brightness uniformity" and "detail retention." Brightness uniformity is measured by calculating the standard deviation of the grayscale values of the corrected image; a standard deviation ≤30 is considered acceptable (typically >80 before correction). Detail retention is measured by comparing the edge gradient values before and after correction (calculated using the Sobel operator); a decrease in edge gradient values ≤10% is acceptable, ensuring that critical details such as text and lines are not lost during the compensation process. For example, a visible light image with a grayscale standard deviation of 92 before correction reduced to 28 after correction; the gradient value of text edges decreased from 120 to 115, a decrease of 4.2%, meeting the quality requirements. The corrected image is stored in TIFF format, and the metadata records key parameters of illumination compensation (such as polynomial order, global average grayscale value, and RMSE) to provide a traceability basis for subsequent processing.
[0064] Extract page edge features from the illumination-corrected dual-band image, detect page contours using Hough transform, calculate perspective transformation matrix, and generate geometric correction parameters.
[0065] This step is crucial for resolving geometric distortions in scanned archival images. By precisely locating the page outline and calculating perspective transformations, the tilted and distorted archival images are restored to standard rectangles, providing a geometrically regular image foundation for subsequent page segmentation. The specific implementation method is as follows:
[0066] Page edge feature extraction is a prerequisite for contour detection. It is necessary to highlight edge information from the illumination-corrected dual-band image and adopt a combination strategy of "edge enhancement + binarization": First, the image is Gaussian blurred (standard deviation σ=1.5, window size 5×5) to smooth noise while preserving edge details; then, the Canny edge detection algorithm is applied to extract edges. The high threshold of the Canny algorithm is set to 150 and the low threshold is set to 50. The high threshold is used to detect strong edges, and the low threshold is used to connect weak edges but exclude noise. For example, after Canny detection, the page edges of a certain archive image form continuous lines, while the text edges inside the paper are partially filtered out; finally, the edge image is binarized (threshold 128), and edge pixels are set to 255 (white) and non-edge pixels are set to 0 (black) to obtain a binary edge image containing only edge information.
[0067] Considering the complementarity of dual-band images, edge feature extraction employs a "dual-band fusion" strategy: the visible light and near-infrared images are processed separately to obtain two binary edge maps. Then, a pixel-wise OR operation is performed (if any band represents an edge pixel, it is retained) to generate a fused edge map. This method can compensate for the shortcomings of single-band edge detection—visible light is clearer for detecting the physical edges of paper, while near-infrared can identify edges obscured by stains, significantly improving the integrity of the fused edge map. For example, if a file has broken edges in the visible light image due to localized mold, but the edges of that area are clear in the near-infrared image, the fused image can form complete edge lines.
[0068] The Hough transform is the core algorithm for detecting straight lines on page outlines. Its principle is to convert straight lines in image space into points in parameter space, and then use an accumulator to count the density of points in parameter space. The density peaks correspond to straight lines in the image. Page outlines are typically rectangular, containing four straight lines (top, bottom, left, and right edges), so the Hough transform focuses on detecting these four lines. The algorithm parameters are set as follows: the polar angle θ ranges from -90° to 90°, with a step size of 1°; the polar radius ρ ranges from 0 to the length of the image diagonal, with a step size of 1 pixel; the accumulator threshold is set to 10% of the total number of edge pixels (ensuring only long straight lines are detected, filtering out short edges such as text and lines). For example, in a 2480×3508 pixel image, the total number of edge pixels is 12000, and the accumulator threshold is set to 1200, meaning only long straight lines like those at the page edges can reach this threshold.
[0069] After the Hough transform, the detected lines are filtered and fitted: First, the polar angle θ of each line is calculated, and lines with polar angles close to 0° or 90° are selected (corresponding to horizontal and vertical edges respectively, with an allowable deviation of ±5°); then, horizontal lines are sorted by their polar radius ρ, and the two with the largest and smallest ρ are taken as the top and bottom edges of the page; vertical lines are similarly sorted by ρ, and the two with the largest and smallest ρ are taken as the left and right edges. If no matching line is detected for a certain edge (such as a missing part of the file edge), it is completed using the neighborhood line extension method. For example, if only the upper half of the line is detected for the left edge, it is extended to the bottom of the image based on the slope of the line (close to 90°) to form a complete left edge.
[0070] The four straight lines of the page outline intersect to form four vertices (top left, top right, bottom left, and bottom right). Let the coordinates of these four vertices in the original distorted image be... The target coordinates in the standard rectangular image are (Typically set to (0,0), (W,0), (0,H), (W,H), where W and H are the width and height of the target image, consistent with the actual size of the file). The perspective transformation matrix M is a 3×3 matrix used to describe the mapping relationship between the original coordinates and the target coordinates. Its elements are obtained by solving a system of linear equations—the coordinates of each vertex satisfy the perspective transformation formula: , The four vertices can be used to construct eight linear equations, which can be solved to obtain the eight unknown elements of matrix M. (It is usually set to 1 to simplify calculations).
[0071] For example, the original vertex coordinates of a file are (50,60), (2430,75), (45,3480), (2420,3490), and the target coordinates are (0,0), (2480,0), (0,3508), (2480,3508). Substituting these coordinates into the formula, the perspective transformation matrix M is obtained as: [[1.02,0.005,-51],[0.003,1.01,-60.5],[0.000001,0.000002,1]]. This matrix is the core geometric correction parameter. At the same time, the target image size (W=2480,H=3508) needs to be recorded as an auxiliary parameter to form a complete set of geometric correction parameters.
[0072] The geometric correction parameters are applied to perform perspective transformation and distortion correction on the illumination-corrected dual-band image, and the visible light and near-infrared images are fused at the pixel level to finally generate a geometrically and color-normalized dual-channel fused archive image.
[0073] This step is the final execution stage of geometric correction and image fusion. It eliminates geometric distortion through perspective transformation and integrates the advantages of dual-band fusion at the pixel level to generate a fused image that combines geometric regularity, realistic color, and rich detail. The specific implementation method is as follows:
[0074] Perspective transformation is the core process for eliminating distortion using geometric correction parameters. Essentially, it maps each pixel (x, y) in the original distorted image to its corresponding position (X, Y) in the target standard image using a perspective transformation matrix M. Since the mapped coordinates may be non-integer, bilinear interpolation is used to calculate the grayscale value of the target pixel. Bilinear interpolation takes the grayscale values of the four neighboring pixels around the target coordinates and averages them by distance to obtain the final value, ensuring image smoothness while avoiding pixel distortion. For example, coordinates (100.2, 200.3) in the original image are mapped to (98, 195) in the target image. The grayscale values of four pixels (100, 200), (101, 200), (100, 201), and (101, 201) in the original image are taken and averaged with weights of 0.8, 0.2, 0.7, and 0.3 respectively, and this average is used as the grayscale value of the target pixel.
[0075] Perspective transformation needs to be performed separately on both visible and near-infrared dual-band images to ensure perfect geometric alignment. The same perspective transformation matrix M and target image size are used during the transformation, ensuring a one-to-one correspondence of pixel coordinates between the transformed dual-band images, laying the foundation for subsequent fusion. After transformation, the image needs to be cropped to remove black borders (areas exceeding the target coordinate range after transformation) caused by the perspective transformation. The cropped image size must strictly be (W, H) to ensure geometric normalization. For example, if an image has 10-20 pixel black borders around its edges after transformation, cropping will yield a standard size of 2480×3508 without any black border interference.
[0076] Pixel-level fusion is the key to integrating the advantages of dual-band images. The core idea is that "visible light contributes color and texture, while near-infrared contributes material and detail." It adopts an adaptive weighted fusion strategy based on regional features, rather than a simple fixed-weight fusion. The values of the fusion weights ω_VIS (visible light weight) and ω_NIR (near-infrared weight) are determined by the features of the region where the pixel is located, satisfying ω_VIS + ω_NIR = 1. The specific rules are as follows: the near-infrared weight is higher for text areas and signature areas (ω_NIR = 0.6-0.7) because the near-infrared band can more clearly present faded text and stamp marks. For example, in the faded handwritten text area of a certain document, the contrast between the text and the background is 30 in the visible light image and 60 in the near-infrared image. When fusing, ω_NIR = 0.7 is taken to highlight the text details; the visible light weight is higher for illustration areas and table line areas (ω_VIS = 0.7-0.8) because visible light can more realistically restore the colors of illustrations and the clarity of table lines; blank areas use equal weight fusion (ω_VIS = ω_NIR = 0.5) to balance the brightness characteristics of the two bands.
[0077] The specific calculations for the fusion process are as follows: For RGB visible light images and single-channel near-infrared images, the near-infrared image is first converted into three-channel data (R=G=B=NIR grayscale value). Then, the fusion value is calculated for each channel's pixels according to their weights, using the formula Fusion_R=ω_VIS×VIS_R+ω_NIR×NIR_R. Fusion_G and Fusion_B are calculated using the same logic. For example, if a pixel in the text area has a visible light R value (VIS_R) of 180 and a near-infrared R value (NIR_R after conversion) of 120, with ω_VIS=0.3 and ω_NIR=0.7, the fused R value is 180×0.3+120×0.7=54+84=138. This value retains the basic brightness of the visible light while incorporating the text detail contrast from the near-infrared.
[0078] The merged images need to undergo geometric and color normalization: Geometric normalization ensures that the size ratio of all archival images is consistent with the actual archives, with A4 archives uniformly set to 2480×3508 pixels and A3 archives to 3508×4960 pixels, avoiding errors caused by size differences in subsequent processing; Color normalization is achieved through histogram matching, aligning the RGB histogram of the merged image with the standard archive color template, ensuring a consistent color style for archival images scanned in different batches. For example, if a batch of scanned archives is generally yellowish, after histogram matching, the color deviation is controlled within ΔE < 5 (CIE1976 color space), meeting the color consistency requirements of archive digitization.
[0079] The final generated dual-channel fused archive image is stored in TIFF format. The image data contains fusion information for the RGB three channels. The metadata records key parameters for geometric correction (perspective transformation matrix, target size) and fusion (region weight allocation). At the same time, it passes quality inspections—geometric distortion rate ≤1% (the deviation of the angles of the four corners of the page from 90° after correction ≤1°), color deviation ΔE < 5, and detail contrast ≥ 50—to ensure that the image fully meets the requirements of subsequent page segmentation and semantic classification.
[0080] S203, Based on the dual-channel fused archive image, a multimodal feature fusion network is used to perform page region segmentation and semantic classification, and output structured page description information including text area, table area, signature area and illustration area;
[0081] Specifically, the dual-channel fused archival image can be input into a multimodal feature fusion network, and the texture features of the visible light channel and the material features of the near-infrared channel can be extracted through a convolutional neural network to generate a dual-modal feature map.
[0082] This step is a fundamental feature-based step in layout analysis. By extracting the core features of the dual-band image—visible light focused visual texture and near-infrared focused material differences—it provides accurate feature support for subsequent fusion and segmentation. The specific implementation method is as follows:
[0083] The dual-channel fused archive image is in RGB three-channel format (visible light contributes color and texture, and near-infrared light is incorporated after channel expansion). Before being input into the network, it needs to be standardized preprocessed: the pixel values from 0-65535 (16-bit) are normalized to the 0-1 range, the global mean of the image is calculated (e.g., the RGB channel mean values are 0.485, 0.456, and 0.406 respectively) and the standard deviation (0.229, 0.224, and 0.225 respectively), and the operation "(pixel value - mean) / standard deviation" is performed on each pixel to eliminate the interference of brightness differences on feature extraction. For example, if the RGB value of a pixel is (192, 168, 120), after normalization it becomes ((192 / 255-0.485) / 0.229≈1.23, (168 / 255-0.456) / 0.224≈1.01, (120 / 255-0.406) / 0.225≈0.35), ensuring the stability of the input feature distribution.
[0084] The dual-branch convolutional structure of the multimodal feature fusion network employs a differentiated design: ResNet-50 is used as the backbone network for feature extraction in the visible light channel. Its residual block structure effectively addresses the gradient vanishing problem in deep networks, focusing on extracting visual features such as text stroke texture and table line edges. For the near-infrared channel, due to the relatively singular dimension of material features, a lightweight version of ResNet-18 is chosen, reducing computation while retaining material differentiation capabilities, such as the difference in near-infrared reflectance between paper fibers and inkpads. The input layers of both branches are adaptively adjusted: the ResNet-50 input is 3 channels (matching the fused image), while the ResNet-18 input is adjusted to 1 channel (extracting near-infrared band information separately, separated from the G channel of the fused image, as near-infrared information is evenly distributed to the RGB channels during fusion).
[0085] The feature extraction process progresses hierarchically from "shallow texture - mid-level structure - deep semantics": the first four layers of ResNet-50 (conv1 to conv4) output feature maps at different scales. The conv2 layer (stride 2, 256 output channels) extracts a 256×620×877 texture feature map, focusing on the thickness of strokes and the clarity of edges. The conv4 layer (stride 2, 1024 output channels) outputs a 1024×155×219 deep feature map, containing semantic differences between text and illustration areas. The conv3 layer of ResNet-18 (256 output channels) outputs a 256×155×219 material feature map, clearly distinguishing between the signature area (inkpad material with strong near-infrared absorption and low feature values) and the blank area (paper with strong reflection and high feature values). For example, the signature area of a document appears as a continuous low-feature-value region in the near-infrared feature map, forming a clear boundary with the surrounding text area, providing a basis for subsequent segmentation.
[0086] The generated bimodal feature maps need to maintain scale consistency. The feature map of the conv2 layer of ResNet-50 is downsampled to a scale of 155×219 by max pooling with a stride of 2, and its size is unified with the feature map of the conv3 layer of ResNet-18 and the feature map of the conv4 layer of ResNet-50. Finally, three types of feature maps are formed: "visible light texture feature map (256×155×219), visible light semantic feature map (1024×155×219), and near-infrared material feature map (256×155×219)," which prepares for subsequent fusion.
[0087] Feature fusion is performed on the dual-modal feature map. The visible light and near-infrared features are weighted and fused using an attention mechanism to generate a unified feature representation after fusion.
[0088] This step is the core of integrating the advantages of dual-modality. It dynamically allocates feature weights through an attention mechanism, allowing the network to focus on the most valuable feature information for page segmentation and avoid interference from invalid features. The specific implementation method is as follows:
[0089] Before feature fusion, channel-dimensional concatenation is required. The three types of feature maps are merged along the channel direction to form a concatenated feature map of 1536×155×219 (256+1024+256=1536). After concatenation, channel compression is performed using a 1×1 convolution kernel to reduce the number of channels to 512. This reduces computational cost while enhancing feature correlation. The number of convolution kernels is set to 512, the stride is 1, and the padding method is "same" (ensuring the feature map size remains unchanged). For example, a 1536-dimensional feature vector of a pixel in the concatenated feature map is output as a 512-dimensional vector after a 1×1 convolution, retaining core correlation information while reducing the dimension by 66%.
[0090] The attention mechanism employs a dual structure of "channel attention + spatial attention." First, the channel attention module (CAM) calculates the importance weights of each feature channel, and then the spatial attention module (SAM) focuses on key regions in the image. The core of the channel attention module is "global average pooling + fully connected layers": the 512×155×219 feature map is globally average pooled by channel to obtain a 512-dimensional channel feature vector; two fully connected layers (the first layer has 256 neurons and uses the ReLU activation function; the second layer has 512 neurons and uses the Sigmoid activation function) calculate the channel weights. The weight values range from 0 to 1, with larger values indicating greater importance for the channel feature. For example, the channel weights corresponding to visible light semantic features are generally between 0.7 and 0.8, the channel weights related to signatures in near-infrared material features reach 0.85, while the weights of channels with high noise levels are only 0.1 to 0.2.
[0091] The spatial attention module focuses on target regions such as text and tables based on the feature maps filtered by channel weights. It performs max pooling and average pooling along the channel dimension to obtain two 1×155×219 feature maps, which are then concatenated and processed by a 3×3 convolution (with sigmoid activation function) to generate a 1×155×219 spatial weight map. Regions with high weight values (such as densely texted areas) receive priority attention. For example, the text area of a document has a weight value of 0.8-0.9 in the spatial weight map, while the blank area has only 0.1-0.2, effectively highlighting the features of the target region.
[0092] The weighted fusion process multiplies the channel weights and spatial weights pixel-by-pixel to obtain the final feature weight matrix, which is then multiplied element-by-element with the original concatenated feature map to generate a unified feature representation of 512×155×219. The fused feature map has multiple advantages: the texture and semantic features of the text area are clear, the line structure features of the table area are prominent, and the material difference features of the signature area are obvious. For example, in a file containing a table, the feature value of the table lines in the fused feature map is twice that of the surrounding text, providing clear feature differences for subsequent pixel-level segmentation.
[0093] Based on a unified feature representation, a fully convolutional network is used to perform pixel-level page region segmentation and generate an initial segmentation mask for the page region.
[0094] This step is crucial for achieving accurate segmentation of the layout area. Through the end-to-end segmentation capability of a fully convolutional network, the fused features are mapped to pixel-level category labels, yielding preliminary region segmentation results. The specific implementation method is as follows:
[0095] The Fully Convolutional Network (FCN) uses the U-Net architecture. Its encoder-decoder structure and skip connection design can extract deep semantic features through encoder downsampling and restore spatial details through decoder upsampling, perfectly meeting the dual requirements of "semantic accuracy" and "boundary precision" for archival layout segmentation. The encoder part of U-Net directly reuses the unified feature representation (512×155×219) output by the multimodal feature fusion network, without repeated downsampling, saving computational resources. The decoder part performs upsampling through transposed convolution. Each upsampling doubles the feature map size and halves the number of channels, and a total of 4 upsampling operations are performed to restore the feature map size to be consistent with the original archival image (2480×3508).
[0096] The purpose of skip connections is to fuse the shallow features (rich in spatial details) of the encoder with the deep features (clear semantic information) of the decoder. For example, the feature map after the first upsampling of the decoder (256×310×438) is concatenated with the feature map of the conv3 layer of ResNet-50 (256×310×438) to supplement details such as text strokes and table lines; the feature map after the last upsampling (64×2480×3508) is concatenated with the feature map of the conv1 layer of ResNet-50 (64×2480×3508) to ensure the accuracy of region boundaries.
[0097] The core of the segmentation process is the class probability mapping. The output layer of U-Net uses a 1×1 convolutional kernel to map the 64-channel feature map output by the decoder into a 4-channel feature map (corresponding to four categories: text area, table area, signature area, and illustration area), with the number of channels matching the number of categories. Then, the Softmax function is used to normalize the 4-channel feature values of each pixel, obtaining the probability of each pixel belonging to each category, using the formula P(c)=exp(x_c) / Σexp(x_i) (where x_c is the feature value of the pixel in category c channel, i=1-4). For example, a pixel with 4-channel feature values (3.2, 0.8, 0.5, 0.3) has a probability of (0.92, 0.05, 0.02, 0.01) after Softmax calculation, clearly belonging to the text area.
[0098] The initial segmentation mask is a visualization of the probability mapping, a single-channel image where pixel values correspond to the category label with the highest probability (text area = 1, table area = 2, signature area = 3, illustration area = 4, background = 0). For example, in the segmentation mask of an A4 document, the pixel values of the 1000×1500 pixel area in the upper left corner are all 1 (text area), the 800×600 pixel area in the middle is 2 (table area), and the 200×200 pixel area in the lower right corner is 3 (signature area). This initially achieves region division, but two types of problems may exist: one is small noise areas (such as isolated label areas below 50 pixels), and the other is blurred transitions at region boundaries (such as uncertain pixel labels at the boundary between the text area and the table area).
[0099] The initial segmentation mask is post-processed, including region merging and boundary optimization, and each region is semantically classified using a classifier. The final output is a structured layout description containing text areas, table areas, signature areas, and illustration areas.
[0100] This step is the final stage for improving segmentation quality and outputting structured results. Post-processing eliminates segmentation noise and optimizes boundaries, followed by classification confirmation to generate standardized page layout description information. The specific implementation method is as follows:
[0101] The core of region merging is to eliminate small, noisy areas, employing an "area threshold + neighborhood voting" strategy: First, the initial segmentation mask is traversed, and the area (number of pixels) of each connected region is counted. An area threshold of 50 pixels is set (approximately equivalent to a physical area of 0.5cm x 0.5cm; areas smaller than this cannot be valid layout areas). Regions with an area less than 50 pixels are marked as "regions to be merged." Then, the pixel category distribution within the 3x3 neighborhood of the region to be merged is analyzed, and the label of the region to be merged is updated to the category with the highest percentage within its neighborhood. For example, if an isolated 10x3 pixel region is labeled 2 (table area), and 80% of its neighboring pixels are labeled 1 (text area), then the label of this region is updated to 1, and it is merged with the surrounding text areas.
[0102] Boundary optimization employs a combination of morphological operations and edge smoothing: First, a morphological dilation operation (3×3 structuring element, 1 iteration) is performed on the segmentation mask to fill in the tiny holes inside the region; then, an erosion operation (also 3×3 structuring element, 1 iteration) is performed to restore the original size of the region, and boundary burrs are eliminated through the opening operation of "dilation-erosion"; finally, Gaussian filtering (standard deviation σ=0.5) is applied to the boundary pixels of the region, and the category labels of the boundary pixels are corrected by neighborhood weighted average. For example, if the original label of a pixel at the boundary is 1 (text area), and 30% of its 5×5 neighborhood is 2 (table area), then it is corrected to a transition label of "text area-table area boundary" through weighted calculation, thereby improving the boundary accuracy.
[0103] A semantic classifier is used to ultimately confirm the region category, avoiding category confusion caused by post-processing. It employs a lightweight convolutional neural network (CNN). The input is the feature vector of each connected region in the segmentation mask (containing 10-dimensional features such as region area, shape factor, and pixel grayscale mean), and the output is the probability of each of the four categories. The shape factor is calculated as S = 4πA / L² (where A is the region area and L is the region perimeter). The shape factor for text areas is typically 0.3-0.5 (irregular polygons), for table areas it's 0.6-0.8 (close to rectangles), and for signature areas it's 0.7-0.9 (circular or elliptical). The shape factor allows for rapid category differentiation. For example, a region with an area of 10,000 pixels and a perimeter of 400 pixels has a shape factor of 4 × 3.14 × 10,000 / (400²) = 0.785. Combined with grayscale features, the classifier outputs a probability of 0.96 for it being a signature area, confirming the correct category.
[0104] The structured layout description information is output in XML format, including core nodes such as file number, image size, and region list. Each region node records information such as region ID, category, coordinate range (pixel coordinates of the top left and bottom right corners), and feature description. This information not only clarifies the location and category of each region, but also provides accurate region positioning basis for subsequent text extraction and information parsing.
[0105] S204. Based on the structured layout description information, the text content is parsed using a joint entity recognition and relation extraction model, and key field entities and their semantic relationships are extracted to generate a preliminary archival information graph.
[0106] Specifically, based on the text area location in the structured layout description information, the text area image can be extracted from the dual-channel fused archive image, and the image can be converted into text content through optical character recognition technology to generate the original text data;
[0107] This step is the basic entry point for digitizing archival content. The core is to accurately locate the text area, optimize the image quality to improve the recognition accuracy, and finally convert the imaged text into editable original text data. The specific implementation methods are as follows:
[0108] The structured layout description information provides the accurate coordinates of the text area in XML format. For example, the text area node of a certain archive is <region id="1" type="text area" coord="(50,60)-(1050,1560)" feature="area 1.5×10 6 pixels" / >, where the two sets of numerical values in the coord field represent the pixel coordinates (x1, y1, x2, y2) of the upper left and lower right corners of the text area respectively. Based on this coordinate, use the image cropping function of OpenCV to extract the text area image. When cropping, a 5-pixel edge redundancy needs to be reserved to avoid losing edge text due to coordinate deviation. Finally, a text area image of 1005×1505 pixels is obtained.
[0109] The preprocessing of the text area image is the key to improving the accuracy of optical character recognition (OCR). It is necessary to perform three operations of "denoising - enhancement - regularization" according to the characteristics of archival text: First, use median filtering (3×3 window) to remove interference such as paper texture and scanning noise, for example, eliminate isolated 2×2 pixel dark spots in the image; Second, enhance the contrast between the text and the background through the CLAHE algorithm, increase the standard deviation of the image gray value from 30 to 60, and make the strokes of faded text clearer; Finally, perform skew correction. Detect the horizontal angle of the text line through the Hough transform. If a skew of more than 0.5° is detected, rotate and correct with the image center as the origin to ensure that the text line is parallel to the horizontal direction. For example, the text area of a certain archive is skewed 1.2° due to scanning offset. After correction, the horizontal deviation of the text line ≤ 0.1°, meeting the OCR recognition requirements.
[0110] The optical character recognition technology adopts a combined scheme of "basic engine + domain dictionary". The basic engine selects Tesseract OCR that supports multiple languages, loads Chinese (chi_sim) and English (eng) bilingual packages, and adapts to the common Chinese-English mixed text in the archive; The domain dictionary constructs a custom word library for archival professional terms (such as "reply", "letter", "operator", etc.), containing more than 5,000 high-frequency words. By modifying the word frequency configuration file of the OCR engine, the recognition accuracy of professional terms is improved. During the recognition process, the preprocessed text area image is segmented into blocks of 1000×1000 pixels, and the results are spliced after block recognition to avoid recognition delay caused by large-size images. For example, the image of a certain text area contains "reply letter regarding the XX water supply project". When the domain dictionary is not loaded, it is recognized as "reply letter regarding the XX water supply item white", and after loading the dictionary, it is corrected to the correct expression, and the accuracy rate is increased to 98%.
[0111] The original text data is a structured storage of the OCR recognition results, using JSON format, and includes information such as file number, text region ID, recognized text, recognition confidence, and text line coordinates. The recognition confidence is output by the OCR engine, ranging from 0 to 1. Text lines with a confidence level < 0.7 are marked as "low confidence regions" and require manual review. For example, a fragment of original text data might be: {"archive_id":"DA_2025_0042","text_region_id":"1","content":"Approval for Project XX\nIssuing Unit: XX City Water Supply Company\nIssuance Date: October 16, 2025\nHandler: Zhang San\nProject Amount: 5 million yuan","confidence":0.92,"low_confidence_lines":[]}, which fully preserves the recognized content and quality indicators, providing a foundation for subsequent text parsing.
[0112] The original text data is input into the joint entity recognition and relation extraction model, which uses a pre-trained language model to encode the text and identify key field entities to generate an entity list.
[0113] This step is crucial for extracting core information from the archives. Leveraging the semantic understanding capabilities of a pre-trained language model, it accurately identifies key field entities in the text (such as organizations, dates, and names), providing an entity foundation for subsequent relation extraction. The specific implementation method is as follows:
[0114] Before inputting raw text data into the model, text cleaning and format standardization are required: First, remove redundant spaces, newlines, and invisible characters (such as control characters in ASCII) from the recognition results, and merge multi-line text into a coherent text sequence. For example, process "Approval for Project XX\nIssuing Unit: XX Company" into "Approval for Project XX Issuing Unit: XX Company". Second, perform sentence segmentation, using punctuation marks such as periods, commas, and colons as delimiters to split long texts into short sentences, with each short sentence limited to 512 characters (to adapt to the input limit of the pre-trained model). Finally, convert the text into a vocabulary index that the model can recognize through the tokenization operation, load the vocabulary of the pre-trained model (such as the Chinese vocabulary of BERT, which contains 21,128 tokens), split "XX City Water Supply Company" into sub-words such as "XX", "City", "Supply", "Water", "Company", and add "[CLS]" (sentence beginning marker) and "[SEP]" (sentence end marker).
[0115] The pre-trained language model selected is BERT-base-Chinese. Its 12-layer Transformer structure has powerful context semantic encoding capabilities and can effectively distinguish the specific meanings of polysemous words in archival texts (e.g., "reply" is a noun in the archival context, not a verb). During model fine-tuning, it is trained based on an archival entity annotation dataset (containing 100,000 annotated samples covering 5 core entities such as organizations, dates, personal names, project names, and amounts). The fine-tuning parameter settings are as follows: learning rate 2e-5 (to avoid destroying pre-trained knowledge with too large a learning rate), batch size 16, 10 training epochs, and the cross-entropy loss function is used to optimize model parameters. For example, for the text "Issuing unit: XX City Water Supply Company", after model encoding, it can capture the semantic association between "Issuing unit" and "XX City Water Supply Company", providing context support for entity recognition.
[0116] For entity recognition, the classic "BIO annotation system + CRF decoding" scheme is adopted. In the BIO annotation system, "B-X" represents the start position of entity X, "I-X" represents the internal position of entity X, and "O" represents a non-entity position, where X corresponds to 5 types of entities (ORG: organization, DATE: date, PER: personal name, PROJ: project name, MONEY: amount). For example, the annotation result of the text "Handler: Zhang San is responsible for the XX Water Supply Project" is "Han / O Ban / O Ren / O: / O Zhang / B-PER San / I-PER / O Fu / O Ze / O XX / B-PROJ Shui / I-PROJ Li / I-PROJ Xiang / I-PROJ Mu / I-PROJ". The CRF decoding layer then corrects the initial label sequence output by the model by learning the transition probabilities of entity labels (e.g., after "B-PER", it is more likely to be followed by "I-PER" rather than "B-ORG"), improving the accuracy of entity boundaries. For example, it corrects the initially output "Zhang / B-PER San / O" by the model to "Zhang / B-PER San / I-PER", avoiding incorrect splitting of personal names.
[0117] The entity list is a structured presentation of the recognition results, in JSON format. Each entity record contains an entity ID (globally unique, such as "E1001"), entity type (such as "ORG"), entity text (such as "XX City Water Supply Company"), text start position (character index, such as 15 - 20), and recognition confidence (the probability value output by CRF, such as 0.96). After generation, entity deduplication needs to be performed. If two entity texts are exactly the same and the positions overlap, the record with a higher confidence is retained to ensure the uniqueness of the entity list.
[0118] Based on entity recognition, the semantic relationships between entities are analyzed through a relation extraction module to generate a set of entity relation pairs;
[0119] This step is the core of uncovering the connections between archival information. Based on the identified entities, semantic analysis is used to determine the logical relationships between entities (such as "person in charge - affiliated institution" and "project - approving unit"), providing the edges to support the construction of the information graph. The specific implementation method is as follows:
[0120] The relation extraction module adopts an architecture of "entity pair semantic encoding + classifier judgment". The input is the entity pairs obtained from entity recognition and their corresponding context text, and the output is the relationship type between the entity pairs. First, a context window for the entity pair is constructed: for any two entities E1 and E2, the smallest text fragment containing these two entities is extracted as the context. The window length is up to 100 characters. If the distance between the two entities exceeds 100 characters, the text of the first and last 50 characters of each entity is concatenated to form the context. For example, for entities "Zhang San" (E1003) and "XX City Water Supply Company" (E1001), the corresponding context is "Issuing Unit: XX City Water Supply Company; Issuance Date: October 16, 2025; Person in Charge: Zhang San", which clearly contains the semantic relationship clues between the two.
[0121] The semantic encoding of entity pairs is based on the output of the BERT model, employing a fusion method of "[CLS] vector + entity 1 vector + entity 2 vector": the [CLS] vector is the sentence-initial vector after BERT encoding, containing the semantic information of the entire context; the entity 1 vector and entity 2 vector are the average vectors of the corresponding words of the entity. For example, "XX City Water Supply Company" corresponds to 5 words, and the average vector of these 5 words is taken as the entity vector. After concatenating the three vectors, the dimension is compressed to 768 dimensions (consistent with the BERT output dimension) through a 1×1 convolution kernel to obtain the final semantic vector of the entity pair. This vector simultaneously contains the entity's own features and contextual information.
[0122] The relation classifier employs a "fully connected layer + Softmax" structure. The input is a 768-dimensional entity pair semantic vector, and the output is a 10-dimensional relation probability vector (corresponding to 10 common relation types in archival scenarios, such as "affiliated institution," "issuing unit," "handler," "project leader," and "approval date"). The classifier is trained on an archival relation annotation dataset (50,000 entity pairs and relation annotation samples), optimized using the cross-entropy loss function. The training parameters are consistent with those of the entity recognition module (learning rate 2e-5, batch size 16). For example, when the semantic vector of the entity pair "Zhang San (E1003) - XX City Water Supply Company (E1001)" is input into the classifier, the probability of outputting the "affiliated institution" relation is 0.95, significantly higher than other relation types, clearly indicating a connection between the two.
[0123] The relationship extraction process needs to handle "one-to-many" and "fuzzy relationship" issues: For cases where one entity corresponds to multiple relationships (e.g., "XX Company" could be both "Issuing Unit" and "Organizing Unit"), the relationships with the highest probabilities are retained and marked as "multi-relationship entity pairs"; for entity pairs where all relationship probabilities are below 0.5, they are marked as "fuzzy relationships" and require manual confirmation. The entity relationship pair set is stored in JSON format. Each relationship record contains a relationship ID (e.g., "R2001"), a head entity ID (e.g., "E1003"), a tail entity ID (e.g., "E1001"), a relationship type (e.g., "Affiliated Institution"), a relationship confidence score (e.g., 0.95), and contextual text (e.g., "Issuing Unit: XX City Water Supply Company...Handler: Zhang San"). For example, a relation pair set fragment is: [{"id":"R2001","head_id":"E1003","tail_id":"E1001","relation":"affiliated organization","confidence":0.95},{"id":"R2002","head_id":"E1004","tail_id":"E1001","relation":"issuing unit","confidence":0.98}], where E1004 is the entity "Approval for Project XX".
[0124] By integrating the entity list and the set of entity relationship pairs, and using entities as nodes and relationships as edges, a graph structure data is constructed, ultimately generating a preliminary archival information graph.
[0125] This step is the core output of archival information structuring. It uses a graph structure to organize discrete entities and relationships into a network-like association model, intuitively presenting the internal information logic of the archives. The specific implementation method is as follows:
[0126] The construction of graph-structured data follows a triplet model of "node-edge-attribute". Nodes correspond to entities in the entity list, and edges correspond to relations in the relation pair set. Both nodes and edges contain rich attribute information to support subsequent applications. In addition to basic information such as entity ID, type, and text, node attributes also include extended attributes such as "source text region ID" (associated with structured layout description information) and "initial entity importance judgment" (based on the frequency of an entity's appearance in the text; if it appears ≥2 times, it is marked as a "core entity"). For example, the node attributes for the entity "XX City Water Supply Company" (E1001) are: {"id":"E1001","type":"ORG","text":"XX City Water Supply Company","start":12,"end":18,"confidence":0.96,"source_region_id":"1","impor
[0127] tance":"core entity"}.
[0128] Based on relation pairs, edge attributes are expanded to include attributes such as "relation direction" (clarifying the logical direction between the head and tail entities; for example, in the "Issuing Unit" relation, the head entity is a document, the tail entity is an organization, and the direction is "document → organization") and "relation evidence" (referring to contextual text fragments as evidence of the relation's validity), enhancing the traceability of the relation. For example, the edge attributes of the relation "R2002" (E1004 → E1001, issuing unit) are: {"id":"R2002","head_id":"E1004","tail_id":"E1001","relation":"Issuing Unit","confidence":0.98,"direction":"document → organization","evidence":"Approval for Project XX issued by: XX City Water Supply Company"}.
[0129] Graph-structured data is stored in the industry-standard JSON-LD format, which supports both machine parsing and human readability. It includes two core fields: "@context" (defining the terminological context) and "@graph" (storing a collection of nodes and edges). The "@context" field defines the semantic descriptions of entity types (e.g., "ORG" corresponds to "organizational entity") and relationship types (e.g., "issuing unit" corresponds to "document issuing body"), improving the graph's understandability. The "@graph" field organizes data in a "node first, edge last" order. Nodes are marked with "@type":"node", and edges are marked with "@type":"edge". The "head" and "tail" fields of each edge are associated with the node ID.
[0130] After the initial archival information graph is generated, two basic checks are required: first, "connectivity check," to ensure that there are valid edges connecting all core entities (such as institutions and documents) to avoid isolated core nodes; second, "consistency check," to check for contradictory relationships (such as the same document being "signed" by two institutions simultaneously). If such contradictions exist, they are marked as "pending verification" and sent for manual review. For example, in a certain archive, "Project XX" (E1005) is associated with both "Zhang San" (E1003) and "Li Si" (E1006) as the "project leader." The system automatically marks this contradictory relationship, awaiting manual confirmation of the actual leader. The graph, once verified, retains the original identification information of entities and relationships and presents a complete association chain of "document-institution-person in charge-project" through a graph structure, providing core data support for subsequent metadata indexing and archiving.
[0131] S205, based on the archival information map and the preset archival metadata specifications, key information is indexed through the content importance assessment module driven by the attention mechanism, and finally a fully structured digital archival object that conforms to the archiving standard is output.
[0132] Specifically, the archival information map can be aligned with the preset archival metadata specifications, mapping entities and relationships in the map to metadata fields, and generating a metadata mapping relationship table;
[0133] This step is a core link in achieving the standardization of archival information. By establishing a link between unstructured graphs and structured metadata, discrete entities and relationships are transformed into metadata elements that meet archiving requirements, providing a clear mapping basis for subsequent indexing. The specific implementation method is as follows:
[0134] The pre-defined archival metadata specifications must comply with national archival digitization standards and be customized to suit the characteristics of archives in different fields. Taking government approval archives as an example, the metadata specifications include two categories: "core fields" and "extended fields." Core fields are mandatory (12 in total), including archive number (unique identifier), archive title (core subject of the document), issuing unit (issuing agency), issuance date (issuance time), handler (specific person in charge), project name (related business project), approval result (approval / rejection), etc. Data types include strings, dates, enumeration values, etc. Extended fields are optional (8 in total), such as project amount, contact information, number of attachments, etc., adapted to specific business needs. The metadata specifications are stored in JSON format, clearly defining the name, data type, whether it is mandatory, value range, and description of each field. For example, the "issuance date" field is defined as: {"Field name":"Issuance date","Data type":"DATE","Required":true,"Value range":"YYYY-MM-DD","Description":"The date the document is officially issued, accurate to the day"}.
[0135] The alignment of the archival information map with the metadata specifications adopts a dual strategy of "direct entity-field mapping + relationship-association mapping". Direct entity-field mapping is based on the matching rule of "entity type - field type". A correspondence table between entity types and metadata fields is pre-established. For example, the organization entity (ORG) corresponds to the "Issuing Unit" and "Organizing Unit" fields, the date entity (DATE) corresponds to the "Issuance Date" and "Archiving Date" fields, the person entity (PER) corresponds to the "Handler" and "Approver" fields, the project entity (PROJ) corresponds to the "Project Name" field, and the amount entity (MONEY) corresponds to the "Project Amount" field. During mapping, the entity list in the graph is first extracted, candidate metadata fields are matched according to the entity type, and then the unique mapping relationship is determined by combining the semantic context of the entity. For example, if the context of “XX City Water Supply Company” (ORG type) in the graph is “Issuing Unit: XX City Water Supply Company”, it is directly mapped to the “Issuing Unit” field corresponding to “Issuance Date”. If the text of a date entity (DATE type) is “October 16, 2025”, and the context is “Archiving Time: October 16, 2025”, it is mapped to the “Archiving Date” field instead of “Issuance Date”.
[0136] Relationship-association mapping is used to supplement the logical relationships between metadata fields. Some fields in the metadata specification have dependencies (e.g., "Project Amount" needs to be associated with "Project Name"), and these relationships need to be clarified through the relationships in the metadata graph. For example, in the graph, the entity "XX Water Supply Project" (PROJ type) and "5 million yuan" (MONEY type) have a "budget amount" relationship. "XX Water Supply Project" is already mapped to the "Project Name" field, and "5 million yuan" is mapped to the "Project Amount" field. Therefore, the association between "Project Name" and "Project Amount" is established through this relationship and marked in the metadata as "Project Amount - Associated Project: XX Water Supply Project". For complex relationships (e.g., "Handler - Affiliated Unit"), the relationship attribute is used as extended information for the metadata field. For example, the extended attribute of the "Handler" field is "Affiliated Unit: XX City Water Supply Company", enriching the relevance of the metadata.
[0137] The alignment process needs to handle "one-to-many" and "many-to-one" mapping conflicts: When an entity can be mapped to multiple metadata fields (e.g., "XX Company" could be both "Issuing Unit" and "Organizing Unit"), the entity is filtered by its "coreness" in the graph—coreness is determined by the number of associations an entity has. Entities with ≥3 associations are preferentially mapped to core metadata fields (e.g., "Issuing Unit"), while those with <3 associations are mapped to extended fields. When multiple entities are mapped to the same metadata field (e.g., two date entities both match "Issuance Date"), they are sorted by their identification confidence, retaining entities with a confidence score ≥0.9. If all meet the criteria, a unique value is determined by combining keywords in the context text (e.g., "Issued" and "Documented"). For example, if two date entities are "October 15, 2025" (confidence score 0.98, context "Documented Date") and "October 16, 2025" (confidence score 0.96, context "Issuance Date"), the latter is mapped to the "Issuance Date" field.
[0138] The metadata mapping table is a structured representation of the alignment results, using XML format, and includes three modules: "Basic Archive Information," "Entity Mapping List," and "Relationship Association List." "Basic Archive Information" records the archive number and metadata specification version; each record in the "Entity Mapping List" includes the entity ID, entity type, entity text, mapping metadata fields, mapping basis (context fragment), and confidence level; each record in the "Relationship Association List" includes the relationship ID, head entity mapping field, tail entity mapping field, relationship type, and association description.
[0139] Based on the metadata mapping table, the attention-driven content importance assessment module calculates the importance score of each information unit and generates the importance assessment results.
[0140] This step is crucial for screening core information from archives. It uses an attention mechanism to focus on the importance of metadata fields and the quality characteristics of information units, quantitatively evaluating the value of each information unit and providing a quantitative basis for screening key information. The specific implementation method is as follows:
[0141] The attention mechanism-driven content importance assessment module adopts a three-stage architecture of "feature input - weight calculation - score fusion". The input information unit features include three core features: "metadata field weight", "information unit quality" and "contextual relevance". Each feature category contains multiple sub-features, forming a 10-dimensional feature vector to ensure the comprehensiveness of the assessment. Metadata field weights are defined by preset metadata specifications. Core fields are weighted at 1.0-1.2 (e.g., "Archive Number" 1.2, "Issuing Unit" 1.1), and extended fields are weighted at 0.6-0.9 (e.g., "Contact Information" 0.7, "Number of Attachments" 0.6). Higher weight values indicate greater importance of the field in the archive. Information unit quality includes three sub-features: entity recognition confidence (0-1), text integrity (1 for no missing characters, 0.5-0.9 for ambiguous characters), and format conformity (1 for conforming to field data type, which needs to be converted to 0.8). Contextual relevance measures the degree of relevance between the information unit and the core theme of the archive. It is obtained by calculating the semantic similarity between the text segment containing the information unit and the archive title (using the cosine similarity algorithm, with values of 0-1). Higher similarity indicates stronger relevance.
[0142] The core of the attention mechanism is to dynamically calculate the weights of various features, employing a combination of self-attention and external knowledge attention. Self-attention captures the intrinsic relationships between features; for example, there is a positive correlation between "metadata field weight" and "contextual relevance"—if the information unit of a core field has a high degree of relevance to the topic, its overall weight should be significantly increased. External knowledge attention introduces a knowledge graph from the archival domain, using the strength of the association between the entity type corresponding to the information unit and the core concepts of the domain (such as "approval," "reply," and "project") as the basis for weight adjustment. For example, if the information unit of the "project name" field contains the core concept of "water supply project," the external knowledge weight increases by 0.2.
[0143] The specific weight calculation process is as follows: First, the 10-dimensional feature vector is input into the attention layer. Then, the correlation matrix between features is calculated through scaled dot product attention, as shown in the formula. , where Q, K, and V are the query, key, and value matrices, respectively, and d_k is the feature dimension (10). This is used to avoid gradient vanishing. For example, the correlation coefficient between the "field weight" feature (1.2) and the "contextual relevance" feature (0.98) of the "archive number" information unit is 0.92, which is significantly higher than the correlation coefficient of 0.75 with the "format standardization" feature (1.0), indicating that the two are more closely related. Then, the self-attention output is fused with the external knowledge weights to obtain the final feature weight vector (10-dimensional, with a weight sum of 1).
[0144] Importance scores are calculated using a weighted summation of "feature value × feature weight", with a score range of 0-100. The formula is Score=Σ(Feature_i×Weight_i)×100, where Feature_i is the normalized value (0-1) of the i-th feature and Weight_i is the weight of the i-th feature. For example, the feature vector of the "archive number" information unit is [1.2 (field weight, 1.0 after normalization), 0.99 (confidence), 1.0 (completeness), 1.0 (normalization), 0.98 (relevance), 0.95 (domain relevance), ...] (all other feature values are above 0.9), and the feature weight vector is [0.25, 0.2, 0.15, 0.1, 0.15, 0.15, ...]. The calculated score is (1.0×0.25+0.99×0.2+1.0×0.15+1.0×0.1+0.98×0.15+0.95×0.15)× 100 ≈ 98.3 points; while the feature vector of the "Number of Attachments" information unit (extended field, weight 0.6) is [0.6 (normalized to 0.5), 0.9 (confidence), 1.0 (completeness), 1.0 (standardization), 0.6 (relevance), 0.5 (domain relevance), ...], and the weight vector is [0.2, 0.2, 0.15, 0.15, 0.2, 0.1, ...], and the calculated score is (0.5×0.2+0.9×0.2+1.0×0.15+1.0×0.15+0.6×0.2+0.5×0.1)×100≈68.5 points.
[0145] The importance assessment results are in JSON format and include information unit ID, corresponding metadata fields, feature details, feature weights, importance score, and score level (core: ≥90 points; important: 80-89 points; average: 70-79 points; minor: <70 points). For example, a fragment of an evaluation result is: [{"info_id":"E1001","metadata_field":"Issuing Unit","features":{"Field Weight":1.1,"Confidence":0.96,"Completeness":1.0,"Relevance":0.95},"weights":{"Field Weight":0.22,"Confidence":0.21,"Completeness":0.16,"Relevance":0.17},"score":92.5,"level":"Core"},{"info_id":"E1005","metadata_field":"Number of Attachments","features":{"Field Weight":0.6,"Confidence":0.9,"Relevance":0.5},"weights":{"Field Weight":0.18,"Confidence":0.2,"Relevance":0.19},"score":65.2,"level":"Secondary"}]. After the evaluation is completed, a consistency check is required to ensure that the information unit scores of the core metadata fields (such as file number and issuance date) are all ≥80 points. If they are lower, they are marked as "needs optimization" and returned to the entity recognition stage for reprocessing. For example, if the "issuance date" information unit score is 75 points, the recognition confidence and contextual relevance need to be checked. If the confidence is low due to OCR recognition errors, the OCR results should be optimized and evaluated again.
[0146] Based on the importance assessment results, key information units are selected and indexed according to metadata specifications to generate indexed structured metadata;
[0147] This step is the core output of archival information standardization. By selecting high-value information units and integrating them according to specifications, scattered information is transformed into metadata with a clear structure that meets archiving requirements, providing core data for final encapsulation. The specific implementation method is as follows:
[0148] The selection of key information units is based on the importance assessment score levels, employing a "tiered selection + core assurance" strategy: First, all information units at the "core" (≥90 points) and "important" (80-89 points) levels are retained. These information units correspond to the core value elements of the archives and are mandatory for archiving. Second, information units at the "general" level (70-79 points) are selected based on the "mandatory" status of the metadata fields. Information units at the "general" level with mandatory fields are retained, while those with optional fields are temporarily stored (for supplementary explanations). Information units at the "minor" level (<70 points) are filtered by default, and are only retained and marked as "low quality pending verification" when the corresponding field is mandatory and there are no other information units. For example, the "handler" field of an archive has a score of 88 (important level) and is directly retained; the "project amount" field (optional) has a score of 72 (general level) and is temporarily stored as supplementary information; the "contact information" field (optional) has a score of 63 (minor level) and is directly filtered.
[0149] After screening, information units need to be standardized and indexed according to the preset metadata specifications. The indexing content includes three parts: "field value standardization", "related information supplementation", and "quality identification". Field value standardization sets a unified format for different data types: date types are uniformly formatted as "YYYY-MM-DD", for example, "October 16, 2025" is indexed as "2025-10-16"; amount types are uniformly retained to two decimal places, with the unit being "yuan", for example, "5 million yuan" is indexed as "5,000,000.00 yuan"; enumeration types strictly follow the specifications for value selection, such as "approval result" only allowing three types of values: "agree", "reject", and "suspend". If the information unit text is "agree in principle", it is standardized to "agree" and the original text is noted in the extended information; redundant spaces and special characters are removed from string types, such as "XX City Water Supply Company" is indexed as "XX City Water Supply Company".
[0150] The supplementary information is based on the "relationship list" in the metadata mapping table. It adds related field information to each indexed field, forming a logical chain between fields. For example, if the "Person in Charge" field is indexed as "Zhang San," the supplementary information would be "Affiliated Unit: XX City Water Supply Company" (from the "Affiliated Institution" relationship); if the "Project Amount" field is indexed as "5,000,000.00 yuan," the supplementary information would be "Related Project: XX Water Supply Project" (from the "Budget Amount" relationship). This related information is appended to the indexed fields as "extended attributes," ensuring both the simplicity of the core fields and the preservation of information relevance.
[0151] Quality indicators are used to record the reliability of indexed fields, determined based on the importance score and confidence level of the information unit: a score ≥ 90 and a confidence level ≥ 0.95 is marked as "Reliable"; a score of 80-89 or a confidence level of 0.9-0.95 is marked as "Relatively Reliable"; a score of 70-79 or a confidence level of 0.8-0.9 is marked as "Pending Verification"; and a score < 70 or a confidence level < 0.8 is marked as "Low Reliability". For example, the "Issuing Unit" field has a score of 92.5 and a confidence level of 0.96, and is marked as "Reliable"; the "Archived Date" field has a score of 78 and a confidence level of 0.85, and is marked as "Pending Verification".
[0152] The indexed structured metadata adopts an XML format conforming to national archival standards (GB / T18894-2021), containing five primary nodes: "Archival Identifier," "Basic Information," "Business Information," "Extended Information," and "Quality Description." Each primary node has secondary nodes set according to metadata fields, with extended information and quality identifiers as node attributes. After the structured metadata is generated, it must be verified through a metadata validation tool to ensure field completeness (no missing core fields), format conformity (meets data type requirements), and logical consistency (no contradictions in related fields). A 100% pass rate is required to proceed to the next step.
[0153] The structured metadata and original archival image data are integrated and encapsulated into digital archival objects that conform to archiving standards, and finally output as fully structured digital archival objects.
[0154] This step is the final stage of archival digitization. Through standardized encapsulation, metadata and image data are merged into a single, long-term digital archival object, ensuring the integrity, usability, and archiving compliance of the archives. The specific implementation method is as follows:
[0155] The core data to be integrated comprises three categories: First, indexed structured metadata (XML format), serving as the "identity information" of the digital archives and recording their core attributes and relationships; second, original archival data, including dual-channel fused archival images (TIFF format, 16-bit depth, lossless compression), text area images (JPEG format, quality factor 0.95), and OCR raw text data (TXT format), serving as the "content carrier" of the archives; and third, process traceability data, including scanning parameter configuration files, illumination correction and geometric correction parameters, entity recognition and relationship extraction logs, and importance assessment reports, serving as the "source traceability basis" of the archives and ensuring the traceability of the digitization process. Before integration, all data must undergo integrity verification using MD5 hash value verification. The MD5 value of each data file is stored in a verification list; for example, the MD5 value of the fused image is "a1b2c3d4e5f67890abcdef1234567890". If the verification fails, the corresponding data is retrieved again.
[0156] The digital archive objects are packaged in PDF / A-2b, an internationally recognized long-term archival format that supports lossless image embedding, scalable metadata storage, and cross-platform compatibility, ensuring that the archives can still be read normally for decades to come. The encapsulation employs a "layered embedding" strategy: structured metadata is converted into PDF XMP (Extensible Metadata Platform) metadata and embedded in the PDF file header. The XMP metadata and structured metadata fields correspond one-to-one; for example, XMP's "dc:title" corresponds to "document title," and "pdfa:issuingUnit" corresponds to "issuing unit." The standardized format of field values is maintained during the conversion process. Dual-channel fused archival images are used as the main content pages of the PDF, embedded in the order of the archival pages. The image resolution is maintained at 300dpi, and JPEG2000 (lossless mode) compression is used to ensure image quality is consistent with the original while controlling file size; a single A4 image is approximately 5-8MB. Text area images, OCR text, and process traceability data are embedded as "attachments" in the PDF, linked through the PDF's attachment dictionary. Users can view or download them using a PDF reader. For example, the OCR text attachment is named "DA_2025_0042_OCR.txt," and the process traceability data is embedded after being packaged into a ZIP archive.
[0157] During the encapsulation process, it is necessary to achieve the correlation and positioning of "metadata-image content". Using the PDF's "annotation" function, key fields in the structured metadata are associated with corresponding areas in the image. For example, if the "Issuing Unit" field is indexed as "XX City Water Supply Company", an annotation is added at the location of this text in the PDF image. Clicking the annotation will jump to the detailed information of the "Issuing Unit" field in the metadata. Conversely, clicking the "Handler" field in the metadata view will automatically locate the text area "Handler: Zhang San" in the image. The implementation of this correlation and positioning relies on the text area coordinates in the structured layout description information. Pixel coordinates are converted to the PDF page's coordinate system (with the lower left corner as the origin, the unit being points, 1 point = 1 / 72 inch). For example, the text area coordinates (50, 60) are converted to PDF coordinates (50×72 / 96, (3508-60)×72 / 96) ≈ (37.5, 2586), ensuring accurate positioning.
[0158] The packaged digital archive object must meet three core requirements of the archiving standard: First, long-term preservation, ensuring compliance with PDF / A-2b format and eliminating dynamic content dependent on specific software (such as JavaScript). All fonts are embedded in the PDF file (using standard fonts such as SimSun and Heiti to avoid garbled characters due to missing fonts). Second, security, adding a digital signature (using the SHA-256 algorithm, generated by the archivist's digital certificate) to ensure the archive content cannot be tampered with. The signature information includes the signer, signature time, and signature algorithm; verifying the signature determines whether the archive has been modified. Third, searchability, embedding OCR-recognized original text in the text layer of the PDF file, supporting full-text search. At the same time, the fields of XMP metadata can be indexed by the archive management system, enabling precise retrieval by fields such as "issuing unit" and "issuance date."
[0159] The final output of a fully structured digital archive object consists of two files: a main file "DA_2025_0042.pdf" (PDF / A-2b format), integrating metadata, images, attachments, and a digital signature; and a verification file "DA_2025_0042_checksum.txt," recording the MD5 hash value and verification time of the main file and all embedded data. Output supports both local storage and cloud archiving. Local storage requires saving to a storage device that meets archival security requirements (such as an encrypted hard drive). Cloud archiving involves uploading to the archive management system, with SSL encryption used during the upload process to ensure data security. For example, after receiving an archive object, an archive management system automatically parses the XMP metadata, enters fields such as "archive number," "issuing unit," and "issuance date" into the database, generates an index, and allows users to quickly retrieve the archive and view its structured metadata and original image content, completing the closed loop of the entire intelligent archive digitization process.
[0160] Another embodiment of the present invention provides an intelligent archive digitization processing system, see [link to relevant documentation]. Figure 3 The system may include:
[0161] The acquisition module 301 is used to simultaneously scan paper archives using a multispectral linear array scanner, acquiring archive image data in the visible light band and near-infrared band respectively.
[0162] The correction module 302 is used to perform non-uniform lighting correction and perspective distortion correction on the archival image data to generate a geometrically and color-standardized dual-channel fused archival image.
[0163] The classification module 303 is used to perform page region segmentation and semantic classification based on the dual-channel fused archive image using a multimodal feature fusion network, and output structured page description information including text area, table area, signature area and illustration area;
[0164] The parsing module 304 is used to parse the text content based on the structured layout description information using a joint entity recognition and relation extraction model, and at the same time extract key field entities and their semantic relationships to generate a preliminary archival information graph.
[0165] The output module 305 is used to index key information based on the archive information map and the preset archive metadata specifications, through an attention mechanism-driven content importance assessment module, and finally output a fully structured digital archive object that conforms to the archiving standards.
[0166] This invention also provides a storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above method embodiments when running.
[0167] This invention also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.
[0168] Specifically, the aforementioned electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the aforementioned processor, and the input / output device is connected to the aforementioned processor.
[0169] The above description, based on the embodiments shown in the figures, details the structure, features, and effects of the present invention. The above description is only a preferred embodiment of the present invention, but the present invention is not limited to the scope of implementation shown in the figures. Any changes made in accordance with the concept of the present invention, or equivalent embodiments modified to have equivalent changes, that do not exceed the spirit covered by the specification and figures, should be within the protection scope of the present invention.
Claims
1. A method for intelligent digitization of archives, characterized in that, The method includes: Paper archives are scanned simultaneously using a multispectral linear array scanner, and archival image data in the visible light band and near-infrared band are collected respectively. Non-uniform lighting correction and perspective distortion correction are performed on the archival image data to generate a geometrically and color-normalized dual-channel fused archival image; Based on the dual-channel fused archive image, a multimodal feature fusion network is used to perform page region segmentation and semantic classification, outputting structured page description information including text area, table area, signature area and illustration area; Based on the structured layout description information, a joint entity recognition and relation extraction model is used to parse the text content, and key field entities and their semantic relationships are extracted to generate a preliminary archival information graph. Based on the aforementioned archival information map and the preset archival metadata specifications, a content importance assessment module driven by an attention mechanism is used to index key information, ultimately outputting a fully structured digital archival object that conforms to archiving standards.
2. The method according to claim 1, characterized in that, The process of simultaneously scanning paper archives using a multispectral linear array scanner to acquire archive image data in both visible and near-infrared bands includes: Initialize the scanning parameters of the multispectral linear array scanner, including sensitivity settings and scanning resolution for the visible and near-infrared bands, and generate a scanning parameter configuration file; According to the scanning parameter configuration file, the scanner is controlled to perform line-by-line synchronous scanning of paper documents, simultaneously acquiring visible light image data and near-infrared image data, and generating the original dual-band image sequence. The original dual-band image sequence is time-aligned to ensure that the visible light image and the near-infrared image are completely synchronized in space and time, generating aligned dual-band image pairs. Image quality metrics are extracted from aligned dual-band image pairs, adaptive image enhancement is performed based on the image quality metrics, and finally archival image data in the visible light and near-infrared bands are generated.
3. The method according to claim 2, characterized in that, The step of performing non-uniform lighting correction and perspective distortion correction on the archival image data to generate a geometrically and color-normalized dual-channel fused archival image includes: Illumination distribution analysis was performed on archival image data in the visible light and near-infrared bands respectively. The non-uniform illumination field was estimated using a polynomial fitting algorithm to generate an illumination distribution model. Based on the illumination distribution model, pixel-by-pixel illumination compensation is performed on visible light images and near-infrared images to eliminate brightness unevenness and generate illumination-corrected dual-band images. Extract page edge features from the illumination-corrected dual-band image, detect page contours using Hough transform, calculate perspective transformation matrix, and generate geometric correction parameters. The geometric correction parameters are applied to perform perspective transformation and distortion correction on the illumination-corrected dual-band image, and the visible light and near-infrared images are fused at the pixel level to finally generate a geometrically and color-normalized dual-channel fused archive image.
4. The method according to claim 3, characterized in that, The method, based on the dual-channel fused archival image, utilizes a multimodal feature fusion network for page region segmentation and semantic classification, outputting structured page layout description information including text areas, table areas, signature areas, and illustration areas, including: The dual-channel fused archival image is input into a multimodal feature fusion network, and the texture features of the visible light channel and the material features of the near-infrared channel are extracted by a convolutional neural network to generate a dual-modal feature map. Feature fusion is performed on the dual-modal feature map. Visible light and near-infrared features are weighted and fused using an attention mechanism to generate a unified feature representation after fusion. Based on a unified feature representation, a fully convolutional network is used to perform pixel-level page region segmentation and generate an initial segmentation mask for the page region. The initial segmentation mask is post-processed, including region merging and boundary optimization, and each region is semantically classified using a classifier. The final output is a structured layout description containing text areas, table areas, signature areas, and illustration areas.
5. The method according to claim 4, characterized in that, The process involves parsing the text content using a joint entity recognition and relation extraction model based on the structured layout description information, simultaneously extracting key field entities and their semantic relationships, and generating a preliminary archival information graph, including: Based on the text area location in the structured layout description information, the text area image is extracted from the dual-channel fused archive image, and the image is converted into text content through optical character recognition technology to generate the original text data; The original text data is input into the joint entity recognition and relation extraction model, which uses a pre-trained language model to encode the text and identify key field entities to generate an entity list. Based on entity recognition, the semantic relationships between entities are analyzed through the relationship extraction module to generate a set of entity relationship pairs; By integrating the entity list and the set of entity relationship pairs, and using entities as nodes and relationships as edges, a graph structure data is constructed, ultimately generating a preliminary archival information graph.
6. The method according to claim 5, characterized in that, Based on the archival information graph and the preset archival metadata specifications, the content importance assessment module driven by an attention mechanism indexes key information, ultimately outputting a fully structured digital archival object that conforms to archiving standards, including: Align the archival information map with the preset archival metadata specifications, map the entities and relationships in the map to metadata fields, and generate a metadata mapping relationship table; Based on the metadata mapping table, the attention-driven content importance assessment module calculates the importance score of each information unit and generates the importance assessment results. Based on the importance assessment results, key information units are selected and indexed according to metadata specifications to generate indexed structured metadata; The structured metadata and original archival image data are integrated and encapsulated into digital archival objects that conform to archiving standards, and finally output as fully structured digital archival objects.
7. An intelligent archival digitization processing system, characterized in that, The system includes: The acquisition module is used to simultaneously scan paper archives using a multispectral linear array scanner, acquiring archive image data in the visible light band and near-infrared band respectively; The correction module is used to perform non-uniform lighting correction and perspective distortion correction on the archival image data, and generate a geometrically and color-normalized dual-channel fused archival image. The classification module is used to perform page region segmentation and semantic classification based on the dual-channel fused archive image using a multimodal feature fusion network, and output structured page description information including text area, table area, signature area and illustration area; The parsing module is used to parse the text content based on the structured layout description information using a joint entity recognition and relation extraction model, while extracting key field entities and their semantic relationships to generate a preliminary archival information graph. The output module is used to index key information based on the archival information map and the preset archival metadata specifications, through an attention mechanism-driven content importance assessment module, and finally output a fully structured digital archival object that conforms to the archiving standards.
8. The system according to claim 7, characterized in that, The acquisition module is specifically used for: Initialize the scanning parameters of the multispectral linear array scanner, including sensitivity settings and scanning resolution for the visible and near-infrared bands, and generate a scanning parameter configuration file; According to the scanning parameter configuration file, the scanner is controlled to perform line-by-line synchronous scanning of paper documents, simultaneously acquiring visible light image data and near-infrared image data, and generating the original dual-band image sequence. The original dual-band image sequence is time-aligned to ensure that the visible light image and the near-infrared image are completely synchronized in space and time, generating aligned dual-band image pairs. Image quality metrics are extracted from aligned dual-band image pairs, adaptive image enhancement is performed based on the image quality metrics, and finally archival image data in the visible light and near-infrared bands are generated.
9. A storage medium, characterized in that, The storage medium stores a computer program, wherein the computer program is configured to execute the method of any one of claims 1-6 when it is run.
10. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to run the computer program to perform the method of any one of claims 1-6.