Archive information processing method and system based on official seal identification

By preprocessing and multi-dimensional feature analysis of official seal archive images, the problem of low accuracy in official seal recognition was solved, and higher quality archive information processing was achieved.

CN121459367APending Publication Date: 2026-02-03FOSHAN POWER SUPPLY BUREAU GUANGDONG POWER GRID
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511626338.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-07
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

In the process of recognizing official seals and documents, existing technologies suffer from insufficient image preprocessing, inadequate feature extraction, and limited compliance verification, resulting in low recognition accuracy.

Method used

By acquiring and preprocessing the original archival images, including filtering, contrast enhancement, and shadow and wrinkle correction, we extract official seal image units and text semantic vectors, perform multi-dimensional feature analysis and compliance integrity verification, and generate structured classified archival data.

Benefits of technology

It improved the accuracy of official seal document recognition, generated more accurate document audit reports, and enhanced the quality of information processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121459367A_ABST
    Figure CN121459367A_ABST
Patent Text Reader

Abstract

The invention discloses an archive information processing method and system based on official seal recognition, and relates to the technical field of information processing.The method includes the steps that an original archive image is preprocessed and optimized to obtain high-quality target image data, and a clear basis is provided for official seal recognition to reduce recognition errors; extracting an official seal image unit, an archive text semantic vector and an official seal surrounding text, capturing information related to the official seal from multiple dimensions, and avoiding identification limitation caused by a single information source; in the feature analysis link, a calibration official seal feature data set is generated, and the accuracy of feature data is improved through calibration, so that feature matching is more accurate; according to the archive information processing scheme of the official seal identification, the technical problem of improving the accuracy rate of the official seal archive identification is solved, so that the generated structured classification archive data and archive inspection report are more accurate, and the accuracy rate of the archive information processing scheme of the official seal identification is improved. And thus, the archive information processing quality based on official seal identification is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information processing technology, and in particular to a method and system for processing archival information based on official seal recognition. Background Technology

[0002] In archival information processing scenarios based on official seal recognition, improving the accuracy of official seal recognition faces a series of technical challenges, which are particularly prominent in the processing of complex archival images. First, during the archival image acquisition and preprocessing stage, factors such as paper aging, wrinkles, and uneven lighting can easily cause blurring and shadow interference in the official seal area of ​​the image. If the preprocessing algorithm cannot specifically optimize noise reduction, contrast enhancement, and shadow / wrinkle correction, subsequent official seal localization will lack a clear image foundation, directly reducing the initial accuracy of official seal recognition. Second, in the official seal extraction and feature analysis stage, the official seal image may suffer from wear, incomplete imprints, or insufficient learning of multi-dimensional features such as shape, texture, and text by the feature extraction algorithm, resulting in low precision of the generated official seal feature dataset. This leads to a low matching degree when compared with the standard dataset, making it difficult to accurately identify the true features of the official seal, thus affecting the recognition accuracy. Furthermore, in the compliance and integrity verification process, relying solely on a single-dimensional verification indicator without comprehensive verification of multiple information such as the official seal image and text semantics can lead to biases in the judgment of the official seal's compliance and integrity, ultimately resulting in inaccurate processing results for archival information based on official seal recognition. These problems, from image preprocessing to feature extraction and verification judgment, accumulate at every stage, severely restricting the improvement of the accuracy of official seal archival recognition. A highly adaptive and fault-tolerant technical solution is urgently needed to address the multiple challenges in complex business scenarios. Summary of the Invention

[0003] This invention provides a method and system for processing archival information based on official seal recognition, solving the technical problem of how to improve the accuracy of official seal recognition in archives.

[0004] The first aspect of this invention provides a method for processing archival information based on official seal recognition, comprising:

[0005] The original image of the target file is acquired and preprocessed to obtain the target image data;

[0006] The target image data is subjected to official seal extraction to obtain official seal image units, document text semantic vectors and surrounding text of the official seal;

[0007] Feature analysis is performed using the official seal image unit, the archive text semantic vector, and the text surrounding the official seal to obtain a calibration official seal feature dataset.

[0008] The compliance and integrity verification is performed using the calibration seal feature dataset, the seal image unit, the seal surrounding text, and the archive text semantic vector to obtain structured classified archive data;

[0009] Based on the structured classification archive data, an archive audit report for the target archive is generated.

[0010] Optionally, the step of acquiring the original file image of the target file and preprocessing it to obtain target image data includes:

[0011] Obtain the original file image of the target file;

[0012] The original archive image is filtered to obtain a denoised archive image;

[0013] The denoised archive image is then contrast-enhanced to obtain a contrast-optimized archive image;

[0014] The contrast-optimized archive image is subjected to adaptive correction for shadows and wrinkles to obtain a corrected archive image;

[0015] The quality of the corrected archive image is checked, and the corrected archive image that passes the quality check is used as the target image data.

[0016] Optionally, the step of extracting the official seal from the target image data to obtain the official seal image unit includes:

[0017] Extract the candidate regions for suspected official seals from the target image data;

[0018] Perform seal boundary edge detection on the suspected candidate seal area to obtain the seal boundary edge image;

[0019] Separate the official seal image units from the boundary edge image of the official seal;

[0020] Extract the archival text semantic vector and the surrounding text of the official seal from the target image data.

[0021] Optionally, the step of performing feature analysis using the official seal image unit, the archival text semantic vector, and the text surrounding the official seal to obtain a calibration official seal feature dataset includes:

[0022] The official seal image unit is subjected to grayscale conversion and pixel normalization to obtain a normalized grayscale official seal image;

[0023] Extract the shape feature vector, texture feature vector, and text content feature vector of the normalized grayscale official seal image;

[0024] The shape feature vector, the texture feature vector, and the text content feature vector are concatenated to obtain the initial official seal feature dataset.

[0025] The shape feature vector and texture feature vector in the initial official seal feature dataset are used for feature comparison. If the feature comparison result does not meet the preset feature comparison conditions, the official seal image unit is enhanced.

[0026] Based on the enhanced official seal image unit, the process jumps to the step of performing grayscale conversion and pixel normalization on the official seal image unit to obtain a normalized grayscale official seal image, until the feature comparison result meets the preset feature comparison condition, and then outputs the adjusted feature data.

[0027] The enhanced official seal image unit, after adjusting the feature data association, is input into a preset text recognition model, and the text recognition result is output.

[0028] The text recognition result is compared with the surrounding text of the official seal to obtain the recognition accuracy.

[0029] When the recognition accuracy is greater than or equal to a preset accuracy threshold, the adjusted feature data is integrated with the text content feature vector to obtain a calibration seal feature dataset.

[0030] Optionally, it also includes:

[0031] If the recognition accuracy is less than the preset accuracy threshold, the model parameters of the preset text recognition model are adjusted.

[0032] Based on the adjusted model parameters, the process jumps to the step of inputting the enhanced official seal image unit associated with the adjusted feature data into the preset text recognition model and outputting the text recognition result, until the recognition accuracy is greater than or equal to the preset accuracy threshold, and then outputs the calibrated official seal feature dataset.

[0033] Optionally, the compliance and integrity verification is performed using the calibration seal feature dataset, the seal image unit, the text surrounding the seal, and the semantic vector of the archival text to obtain structured classified archival data, including:

[0034] The official seal feature dataset, the text surrounding the official seal, and the semantic vector of the archive text are used for classification and identification to obtain the official seal classification identification data of the target archive;

[0035] The compliance determination is performed using the calibration seal feature dataset and the preset seal feature specification database to obtain the seal compliance result of the target file;

[0036] The integrity of the official seal image unit is verified to obtain the integrity verification result;

[0037] The official seal classification identifier data, the calibration official seal feature dataset, the official seal compliance result, and the integrity verification result are imported into a preset structured data template to obtain structured classification archive data.

[0038] Optionally, the step of classifying and labeling the target archive using the calibration seal feature dataset, the text surrounding the seal, and the semantic vector of the archive text to obtain the seal classification label data includes:

[0039] Extract multiple surrounding text fragments from the text surrounding the official seal;

[0040] Calculate the text similarity between the text content feature vector in the calibration seal feature dataset and multiple surrounding text fragments;

[0041] Select a preset number of surrounding text fragments as target surrounding text based on text similarity from highest to lowest.

[0042] The text surrounding each target is used as input to a preset word vector model, and multiple word semantic vectors are output.

[0043] Calculate the semantic matching degree between the semantic vector of the archive text and the semantic vector of each word;

[0044] The semantic vectors of words associated with semantic matching degrees greater than or equal to a preset matching degree threshold are retained.

[0045] Based on the archival domain dictionary and the calibration seal feature dataset, the semantic vectors of words associated with semantic matching degrees less than the preset matching degree threshold are corrected;

[0046] By integrating the corrected word semantic vectors with the retained word semantic vectors, the semantic classification results are obtained;

[0047] The archival text semantic vectors are input into a pre-trained archival classification model, which outputs the corresponding archival category labels.

[0048] The semantic classification results are used to map the archive category labels to obtain the official seal classification identifier data of the target archive.

[0049] Optionally, the step of using the calibrated official seal feature dataset and the preset official seal feature specification database to perform compliance determination and obtain the official seal compliance result of the target file includes:

[0050] The calibration seal feature dataset is compared with the preset seal feature specification database to obtain feature values ​​of multiple different feature types, including shape features, texture features and text content features.

[0051] When all types of feature values ​​meet the associated preset feature standard conditions, the calibration seal feature dataset is marked as compliant, the target file is determined to be compliant, and a seal compliance result is generated.

[0052] If any type of feature value does not meet the associated preset feature standard condition, the calibration seal feature dataset is marked as to be evaluated.

[0053] Using the feature values ​​described in the calibration seal feature dataset to be evaluated, the degree of difference of multiple different feature types is calculated;

[0054] The suspiciousness score of the target file is obtained by weighting the calculation using a preset coupling factor that combines the differences and correlations.

[0055] Based on the preset risk threshold range to which the comprehensive suspicion score belongs, the official seal compliance result of the target file is determined. The official seal compliance result includes compliance level, compliance pending verification level, compliance defect level, and compliance violation level.

[0056] Optionally, generating the archival audit report for the target archive based on the structured classified archival data includes:

[0057] Data cleaning is performed on the structured classification archive data;

[0058] The consistency of the structured classification archive data and the cleaned structured classification archive data is verified.

[0059] When the consistency verification result is passed, the target key is generated by using the official seal compliance result and the integrity verification result in the cleaned structured classification archive data.

[0060] The target key is used to retrieve a preset file risk key-value pair table and match the corresponding file risk level of the target file;

[0061] By integrating the cleaned structured and classified archive data with the archive risk level, an archive audit report for the target archive is generated.

[0062] A second aspect of the present invention provides a document information processing system based on official seal recognition, comprising:

[0063] The official seal data acquisition and image preprocessing module is used to acquire the original image of the target document and preprocess it to obtain the target image data;

[0064] The official seal image unit feature extraction module is used to extract the official seal from the target image data to obtain the official seal image unit, the document text semantic vector, and the text surrounding the official seal.

[0065] The official seal image unit feature processing module is used to perform feature analysis using the official seal image unit, the archive text semantic vector and the text surrounding the official seal to obtain a calibration official seal feature dataset.

[0066] The official seal compliance judgment and integrity verification module is used to perform compliance and integrity verification using the calibration official seal feature dataset, the official seal image unit, the text surrounding the official seal and the semantic vector of the archive text, to obtain structured classified archive data;

[0067] The official seal suspiciousness analysis, assessment and early warning module is used to generate an archive audit report for the target archive based on the structured classified archive data.

[0068] As can be seen from the above technical solutions, the present invention has the following advantages:

[0069] This invention provides a method and system for processing archival information based on official seal recognition. The method involves acquiring and preprocessing the original image of the target archive to obtain target image data. From the target image data, official seal image units, archival text semantic vectors, and surrounding text are extracted. These features are then used for feature analysis to obtain a calibration official seal feature dataset. Finally, the calibration official seal feature dataset, official seal image units, surrounding text, and archival text semantic vectors are combined to perform compliance and integrity checks to obtain structured and classified archival data. Based on this data, an archival audit report is generated. This invention obtains high-quality target image data by preprocessing and optimizing the original archival images, providing a clear foundation for official seal recognition and reducing recognition errors. It extracts official seal image units, archival text semantic vectors, and surrounding text to capture information related to the official seal from multiple dimensions, avoiding the limitations of recognition caused by a single information source. The feature analysis stage generates a calibrated official seal feature dataset, improving the accuracy of feature data and making feature matching more precise. Compliance and integrity verification, combined with multiple aspects of information, further ensures the reliability of the recognition results. This archival information processing scheme for official seal recognition solves the technical problem of improving the accuracy of official seal archival recognition, making the generated structured classified archival data and archival audit reports more accurate, thereby effectively improving the quality of archival information processing based on official seal recognition. Attached Figure Description

[0070] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0071] Figure 1 A flowchart illustrating the steps of a method for processing archival information based on official seal recognition, provided in an embodiment of the present invention;

[0072] Figure 2 This is a schematic diagram of the structure of an archival information processing system based on official seal recognition, provided in an embodiment of the present invention. Detailed Implementation

[0073] This invention provides a method and system for processing archival information based on official seal recognition, which addresses the technical problem of how to improve the accuracy of official seal recognition in archival documents.

[0074] To make the objectives, features, and advantages of this invention more apparent and understandable, the technical solutions of the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described below are only some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0075] Please see Figure 1 , Figure 1 A flowchart illustrating the steps of a method for processing archival information based on official seal recognition, provided in an embodiment of the present invention.

[0076] This invention provides a method for processing archival information based on official seal recognition, comprising:

[0077] Step 101: Obtain the original file image of the target file and perform preprocessing to obtain the target image data.

[0078] Further, step 101 may include the following sub-steps:

[0079] S11. Obtain the original file image of the target file.

[0080] In this embodiment of the invention, an industrial-grade image acquisition device, such as a high-precision CCD camera equipped with a uniform LED light source, is used to photograph or scan the target paper archive, converting the visual information of the archive into a digital image format. The target archive refers to the paper or electronic archive carrier to be used for seal recognition and information processing. The original archive image is an initial digital image without any image processing operations, containing the original visual information of the archive, such as paper texture, seal details, and text content. It is necessary to ensure that the archive is not significantly shifted and the lighting is uniform during the image acquisition process, so as to preserve as much detail information of the seal area as possible, thereby improving the accuracy of seal archive recognition.

[0081] S12. Filter the original archive image to obtain a denoised archive image.

[0082] In this embodiment of the invention, an adaptive median filtering algorithm is used to dynamically adjust the size of the filtering window according to the type of noise in the original archival image (such as salt and pepper noise or Gaussian noise) to smooth the image pixels and remove noise interference. Herein, filtering is an image processing technique that uses a specific algorithm to operate on image pixels to eliminate noise in the image. The denoised archival image is an image with effective noise suppression and clearer image details after filtering. Targeted filtering operations reduce the impact of noise on the location and feature extraction of the official seal area.

[0083] S13. Enhance the contrast of the denoised archive image to obtain a contrast-optimized archive image.

[0084] In this embodiment of the invention, a retinal cortex theory algorithm is used to decompose the denoised archival image into an illumination component and a reflection component. The illumination component is smoothed to eliminate the influence of uneven illumination, and the reflection component is grayscale stretched to make the grayscale difference between the official seal area and the background in the image more significant. Among them, contrast enhancement is an image processing technique that increases the grayscale difference between the official seal and the background in the image by adjusting the grayscale range of the image pixels. The contrast-optimized archival image is an image with clearer and more distinct details (such as text and texture) of the official seal after contrast enhancement. By improving the contrast of the official seal area, the edge detection and feature extraction of the official seal can capture details more accurately, further improving the accuracy of official seal document recognition.

[0085] S14. Perform adaptive correction of shadows and wrinkles on the contrast-optimized archive image to obtain the corrected archive image.

[0086] In this embodiment of the invention, the contrast-optimized archive image is converted to the YUV color space (YUV color space, a color encoding method that separates luminance and chrominance signals, where Y represents the luminance channel and U and V represent the chrominance channels). The luminance channel representing the brightness of the image is extracted from the YUV color space, and the shadow region is automatically segmented using the maximum inter-class variance method. Local mean smoothing is performed on the shadow region using a dynamic window, while the original pixel values ​​are preserved in the non-shadow region, thereby eliminating shadow interference. Simultaneously, Canny edge detection (CannyEdgeDetection, a multi-stage edge detection algorithm) is used to identify wrinkle edges, and the wrinkle pixel coordinates are recorded. Based on the wrinkle geometric features, a polynomial deformation model (a model that uses a polynomial function y=ax) is constructed. 2 (A mathematical model for fitting the wrinkle curve using +bx+c to achieve pixel position remapping) After fitting the wrinkle curve using the least squares method, the pixels in the wrinkle area are remapped and stretched. The process is iterated and optimized 5 times to ensure texture continuity, restoring the wrinkle area to a smooth surface, thus obtaining the corrected archival image. Among them, the adaptive correction of shadows and wrinkles refers to an image processing technique that adaptively adjusts processing parameters according to the distribution characteristics and geometric shape of shadows and wrinkles in the image. It can specifically eliminate different types of shadow interference and repair wrinkles of different degrees. The corrected archival image is an image after shadow elimination and wrinkle correction, in which the official seal area has no obvious shadow occlusion, the paper texture is continuous and smooth, and the details of the official seal are fully presented. By eliminating the grayscale distortion of the official seal area caused by shadows and the deformation of the official seal shape caused by wrinkles, the positioning offset of the official seal and the feature extraction deviation are avoided.

[0087] S15. Perform quality inspection on the corrected archive images, and use the corrected archive images that pass the quality inspection as the target image data.

[0088] In this embodiment of the invention, the image is corrected by calculating the sharpness (average gradient value calculated using the Laplacian gradient operator, required to be ≥50), noise-free rate (statistical proportion of noise pixels, required to be ≤0.5%, noise pixels are defined as pixels whose grayscale value differs from the surrounding 8 neighboring pixels by ≥30), and shadow-free rate (statistical proportion of shadow area pixels, required to be ≤0.1%). If all three indicators meet the standards, the image is deemed to be of acceptable quality; otherwise, the process returns to the preprocessing steps corresponding to S12-S14 for reprocessing until the image passes the inspection, thereby determining the corrected image of acceptable quality as the target image data. Here, quality inspection refers to the process of quantitatively evaluating the image using preset image quality evaluation indicators (such as sharpness, noise ratio, and shadow ratio) to determine whether it meets the requirements for subsequent processing. The target image data is a standardized image of the archive that has been preprocessed and passed the quality inspection, and can be directly used for subsequent seal extraction and feature analysis. The target image data fully preserves the detailed features of the seal area, avoiding a decrease in the accuracy of seal recognition due to substandard image quality.

[0089] Step 102: Extract the official seal from the target image data to obtain the official seal image unit, the semantic vector of the archive text, and the text surrounding the official seal.

[0090] Furthermore, step 102 may include the following sub-steps:

[0091] S21. Extract the candidate region of the suspected official seal from the target image data.

[0092] In this embodiment of the invention, the target image data is first converted into the HSV color space (HSV color space, a color encoding method based on hue, saturation, and brightness, where H represents hue, S represents saturation, and V represents brightness). Based on the common red characteristics of official seals, the HSV red threshold range is set as H∈[0, 10]∪[160, 180], S∈[40, 255], V∈[40, 255]. Pixel regions in the image that meet the red characteristics are filtered out and a binary color mask is generated. Then, the connected regions in the mask are first dilated and then eroded to eliminate small interference regions. Subsequently, the circularity of each connected region is calculated (circularity = 4π × area / perimeter). 2Official seals are mostly circular or near-circular, with a circularity threshold set to 0.7-1.0. The area (the actual size of the seal corresponds to an image pixel area of ​​typically 500-5000 pixels, hence the area threshold) is used to retain connected regions that simultaneously satisfy both the circularity and area thresholds. Finally, regions within 10 pixels of the image edge are excluded to avoid edge interference, thus obtaining potential seal candidate regions. These potential seal candidate regions refer to areas selected from the target image data that match the seal features in color, shape, and area. These regions may contain real seals and require further verification. Through multi-dimensional filtering based on color, shape, and position, potential seal regions are initially identified, narrowing the scope for accurate extraction of seal image units and effectively reducing interference from irrelevant regions in subsequent recognition.

[0093] S22. Perform seal boundary edge detection on the suspected seal candidate area to obtain the seal boundary edge image.

[0094] In this embodiment of the invention, the suspected official seal candidate region is first subjected to a 3×3 Gaussian filter, and then the Sobel operator (a Sobel operator for calculating image gradients, which calculates the gradients in the horizontal and vertical directions respectively) is used to calculate the magnitude and direction of the image gradient. Non-edge pixels are suppressed by double thresholding (high threshold 100-150, low threshold 50-80), high threshold edges are retained and low threshold edges are connected to form an initial edge image. Subsequently, morphological dilation and erosion are performed on the initial edge image to finally obtain a continuous and complete official seal boundary contour, thus obtaining the official seal boundary edge image. Here, official seal boundary edge detection refers to the process of extracting the boundary line between the official seal and the background in the suspected official seal candidate region through image processing algorithms to determine the official seal contour. The official seal boundary edge image is a binary image that retains only the pixels of the official seal boundary contour, clearly presenting the shape contour of the official seal. By enhancing the continuity and integrity of the edges, the accurate positioning of the official seal boundary is ensured, avoiding the deviation in official seal extraction caused by edge breakage or noisy edges.

[0095] S23. Separate the official seal image unit from the boundary edge image of the official seal.

[0096] In this embodiment of the invention, gradient enhancement is first performed on the edge image of the official seal boundary. Then, starting from the first edge pixel at the top left corner of the edge image, an eight-neighborhood tracking algorithm (detecting neighboring pixels sequentially in a clockwise direction, and continuing to track edge pixels) is used to record the coordinates of all pixels on the contour, forming a closed official seal contour. The minimum bounding rectangle is calculated based on this closed contour, and it is expanded outward by 2-3 pixels to ensure complete coverage of the official seal. The rectangular area is then cropped to obtain the original official seal image. Finally, a bilinear interpolation algorithm is used to uniformly adjust the image to a 256×256 pixel grayscale image, thereby obtaining the official seal image unit. Here, the official seal image unit refers to an independent image region containing complete official seal information separated from the target image. After normalization processing, it can be directly used for subsequent feature extraction. Through accurate extraction and normalization processing, the official seal region is separated from the background and standardized, providing a uniform image basis and avoiding feature extraction errors caused by differences in image size or incomplete edge cropping.

[0097] S24. Extract the semantic vector of the archive text and the surrounding text of the official seal from the target image data.

[0098] In this embodiment of the invention, using the center coordinates of the official seal image unit obtained in S23 as the origin, a 200-pixel × 200-pixel rectangular area is delineated in the target image data as the text extraction range, ensuring coverage of the text surrounding the official seal. Then, text recognition is performed on this rectangular area, removing special characters (such as @, #, &, etc.) from the recognition results, and processing is done using Jieba word segmentation (for Chinese scenarios) or space splitting (for English scenarios) to obtain a word segmentation list. This list represents the text surrounding the official seal. Subsequently, the word segmentation list is input into a pre-trained Word2Vec model (Word2Vec...). Model, a language model that maps words to low-dimensional vectors, uses a training corpus containing 100,000 archival texts (with word vector dimensions set to 100). It obtains a 100-dimensional word vector for each word and then calculates the weighted average of all word vectors based on word frequency normalization weights (number of occurrences of the word / total number of words) to obtain the archival text semantic vector. The archival text semantic vector refers to the low-dimensional numerical vector that transforms the semantic information of the archival text, which can be used to quantify and represent the semantics of the text. The text surrounding the official seal refers to the text content in the area surrounding the official seal in the target image data, typically containing key information such as the unit name and document type associated with the seal.

[0099] Step 103: Use the official seal image unit, the semantic vector of the archive text and the surrounding text of the official seal to perform feature analysis to obtain the calibration official seal feature dataset.

[0100] Furthermore, step 103 may include the following sub-steps:

[0101] S31. Perform grayscale conversion and pixel normalization on the official seal image unit to obtain a normalized grayscale official seal image.

[0102] In this embodiment of the invention, it is first determined whether the official seal image unit is a color image. If it is a color image, a weighted average method (grayscale value = 0.299 × R + 0.587 × G + 0.114 × B, where R, G, and B are the red, green, and blue channel pixel values ​​of the color image, respectively) is used to convert it into a single-channel grayscale image. If it is already a grayscale image, the process proceeds directly to the next step. Then, the pixel values ​​of the grayscale image are linearly normalized to map all pixel values ​​to a uniform range of 0-1, eliminating the influence of differences in pixel value magnitude under different lighting conditions. Subsequently, the average grayscale value of the normalized image is calculated. If the average grayscale value is between 0.3 and 0.114, the process is complete. If the value is within the range of 0.7, the processing is considered qualified, resulting in a normalized grayscale seal image. Grayscale conversion refers to the image processing process of merging the three-channel pixel values ​​of a color image into a single-channel grayscale value using a specific algorithm, removing color information and retaining brightness information. Pixel normalization refers to the preprocessing operation of mapping image pixel values ​​to the range of 0-1, eliminating differences in pixel value magnitudes and unifying data scales. The normalized grayscale seal image is a seal image that retains only brightness information and has a unified pixel value scale after grayscale and pixel normalization processing. By unifying the image format and pixel scale, the influence of color interference and pixel value differences on feature extraction is avoided.

[0103] S32. Extract the shape feature vector, texture feature vector, and text content feature vector from the normalized grayscale official seal image.

[0104] In this embodiment of the invention, for shape features, a shape convolution kernel of size K×K (e.g., 5×5) is used, and the coordinates are calculated using the following formula. The shape feature response values ​​at the specified locations are used to generate a shape feature map. Then, the ReLU activation function is applied to the shape feature map to suppress negative response values. Next, a 2×2 max pooling operation is performed to compress the feature map size. Finally, global average pooling is performed on the pooled feature map, which calculates the average value of all pixels in the feature map, resulting in a 1×32 dimensional shape feature vector. .

[0105]

[0106] In the formula, Indicates coordinates Shape feature response value at the location, Indicates the size of the convolution kernel. Indicates the convolution kernel in coordinates The weight parameters at that location, This represents the pixel value at the corresponding position in the input image. This represents the bias term for shape feature extraction.

[0107] For texture features, a texture convolution kernel of size K×K (e.g., 5×5) is used, and the coordinates are calculated using the following formula. The texture feature response values ​​at each location are used to generate a texture feature map. Then, the ReLU activation function is applied to the texture feature map to suppress negative response values. Next, a 2×2 max pooling operation is performed to compress the feature map size. Finally, global average pooling is performed on the pooled feature map, which calculates the average value of all pixels in the feature map, resulting in a 1×64 dimensional texture feature vector. .

[0108]

[0109] In the formula, Indicates coordinates Texture feature response value at that location, This represents the bias term for shape feature extraction.

[0110] For text content features, Canny edge detection is first applied to the normalized grayscale seal image to enhance text edges. Then, a text convolution kernel of size K×K (e.g., 3×3) is used to calculate the coordinates using the following formula. The text content feature response values ​​at each location are used to generate a text content feature map. This map is then flattened into a one-dimensional vector and input into a fully connected layer with 128 neurons. The output of the fully connected layer is L2 normalized to obtain a 1×128 dimensional text content feature vector. .

[0111]

[0112] In the formula, Indicates coordinates The text content feature response value at that location, The confidence factor represents the text feature extraction. This parameter represents the text content at the corresponding position in the input image.

[0113] The shape feature vector is a low-dimensional numerical vector used to characterize the shape and outline features of the official seal; the texture feature vector is a low-dimensional numerical vector used to characterize the surface texture details of the official seal; and the text content feature vector is a low-dimensional numerical vector used to characterize the semantic features of the text content on the official seal.

[0114] By utilizing convolutional neural network feature modeling technology, multi-dimensional data analysis and deep feature extraction are performed on the shape, texture, and text details of official seals, providing a reliable data foundation for the identification of official seal archives.

[0115] S33. Concatenate the shape feature vector, texture feature vector, and text content feature vector to obtain the initial official seal feature dataset.

[0116] In this embodiment of the invention, the obtained 1×32-dimensional shape feature vector, 1×64-dimensional texture feature vector, and 1×128-dimensional text content feature vector are horizontally concatenated in sequence to integrate them into a 1×224-dimensional feature vector set. This set contains complete feature information of the official seal in the three dimensions of shape, texture, and text content, thus obtaining the initial official seal feature dataset. Vector concatenation refers to the operation of combining multiple feature vectors of different dimensions into a higher-dimensional feature vector in a preset order, enabling the fusion of multi-dimensional features. By integrating multi-dimensional features through vector concatenation, the limitations of a single feature dimension are avoided.

[0117] S34. Use the shape feature vector and texture feature vector in the initial official seal feature dataset to perform feature comparison. If the feature comparison result does not meet the preset feature comparison conditions, then perform image enhancement on the official seal image unit.

[0118] In this embodiment of the invention, the cosine similarity between the texture feature vector in the initial official seal feature dataset and the standard texture feature vector is calculated:

[0119]

[0120] In the formula, Represents cosine similarity. Represents texture feature vectors, This represents the standard texture feature vector.

[0121] Calculate the Euclidean distance between the shape feature vector in the initial official seal feature dataset and the standard shape feature vector:

[0122]

[0123] In the formula, Represents Euclidean distance. The first eigenvector representing the shape feature vector One element, This represents the first feature vector of the shape of a standard official seal of the same type in the official seal specification database. Each element.

[0124] Preset feature comparison conditions are ≤0.5 and ≥0.8, if >0.5 or If the value is less than 0.8, meaning the preset conditions are not met, image enhancement is performed on the official seal image unit. Then, the shape and texture feature vectors of the enhanced image are extracted again and compared again until the comparison result meets the threshold requirements, thereby completing the feature calibration process.

[0125] Feature matching refers to the process of determining whether the feature matching degree meets the standard by calculating the distance or similarity between the initial feature vector and the standard feature vector; preset feature matching conditions refer to the Euclidean distance and cosine similarity thresholds set based on the standard official seal features, which are used to define whether the features are qualified; image enhancement can improve the local contrast and details of the image; by comparing with standard features and performing image enhancement and feature re-extraction for cases that do not meet the conditions, the accuracy of shape and texture features is ensured.

[0126] Since shape and texture features are the most stable visual intrinsic attributes of official seals and are not affected by changes in text content, while text content features are prone to fluctuations due to differences in font and layout, prioritizing the verification of the former two can capture the core stable features of official seal recognition. Euclidean distance is used to measure the difference in shape features because it can accurately quantify the global deviation of the shape contour in the vector space, which meets the requirement of shape features for overall shape consistency. Cosine similarity is more suitable for evaluating the directional consistency of texture features and can effectively capture the similarity of texture distribution. The combination of the two not only ensures the targeting of feature comparison, but also improves the stability and efficiency of verification by avoiding the variability of text features.

[0127] S35. Based on the enhanced official seal image unit, jump to execute the steps of performing grayscale conversion and pixel normalization on the official seal image unit to obtain a normalized grayscale official seal image, until the feature comparison result meets the preset feature comparison conditions, and output the adjusted feature data.

[0128] In this embodiment of the invention, after completing the image enhancement of the official seal image unit, the process jumps to perform grayscale conversion and pixel normalization on the official seal image unit to obtain a normalized grayscale official seal image, until the feature comparison result meets the preset feature comparison conditions, and then the adjusted feature data is output.

[0129] Adjusted feature data refers to the set of multi-dimensional features of the official seal that meets the preset feature comparison conditions after multiple feature comparisons and image enhancement optimizations, and is more in line with the standard features than the initial features.

[0130] S36. Input the enhanced official seal image unit after adjusting the feature data association into the preset text recognition model, and output the text recognition result.

[0131] In this embodiment of the invention, an image unit of an official seal, after image enhancement and correlation adjustment of feature data, is input into a pre-trained CRNN text recognition model. The model structure is CNN feature extraction + RNN temporal modeling + CTC decoding. The training set includes common Song and Li fonts for official seals. After the model extracts features and performs temporal modeling on the text content in the official seal image unit, it outputs the text recognition result, such as "XX Company Official Seal," through CTC decoding, thus obtaining the text recognition result. The pre-set text recognition model refers to a pre-trained deep learning model specifically optimized for official seal text recognition, which can accurately recognize the text content on the official seal; the text recognition result refers to the recognized text of the text content in the official seal image unit output by the model.

[0132] S37. Compare the text recognition results with the surrounding text of the official seal to obtain the recognition accuracy.

[0133] In this embodiment of the invention, the text recognition result (e.g., "XX Company Official Seal") and the surrounding text of the official seal (e.g., "XX Co., Ltd. Document Seal" identified from a 200-pixel × 200-pixel area around the official seal) are first preprocessed to remove punctuation marks, spaces, and other irrelevant characters, and are uniformly converted to lowercase (for English) or retain the original characters (for Chinese). Then, the difference between the two is calculated using the edit distance algorithm (LevenshteinDistance, which calculates the minimum number of editing operations between two strings, including insertion, deletion, and replacement). The matching degree is quantified by the formula "Recognition Accuracy = (1 - Edit Distance / Length of Longer String) × 100%". For example, if the text recognition result is "XX Company Official Seal", the surrounding text is "XX Co., Ltd. Official Seal", the edit distance is 2 (requiring the insertion of the word "Limited"), and the length of the longer string is 6, then the recognition accuracy = (1 - 2 / 6) × 100% ≈ 66.7%, thus obtaining the recognition accuracy. Here, the recognition accuracy refers to the degree of matching between the text recognition result and the surrounding text of the official seal, expressed as a percentage, and is used to measure the accuracy of text recognition.

[0134] S38. When the recognition accuracy is greater than or equal to the preset accuracy threshold, the adjustment feature data and the text content feature vector will be integrated to obtain the calibration seal feature dataset.

[0135] In this embodiment of the invention, the obtained recognition accuracy is compared with a preset accuracy threshold (set to 85% based on the official seal text recognition scenario to ensure that the text content matching degree reaches a practical standard). If the recognition accuracy is ≥85%, the text recognition result is determined to be reliable. At this time, the adjusted feature data output by S35 (including optimized shape feature vector and texture feature vector) and the text content feature vector extracted by S32 are weighted and integrated. The weights are allocated as follows: shape feature vector accounts for 30%, texture feature vector accounts for 40%, and text content feature vector accounts for 30%. The fused vector set is obtained through weighted operation, which is the calibrated official seal feature dataset. The preset accuracy threshold refers to the minimum accuracy standard used to judge whether the text recognition result is reliable. The calibrated official seal feature dataset refers to the official seal feature set that includes shape, texture, and text features and meets the reliability requirements after multi-dimensional feature optimization and text content verification.

[0136] S39. When the recognition accuracy is less than the preset accuracy threshold, adjust the model parameters of the preset text recognition model.

[0137] In this embodiment of the invention, if the recognition accuracy obtained in S37 is <85%, the differences between the text recognition result and the surrounding text of the official seal (such as misrecognized or missed characters) are marked as difficult examples. The difficult examples and the corresponding image-enhanced official seal image units are combined to form a fine-tuning dataset, which is input into a preset CRNN text recognition model for parameter fine-tuning. The model parameters are adjusted by iteratively training for 5-10 rounds until the loss value converges. The model parameter adjustment refers to the process of fine-tuning the weight parameters of some layers of the model to adapt the model to difficult examples and improve the recognition ability in specific scenarios. This step follows the recognition accuracy obtained in S37 and the preset text recognition model used in S36. By specifically fine-tuning the model parameters corresponding to the difficult examples, the text recognition deviation problem is solved and the model's adaptability to the recognition of official seal text is enhanced.

[0138] S310. Based on the adjusted model parameters, jump to execute the step of inputting the enhanced official seal image unit associated with the adjusted feature data into the preset text recognition model and outputting the text recognition result until the recognition accuracy is greater than or equal to the preset accuracy threshold, and output the calibrated official seal feature dataset.

[0139] In this embodiment of the invention, after adjusting the model parameters in step S39, the enhanced seal image unit associated with the adjusted feature data is re-input into the CRNN text recognition model after parameter fine-tuning, and a new text recognition result is output. Then, step S37 is executed to compare the new text recognition result with the surrounding text of the seal to obtain the recognition accuracy. If it is still <85%, the process of "model parameter fine-tuning → text recognition → accuracy comparison" is repeated, iterating up to 3 times to avoid overfitting, until the recognition accuracy is ≥85%. At this time, the adjusted feature data and the corresponding text content feature vector are weighted and integrated according to the weights of shape, texture, and text features at 30%, 40%, and 30%, respectively, to obtain the calibrated seal feature dataset. The calibrated seal feature dataset is the final seal feature set that integrates multi-dimensional features of shape, texture, and text and meets the recognition accuracy requirements. Through the iterative optimization mechanism, it is ensured that the output feature dataset meets the reliability standard at the text recognition level.

[0140] Step 104: Compliance and integrity verification is performed using the calibration seal feature dataset, seal image unit, seal surrounding text and archive text semantic vector to obtain structured classified archive data.

[0141] Furthermore, step 104 may include the following sub-steps:

[0142] S41. Using the calibration of the official seal feature dataset, the semantic vectors of the text surrounding the official seal and the archival text, classification and labeling are performed to obtain the official seal classification label data of the target archive.

[0143] Furthermore, S41 may include the following sub-steps:

[0144] S411. Extract multiple surrounding text fragments from the text surrounding the official seal.

[0145] In this embodiment of the invention, the text surrounding the official seal extracted in S24 is first processed into sentences, using periods, semicolons, and line breaks as delimiters to obtain an initial set of sentences. Then, through keyword matching (filtering sentences containing keywords highly related to the official seal, such as "company," "seal," "date," and "number") and length filtering (retaining effective text of 5-30 characters and removing overly short, meaningless fragments or overly long, redundant content), 8-12 candidate text fragments are extracted from the initial set of sentences. Finally, a sliding window method is used to perform a second screening of the candidate fragments to ensure that the extracted text fragments cover the core information surrounding the official seal while avoiding duplication and redundancy, thereby obtaining multiple surrounding content text fragments. Among them, the surrounding content text fragments refer to local text units extracted from the text surrounding the official seal that have a semantic relationship with the official seal and can reflect key information such as the subject to which the official seal belongs and the usage scenario.

[0146] S412. Calculate the text similarity between the text content feature vector in the calibrated official seal feature data set and multiple surrounding content text segments.

[0147] In an embodiment of the present invention, first perform L2 normalization on the 1×128-dimensional text content feature vector in the calibrated official seal feature data set, and then convert each surrounding content text segment into a 1×100-dimensional segment semantic vector through the Word2Vec model, and perform the same normalization operation on it; subsequently, use the cosine similarity formula to calculate the similarity between the text content feature vector and each surrounding content text segment, obtaining a set of similarity values between 0 and 1, thereby completing the calculation of text similarity; where the text similarity refers to the matching degree between the text content feature vector and the surrounding content text segment at the semantic level, and is used to measure the closeness of the association between the two.

[0148] S413. Select a preset number of surrounding content text segments as the target surrounding content text according to the text similarity from large to small.

[0149] In an embodiment of the present invention, sort the text similarity values calculated in S412 in descending order, extract the top 3 surrounding content text segments in the sorting, and determine them as the target surrounding content text; if the total number of surrounding content text segments is less than 3, then all are selected as the target surrounding content text; where the target surrounding content text refers to the key text segments screened from the surrounding content text segments and having the most matching semantics with the official seal text content features, which can effectively reflect the associated semantics between the file and the official seal.

[0150] S414. Input each target surrounding content text into a preset word vector model to output multiple word semantic vectors.

[0151] In an embodiment of the present invention, perform word segmentation on each target surrounding content text using a Chinese word segmentation tool, remove stop words such as "of", "in", etc., to obtain a word sequence composed of core words; input the word sequence into a pre-trained Word2Vec word vector model, and after the model generates a 1×100-dimensional word vector for each word, by taking the average value of all word vectors in the word sequence, obtain a 1×100-dimensional word semantic vector corresponding to each target surrounding content text; where the preset word vector model refers to a pre-trained model that can convert text vocabulary into low-dimensional semantic vectors and can quantify the semantic information of the vocabulary; the word semantic vector refers to a low-dimensional vector generated by the word vector model and representing the semantics of the target surrounding content text; this step converts the text information into a computable vector form.

[0152] S415. Calculate the semantic matching degree between the file text semantic vector and each word semantic vector.

[0153] In this embodiment of the invention, the weighted average semantic similarity between the archive text and surrounding content is obtained through the following formula, thereby completing the calculation of the semantic matching degree:

[0154]

[0155] In the formula, Indicates semantic matching degree. Indicates the first The weight of each surrounding text fragment. Represents the semantic vector of the archival text. Indicates the first Word semantic vectors, This indicates the number of text fragments surrounding the selected target.

[0156] S416. Retain the semantic vectors of words associated with semantic matching degree greater than or equal to the preset matching degree threshold.

[0157] In this embodiment of the invention, a preset matching degree threshold is set to 0.7. If the value is ≥0.7, the semantic vector of the currently associated word and the corresponding text fragment are directly retained as valid data for semantic matching.

[0158] S417. Based on the archive domain dictionary and the calibration seal feature dataset, correct the semantic vectors of words associated with semantic matching degree less than the preset matching degree threshold.

[0159] In an embodiment of the present invention, if If the score is less than 0.7, then the sentence structure of the archival text is parsed using dependency parsing, such as extracting core subject-verb-object components like "Party A (subject), signing (predicate), contract (object)". Simultaneously, an archival domain dictionary, such as one containing industry terms like "official seal", "contract", and "approval", is used to correct semantic comprehension biases and regenerate the archival text's semantic vector. Then, based on the text content feature vectors in the calibrated official seal feature dataset, the generation logic of the word semantic vectors is adjusted, and the semantic matching degree is recalculated until... ≥0.7.

[0160] S418. Integrate the corrected word semantic vectors with the retained word semantic vectors to obtain the semantic classification results.

[0161] In this embodiment of the invention, if there is a case where the semantic matching degree is less than the preset matching degree threshold, the semantic vector of the corresponding word is first corrected based on the archive domain dictionary and the calibration official seal feature dataset. Then, the semantic vector of the word retained in S416 is merged with the semantic vector of the word corrected in S417. Based on their semantic matching degree, weight and text content, the semantic association category of archives and official seals (such as "contract official seal", "approval official seal", etc.) is summarized to form a semantic classification result.

[0162] Semantic classification results refer to the semantic association category identifier between archives and official seals, which is obtained by performing a series of operations such as weighted semantic matching, threshold filtering, and bias correction on the semantic vectors of the archive text and the semantic vectors of the words surrounding the official seal.

[0163] S419. Use the semantic vector of the archival text as input to the pre-trained archival classification model and output the corresponding archival category label.

[0164] In this embodiment of the invention, the semantic vector of the archival text is input into an archival classification model based on a pre-trained architecture such as BERT or TextCNN. This model has been trained on a massive amount of archival corpus and can identify categories such as contract archives, approval archives, and meeting minutes archives. The model directly outputs the category label to which the target archive belongs by extracting features from the semantic vector and performing classification reasoning.

[0165] Archival classification models based on pre-trained architectures such as BERT or TextCNN refer to models specifically designed for classifying archival texts. These models are based on mature natural language processing architectures such as BERT (Bidirectional Transformer Encoder) and TextCNN (Text Convolutional Neural Network). They are pre-trained and fine-tuned on archival corpora and are specifically designed for classifying archival texts. The BERT architecture relies on the bidirectional Transformer mechanism to capture the semantic relationships between text contexts, enabling a deep understanding of complex sentence structures in archives, such as contract terms and approval process descriptions. TextCNN, on the other hand, extracts key local features of the text through multi-size convolutional kernels, such as archive-specific fields like official seal numbers and filing dates, balancing classification efficiency and feature capture capabilities. These models take the semantic vector of the archival text as input, and use the weight parameters formed by pre-training to perform feature mapping and category inference. Finally, they output standardized labels such as contract archives and approval archives, achieving automated classification of archives and providing a business-dimensional category benchmark for official seal classification and archive verification.

[0166] Archive category labels are category identifiers output by the archive classification model after classifying and reasoning the semantic vectors of the archive text, such as contract archives and approval archives, and are used to identify the business category to which the archives belong.

[0167] S4110. Use semantic classification results and archive category labels to perform label mapping to obtain the official seal classification identifier data of the target archive.

[0168] In this embodiment of the invention, a mapping relationship table of "semantic classification result - archive category label" is established, such as "contract-type official seal" mapping to "contract archive". The semantic classification result of S418 is matched with the archive category label of S419, and the structured official seal classification identification data containing "official seal type, archive category and semantic matching degree" is output. This step integrates and maps multi-dimensional information to finally form an official seal classification identification conclusion that can be used for the verification of archive compliance and integrity.

[0169] Official seal classification identification data is structured data formed by mapping semantic classification results with archive category labels. It contains information such as "official seal type, archive category, semantic matching degree" and is used for compliance and integrity verification of archives, clarifying the classification relationship between target archives and official seals.

[0170] S42. Compliance determination is performed using the calibrated official seal feature dataset and the preset official seal feature specification database to obtain the official seal compliance result of the target file.

[0171] Furthermore, S42 may include the following sub-steps:

[0172] S421. The feature comparison is performed using the calibration seal feature dataset and the preset seal feature specification database to obtain feature values ​​of multiple different feature types, including shape features, texture features and text content features.

[0173] In this embodiment of the invention, firstly, based on the unit to which the target file belongs, the corresponding standard official seal data is queried from the official seal feature specification database (stored in a MySQL database) which stores standard shape vectors, texture vectors, text feature vectors, and size, color, and font standards; then, multi-dimensional comparisons are performed on the calibrated official seal feature dataset: when comparing shape features, the Euclidean distance between the calibrated shape vector and the standard shape vector is calculated (required ≤0.4); when comparing texture features, the cosine similarity between the calibrated texture vector and the standard texture vector is calculated (required ≥0.85); when comparing text content features, the edit distance between the calibrated text vector and the standard text vector is calculated (required ≤3); through the above operations, the feature values ​​corresponding to the three types of features—shape, texture, and text content—are obtained.

[0174] The calibration seal feature dataset refers to a structured dataset formed after feature extraction and error correction of the seal image units in the target archive. The core of it includes the shape feature vector, texture feature vector, text content feature vector of the seal, as well as calibration parameters after noise and geometric distortion are eliminated during the feature extraction process.

[0175] The pre-built official seal feature specification database refers to a structured database that is pre-built and stored in databases such as MySQL, containing standard features of various compliant official seals. The core data includes: standard shape parameters, standard texture templates, standard text content specifications of compliant official seals in various units / scenarios, as well as the qualified thresholds of the corresponding features, such as shape Euclidean distance ≤ 0.4 and texture cosine similarity ≥ 0.85. These serve as compliance benchmarks for official seal feature comparison and are used to determine whether the calibrated official seal features meet the specifications.

[0176] Shape features refer to the core features that characterize the geometric form of an official seal. Specifically, these include the seal's outline type (circular, square, elliptical, etc.), outline parameters (such as the circumcircle radius of a circular seal and the interior angles of a square seal), and the layout of internal geometric elements (such as the position of the pentagram and the width of the ring edge). These features are usually stored in vector form (such as outline coordinate vectors). By calculating the Euclidean distance with standard shape parameters, the shape of the official seal is quantitatively judged to determine whether it is compliant. This is a key indicator for verifying the compliance of the physical form of the official seal.

[0177] Texture features refer to the characteristics that characterize the surface texture and pattern details of an official seal. Specifically, they include the contrast, correlation, and energy of the seal's anti-counterfeiting texture (extracted through the gray-level co-occurrence matrix), the pixel distribution pattern of internal patterns (such as five-pointed stars and floral patterns), and edge clarity. These features are stored in the form of feature vectors. By calculating the cosine similarity with the standard texture template, the consistency of the seal's texture with the standard is quantitatively determined and used to identify the compliance of the seal's anti-counterfeiting features.

[0178] Text content features refer to the characteristics that characterize the text information on an official seal. Specifically, these include the text content (such as the full name of the organization, the word "official seal," and the seal number), text format (font type, font size, character spacing), and text layout (circular arrangement, centered arrangement), etc. They are usually converted into text content feature vectors (such as those generated by Word2Vec). By calculating the edit distance and cosine similarity with standard text vectors, the text on the official seal is quantitatively judged to determine whether it conforms to the specifications. This is the core verification basis for avoiding "fake seals" and "seals with misspelled characters."

[0179] Feature values ​​refer to numerical values ​​that are output after the calibration official seal feature dataset is compared with the preset official seal feature specification database, and are used to quantify the degree of matching of a certain type of feature. Different feature types correspond to feature values ​​with different meanings: for example, the shape feature value is "Euclidean distance" (the smaller the value, the more compliant the shape), the texture feature value is "cosine similarity" (the closer the value is to 1, the more matched the texture), and the text content feature value is "edit distance" (the smaller the value, the smaller the text difference).

[0180] S422. When all types of feature values ​​meet the associated preset feature standard conditions, the calibration seal feature dataset is marked as compliant, the target file is determined to be compliant, and the seal compliance result is generated.

[0181] In this embodiment of the invention, various feature qualification thresholds stored in the preset official seal feature specification database are called, such as Euclidean distance of shape features ≤0.4, cosine similarity of texture features ≥0.85, and edit distance of text content features ≤3. The shape, texture, and text content feature values ​​obtained in S421 are checked one by one to see if they fall within the corresponding threshold range. If all three feature values ​​meet the preset standard conditions, the calibrated official seal feature dataset is marked as "feature qualified", the target file is simultaneously determined as "compliance qualified level", and an official seal compliance result containing "compliance level, feature comparison details (such as each feature value and threshold), and judgment basis" is generated.

[0182] Preset feature standard conditions refer to the qualified threshold ranges set for various official seal features in the preset official seal feature specification database, such as the upper limit of Euclidean distance for shape features and the lower limit of cosine similarity for texture features. These are the benchmark conditions for judging whether the feature values ​​meet the standards.

[0183] The compliance level is one of the classification results of the compliance of the official seal of the target archive. It means that the official seal contained in the archive meets the preset standards in all feature dimensions such as shape, texture and text content, and belongs to the fully compliant archive type.

[0184] The official seal compliance result is structured conclusion data generated after determining the compliance of the official seal in the target document. It includes information such as compliance level, specific numerical values ​​and standard thresholds for each feature comparison, and explanation of the judgment logic, and is used to intuitively present whether the official seal complies with regulations and the specific basis.

[0185] S423. If any type of feature value does not meet the associated preset feature standard conditions, the calibration seal feature dataset will be marked as pending evaluation.

[0186] In this embodiment of the invention, if any of the shape, texture, and text content feature values ​​obtained in S421 do not fall within the preset feature standard conditions, such as shape Euclidean distance > 0.4, texture cosine similarity < 0.85, and text editing distance > 3, then the calibration seal feature dataset is marked as "to be evaluated".

[0187] S424. Using the feature values ​​in the calibration seal feature dataset to be evaluated, calculate the degree of difference for multiple different feature types.

[0188] Difference degree is a quantitative value of the degree of deviation between the feature value to be evaluated and the preset standard conditions. It is divided into three categories: shape difference degree, texture difference degree, and text difference degree, which respectively reflect the degree of deviation between the corresponding feature and the standard. The larger the value, the more serious the deviation.

[0189] In this embodiment of the invention, for the dataset of calibration seal features labeled as to be evaluated, three types of feature differences are calculated respectively:

[0190]

[0191]

[0192]

[0193] In the formula, This represents the Euclidean distance between the calibrated shape vector of the official seal and the standard shape vector. This represents the cosine similarity between the calibrated official seal texture vector and the standard texture vector. This represents the edit distance between the calibrated official seal text vector and the standard text vector. Indicates the degree of shape difference. Indicates texture difference. Indicates the degree of textual difference.

[0194] S425. The target file’s suspicion score is obtained by weighting the calculation using the preset coupling factors of each degree of difference and correlation.

[0195] The preset coupling factor is a weighting coefficient set for different feature differences. It is used to reflect the difference in the degree of influence of each feature on the compliance of the official seal in the comprehensive suspiciousness score. For example, the texture difference weight is 0.4, which indicates that its influence on suspiciousness is relatively greater.

[0196] The comprehensive suspicion score is a quantitative score (range 0-1) obtained by weighting the difference of each feature with a preset coupling factor. The higher the score, the stronger the suspicion of the official seal of the target document.

[0197] In this embodiment of the invention, a preset coupling factor is invoked, with shape difference weighting at 0.3, texture difference weighting at 0.4, and text difference weighting at 0.3. A weighted average formula is then used to calculate the comprehensive suspiciousness score of the target file. :

[0198]

[0199] S426. Based on the preset risk threshold range to which the comprehensive doubt score belongs, determine the official seal compliance result of the target file. The official seal compliance result includes compliance level, compliance pending verification level, compliance defect level, and compliance violation level.

[0200] In this embodiment of the invention, preset risk threshold ranges are invoked, such as 0≤score<0.2 corresponding to compliance compliance level, 0.2≤score<0.5 corresponding to compliance pending review level, 0.5≤score<0.8 corresponding to compliance defect level, and score≥0.8 corresponding to compliance violation level. The comprehensive suspiciousness score obtained in S425 is matched to the corresponding threshold range. If the score<0.2, it is judged as compliance compliance level, although the previous features did not fully meet the standards, the overall difference is minimal. If 0.2≤score<0.5, it is judged as compliance pending review level, which requires manual review of feature differences. If 0.5≤score<0.8, it is judged as compliance defect level, which has obvious feature deviations but is not malicious violation. If the score≥0.8, it is judged as compliance violation level, which has features that seriously deviate from the standard and is suspected of being forged. Finally, the seal compliance result containing the compliance level, score value and corresponding threshold range is output.

[0201] S43. Perform integrity verification on the official seal image unit and obtain the integrity verification result.

[0202] Integrity verification is a process of checking whether the outline boundary, internal core elements and layout structure of the official seal image unit are complete and without missing parts. It uses quantitative indicators to determine whether the official seal has physical damage or defects.

[0203] The integrity verification result is a judgment on the integrity of the official seal image unit, including a status indicator of whether it is complete or missing, as well as a specific description of the missing location (such as the edge, pentagram, or text area), which is used to reflect the integrity of the physical form of the official seal.

[0204] In this embodiment of the invention, the outline boundary and internal core elements of the official seal image unit, such as a five-pointed star, text ring, and anti-counterfeiting pattern, are first extracted. The integrity of the seal edge is determined by calculating the outline closure (≥95%). The integrity of the internal structure is determined by detecting the pixel coverage of the core elements (e.g., the pixel ratio of the five-pointed star is ≥80%, and the pixel loss rate of the text ring is ≤5%). At the same time, the element layout of the preset complete official seal template is compared (e.g., the relative position deviation between the text ring and the five-pointed star is ≤2 pixels). If all the above indicators meet the preset threshold, it is determined to be "complete". Otherwise, it is marked as "missing" (e.g., broken edges, incomplete five-pointed stars, missing text, etc.). Finally, the integrity verification result containing the integrity status and the specific missing location is output.

[0205] S44. Import the official seal classification identification data, the calibration official seal feature dataset, the official seal compliance results, and the integrity verification results into the preset structured data template to obtain structured classification archive data.

[0206] In this embodiment of the invention, a preset structured data template, such as JSON format, containing fields for official seal classification, features, compliance, and integrity, is invoked. This template includes fields such as "Official Seal Type," "Shape Feature Value," "Compliance Level," and "Integrity Status." The official seal classification identifier data from S4110, the calibration official seal feature dataset from S38, the official seal compliance results from S422 or S426, and the integrity verification results from S43 are filled into the corresponding fields. After format verification, structured classified archive data integrating official seal classification, features, compliance, and integrity information is output. This step is the final integration stage of the archive's full-dimensional official seal information, providing standardized data support for archive storage, retrieval, and auditing.

[0207] Step 105: Generate an archive audit report for the target archive based on the structured classification archive data.

[0208] Furthermore, step 105 may include the following sub-steps:

[0209] S51. Perform data cleaning on structured classified archive data.

[0210] Data cleaning is the process of identifying, correcting, or removing outliers, missing values, duplicate data, and other non-compliant data in structured classified archives, with the aim of improving data quality.

[0211] In this embodiment of the invention, outliers, missing values, and duplicate data in structured classified archive data are identified through data verification rules such as field non-empty verification, data format matching verification, and numerical range reasonableness verification. Outliers are replaced with the mean or truncated. Missing values ​​are filled with default values ​​or marked as to be supplemented according to the importance of the field. Duplicate data retains the latest record and deletes redundant items, and finally, cleaned and standardized structured data is obtained.

[0212] S52. Perform consistency verification between the structured classification archive data and the cleaned structured classification archive data.

[0213] In this embodiment of the invention, key information of core fields (such as official seal type, compliance level, suspiciousness score, integrity status, etc.) is first extracted from the two types of data. By comparing the fields one by one, it is checked whether the cleaned data retains the core semantics of the original data (such as the compliance level changing from "compliance pending verification" to "compliance met" and marking the difference). At the same time, the field matching rate (≥95%) and numerical deviation rate (such as feature value deviation ≤0.01) of the two types of data are calculated. If the field matching rate and numerical deviation rate both meet the preset thresholds and the core information has not been tampered with, the consistency verification is determined to be passed and the result "verification passed" is output. If there are cases where core fields do not match or the numerical deviation exceeds the standard, the difference fields are marked (such as the original value of "suspiciousness score" is 0.35, and after cleaning it is 0.42) and the result "difference needs to be reviewed" is output.

[0214] Consistency verification refers to the process of comparing the core information of the original structured classification archive data with the cleaned structured classification archive data to verify whether the cleaned data is consistent with the core semantics and key values ​​of the original data. The purpose is to avoid introducing new errors during data cleaning.

[0215] The consistency verification result is the judgment conclusion after consistency verification, including two categories: "verification passed" (core information is consistent and there are no abnormal deviations) and "differences need to be reviewed" (there are core field mismatches or numerical deviations exceeding the standard). It is used to determine whether the cleaned data can be used for report generation.

[0216] S53. When the consistency verification result is passed, the target key is generated by using the official seal compliance result and the integrity verification result in the cleaned structured classification archive data.

[0217] In this embodiment of the invention, the compliance results (such as compliance level, compliance pending review level, etc.) and integrity verification results (such as complete, missing) of the official seals are extracted from the cleaned data. The target key is generated by combining them in the format of "compliance level + integrity status" (such as "compliance level + complete", "compliance violation level + missing, etc.). The target key serves as an index for retrieving risk levels, inherits the consistency verification results, and provides accurate query conditions for risk level matching.

[0218] S54. Use the target key to retrieve the preset file risk key-value pair table and match the corresponding file risk level of the target file.

[0219] In this embodiment of the invention, a preset file risk key-value pair table is invoked, as shown in Table 1 below, and precise matching is performed using the target key: if the target key is "Compliance Compliance Level + Complete", it matches "Low Risk"; "Compliance Compliance Level + Missing" or "Compliance Pending Review Level + Complete" matches "Low to Medium Risk"; "Compliance Pending Review Level + Missing" or "Compliance Flaw Level + Complete" matches "Medium Risk"; "Compliance Flaw Level + Missing" matches "Low to Medium Risk"; and "Compliance Violation Level + Complete" or "Compliance Violation Level + Missing" matches "High Risk". The final output is the matched file risk level. This step achieves automated risk level determination through key-value pair mapping, providing a risk assessment basis for the file audit report.

[0220] Table 1. Preset Archive Risk Key-Value Pair Table

[0221]

[0222] The target key is a string generated by combining the official seal compliance result and the integrity verification result. It is used to uniquely identify the overall status of the file in terms of compliance and integrity, and is a key index for retrieving risk levels.

[0223] The preset file risk key-value pair table is a pre-built mapping table that stores the correspondence between target keys and file risk levels. It includes all possible "compliance-integrity" combinations and their corresponding risk levels, serving as a standard reference for risk level matching.

[0224] Archive risk level: This is a classification result that represents the overall risk level of archives based on target key matching. It includes five levels: low risk, low-medium risk, medium risk, medium-high risk, and high risk. It is used to quantify the risk level of archives and provide a basis for audit decisions.

[0225] S55. Integrate the cleaned and structured classified archive data with the archive risk level to generate an archive audit report for the target archive.

[0226] In this embodiment of the invention, a preset report template is invoked, which includes modules for basic archival information, official seal classification details, feature comparison data, compliance conclusions, integrity status, risk level assessment, and disposal suggestions. The cleaned structured classified archival data is filled into the module blocks, and the archival risk level and corresponding risk description matched by S54 are embedded simultaneously. A standardized archival audit report containing text descriptions, data tables, and key indicator visualization charts is generated through formatting. This report fully presents the target archival information in all dimensions, such as official seal classification, feature compliance, integrity, and risk level, providing the archival management department with structured conclusions that can be directly used for audit decisions.

[0227] An archive audit report is a standardized report generated by integrating structured and classified archive data with archive risk levels. It includes basic archive information, official seal classification details, feature comparison results, compliance conclusions, integrity status, risk level, and disposal recommendations. It is used to systematically present the full-dimensional conclusions of archive audits and support audit decisions.

[0228] Please see Figure 2 , Figure 2 This is a schematic diagram of the structure of an archival information processing system based on official seal recognition, provided in an embodiment of the present invention.

[0229] This invention provides a document information processing system based on official seal recognition, comprising:

[0230] The official seal data acquisition and image preprocessing module is used to acquire the original image of the target document and preprocess it to obtain the target image data;

[0231] The official seal image unit feature extraction module is used to extract the official seal from the target image data, and obtain the official seal image unit, the document text semantic vector and the text surrounding the official seal;

[0232] The official seal image unit feature processing module is used to perform feature analysis using the official seal image unit, the semantic vector of the archive text and the text surrounding the official seal to obtain the calibration official seal feature dataset.

[0233] The official seal compliance judgment and integrity verification module is used to perform compliance and integrity verification by using the calibrated official seal feature dataset, official seal image unit, official seal surrounding text and archive text semantic vectors to obtain structured classified archive data;

[0234] The official seal suspiciousness analysis, assessment and early warning module is used to generate an archive audit report for the target archive based on structured classified archive data.

[0235] Furthermore, the official seal data acquisition and image preprocessing module includes:

[0236] The original archive image submodule is used to obtain the original archive image of the target archive;

[0237] The denoised archival image submodule is used to filter the original archival image to obtain a denoised archival image.

[0238] The contrast-optimized archive image submodule is used to enhance the contrast of denoised archive images to obtain contrast-optimized archive images.

[0239] The document image correction submodule is used to perform adaptive correction of shadows and wrinkles on contrast-optimized document images to obtain corrected document images.

[0240] The target image data submodule is used to perform quality inspection on the calibration archive images and use the calibration archive images that pass the quality inspection as the target image data.

[0241] Furthermore, the feature extraction module for official seal image units includes:

[0242] The suspected official seal candidate region submodule is used to extract suspected official seal candidate regions from the target image data.

[0243] The official seal boundary edge image submodule is used to detect the official seal boundary edge of suspected candidate regions and obtain the official seal boundary edge image.

[0244] The official seal image unit submodule is used to separate the official seal image unit from the official seal boundary edge image;

[0245] The text extraction submodule is used to extract the archival text semantic vector and the text surrounding the official seal from the target image data.

[0246] Furthermore, the official seal image unit feature processing module includes:

[0247] The normalized grayscale seal image submodule is used to perform grayscale conversion and pixel normalization on the seal image units to obtain a normalized grayscale seal image;

[0248] The feature vector submodule is used to extract the shape feature vector, texture feature vector, and text content feature vector of the normalized grayscale official seal image;

[0249] The initial official seal feature dataset submodule is used to concatenate the shape feature vector, texture feature vector, and text content feature vector to obtain the initial official seal feature dataset.

[0250] The image enhancement submodule is used to perform feature comparison using the shape feature vector and texture feature vector in the initial official seal feature dataset. If the feature comparison result does not meet the preset feature comparison conditions, the image unit of the official seal is enhanced.

[0251] The first jump rotor module is used to jump to execute the steps of performing grayscale conversion and pixel normalization on the official seal image unit based on the image enhancement of the official seal image unit, so as to obtain a normalized grayscale official seal image, until the feature comparison result meets the preset feature comparison conditions, and then outputs the adjusted feature data.

[0252] The text recognition result submodule is used to input the enhanced official seal image unit after adjusting the feature data association into the preset text recognition model and output the text recognition result;

[0253] The recognition accuracy submodule is used to compare the text recognition results with the surrounding text of the official seal to obtain the recognition accuracy.

[0254] The calibration seal feature dataset submodule is used to integrate the adjusted feature data with the text content feature vector to obtain the calibration seal feature dataset when the recognition accuracy is greater than or equal to a preset accuracy threshold.

[0255] Furthermore, the official seal image unit feature processing module also includes:

[0256] The model parameter submodule is used to adjust the model parameters of the preset text recognition model when the recognition accuracy is less than the preset accuracy threshold.

[0257] The second jump rotor module is used to jump to execute the steps of inputting the image-enhanced official seal image unit associated with the adjusted feature data into the preset text recognition model and outputting the text recognition result, based on the adjusted model parameters, until the recognition accuracy is greater than or equal to the preset accuracy threshold, and outputting the calibrated official seal feature dataset.

[0258] Furthermore, the module for judging the compliance and integrity of official seals includes:

[0259] The official seal classification and identification data submodule is used to classify and identify the target archive by using a calibrated official seal feature dataset, the text surrounding the official seal and the semantic vector of the archive text.

[0260] The official seal compliance result submodule is used to make compliance judgments by using a calibrated official seal feature dataset and a preset official seal feature specification database to obtain the official seal compliance result of the target file;

[0261] The integrity verification result submodule is used to perform integrity verification on the official seal image unit and obtain the integrity verification result.

[0262] The structured classification archive data submodule is used to import official seal classification identification data, calibration official seal feature dataset, official seal compliance results, and integrity verification results into a preset structured data template to obtain structured classification archive data.

[0263] Furthermore, the official seal classification identifier data submodule includes:

[0264] The surrounding text fragment unit is used to extract multiple surrounding text fragments from the surrounding text of the official seal.

[0265] The text similarity unit is used to calculate the text similarity between the text content feature vector in the calibration seal feature dataset and multiple surrounding text fragments.

[0266] The target surrounding content text unit is used to select a preset number of surrounding content text fragments as target surrounding content text based on text similarity from highest to lowest.

[0267] The word semantic vector unit is used to take the surrounding text of each target as input to a preset word vector model and output multiple word semantic vectors.

[0268] The semantic matching degree unit is used to calculate the semantic matching degree between the semantic vector of the archive text and the semantic vector of each word.

[0269] The retention unit is used to retain the semantic vectors of words associated with a semantic matching degree greater than or equal to a preset matching degree threshold;

[0270] The correction unit is used to correct the semantic vectors of words whose semantic matching degree is less than a preset matching degree threshold based on the archive domain dictionary and the calibration seal feature dataset;

[0271] The semantic classification result unit is used to integrate the corrected word semantic vectors with the retained word semantic vectors to obtain the semantic classification result;

[0272] The archive category label unit is used to input the archive text semantic vector into the pre-trained archive classification model and output the corresponding archive category label;

[0273] The label mapping unit is used to map the semantic classification results with the archive category labels to obtain the official seal classification identifier data of the target archive.

[0274] Furthermore, the official seal compliance result submodule includes:

[0275] The feature value unit is used to compare the features of the calibrated official seal feature dataset with the preset official seal feature specification database to obtain feature values ​​of multiple different feature types, including shape features, texture features and text content features.

[0276] The compliance unit is used to mark the calibration seal feature dataset as compliant when all types of feature values ​​meet the associated preset feature standard conditions, determine the target file as compliant, and generate the seal compliance result.

[0277] The unit to be evaluated is used to mark the calibration seal feature dataset as to be evaluated if any type of feature value does not meet the associated preset feature standard conditions;

[0278] The difference unit is used to calculate the difference of multiple different feature types using the feature values ​​in the calibration seal feature dataset to be evaluated;

[0279] The comprehensive suspicion scoring unit is used to perform weighted calculations using preset coupling factors of various degrees of difference and correlation to obtain a comprehensive suspicion score for the target file;

[0280] The processing unit is used to determine the official seal compliance result of the target file based on the preset risk threshold range to which the comprehensive doubt score belongs. The official seal compliance result includes compliance level, compliance pending verification level, compliance defect level, and compliance violation level.

[0281] Furthermore, the official seal suspiciousness analysis, assessment, and early warning module includes:

[0282] The structured classification archive data cleaning submodule is used to clean structured classification archive data.

[0283] The consistency verification submodule is used to perform consistency verification between structured classification archive data and cleaned structured classification archive data.

[0284] The target key submodule is used to generate a target key by combining the compliance and integrity verification results of the official seals in the cleaned structured classified archive data when the consistency verification result is passed.

[0285] The archive risk level submodule is used to retrieve a preset archive risk key-value pair table using the target key and match the corresponding archive risk level of the target archive.

[0286] The Archive Audit Report submodule is used to integrate cleaned, structured, categorized archive data with archive risk levels to generate an archive audit report for the target archive.

[0287] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0288] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.

[0289] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0290] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0291] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0292] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for processing archival information based on official seal recognition, characterized in that, include: The original image of the target file is acquired and preprocessed to obtain the target image data; The target image data is subjected to official seal extraction to obtain official seal image units, document text semantic vectors and surrounding text of the official seal; Feature analysis is performed using the official seal image unit, the archive text semantic vector, and the text surrounding the official seal to obtain a calibration official seal feature dataset. The compliance and integrity verification is performed using the calibration seal feature dataset, the seal image unit, the seal surrounding text, and the archive text semantic vector to obtain structured classified archive data; Based on the structured classification archive data, an archive audit report for the target archive is generated.

2. The method for processing archival information based on official seal recognition according to claim 1, characterized in that, The process of acquiring the original image of the target file and preprocessing it to obtain target image data includes: Obtain the original file image of the target file; The original archive image is filtered to obtain a denoised archive image; The denoised archive image is then contrast-enhanced to obtain a contrast-optimized archive image; The contrast-optimized archive image is subjected to adaptive correction for shadows and wrinkles to obtain a corrected archive image; The quality of the corrected archive image is checked, and the corrected archive image that passes the quality check is used as the target image data.

3. The method for processing archival information based on official seal recognition according to claim 1, characterized in that, The step of extracting the official seal from the target image data to obtain an official seal image unit includes: Extract the candidate regions for suspected official seals from the target image data; Perform seal boundary edge detection on the suspected candidate seal area to obtain the seal boundary edge image; Separate the official seal image units from the boundary edge image of the official seal; Extract the archival text semantic vector and the surrounding text of the official seal from the target image data.

4. The method for processing archival information based on official seal recognition according to claim 1, characterized in that, The step involves using the official seal image unit, the semantic vector of the archive text, and the surrounding text of the official seal to perform feature analysis, resulting in a calibrated official seal feature dataset, including: The official seal image unit is subjected to grayscale conversion and pixel normalization to obtain a normalized grayscale official seal image; Extract the shape feature vector, texture feature vector, and text content feature vector of the normalized grayscale official seal image; The shape feature vector, the texture feature vector, and the text content feature vector are concatenated to obtain the initial official seal feature dataset. The shape feature vector and texture feature vector in the initial official seal feature dataset are used for feature comparison. If the feature comparison result does not meet the preset feature comparison conditions, the official seal image unit is enhanced. Based on the enhanced official seal image unit, the process jumps to the step of performing grayscale conversion and pixel normalization on the official seal image unit to obtain a normalized grayscale official seal image, until the feature comparison result meets the preset feature comparison condition, and then outputs the adjusted feature data. The enhanced official seal image unit, after adjusting the feature data association, is input into a preset text recognition model, and the text recognition result is output. The text recognition result is compared with the surrounding text of the official seal to obtain the recognition accuracy. When the recognition accuracy is greater than or equal to a preset accuracy threshold, the adjusted feature data is integrated with the text content feature vector to obtain a calibration seal feature dataset.

5. The method for processing archival information based on official seal recognition according to claim 4, characterized in that, Also includes: If the recognition accuracy is less than the preset accuracy threshold, the model parameters of the preset text recognition model are adjusted. Based on the adjusted model parameters, the process jumps to the step of inputting the enhanced official seal image unit associated with the adjusted feature data into the preset text recognition model and outputting the text recognition result, until the recognition accuracy is greater than or equal to the preset accuracy threshold, and then outputs the calibrated official seal feature dataset.

6. The method for processing archival information based on official seal recognition according to claim 1, characterized in that, The compliance and integrity verification is performed using the calibration seal feature dataset, the seal image unit, the text surrounding the seal, and the semantic vector of the archive text to obtain structured classified archive data, including: The official seal feature dataset, the text surrounding the official seal, and the semantic vector of the archive text are used for classification and identification to obtain the official seal classification identification data of the target archive; The compliance determination is performed using the calibration seal feature dataset and the preset seal feature specification database to obtain the seal compliance result of the target file; The integrity of the official seal image unit is verified to obtain the integrity verification result; The official seal classification identifier data, the calibration official seal feature dataset, the official seal compliance result, and the integrity verification result are imported into a preset structured data template to obtain structured classification archive data.

7. The method for processing archival information based on official seal recognition according to claim 6, characterized in that, The step of classifying and labeling the target archive's official seal by using the calibration seal feature dataset, the surrounding text of the official seal, and the semantic vector of the archive text, includes: Extract multiple surrounding text fragments from the text surrounding the official seal; Calculate the text similarity between the text content feature vector in the calibration seal feature dataset and multiple surrounding text fragments; Select a preset number of surrounding text fragments as target surrounding text based on text similarity from highest to lowest. The text surrounding each target is used as input to a preset word vector model, and multiple word semantic vectors are output. Calculate the semantic matching degree between the semantic vector of the archive text and the semantic vector of each word; The semantic vectors of words associated with semantic matching degrees greater than or equal to a preset matching degree threshold are retained. Based on the archival domain dictionary and the calibration seal feature dataset, the semantic vectors of words associated with semantic matching degrees less than the preset matching degree threshold are corrected; By integrating the corrected word semantic vectors with the retained word semantic vectors, the semantic classification results are obtained; The archival text semantic vectors are input into a pre-trained archival classification model, which outputs the corresponding archival category labels. The semantic classification results are used to map the archive category labels to obtain the official seal classification identifier data of the target archive.

8. The method for processing archival information based on official seal recognition according to claim 6, characterized in that, The compliance determination, which uses the calibrated official seal feature dataset and the preset official seal feature specification database to obtain the official seal compliance result of the target file, includes: The calibration seal feature dataset is compared with the preset seal feature specification database to obtain feature values ​​of multiple different feature types, including shape features, texture features and text content features. When all types of feature values ​​meet the associated preset feature standard conditions, the calibration seal feature dataset is marked as compliant, the target file is determined to be compliant, and a seal compliance result is generated. If any type of feature value does not meet the associated preset feature standard condition, the calibration seal feature dataset is marked as to be evaluated. Using the feature values ​​described in the calibration seal feature dataset to be evaluated, the degree of difference of multiple different feature types is calculated; The suspiciousness score of the target file is obtained by weighting the calculation using a preset coupling factor that combines the differences and correlations. Based on the preset risk threshold range to which the comprehensive suspicion score belongs, the official seal compliance result of the target file is determined. The official seal compliance result includes compliance level, compliance pending verification level, compliance defect level, and compliance violation level.

9. The method for processing archival information based on official seal recognition according to any one of claims 1-8, characterized in that, The process of generating an archival audit report for the target archive based on the structured classified archival data includes: The structured classification archive data is cleaned. The structured classification archive data and the cleaned structured classification archive data are used to perform consistency verification. When the consistency verification result is passed, the target key is generated by using the official seal compliance result and the integrity verification result in the cleaned structured classification archive data. The target key is used to retrieve a preset file risk key-value pair table and match the corresponding file risk level of the target file; By integrating the cleaned structured and classified archive data with the archive risk level, an archive audit report for the target archive is generated.

10. A system for processing archival information based on official seal recognition, characterized in that, include: The official seal data acquisition and image preprocessing module is used to acquire the original image of the target document and preprocess it to obtain the target image data; The official seal image unit feature extraction module is used to extract the official seal from the target image data to obtain the official seal image unit, the document text semantic vector, and the text surrounding the official seal. The official seal image unit feature processing module is used to perform feature analysis using the official seal image unit, the archive text semantic vector and the text surrounding the official seal to obtain a calibration official seal feature dataset. The official seal compliance judgment and integrity verification module is used to perform compliance and integrity verification using the calibration official seal feature dataset, the official seal image unit, the text surrounding the official seal and the semantic vector of the archive text, to obtain structured classified archive data; The official seal suspiciousness analysis, assessment and early warning module is used to generate an archive audit report for the target archive based on the structured classified archive data.