Seal image segmentation method and system for Chinese painting and calligraphy seal analysis

By combining object detection and a prompt-guided image segmentation model, this method solves the problem of independent task execution in traditional seal image segmentation methods, achieving accurate segmentation and efficient matching of seal images. It is applicable to fields such as museums, auction houses, and calligraphy and painting research.

CN121353684APending Publication Date: 2026-01-16ZHEJIANG UNIV CITY COLLEGE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511941827.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-22
Publication Date
2026-01-16

AI Technical Summary

Technical Problem

Traditional seal image segmentation methods lack coordinated optimization. Seal detection, segmentation, matching, and reconstruction tasks are executed independently, resulting in the inability to share intermediate information. This makes it difficult to handle seals with blurred edges, irregular shapes, or defects, thus affecting the analysis results.

Method used

The target detection algorithm is used to obtain the seal bounding box, and the prompt-guided image segmentation model is used for feature extraction and segmentation. Multi-scale semantic features and global features are fused together to perform feature matching and super-resolution reconstruction, forming a closed-loop processing of detection-segmentation-matching-reconstruction.

Benefits of technology

It achieves accurate segmentation and matching in cases where the seal boundary is blurred, stuck, irregular in shape, or incomplete, improving the visual effect and processing efficiency of seal images and supporting automated analysis without human intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121353684A_ABST
    Figure CN121353684A_ABST
Patent Text Reader

Abstract

The invention discloses a seal image segmentation method and system for Chinese painting and calligraphy seal analysis. According to the seal image segmentation method, the seal detection result is innovatively used as prompt guide input, and the prompt guide type image segmentation model is guided to perform mask-level accurate segmentation on the basis of the detection result, so that compared with a seal image segmentation method in which mask generation is performed through direct cutting of a detection frame or simple form hypothesis; the method can effectively deal with complex conditions such as fuzzy and sticky seal boundaries, irregular shapes or incompleteness and the like, and provides reliable input for subsequent matching and image reconstruction. Besides, the multi-scale semantic features extracted by the image segmentation model and the global features extracted in the feature extraction stage are fused, feature fusion with consistent semantics is realized, matching errors caused by inconsistent feature spaces are avoided, independence between two subtasks of seal segmentation and feature extraction is broken, and the accuracy of seal segmentation and feature extraction is improved. And feature-level collaborative optimization is realized, and the matching robustness of the seal under complex conditions is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to a method and system for segmenting seal images for the analysis of Chinese calligraphy and painting seals. Background Technology

[0002] In recent years, deep learning technology has developed rapidly in the field of computer vision, and the automation level of tasks such as object detection, image segmentation, and image matching has been continuously improved, providing a new technical path for intelligent seal analysis. At the same time, the Segment Anything Model (SAM) and its improved version SAM2 have provided powerful capabilities for the automatic segmentation of arbitrary objects in images. In particular, SAM2 performs exceptionally well in multi-scale fusion, cue adaptability, and interactive performance, and has the potential to integrate traditional image processing tasks into an end-to-end workflow.

[0003] Traditional seal image segmentation methods typically break down the seal segmentation task into multiple sub-tasks, such as seal detection, seal image segmentation, feature extraction, seal matching, and seal image reconstruction, which are performed independently. These methods generally suffer from two problems: (1) Each subtask is executed independently, lacking linkage optimization, resulting in the inability to share intermediate information and lack of semantic collaboration between tasks, which affects the overall analysis effect.

[0004] (2) In the stage of seal image segmentation, the detection box is often directly cropped or a mask is generated by simple morphological assumptions. There is a lack of real boundary modeling, making it difficult to handle seals with blurred edges, irregular shapes or incomplete parts, which seriously affects the quality of subsequent matching. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a method and system for analyzing Chinese calligraphy and painting seals.

[0006] To solve the above-mentioned technical problems, the present invention adopts the following technical solution: This invention provides a seal image segmentation method for analyzing Chinese calligraphy and painting seals, comprising the following steps: using an object detection algorithm to detect seals in Chinese calligraphy and painting images, obtaining the bounding box coordinates of each seal; using the seal bounding box coordinates as a cue, using a cue-guided image segmentation model to extract features and segment the Chinese calligraphy and painting images, obtaining multi-scale semantic features of the Chinese calligraphy and painting images and all seal images in the Chinese calligraphy and painting images; extracting features from the seal images output by the image segmentation model, obtaining global feature vectors of the seal images; fusing the multi-scale semantic features extracted by the image segmentation model with the global features of the seal images, obtaining fused features; using the fused features to perform feature matching with feature vectors in a standard seal database; and performing super-resolution reconstruction on the successfully matched seal images to obtain the final seal.

[0007] Another aspect of this invention provides a seal image segmentation system for analyzing Chinese calligraphy and painting seals, comprising: a seal detection module for detecting seals in Chinese calligraphy and painting images using a target detection algorithm to obtain the bounding box coordinates of each seal; a seal segmentation module for using the seal bounding box coordinates as prompts and employing a prompt-guided image segmentation model to extract features and segment the Chinese calligraphy and painting images, obtaining multi-scale semantic features of the Chinese calligraphy and painting images and all seal images in the Chinese calligraphy and painting images; a feature extraction module for extracting features from the seal images output by the image segmentation model to obtain global feature vectors of the seal images; a feature fusion module for fusing the multi-scale semantic features extracted by the image segmentation model with the global features of the seal images to obtain fused features; a seal matching module for performing feature matching using the fused features and feature vectors in a standard seal database; and a seal image reconstruction module for performing super-resolution reconstruction on the successfully matched seal images to obtain the final seal.

[0008] The beneficial technical effects of this invention are as follows: (1) The stamp detection result is used as a prompting input. The prompting-guided image segmentation model performs precise segmentation at the mask level based on the detection result. Compared with stamp image segmentation methods that directly crop the detection box or generate a mask through simple morphological assumptions, it can effectively deal with complex situations such as blurred, stuck, irregular or incomplete stamp boundaries, and provide reliable input for subsequent matching and image reconstruction.

[0009] (2) The multi-scale semantic features extracted by the image segmentation model are fused with the global features extracted in the feature matching stage to achieve semantically consistent feature fusion, avoid matching errors caused by inconsistent feature spaces, and break the independence between the two sub-tasks of seal segmentation and feature extraction, realize feature-level collaborative optimization, and effectively improve the matching robustness of seals under complex conditions.

[0010] (3) The seal detection result is used as a prompt to guide the image segmentation model to perform accurate segmentation, so as to achieve seamless connection between seal detection and image segmentation and avoid repeated calculation and boundary error propagation. The multi-scale semantic features extracted by the image segmentation model are not only used for segmentation mask generation, but also used to fuse with the global features extracted in the feature matching stage to achieve semantically consistent feature fusion, so as to avoid matching errors caused by inconsistent feature space. The high-quality seal image obtained by image segmentation is used as the accurate input area, and the seal super-resolution reconstruction is performed in combination with the standard image after matching to generate a clearer and more recognizable seal image, improve the image visual effect, and form a closed-loop processing chain of "detection-segmentation-matching-reconstruction". This realizes the collaborative optimization between sub-tasks, effectively reduces information redundancy, and improves end-to-end processing efficiency and accuracy.

[0011] (4) It supports automatic seal image segmentation process without human intervention, which has a high degree of automation and practical application value, and is particularly suitable for museums, auction houses, calligraphy and painting research and digital humanities. Attached Figure Description

[0012] Figure 1 This is a flowchart illustrating the seal image segmentation method for analyzing Chinese calligraphy and painting seals in Embodiment 1 of the present invention. Figure 2 This is a flowchart illustrating the seal image segmentation method for analyzing Chinese calligraphy and painting seals in Embodiment 2 of the present invention. Figure 3 This is a flowchart illustrating the seal image segmentation method for analyzing Chinese calligraphy and painting seals in Embodiment 3 of the present invention. Figure 4 This is a flowchart illustrating the seal image segmentation method for analyzing Chinese calligraphy and painting seals in Embodiment 4 of the present invention. Figure 5 This is a flowchart illustrating the seal image segmentation method for analyzing Chinese calligraphy and painting seals in Embodiment 5 of the present invention. Figure 6 This is a schematic diagram of the seal image segmentation system for analyzing Chinese calligraphy and painting seals according to the present invention. Detailed Implementation

[0013] To enable those skilled in the art to more clearly understand the purpose, technical solution, and advantages of the present invention, the present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0014] This invention provides a seal image segmentation method and system for analyzing Chinese calligraphy and painting seals. The seal image segmentation method includes the following steps: using a target detection algorithm to detect seals in Chinese calligraphy and painting images, obtaining the bounding box coordinates of each seal; using the seal bounding box coordinates as a cue, using a cue-guided image segmentation model to extract features and segment the Chinese calligraphy and painting images, obtaining multi-scale semantic features of the Chinese calligraphy and painting images and all seal images in the Chinese calligraphy and painting images; extracting features from the seal images output by the image segmentation model to obtain global feature vectors of the seal images; fusing the multi-scale semantic features extracted by the image segmentation model with the global features of the seal images to obtain fused features; using the fused features to perform feature matching with feature vectors in a standard seal database; and performing super-resolution reconstruction on the successfully matched seal images to obtain the final seal.

[0015] This invention innovatively uses the seal detection result as a prompting input. The prompt-guided image segmentation model performs precise segmentation at the mask level based on the detection result. Compared to seal image segmentation methods that directly crop the detection box or generate masks based on simple morphological assumptions, this method effectively handles complex situations such as blurred, adhered, irregular, or incomplete seal boundaries, providing reliable input for subsequent matching and image reconstruction. Furthermore, this invention fuses the multi-scale semantic features extracted by the image segmentation model with the global features extracted during the feature extraction stage, achieving semantically consistent feature fusion. This avoids matching errors caused by inconsistent feature spaces and breaks the independence between the seal segmentation and feature extraction subtasks, achieving feature-level collaborative optimization and effectively improving the robustness of seal matching under complex conditions.

[0016] Example 1: like Figure 1 As shown, in this embodiment of the invention, the seal image segmentation method for analyzing Chinese calligraphy and painting seals includes steps S10 to S60: S10: Use YOLOv10 to detect seals in Chinese calligraphy and painting images and obtain the bounding box coordinates of each seal.

[0017] Object detection is a computer vision task that aims to identify the location and category label of target objects in an image. This step uses an object detection algorithm to detect seals in a Chinese calligraphy and painting image, aiming to identify the seal regions in the image and obtain the bounding box coordinates of each seal. YOLOv10 is an advanced object detection algorithm based on the YOLO series, possessing high accuracy and high real-time performance, suitable for object detection tasks in complex images. This embodiment selects YOLOv10 for seal detection in Chinese calligraphy and painting images. Compared with traditional seal detection algorithms, it has better performance in detection accuracy and boundary adaptation, and can handle adhered seals and complex backgrounds.

[0018] To enhance YOLOv10's ability to recognize seal regions, this embodiment of the invention requires training the YOLOv10 model using a collected seal dataset (hereinafter referred to as the dataset) before using YOLOv10 to detect seals in Chinese calligraphy and painting images. The specific training process is as follows: Step 1: Construct the dataset and divide it into training, validation, and test sets. In this embodiment of the invention, web crawlers are used to collect publicly available data. The dataset is constructed by collecting images of calligraphy and paintings from the digital public collections of museums. The dataset contains 4,000 images of calligraphy and paintings. For each image, irrelevant factors such as borders are first manually removed. Then, the seals are manually annotated in detail, including the frame coordinates and mask of each seal. A total of 13,002 seals are annotated. The dataset is then divided into training, validation, and test sets in an 8:1:1 ratio.

[0019] The second step is to preprocess the training data. Preprocessing includes standardizing the image size of the training set to 640×640 pixels, and performing enhancement operations such as random scaling, flipping, cropping, and color perturbation on the images in the training set to improve the model's generalization ability. In addition, the images and annotations need to be processed simultaneously to ensure the consistency of the augmented bounding box coordinates.

[0020] Step 3: Initialize the YOLOv10 model structure. Specifically, load the YOLOv10-S model structure released by the official YOLOv10 as the base model, and adjust the number of categories to 1 (stamp class) as needed.

[0021] Step 4: Set training parameters. This embodiment uses the following configuration for training: Total training epochs: 50; Optimizer: AdamW; Initial learning rate: 0.01; Batch size: 32; Loss function: YOLOv10 default configuration; Data augmentation strategy: Mosaic, MixUp, HSV perturbation and other combined augmentation strategies.

[0022] Step 5: Use the training script provided by the official YOLOv10 software to train on a single RTX 4090 graphics card. The training process automatically records the loss curve and mAP metric, and saves the optimal weights on the validation set.

[0023] Step 6: After training is complete, save the optimal model weight file. The trained YOLOv10 model enhances the ability to recognize stamp regions.

[0024] Step S10, which involves using YOLOv10 to detect seals in Chinese calligraphy and painting images and obtaining the bounding box coordinates of each seal, includes the following steps: S11: Preprocess the Chinese calligraphy and painting image input by the user. Specifically, preprocess the Chinese calligraphy and painting image containing seals input by the user, adjust the image format to RGB color image, and adjust the image size to 640×640.

[0025] S12: Input the preprocessed Chinese calligraphy and painting images into the trained YOLOv10 model for seal detection, and output the bounding box coordinates of each seal. The output seal bounding box coordinates serve as preliminary seal candidate regions and SAM2 hints for subsequent segmentation and analysis.

[0026] The YOLOv10 model comprises the Stem module, a spatial-channel decoupled downsampling module, multiple rank-guided stages, a partially self-attention (PSA) module, and a head. The Stem module consists of lightweight convolutional layers for rapidly extracting initial low-level image features while maintaining low computational overhead. Each stage of YOLOv10 consists of multiple basic building blocks employing a rank-guided compact inverted block (CIB) structure, achieving efficient mixing of spatial and channel features through depthwise and one-dimensional convolutions. Furthermore, different stages utilize adaptive block structure designs based on their feature redundancy to achieve a balance between efficiency and performance. The partially self-attention module, inserted after the last stage, divides the feature channels into two parts: one part enhances global modeling capabilities through a multi-head self-attention mechanism, while the other part maintains an efficient convolutional structure, thereby introducing long-range dependency modeling capabilities with low computational cost. The YOLOv10 header includes a lightweight classification header and a regression header. The lightweight classification header consists of two depthwise separable convolutions and a 1×1 convolution, which can significantly reduce computational overhead. The regression header is used to regress the location information of the bounding boxes.

[0027] After the preprocessed Chinese calligraphy and painting images are input into the trained YOLOv10 model, they first undergo initial feature extraction through the Stem module. The feature map output by the Stem module is then fed into the spatial-channel decoupled downsampling module. This module first adjusts the number of channels using a 1×1 point convolution, and then performs spatial downsampling using a 3×3 depthwise convolution, effectively reducing computation and improving information retention. The downsampled feature map is then sequentially fed into multiple stages and a portion of the self-attention module for multi-stage feature extraction. The multi-scale feature map output by the partial self-attention module is then fed into the YOLOv10 head, which outputs the class probability, bounding box coordinates, and target confidence at each detection scale. Finally, the YOLOv10 model outputs the bounding box coordinates of each seal as preliminary seal candidate regions and SAM2 hints, providing structural priors for subsequent accurate segmentation and achieving an end-to-end segmentation mechanism driven by automatic hints.

[0028] S20: Using the coordinates of the seal bounding box as a cue, use SAM2 to extract features and segment the Chinese calligraphy and painting images to obtain multi-scale semantic features of the Chinese calligraphy and painting images and all seal images in the Chinese calligraphy and painting images.

[0029] This step uses the bounding box coordinates of the seals output by the YOLOv10 model as a cue input to guide the SegmentAnything Model v2 (SAM2) for cue-driven segmentation. The SAM2 model simultaneously receives the original Chinese calligraphy and painting image and the bounding box cue. Utilizing its powerful cross-scale attention mechanism and mask decoding capability, it outputs a precise segmentation mask for each seal, capturing edge details and the actual shape of the seal.

[0030] To enhance SAM2's ability to segment stamp regions, the SAM2 model needs to be fine-tuned. The specific fine-tuning process is as follows: Step 1: Prepare the segmentation training dataset. Based on the seal dataset constructed above, a segmentation training dataset is built using the mask corresponding to each seal. Each sample in the segmentation training dataset includes: the original Chinese calligraphy and painting image, the seal bounding box output by YOLOv10 (as a cue box), and the corresponding precise seal mask. This segmentation training dataset is used to fine-tune the SAM2 model.

[0031] The second step is to construct the SAM2 model structure. The SAM2 model consists of three parts: an image encoder, a cue encoder, and a mask decoder. The image encoder uses an efficient Hiera Transformer structure to extract multi-scale semantic features of the image. The cue encoder is responsible for converting bounding box cues into high-dimensional embedding vectors. The mask decoder predicts the final segmentation mask through multi-layer cross-attention and dynamic mask generation modules.

[0032] Step 3: Fine-tune the configuration. In this embodiment, the following parameters are used to fine-tune the SAM2 model: The total number of training rounds is 50; the optimizer is AdamW; the initial learning rate is set to 1e-5, using a linear warmup plus cosine decay strategy; the weight decay is set to 1e-7; the batch size is 8; the input image size is uniformly scaled to 512×512; the loss function is a weighted combination of masked binary cross-entropy and Dice Loss.

[0033] The fourth step involves using the training script provided by the official SAM2 team to fine-tune the training on a single RTX 4090 graphics card. During training, the image encoder parameters are frozen, and only the cue encoder and mask decoder are trained. The loss curve and mAP metric are automatically recorded during training, and the optimal weights are saved on the validation set.

[0034] Fifth, after completing the fine-tuning training, save the optimal model weight file. The fine-tuned SAM2 model enhances its ability to segment the stamp region.

[0035] Step S20, which uses the seal bounding box coordinates as a prompt, employs SAM2 to extract features and segment the Chinese calligraphy and painting image, obtaining multi-scale semantic features of the Chinese calligraphy and painting image and all seal images within the image, includes the following steps: S21: Input the Chinese calligraphy and painting image and the coordinates of the seal bounding boxes into the trained SAM2 model. Specifically, input a Chinese calligraphy and painting image containing seals provided by the user and the coordinates of the seal bounding boxes detected by YOLOv10 into the trained SAM2 model. Each seal bounding box represents a candidate region of seals to be segmented.

[0036] S22: The SAM2 model is used as an image encoder to extract features from Chinese calligraphy and painting images and output multi-scale semantic features. The SAM2 model uses Hiera Transformer to build its multi-scale image encoder. It first divides the image into several non-overlapping patches and extracts semantic features from local to global through a multi-level attention structure. Finally, it outputs semantic feature representations containing multiple resolution scales.

[0037] S23: The cue encoder of the SAM2 model encodes the coordinates of the stamp bounding boxes into high-dimensional cue embedding vectors. The cue encoder receives the coordinate information of each stamp bounding box and encodes it into a high-dimensional cue embedding vector, which will guide the segmentation attention to the target region in subsequent stages.

[0038] S24: The high-dimensional cue embedding vector is cross-fused with the image features output by the image encoder. The mask decoder of the SAM2 model uses a cross-attention mechanism, taking the high-dimensional cue embedding vector as the query and the image features output by the image encoder as the key and value, to perform multi-head attention calculation, thereby focusing on the stamp area within the cue box and enhancing the expression of target boundary features.

[0039] S25: Based on the fused features, the SAM2 model's mask decoder generates a segmentation mask corresponding to each cue box. The fused features are passed through a series of feedforward layers and a multi-scale fusion module in the mask decoder to finally generate a high-precision segmentation mask for each cue box. This segmentation mask is a single-channel binary image of the same size as the input image.

[0040] S26: Integrate and output the segmentation masks corresponding to all prompt boxes, and output the multi-scale semantic features extracted by the image encoder. Integrate and output the segmentation masks corresponding to all prompt boxes to obtain the segmentation results for all stamps in the entire image. Furthermore, while generating the segmentation masks, the multi-scale semantic features extracted by the image encoder will be retained, and this multi-scale semantic feature output can be used for subsequent feature fusion.

[0041] Through step S20, a prompt-enhanced SAM2 segmentation structure based on object detection prompts is adopted in the field of traditional Chinese painting to replace the traditional boundary speculation or full image segmentation method for seal segmentation, making the seal outline extraction more accurate and improving the subsequent matching and reconstruction effect.

[0042] S30: Use DINOv2 to extract features from the stamp image output by SAM2 to obtain the global feature vector of the stamp image.

[0043] DINOv2 is a self-supervised contrastive learning framework for generating high-quality image feature representations. This step uses the stamp image output from the SAM2 model as the main input, feeding it into the pre-trained DINOv2 model. DINOv2 extracts deep semantic features of the stamp that remain consistent despite deformation, noise, and style changes by constructing positive and negative sample pairs.

[0044] To enhance DINOv2's ability to extract seal features, the DINOv2 model needs to be trained first. The specific training process is as follows: The first step is to construct a training set of seal images. Using the dataset constructed above, the images are uniformly cropped to contain only the seal area, and uniformly resized to 224×224, then saved in RGB format.

[0045] The second step involves image enhancement to generate multiple views. For each stamp image, a random enhancement strategy is used to generate multiple image views at different scales, including: Two global views (cropping scale of 0.4–1.0); Six partial views (cropping scale 0.05–0.4); At the same time, the view is enhanced with color perturbation, blurring, highlights, and inversion to improve the model's style robustness.

[0046] The third step is to construct the DINOv2 training model. The training model contains a pair of structurally identical ViT encoders, serving as the student network and the teacher network respectively. Both employ the ViT-Small architecture of the Vision Transformer, extracting global representation features from the input image. The parameters of the teacher network do not participate in backpropagation but are updated synchronously from the student network via exponential moving average.

[0047] The fourth step is to perform self-supervised contrastive learning. All enhanced image views are input into two networks, and the feature similarity between each view is calculated. The self-supervised contrastive loss function built into DINOv2 is used to optimize the output of the student network, so that different views from the same image are closer in the feature space, while different images are farther apart.

[0048] Fifth, configure the training parameters as follows: Network architecture: ViT-Small; Input size: 224×224; Batch size: 256; Optimizer: AdamW; Initial learning rate: 1e-4, using Cosine decay; Weight decay: Initially 0.04, eventually increasing to 0.4; Number of training epochs: 300; Teacher network parameter updates use EMA strategy, with an initial momentum value of 0.996.

[0049] Step 6: After the model training is complete, save the weights of the student network for subsequent feature extraction tasks. The trained DINOv2 model enhances its ability to extract seal features.

[0050] Step S30, which involves using DINOv2 to extract features from the stamp image output by SAM2 to obtain the global feature vector of the stamp image, includes the following steps: S31: Prepare the stamp image from which features are to be extracted. The stamp image from which features are to be extracted is the stamp image obtained by image segmentation using the SAM2 model. The stamp images are uniformly scaled to 224×224 pixels and converted to RGB format.

[0051] S32: Perform standard preprocessing on the stamp image from which features are to be extracted. Standard preprocessing includes image normalization and standardization to ensure that the image input during the feature extraction stage of the DINOv2 model is consistent with the image input during the training stage.

[0052] S33: Input the preprocessed seal image into the trained DINOv2 model for feature extraction to obtain the global feature vector of the current seal image. DINOv2 uses the Vision Transformer architecture, and the output of the last layer contains a [CLS] token, representing the global semantic information of the entire image. Extract this [CLS] token output as the global feature vector of the current seal image.

[0053] S34: Save and record the global feature vector of the current stamp image. Save the extracted [CLS] token vector as a fixed-length real number array with dimension 384 for use in subsequent tasks.

[0054] S40: The multi-scale semantic features extracted by SAM2 are fused with the global features extracted by DINOv2 to obtain fused features.

[0055] This step fuses the multi-scale semantic features (f1) extracted by the SAM2 image encoder with the global features (f2) extracted by DINOv2 to obtain fused features. First, a 1x1 convolution is used to map the two features to the same channel dimension, and then cross-attention is used to fuse the features, where f1 is the query and f2 is the key and value. The unified feature (fused feature) obtained after fusion is used for downstream stamp matching.

[0056] Through step S40, a dual-stream fusion structure of "segmentation encoding semantics + matching embedding semantics" is introduced for the first time in the field of traditional Chinese painting. This fully utilizes mask structure information and full-image context semantics to improve the discriminative power and robustness of seal recognition.

[0057] S50: Use fused features to perform feature matching with feature vectors in the standard seal database.

[0058] This step is the seal matching step. Seal matching refers to performing feature matching between the detected seal and the image in the standard seal database to determine ownership and authenticity. Specifically, the fused feature is input into the comparison module and matched with the feature vectors in the standard seal database for similarity. This involves comparing the similarity of feature vectors by calculating the cosine distance between them, and using a locality-sensitive hashing algorithm for fast nearest neighbor search. The algorithm searches for the vector in the standard seal database that is most similar to the feature vector (fused feature) of the input query image. If the similarity of the most similar vector exceeds a matching threshold, the match is successful; otherwise, the match fails.

[0059] S60: Use MambaIRv2 to perform super-resolution reconstruction on the successfully matched seal images to obtain the final seal.

[0060] Super-resolution reconstruction is the process of restoring a low-resolution image to a high-resolution image, aiming to recover image details and enhance edge quality. It is commonly used in fields such as cultural relic image restoration and medical imaging. MambaIRv2 is an image restoration network based on Mamba that supports image magnification, deblurring, and other functions. This step uses MambaIRv2 to perform super-resolution reconstruction on the successfully matched seal images. The successfully matched seal images are input into the MambaIRv2 model, and the output is a high-quality, high-resolution final seal image, which is convenient for expert identification and visual display.

[0061] Step S60, which involves using MambaIRv2 to perform super-resolution reconstruction on the successfully matched seal images to obtain the final seal, includes the following steps: S61: Preprocess the successfully matched seal images. Specifically, adjust the successfully matched seal images to RGB format and uniformly adjust the size to the preset size of 256×256.

[0062] S62: Perform shallow feature extraction on the preprocessed stamp image. The preprocessed stamp image is input into the MambaIRv2 model. First, a lightweight convolutional layer is used to perform shallow feature extraction on the image, capturing the local texture information of the image and providing initial feature representations for subsequent deep modeling.

[0063] S63: Input shallow features into the MambaIRv2 backbone network and output multi-layer intermediate features.

[0064] The MambaIRv2 backbone network consists of multiple stacked Mamba blocks, each layer containing non-linear activation and normalization operations to enhance the model's expressive power. MambaIRv2 uses improved Mamba blocks for deep information modeling, unlike traditional convolutional or Transformer networks. Its core includes: Non-causal modeling structure: breaks the causal constraints in the original Mamba, allowing each token to consider both preceding and following contextual information.

[0065] Multi-branch structure: Each Mamba block contains multiple state streams and dynamic parameter update paths to capture image dependencies at different scales; State mixing with gating mechanism: selectively fuse local and global features through gating to improve modeling accuracy and control computational load.

[0066] S64: Local enhancement of the multi-layer intermediate features output. To address the weak edge and texture recovery capabilities of non-convolutional structures, MambaIRv2 integrates a shallow attention module to enhance the perception and recovery of areas such as stamp edges, text structures, and detailed lines. The intermediate features output from each Mamba block are input into the shallow attention module to enhance local details, and the enhanced intermediate features are output.

[0067] S65: Fuse the multi-layer intermediate features after local enhancement and generate a reconstructed image. At the end of the network, the intermediate features output from the multi-layer Mamba blocks are fused at multiple scales, and through a series of upsampling and convolution operations, the image is gradually restored to the original image size to generate a high-quality output image, resulting in the final stamp.

[0068] Through step S60, a Mamba-based model is used for the first time in the field of traditional Chinese painting. The reconstruction quality is improved by using high-quality boundaries and structural priors, thus making up for the problems of blurred and damaged seals in traditional images.

[0069] The beneficial technical effects of the embodiments of the present invention are as follows: (1) Using the stamp detection box of YOLOv10 as a prompting input, SAM2 is guided to perform precise segmentation at the mask level based on the detection results. Compared with the stamp image segmentation method that directly crops the detection box or generates a mask through simple morphological assumptions, it can effectively deal with complex situations such as blurred, stuck, irregular or incomplete stamp boundaries, and provide reliable input for subsequent matching and image reconstruction.

[0070] (2) The multi-scale semantic features extracted by the SAM2 encoder are fused with the global features extracted by DINOv2 to enhance the feature representation of DINOv2, realize semantically consistent feature fusion, avoid matching errors caused by inconsistent feature spaces, and break the independence between the two sub-tasks of seal segmentation and feature extraction, realize feature-level collaborative optimization, and effectively improve the matching robustness of seals under complex conditions.

[0071] (3) Using the stamp detection bounding box of YOLOv10 as a prompting input, SAM2 is guided to perform accurate segmentation, achieving seamless connection between stamp detection and image segmentation, avoiding repeated calculations and boundary error propagation; the multi-scale semantic features extracted by SAM2 are not only used for segmentation mask generation, but also used to fuse with the global features extracted by DINOv2 to achieve semantically consistent feature fusion, so as to avoid matching errors caused by inconsistent feature spaces; the high-quality segmentation mask obtained by stamp segmentation is used as the accurate input area, and stamp super-resolution reconstruction is performed in combination with the matched standard image to generate clearer and more recognizable stamp images, improve the image visual effect, and form a closed-loop processing chain of "detection-segmentation-matching-reconstruction", realizing the collaborative optimization between sub-tasks, effectively reducing information redundancy, and improving end-to-end processing efficiency and accuracy.

[0072] (4) It supports automatic seal image segmentation process without human intervention, which has a high degree of automation and practical application value, and is particularly suitable for museums, auction houses, calligraphy and painting research and digital humanities.

[0073] Example 2: like Figure 2 As shown, in this embodiment of the invention, the seal image segmentation method for analyzing Chinese calligraphy and painting seals includes steps S110 to S160: S110: Use the trained DETR model to detect seals in Chinese calligraphy and painting images and obtain the bounding box coordinates of each seal; S120: Using the coordinates of the seal bounding box as a cue, SAM2 is used to extract features and segment images of Chinese calligraphy and painting images to obtain multi-scale semantic features of Chinese calligraphy and painting images and all seal images in Chinese calligraphy and painting images. S130: Use DINOv2 to extract features from the stamp image output by SAM2 to obtain the global feature vector of the stamp image; S140: The multi-scale semantic features extracted by SAM2 are fused with the global features extracted by DINOv2 to obtain fused features; S150: Perform feature matching using fused features and feature vectors in a standard seal database; S160: Use MambaIRv2 to perform super-resolution reconstruction on the successfully matched seal images to obtain the final seal.

[0074] Steps S120 to S160 in this embodiment are the same as steps S20 to S60 in Embodiment 1. The difference is that step S110 uses a trained DETR model to detect seals in Chinese calligraphy and painting images to obtain the bounding box coordinates of each seal.

[0075] Before using the trained DETR model to detect seals on Chinese calligraphy and painting images, DETR needs to be trained first. The specific training steps are as follows: Step 1: Constructing the Object Detection Dataset. We used web scraping to collect publicly available data, specifically images of calligraphy and paintings from museums' digital collections, to construct the object detection dataset. The dataset contains 4000 images. For each image, we first manually removed irrelevant elements such as borders, and then manually performed detailed annotations on the seals, including the bounding box coordinates and mask for each seal. A total of 13002 seals were annotated. The dataset was then divided into training, validation, and test sets in an 8:1:1 ratio.

[0076] Step 2: Data augmentation and preprocessing of images in the object detection dataset. During training, data augmentation is performed using methods such as random scaling, flipping, and color perturbation to improve model robustness. Images are normalized to [0,1] and then normalized using standard mean and standard deviation.

[0077] Step 3: Construct the DETR model structure. DETR consists of three main parts: CNN backbone (ResNet-101): used to extract image features; Transformer encoder-decoder architecture: models the global relationships between objects within an image; Prediction Header: Outputs the category and bounding box coordinates of each target.

[0078] The feature map output by the backbone is positionally encoded and then input into the Transformer encoder. The decoder uses a fixed number of learnable object queries and outputs the prediction result corresponding to each query.

[0079] Step 4: Configure training parameters. Use the following training settings: Optimizer: AdamW; Initial learning rate: 1e-4; Learning rate scheduling: Cosine Annealing; Weight decay: 1e-6; Number of training epochs: 50; Batch size: 16; Loss function: DETR uses Hungarian loss, including: Matching loss: based on the Hungarian matching algorithm to establish a one-to-one correspondence between the prediction result and the ground truth; Classification loss: using cross-entropy; Bounding box loss: a weighted combination of L1 loss and GIoU loss.

[0080] Step 5: During each iteration of the training process, the loss is calculated and backpropagated to update the model parameters. The model performance is evaluated on the validation set at regular intervals, using mAP (mean Average Precision) and IoU (Intersection over Union) as metrics.

[0081] Step 6: Save the model parameters and weights after training is complete.

[0082] After training the DETR model, Chinese calligraphy and painting images are input into the trained DETR model for seal detection, and the bounding box coordinates of each seal can be obtained.

[0083] It should be noted that in other embodiments of the present invention, a trained lightweight version of DETR (such as DINO-DETR, RT-DETR) can also be used to detect seals in Chinese calligraphy and painting images. The training process of the model can refer to the training process of DETR, which will not be repeated here. These trained models have stronger contextual understanding capabilities in dense scenes and can effectively detect overlapping seals in complex calligraphy and painting images.

[0084] Example 3: like Figure 3 As shown, in this embodiment of the invention, the seal image segmentation method for analyzing Chinese calligraphy and painting seals includes steps S210 to S260: S210: Use YOLOv10 to detect seals in Chinese calligraphy and painting images and obtain the bounding box coordinates of each seal; S220: Using the coordinates of the seal bounding box as a cue, MedSAM is used to extract features and segment images of Chinese calligraphy and painting images to obtain multi-scale semantic features of Chinese calligraphy and painting images and all seal images in Chinese calligraphy and painting images. S230: Use DINOv2 to extract features from the stamp image output by MedSAM to obtain the global feature vector of the stamp image; S240: The multi-scale semantic features extracted by MedSAM are fused with the global features extracted by DINOv2 to obtain fused features; S250: Use fused features to perform feature matching with feature vectors in a standard seal database; S260: Use MambaIRv2 to perform super-resolution reconstruction on the successfully matched seal images to obtain the final seal.

[0085] Steps S210 to S260 in this embodiment correspond one-to-one with steps S10 to S60 in Embodiment 1. The difference is that step S220 uses MedSAM instead of SAM2 to extract features and segment images of Chinese calligraphy and painting.

[0086] Specifically, in step S220, the coordinates of the seal bounding box output by the YOLOv10 model are used as a cue input to guide MedSAM to perform cue-driven segmentation. The MedSAM model simultaneously receives the original Chinese calligraphy and painting image and the bounding box cue. Utilizing its powerful cross-scale attention mechanism and mask decoding capability, it outputs a precise segmentation mask corresponding to each seal to capture edge details and the actual shape of the seal.

[0087] To enhance the MedSAM model's ability to segment seal regions, it is necessary to fine-tune the MedSAM model. The training steps for the MedSAM model are the same as those for the SAM2 model in Example 1. Furthermore, the steps for using MedSAM to extract features from Chinese calligraphy and painting images and segment seal images are the same as those for using SAM2 to extract features from Chinese calligraphy and painting images and segment images in Example 1, and will not be repeated here.

[0088] It should be noted that in other embodiments of the present invention, traditional semantic segmentation models (such as U-Net) can also be used in conjunction with detection masks for weakly supervised segmentation. These models are all encoder-decoder structures. Using these models, feature extraction and image segmentation can be performed on Chinese calligraphy and painting images to obtain multi-scale semantic features of Chinese calligraphy and painting images as well as all seal images in Chinese calligraphy and painting images.

[0089] Example 4: like Figure 4As shown, in this embodiment of the invention, the seal image segmentation method for analyzing Chinese calligraphy and painting seals includes steps S310 to S360: S310: Use YOLOv10 to detect seals in Chinese calligraphy and painting images and obtain the bounding box coordinates of each seal; S320: Using the coordinates of the seal bounding box as a cue, SAM2 is used to extract features and segment images of Chinese calligraphy and painting images to obtain multi-scale semantic features of Chinese calligraphy and painting images and all seal images in Chinese calligraphy and painting images. S330: Use DINOv2 to extract features from the stamp image output by SAM2 to obtain the global feature vector of the stamp image; S340: The multi-scale semantic features extracted by SAM2 are fused with the global features extracted by DINOv2 through multilayer perceptron fusion or attention-weighted fusion to obtain fused features; S350: Use fused features to perform feature matching with feature vectors in a standard seal database; S360: Use MambaIRv2 to perform super-resolution reconstruction on the successfully matched seal images to obtain the final seal.

[0090] Steps S310 to S360 in this embodiment correspond one-to-one with steps S10 to S60 in Embodiment 1. The difference is that step S340 fuses the multi-scale semantic features extracted by SAM2 with the global features extracted by DINOv2 through multilayer perceptron fusion or attention-weighted fusion to obtain fused features.

[0091] Specifically, when using a multilayer perceptron for fusion, the two features are first concatenated and then flattened, then passed through a multilayer perceptron, and finally their shape is restored. When using attention-weighted fusion, the two features are first averaged, then passed through several fully connected layers to obtain attention weights, and finally the obtained attention weights are used to weight and sum the original two features.

[0092] Example 5: like Figure 5 As shown, in this embodiment of the invention, the seal image segmentation method for analyzing Chinese calligraphy and painting seals includes steps S410 to S460: S410: Use YOLOv10 to detect seals in Chinese calligraphy and painting images and obtain the bounding box coordinates of each seal; S420: Using the coordinates of the seal bounding box as a cue, SAM2 is used to extract features and segment images of Chinese calligraphy and painting images to obtain multi-scale semantic features of Chinese calligraphy and painting images and all seal images in Chinese calligraphy and painting images. S430: Use DINOv2 to extract features from the stamp image output by SAM2 to obtain the global feature vector of the stamp image; S440: The multi-scale semantic features extracted by SAM2 are fused with the global features extracted by DINOv2 to obtain fused features; S450: Use fused features to perform feature matching with feature vectors in a standard seal database; S460: Use SwinIR to perform super-resolution reconstruction on the successfully matched seal images to obtain the final seal.

[0093] Steps S410 to S460 in this embodiment correspond one-to-one with steps S10 to S60 in Embodiment 1. The difference is that step S460 uses SwinIR instead of MambaIRv2 to perform super-resolution reconstruction on the successfully matched seal image to obtain the final seal.

[0094] The steps for super-resolution reconstruction using SwinIR are as follows: S461: Given a low-quality stamp image I0 that has been successfully matched, use a 3×3 convolution H1 to extract shallow features, then the shallow features F0=H1(I0).

[0095] S462: Extract the shallow features and then perform deep feature extraction. The deep feature extraction module consists of a residual SwinTransformer module and a 3×3 convolution. The feature map is divided into local windows. Each Transformer block uses multiple Swin Transformers to implement a local attention mechanism. Then, the deep feature F1 = H2(F0), and H3 represents the deep feature extraction module.

[0096] S463: Fuse shallow features F0 and deep features F1 to reconstruct a high-resolution image I1, then I1=H3(F0+F1), where H3 represents the high-resolution image reconstruction module.

[0097] Example 6: like Figure 6As shown in this embodiment of the invention, the seal image segmentation system for analyzing Chinese calligraphy and painting seals includes: a seal detection module 10, used to detect seals in Chinese calligraphy and painting images using YOLOv10 to obtain the bounding box coordinates of each seal; a seal segmentation module 20, used to extract features and segment images of Chinese calligraphy and painting images using SAM2, using the seal bounding box coordinates as prompts to obtain multi-scale semantic features of Chinese calligraphy and painting images and all seal images in Chinese calligraphy and painting images; a feature extraction module 30, used to extract features from the seal images output by SAM2 using DINOv2 to obtain global feature vectors of seal images; a feature fusion module 40, used to fuse the multi-scale semantic features extracted by SAM2 with the global features extracted by DINOv2 to obtain fused features; a seal matching module 50, used to perform feature matching using the fused features and feature vectors in a standard seal database; and a seal image reconstruction module 60, used to perform super-resolution reconstruction of the successfully matched seal images using MambaIRv2 to obtain the final seal.

[0098] Example 7: See you again Figure 6 In this embodiment of the invention, the seal image segmentation system for analyzing Chinese calligraphy and painting seals includes: a seal detection module 10, used to detect seals in Chinese calligraphy and painting images using the DETR model to obtain the bounding box coordinates of each seal; a seal segmentation module 20, used to extract features and segment images of Chinese calligraphy and painting images using SAM2, with the seal bounding box coordinates as a cue, to obtain multi-scale semantic features of the Chinese calligraphy and painting images and all seal images in the Chinese calligraphy and painting images; a feature extraction module 30, used to extract features from the seal images output by SAM2 using DINOv2 to obtain global feature vectors of the seal images; a feature fusion module 40, used to fuse the multi-scale semantic features extracted by SAM2 with the global features extracted by DINOv2 to obtain fused features; a seal matching module 50, used to perform feature matching using the fused features and feature vectors in a standard seal database; and a seal image reconstruction module 60, used to perform super-resolution reconstruction of the successfully matched seal images using MambaIRv2 to obtain the final seal.

[0099] Example 8: See you again Figure 6In this embodiment of the invention, the seal image segmentation system for analyzing Chinese calligraphy and painting seals includes: a seal detection module 10, used to detect seals in Chinese calligraphy and painting images using YOLOv10 to obtain the bounding box coordinates of each seal; a seal segmentation module 20, used to extract features and segment images of Chinese calligraphy and painting images using MedSAM, with the seal bounding box coordinates as prompts, to obtain multi-scale semantic features of Chinese calligraphy and painting images and all seal images in Chinese calligraphy and painting images; a feature extraction module 30, used to extract features from the seal images output by MedSAM using DINOv2 to obtain global feature vectors of seal images; a feature fusion module 40, used to fuse the multi-scale semantic features extracted by MedSAM with the global features extracted by DINOv2 to obtain fused features; a seal matching module 50, used to perform feature matching using the fused features and feature vectors in a standard seal database; and a seal image reconstruction module 60, used to perform super-resolution reconstruction of the successfully matched seal images using MambaIRv2 to obtain the final seal.

[0100] Example 9: See you again Figure 6 In this embodiment of the invention, the seal image segmentation system for analyzing Chinese calligraphy and painting seals includes: a seal detection module 10, used to detect seals in Chinese calligraphy and painting images using YOLOv10 to obtain the bounding box coordinates of each seal; a seal segmentation module 20, used to extract features and segment images of Chinese calligraphy and painting images using SAM2, using the seal bounding box coordinates as prompts, to obtain multi-scale semantic features of Chinese calligraphy and painting images and all seal images in Chinese calligraphy and painting images; a feature extraction module 30, used to extract features from the seal images output by SAM2 using DINOv2 to obtain global feature vectors of seal images; a feature fusion module 40, used to fuse the multi-scale semantic features extracted by SAM2 with the global features extracted by DINOv2 through multilayer perceptron fusion or attention-weighted fusion to obtain fused features; a seal matching module 50, used to perform feature matching using the fused features and feature vectors in a standard seal database; and a seal image reconstruction module 60, used to perform super-resolution reconstruction of the successfully matched seal images using MambaIRv2 to obtain the final seal.

[0101] Example 10: See you again Figure 6In this embodiment of the invention, the seal image segmentation system for analyzing Chinese calligraphy and painting seals includes: a seal detection module 10, used to detect seals in Chinese calligraphy and painting images using YOLOv10 to obtain the bounding box coordinates of each seal; a seal segmentation module 20, used to extract features and segment images of Chinese calligraphy and painting images using SAM2, using the seal bounding box coordinates as prompts, to obtain multi-scale semantic features of Chinese calligraphy and painting images and all seal images in Chinese calligraphy and painting images; a feature extraction module 30, used to extract features from the seal images output by SAM2 using DINOv2 to obtain global feature vectors of seal images; a feature fusion module 40, used to fuse the multi-scale semantic features extracted by SAM2 with the global features extracted by DINOv2 to obtain fused features; a seal matching module 50, used to perform feature matching using the fused features and feature vectors in a standard seal database; and a seal image reconstruction module 60, used to perform super-resolution reconstruction of the successfully matched seal images using SwinIR to obtain the final seal.

[0102] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention in any way. Those skilled in the art can make various equivalent changes and improvements based on the above embodiments, and all equivalent variations or modifications made within the scope of the claims should fall within the protection scope of the present invention.

Claims

1. A seal image segmentation method for Chinese calligraphy and painting seal analysis, characterized in that, The method comprises the following steps: using a target detection algorithm to detect seals in Chinese calligraphy and painting images, and obtaining the boundary box coordinates of each seal; using the seal boundary box coordinates as a hint, using a hint-guided image segmentation model to extract features and segment the Chinese calligraphy and painting image, and obtaining the multi-scale semantic features of the Chinese calligraphy and painting image and all seal images in the Chinese calligraphy and painting image; extracting features from the seal images output by the image segmentation model to obtain a global feature vector of the seal image; fusing the multi-scale semantic features extracted by the image segmentation model and the global features of the seal image to obtain fused features; performing feature matching on the fused features and the feature vectors in the standard seal database; performing super-resolution reconstruction on the seal images that match successfully to obtain the final seal.

2. The seal image segmentation method for Chinese seal for painting and calligraphy analysis according to claim 1, characterized in that, The seal image segmentation method for Chinese calligraphy and painting seal analysis comprises the following steps: using YOLOv10 to detect seals in Chinese calligraphy and painting images, and obtaining the boundary box coordinates of each seal; using the seal boundary box coordinates as a hint, using SAM2 to extract features and segment the Chinese calligraphy and painting image, and obtaining the multi-scale semantic features of the Chinese calligraphy and painting image and all seal images in the Chinese calligraphy and painting image; using DINOv2 to extract features from the seal images output by SAM2 to obtain a global feature vector of the seal image; fusing the multi-scale semantic features extracted by SAM2 and the global features extracted by DINOv2 to obtain fused features; performing feature matching on the fused features and the feature vectors in the standard seal database; using MambaIRv2 to perform super-resolution reconstruction on the seal images that match successfully to obtain the final seal.

3. The seal image segmentation method for Chinese seal analysis according to claim 2, wherein, The method comprises the following steps: preprocessing the Chinese calligraphy and painting image input by the user; inputting the preprocessed Chinese calligraphy and painting image into the trained YOLOv10 model to detect seals and output the boundary box coordinates of each seal.

4. The seal image segmentation method for Chinese seal analysis according to claim 2, wherein, The method comprises the following steps: inputting the Chinese calligraphy and painting image and the seal boundary box coordinates into the trained SAM2 model; using the image encoder of the SAM2 model to extract features from the Chinese calligraphy and painting image and output multi-scale semantic features; using the hint encoder of the SAM2 model to encode the seal boundary box coordinates into a high-dimensional hint embedding vector; cross-fusing the high-dimensional hint embedding vector and the image features output by the image encoder; based on the fused features, the mask decoder of the SAM2 model generates a segmentation mask corresponding to each hint box; integrating and outputting all the segmentation masks corresponding to the hint boxes, and outputting the multi-scale semantic features extracted by the image encoder.

5. The seal image segmentation method for Chinese seal analysis according to claim 2, wherein, The method comprises the following steps: preparing the seal image to be extracted; performing standard preprocessing on the seal image to be extracted; The pre-processed seal image is input into the DINOv2 network for feature extraction, and a global feature vector of the current seal image is obtained. The global feature vector of the current seal image is saved and recorded.

6. The seal image segmentation method for Chinese seal analysis according to claim 2, wherein, The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded.

7. The seal image segmentation method for Chinese seal analysis according to claim 2, wherein, The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded.

8. The seal image segmentation method for Chinese seal for painting and calligraphy according to claim 1, characterized in that, The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded.

9. A seal image segmentation system for Chinese calligraphy and painting seal analysis, characterized in that, The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded.

10. The seal image segmentation system for Chinese seal for painting and calligraphy analysis according to claim 9, characterized in that, The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The global feature vector of the current seal image is saved and recorded. The seal detection module is configured to perform seal detection on the Chinese calligraphy and painting image using YOLOv10 to obtain a bounding box coordinate of each seal; The seal segmentation module is configured to use SAM2 to perform feature extraction and image segmentation on the Chinese calligraphy and painting image by taking the seal bounding box coordinate as a hint, to obtain multi-scale semantic features of the Chinese calligraphy and painting image and all seal images in the Chinese calligraphy and painting image; The feature extraction module is configured to use DINOv2 to perform feature extraction on the seal images output by SAM2 to obtain global feature vectors of the seal images; The feature fusion module is configured to fuse the multi-scale semantic features extracted by SAM2 and the global features extracted by DINOv2 to obtain fused features; The seal matching module is configured to perform feature matching on the fused features and feature vectors in a standard seal database; The seal image reconstruction module is configured to use MambaIRv2 to perform super-resolution reconstruction on the seal images for which the matching is successful to obtain final seals.

Citation Information

Patent Citations

  • Seal character detection and identification method and device

    CN116682115A

  • Remote sensing image high-quality automatic instance segmentation method based on SAM large model fine tuning

    CN118691815A