Seal image recognition method, device and equipment and medium

Through the preprocessing of seal images and detection and recognition of neural network models, the background interference, defects and blurring problems of seal images in the trade context are solved, and accurate recognition of seal images to text is achieved, thereby improving the security and efficiency of supply chain financing.

CN120673420APending Publication Date: 2025-09-19AGRICULTURAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510718639.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

In the trade context, seal images have problems such as background interference, defects and blur, which makes it difficult to accurately extract effective text information, affecting the security and efficiency of supply chain financing.

Method used

A preprocessing algorithm is used to clarify the seal image, and the seal area is detected and identified by combining target detection and neural network model. The target image recognition model is trained through self-supervised learning and transfer learning to achieve accurate recognition of seal images to text sequences.

Benefits of technology

In the presence of background interference, defects and blur, it can accurately extract text information from seal images, improve the efficiency and accuracy of authenticity review of trade background information, and reduce the demand for annotated data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673420A_ABST
    Figure CN120673420A_ABST
Patent Text Reader

Abstract

The invention discloses a seal image recognition method and device, equipment and a medium. The method comprises the following steps: acquiring a to-be-processed image; wherein the to-be-processed image comprises seal information; preprocessing the to-be-processed image to obtain a target processed image; performing seal area detection on the target processing image to obtain a seal image; and inputting the seal image into a target image recognition model to obtain character sequence data. According to the technical scheme, in a trade background scene, accurate information extraction can be carried out on the seal image with background interference, defect, blurring and other problems, and accurate recognition from the seal image to the character is completed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a seal image recognition method, device, equipment and medium. Background Art

[0002] Supply chain finance refers to providing financing services to upstream and downstream enterprises in the supply chain based on supply chain relationships to increase supply chain liquidity and optimize supply chain capital operations, thereby improving the efficiency and profitability of the entire supply chain.

[0003] Supply chain financing provides financing services based on the actual transaction processes and business relationships in trade activities. Trade background provides the physical support for financing. For example, documents such as purchase contracts, invoices, and delivery notes record basic transaction information. Therefore, the authenticity of trade background is a key foundation for ensuring the security, effectiveness, and sustainable development of supply chain financing transactions.

[0004] Verifying the authenticity of trade background documents involves examining the signatures and seals of relevant parties. Seals on trade background documents are complex and disruptive, with two key characteristics: First, the seal image itself contains significant blank space, while the text area itself occupies a significant portion. Second, the seal image often appears alongside background text, resulting in background interference in some areas. Factors such as scanning quality, photographic conditions, and paper texture can also cause smudges, blurring, or partial obscuration in the image. How to quickly and accurately extract valid textual information from trade background seal images is a major pain point in supply chain financing. Summary of the Invention

[0005] The present invention provides a seal image recognition method, device, equipment and medium, which can accurately extract information from seal images with background interference, defects and blurring in trade background scenarios, and complete precise recognition of seal images into text.

[0006] According to one aspect of the present invention, a seal image recognition method is provided, comprising:

[0007] Acquire an image to be processed; wherein the image to be processed contains seal information;

[0008] Preprocessing the image to be processed to obtain a target processed image;

[0009] Performing seal area detection on the target processed image to obtain a seal image;

[0010] The seal image is input into a target image recognition model to obtain text sequence data.

[0011] According to another aspect of the present invention, there is provided a seal image recognition device, comprising:

[0012] The original image acquisition module is used to acquire the image to be processed; wherein the image to be processed contains seal information;

[0013] A preprocessing module, configured to preprocess the image to be processed to obtain a target processed image;

[0014] A seal image acquisition module is used to perform seal area detection on the target processed image to obtain a seal image;

[0015] The image recognition module is used to input the seal image into a target image recognition model to obtain text sequence data.

[0016] According to another aspect of the present invention, an electronic device is provided, comprising:

[0017] at least one processor; and

[0018] a memory communicatively connected to the at least one processor; wherein,

[0019] The memory stores a computer program executable by the at least one processor. The computer program is executed by the at least one processor so that the at least one processor can execute the seal image recognition method according to any embodiment of the present invention.

[0020] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the seal image recognition method according to any embodiment of the present invention when executed.

[0021] The technical solution of the embodiment of the present invention obtains an image to be processed, wherein the image to be processed contains seal information; preprocesses the image to obtain a target processed image; performs seal region detection on the target processed image to obtain a seal image; and inputs the seal image into a target image recognition model to obtain text sequence data. This technical solution, specifically designed for trade scenarios, can accurately extract information from seal images that suffer from background interference, defects, and blur, achieving precise recognition of seal images into text.

[0022] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0024] Figure 1 This is a flow chart of a seal image recognition method provided according to the first embodiment of the present invention;

[0025] Figure 2 This is a flow chart of a seal image recognition method provided according to the second embodiment of the present invention;

[0026] Figure 3 is an example diagram of the processing process of the first image recognition model provided according to the second embodiment of the present invention;

[0027] Figure 4 is an example diagram of the processing process of the second image recognition model provided according to the second embodiment of the present invention;

[0028] Figure 5 is an example diagram of the processing process of the target image recognition model provided according to the second embodiment of the present invention;

[0029] Figure 6 2 is a schematic structural diagram of a seal image recognition device provided according to a third embodiment of the present invention;

[0030] Figure 7 It is a structural diagram of an electronic device provided according to the fourth embodiment of the present invention. DETAILED DESCRIPTION

[0031] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0032] It should be noted that the terms "first", "second", "original" and "target" in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0033] Example 1

[0034] Figure 1 This is a flow chart of a seal image recognition method provided according to the first embodiment of the present invention. This embodiment is applicable to the recognition of seal images in trade scenarios. The method can be executed by a seal image recognition device. The seal image recognition device can be implemented in the form of hardware and / or software. The seal image recognition device can be configured in an electronic device with data processing capabilities. Figure 1 As shown, the method includes:

[0035] S110: Acquire an image to be processed.

[0036] Among them, the image to be processed contains seal information. Seal information can be understood as the signature and seal information of the relevant parties in the trade background document information. The image to be processed can be understood as referring to various asset document images generated between enterprises in the process of commodity and service transactions. In this embodiment, the image to be processed can be a data file in the trade background. Exemplarily, the image to be processed may include document images such as purchase contracts, invoices, and delivery notes. It is understandable that the seal information in the image to be processed may be affected by background interference or the document or scanning factors, resulting in the presence of stains, blurring, or partial occlusion in the seal image. In this embodiment, the original resource file image containing seal information of different users can be obtained.

[0037] S120 , pre-processing the image to be processed to obtain a target processed image.

[0038] The preprocessing may be performed to sharpen the image. In this embodiment, different algorithms may be used for preprocessing based on actual conditions. For example, the preprocessing in this embodiment may include using algorithms such as spatial domain filtering, channel filtering, and high-contrast overlay. The target processed image may be understood as a clear image obtained through the image preprocessing operation.

[0039] In this embodiment, the image to be processed can be pre-processed using algorithms such as spatial domain filtering, channel filtering, and high-contrast overlay to obtain a target processed image that is more suitable for recognition.

[0040] S130 , performing seal area detection on the target processed image to obtain a seal image.

[0041] The seal area detection may be a seal position area detected by target detection, and the seal image may be a seal image obtained by cropping the seal position area detected in the target processed image.

[0042] In this embodiment, a target detection tool may be used to perform seal region detection on the target processing image to obtain a corresponding seal position region, and then the seal position region may be cropped to obtain a seal image.

[0043] S140: Input the seal image into the target image recognition model to obtain text sequence data.

[0044] The target image recognition model may be a model for recognizing text on a seal image. In this embodiment, the target image recognition model may be a pre-trained neural network model. The text sequence data may be a text sequence obtained by recognizing text on a seal image. It is understood that the text sequence data in this embodiment may include text information such as the company entity corresponding to the seal information and the seal type. For example, the text sequence information may include text information such as "XX Co., Ltd." or "Contract Seal."

[0045] In this embodiment, the seal image can be input into a pre-trained target image recognition model, and the output is the text sequence data contained in the seal image. The recognized text sequence data in this embodiment can be used to verify the authenticity of the established trade related party information and can also be used in other subsequent key factor analysis models.

[0046] The technical solution of the embodiment of the present invention obtains an image to be processed, wherein the image to be processed contains seal information; preprocesses the image to obtain a target processed image; performs seal region detection on the target processed image to obtain a seal image; and inputs the seal image into a target image recognition model to obtain text sequence data. This technical solution, specifically designed for trade scenarios, can accurately extract information from seal images that suffer from background interference, defects, and blur, achieving precise recognition of seal images into text.

[0047] Example 2

[0048] Figure 2This is a flow chart of a seal image recognition method provided by the second embodiment of the present invention. This embodiment is optimized based on the above embodiment. The specific optimization is: performing seal area detection on the target processing image to obtain the seal image, including: inputting the target processing image into the target detection model to detect the seal position area; performing cropping processing on the seal position area in the target processing image to obtain the seal image. Figure 2 As shown, the method includes:

[0049] S210: Acquire an image to be processed.

[0050] The image to be processed contains seal information.

[0051] S220 , pre-processing the image to be processed to obtain a target processed image.

[0052] S230: Input the processed target image into the target detection model to detect and obtain the seal position area.

[0053] The target detection model may be a pre-trained neural network model that can be used to detect target objects. The seal location region may refer to the location region of the seal contained in the target processed image. In this embodiment, the target processed image can be input into the target detection model, and the target detection model can detect the seal information in the target processed image to obtain the corresponding seal location region.

[0054] S240 , performing a cropping process on the seal position area in the target processed image to obtain a seal image.

[0055] The cropping process may be a process operation for cropping the area where the seal position is located. The seal image may refer to an image containing seal information. In this embodiment, the seal position area can be detected in the target processing image, and the corresponding seal position area can be cropped to obtain a seal image containing the seal information.

[0056] S250: Input the seal image into the target image recognition model to obtain text sequence data.

[0057] In this embodiment, optionally, the training steps of the target image recognition model include: obtaining a sample training set and an initial image recognition model; wherein the sample training set includes a general image sample training set and a seal image sample training set; training the initial image recognition model based on the general image sample training set and a self-supervised learning method to obtain a first image recognition model; training the first image recognition model based on the seal image sample training set and a self-supervised learning method to obtain a second image recognition model; training the second image recognition model based on transfer learning and the seal image sample training set to obtain a trained target image recognition model.

[0058] The initial image recognition model may refer to a neural network model that has not been trained with sample data. In this embodiment, the initial image recognition model may include an initial image encoder, a first initial decoder, and a second initial decoder. The encoder may be a ViT encoder; the second initial decoder may be a standard Transformer decoder. The general image sample training set may refer to a dataset containing a large number of image samples. In this embodiment, due to the limited number of seal image samples, pre-training can be performed using a large number of general image samples, thereby improving model accuracy. The seal image sample training set may be a dataset containing a variety of seal image samples. Self-supervised learning may be a method that replaces human annotation by setting pseudo-supervised tasks using certain attributes of the sample data, thereby converting an unsupervised learning problem into a supervised problem. In this embodiment, a script can be used to generate virtually unlimited training data from the millions of images in the general image sample training set. Furthermore, after the model learns feature representations, transfer learning can be used to fine-tune certain supervised tasks. The first image recognition model may be a network model obtained through self-supervised learning using the general image sample training set. The second image recognition model may be a network model obtained through self-supervised learning using the seal image sample training set. Transfer learning is a machine learning method that uses knowledge learned from one task to improve the learning process of another related task, reducing the computational resources and time required to learn the new task while improving model performance. Typically, transfer learning occurs when there is some degree of correlation or similarity between two or more tasks. The target image recognition model can be a network model obtained by training a previously trained second seal image model using transfer learning on a training set of seal image samples.

[0059] In this embodiment, a general image sample training set and a seal image sample training set can be obtained, and then the self-supervised learning mechanism of the masked autoencoder MAE is used to train the initial image recognition model on a large number of general image sample training sets, with the training goal of reconstructing the original sample image from the incomplete image, and the general image feature extraction and reconstruction capabilities of the initial encoder in the initial image recognition model are trained, that is, a first image recognition model is obtained; then the first image recognition model is trained using the disturbed seal image sample training set to fine-tune the model, with the training goal of reconstructing a clean and complete seal image from the disturbed seal sample image, and the image reconstruction capability of the encoder in the first image recognition model for the seal image is trained, that is, a second image recognition model is obtained; finally, the ViT encoder parameters contained in the trained second image recognition model can be fixed, and the second initial decoder can be connected by migration, and the second image recognition model is continued to be trained using the seal image sample training set, so that the high-level semantic features extracted by the encoder can be directly translated, and end-to-end training from the disturbed seal image to the text sequence is performed until a target image recognition model that meets the requirements is obtained.

[0060] In this embodiment, the operation of first training the initial image recognition model with a general image sample training set and then training the model with a seal image sample training set is primarily based on the amount of training data. Since the general image sample training set has a large amount of data, the model can first be subjected to an operation equivalent to pre-training on a very large sample training set, which can enable the entire model to have a certain general completion capability. Moreover, since the seal image sample training set is generally a relatively small data set, training with the seal image sample training set can achieve a fine-tuning effect on the model's completion capability, thereby improving the model's recognition effect.

[0061] In this embodiment, the initial image recognition model uses the Vision Transformer (ViT) as the backbone network model. First, a masked autoencoder (MAE) learning mechanism is used to train the encoder in the initial image recognition model's general image feature extraction and reasoning capabilities. Then, fine-tuning is performed on samples of disturbed seal images to train the encoder's image reconstruction capabilities for seal images. Finally, based on transfer learning, the trained ViT encoder parameters are fixed and transferred to a standard Transformer decoder. An end-to-end training process is performed from disturbed seal images to text sequences, resulting in a trained target image recognition model. In this embodiment, the target image recognition model uses the ViT model as the encoder for transfer, and uses the Transformer as the decoder to generate the recognized text sequence. The ViT model is a version of the Transformer model applied to visual tasks. It uses a self-attention mechanism to enable the model to grasp global information, effectively extract features of important parts of the image, and make the information of each part of the image correlated, thereby facilitating the reconstruction of image information. The Transformer model itself can be directly used to output text sequences, which is consistent with the text recognition task. In this embodiment, both the decoder and encoder are composed of multiple Transformer blocks stacked together.

[0062] Through such a setting in this embodiment, the seal image text information in scenarios such as background text interference, incomplete seals or blurred seals can be effectively identified, providing more convenient and accurate data support for the authenticity review of trade background information, and also improving the reliability and accuracy of model training.

[0063] In this embodiment, optionally, the initial image recognition model includes an initial image encoder and a first initial decoder; the general image sample training set includes multiple image samples; accordingly, the initial image recognition model is trained based on the general image sample training set and the self-supervised learning method to obtain a first image recognition model, including: for each image sample, the image sample is split and masked to obtain a masked image block and an unmasked image block; wherein the masked image block and the unmasked image block both contain position coding information; the unmasked image block is input into the initial image encoder in the initial image recognition model for feature extraction to obtain a first feature vector; the first feature vector and the masked image block are input into the first initial decoder in the initial image recognition model to obtain an output image; a first set loss function is used to combine the output image and the image sample to determine a first loss function value; the encoding parameters and the first decoding parameters in the initial image encoder and the first initial decoder are updated according to the first loss function value, and the first feature vector is returned to be determined again until the first training end condition is met to obtain a trained first image recognition model.

[0064] The splitting process may refer to splitting an image sample into image patches of fixed size. The masking process may refer to randomly masking most of the split image patches, thereby obtaining masked image patches and unmasked image patches. The position encoding may be an encoding using learnable parameters to represent the position information of each image patch. In this embodiment, each image patch obtained after the splitting process carries a corresponding position encoding. The first feature vector may be a feature vector obtained by extracting features from an input unmasked image patch by an initial image encoder. The output image may be a reconstructed image obtained by decoding the input first feature vector and masked image patch by the initial encoder. The first set loss function may be a pre-set function. In this embodiment, the first loss function may be a mean squared error loss function. The first loss function value may be data obtained by calculating pixel loss based on the masked original image patch and each image patch of the generated output image. The encoding parameters may refer to the various parameter information of the encoder. The first decoding parameters may refer to the various parameters of the first initial decoder. The first training end condition may be a pre-set condition. In this embodiment, the first training end condition may be setting the number of model iterations or setting the requirements that the model output image must meet, and may be set according to actual needs.

[0065] In this embodiment, for each image sample, after cropping the input image sample to obtain multiple image patch blocks of a fixed size, the image blocks are flattened and converted into multiple fixed-length vectors. Therefore, each image block can be represented as a vector. Then, a random masking operation is performed on most of the image block vectors to obtain masked image blocks and unmasked image blocks, thereby forming paired input and output data with the original image and converting it into a self-supervised learning task. In addition, because the image sample is split into multiple image patch blocks, the position information is lost, and the self-attention layer in the network does not contain position information. Therefore, a learnable parameter is used as the position encoding to be superimposed with the vector, and the superimposed image block vector is used as the input of the encoder.

[0066] Specifically, this embodiment adopts the self-supervised learning mechanism of MAE and conducts training on a large-scale general image dataset, with the training goal of reconstructing the original image from the incomplete patch, to train the encoder's general image feature extraction and reconstruction capabilities. The processing example of the first image recognition model in this embodiment is shown in the figure below. Figure 3As shown. For visual images, information is highly redundant. Missing a small number of pixels may not confuse the model, and the model can infer based on the surrounding pixel information. Therefore, randomly creating a high proportion of occlusion can create a task that cannot be easily solved by reasoning. The decoder trained on this task can acquire similar reasoning capabilities, which is conducive to learning useful representations. Since there is no occlusion in the downstream task after migration, in order to reduce the input data difference caused by migration, in this embodiment, only the unmasked image patch can be input into the initial image encoder in the initial image recognition model for feature extraction to obtain the first feature vector; after the encoder outputs the first feature vector, the vector of the masked image patch is added and superimposed with the position code to incorporate the position information of the original image sample. At the same time, different masked patches can be distinguished by the position code. This information is then input into the first initial decoder in the initial image recognition model. Finally, the vector output by the first initial decoder is linearly mapped and reconstructed into an image through shape transformation to obtain the output image. The trained model can effectively extract important features from the image and has a certain degree of reasoning ability. In this embodiment, the mean square error loss function can be used in the training process to determine the first loss function value by combining the output image and the original image sample. The pixel loss can be calculated only for the masked original image patch and the generated patch. The calculation of the first loss function value is shown in the following formula:

[0067]

[0068] Where D represents the total number of pixels; N represents the total number of samples; n represents the nth sample; Y'nd represents the value of the output of sample n at the dth pixel predicted by the model; Ynd represents the correct label of the value of sample n at the dth pixel.

[0069] In this embodiment, the encoding parameters in the initial image encoder and the first decoding parameters in the first initial decoder can be updated respectively according to the first loss function value, and the unmasked image blocks of each image sample are input into the initial image encoder for feature extraction to obtain a first feature vector in a loop; the first feature vector and the masked image block are input into the first initial decoder in the initial image recognition model to obtain an output image, and the output image and the image sample are combined based on the set loss function formula to determine the first loss function value, and the encoding parameters and the first decoding parameters are continuously updated by the first loss function value, and the model is iterated until the first training end condition of the model is met, thereby obtaining a trained first image recognition model.

[0070] Through such a setting in this embodiment, in the image reconstruction training task, the output of the encoder can be directly used as the input of the decoder, and the output of the decoder is mapped by the linear layer and then reconstructed through shape transformation, thereby completing the training operation of the first image recognition model and improving the reliability of model training.

[0071] In this embodiment, optionally, for each image sample, the image sample is split and masked to obtain a masked image block and an unmasked image block, including: for each image sample, the image sample is split according to a set size to obtain multiple image block vectors; and the multiple image block vectors are randomly masked according to a set ratio to obtain a masked image block and an unmasked image block.

[0072] Among them, the set size can be a pre-set image block size. In this embodiment, it can be set according to actual needs, and this embodiment does not limit this. The set ratio can be a pre-set ratio and can be set according to actual needs. For example, the set ratio in this embodiment can be 80%. The random masking operation can be an operation of assigning a set number to each image block of the set ratio to achieve the purpose of masking. In this embodiment, the set number can be 1 or 0. The masked image block can refer to the image block obtained after the assignment process. The unmasked image block can be the original image block obtained after the image sample is split.

[0073] In this embodiment, for each image sample, the image sample can be split according to a set fixed size to obtain multiple fixed-size image block patches, and the image block patches are flattened and converted into multiple fixed-length vectors; the vectors are then sent to a linear layer for linear mapping, thereby obtaining multiple image block vectors; in this embodiment, most image block vectors can be randomly masked according to a pre-set ratio, thereby obtaining masked image block vectors and unmasked image block vectors.

[0074] Through such a setting in this embodiment, each image sample can be split and masked, so that the input object and output object of the model can be formed based on the original image, so as to facilitate the training of the model for self-supervised learning tasks, thereby improving the efficiency of model training.

[0075] Furthermore, in this embodiment, the first image recognition model is trained based on the seal image sample training set and the self-supervised learning method to obtain the specific process of the second image recognition model as follows: on the disturbed seal image sample training set, the first image recognition model is fine-tuned, with the reconstruction of a clean and complete seal image from the disturbed seal image as the training goal, and the corresponding encoder in the first image recognition model is trained for the image reconstruction capability of the seal image. The processing process example of the second image recognition model in this embodiment is shown in the figure below. Figure 4 As shown. In this embodiment, each seal image sample can be cut into image patches, and all image patches are sent to the first image recognition model trained in the first stage for further training. The output of the encoder is directly used as the input of the first decoder. Finally, the vector output by the first decoder is linearly mapped and transformed into the reconstructed seal image through shape transformation. In this embodiment, the training of the first image recognition model has been able to effectively extract features and has a certain degree of reasoning ability. On this basis, a smaller learning rate is used to fine-tune the model and apply it to the image reconstruction task. For disturbed seal images, it will be able to better learn to reconstruct seal images. In this embodiment, the training of the second image recognition model can use the mean square error loss function, and combine it with the total variation loss as a regularization term to calculate the reconstructed seal image and the original seal image. The role of the total variation loss in image reconstruction and denoising is to maintain the smoothness of the image and eliminate the artifacts that may be caused by image restoration. In some images with more details, it will make the restored image too smooth, but it is well suited for seal images with almost no detailed texture. For a certain sample, the total variation loss is calculated as follows:

[0076]

[0077] Where Xi,j represents the pixel value of the i-th row and j-th column of the image; β is a hyperparameter. The larger the value of β, the smoother the image, but the more serious the loss of details. It is usually set to 2.

[0078] In this embodiment, the first image recognition model is iteratively trained using a seal image sample training set and self-supervised learning, and the parameters of the corresponding first encoder in the first image recognition model are updated using the loss function value until the training conditions are met, thereby obtaining a trained second image recognition model.

[0079] Furthermore, in this embodiment, when a training set of seal image samples is insufficient, seal images from trade background materials can be simulated using clear original seal images through preprocessing methods such as synthesizing background text, filter blurring, rotation, and partial area fading. This generates paired input images and label images, transforming the task into self-supervised learning. When a sufficient dataset is available, seal images from real trade background materials and original seal images can be directly used as training data.

[0080] In this embodiment, optionally, the second image recognition model includes a trained image encoder and a second initial decoder; the seal image sample training set includes multiple seal image samples; the second image recognition model is trained using transfer learning and the seal image sample training set to obtain a trained target image recognition model, including: for each seal image sample, determining its corresponding input image block and label text sequence; fixing the corresponding image encoder parameters in the second image recognition model, and connecting the second initial decoder through migration; inputting the input image block into the image encoder in the second image recognition model for feature extraction to obtain a second feature vector; inputting the second feature vector into the second initial decoder to obtain output text data; using a second set loss function formula to combine the output text data and the label text sequence to determine the second loss function value; updating the second decoding parameters in the second initial decoder according to the second loss function value, returning to re-determine the text sequence data until the second training end condition is met, and obtaining a trained target image recognition model.

[0081] The input image block may be a plurality of image blocks obtained by splitting a seal image sample. The label text sequence may refer to corresponding text sequence data obtained by identifying the seal image sample. The second feature vector may be a feature vector obtained by feature extraction through an image encoder in the second image recognition model. It is understood that the image encoder of the second image recognition model in this embodiment is a trained encoder. The output text data may refer to text information output by the second feature vector after passing through the second initial decoder. The output text data in this embodiment may be text sequence data. The second set loss function formula may be a pre-set loss function relationship formula. The second set loss function formula in this embodiment may be a cross-entropy loss function relationship formula. The second loss function value may be data obtained by calculating the difference between the output text data and the label text sequence using the second set loss function formula. The second decoding parameter may refer to various parameters in the second initial decoder. The second training end condition may be a pre-set condition. The second training end condition in this embodiment may be a setting of the number of model iterations or a setting of requirements for the text information output by the model, and may be set according to actual needs.

[0082] Specifically, in this embodiment, for each seal image sample, the input image block and the corresponding label text sequence corresponding to each impression image sample are determined, and then the ViT encoder parameters in the trained second image recognition model can be fixed, and the second initial decoder can be connected through transfer learning to directly translate the high-level semantic features extracted by the image encoder, and perform end-to-end training from the disturbed seal image to the text sequence. The processing example of the target image recognition model in this embodiment is shown in the figure below. Figure 5 As shown. In this embodiment, the seal image sample training set in the training process of the second image recognition model can be reused, but the training label is the seal text content. If the seal image contains multiple paragraphs of text, they can be distinguished by manually setting delimiters. In this embodiment, for each seal image sample, the seal image sample with interference is cut into corresponding image block patches, and all image block patches are sent to the image encoder in the trained second image recognition model for feature extraction to obtain a second eigenvector. The second eigenvector output by the encoder is then used as the intermediate layer input of the second initial decoder, and the output of each decoder is used as the initial input of the next decoder to calculate the self-attention together, and finally output the probability value vector of the current word, that is, the output text sequence data can be obtained. Then in the model training, the cross entropy loss function can be used to calculate the difference between the probability of the model output text data and the label text sequence to determine the value of the second loss function. The calculation method is shown in the following formula:

[0083]

[0084] Among them, K represents the total number of categories contained in the label; N represents the total number of samples; n represents the nth sample; Pnc represents the probability that sample n belongs to category c; lnc represents the correct probability label of category c for sample n.

[0085] In this embodiment, the second decoding parameters in the second initial decoder can be adjusted according to the second loss function value, and the input image block can be input into the image encoder in the second image recognition model for feature extraction in a loop to obtain a second feature vector; the second feature vector is input into the second initial decoder to obtain output text data; the second loss function value is determined by combining the output text data and the label text sequence data using the second set loss function formula; the second decoding parameters are continuously updated by the second loss function value, and the model is iterated in a loop until the second training end condition of the model is met, thereby obtaining a trained target image recognition model.

[0086] In this embodiment, the encoder trained with the first and second image recognition models in the first two steps can accurately extract features from key areas of the seal image and has a high degree of noise resistance. The feature vectors extracted by the encoder can be considered to have been reconstructed within a high-level semantic space. By connecting to the decoder and retranslating them into natural language understandable to humans, the trained model in this embodiment can accurately extract information from seal images with interference, defects, and blur in trade contexts, completing end-to-end inference and recognition from seal images to text, providing more convenient and accurate data support for the authenticity review of trade background materials.

[0087] Through such a setting in this embodiment, the self-supervised learning mechanism and transfer learning of the masked autoencoder MAE are introduced into the seal text recognition task. While training the network's seal image reconstruction and feature extraction capabilities, the generalization ability of the model is improved, and the demand for seal image annotation data during end-to-end training of the deep learning network model is greatly reduced.

[0088] In this implementation, by reusing the encoder of the image reconstruction task and connecting it to the decoder of the seal recognition task, the image reconstruction and denoising in the high-level semantic space are completed, which can improve the performance of the model in interference scenarios and achieve end-to-end reasoning.

[0089] In this embodiment, optionally, the output text data is text sequence data; the second feature vector is input into the second initial decoder to obtain the output text data, including: inputting the second feature vector and the output data of the previous initial decoder into the second initial decoder until the text sequence data is obtained.

[0090] In this embodiment, during the text recognition training task, the output of the previous decoder is used as the initial input for this step. The output of the encoder is added to the intermediate layer to jointly calculate the multi-head self-attention. The final output vector undergoes linear mapping and Softmax processing to become the word probability vector for this step. Since each decoder output data represents a single character during the text recognition process of the second initial decoder, the output data of the previous initial decoder must also be input into the current second initial decoder to obtain the entire text sequence data.

[0091] In this embodiment, for each input to the second initial decoder, the second feature vector and the output data of the previous initial decoder are used as input objects to the second initial decoder, so that text sequence data can be obtained.

[0092] In this embodiment, such a setting can improve the text recognition function in complex seal image scenes and improve the accuracy of text recognition.

[0093] The introduction of transfer learning in this embodiment significantly reduces the demand for seal image annotated data and improves the model's generalization ability. The MAE training method can train the network's image reconstruction and feature extraction capabilities, enhancing the model's performance. Furthermore, the use of MAE and transfer learning training methods improves the network model's ability to recognize seal images in interference scenarios, reducing the demand for annotated data. Furthermore, the image encoder trained through the image reconstruction task can accurately extract features from key areas of the seal image and has a high degree of noise resistance. The extracted feature vectors are reconstructed within a high-level semantic space, which can improve the model's performance in interference scenarios. While the step of reconstructing the complete image from high-level semantic features is practical for the human eye, it is redundant for the neural network model, and subsequent feature extraction of the restored image may even result in information loss. Furthermore, considering that the text area itself occupies a large area of ​​the seal image, the process of detecting the position of the text in the seal image and performing image correction has no practical business significance and may even result in additional computational overhead and information loss. Therefore, we chose to directly perform end-to-end inference training on the model, omitting the steps of reconstructing the human eye visual image in the preprocessing, as well as the steps of detecting the position of the seal text and performing image correction, thereby reducing information loss and speeding up the detection speed.

[0094] The technical solution of the embodiment of the present invention obtains a to-be-processed image containing seal information; pre-processes the to-be-processed image to obtain a target processed image; inputs the target processed image into a target detection model to detect the seal location region; crops the seal location region in the target processed image to obtain a seal image; and inputs the seal image into a target image recognition model to obtain text sequence data. This technical solution, specifically designed for trade scenarios, can accurately extract information from seal images that suffer from background interference, defects, and blur, achieving precise recognition of seal images into text.

[0095] Example 3

[0096] Figure 6 Schematic diagram of a seal image recognition device according to the third embodiment of the present invention. Figure 6 As shown, the device includes:

[0097] The original image acquisition module 610 is used to acquire the image to be processed; wherein the image to be processed includes seal information;

[0098] A preprocessing module 620 is used to preprocess the image to be processed to obtain a target processed image;

[0099] The seal image acquisition module 630 is used to perform seal area detection on the target processed image to obtain a seal image;

[0100] The image recognition module 640 is used to input the seal image into the target image recognition model to obtain text sequence data.

[0101] Optionally, the seal image acquisition module 630 is specifically configured to:

[0102] Input the processed target image into the target detection model to detect the seal location area;

[0103] The seal position area in the target processed image is cropped to obtain the seal image.

[0104] Optionally, the seal image recognition device can also obtain a target image recognition model through the following model training:

[0105] A data acquisition module is used to acquire a sample training set and an initial image recognition model; wherein the sample training set includes a general image sample training set and a seal image sample training set;

[0106] A first model training module is used to train the initial image recognition model based on a general image sample training set and a self-supervised learning method to obtain a first image recognition model;

[0107] A second model training module is used to train the first image recognition model based on the seal image sample training set and a self-supervised learning method to obtain a second image recognition model;

[0108] The third model training module is used to train the second image recognition model based on transfer learning and the seal image sample training set to obtain a trained target image recognition model.

[0109] Optionally, the initial image recognition model includes an initial image encoder and a first initial decoder; the general image sample training set includes multiple image samples;

[0110] Accordingly, the first model training module includes:

[0111] a sample processing unit, configured to perform a splitting process and a masking process on each image sample to obtain a masked image block and an unmasked image block; wherein both the masked image block and the unmasked image block contain position coding information;

[0112] A first feature extraction unit is configured to input the unmasked image block into an initial image encoder in an initial image recognition model for feature extraction to obtain a first feature vector;

[0113] an output image acquisition unit, configured to input the first feature vector and the masked image block into a first initial decoder in the initial image recognition model to obtain an output image;

[0114] a loss function value determining unit, configured to determine a loss function value by combining the output image and the image sample using a first set loss function formula;

[0115] A parameter updating unit is used to update the encoding parameters and the first decoding parameters in the initial image encoder and the first initial decoder according to the loss function value, return to re-determine the first feature vector until the first training end condition is met, and obtain a trained first image recognition model.

[0116] Optionally, the sample processing unit is specifically used to:

[0117] For each image sample, the image sample is split according to the set size to obtain multiple image block vectors;

[0118] A random masking operation is performed on multiple image block vectors according to a set ratio to obtain masked image blocks and unmasked image blocks.

[0119] Optionally, the second image recognition model includes a trained image encoder and a second initial decoder; the seal image sample training set includes a plurality of seal image samples; and the third model training module is specifically used to:

[0120] The second image recognition model is trained using transfer learning and a seal image sample training set to obtain a trained target image recognition model, including:

[0121] For each seal image sample, determine its corresponding input image block and label text sequence;

[0122] Fix the corresponding image encoder parameters in the second image recognition model and connect the second initial decoder through migration;

[0123] Inputting the input image block into the image encoder in the second image recognition model to perform feature extraction to obtain a second feature vector;

[0124] Inputting the second feature vector into the second initial decoder to obtain output text data;

[0125] The second set loss function is used to combine the output text data and the label text sequence to determine the loss function value;

[0126] The second decoding parameter in the second initial decoder is updated according to the loss function value, and the text sequence data is returned to be re-determined until the second training end condition is met to obtain a trained target image recognition model.

[0127] Optionally, the output text data is text sequence data; the third model training module is specifically used to:

[0128] The second feature vector and the output data of the previous initial decoder are input into the second initial decoder until the text sequence data is obtained.

[0129] A seal image recognition device provided by an embodiment of the present invention can execute a seal image recognition method provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects of the execution method.

[0130] Example 4

[0131] Figure 7 1 is a schematic diagram of the structure of an electronic device provided according to embodiment four of the present invention. The electronic device 10 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.

[0132] like Figure 7 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0133] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0134] The processor 11 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any other suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the seal image recognition method.

[0135] In some embodiments, the seal image recognition method can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the seal image recognition method described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to perform the seal image recognition method in any other appropriate manner (for example, by means of firmware).

[0136] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0137] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0138] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0139] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0140] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0141] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.

[0142] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.

[0143] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. A seal image recognition method, characterized in that: include: Acquire an image to be processed; wherein the image to be processed contains seal information; Preprocessing the image to be processed to obtain a target processed image; Performing seal area detection on the target processed image to obtain a seal image; The seal image is input into a target image recognition model to obtain text sequence data.

2. The method according to claim 1, characterized in that Performing seal area detection on the target processed image to obtain a seal image includes: Inputting the target processed image into a target detection model to detect and obtain a seal position area; The seal position area in the target processed image is cropped to obtain a seal image.

3. The method according to claim 1, characterized in that The training steps of the target image recognition model include: Obtaining a sample training set and an initial image recognition model; wherein the sample training set includes a general image sample training set and a seal image sample training set; Training the initial image recognition model based on the general image sample training set and a self-supervised learning method to obtain a first image recognition model; Training the first image recognition model based on the seal image sample training set and a self-supervised learning method to obtain a second image recognition model; The second image recognition model is trained based on transfer learning and the seal image sample training set to obtain a trained target image recognition model.

4. The method according to claim 3, characterized in that The initial image recognition model includes an initial image encoder and a first initial decoder; the general image sample training set includes a plurality of image samples; Accordingly, the initial image recognition model is trained based on the general image sample training set and the self-supervised learning method to obtain a first image recognition model, including: For each image sample, performing a splitting process and a masking process on the image sample to obtain a masked image block and an unmasked image block; wherein the masked image block and the unmasked image block both contain position coding information; Inputting the unmasked image block into the initial image encoder in the initial image recognition model to perform feature extraction to obtain a first feature vector; Inputting the first feature vector and the masked image block into a first initial decoder in the initial image recognition model to obtain an output image; Determine a first loss function value by combining the output image and the image sample using a first set loss function formula; The encoding parameters and the first decoding parameters in the initial image encoder and the first initial decoder are updated according to the first loss function value, and the first feature vector is determined again until the first training end condition is met to obtain a trained first image recognition model.

5. The method according to claim 4, characterized in that For each image sample, performing a splitting process and a masking process on the image sample to obtain a masked image block and an unmasked image block, including: For each image sample, split the image sample according to a set size to obtain multiple image block vectors; A random masking operation is performed on multiple image block vectors according to a set ratio to obtain masked image blocks and unmasked image blocks.

6. The method according to claim 3, characterized in that The second image recognition model includes a trained image encoder and a second initial decoder; the seal image sample training set includes a plurality of seal image samples; The second image recognition model is trained using transfer learning and the seal image sample training set to obtain a trained target image recognition model, including: For each seal image sample, determine its corresponding input image block and label text sequence; Fixing the corresponding image encoder parameters in the second image recognition model and connecting the second initial decoder via migration; Inputting the input image block into the image encoder in the second image recognition model to perform feature extraction to obtain a second feature vector; Inputting the second feature vector into the second initial decoder to obtain output text data; Using a second set loss function formula to combine the output text data and the label text sequence to determine a second loss function value; The second decoding parameter in the second initial decoder is updated according to the second loss function value, and the text sequence data is returned to be re-determined until the second training end condition is met to obtain a trained target image recognition model.

7. The method according to claim 6, characterized in that The output text data is text sequence data; Inputting the second feature vector into the second initial decoder to obtain output text data includes: The second feature vector and the output data of the last initial decoder are input into the second initial decoder until the text sequence data is obtained.

8. A seal image recognition device, characterized in that: include: The original image acquisition module is used to acquire the image to be processed; wherein the image to be processed contains seal information; A preprocessing module, configured to preprocess the image to be processed to obtain a target processed image; A seal image acquisition module is used to perform seal area detection on the target processed image to obtain a seal image; The image recognition module is used to input the seal image into a target image recognition model to obtain text sequence data.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor. The computer program is executed by the at least one processor so that the at least one processor can execute the seal image recognition method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the seal image recognition method according to any one of claims 1 to 7 when executed.