Image processing method and device, equipment, storage medium and program product

Through feature conversion and feature extraction of the target detection model, combined with the exponential moving average attention mechanism and color gamut range separation, the accuracy of seal and signature recognition in paper contracts is solved, and the precise separation and identification of seal and signature are achieved.

CN120340053APending Publication Date: 2025-07-18LIAONING MOBILE COMM +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510376688.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, the scanned or photographed paper contracts are unable to clearly and accurately identify the seal and signature due to overlapping seals and signatures, background creases, stains, fading, patterns and watermarks.

Method used

The object detection model is used for feature conversion and feature extraction, and the feature extraction network of the exponential moving average attention mechanism captures long-distance dependencies, combines global and local feature detection seals and signatures, and pixel separation is performed through preset color gamut range.

Benefits of technology

It realizes accurate identification and separation of seals and signatures, reduces the number of model parameters and calculation complexity, and improves the accuracy and efficiency of identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340053A_ABST
    Figure CN120340053A_ABST
Patent Text Reader

Abstract

The invention discloses an image processing method and device, equipment, a storage medium and a program product. The method comprises the following steps: acquiring a to-be-processed image, wherein the to-be-processed image comprises a seal and a signature; inputting the to-be-processed image into a target detection model, and performing feature conversion on the to-be-processed image by using a feature conversion network of the target detection model to obtain a plurality of feature maps; a feature extraction network of the target detection model is used to extract global features and local features of the plurality of feature maps, and the feature extraction network is configured with an index moving average attention mechanism; using a target detection network of the target detection model to detect a seal and a signature in the to-be-processed image based on the global features and the local features to obtain a detection result; and carrying out pixel separation on the seal and the signature in the detection result according to a preset color gamut range to obtain the seal and / or the signature. According to the embodiment of the invention, the seal and the signature on the file are accurately identified and separated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of image processing, and particularly relates to an image processing method, apparatus, device, storage medium, and program product. Background Art

[0002] Currently, the recognition of seals and signatures is becoming increasingly important. For example, during the inspection of business contracts, it is usually necessary to recognize the seals and legal person signatures on the business contracts to complete the inspection.

[0003] In the related art, the OCR (Optical Character Recognition) method is usually used to extract the seals and signatures in business contracts to complete the inspection. However, since the current electronic contracts entered into the system are all scanned or photographed copies of paper contracts, and there are problems such as overlapping of seals and signatures, background creases, stains, fading, patterns, and watermarks on the paper contracts, it is impossible to clearly and accurately recognize the seals and signatures on the scanned or photographed copies of paper contracts when using the OCR method. Summary of the Invention

[0004] Embodiments of this application provide an image processing method, apparatus, device, storage medium, and program product, which can accurately recognize and separate the seals and signatures on a document.

[0005] On the one hand, embodiments of this application provide an image processing method, the method including:

[0006] Obtain an image to be processed, where the image to be processed includes a seal and a signature;

[0007] Input the image to be processed into a target detection model, and use the feature transformation network of the target detection model to perform feature transformation on the image to be processed to obtain a plurality of feature maps;

[0008] Use the feature extraction network of the target detection model to extract the global features and local features of the plurality of feature maps, where the feature extraction network is configured with an exponential moving average attention mechanism;

[0009] Use the target detection network of the target detection model to detect the seals and signatures in the image to be processed based on the global features and the local features to obtain a detection result;

[0010] Perform pixel separation on the seals and signatures in the detection result according to a preset color gamut range to obtain the seal and / or signature.

[0011] Optionally, before inputting the image to be processed into the target detection model, the method further includes:

[0012] Obtain an image training sample set, where the image training sample set includes a plurality of image training samples, and each image training sample includes a sample image and an image marking box corresponding to the sample image; the sample image includes a first seal image and a first signature image; the image marking box is used to characterize the position area of the first seal image and the first signature image;

[0013] Use the image training sample set to perform model training to obtain the target detection model.

[0014] Optionally, the using the image training sample set to perform model training to obtain the target detection model includes:

[0015] Input the sample image into a preset detection model to obtain a predicted target detection box;

[0016] Determine a first loss function of the preset detection model according to the predicted target detection box and the image marking box;

[0017] In the case that the first loss function does not meet the preset convergence condition, adjust the model parameters of the preset detection model, and return to the step of inputting the sample image into the preset detection model to obtain a predicted target detection box until the first loss function meets the preset convergence condition.

[0018] Optionally, the adjusting the model parameters of the preset detection model includes:

[0019] Perform exponential moving average smoothing processing on the model parameters to adjust the model parameters.

[0020] Optionally, the preset detection model includes: a feature conversion network, a feature extraction network, and a target detection network. The number of pooling layers in the feature conversion network is greater than a preset number of pooling layers, and the convolution kernel size of the preset detection model is smaller than a preset convolution kernel size;

[0021] The feature extraction network includes a global feature processing unit and a local feature processing unit, and the global feature processing unit is configured with an exponential moving average attention mechanism;

[0022] The target detection network includes a feature fusion sub-network and a feature classification sub-network;

[0023] The inputting the sample image into the preset detection model to obtain a predicted target detection box includes:

[0024] Perform feature conversion on the sample image through the feature conversion network to obtain a plurality of sample feature maps;

[0025] The global feature processing unit performs global feature extraction on the multiple sample feature maps to obtain multiple predicted global features;

[0026] The local feature processing unit performs local feature extraction on the multiple sample feature maps to obtain multiple predicted local features;

[0027] The feature fusion sub-network performs feature fusion on the multiple predicted global features and the multiple predicted local features to obtain multiple predicted candidate detection frames;

[0028] The feature classification sub-network calculates the confidence of the multiple candidate predicted detection frames to obtain the predicted confidence of each predicted candidate detection frame;

[0029] The feature classification sub-network outputs the predicted target detection frame, and the predicted target detection frame includes the predicted candidate detection frame corresponding to the predicted confidence greater than the preset confidence threshold.

[0030] Optionally, before the feature classification sub-network outputs the predicted target detection frame, the method further includes:

[0031] Dilate the predicted target detection frame outward by a preset number of pixels to obtain the dilated predicted detection frame;

[0032] The output of the predicted target detection frame by the feature classification sub-network includes:

[0033] The feature classification sub-network outputs the dilated predicted target detection frame.

[0034] Optionally, the feature extraction network includes a global feature sub-network and a local feature sub-network, wherein the global feature sub-network is configured with an exponential moving average attention mechanism;

[0035] The extraction of the global features and local features of the multiple feature maps by using the feature extraction network of the object detection model includes:

[0036] The global feature extraction sub-network performs global feature extraction on the multiple feature maps to obtain multiple global features;

[0037] The local feature extraction sub-network performs local feature extraction on the multiple feature maps to obtain multiple local features.

[0038] Optionally, the object detection network includes a feature fusion sub-network and a feature classification sub-network;

[0039] The use of the object detection network of the object detection model to detect the seals and signatures in the image to be processed based on the global features and the local features to obtain the detection result includes:

[0040] Fusing multiple global features and multiple local features through the feature fusion sub-network to obtain multiple first candidate detection frames;

[0041] Calculating the confidence of multiple first candidate detection frames through the feature classification sub-network to obtain the first confidence of each first candidate detection frame;

[0042] Outputting the detection result through the feature classification sub-network, where the detection result includes the first candidate detection frames corresponding to the first confidence being greater than the preset confidence threshold.

[0043] Optionally, after separating the seal and signature in the detection result according to the preset color gamut range to obtain the seal and / or signature, the method further includes:

[0044] Inputting the seal and / or the signature into the backbone extraction module in the image enhancement model to extract features of the seal and / or the signature, obtaining multiple features to be enhanced, where the features to be enhanced are used to represent the detailed features of the attributes of the seal and / or the signature;

[0045] Inputting multiple features to be enhanced into the separation and distillation module in the image enhancement model to perform detail enhancement on the features to be enhanced, obtaining the seal and / or the signature after detail enhancement; wherein, the number of the separation and distillation modules is less than the preset number.

[0046] Optionally, the separation and distillation module includes a feature distillation unit, a feature concentration unit, and a feature enhancement unit;

[0047] The step of inputting multiple features to be enhanced into the separation and distillation module in the image enhancement model to perform detail enhancement on the features to be enhanced, obtaining the seal and / or the signature after detail enhancement, includes:

[0048] Performing feature distillation on multiple features to be enhanced through the feature distillation unit to obtain multiple first key features, where the first key features are used to represent the correlation information between the features to be enhanced;

[0049] Performing feature concentration on the first key features through the feature concentration unit to obtain second key features, where the second key features are used to represent the first key features after enhanced feature expression;

[0050] Performing feature enhancement on the second key features through the feature enhancement unit to obtain the seal and / or the signature after detail enhancement.

[0051] Optionally, the feature enhancement unit includes a convolutional channel sub-unit and a super-resolution sub-unit;

[0052] Performing feature enhancement on the second key feature through the feature enhancement unit to obtain the seal and / or the signature with enhanced details includes:

[0053] Performing differential feature extraction on the second key feature through the convolutional channel subunit to obtain differential features;

[0054] Performing detail enhancement on the differential features to obtain enhanced differential features;

[0055] Performing image restoration on the enhanced differential features through the super-resolution subunit to obtain the seal and / or the signature with enhanced details.

[0056] On the other hand, an embodiment of the present application provides an image processing apparatus, which includes:

[0057] An acquisition module, configured to acquire an image to be processed, where the image to be processed includes a seal and a signature;

[0058] A conversion module, configured to input the image to be processed into a target detection model, and perform feature conversion on the image to be processed by using the feature conversion network of the target detection model to obtain a plurality of feature maps

[0059] An extraction module, configured to extract the global features and local features of the plurality of feature maps by using the feature extraction network of the target detection model, where the feature extraction network is configured with an exponential moving average attention mechanism;

[0060] A detection module, configured to detect the seal and the signature in the image to be processed based on the global features and the local features by using the target detection network of the target detection model to obtain a detection result;

[0061] A pixel separation module, configured to perform pixel separation on the seal and the signature in the detection result according to a preset color gamut range to obtain the seal and / or the signature.

[0062] On yet another aspect, an embodiment of the present application provides an electronic device, which includes: a processor and a memory storing computer program instructions;

[0063] When the processor executes the computer program instructions, the image processing method described in the first aspect is implemented.

[0064] On yet another aspect, an embodiment of the present application provides a computer storage medium, where computer program instructions are stored on the computer-readable storage medium, and when the computer program instructions are executed by a processor, the image processing method described in the first aspect is implemented.

[0065] In another aspect, an embodiment of the present application further provides a computer program product. When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device is caused to execute the image processing method as described in the first aspect.

[0066] For the image processing method, apparatus, device, storage medium and program product according to the embodiments of the present application, in the embodiments of the present application, when processing an image, the image to be processed can be input into a target detection model, and the target detection model performs feature transformation on the image to be processed to obtain a plurality of feature maps corresponding to the image to be processed. Then, a feature extraction network configured with only an exponential moving average attention mechanism is used for feature extraction, so as to better capture long-range dependencies in the image to be processed, enhance the effect of extracting global features and local features, reduce the number of model parameters and computational complexity, and finally perform detection based on the extracted global features and local features to obtain a detection result, so as to more clearly extract the seal and signature. Then, the background color interference in the detection result is removed by using the color range to obtain the completely separated seal and / or signature. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments of the present application. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0068] Figure 1 is a schematic flowchart of an image processing method provided by an embodiment of the present application;

[0069] Figure 2 is a schematic flowchart of an image processing method provided by another embodiment of the present application;

[0070] Figure 3 is a schematic flowchart of an image processing method provided by another embodiment of the present application;

[0071] Figure 4 is a schematic flowchart of an image processing method provided by yet another embodiment of the present application;

[0072] Figure 5 is a schematic flowchart of an image processing method provided by yet another embodiment of the present application;

[0073] Figure 6 is a schematic flowchart of an image processing method provided by still another embodiment of the present application;

[0074] Figure 7 is a schematic flowchart of an image processing method provided by still another embodiment of the present application;

[0075] Figure 8 It is a schematic structural diagram of an image processing apparatus provided by another embodiment of the present application;

[0076] Figure 9 It is a schematic structural diagram of an electronic device provided by another embodiment of the present application. Detailed implementation manners

[0077] The features and exemplary embodiments of various aspects of the present application will be described in detail below. To make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without some of these specific details. The following description of the embodiments is only intended to provide a better understanding of the present application by showing examples of the present application.

[0078] It should be noted that, in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, the elements defined by the statement "comprising..." do not exclude the presence of additional identical elements in the process, method, article or device comprising the said elements.

[0079] To solve the problems of the prior art, embodiments of the present application provide an image processing method, apparatus, device, storage medium and program product. In the embodiments of the present application, when processing an image, the image to be processed can be input into a target detection model, and the target detection model performs feature transformation on the image to be processed to obtain a plurality of feature maps corresponding to the image to be processed. Then, a feature extraction network configured with only an exponential moving average attention mechanism is used for feature extraction, so as to better capture the long-range dependencies in the image to be processed, enhance the effect of extracting global features and local features, reduce the number of model parameters and computational complexity. Finally, detection is performed based on the extracted global features and local features to obtain a detection result, so as to more clearly extract the seal and signature. Then, the background color interference in the detection result is removed by using the color range to obtain the completely separated seal and / or signature.

[0080] The image processing method provided by the embodiments of the present application will be introduced first below.

[0081] Figure 1 The flowchart shows the image processing method provided by an embodiment of the present application. As Figure 1 shown, the image processing method may include S101 - S105:

[0082] S101, Obtain the image to be processed.

[0083] In some embodiments, the electronic device may obtain the image to be processed uploaded by the user, and may also obtain the image to be processed stored on the cloud platform through the cloud platform for image processing. Among them, the image to be processed includes a seal and a signature. When performing image processing, the seal and signature can be recognized to facilitate subsequent processing of the seal and signature.

[0084] S102, Input the image to be processed into the target detection model, and use the feature transformation network of the target detection model to perform feature transformation on the image to be processed to obtain multiple feature maps.

[0085] In some embodiments, after obtaining the image to be processed, the image to be processed is input into the target detection model. Then, the feature transformation network in the target detection model will convert the image to be processed into multiple feature maps, that is, vectorize it, to facilitate target recognition of the image to be processed.

[0086] S103, Use the feature extraction network of the target detection model to extract the global features and local features of the multiple feature maps.

[0087] In some embodiments, after obtaining multiple feature maps, the feature extraction network will extract multiple global features and local features from the multiple feature maps. Among them, in order to better extract local features and global features, the feature extraction network is configured with an Exponential Moving Average (EMA) mechanism to improve the model training efficiency, reduce the number and computational complexity of the models, make the model more efficient and suitable for running on devices with limited resources. At the same time, it can also capture target features at different scales to extract local features and global features.

[0088] S104, Use the target detection network of the target detection model to detect the seal and signature in the image to be processed based on the global features and local features to obtain the detection result.

[0089] In some embodiments, the target detection network first fuses the global features and local features, and then performs seal and signature detection based on the fused global features and local features to obtain the detection result.

[0090] It should be noted that the detection result can include the position areas of the seal and signature, which are marked with image marking frames, that is, the target detection model can accurately identify the position areas of the seal and signature in the image to be processed through global features and local features, and mark them with image marking frames.

[0091] S105, perform pixel separation on the seal and signature in the detection result according to a preset color gamut range to obtain the seal and / or signature.

[0092] In some embodiments, since the colors of the seal and signature are inconsistent, the seal and signature can be pixel-separated through a preset color gamut range to obtain the seal and / or signature.

[0093] In the embodiments of the present application, when processing an image, the image to be processed can be input into the target detection model. The target detection model performs feature transformation on the image to be processed to obtain multiple feature maps corresponding to the image to be processed. Then, a feature extraction network configured with only an exponential moving average attention mechanism is used for feature extraction, so as to better capture the long-range dependencies in the image to be processed, enhance the effect of extracting global features and local features, reduce the number of model parameters and computational complexity, and finally perform detection based on the extracted global features and local features to obtain the detection result, so as to more clearly extract the seal and signature. Then, the background color interference in the detection result is removed using the color range to obtain the completely separated seal and / or signature.

[0094] Refer to Figure 2 , in some embodiments, before S102, the method may further include:

[0095] S201, obtain an image training sample set;

[0096] S202, perform model training using the image training sample set to obtain a target detection model.

[0097] In this embodiment, in order to enable the target detection model to better perform image processing on the image to be processed, the target detection model needs to be trained. Therefore, the target detection model can be obtained through the image training sample set and model training using the image training samples. Among them, the image training sample set includes multiple image training samples, and each image training sample includes a sample image and an image marking frame corresponding to the sample image; the sample image includes a first seal image and a first signature image; the image marking frame is used to represent the position areas of the first seal image and the first signature image.

[0098] Among them, it is worth noting that the image training sample set can be a data set of contract scanned copies or a data set of contract photos. The image training sample set covers various scenarios such as the overlap of seals and signatures and background noise interference. Then, the positions of the seals and signatures in the image sample set are marked using an image annotation tool to obtain the image training sample set.

[0099] Referring to Figure 3 , in some embodiments, S202 may include:

[0100] S2021, input the sample image into a preset detection model to obtain a predicted target detection box;

[0101] S2022, determine the first loss function of the preset detection model according to the predicted target detection box and the image marking box;

[0102] S2023, in the case that the first loss function does not meet the preset convergence condition, adjust the model parameters of the preset detection model, and return to the step of inputting the sample image into the preset detection model to obtain a predicted target detection box until the first loss function meets the preset convergence condition.

[0103] In this embodiment, in order to enable the trained target detection model to better perform image processing, the preset detection model adopts the YOLOV8 model, so that the trained target detection model has the ability of image recognition and processing. During the training process, the sample image is input into the preset detection model, and the preset detection model continuously captures the long-distance dependence relationships in the image and the target features at different scales, improving the ability of the preset detection model to learn seals and signatures of different sizes, obtaining a predicted target detection box, then determining the first loss function according to the target detection box and the image marking box, and then adjusting the model parameters to make the first loss function converge continuously to complete the training of the preset detection model.

[0104] In some other embodiments, in order to adjust the model parameters more quickly, S2023 may specifically include:

[0105] Perform exponential moving average smoothing on the model parameters to adjust the model parameters.

[0106] In this embodiment, when adjusting the model parameters, the model parameters can be adjusted by exponential moving average smoothing, that is, the weights in the model parameters are weighted averaged by exponential moving average to achieve smoothing of the weights, making the weights more continuous and stable before and after adjustment, reducing the fluctuations of the weight changes, and thus better realizing the training of the preset detection model.

[0107] In some other embodiments, the preset detection model may include: a feature transformation network, a feature extraction network, and an object detection network. The number of pooling layers in the feature transformation network is greater than the preset number of pooling layers, and the convolution kernel size of the preset detection model is smaller than the preset convolution kernel size.

[0108] In this embodiment, the feature transformation network may be a lightweight CloFormer network, which can use the self-attention mechanism to enable the original YOLOV8 model to better capture long-range dependencies in images. At the same time, it can reduce the number of model parameters and computational complexity, making the model more efficient and suitable for running on devices with limited resources.

[0109] The feature extraction network includes a global feature processing unit and a local feature processing unit. The global feature processing unit is configured with an exponential moving average attention mechanism.

[0110] In this embodiment, the downsampling combined with the attention mechanism used in the global feature processing unit of the original YOLOV8 can be replaced with the EMA attention mechanism. Using the EMA attention mechanism can more effectively capture target features at different scales, which can promote the ability of the preset detection model to learn seals and signatures of different sizes.

[0111] The object detection network includes a feature fusion sub-network and a feature classification sub-network;

[0112] Based on the above structural adjustment of the preset detection model, S2021 may specifically include:

[0113] Performing feature transformation on the sample image through the feature transformation network to obtain multiple sample feature maps;

[0114] Performing global feature extraction on the multiple sample feature maps through the global feature processing unit to obtain multiple predicted global features;

[0115] Performing local feature extraction on the multiple sample feature maps through the local feature processing unit to obtain multiple predicted local features;

[0116] Performing feature fusion on the multiple predicted global features and the multiple predicted local features through the feature fusion sub-network to obtain multiple predicted candidate detection frames;

[0117] Performing confidence calculation on the multiple candidate predicted detection frames through the feature classification sub-network to obtain the predicted confidence of each predicted candidate detection frame;

[0118] Outputting the predicted object detection frame through the feature classification sub-network. The predicted object detection frame includes the predicted candidate detection frame corresponding to the predicted confidence greater than the preset confidence threshold.

[0119] In this embodiment, during the training of the preset detection model, since the preset detection model is optimized, all networks of the preset detection model need to be trained during the training process. First, the feature transformation network transforms the sample image, that is, converts the sample image into multiple sample feature maps, and then inputs the multiple feature maps into the global feature processing unit and the local feature processing unit. The local feature processing unit extracts high-frequency local features through convolution operations, and the global processing unit extracts global features through the configured exponential moving average attention mechanism. Then, the extracted global features and local features are fused in the feature fusion sub-network to further strengthen the context information, thereby obtaining multiple predicted candidate detection frames. To improve the recognition accuracy, in the feature classification sub-network, the confidence of each candidate predicted detection frame is calculated to obtain the predicted confidence of each predicted candidate detection frame. Then, the candidate predicted detection frames are screened through a preset confidence threshold to obtain the target detection frame, that is, the predicted target detection frame is the predicted candidate detection frame whose predicted confidence is greater than the preset confidence threshold.

[0120] In this embodiment, in order to reduce the noise and fluctuations during the training process and reduce the computational complexity of the model, pruning operations can be performed on each network, that is, reducing the number of model parameters by reducing the size of the convolution kernel. At the same time, a pooling layer can be added to reduce the size of the feature map and reduce the computational complexity, so as to improve the detection effect while reducing the computational consumption of the original preset detection model.

[0121] In some other embodiments, according to the computing resources of the electronic device, the batch processing size during model training can be set to 64, the initial value of the learning rate can be set to 0.001, the training period can be set to 200, the weight decay can be set to 0.0005, and the optimizer can be selected as Adam, with other parameters being default values.

[0122] In some other embodiments, since during the model training process, the edge of the predicted candidate box output by the model may exactly coincide with the edge of the seal, resulting in part of the seal being missing and incomplete extraction, before the predicted target detection box is output through the feature classification sub-network, the method may further include:

[0123] Expanding the predicted target detection box outward by a preset number of pixels to obtain an expanded predicted detection box;

[0124] Outputting the predicted target detection box through the feature classification sub-network includes:

[0125] Outputting the expanded predicted target detection box through the feature classification sub-network.

[0126] In this embodiment, the predicted target detection box can be expanded outward by a preset number of pixels to obtain an expanded predicted detection box, so as to expand the area of the predicted target detection box, thereby reducing the possibility of the predicted target detection box overlapping with the seal.

[0127] As an example, the preset number of pixels can be 5 pixels, and the preset confidence threshold can be 0.7.

[0128] In some other embodiments, after the model is trained, a test set can be used to test the target detection model, and the test method adopts a conventional test method, which will not be elaborated here.

[0129] Referring to Figure 4 , in some embodiments, after the model is trained by adopting the above method, S103 can specifically include:

[0130] S1031, the global feature extraction sub-network performs global feature extraction on multiple feature maps to obtain multiple global features;

[0131] S1032, the local feature extraction sub-network performs local feature extraction on multiple feature maps to obtain multiple local features.

[0132] In this embodiment, after the model is trained, the feature extraction network can include a global feature sub-network and a local feature sub-network, that is, the global feature sub-network is the global feature processing unit after training, and the local feature sub-network is the local feature processing unit after training. Then, the global feature sub-network and the local feature sub-network perform global feature extraction and local feature extraction on multiple feature maps to obtain multiple global features and multiple local features.

[0133] Referring to Figure 5 , in some other embodiments, S104 can specifically include:

[0134] S1041, the feature fusion sub-network performs feature fusion on multiple global features and multiple local features to obtain multiple first candidate detection boxes;

[0135] S1042, the feature classification sub-network calculates the confidence of multiple first candidate detection boxes to obtain the first confidence of each first candidate detection box;

[0136] S1043, the feature classification sub-network outputs the detection result.

[0137] In this embodiment, when using the trained object detection model to perform image processing on the image to be processed, the object detection network of the object detection model may include a feature fusion sub-network and a feature classification sub-network, where the feature fusion sub-network is a trained feature fusion sub-network, and the feature classification sub-network is a trained feature classification sub-network.

[0138] The feature fusion sub-network performs feature fusion on multiple global features and multiple local features to obtain multiple first candidate detection frames. Then, the feature classification sub-network calculates the confidence of each first candidate detection frame to obtain the first confidence of each first candidate detection frame, and outputs the detection result through the output layer of the feature classification sub-network, where the detection result includes the first candidate detection frame corresponding to the first confidence greater than the preset confidence threshold.

[0139] At this time, the positions of the seal and signature in the image to be processed can be initially determined through the object detection network.

[0140] In some embodiments, in S105, after the positions of the seal and signature are determined by the object detection model, the seal and signature can be separated by a preset color gamut range, that is, pixel-level separation of the seal and signature is achieved based on the RGB color gamut of the red and blue seals, where the red color gamut is [0, 0, 100] to [100, 100, 255], and the blue color gamut is [100, 0, 0] to [255, 100, 100].

[0141] Refer to Figure 6 , in some other embodiments, in order to better process the image in the future, after S105, the method may further include:

[0142] S601, input the seal and / or signature into the backbone extraction module in the image enhancement model to extract features of the seal and / or signature to obtain multiple features to be enhanced, and the features to be enhanced are used to represent the detailed features of the attributes of the seal and / or signature;

[0143] S602, input the multiple features to be enhanced into the separation and distillation module in the image enhancement model to perform detail enhancement on the features to be enhanced to obtain the seal and / or signature after detail enhancement.

[0144] In this embodiment, in order to improve the accuracy of the OCR (Optical Character Recognition) method, after obtaining the seal and signature, the image enhancement model can be used to perform detail enhancement on the seal and / or signature to obtain the seal and / or signature after detail enhancement, thereby improving the recognition accuracy of the OCR method.

[0145] Specifically, the image enhancement model selects the BSRN (Backbone Super-Resolution Network) model. Among them, the image enhancement model can include a backbone extraction module and a separation and distillation module. The backbone extraction module extracts high-level features of the seal and / or signature to obtain multiple features to be enhanced, which can include structural information, edge information, and texture information. Then, the features to be enhanced are input into the separation and distillation module for detail enhancement.

[0146] Among them, the number of separation and distillation modules is less than a preset number.

[0147] It should be noted that since the seals and signatures identified by the object detection model belong to small-size images with few levels of details, excessive separation and distillation modules will lead to overfitting. Therefore, to reduce the possibility of overfitting, the number of separation and distillation modules can be set to be less than the preset number, simplifying the model structure, reducing the number of parameters, thereby reducing the computational complexity of the image enhancement model. While reducing the consumption of computing resources, the model can better generalize to new data.

[0148] Refer to Figure 7 , in some other embodiments, the separation and distillation module can include a feature distillation unit, a feature concentration unit, and a feature enhancement unit. S602 can specifically include:

[0149] S6021, performing feature distillation on multiple features to be enhanced through the feature distillation unit to obtain multiple first key features;

[0150] S6022, performing feature concentration on the first key features through the feature concentration unit to obtain second key features;

[0151] S6023, performing feature enhancement on the second key features through the feature enhancement unit to obtain the seal and / or signature with enhanced details.

[0152] In this embodiment, when processing the seal and signature, the feature distillation unit can be used to perform feature distillation on the features to be enhanced to obtain multiple first key features, where the first key features are used to represent the association information between the features to be enhanced. Then, feature concentration is performed on the first key features to obtain second key features, where the second key features are related to the first key features after enhanced feature expression. Then, feature enhancement is performed on the second key features to obtain the seal and / or signature with enhanced details. Through the above feature distillation, feature concentration, and feature enhancement, the feature information of the seal and / or signature can be obtained to better enhance the details of the seal and / or signature.

[0153] In some other embodiments, the feature enhancement unit may include a convolutional channel subunit and a super-resolution subunit; S6023 may specifically include:

[0154] Extract differential features from the second key feature through the convolutional channel subunit to obtain differential features;

[0155] Enhance the details of the differential features to obtain enhanced differential features;

[0156] Restore the enhanced differential features through the super-resolution subunit to obtain a seal and / or signature with enhanced details.

[0157] In this embodiment, a contrast channel perception attention subunit is provided in the feature enhancement unit of the original BSRN model to enhance the model's ability to capture image features in the channel dimension. However, since the information interaction between the seal and / or signature channels is not complex, the original contrast channel perception attention subunit is adjusted to a convolutional channel subunit. The convolutional channel subunit extracts differential features from the second key feature to obtain differential features, and then the super-resolution subunit is used to restore the enhanced differential features to obtain a seal and / or signature with enhanced details.

[0158] It should be noted that the training method of the image enhancement model is the same as that of the above-mentioned object detection model, and will not be elaborated here.

[0159] In this embodiment, in the electronic device, a communication interface for providing services externally can be set. The inputs of the communication interface are all the images to be processed and the file paths. They are deployed on the background server for automatic processing and the results are returned. According to the different numbers of seals and signatures in the images to be processed, the images after intercepting the seals and signatures, the images after pixel separation, and the images with enhanced details can be returned according to the corresponding communication interfaces respectively and displayed on the electronic device.

[0160] Based on the image processing method provided in the above embodiments, correspondingly, the present application also provides a specific implementation manner of an image processing device. Please refer to the following embodiments.

[0161] First, refer to Figure 8 , the image processing device 800 provided in the embodiment of the present application may include:

[0162] An acquisition module 801, configured to acquire an image to be processed, where the image to be processed includes a seal and a signature;

[0163] A conversion module 802, configured to input the image to be processed into the object detection model, and use the feature conversion network of the object detection model to perform feature conversion on the image to be processed to obtain a plurality of feature maps;

[0164] An extraction module 803 is configured to extract global features and local features of multiple feature maps by using a feature extraction network of an object detection model, wherein the feature extraction network is configured with an exponential moving average attention mechanism;

[0165] A detection module 804 is configured to detect seals and signatures in an image to be processed based on the global features and local features by using an object detection network of the object detection model, and obtain a detection result;

[0166] A pixel separation module 805 is configured to perform pixel separation on the seals and signatures in the detection result according to a preset color gamut range, so as to obtain seals and / or signatures.

[0167] As an optional implementation manner, the conversion module 802 may further be specifically configured to:

[0168] Obtain an image training sample set, where the image training sample set includes multiple image training samples, and each image training sample includes a sample image and an image marking frame corresponding to the sample image; the sample image includes a first seal image and a first signature image; the image marking frame is used to represent the position area of the first seal image and the first signature image;

[0169] Perform model training by using the image training sample set to obtain an object detection model.

[0170] As an optional implementation manner, the conversion module 802 may further be specifically configured to:

[0171] Input the sample image into a preset detection model to obtain a predicted object detection frame;

[0172] Determine a first loss function of the preset detection model according to the predicted object detection frame and the image marking frame;

[0173] In the case that the first loss function does not meet the preset convergence condition, adjust the model parameters of the preset detection model, and return to the step of inputting the sample image into the preset detection model to obtain a predicted object detection frame until the first loss function meets the preset convergence condition.

[0174] As an optional implementation manner, the conversion module 802 may further be specifically configured to:

[0175] Perform exponential moving average smoothing processing on the model parameters to adjust the model parameters.

[0176] As an alternative implementation, the preset detection model includes: a feature transformation network, a feature extraction network, and an object detection network. The number of pooling layers in the feature transformation network is greater than the preset number of pooling layers, and the convolution kernel size of the preset detection model is smaller than the preset convolution kernel size. The feature extraction network includes a global feature processing unit and a local feature processing unit, and the global feature processing unit is configured with an exponential moving average attention mechanism. The object detection network includes a feature fusion sub-network and a feature classification sub-network.

[0177] The conversion module 802 can also be specifically used for:

[0178] Performing feature transformation on the sample image through the feature transformation network to obtain multiple sample feature maps;

[0179] Performing global feature extraction on the multiple sample feature maps through the global feature processing unit to obtain multiple predicted global features;

[0180] Performing local feature extraction on the multiple sample feature maps through the local feature processing unit to obtain multiple predicted local features;

[0181] Performing feature fusion on the multiple predicted global features and the multiple predicted local features through the feature fusion sub-network to obtain multiple predicted candidate detection frames;

[0182] Calculating the confidence of the multiple candidate predicted detection frames through the feature classification sub-network to obtain the predicted confidence of each predicted candidate detection frame;

[0183] Outputting the predicted object detection frame through the feature classification sub-network, where the predicted object detection frame includes the predicted candidate detection frame corresponding to the predicted confidence greater than the preset confidence threshold.

[0184] As an alternative implementation, the conversion module 802 can also be used for:

[0185] Expanding the predicted object detection frame outward by a preset number of pixels to obtain an expanded predicted detection frame;

[0186] Outputting the expanded predicted object detection frame through the feature classification sub-network.

[0187] As an alternative implementation, the feature extraction network includes a global feature sub-network and a local feature sub-network, where the global feature processing unit is configured with an exponential moving average attention mechanism. The extraction module 803 can also be specifically used for:

[0188] The global feature extraction sub-network performs global feature extraction on the multiple feature maps to obtain multiple global features;

[0189] The local feature extraction sub-network performs local feature extraction on the multiple feature maps to obtain multiple local features.

[0190] As an alternative implementation, the object detection network includes a feature fusion sub-network and a feature classification sub-network; the detection module 804 can be specifically configured to:

[0191] Perform feature fusion on multiple global features and multiple local features through the feature fusion sub-network to obtain multiple first candidate detection frames;

[0192] Calculate the confidence of multiple first candidate detection frames through the feature classification sub-network to obtain the first confidence of each first candidate detection frame;

[0193] Output the detection result through the feature classification sub-network, where the detection result includes the first candidate detection frames corresponding to the first confidence greater than the preset confidence threshold.

[0194] As an alternative implementation, the pixel separation module 805 can also be specifically configured to:

[0195] Input the seal and / or signature into the backbone extraction module in the image enhancement model to extract features of the seal and / or signature, obtaining multiple features to be enhanced, where the features to be enhanced are detailed features representing the attributes of the seal and / or signature;

[0196] Input the multiple features to be enhanced into the separation and distillation module in the image enhancement model to perform detail enhancement on the features to be enhanced, obtaining the seal and / or signature after detail enhancement; where the number of separation and distillation modules is less than the preset number.

[0197] As an alternative implementation, the separation and distillation module includes a feature distillation unit, a feature concentration unit, and a feature enhancement unit; the pixel separation module 805 can also be specifically configured to:

[0198] Perform feature distillation on multiple features to be enhanced through the feature distillation unit to obtain multiple first key features, where the first key features are used to represent the association information between the features to be enhanced;

[0199] Perform feature concentration on the first key features through the feature concentration unit to obtain second key features, where the second key features are the first key features after enhanced feature expression;

[0200] Perform feature enhancement on the second key features through the feature enhancement unit to obtain the seal and / or signature after detail enhancement.

[0201] As an alternative implementation, the feature enhancement unit includes a convolutional channel sub-unit and a super-resolution sub-unit; the pixel separation module 805 can also be specifically configured to:

[0202] Extract differential features from the second key features through the convolutional channel sub-unit to obtain differential features;

[0203] Perform detail enhancement on the differential features to obtain enhanced differential features;

[0204] Perform image restoration on the enhanced differential features through a super-resolution subunit to obtain a seal and / or signature with enhanced details.

[0205] Figure 9 The figure shows a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application.

[0206] The electronic device may include a processor 901 and a memory 902 storing computer program instructions.

[0207] Specifically, the above-mentioned processor 901 may include a Central Processing Unit (CPU), or an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.

[0208] The memory 902 may include a mass memory for data or instructions. By way of example and not limitation, the memory 902 may include a Hard Disk Drive (HDD), a floppy disk drive, a flash memory, an optical disc, a magneto-optical disc, a magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. In one example, the memory 902 may include removable or non-removable (or fixed) media, or the memory 902 is a non-volatile solid-state memory. The memory 902 may be inside or outside the integrated gateway disaster recovery device.

[0209] In one example, the memory 902 may be a Read Only Memory (ROM). In one example, the ROM may be a mask-programmed ROM, a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically Erasable PROM (EEPROM), an Electrically Rewritable ROM (EAROM), or a flash memory, or a combination of two or more of these.

[0210] The memory 902 may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage media device, an optical storage media device, a flash memory device, an electrical, optical, or other physical / tangible memory storage device. Thus, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the image processing method according to the first aspect of the present disclosure.

[0211] The processor 901 reads and executes the computer program instructions stored in the memory 902 to implement Figure 1 an image processing method in the illustrated embodiment.

[0212] In one example, the electronic device may further include a communication interface 903 and a bus 904. Among them, as Figure 9 shown, the processor 901, the memory 902, and the communication interface 903 are connected through the bus 904 and complete communication with each other.

[0213] The communication interface 903 is mainly used to implement communication between each module, device, unit, and / or device in the embodiments of the present application.

[0214] The bus 904 includes hardware, software, or both, and couples the components of the electronic device to each other. By way of example and not limitation, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses or a combination of two or more of these. In a suitable case, the bus 504 may include one or more buses. Although the embodiments of the present application describe and illustrate a specific bus, the present application contemplates any suitable bus or interconnect.

[0215] The electronic device can execute the image processing method in the embodiments of the present application, thereby implementing in combination with Figures 1 - 8The described image processing method and apparatus.

[0216] In addition, in combination with the image processing method in the above embodiments, an embodiment of the present application can provide a computer storage medium to implement. Computer program instructions are stored on the computer storage medium; when the computer program instructions are executed by a processor, any one of the image processing methods in the above embodiments is implemented.

[0217] In an optional embodiment, in combination with the image processing method in the above embodiments, an embodiment of the present application can provide a computer program product to implement. The instructions in the computer program product are executed by a processor of an electronic device, so that the electronic device can implement any one of the image processing methods in the above embodiments.

[0218] It should be clear that the present application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present application.

[0219] It should also be noted that the functional blocks shown in the above structural block diagrams can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, etc. When implemented in software, the elements of the present application are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted via a data signal carried in a carrier wave on a transmission medium or a communication link. A "machine-readable medium" can include any medium that can store or transmit information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical discs, hard disks, fiber optic media, radio frequency (RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.

[0220] It should also be noted that the exemplary embodiments mentioned in the present application describe some methods or systems based on a series of steps or devices. However, the present application is not limited to the order of the above steps. That is, the steps can be executed in the order mentioned in the embodiments, or different from the order in the embodiments, or several steps can be executed simultaneously.

[0221] It should be clear that the present application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present application.

[0222] The functional blocks shown in the above structural block diagrams can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, and so on. When implemented in software, the elements of the present application are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted via a data signal carried in a carrier wave over a transmission medium or a communication link. A "machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical discs, hard disks, fiber optic media, radio frequency (RF) links, and so on. The code segment can be downloaded via a computer network such as the Internet, an intranet, and so on.

[0223] It should also be noted that the exemplary embodiments mentioned in the present application describe some methods or systems based on a series of steps or devices. However, the present application is not limited to the order of the above steps, that is, the steps can be executed in the order mentioned in the embodiments, or different from the order in the embodiments, or several steps can be executed simultaneously.

[0224] Aspects of the present disclosure have been described above with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block in the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing devices to produce a machine such that the instructions executed by the processor of the computer or other programmable data processing devices enable the implementation of the functions / actions specified in one or more blocks of the flowcharts and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and the combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by dedicated hardware that performs the specified functions or actions, or by a combination of dedicated hardware and computer instructions.

[0225] As described above, this is only the specific implementation manner of the present application. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein. It should be understood that the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application.

Claims

1. An image processing method, characterized in that, Including: Obtain an image to be processed, where the image to be processed includes a seal and a signature; Input the image to be processed into a target detection model, and use the feature transformation network of the target detection model to perform feature transformation on the image to be processed to obtain multiple feature maps; Use the feature extraction network of the target detection model to extract the global features and local features of the multiple feature maps, where the feature extraction network is configured with an exponential moving average attention mechanism; Use the target detection network of the target detection model to detect the seal and signature in the image to be processed based on the global features and the local features to obtain a detection result; Perform pixel separation on the seal and signature in the detection result according to a preset color gamut range to obtain the seal and / or signature.

2. The method according to claim 1, wherein Before the step of inputting the image to be processed into the target detection model, the method further includes: Obtain an image training sample set, where the image training sample set includes multiple image training samples, and each image training sample includes a sample image and an image marking frame corresponding to the sample image; the sample image includes a first seal image and a first signature image; the image marking frame is used to represent the position area of the first seal image and the first signature image; Use the image training sample set to perform model training to obtain the target detection model.

3. The method according to claim 2, wherein, The step of using the image training sample set to perform model training to obtain the target detection model includes: Input the sample image into a preset detection model to obtain a predicted target detection frame; Determine a first loss function of the preset detection model according to the predicted target detection frame and the image marking frame; In the case that the first loss function does not meet the preset convergence condition, adjust the model parameters of the preset detection model, and return to the step of inputting the sample image into the preset detection model to obtain a predicted target detection frame until the first loss function meets the preset convergence condition.

4. The method according to claim 3, wherein The step of adjusting the model parameters of the preset detection model includes: Perform exponential moving average smoothing processing on the model parameters to adjust the model parameters.

5. The method according to claim 3, characterized in that, The preset detection model includes a feature transformation network, a feature extraction network, and a target detection network. The number of pooling layers in the feature transformation network is greater than a preset number of pooling layers, and the convolution kernel size of the preset detection model is smaller than a preset convolution kernel size; The feature extraction network includes a global feature processing unit and a local feature processing unit, and the global feature processing unit is configured with an exponential moving average attention mechanism; The target detection network includes a feature fusion sub-network and a feature classification sub-network; The step of inputting the sample image into the preset detection model to obtain a predicted target detection frame includes: Perform feature transformation on the sample image through the feature transformation network to obtain multiple sample feature maps; Extract global features of the multiple sample feature maps through the global feature processing unit to obtain multiple predicted global features; Extract local features of the multiple sample feature maps through the local feature processing unit to obtain multiple predicted local features; Feature fusion is performed on multiple predicted global features and multiple predicted local features through the feature fusion sub-network to obtain multiple predicted candidate detection frames; Confidence calculation is performed on multiple candidate predicted detection frames through the feature classification sub-network to obtain the predicted confidence of each predicted candidate detection frame; The predicted target detection frame is output through the feature classification sub-network, and the predicted target detection frame includes the predicted candidate detection frame corresponding to the predicted confidence greater than the preset confidence threshold.

6. The method according to claim 5, wherein Before the predicted target detection frame is output through the feature classification sub-network, the method further includes: Expanding the predicted target detection frame outward by a preset number of pixels to obtain an expanded predicted detection frame; The output of the predicted target detection frame through the feature classification sub-network includes: Outputting the expanded predicted target detection frame through the feature classification sub-network.

7. The method according to any one of claims 1-6, characterized in that, The feature extraction network includes a global feature sub-network and a local feature sub-network, wherein the global feature sub-network is configured with an exponential moving average attention mechanism; The extraction of the global features and local features of multiple feature maps by using the feature extraction network of the object detection model includes: The global feature extraction sub-network extracts global features from multiple feature maps to obtain multiple global features; The local feature extraction sub-network extracts local features from multiple feature maps to obtain multiple local features.

8. The method according to claim 7, wherein The object detection network includes a feature fusion sub-network and a feature classification sub-network; The detection of the seal and signature in the image to be processed based on the global feature and the local feature by using the object detection network of the object detection model to obtain a detection result includes: Feature fusion is performed on multiple global features and multiple local features through the feature fusion sub-network to obtain multiple first candidate detection frames; Confidence calculation is performed on multiple first candidate detection frames through the feature classification sub-network to obtain the first confidence of each first candidate detection frame; The detection result is output through the feature classification sub-network, and the detection result includes the first candidate detection frame corresponding to the first confidence greater than the preset confidence threshold.

9. The method according to claim 1, wherein After pixel separation of the seal and signature in the detection result according to the preset color gamut range to obtain the seal and / or signature, the method further includes: Inputting the seal and / or the signature into the backbone extraction module in the image enhancement model to extract features of the seal and / or the signature to obtain multiple features to be enhanced, and the features to be enhanced are used to represent the detailed features of the attributes of the seal and / or the signature; Inputting multiple features to be enhanced into the separation and distillation module in the image enhancement model to perform detail enhancement on the features to be enhanced to obtain the seal and / or the signature after detail enhancement; wherein the number of the separation and distillation modules is less than the preset number.

10. The method according to claim 9, wherein The separation and distillation module includes a feature distillation unit, a feature concentration unit, and a feature enhancement unit; Inputting the multiple features to be enhanced into a separation and distillation module of the image enhancement model to perform detail enhancement on the features to be enhanced, and obtaining the seal and / or the signature after detail enhancement, includes: Performing feature distillation on the multiple features to be enhanced through the feature distillation unit to obtain multiple first key features, where the first key features are used to represent the association information between the features to be enhanced; Performing feature concentration on the first key features through the feature concentration unit to obtain second key features, where the second key features are used to represent the first key features after enhanced feature expression; Performing feature enhancement on the second key features through the feature enhancement unit to obtain the seal and / or the signature after detail enhancement.

11. The method according to claim 10, wherein The feature enhancement unit includes a convolutional channel sub-unit and a super-resolution sub-unit; The performing feature enhancement on the second key features through the feature enhancement unit to obtain the seal and / or the signature after detail enhancement, includes: Performing differential feature extraction on the second key features through the convolutional channel sub-unit to obtain differential features; Performing detail enhancement on the differential features to obtain enhanced differential features; Performing image restoration on the enhanced differential features through the super-resolution sub-unit to obtain the seal and / or the signature after detail enhancement.

12. An image processing apparatus, characterized in that, Includes: An acquisition module, configured to acquire a to-be-processed image, where the to-be-processed image includes a seal and a signature; A conversion module, configured to input the to-be-processed image into a target detection model, and perform feature conversion on the to-be-processed image through a feature conversion network of the target detection model to obtain multiple feature maps; An extraction module, configured to extract global features and local features of the multiple feature maps through a feature extraction network of the target detection model, where the feature extraction network is configured with an exponential moving average attention mechanism; A detection module, configured to detect the seal and the signature in the to-be-processed image based on the global features and the local features through a target detection network of the target detection model to obtain a detection result; A pixel separation module, configured to perform pixel separation on the seal and the signature in the detection result according to a preset color gamut range to obtain the seal and / or the signature.

13. An electronic device, characterized in that, The device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, the image processing method according to any one of claims 1-11 is implemented.

14. A computer-readable storage medium, characterized in that, Computer program instructions are stored on a computer-readable storage medium, and when the computer program instructions are executed by a processor, the image processing method according to any one of claims 1-11 is implemented.

15. A computer program product, characterized in that, When instructions in the computer program product are executed by a processor of an electronic device, the electronic device is caused to execute the image processing method according to any one of claims 1-11.