Seal segmentation method, device and computer equipment based on UNet-S network structure model
By introducing multi-scale depth separation residual modules and jump connections into the U-Net network structure model, the problem of insufficient seal edge segmentation accuracy in the Republic of China archives is solved, and high-precision and efficient seal segmentation effect is achieved.
Patent Information
- Application Number
- CN202210566573.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-23
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-05-23
AI Technical Summary
In the prior art, image segmentation networks generally lack accuracy in the archives field of the Republic of China, and cannot effectively segment the edge of the seal.
The seal segmentation method based on the U-Net network structure model is adopted, and the residual module and jump connection mechanism can be separated by multi-scale depth, multi-scale features are extracted and up-down sampling is performed, and high-precision seal segmentation diagram is output.
It realizes the precise segmentation of the edges of seals in archives of the Republic of China, improves the processing accuracy and speed, and solves the shortcomings of traditional methods in noise and edge segmentation.
Smart Images

Figure CN115205862B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of seal segmentation, and in particular to a seal segmentation method, device and computer equipment based on a U-Net network structure model. Background Art
[0002] The digitization and dataization of the archives of the Republic of China are of great significance. On the one hand, the archives of the Republic of China are objective records of modern historical activities, occupy a special position in the national archives, and are extremely valuable archives; on the other hand, digital processing lays the foundation for the application of artificial intelligence technology. Content-based image retrieval and image semantic annotation are two typical application scenarios of archival images, and identifying and segmenting seals in images creates favorable conditions for these two applications. The particularity of the archives of the Republic of China determines the difficulty of seal segmentation. Its particularity is mainly reflected in the following three points: (1) Challenges brought by the limitations of the times. Poor paper quality brings problems of print invasion and excessive noise, and manual brush writing brings variable fonts and different text sizes. (2) Challenges brought by complex layout structure. There are archive pages that only contain seals, there are also archive pages composed of seals or even multiple seals and tables, and there are archive pages with multiple seals overlapping. In addition, the vertical arrangement of text and the writing order from right to left also make layout analysis and understanding difficult. (3) Challenges brought by the particularity of seals. The diversity of seal shapes and the non-fixed positions lead to a lack of prior knowledge; the color features brought by the black and white seals cannot be utilized; and the uneven distribution of seals in archival data leads to an imbalance of positive and negative samples.
[0003] Traditional seal detection is based on the three main features of the image: color, texture, and shape. The color feature-based method determines the position of the seal based on the red color of the seal to achieve seal detection. Because the red color of the seal is significantly different from the color of paper and text, it occupies an important position in traditional methods. However, the archives of the Republic of China are all written in brush and black seals, so this type of method cannot be used. The texture feature-based method uses the grayscale spatial distribution law of the pixel neighborhood. The disadvantage of this method is that it needs to be combined with other image features, so it is not suitable for the processing of handwritten text archives. The shape feature-based method can accurately express the characteristics of the target features, but its anti-interference ability is weak. The Republic of China archives seal data set has a lot of noise and the seal border is seriously discontinuous, so this method is difficult to achieve the Republic of China archives seal segmentation.
[0004] Traditional methods cannot effectively solve the problem of seal segmentation in archives of the Republic of China, and deep learning methods have made it possible to solve this problem. Therefore, in the prior art, fully convolutional networks (FCN) are used for image semantic segmentation to understand and identify image content at the pixel level, and segmentation is achieved based on semantic information, which improves the processing speed while ensuring segmentation accuracy; however, image segmentation networks are mostly used in the fields of medical image segmentation and urban scene semantic segmentation, and have high segmentation accuracy in the above fields, while the particularity of archives of the Republic of China is a challenge. Therefore, image segmentation networks generally have insufficient accuracy in this field and cannot effectively segment the edges of seals. Therefore, effective segmentation of seals in archives of the Republic of China has become a problem that needs to be solved urgently. Summary of the invention
[0005] The main purpose of this application is to provide a seal segmentation method based on the U-Net network structure model, aiming to solve the technical problem in the prior art that the image segmentation network is generally insufficient in accuracy, resulting in the inability to accurately and effectively segment the seal edge.
[0006] This application proposes a seal segmentation method based on the UNet-S network structure model, including:
[0007] S1, obtaining an archive image file in a Republic of China archive data set, wherein the archive image file includes a plurality of archive images;
[0008] S2, inputting the plurality of archive images into the UNet-S network structure model, and outputting seal segmentation maps of the plurality of archive images;
[0009] The step S2 of inputting the plurality of archive images into the UNet-S network structure model and outputting the seal segmentation maps of the plurality of archive images comprises:
[0010] S21, performing a common convolution operation on the plurality of archive images to extract features and obtain a plurality of feature images, wherein the plurality of archive images have the same size;
[0011] S22, inputting the plurality of feature images into a multi-scale depth-separable residual module to perform multi-scale feature extraction on the plurality of feature images to obtain a plurality of multi-scale feature images;
[0012] S23, performing a maximum pooling operation on the plurality of multi-scale feature images to obtain a plurality of first images, and using the first images as the feature images;
[0013] S24, repeating step S22 and step S23 according to a first preset number of times to complete downsampling and obtain multiple second images;
[0014] S25, enlarging the sizes of the plurality of second images based on bilinear interpolation to obtain a plurality of third images;
[0015] S26, performing jump connection on the plurality of second images and the plurality of third images, and inputting the plurality of third images into the multi-scale depth-separable residual module to obtain a plurality of fourth images;
[0016] S27, performing a common convolution operation on the plurality of fourth images to obtain a plurality of fifth images, and using the fifth images as the second images;
[0017] S28. Repeat steps S25 to S27 according to a second preset number of times to complete upsampling and output a plurality of seal segmentation images.
[0018] Preferably, the step S22 of inputting the plurality of feature images into a multi-scale depth separable residual module to perform multi-scale feature extraction on the plurality of feature images to obtain a plurality of multi-scale feature images comprises:
[0019] S221, performing a depth-separable convolution operation and a normalization process on the multiple feature images according to the sizes of the multiple feature images, and introducing a ReLu activation function to generate multiple first multi-scale feature maps with different receptive fields;
[0020] S222, cascading the feature image and the first multi-scale feature map along the channel dimension based on a cascade operation to generate a cascade feature map;
[0021] S223, performing a depth-separable convolution operation and normalization processing on the cascade feature map, and introducing a ReLu activation function to generate a second multi-scale feature map;
[0022] S224, performing dimension upscaling or dimension downscaling processing on the feature image according to the number of output channels of the second multi-scale feature map, so that the input channel of the feature image is the same as the output channel of the second multi-scale feature map;
[0023] S225, adding the feature image that has undergone dimension increase or dimension decrease processing to the second multi-scale feature image, and outputting the addition result as a multi-scale feature image.
[0024] Preferably, the step S221 of performing a depth-wise separable convolution operation on the plurality of feature images according to sizes of the plurality of feature images comprises:
[0025] S2211, obtaining the coordinate value of the feature image;
[0026] S2212, obtaining a convolution kernel;
[0027] S2213, obtaining a bias value;
[0028] S2214, calculating the output result of the feature image according to the coordinate value of the feature image, the convolution kernel, and the offset value, wherein the calculation formula is:
[0029]
[0030] Wherein, (x, y) represents the horizontal and vertical coordinate values of the feature image in the multi-scale depth separable residual module, k represents the offset of the feature image in the horizontal coordinate, and l represents the offset of the feature image in the vertical coordinate; represents the output result of the jth feature image of this layer, W j is the convolution kernel of the jth feature image, σ is the activation function, b j is the bias value;
[0031] Preferably, in the step S22 of inputting the plurality of feature images into a multi-scale depth separable residual module to perform multi-scale feature extraction on the plurality of feature images to obtain a plurality of multi-scale feature images, the multi-scale separable residual module includes the following function:
[0032]
[0033] F′=[F n-1 ,F′];
[0034]
[0035] Among them, F n-1 is the input feature image; W is the convolution kernel, its subscript indicates the convolution layer in the module from which it comes, and the superscript indicates the convolution kernel size; b indicates the bias of the convolution layer, and its subscript indicates the convolution layer in the module from which it comes; F' and F' n is the intermediate output of the multi-scale separable residual module; F n is the output result of the multi-scale separable residual module; is the convolution operation; σ is the activation function.
[0036] Preferably, the step of obtaining the archive image file in the archive data set of the Republic of China, wherein the archive image file includes a plurality of archive images, further comprises:
[0037] S11. Mark the positions of the seal parts of the multiple archive images to obtain a JSON file corresponding to the positions of the seal parts and a mask image of each archive image.
[0038] Preferably, the loss function of the UNet-S network structure model is a BCEDiceLoss function, and the BCEDiceLoss function is a fusion of a BCELoss function and a Sigmoid function;
[0039] The BCELoss function is:
[0040] L BCE = -tlogp-(1-t)log(1-p);
[0041] Among them, t is the label value of the mask image of the archive image, and p is the probability predicted by the UNet-S network structure model for the category label;
[0042] The Sigmoid function is:
[0043]
[0044] Among them, TP represents true positive, TN represents true negative, FP represents false positive, and FN represents false negative;
[0045] L BCEDiceLoss =λL BCE +L Sigmoid ;
[0046] Among them, λ is the weight factor of the loss function.
[0047] Preferably, after the step S11 of marking the positions of the seal parts of the plurality of archive images to obtain the JSON file corresponding to the positions of the seal parts and the mask image of each archive image, the method further comprises:
[0048] S12, calculating the accuracy of each pixel in the seal segmentation map according to the mask map and the seal segmentation map, and outputting the calculation result.
[0049] The present application also provides a seal segmentation device based on the UNet-S network structure model, comprising:
[0050] An acquisition module, used to acquire an archive image file in the archive data set of the Republic of China, wherein the archive image file includes a plurality of archive images;
[0051] An input module, used for inputting a plurality of the archive images into the UNet-S network structure model, and outputting seal segmentation maps of the plurality of archive images;
[0052] Wherein, the input module includes:
[0053] A first common convolution unit is used to perform a common convolution operation on the plurality of archive images to extract features and obtain a plurality of feature images, wherein the plurality of archive images have the same size;
[0054] A multi-scale feature extraction unit, used for inputting the plurality of feature images into a multi-scale depth separable residual module to perform multi-scale feature extraction on the plurality of feature images to obtain a plurality of multi-scale feature images;
[0055] A pooling unit, configured to perform a maximum pooling operation on the plurality of multi-scale feature images to obtain a plurality of first images, and use the first images as the feature images;
[0056] A down-sampling unit, configured to repeat step S22 and step S32 according to a first preset number of times to complete down-sampling and obtain a plurality of second images;
[0057] an enlarging unit, configured to enlarge the sizes of the plurality of second images based on bilinear interpolation to obtain a plurality of third images;
[0058] a jump connection unit, configured to perform jump connections between the plurality of second images and the plurality of third images, and input the plurality of third images into the multi-scale depth-separable residual module to obtain a plurality of fourth images;
[0059] a second common convolution unit, configured to perform a (1*1) common convolution operation on the plurality of fourth images to obtain a plurality of fifth images, and use the fifth images as the second images;
[0060] The up-sampling unit is used to repeat step S25 to step S27 according to a second preset number of times to complete up-sampling and output a plurality of seal segmentation images.
[0061] The present application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the seal segmentation method based on the UNet-S network structure model are implemented.
[0062] The present application also provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the seal segmentation method based on the UNet-S network structure model are implemented.
[0063] The beneficial effects of the present application are: in order to better utilize the shallow information of the UNet-S network and fuse the feature images, the UNet-S network structure model is 5 layers. In the UNet-S network structure model, after the ordinary convolution operation is performed on the archive image, the feature image is input into the multi-scale depth separable residual module, so that the UNet-S network can not only effectively fuse semantic information and edge information, reduce the noise of the feature image, reduce network parameters, but also effectively avoid network degradation, gradient explosion and other problems. The seal segmentation map output by the Unet-S network has a high accuracy and can achieve the technical effect of effectively segmenting the archives of the Republic of China. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 This is a flow chart of a seal segmentation method based on the UNet-S network structure model according to an embodiment of the present application.
[0065] Figure 2 This is a schematic diagram of a method flow in which a plurality of archival images are input into a UNet-S network structure model and seal segmentation maps of the plurality of archival images are output in a seal segmentation method process based on a UNet-S network structure model in an embodiment of the present application.
[0066] Figure 3 A schematic diagram of the internal structure of a computer device according to an embodiment of the present application.
[0067] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0068] It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0069] like Figure 1-Figure 3 As shown, this application proposes a seal segmentation method based on the UNet-S network structure model, including:
[0070] S1, obtaining an archive image file in a Republic of China archive data set, wherein the archive image file includes a plurality of archive images;
[0071] S2, inputting the plurality of archive images into the UNet-S network structure model, and outputting seal segmentation maps of the plurality of archive images;
[0072] The step S2 of inputting the plurality of archive images into the UNet-S network structure model and outputting the seal segmentation maps of the plurality of archive images comprises:
[0073] S21, performing a (3*3) common convolution operation on the plurality of archive images to extract features and obtain a plurality of feature images, wherein the plurality of archive images have the same size;
[0074] S22, inputting the plurality of feature images into a multi-scale depth-separable residual module to perform multi-scale feature extraction on the plurality of feature images to obtain a plurality of multi-scale feature images;
[0075] S23, performing a maximum pooling operation on the plurality of multi-scale feature images to obtain a plurality of first images, and using the first images as the feature images;
[0076] S24, repeating step S22 and step S32 according to a first preset number of times to complete downsampling and obtain multiple second images;
[0077] S25, enlarging the sizes of the plurality of second images based on bilinear interpolation to obtain a plurality of third images;
[0078] S26, performing jump connection on the plurality of second images and the plurality of third images, and inputting the plurality of third images into the multi-scale depth-separable residual module to obtain a plurality of fourth images;
[0079] S27, performing a (1*1) common convolution operation on the plurality of fourth images to obtain a plurality of fifth images, and using the fifth images as the second images;
[0080] S28. Repeat steps S25 to S27 according to a second preset number of times to complete upsampling and output a plurality of seal segmentation images.
[0081] As described in the above steps S1 and S2, the UNet-S network is composed of a U-shaped structure combining downsampling and upsampling. The main contributions in the UNet-S network are downsampling, upsampling, and jump connections. In the downsampling stage, the input image is subjected to alternating and stacking convolution and pooling operations to extract abstract features. In the upsampling stage, the abstract features are deconvolved. In the jump connection part, the deep semantic information and the shallow position information are connected together by depth to restore the fine-grained details of the target object. Since the archival image files have large noise when they are input into the UNet-S network structure model, in the UNet-S network structure model, after ordinary convolution operations are performed on the archival images, in order to reduce the noise of the feature images, the feature images can be input into the multi-scale deeply separable residual module. In this way, the UNet-S network can not only effectively integrate the advantages of semantic information and edge information, thereby reducing the noise of the feature images, but also effectively avoid problems such as network degradation and gradient explosion. In order to better utilize the shallow information of the UNet-S network and fuse the feature images, the first preset number and the second preset number are 5, that is, the UNet-S network structure model has 5 layers; specifically, in the encoder stage, firstly, ordinary convolution operations are performed on multiple archival images to extract features and obtain multiple feature images, and then the feature images are input into the multi-scale deeply separable residual module to obtain multiple multi-scale feature images, and then the multiple images are input into the multi-scale deeply separable residual module to obtain multiple multi-scale feature images. The multi-scale feature image is subjected to the maximum pooling operation, so that the image size of the multiple multi-scale feature images becomes half of the original, and then steps S22 and S32 are repeated four times, so that the feature image is alternately subjected to residual learning and maximum pooling operations to complete downsampling; then the decoder stage is entered through 1*1 ordinary convolution, and in the decoder stage, bilinear interpolation with a 2-fold magnification factor can be first used as upsampling, so as to enlarge the image size of the multiple second images, so that the image size of the second image is doubled, and then the second image and the third image are jump-learned, so that the second image and the third image can be feature-fused to obtain multiple fourth images, and ordinary convolution operations are performed on the multiple fourth images, and steps S25-S27 are repeated four times to complete the upsampling operation, and finally multiple seal segmentation maps are output, so that the output seal segmentation map has a high precision, and the technical effect of effectively segmenting the archives of the Republic of China can be achieved. It should be noted that the "multiple" in this application is not an actually determined value, and any number of images can be put in according to actual needs, for example, hundreds of images, thousands of images or tens of thousands of images.
[0082] In one embodiment, the step S22 of inputting the plurality of feature images into a multi-scale depth separable residual module to perform multi-scale feature extraction on the plurality of feature images to obtain a plurality of multi-scale feature images includes:
[0083] S221, performing a depth-separable convolution operation and a normalization process on the multiple feature images according to the sizes of the multiple feature images, and introducing a ReLu activation function to generate multiple first multi-scale feature maps with different receptive fields;
[0084] S222, cascading the feature image and the first multi-scale feature map along the channel dimension based on a cascade operation to generate a cascade feature map;
[0085] S223, performing a depth-separable convolution operation and normalization processing on the cascade feature map, and introducing a ReLu activation function to generate a second multi-scale feature map;
[0086] S224, performing dimension upscaling or dimension downscaling processing on the feature image according to the number of output channels of the second multi-scale feature map, so that the input channel of the feature image is the same as the output channel of the second multi-scale feature map;
[0087] S225, adding the feature image that has undergone dimension increase or dimension decrease processing to the second multi-scale feature image, and outputting the addition result as a multi-scale feature image.
[0088] As described in the above steps S221-S225, in order to more effectively extract features on the feature image, the present embodiment adds a multi-scale feature extraction block to the residual module, that is, generates a multi-scale depth-separable residual module. Specifically, after feature extraction of the same size of the archive image, feature images of different sizes can be generated. Since the sizes of the feature images are different, a 3*3 depth-separable convolution operation can be performed on multiple feature images, batch normalization processing can be performed, and a ReLu activation function can be introduced to generate multiple first multi-scale feature maps with different receptive fields, wherein the first multi-scale feature maps of the receptive fields generated by feature images of different sizes are different. For example, the feature image can be divided into a large image, a medium image, and a small image according to a preset image scale rule, so that a 1*1 receptive field is used for the small image, a 3*3 receptive field is used for the medium image, and a 5*5 receptive field is used for the large image; then the feature image and the first multi-scale feature map are cascaded along the channel dimension based on a cascade operation to generate a cascade feature map, and a 3*3 depth-separable convolution operation is performed on the cascade feature map. Separable convolution operation, normalization processing, and introduction of ReLu activation function to generate a second multi-scale feature map; since the receptive field of two 3*3 depth separable convolutions is the same as the receptive field of a 5*5 degree separable convolution, and the cascaded cascade feature map contains feature maps with a receptive field of 3*3 and a receptive field of 1*1, after performing a 3*3 depth separable convolution on it, a second multi-scale feature map with a receptive field of 5*5 and 3*3 can be obtained, thereby realizing multi-scale feature extraction; according to the number of output channels of the second multi-scale feature map, the feature image is dimensionally upgraded or reduced so that the input channel of the feature image is the same as the output channel of the second multi-scale feature map, and finally the feature image that has been dimensionally upgraded or reduced is added to the second multi-scale feature map, and the addition result is output as a multi-scale feature image, so that the multi-scale features of the image can be effectively extracted, thereby improving the network performance. At the same time, the convolution layer in the residual module of the prior art is replaced by a depth separable convolution, which can reduce the number of network parameters.
[0089] In one embodiment, the step S221 of performing a depth-wise separable convolution operation on the plurality of feature images according to the sizes of the plurality of feature images comprises:
[0090] S2211, obtaining the coordinate value of the feature image;
[0091] S2212, obtaining a convolution kernel;
[0092] S2213, obtaining a bias value;
[0093] S2214, calculating the output result of the feature image according to the coordinate value of the feature image, the convolution kernel, and the offset value, wherein the calculation formula is:
[0094]
[0095] Wherein, (x, y) represents the horizontal and vertical coordinate values of the feature image in the multi-scale depth separable residual module, k represents the offset of the feature image in the horizontal coordinate, and l represents the offset of the feature image in the vertical coordinate; represents the output result of the jth feature image of this layer, W j is the convolution kernel of the jth feature image, σ is the activation function, b j is the bias value;
[0096] In one embodiment, in the step S22 of inputting the plurality of feature images into a multi-scale depth separable residual module to perform multi-scale feature extraction on the plurality of feature images to obtain a plurality of multi-scale feature images, the multi-scale separable residual module includes the following function:
[0097]
[0098] F′=[F n-1 ,F′].......(2);
[0099]
[0100]
[0101] Among them, F n-1 is the input feature image; W is the convolution kernel, its subscript indicates the convolution layer in the module from which it comes, and the superscript indicates the convolution kernel size; b indicates the bias of the convolution layer, and its subscript indicates the convolution layer in the module from which it comes; F' and F' n is the intermediate output of the multi-scale separable residual module; F n is the output result of the multi-scale separable residual module; is the convolution operation; σ is the activation function.
[0102] First, enter F n-1 The intermediate output F' is obtained by formula (1), and F is converted to n-1 After cascading with F′ along the channel dimension, F′ is updated, and then the features are further extracted through formula (3) to obtain the multi-scale feature map F′ n Finally, through formula (4), a 1×1 convolution is used to transform F n-1 Perform dimension reduction operations to make the number of channels equal to F′ n After they are consistent, add the two together to get the final output F n .
[0103] In one embodiment, the step of obtaining an archival image file in the Republic of China archival data set, wherein the archival image file includes a plurality of archival images, further comprises:
[0104] S11. Mark the positions of the seal parts of the multiple archive images to obtain a JSON file corresponding to the positions of the seal parts and a mask image of each archive image.
[0105] In one embodiment, in order to solve the problem of imbalance between positive and negative samples in the Republic of China archives data set, the loss function of the UNet-S network structure model is a BCEDiceLoss function that introduces a weight factor, and the BCEDiceLoss function is a fusion of the BCELoss function and the Sigmoid function;
[0106] The BCELoss function is:
[0107] L BCE = -tlogp-(1-t)log(1-p);
[0108] Among them, t is the label value of the mask image of the archive image, and p is the probability predicted by the UNet-S network structure model for the category label;
[0109] The Sigmoid function is:
[0110]
[0111] Among them, TP represents true positive, TN represents true negative, FP represents false positive, and FN represents false negative;
[0112] LBCEDiceLoss=λLBCE+LSigmoid;
[0113] Among them, λ is the weight factor of the loss function.
[0114] In the Republic of China archives dataset, half of the images are without seals, while in the other half of the images with seals, the proportion of seal pixels to total pixels is not large, resulting in the problem of imbalance between positive and negative samples. In the case of an unbalanced dataset, when λ=0.9, the problem of low accuracy of the output seal segmentation map caused by the imbalance of the dataset can be reduced.
[0115] In one embodiment, after the step S11 of marking the positions of the seal parts of the plurality of archive images to obtain the JSON file corresponding to the positions of the seal parts and the mask image of each archive image, the method further includes:
[0116] S12, calculating the accuracy of each pixel in the seal segmentation map according to the mask map and the seal segmentation map, and outputting the calculation result.
[0117] As described in the above step S12, in order to verify the segmentation accuracy of the seal segmentation map after the output of the UNet-S network structure model, after obtaining the archival image, the seal part in the archival image can be marked to obtain a mask map of each archival image, and then the accuracy of each pixel in the seal segmentation map is calculated according to the mask map and the seal segmentation map, and the calculation result is output, so that the seal segmentation accuracy value of the UNet-S network structure model can be understood according to the output calculation result; better, before the archival image is trained, the marked mask map can be placed in the UNet-S network structure model to help the UNet-S network structure model to better learn and improve the segmentation accuracy of the UNet-S network structure model.
[0118] The present application also provides a seal segmentation device based on the UNet-S network structure model, comprising:
[0119] An acquisition module, used to acquire an archive image file in the archive data set of the Republic of China, wherein the archive image file includes a plurality of archive images;
[0120] An input module, used for inputting a plurality of the archive images into the UNet-S network structure model, and outputting seal segmentation maps of the plurality of archive images;
[0121] Wherein, the input module includes:
[0122] A first common convolution unit is used to perform a common convolution operation on the plurality of archive images to extract features and obtain a plurality of feature images, wherein the plurality of archive images have the same size;
[0123] A multi-scale feature extraction unit, used for inputting the plurality of feature images into a multi-scale depth separable residual module to perform multi-scale feature extraction on the plurality of feature images to obtain a plurality of multi-scale feature images;
[0124] A pooling unit, configured to perform a maximum pooling operation on the plurality of multi-scale feature images to obtain a plurality of first images, and use the first images as the feature images;
[0125] A down-sampling unit, configured to repeat step S22 and step S32 according to a first preset number of times to complete down-sampling and obtain a plurality of second images;
[0126] an enlarging unit, configured to enlarge the sizes of the plurality of second images based on bilinear interpolation to obtain a plurality of third images;
[0127] a jump connection unit, configured to perform jump connections between the plurality of second images and the plurality of third images, and input the plurality of third images into the multi-scale depth-separable residual module to obtain a plurality of fourth images;
[0128] a second common convolution unit, configured to perform a (1*1) common convolution operation on the plurality of fourth images to obtain a plurality of fifth images, and use the fifth images as the second images;
[0129] The up-sampling unit is used to repeat step S25 to step S27 according to a second preset number of times to complete up-sampling and output a plurality of seal segmentation images.
[0130] In one embodiment, the multi-scale feature extraction unit includes:
[0131] Generate a first multi-scale feature map unit, which is used to perform a depth-separable convolution operation and a normalization process on the multiple feature images according to the sizes of the multiple feature images, and introduce a ReLu activation function to generate multiple first multi-scale feature maps with different receptive fields;
[0132] A cascade feature map generating unit, configured to cascade the feature image and the first multi-scale feature map along a channel dimension based on a cascade operation to generate a cascade feature map;
[0133] Generate a second multi-scale feature map unit, which is used to perform a depth-separable convolution operation and normalization processing on the cascade feature map, and introduce a ReLu activation function to generate a second multi-scale feature map;
[0134] a dimension up-down processing unit, configured to perform dimension up-down or dimension down-down processing on the feature image according to the number of output channels of the second multi-scale feature map, so that the input channel of the feature image is the same as the output channel of the second multi-scale feature map;
[0135] The multi-scale feature image output unit is used to add the feature image that has been subjected to dimensionality increase or dimensionality reduction processing to the second multi-scale feature image, and output the addition result as the multi-scale feature image.
[0136] In one embodiment, generating a first multi-scale feature map unit includes:
[0137] A first acquisition subunit, used to acquire the value of the feature image;
[0138] A second acquisition subunit is used to acquire a convolution kernel;
[0139] A third acquisition subunit, used to acquire a bias value;
[0140] The fourth acquisition subunit is used to calculate the output result of the feature image according to the value of the feature image, the convolution kernel, and the bias value, wherein the calculation formula is:
[0141]
[0142] Wherein, (x, y) represents the horizontal and vertical coordinate values of the feature image in the multi-scale depth separable residual module, k represents the offset of the feature image in the horizontal coordinate, and l represents the offset of the feature image in the vertical coordinate; represents the output result of the jth feature image of this layer, is the value of the feature image, W j is the convolution kernel of the jth feature image, σ is the activation function, b j is the bias value;
[0143] In one embodiment, the multi-scale separable residual module includes the following functions:
[0144]
[0145] F′=[F n-1 ,F′];
[0146]
[0147] Among them, F n-1 is the input feature image; W is the convolution kernel, its subscript indicates the convolution layer in the module from which it comes, and the superscript indicates the convolution kernel size; b indicates the bias of the convolution layer, its subscript indicates the convolution layer in the module from which it comes; F' and F n ' is the intermediate output of the multi-scale separable residual module; F n is the output result of the multi-scale separable residual module; is the convolution operation; σ is the activation function.
[0148] In one embodiment, the seal segmentation device based on the UNet-S network structure model further includes:
[0149] The annotation module is used to annotate the positions of the seal parts of multiple archive images, and obtain a JSON file corresponding to the positions of the seal parts and a mask image of each archive image.
[0150] In one embodiment, the loss function of the UNet-S network structure model is a BCEDiceLoss function, and the BCEDiceLoss function is a fusion of a BCELoss function and a Sigmoid function;
[0151] The BCELoss function is:
[0152] L BCE = -tlogp-(1-t)log(1-p);
[0153] Among them, t is the label value of the mask image of the archive image, and p is the probability predicted by the UNet-S network structure model for the category label;
[0154] The Sigmoid function is:
[0155]
[0156] Among them, TP represents true positive, TN represents true negative, FP represents false positive, and FN represents false negative;
[0157] L BCEDiceLoss =λL BCE +L Sigmoid ;
[0158] Among them, λ is the weight factor of the loss function.
[0159] In one embodiment, the seal segmentation device based on the UNet-S network structure model further includes:
[0160] The calculation module is used to calculate the accuracy of each pixel in the seal segmentation map according to the mask map and the seal segmentation map, and output the calculation result.
[0161] like Figure 3 As shown, the present application also provides a computer device, which may be a server, and its internal structure may be as shown in Figure 3 As shown. The computer device includes a processor, a memory, a display screen, an input device, a network interface and a database connected via a system bus. Among them, the processor designed by the computer is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store all data required for the process of the seal segmentation method based on the UNet-S network structure model. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the seal segmentation method based on the UNet-S network structure model is implemented.
[0162] Those skilled in the art will understand that Figure 3 The structure shown in is merely a block diagram of a portion of the structure related to the present application solution and does not constitute a limitation on the computer device to which the present application solution is applied.
[0163] An embodiment of the present application also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, any one of the above-mentioned seal segmentation methods based on the UNet-S network structure model is implemented.
[0164] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing related hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media provided in this application and used in the embodiments may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0165] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, device, article or method including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, device, article or method. In the absence of further restrictions, an element defined by the sentence "includes a ..." does not exclude the presence of other identical elements in the process, device, article or method including the element.
[0166] The above description is only a preferred embodiment of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A seal segmentation method based on the UNet-S network structure model, characterized in that: include: S1, obtaining an archive image file in a Republic of China archive data set, wherein the archive image file includes a plurality of archive images; S2, inputting the plurality of archive images into the UNet-S network structure model, and outputting seal segmentation maps of the plurality of archive images; The step S2 of inputting the plurality of archive images into the UNet-S network structure model and outputting the seal segmentation maps of the plurality of archive images comprises: S21, performing a common convolution operation on the plurality of archive images to extract features and obtain a plurality of feature images, wherein the plurality of archive images have the same size; S22, inputting the plurality of feature images into a multi-scale depth-separable residual module to perform multi-scale feature extraction on the plurality of feature images to obtain a plurality of multi-scale feature images; S23, performing a maximum pooling operation on the plurality of multi-scale feature images to obtain a plurality of first images, and using the first images as the feature images; S24, repeating step S22 and step S23 according to a first preset number of times to complete downsampling and obtain multiple second images; S25, enlarging the sizes of the plurality of second images based on bilinear interpolation to obtain a plurality of third images; S26, performing jump connection on the plurality of second images and the plurality of third images, and inputting the plurality of third images into the multi-scale depth-separable residual module to obtain a plurality of fourth images; S27, performing a common convolution operation on the plurality of fourth images to obtain a plurality of fifth images, and using the fifth images as the second images; S28. Repeat steps S25 to S27 according to a second preset number of times to complete upsampling and output a plurality of seal segmentation images.
2. The seal segmentation method based on the UNet-S network structure model according to claim 1 is characterized in that: The step S22 of inputting the plurality of feature images into a multi-scale depth separable residual module to perform multi-scale feature extraction on the plurality of feature images to obtain a plurality of multi-scale feature images includes: S221, performing a depth-separable convolution operation and a normalization process on the multiple feature images according to the sizes of the multiple feature images, and introducing a ReLu activation function to generate multiple first multi-scale feature maps with different receptive fields; S222, cascading the feature image and the first multi-scale feature map along the channel dimension based on a cascade operation to generate a cascade feature map; S223, performing a depth-separable convolution operation and normalization processing on the cascade feature map, and introducing a ReLu activation function to generate a second multi-scale feature map; S224, performing dimension upscaling or dimension downscaling processing on the feature image according to the number of output channels of the second multi-scale feature map, so that the input channel of the feature image is the same as the output channel of the second multi-scale feature map; S225, adding the feature image that has undergone dimension increase or dimension decrease processing to the second multi-scale feature map, and outputting the addition result as a multi-scale feature image.
3. The seal segmentation method based on the UNet-S network structure model according to claim 2 is characterized in that: The step S221 of performing a depth-wise separable convolution operation on the plurality of feature images according to the sizes of the plurality of feature images comprises: S2211, obtaining the coordinate value of the feature image; S2212, obtaining a convolution kernel; S2213, obtaining a bias value; S2214, calculating the output result of the feature image according to the coordinate value of the feature image, the convolution kernel, and the offset value, wherein the calculation formula is: Wherein, (x, y) represents the horizontal and vertical coordinate values of the feature image in the multi-scale depth separable residual module, k represents the offset of the feature image in the horizontal coordinate, and l represents the offset of the feature image in the vertical coordinate; represents the output result of the jth feature image of this layer, W j is the convolution kernel of the jth feature image, σ is the activation function, b j is the bias value.
4. The seal segmentation method based on the UNet-S network structure model according to claim 2 is characterized in that: In the step S22 of inputting the plurality of feature images into a multi-scale depth separable residual module to perform multi-scale feature extraction on the plurality of feature images to obtain a plurality of multi-scale feature images, the multi-scale separable residual module includes the following function: Among them, F n-1 is the input feature image; W is the convolution kernel, its subscript indicates the convolution layer in the module from which it comes, and the superscript indicates the convolution kernel size; b indicates the bias of the convolution layer, and its subscript indicates the convolution layer in the module from which it comes; F' and F' n is the intermediate output of the multi-scale separable residual module; F n is the output result of the multi-scale separable residual module; is the convolution operation; σ is the activation function.
5. The seal segmentation method based on the UNet-S network structure model according to claim 1 is characterized in that: The step of obtaining the archive image file in the archive data set of the Republic of China, wherein the archive image file includes a plurality of archive images, further comprises: S11. Mark the positions of the seal parts of the multiple archive images to obtain a JSON file corresponding to the positions of the seal parts and a mask image of each archive image.
6. The seal segmentation method based on the UNet-S network structure model according to claim 5 is characterized in that: The loss function of the UNet-S network structure model is the BCEDiceLoss function, which is a fusion of the BCELoss function and the Sigmoid function; The BCELoss function is: L BCE =-tlogp-(1-t)log(1-p); Among them, t is the label value of the mask image of the archive image, and p is the probability predicted by the UNet-S network structure model for the category label; The Sigmoid function is: Among them, TP represents true positive, TN represents true negative, FP represents false positive, and FN represents false negative; THE BCEDiceLoss =λL BCE +L Sigmoid ; Among them, λ is the weight factor of the loss function.
7. The seal segmentation method based on the UNet-S network structure model according to claim 5 is characterized in that: After the step S11 of marking the positions of the seal parts of the plurality of archive images to obtain the JSON file corresponding to the positions of the seal parts and the mask image of each archive image, the method further includes: S12, calculating the accuracy of each pixel in the seal segmentation map according to the mask map and the seal segmentation map, and outputting the calculation result.
8. A seal segmentation device based on a UNet-S network structure model, based on the seal segmentation method based on a UNet-S network structure model according to any one of claims 1 to 7, characterized in that: include: An acquisition module, used to acquire an archive image file in the archive data set of the Republic of China, wherein the archive image file includes a plurality of archive images; An input module, used for inputting a plurality of the archive images into the UNet-S network structure model, and outputting seal segmentation maps of the plurality of archive images; Wherein, the input module includes: A first common convolution unit is used to perform a common convolution operation on the plurality of archive images to extract features and obtain a plurality of feature images, wherein the plurality of archive images have the same size; A multi-scale feature extraction unit, used for inputting the plurality of feature images into a multi-scale depth separable residual module to perform multi-scale feature extraction on the plurality of feature images to obtain a plurality of multi-scale feature images; A pooling unit, configured to perform a maximum pooling operation on the plurality of multi-scale feature images to obtain a plurality of first images, and use the first images as the feature images; A down-sampling unit, configured to repeat step S22 and step S32 according to a first preset number of times to complete down-sampling and obtain a plurality of second images; an enlarging unit, configured to enlarge the sizes of the plurality of second images based on bilinear interpolation to obtain a plurality of third images; a jump connection unit, configured to perform jump connections between the plurality of second images and the plurality of third images, and input the plurality of third images into the multi-scale depth-separable residual module to obtain a plurality of fourth images; a second common convolution unit, configured to perform a (1*1) common convolution operation on the plurality of fourth images to obtain a plurality of fifth images, and use the fifth images as the second images; The up-sampling unit is used to repeat step S25 to step S27 according to a second preset number of times to complete up-sampling and output a plurality of seal segmentation images.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the seal segmentation method based on the UNet-S network structure model described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the seal segmentation method based on the UNet-S network structure model described in any one of claims 1 to 7 are implemented.