Two-dimensional code image segmentation method based on frequency characteristic guidance
By combining high-frequency and low-frequency features in image segmentation, the problems of incomplete segmentation targets and blurred edges in QR code image segmentation in the prior art are solved, and a more accurate and complete image segmentation effect is achieved.
Patent Information
- Application Number
- CN202510158931.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-06-03
AI Technical Summary
The existing deep learning-based image segmentation method has limited encoder-decoder's ability to extract image features in QR code image segmentation, resulting in incomplete segmentation target area and blurred edges.
The FrequencyFeature Guidance Unet (FFG-Unet) method is designed, and through the combination of high-frequency and low-frequency features, a high-frequency feature extraction module (HFEB), a multi-scale high-frequency feature fusion module (MHFB) and a frequency feature fusion module (FFFB) are proposed to improve the model's ability to capture image details and edge features.
By integrating high-frequency and low-frequency characteristics, FFG-Unet can accurately outline the overall structure of the target area while maintaining edge clarity, significantly improving the accuracy and integrity of QR code image segmentation.
Smart Images

Figure CN120087389A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to a QR code image segmentation method guided by frequency features. Background Art
[0002] The QR code segmentation technology is of great significance to modern image recognition technology. With the rapid development of intelligent industry, industrial applications centered on vision are becoming increasingly important. Among them, improving the quality of input images is the key to ensuring the stable operation of subsequent machine vision tasks or user operations. Currently, image segmentation algorithms centered on deep learning have made remarkable progress. Based on large-scale low-quality - high-quality paired data, these methods enable neural networks to learn complex transformations for image restoration through data-driven means. Traditional QR code image segmentation technologies include threshold segmentation, region growing algorithms, and edge detection algorithms, etc. Although these technologies are simple and efficient, they are limited to manually performing specific tasks. Fortunately, with the rapid development of artificial intelligence, algorithms based on deep learning have made significant progress in image segmentation. For example, the unique encoder-decoder architecture of Unet shows excellent performance in image segmentation. TransUnet combines Transformer with Unet and is applied to the segmentation fields of medical images and natural images, which all demonstrate the ability of domain transfer and provide valuable insights for subsequent research. nnUnet performs excellently in complex image segmentation tasks, especially in multi-object segmentation tasks, through its innovative multi-scale feature fusion strategy. UNeXt adopts a grouped extraction strategy for high-dimensional feature extraction, which enhances the model's ability to recognize fine textures in images. However, in the research of existing image segmentation methods, due to the limited ability of the encoder-decoder to extract image features, phenomena such as incomplete segmentation of the target area and blurred edges often occur.
[0003] In summary, the challenges of current deep learning-based image segmentation methods are as follows: (1) The encoder-decoder has limited ability to extract image features and lacks attention to high-frequency, low-frequency, and detail information in QR code images. (2) The situation of lost segmentation targets and blurred edges often occurs. Therefore, it is necessary to enable the model to pay attention to different high-frequency and low-frequency feature information in the image according to different situations when segmenting contaminated QR code pictures to achieve accurate segmentation. Summary of the Invention
[0004] The purpose of the present invention is to overcome the problems existing in the prior art, provide a QR code image segmentation method guided by frequency features, propose Frequency Feature Guidance Unet (FFG-Unet), and alleviate the limitations of existing methods by combining high-frequency and low-frequency features. Specifically, a high-frequency feature extraction module High-Frequency Feature Extraction Block (HFEB) is designed. First, obtain the high-frequency feature map of the image through Rich Models for Steganalysis (RMS), and input it into the feature fusion module (HFEB). HFEB consists of multiple Swin Transformer Encoder (STE). Then, a Multi-Scale High-Frequency Feature Fusion Block (MHFB) is proposed. We input the high-frequency features and the image features in the encoder into MHFB for multi-scale feature fusion. The high-frequency features will guide the image features to pay more attention to the image details, thereby improving the network's attention to the detailed features of the target area. In addition, a frequency feature fusion module (FFFB) is proposed. This module first extracts the low-frequency features of the encoded layer, and then deeply fuses the low-frequency features with the high-frequency features. These fused features are more sensitive to the smoothing and detailed areas of the image, improving the model's perception ability of complex features during the decoding process. Through the proposed high-low frequency feature guidance optimization method, the problem of information loss in the encoding-decoding process of the model is effectively alleviated, and the ability to capture the detailed and edge features of the target is improved. Finally, the ASF and Tok-KAN modules are introduced in the skip connection and high-dimensional feature extraction stages respectively to further enhance the model's perception ability of the context.
[0005] To achieve the above object, the technical solution of the present invention is: A QR code image segmentation method guided by frequency features, including:
[0006] Design a high-frequency feature extraction module HFEB. First, obtain the high-frequency feature map of the image through digital image steganalysis RMS, and input it into HFEB;
[0007] Propose a multi-scale high-frequency feature fusion module MHFB. Input the high-frequency features extracted by HFEB and the image features extracted by the encoder into MHFB for multi-scale feature fusion, and pass the final result to the next-layer encoder through a residual connection;
[0008] Propose a frequency feature fusion module FFFB. First, extract the low-frequency features of the image features extracted by the encoder, and then deeply fuse the low-frequency features with the high-frequency features extracted by HFEB, and finally generate weights to act on the decoder;
[0009] In the skip connection and deep feature extraction stages, an attention-aware module ASF and a tokenized KAN network block Tok-KAN are introduced respectively, and finally QR code image segmentation is achieved.
[0010] In an embodiment of the present invention, the HFEB is composed of multiple shifted window transform encoders STE, and each STE is composed of a shifted window transformer Swin Transformer Block and a module merging operation Patch Merging.
[0011] In an embodiment of the present invention, the process of extracting high-frequency features by the high-frequency feature extraction module HFEB is summarized as:
[0012] HFM = RMS(M) (1)
[0013] HF k = BN(STE k (STE k-1 …(STE 2 (STE 1 (HFM))))) (2)
[0014] Among them, HFM represents the high-frequency feature map, M represents the input image, represents the high-frequency feature output by the k-th STE, and BN(·) represents normalization.
[0015] In an embodiment of the present invention, the feature fusion process of the l-th layer of the multi-scale high-frequency feature fusion module MHFB is expressed as follows:
[0016]
[0017] Among them, is the image feature output by the l-th layer encoder, F' l is the result of mixing high-frequency features and image features, CAT(·) is a splicing operation, Conv3 is a convolutional layer with a convolutional kernel size of 3×3, MLP represents a multi-layer perceptron, and HF l represents the high-frequency feature corresponding to the image feature output by the l-th layer encoder.
[0018] In an embodiment of the present invention, the frequency feature fusion module FFFB extracts the low-frequency features of the image features by the encoder through wavelet row transformation and wavelet column transformation.
[0019] In an embodiment of the present invention, the low-frequency feature extraction process of the l-th layer of the frequency feature fusion module FFFB is summarized as:
[0020]
[0021] LF = UpSamples(LF small ) (7)
[0022] where x and y represent spatial positions, k and n represent the indices of feature channels, and F en (x, k) represents the pixel point of the image feature, and N is the total number of feature channels, ψ j (·) represents the wavelet function, A j (x, n) is the matrix output after the row transformation, and φ j (·) represents the mother function of the wavelet, which contains the low-frequency component obtained after the image undergoes two-dimensional discrete wavelet transform, and then through upsampling UpSamples, the size of the low-frequency feature LF is restored;
[0023] After obtaining the low-frequency feature, it is fused with the high-frequency feature to guide and optimize the decoding process; the mixed-frequency feature MF is respectively passed through average pooling and max pooling to extract feature information at different scales, and then the results of the two poolings are concatenated in the feature dimension. Then, the feature is dimensionally reduced through the mean function Mean(·), and then further transformed through the multi-layer perceptron MLP for the concatenated feature. Finally, the weight w is generated through the Sigmoid(·) function and element-wise multiplied with the corresponding decoder output feature. The overall process is as follows:
[0024] MF l = ADD(LF l , HF l ) (8)
[0025] MF′ l = CAT(AvgPool(MF l ), MaxPool(MF l )) (9)
[0026] w = Reshape(Sigmoid(MLP(Conv1(Mean(MF′ l ))))) (10)
[0027]
[0028] where ADD represents pixel-wise addition, CAT represents the concatenation operation, AvgPool and MaxPool respectively represent average pooling and max pooling, Reshape represents size transformation, and w represents the generated weight, which are the low-frequency feature and high-frequency feature of the l-th layer respectively; represents the decoder output feature of the l-th layer, and Conv1 represents a convolutional layer with a kernel size of 1×1.
[0029] In one embodiment of the present invention, the overall ASF process is summarized as follows:
[0030]
[0031]
[0032] Among them, F' represents the result of the convolutional transformation, CAT represents the concatenation operation, and Transformer represents the self-attention mechanism processing module, is the image feature output by the l-th layer encoder; Conv1 and Conv3 respectively represent the convolutional layers with convolutional kernel sizes of 1×1 and 3×3;
[0033] The image feature output by the l-th layer encoder is first processed by the ASF attention mechanism in each skip connection layer, and then fused with the corresponding decoder output feature, as follows:
[0034]
[0035] Among them, ADD represents pixel-by-pixel addition, and respectively represent the image feature output by the encoder of the (l - 1)-th layer and the decoder output feature.
[0036] In one embodiment of the present invention, the specific implementation of Tok-KAN is as follows:
[0037] Each convolutional block consists of a convolutional layer Conv, a batch normalization layer BN, and a ReLU(·) activation function, with a kernel size of 3×3, a stride of 1, and a padding of 1 applied; a max pooling layer with a size of 2x2 is integrated in the convolutional block within the encoder; formally, given an image The output of the convolutional block is elaborated as follows:
[0038] X l = Pool(Conv(X l-1 )) (15)
[0039] Among them, represents the output feature of the l-th layer;
[0040] In Tok-KAN, first, positional encoding is performed on the output feature X l to obtain specific positional tokens, which are passed into the feature extraction module composed of a KAN layer, a depth convolutional layer DwConv, BN, and ReLU(·) activation. Residual connections are used, and the original tokens are added as residuals. Finally, the output feature of the normalized LN result is passed to the next module. Formally, the output of the k-th Tok-KAN block is summarized as:
[0041]
[0042] wherein is the output feature of the k-th Tok-KAN block.
[0043] In an embodiment of the present invention, the method uses a combination of binary cross-entropy BCE and Dice loss to train the model, and the loss between the predicted value and the target value y is as follows:
[0044]
[0045] where α and β represent the weights of the loss.
[0046] The present invention also provides a computer-readable storage medium, on which computer program instructions capable of being run by a processor are stored. When the processor runs the computer program instructions, the method steps as described above can be implemented.
[0047] Compared with the prior art, the present invention has the following beneficial effects: The present invention proposes Frequency Feature Guidance Unet (FFG-Unet), which alleviates the limitations of existing methods by combining high-frequency and low-frequency features. The core idea of FFG-Unet is to make full use of the high-frequency and low-frequency information in the image to achieve more accurate segmentation. High-frequency information is rich in edge and detail features of the target region, while low-frequency information provides important clues for the image background and smooth regions. By integrating these two types of information, FFG-Unet can accurately outline the overall structure of the target region while maintaining edge sharpness. At the same time, the present invention emphasizes the importance of the decoding layer, pays more attention to it, and improves the integrity of the final segmentation target.
[0048] In the FFG-Unet, the present invention designs a high-frequency feature extraction module (HFEB). First, a high-frequency feature map (HFM) of the QR code image is obtained through Rich Models for Steganalysis (RMS), and it is input into the high-frequency feature extraction module (HFEB). Then, the present invention proposes a multi-scale feature fusion module (MHFB). The high-frequency features and the image features in the encoder are input into the MHFB for multi-scale feature fusion. The high-frequency features will guide the image features to pay more attention to the image details, thereby enhancing the network's attention to the detailed features of the target area. In addition, the present invention proposes a frequency feature fusion module (FFFB). This module first extracts the low-frequency features of the encoding layer, and then deeply fuses the low-frequency features with the high-frequency features. These fused features are more sensitive to the smoothness and detailed areas of the image, enhancing the model's perception ability of complex features during the decoding process. Finally, the decoder integrates these high-dimensional features to generate an accurate segmentation result of the image, completing the entire process from initial feature extraction to final segmentation output. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 It is a structural diagram of the method of the present invention.
[0050] Figure 2 The size of the image features in the encoding layer is the same as that of the high-frequency features, and the image features in the encoding layer are used as residuals and added to the final result.
[0051] Figure 3 In the FFFB, the features of the picture encoding layer have the same size as both the high-frequency and low-frequency features. The darker color of the feature channels indicates that the channels have obtained larger weights, while the lighter color indicates that the channels have obtained smaller weights.
[0052] Figure 4 It is the visualization result of the method proposed by the present invention on two contaminated QR code datasets. DETAILED DESCRIPTION OF THE INVENTION
[0053] Next, in conjunction with the drawings, the technical solution of the present invention will be specifically described.
[0054] The present invention proposes a QR code image segmentation method based on frequency feature guidance, including:
[0055] Design a high-frequency feature extraction module HFEB. First, obtain the high-frequency feature map of the image through digital image steganalysis (RMS), and input it into the HFEB;
[0056] Propose a multi-scale high-frequency feature fusion module MHFB. Input the high-frequency features extracted by the HFEB and the image features extracted by the encoder into the MHFB for multi-scale feature fusion, and through residual connection, transfer the final result to the next layer of the encoder;
[0057] A frequency feature fusion module FFFB is proposed. First, the low-frequency features of the image features extracted by the encoder are extracted, and then the low-frequency features are deeply fused with the high-frequency features extracted by the HFEB, and finally weights are generated to act on the decoder;
[0058] An attention perception module (ASF) and a tokenized KAN network block (Tok-KAN) are introduced in the skip connection and deep feature extraction stages respectively, and finally QR code image segmentation is realized.
[0059] The following is the specific implementation process of the present invention.
[0060] 1 Method Structure
[0061] The method proposed by the present invention is outlined as Figure 1 shown. A QR code image segmentation framework FFG-Unet based on frequency feature guidance is proposed. The image is input into RMS and the encoder respectively to generate a high-frequency feature map and image features. HFEB is used to extract high-frequency features. MHFB is used for multi-scale depth fusion of high-frequency features and image features. FFFB is used to extract low-frequency features and fuse them with high-frequency features, and finally weights are generated to act on the decoding layer. ASF and Tok-KAN act on the skip connection and deep feature extraction stages respectively.
[0062] 1.1 High-Frequency Feature Extraction Module (HFEB)
[0063] In the HFEB, the high-frequency feature map (HFM) of the picture is extracted through RMS and passed into the high-frequency feature extractor. Different from the module used by the image encoder, the high-frequency feature extractor consists of a Swin Transformer Encoder (STE). Such a heterogeneous design can capture high-frequency features (HF) in the image more comprehensively and effectively. Each STE consists of a Swin Transformer Block and Patch Merging. The high-frequency feature extraction process can be summarized as:
[0064] HFM = RMS(M) (1)
[0065] HF k = BN(STE k (STE k-1 …(STE 2 (STE 1 (HFM))))) (2)
[0066] Among them, HFM represents the high-frequency feature map, M represents the input image, represents the high-frequency feature output by the kth STE, and BN(·) represents normalization.
[0067] 1.2 Multi-scale High-frequency Feature Fusion Block (MHFB)
[0068] The structure of MHFB is shown as Figure 2 follows. The extracted high-frequency features and image features are multi-scale deeply fused through MHFB, and the final result is passed to the next layer of the encoder through residual connection. Guided by the high-frequency features, the model's perception ability of image details is significantly enhanced, making the processing of edge and detail information more accurate, thus improving the overall accuracy of the segmentation task. The feature fusion process of the l-th layer is as follows:
[0069]
[0070] where is the image feature output by the l-th layer encoder, F' l is the result of mixing high-frequency features and image features, CAT(·) represents the concatenation operation, Conv3 is a convolutional layer with a convolutional kernel size of 3×3, MLP represents a multi-layer perceptron, and HF l represents high-frequency features.
[0071] 1.3 Frequency Feature Fusion Block (FFFB)
[0072] The structure of FFFB is shown as Figure 3 follows. First, the output features of the image encoding layer are used to extract the Low-Features (LF) through wavelet row transformation and wavelet column transformation. The low-frequency feature extraction process of the l-th layer can be summarized as:
[0073]
[0074] LF = UpSamples(LF small ) (7)
[0075] where x, y represent spatial positions, and k, n represent the indices of feature channels. F en (x,k) represents the pixel point of the image feature, and N is the total number of feature channels, ψ j (·) represents the wavelet function, A j (x,n) is the output matrix after row transformation, φ j (·) represents the mother function of the wavelet, which contains the low-frequency components obtained after the image undergoes two-dimensional discrete wavelet transform, and then through upsampling UpSamples, the size of the low-frequency feature LF is restored;
[0076] Emphasize the importance of the decoder for the segmentation model and pay more attention to it. After obtaining the low-frequency features, fuse them with the high-frequency features to guide and optimize the decoding process. The core of this fusion strategy lies in combining the characteristics of high-frequency and low-frequency features to improve the model's sensitivity to detailed features during image decoding. Specifically, these MixFeatures (MF) are respectively subjected to average pooling and max pooling to extract feature information at different scales. Then, the results of these two poolings are concatenated in the feature dimension to enhance the feature representation ability. Next, the features are dimensionally reduced through the mean function Mean(·), and then the concatenated features are further transformed through MLP so that they can more accurately capture the key information. Finally, the weight (w) is generated through the Sigmoid(·) function and element-wise multiplied with the corresponding decoder layer features. This weight adjustment mechanism enables the model to more flexibly focus on key features and improve the accuracy of the decoder layer during image prediction, as follows:
[0077] MF l = ADD(LF l , HF l ) (8)
[0078] MF' l = CAT(AvgPool(MF l ), MaxPool(MF l )) (9)
[0079] w = Reshhape(Sigmoid(MLP(Conv1(Mean(MF' l ))))) (10)
[0080]
[0081] Among them, ADD represents element-wise addition, CAT represents the concatenation operation, AvgPool and MaxPool respectively represent average pooling and max pooling. Reshape represents size transformation, and w represents the weight. are the low-frequency features and high-frequency features of the l-th layer respectively; represents the decoder output features of the l-th layer, and Conv1 represents a convolutional layer with a kernel size of 1×1.
[0082] 1.4 Skip Connection
[0083] In the skip connection layer, AssemFormer (ASF) is introduced, which combines the convolutional structure and the Transformer structure. Ordinary convolutional operations are responsible for extracting local features in the image, such as edges and textures, while the Transformer captures long-range dependencies through the self-attention mechanism, enabling the model to have stronger context awareness when dealing with complex scenarios. The overall process of ASF can be summarized as follows:
[0084]
[0085] Among them, F' represents the result of convolutional transformation, CAT represents the concatenation operation, and Transformer represents the self-attention mechanism processing module. is the feature of the encoding layer, and Conv1 and Conv3 respectively represent convolutional layers with a kernel size of 1×1 and 3×3.
[0086] The encoder features are first processed by the ASF attention mechanism in each skip connection layer and then fused with the corresponding decoder layer features as follows:
[0087]
[0088] Among them, ADD represents element-wise addition. and represent the encoder and decoder features of the l-1 layer respectively.
[0089] 1.5 Convolution Stages
[0090] Each convolutional block consists of the following components: a convolutional layer (Conv), a batch normalization layer (BN), and a ReLU(·) activation function. The applied kernel size is 3×3, the stride is 1, and the padding is 1. The convolutional block inside the encoder integrates a max pooling layer of size 2x2. Formally, given an image The output of the convolutional block can be elaborated as:
[0091] l = Pool(Conv(X l-1 )) (15)
[0092] Among them, represents the output feature map of the l-th layer.
[0093] In TokenizedKAN (Tok-KAN), first, the output feature X lPerform position encoding to obtain position tokens. Pass them into a feature extraction module composed of a KAN layer, a depthwise convolutional layer (DwConv), Batch-Normalization (BN), and the ReLU(·) activation, and use residual connections, adding the original tokens as residuals. Finally, output the normalized (LN) result features to the next module. Formally, the output of the k-th TokenizedKAN block can be summarized as:
[0094] T k = LN(T k-1 + Relu(BN(DwConv(KAN(T k-1 ))))) (16)
[0095] where is the output feature of the k-th.
[0096] 1.6 Loss Function
[0097] Use a combination of binary cross-entropy (BCE) and Dice loss to train the model. The loss between the predicted value and the target value y is as follows:
[0098]
[0099] where α, β represent the weights of the loss.
[0100] 2 Experiments
[0101] To ensure the applicability of the model, the present invention conducts experiments on two public datasets, including Data-2-1 and Data-2-2 respectively.
[0102] For Data-2-1, the target of the QR-DN1.0 dataset is used as the target image, and five types of noise processing, such as applying Gaussian noise and blur effects to the original image, are performed to simulate the QR code pollution effect in reality. In the segmentation algorithm, 800 images are randomly used as the training set of the model, 200 images are used as validation images, and 300 images are used for testing.
[0103] For Data-2-2, to evaluate the performance of the model, in the degraded images generated using the QR-DN1.0 dataset and high-quality generated QR code images, the degradation methods include 7 types such as the blur effect of real image degradation and simulated scratches. Finally, 800 of them are randomly selected as the training set, 200 images are used as validation images, and 300 images are used for testing.
[0104] All methods are implemented using Python 3.9 and PyTorch 2.0.1. Training and testing are conducted on a 12-core machine equipped with an Intel Xeon Silver 4214R 2.4 GHz CPU (128 GB RAM) and an NVIDIA GeForce RTX 3090 GPU with 24 GB VRAM. During the training phase, the Adam optimizer is used with a learning rate of 1e-3 and a momentum of 0.9. A cosine annealing learning rate scheduler is also used, with a minimum learning rate of 1e-5. The batch size is set to 8. The FFG-Unet is trained for a total of 400 epochs, and the seed is fixed at 1234 during training. The image size is uniformly 256×256.
[0105] To ensure a comprehensive evaluation, the average Dice coefficient (Dice) and the average intersection over union (IoU) are used as evaluation metrics. The IoU metric mainly evaluates the overlap between the predicted segmentation and the ground truth segmentation label, and is defined as follows:
[0106]
[0107] where TP is the true positive, FP is the false positive, and FN is the false negative.
[0108] The Dice coefficient, also known as the F1 score, is a metric used to measure the similarity between two sets in binary classification problems. Compared to IoU, Dice is more sensitive to small regions. It is defined as follows:
[0109]
[0110] 2.1 Qualitative Comparison
[0111] To qualitatively compare the segmentation results of the QR codes, the results were visualized. As Figure 4 shown, in the segmentation results of the two contaminated QR code datasets, Data-2-1 and Data-2-2, the segmentation results of the method of the present invention are almost similar to the baseline.
[0112] 2 Quantitative Comparison
[0113] Table 1 Quantitative comparison of the present invention with state-of-the-art algorithms on real datasets
[0114]
[0115] Table 1 shows the quantitative experimental results of FFG-Unet on two datasets, and the best results are shown in bold. The experimental results indicate that FFG-Unet outperforms all other algorithms in the contaminated QR code dataset. In Data-2-1, the IoU and Dice metrics of FFG-Unet reached 99.12% and 99.53% respectively. Compared with the second and third methods, Rolling-Unet and DS-Unet, the IoU metric of FFG-Unet is 1.01% and 2.25% higher respectively, and the Dice metric is 0.53% and 1.13% higher respectively. The reason is that the present invention specifically utilizes the high and low frequency features of the image to enhance the model's perception ability of more image detail information and global information. On the Data-2-2 dataset, the present invention also obtained the highest IoU value of 96.57% and Dice value of 98.21%. Rolling-Unet and EGE-Unet tend to prioritize global features at the expense of local details, which limits their segmentation ability. The present invention solves the above problems by means of a multi-stage and flexible feature extraction strategy, which enhances the model's attention to high-low frequency features.
[0116] The present invention also provides a computer-readable storage medium, on which computer program instructions capable of being run by a processor are stored. When the processor runs the computer program instructions, the method steps as described above can be implemented.
[0117] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0118] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one or more of the processes or multiple processes and / or blocks Figure 1 one or more of the blocks or multiple blocks.
[0119] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in one or more of the processes and / or blocks Figure 1 one or more of the processes and / or blocks Figure 1 specified in the block or blocks.
[0120] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the processes and / or blocks Figure 1 one or more of the processes and / or blocks Figure 1 specified in the block or blocks.
[0121] As described above, the above are only the preferred embodiments of the present invention, and are not intended to limit the present invention in any other form. Any person skilled in the art may use the technical content disclosed above to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the technical content of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. A two-dimensional code image segmentation method based on frequency feature guidance, characterized in that: include: Design a high-frequency feature extraction module HFEB. First, obtain the high-frequency feature map of the image through digital image steganalysis RMS and input it into HFEB. A multi-scale high-frequency feature fusion module MHFB is proposed. The high-frequency features extracted by HFEB and the image features extracted by the encoder are passed to MHFB for multi-scale feature fusion. The final result is passed to the next encoder layer through residual connection. A frequency feature fusion module FFFB is proposed, which first extracts the low-frequency features of the image features extracted by the encoder, and then deeply fuses the low-frequency features with the high-frequency features extracted by HFEB, and finally generates weights to act on the decoder; The attention perception module ASF and the tokenized KAN network block Tok-KAN are introduced in the skip connection and deep feature extraction stages respectively, and finally the QR code image segmentation is realized.
2. The frequency feature-guided two-dimensional code image segmentation method according to claim 1, characterized in that: HFEB consists of multiple shift window transform encoders STE, each STE consists of a shift window transformer Swin Transformer Block and a module merging operation Patch Merging.
3. The frequency feature-guided two-dimensional code image segmentation method according to claim 2, characterized in that: The process of high-frequency features extracted by the high-frequency feature extraction module HFEB can be summarized as follows: HFM=RMS(M) (1) HF k =BN(STE k (You k-1 …(STE2(STE1(HFM))))) (2) Among them, HFM represents the high-frequency feature map, M represents the input image, represents the high-frequency features of the k-th STE output, and BN(·) represents normalization.
4. The frequency feature-guided two-dimensional code image segmentation method according to claim 1, characterized in that: The feature fusion process of the first layer of the multi-scale high-frequency feature fusion module MHFB is expressed as follows: in, is the image feature output by the encoder at layer l, F' l is the result of mixing high-frequency features with image features, CAT(·) is a concatenation operation, Conv3 is a convolutional layer with a convolution kernel size of 3×3, MLP represents a multi-layer perceptron, and HF l Represents the high-frequency features corresponding to the image features output by the l-th layer encoder.
5. The frequency feature-guided two-dimensional code image segmentation method according to claim 1, characterized in that: The frequency feature fusion module FFFB extracts the low-frequency features of the image features through wavelet row changes and wavelet column changes for the encoder.
6. According to the frequency feature guided two-dimensional code image segmentation method of claim 1, the frequency feature fusion module FFFB, the first layer low-frequency feature extraction process, can be summarized as follows: LF=UpSamples(LF small ) (7) in, x, y represent the spatial position, k, n represent the index of the feature channel, F en (x, k) represents the pixel point of the image feature, and N is the total number of feature channels, ψ j (·) represents the wavelet function, A j (x,n) is the matrix output after row transformation, φ j (·) represents the mother function of the wavelet, It contains the low-frequency components of the image after two-dimensional discrete wavelet transform, and then restores the size of the low-frequency features LF through upsampling UpSamples; After obtaining the low-frequency features, they are integrated with the high-frequency features to guide and optimize the decoding process; the mixed frequency features MF are subjected to average pooling and maximum pooling respectively to extract feature information of different scales, and then the results of the two poolings are spliced in the feature dimension. Then, the features are reduced in dimension by the mean function Mean(·), and the spliced features are further transformed by the multi-layer perceptron MLP. Finally, the weight w is generated by the Sigmoid(·) function and multiplied element-by-element with the corresponding decoder output feature. The overall process is as follows: MF l =ADD(LF l ,HG l ) (8) MF′ l =CAT(AvgPool(MF l ),MaxPool(MF l )) (9) w=Reshape(Sigmoid(MLP(Conv1(Mean(MF′ l ))))) (10) Among them, ADD means pixel-by-pixel addition, CAT means concatenation operation, AvgPool and MaxPool mean average pooling and maximum pooling respectively, Reshape means size transformation, and w means the generated weight. are the low-frequency features and high-frequency features of the lth layer respectively; Represents the decoder output features of the lth layer, and Conv1 represents the convolution layer with a convolution kernel size of 1×1.
7. The frequency feature-guided two-dimensional code image segmentation method according to claim 1, characterized in that: The overall process of ASF can be summarized as follows: Among them, F' represents The result of the convolution transformation, CAT represents the concatenation operation, and Transformer represents the self-attention mechanism processing module. is the image feature output by the encoder at the lth layer, Conv1 and Conv3 represent the convolutional layer with a convolutional kernel size of 1×1 and 3×3 respectively; The image features output by the l-th layer encoder are first processed by the ASF attention mechanism in each skip connection layer, and then fused with the corresponding decoder output features, as shown below: Among them, ADD means pixel-by-pixel addition, and They represent the image features output by the encoder and the decoder output features of the l-1 layer respectively.
8. The frequency feature-guided two-dimensional code image segmentation method according to claim 1, characterized in that: The specific implementation of Tok-KAN is as follows: Each convolutional block consists of a convolutional layer Conv, a batch normalization layer BN, and a ReLU (·) activation function, with a kernel size of 3×3, a stride of 1, and a padding of 1. The convolutional block in the encoder integrates a maximum pooling layer of size 2x2. Formally, given an image The output of the convolutional block is elaborated as: X l =Pool(Conv(X l-1 )) (15) in, Represents the output features of the lth layer; In Tok-KAN, we first perform the output feature X l Perform position encoding to obtain specific position tokens, pass them to the feature extraction module consisting of a KAN layer, a deep convolutional layer DwConv, BN and ReLU (·) activation, use residual connection, and add the original token as residual, and finally pass the normalized LN result output features to the next module. Formally speaking, the output of the kth Tok-KAN block is summarized as: in is the output feature of the kth Tok-KAN block.
9. The frequency feature-guided two-dimensional code image segmentation method according to claim 1, characterized in that: The method uses a combination of binary cross entropy (BCE) and Dice loss to train the model. The loss between and the target value y is as follows: Among them, α and β represent the weights of the loss.
10. A computer-readable storage medium having stored thereon computer program instructions that can be executed by a processor, and when the processor executes the computer program instructions, the method steps according to any one of claims 1 to 9 can be implemented.
Citation Information
Cited By
Leaf-shaped scarp recognition method and device based on multi-modal data and electronic equipment
CN121527556A
Leaf-shaped steepness identification method and device based on multi-modal data and electronic equipment
CN121527556B