Grating visual illusion character recognition method
By building a DNNs text recognition model with raster vision illusion perception ability, using technical means such as multi-scale feature mapping module and attention module, the problem of insufficient raster vision illusion perception ability in DNNs is solved, and the accuracy and robustness of text recognition and camouflage object detection are significantly improved.
Patent Information
- Application Number
- CN202510177906.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-18
AI Technical Summary
Existing deep neural networks (DNNs) are weak in raster vision illusion perception, resulting in poor performance in tasks such as text recognition, especially when dealing with distortion or occlusion scenarios.
A DNNs text recognition model with raster vision illusion perception ability is built. By guiding the model to learn global shape preferences rather than local features during training, it adopts modules such as multi-scale feature mapping module (MFP), feedback interactive attention module (FBIAM), feedforward interactive attention module (FFIAM) and edge fusion module (EFM), combined with the fusion of edge loss and identification loss, the model's raster vision illusion perception ability is improved.
It significantly improves the accuracy and robustness of text recognition, can more effectively identify text presented in raster vision illusion on printed materials or billboards, and is applied to camouflage object detection tasks, improving the accuracy of camouflage object detection.
Smart Images

Figure CN120107749A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and deep learning, and specifically relates to a grating visual illusion text recognition method. Background Art
[0002] Deep Neural Networks (DNNs) have the ability to express strong semantic-level features and perform well in computer vision tasks such as image recognition. However, compared with the human visual system, distortions in the input image, such as blurring and rasterization, can easily cause a sharp drop in the performance of the neural network. In contrast, the human visual system has evolved to use limited sensory information to extract features in complex scenes and complete recognition and decision-making. Illusory contour perception is crucial to this process. Even if there is a continuity break between adjacent features, global contour perception still occurs, which is manifested in the robustness of the human visual system to distorted or occluded scenes. As a classic illusory contour, grating illusion is widely used in psychology and neurophysiology research, but the grating illusion perception ability in DNNs has rarely been studied and applied. Therefore, constructing a brain-inspired DNNs text recognition model with grating illusion perception ability can improve the robustness of text recognition, and DNNs with illusion perception ability are expected to be applied to computer vision tasks such as scene character recognition and camouflaged target detection to improve their performance and promote the development of application scenarios such as autonomous driving and rescue robots.
[0003] Grating visual illusion perception refers to the clear boundaries that the biological visual system can still perceive in areas without color contrast or brightness gradient. In addition to the human visual system, this perception ability is also widely found in non-human primates, fish, birds and other organisms. This shows that the ability to perceive grating visual illusions plays a fundamental and key role in the biological visual system. In theory, this is also one of the visual perception abilities that DNNs should have. However, current DNNs based on Convolutional Neural Network (CNN) and Transformer architectures basically adopt a pure feedforward architecture and have almost no grating visual illusion perception ability. In addition, the existing methods for improving DNNs' visual illusion perception have cumbersome training paradigms and lengthy iteration steps, making it difficult to directly generalize and apply them to real complex tasks. Summary of the invention
[0004] In order to solve the above technical problems, the present invention provides a grating visual illusion text recognition method, which guides DNNs to learn global shape preferences rather than local features during training, so as to improve the DNNs grating visual illusion perception ability and thereby enhance the robustness of text recognition.
[0005] The technical solution adopted by the present invention is: a grating visual illusion text recognition method, the specific steps are as follows:
[0006] S1. Construct a DNNs text recognition model with grating visual illusion perception;
[0007] The DNNs text recognition model includes: a text recognition model and an edge detection model.
[0008] The edge detection model is used to provide edge map pseudo labels; the text recognition model includes: 1 MFP module, 4 cortical column modules, 3 FFIAM modules, 3 EFM modules, and 1 classification layer module.
[0009] Among them, the four cortical column modules are the primary visual cortex V1 cortical column, the secondary visual cortex V2 cortical column and two fourth visual area V4 cortical columns.
[0010] The primary visual cortex V1 cortical column includes: 1 primary first-stage module, 1 FBIAM module, 1 primary second-stage module, and 1 primary third-stage module. The secondary visual cortex V2 cortical column includes: 1 pooling layer, 1 secondary first-stage module, 1 FBIAM module, 1 secondary second-stage module, and 1 secondary third-stage module. The first visual area IV V4 cortical column includes: 1 pooling layer, 1 area IV first-stage module, 1 FBIAM module, 1 area IV second-stage module, and 1 area IV third-stage module. The second visual area IV V4 cortical column includes: 1 pooling layer, 1 area IV first-stage module, 1 area IV second-stage module, and 1 area IV third-stage module.
[0011] S2. Use the MNIST training set to train the DNNs text recognition model. First, input the training image into the edge detection model and the text recognition model respectively. The edge detection model obtains the edge map pseudo label, and the text recognition model obtains four sets of side output features and a set of multi-scale fusion features obtained based on the MFP module.
[0012] S3, passing the four sets of side output features in step S2 into the FFIAM module to obtain three sets of attention-modulated side output features;
[0013] S4, passing a group of the smallest feature size among the attention-modulated side output features obtained in step S3 into the classification layer module, and sequentially passing through a global average pooling layer, an LN layer, and a fully connected linear layer to obtain a probability output of text recognition;
[0014] Among them, the probability output of text recognition is the probability of belonging to each category, and the MNIST data set used has a total of 10 categories.
[0015] S5, calculating the recognition loss by using the cross entropy loss function on the text recognition probability output in step S4 and the text label;
[0016] S6, the attention modulated side output features obtained in step S3 and the multi-scale fusion features obtained by the MFP module in step S2 are passed to the EFM module and integrated to obtain an edge prediction map, and then the edge loss is calculated by the edge loss function with the edge map pseudo label in step S2;
[0017] S7, performing weighted summation of the recognition loss in step S5 and the edge loss in step S6 to obtain a fusion loss function, and then iterating the training cycle to a preset maximum number of training times to complete the training of the DNNs text recognition model and obtain a trained model;
[0018] S8. Based on the DNNs text recognition model trained in step S7, the MNIST test set is used for testing, that is, the test image is input into the trained model for testing to obtain the text recognition result, thereby completing the grating visual illusion text recognition.
[0019] Furthermore, the step S2 is specifically as follows:
[0020] S21, passing the training image into the edge detection model to obtain the edge map pseudo label;
[0021] S22, passing the training image into the MFP module to obtain a set of multi-scale fusion features;
[0022] The input image is passed in parallel to the convolution kernel with dilation rates of 2, 4, 6, and 7, with a size of 3×3 and the number of output channels of C. 1 It uses four convolutional layers and adds the four output results. After an LN layer and channel shuffle, it outputs multi-scale fusion features.
[0023] Among them, C 1 =64.
[0024] S23, sequentially passing the multi-scale fusion features in step S22 through four sub-modules consisting of depthwise separable convolutions with different receptive field sizes, i.e., four cortical column modules, to obtain four sets of side output features;
[0025] 1) The first submodule is the primary visual cortex V1 cortical column;
[0026] First, the multi-scale fusion features are passed into the primary first-stage module to obtain the output features of the first stage, that is, the multi-scale fusion features are first passed into a receptive field size of 5×5 and the number of output channels is C. 1 The depth of the separable convolutional layer is then passed through a receptive field size of 3×3, a void rate of 2, and an output channel number of 2×C 1The depth-separable convolutional layer is passed through a receptive field size of 1×1 and the number of output channels is C. 1 The convolution layer is constructed, and the obtained output is added to the multi-scale fusion feature in step S22 to obtain the output feature of the first stage.
[0027] Then the output features of the first stage are passed into the FBIAM module, and the second set of side output features are fed back to the FBIAM module to obtain the output features of the FBIAM module. That is, the two input features are first fused through a feedback channel attention module to obtain the channel attention result, and then the channel attention result is passed into the multi-scale spatial attention module to obtain the output features of the FBIAM module, as follows:
[0028] The second set of side output features are first upsampled to twice the original resolution through deconvolution, and then passed to a convolution layer with a receptive field size of 1×1 and an output channel number of C and an LN layer in sequence to obtain the alignment results of the second set of side output features. The output features of the first stage are passed to spatial average pooling and spatial maximum pooling in parallel to obtain the output results of spatial average pooling and the output results of spatial maximum pooling, respectively. The two output results are then passed to a shared convolution layer with a receptive field size of 1×1 and an output channel number of k×C, a GELU, and a shared convolution layer with a receptive field size of 1×1 and an output channel number of C. The two processed results are then added and passed through the sigmoid function to perform the Hadamard product with the output features of the first stage to obtain the processed results of the output features of the first stage, which are then added to the alignment results of the second set of side output features to obtain the channel attention results. Then the channel attention results are passed in parallel to the channel average pooling and channel maximum pooling and the two results are spliced. The spliced results are passed in parallel to a convolutional layer with a receptive field size of 3×7 and an output channel of 1, a convolutional layer with a receptive field size of 5×13 and an output channel of 1, a convolutional layer with a receptive field size of 7×3 and an output channel of 1, and a convolutional layer with a receptive field size of 13×5 and an output channel of 1. The four results are added and passed through the sigmoid function, and then the Hadamard product is performed with the channel attention result. The product result is added to the channel attention result and after channel shuffle, the multi-scale spatial attention result, i.e., the output feature of the FBIAM module, is obtained.
[0029] Here, C represents the smaller number of channels in the two inputs, and k is set to 4.
[0030] Finally, the output features of the FBIAM module are passed to the primary second-stage module, and the output is added to the output features of the FBIAM module to obtain the output features of the second stage, and this output feature is passed to the primary third-stage module to obtain the first set of side output features, that is, the output features of the FBIAM module are passed to a receptive field size of 4×4, a void rate of 2, and an output channel number of 2×C. 1 The depth of the separable convolutional layer is then passed through a receptive field size of 1×1 and the number of output channels is C 1 The convolutional layer is added to the output features of the FBIAM module to obtain the output features of the second stage, and then the output features of the second stage are passed to a receptive field size of 4×4, a void rate of 2, and the number of output channels C. 1 A depth-separable convolutional layer is constructed to obtain the first set of side output features.
[0031] 2) The second submodule is the cortical column of the secondary visual cortex V2;
[0032] First, the output features of the second stage in the primary visual cortex V1 cortical column are passed through a pooling layer and then passed to the secondary first stage module to obtain the output features of the first stage. That is, the output features of the second stage in the first submodule are passed through a pooling layer and downsampled to half of the original resolution. Then, they are passed to a receptive field size of 6×6, a void rate of 2, and an output channel number of C. 2 The depth of the separable convolutional layer is used to obtain the output features of the first stage.
[0033] Among them, C 2 =128.
[0034] Then, the output features of the first stage are passed into the FBIAM module, and the third set of side output features are fed back to the FBIAM module to obtain the output features of the FBIAM module.
[0035] Among them, the FBIAM module structure and processing flow are consistent in the primary visual cortex V1 cortical column, the secondary visual cortex V2 cortical column and the first fourth visual area V4 cortical column.
[0036] Finally, the output features of the FBIAM module are passed to the secondary second-stage module, and the output is added to the output features of the FBIAM module to obtain the output features of the second stage, and this output feature is passed to the secondary third-stage module, and the output is added to the output features of the second stage to obtain the second set of side output features, that is, the output features of the FBIAM module are passed to a receptive field size of 7×7, a void rate of 2, and an output channel number of C. 2 The depth of the separable convolutional layer is then passed through a receptive field size of 3×3, a dilation rate of 6, and an output channel number of 2×C2 The depth-separable convolutional layer is passed through a receptive field size of 1×1 and the number of output channels is C. 2 The convolutional layer is added to the output features of the FBIAM module to obtain the output features of the second stage. The output features of the second stage are then passed to a receptive field size of 7×7, a void rate of 2, and an output channel number of 2×C 2 The depth of the separable convolutional layer is then passed through a receptive field size of 1×1 and the number of output channels is C 2 The convolution layer is added to the output features of the second stage to obtain the second set of side output features.
[0037] 3) The third submodule is the first visual area V4 cortical column;
[0038] First, the output features of the second stage in the secondary visual cortex V2 cortical column are passed through a pooling layer and then passed to the first stage module of the fourth area to obtain the output features of the first stage, that is, the output features of the second stage in the second submodule are passed through a pooling layer, downsampled to half of the original resolution, and then passed to a receptive field size of 7×7, a void rate of 3, and an output channel number of C. 3 The depth of the separable convolutional layer is used to obtain the output features of the first stage.
[0039] Among them, C 3 =256.
[0040] Then, the output features of the first stage are passed into the FBIAM module, and the fourth set of side output features are fed back to the FBIAM module to obtain the output features of the FBIAM module.
[0041] Finally, the output features of the FBIAM module are passed to the second stage module in the fourth zone, and the output is added to the output features of the FBIAM module to obtain the output features of the second stage, and this output feature is passed to the third stage module in the fourth zone, and the output is added to the output features of the second stage to obtain the third set of side output features, that is, the output features of the FBIAM module are passed to a receptive field size of 9×9, a void rate of 3, and an output channel number of C. 3 The depth of the separable convolutional layer is then passed through a receptive field size of 3×3, a dilation rate of 12, and an output channel number of 2×C 3 The depth-separable convolutional layer is passed through a receptive field size of 1×1 and the number of output channels is C. 3 The convolutional layer is added to the output features of the FBIAM module to obtain the output features of the second stage. The output features of the second stage are then passed to a receptive field size of 9×9, a void rate of 3, and an output channel number of 2×C 3The depth of the separable convolutional layer is then passed through a receptive field size of 1×1 and the number of output channels is C 3 The convolution layer is added to the output features of the second stage to obtain the third set of side output features.
[0042] 4) the fourth submodule is the second visual area V4 cortical column;
[0043] First, the output features of the second stage in the first visual area V4 cortical column are passed through a pooling layer and then passed to the first stage module of the fourth area to obtain the output features of the first stage. That is, the output features of the second stage in the third submodule are passed through a pooling layer and downsampled to half of the original resolution. Then, they are passed to a receptive field size of 7×7, a hole rate of 3, and an output channel number of C. 4 The depth of the separable convolutional layer is used to obtain the output features of the first stage.
[0044] Among them, C 4 =512.
[0045] Then the output features of the first stage are passed to the second stage module of the fourth zone, and the output is added to the output features of the first stage to obtain the output features of the second stage. Finally, the output features of the second stage are passed to the third stage module of the fourth zone, and the output is added to the output features of the second stage to obtain the fourth set of side output features, that is, the output features of the first stage are passed to a receptive field size of 9×9, a void rate of 3, and an output channel number of C. 4 The depth of the separable convolutional layer is then passed through a receptive field size of 3×3, a dilation rate of 12, and an output channel number of 2×C 4 The depth-separable convolutional layer is passed through a receptive field size of 1×1 and the number of output channels is C. 4 The convolutional layer is added to the output features of the first stage to obtain the output features of the second stage. The output features of the second stage are then passed to a receptive field size of 9×9, a void rate of 3, and an output channel number of 2×C. 4 The depth of the separable convolutional layer is then passed through a receptive field size of 1×1 and the number of output channels is C 4 The convolution layer is added to the output features of the second stage to obtain the fourth set of side output features.
[0046] Furthermore, the step S3 is specifically as follows:
[0047] S31, arranging the four groups of side output features obtained in step S2 from large to small according to feature size;
[0048] S32, based on the sorting result of step S31, passing the first set of features and the second set of features into the first FFIAM to obtain the first set of attention-modulated side output features;
[0049] The first set of features and the second set of features are first fused through a feedforward channel attention module to obtain a channel attention result, and then the channel attention result is passed to the multi-scale spatial attention module to obtain the first set of attention-modulated side output features, as follows:
[0050] The first set of features is sequentially subjected to a downsampling, an LN layer, and an output channel C with a receptive field size of 1×1. 5 The convolution layer is used to obtain the alignment result of the first set of features, and the alignment result of the first set of features is passed to the spatial average pooling and the spatial maximum pooling in parallel to obtain the output results of the spatial average pooling and the output results of the spatial maximum pooling respectively. Then the two output results are successively passed to an output channel with a receptive field size of 1×1 and the number of channels is k×C 5 The convolutional layer, a GELU, a receptive field size of 1×1, and the number of output channels is C 5 The convolution layer is passed through a sigmoid function to obtain two processed results. The two processed results are then added and passed through the sigmoid function to perform Hadamard product with the alignment result of the first set of features, and the result is added to the second set of features to obtain the channel attention result. The channel attention result is then passed in parallel to the channel average pooling and channel maximum pooling and the two results are concatenated. The concatenated result is passed in parallel to a convolution layer with a receptive field size of 3×7 and an output channel number of 1, a convolution layer with a receptive field size of 5×13 and an output channel number of 1, a convolution layer with a receptive field size of 7×3 and an output channel number of 1, and a convolution layer with a receptive field size of 13×5 and an output channel number of 1. The four results are added and passed through the sigmoid function to perform Hadamard product with the channel attention result. The product result is added to the channel attention result and shuffled through the channel to obtain the multi-scale spatial attention result, that is, the side output features of the first set of attention modulation.
[0051] Among them, C 5 Indicates the larger number of channels of the two inputs, and k is set to 4. In the text recognition model, the three FFIAM module structures are consistent with the processing flow.
[0052] S33, passing the obtained first group of attention-modulated side output features and the third group of features into the second FFIAM to obtain the second group of attention-modulated side output features;
[0053] S34. The second set of attention-modulated side output features and the fourth set of features are passed into the third FFIAM to obtain the third set of attention-modulated side output features.
[0054] Furthermore, the step S5 is specifically as follows:
[0055] The probability output of the character recognition in step S4 is set to P cls , whose size is N×M; the text label is represented by Y cls , whose size is N×M.
[0056] Among them, N represents the number of samples in a batch, and M represents the total number of recognized categories.
[0057] Then use the cross entropy loss function to calculate P cls With Y cls The recognition loss L cls , the calculation expression is as follows:
[0058]
[0059] Among them, y ij represents the true label of sample i in category j, p ij represents the predicted probability of sample i in category j.
[0060] Furthermore, the step S6 is specifically as follows:
[0061] S61, passing the attention-modulated side output features obtained in step S3 and the multi-scale fusion features obtained by the MFP module in step S2 into the EFM module and integrating them to obtain an edge prediction map;
[0062] First, the three groups of attention-modulated side output features obtained in step S3 and the multi-scale fusion features obtained in step S2 are arranged in ascending order of size, and then the first and second groups of features are passed into the first EFM to obtain the output result of the first EFM, as follows:
[0063] The two sets of input features are respectively passed through a receptive field size of 5×5, and the number of output channels is C. 6 After a convolutional layer, an LN layer, and a GELU, the smaller feature is upsampled to twice the original resolution and added to the other to get the output of the first EFM.
[0064] Among them, C 6 Represents the smaller number of channels of the two inputs.
[0065] Then the output result of the first EFM and the third set of features are passed into the second EFM to obtain the output result of the second EFM. The specific processing flow is the same as that of the first EFM. The output result of the second EFM and the fourth set of features are passed into the third EFM to obtain the output result of the third EFM, that is, the edge prediction map.
[0066] S62, calculating the edge loss through the edge loss function based on the edge prediction image obtained in step S61 and the edge image pseudo label in step S2;
[0067] Set the edge map pseudo label obtained in step S2 to be Y edge =(y j ,j=1,…,|Y edge |),y j ∈[0,1].
[0068] Among them, y j Represents the pixel value at the jth pixel.
[0069] Then Y edge Classified into positive sample set and negative sample set, where the positive sample set and negative sample set are represented as Y + ={y j ,y j >η} and Y - ={y j ,y j =0}, all other pixel values are ignored, η represents the pixel value threshold, which is set to 0.2;
[0070] Then set the edge prediction map obtained in step S61 to be represented as P edge =(p j ,j=1,…,|P edge |),p j ∈[0,1].
[0071] Among them, p j Represents the value after a sigmoid function is processed at the jth pixel.
[0072] Finally, the edge loss function is used to calculate the edge map pseudo label Y edge and edge prediction graph P edge The edge loss L edge , the calculation expression is as follows:
[0073]
[0074]
[0075] Among them, α and β represent weight coefficients, and λ represents the weight of the control coefficient size, which is set to 1.
[0076] Furthermore, in step S7, the weighted sum of the recognition loss in step S5 and the edge loss in step S6 is as follows:
[0077] Denote the total loss as L total , the calculation expression is as follows:
[0078] L total =L cls +γ·L edge
[0079] Among them, γ represents the weight coefficient, which is set to 0.01.
[0080] The beneficial effects of the present invention are as follows: the method of the present invention first constructs a DNNs text recognition model with grating visual illusion perception, uses the MNIST training set to train the DNNs text recognition model, then uses the test set to test the trained model, inputs the test image into the text recognition model to obtain side output features and multi-scale fusion features, obtains attention-modulated side output features from the side output features through the FFIAM module, and passes a group of the smallest feature sizes in the attention-modulated side output features into the classification layer to obtain the probability output of text recognition, and finally obtains the text recognition result to complete the grating visual illusion text recognition. The method of the present invention constructs a DNNs text recognition model with grating visual illusion perception based on the relevant neural mechanism of grating visual illusion perception, and constructs MFP, FBIAM, FFIAM, and EFM modules at the same time, and uses the fusion of edge loss and recognition loss to guide the DNNs model to learn global contour perception ability during training, that is, to learn global shape preferences rather than local features, and give the model grating visual illusion perception ability, thereby improving the accuracy and robustness of text recognition, helping to improve the recognition accuracy of text presented in a grating visual illusion manner on printed materials and billboards in scene character recognition tasks, and improving system reliability. At the same time, the method of the present invention can also be applied to camouflage target detection tasks, using the global contour perception attribute of visual illusion to improve the accuracy of camouflage target detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0081] Figure 1 The present invention is a flowchart of a grating visual illusion text recognition method.
[0082] Figure 2 Schematic diagram of the DNNs text recognition model described in an embodiment of the present invention.
[0083] Figure 3 Schematic diagram of a multi-scale feature projection module (MFP) in an embodiment of the present invention.
[0084] Figure 4 Schematic diagram of a Feedback Interaction Attention Module (FBIAM) in an embodiment of the present invention.
[0085] Figure 5 Schematic diagram of a feedforward interaction attention module (FFIAM) in an embodiment of the present invention.
[0086] Figure 6 FIG. 1 is a schematic diagram of an edge fusion module (EFM) in an embodiment of the present invention.
[0087] Table 1 is a comparison of the recognition results of the present invention and other methods. DETAILED DESCRIPTION
[0088] The method of the present invention is further described below in conjunction with the accompanying drawings and embodiments.
[0089] like Figure 1 As shown, a flow chart of a grating visual illusion text recognition method of the present invention, the specific steps are as follows:
[0090] S1. Construct a DNNs text recognition model with grating visual illusion perception;
[0091] like Figure 2 As shown, the DNNs text recognition model includes: a text recognition model and an edge detection model.
[0092] The edge detection model is used to provide edge map pseudo labels; the text recognition model includes: 1 MFP module, 4 cortical column modules, 3 FFIAM modules, 3 EFM modules, 1 classification layer module (i.e. Figure 2 in the .
[0093] Among them, the four cortical column modules are the primary visual cortex V1 cortical column, the secondary visual cortex V2 cortical column and two fourth visual area V4 cortical columns.
[0094] The primary visual cortex V1 cortical column includes: 1 primary first-stage module, 1 FBIAM module, 1 primary second-stage module, and 1 primary third-stage module. The secondary visual cortex V2 cortical column includes: 1 pooling layer, 1 secondary first-stage module, 1 FBIAM module, 1 secondary second-stage module, and 1 secondary third-stage module. The first visual area IV V4 cortical column includes: 1 pooling layer, 1 area IV first-stage module, 1 FBIAM module, 1 area IV second-stage module, and 1 area IV third-stage module. The second visual area IV V4 cortical column includes: 1 pooling layer, 1 area IV first-stage module, 1 area IV second-stage module, and 1 area IV third-stage module.
[0095] S2. Use the MNIST training set to train the DNNs text recognition model. First, input the training image into the edge detection model and the text recognition model respectively. The edge detection model obtains the edge map pseudo label, and the text recognition model obtains four sets of side output features and a set of multi-scale fusion features obtained based on the MFP module.
[0096] S3, passing the four sets of side output features in step S2 into the FFIAM module to obtain three sets of attention-modulated side output features;
[0097] S4, passing a group of the smallest feature size among the attention-modulated side output features obtained in step S3 into the classification layer module, and sequentially passing through a global average pooling layer, an LN layer, and a fully connected linear layer to obtain a probability output of text recognition;
[0098] Among them, the probability output of text recognition is the probability of belonging to each category, and the MNIST data set used has a total of 10 categories.
[0099] S5, calculating the recognition loss by using the cross entropy loss function on the text recognition probability output in step S4 and the text label;
[0100] S6, the attention modulated side output features obtained in step S3 and the multi-scale fusion features obtained by the MFP module in step S2 are passed to the EFM module and integrated to obtain an edge prediction map, and then the edge loss is calculated by the edge loss function with the edge map pseudo label in step S2;
[0101] S7, performing weighted summation of the recognition loss in step S5 and the edge loss in step S6 to obtain a fusion loss function, and then iterating the training cycle to a preset maximum number of training times to complete the training of the DNNs text recognition model and obtain a trained model;
[0102] Among them, the maximum number of training times in this embodiment is 100.
[0103] S8. Based on the DNNs text recognition model trained in step S7, the MNIST test set is used for testing, that is, the test image (corresponding grating optical illusion image) is input into the trained model for testing to obtain the text recognition result, thereby completing the grating optical illusion text recognition.
[0104] In this embodiment, step S2 is specifically as follows:
[0105] S21, passing the training image into the edge detection model to obtain the edge map pseudo label;
[0106] like Figure 2 As shown in the left part, this embodiment uses the edge detection model LVP-Net pre-trained on the BSDS500-VOC dataset to generate an edge map, and uses the edge map after non-maximum suppression as the edge map pseudo label.
[0107] S22, passing the training image into the MFP module to obtain a set of multi-scale fusion features;
[0108] like Figure 3 As shown, the input image is passed in parallel to the convolution kernel with dilation rates of 2, 4, 6, and 7, with a size of 3×3 and the number of output channels of C. 1 It uses four convolutional layers and adds the four output results. After a LN (Layer Normalization) layer and channel shuffle, it outputs the multi-scale fusion features.
[0109] Among them, C 1 =64.
[0110] S23, sequentially passing the multi-scale fusion features in step S22 through four sub-modules consisting of depthwise separable convolutions with different receptive field sizes, i.e., four cortical column modules, to obtain four sets of side output features;
[0111] Inspired by the neural mechanism of grating visual illusion, in addition to the feedforward connection with pooling, the submodules also interact with each other through feedback connections. The size of the receptive field gradually increases with the deepening of the layer. The specific structure is as follows: Figure 2 As shown, DW represents the depth-wise separable convolution, C 1 , C 2 , C 3 and C 4 They are 64, 128, 256, and 512 respectively.
[0112] 1) The first submodule is the primary visual cortex V1 cortical column;
[0113] First, the multi-scale fusion features are passed into the primary first-stage module to obtain the output features of the first stage, that is, the multi-scale fusion features are first passed into a receptive field size of 5×5 and the number of output channels is C. 1 The depth of the separable convolutional layer is then passed through a receptive field size of 3×3, a void rate of 2, and an output channel number of 2×C 1 The depth-separable convolutional layer is passed through a receptive field size of 1×1 and the number of output channels is C. 1 The convolution layer is constructed, and the obtained output is added to the multi-scale fusion feature in step S22 to obtain the output feature of the first stage.
[0114] like Figure 4 As shown in the figure, the output features of the first stage are passed into the Feedback Interaction Attention Module (FBIAM), and the second set of side output features are fed back to the FBIAM module to obtain the output features of the FBIAM module. That is, the two input features are first fused through a feedback channel attention module to obtain the channel attention result, and then the channel attention result is passed into the multi-scale spatial attention module to obtain the output features of the FBIAM module, as shown below:
[0115] The second set of side output features are first upsampled to twice the original resolution through deconvolution, and then passed to a convolution layer with a receptive field size of 1×1 and an output channel number of C and an LN layer in sequence to obtain the alignment results of the second set of side output features. The output features of the first stage are passed to spatial average pooling and spatial maximum pooling in parallel to obtain the output results of spatial average pooling and the output results of spatial maximum pooling, respectively. The two output results are then passed to a shared convolution layer with a receptive field size of 1×1 and an output channel number of k×C, a Gaussian error linear unit (GELU), and a shared convolution layer with a receptive field size of 1×1 and an output channel number of C. The two processed results are then added and passed through the sigmoid function to perform the Hadamard product with the output features of the first stage to obtain the processed results of the output features of the first stage, which are added to the alignment results of the second set of side output features to obtain the channel attention results. Then the channel attention results are passed in parallel to the channel average pooling and channel maximum pooling and the two results are spliced. The spliced results are passed in parallel to a convolutional layer with a receptive field size of 3×7 and an output channel of 1, a convolutional layer with a receptive field size of 5×13 and an output channel of 1, a convolutional layer with a receptive field size of 7×3 and an output channel of 1, and a convolutional layer with a receptive field size of 13×5 and an output channel of 1. The four results are added and passed through the sigmoid function, and then the Hadamard product is performed with the channel attention result. The product result is added to the channel attention result and after channel shuffle, the multi-scale spatial attention result, i.e., the output feature of the FBIAM module, is obtained.
[0116] Here, C represents the smaller number of channels in the two inputs, and k is set to 4.
[0117] Finally, the output features of the FBIAM module are passed to the primary second-stage module, and the output is added to the output features of the FBIAM module to obtain the output features of the second stage, and this output feature is passed to the primary third-stage module to obtain the first set of side output features, that is, the output features of the FBIAM module are passed to a receptive field size of 4×4, a void rate of 2, and an output channel number of 2×C. 1 The depth of the separable convolutional layer is then passed through a receptive field size of 1×1 and the number of output channels is C 1 The convolutional layer is added to the output features of the FBIAM module to obtain the output features of the second stage, and then the output features of the second stage are passed to a receptive field size of 4×4, a void rate of 2, and the number of output channels C. 1 A depth-separable convolutional layer is constructed to obtain the first set of side output features.
[0118] 2) The second submodule is the cortical column of the secondary visual cortex V2;
[0119] First, the output features of the second stage in the primary visual cortex V1 cortical column are passed through a pooling layer and then passed to the secondary first stage module to obtain the output features of the first stage. That is, the output features of the second stage in the first submodule are passed through a pooling layer and downsampled to half of the original resolution. Then, they are passed to a receptive field size of 6×6, a void rate of 2, and an output channel number of C. 2 The depth of the separable convolutional layer is used to obtain the output features of the first stage.
[0120] Among them, C 2 =128.
[0121] Then, the output features of the first stage are passed into the FBIAM module, and the third set of side output features are fed back to the FBIAM module to obtain the output features of the FBIAM module.
[0122] Among them, the FBIAM module structure and processing flow are consistent in the primary visual cortex V1 cortical column, the secondary visual cortex V2 cortical column and the first fourth visual area V4 cortical column.
[0123] Finally, the output features of the FBIAM module are passed to the secondary second-stage module, and the output is added to the output features of the FBIAM module to obtain the output features of the second stage, and this output feature is passed to the secondary third-stage module, and the output is added to the output features of the second stage to obtain the second set of side output features, that is, the output features of the FBIAM module are passed to a receptive field size of 7×7, a void rate of 2, and an output channel number of C. 2 The depth of the separable convolutional layer is then passed through a receptive field size of 3×3, a dilation rate of 6, and an output channel number of 2×C 2 The depth-separable convolutional layer is passed through a receptive field size of 1×1 and the number of output channels is C. 2 The convolutional layer is added to the output features of the FBIAM module to obtain the output features of the second stage. The output features of the second stage are then passed to a receptive field size of 7×7, a void rate of 2, and an output channel number of 2×C 2 The depth of the separable convolutional layer is then passed through a receptive field size of 1×1 and the number of output channels is C 2 The convolution layer is added to the output features of the second stage to obtain the second set of side output features.
[0124] 3) The third submodule is the first visual area V4 cortical column;
[0125] First, the output features of the second stage in the secondary visual cortex V2 cortical column are passed through a pooling layer and then passed to the first stage module of the fourth area to obtain the output features of the first stage, that is, the output features of the second stage in the second submodule are passed through a pooling layer, downsampled to half of the original resolution, and then passed to a receptive field size of 7×7, a void rate of 3, and an output channel number of C. 3 The depth of the separable convolutional layer is used to obtain the output features of the first stage.
[0126] Among them, C 3 =256.
[0127] Then, the output features of the first stage are passed into the FBIAM module, and the fourth set of side output features are fed back to the FBIAM module to obtain the output features of the FBIAM module.
[0128] Finally, the output features of the FBIAM module are passed to the second stage module in the fourth zone, and the output is added to the output features of the FBIAM module to obtain the output features of the second stage, and this output feature is passed to the third stage module in the fourth zone, and the output is added to the output features of the second stage to obtain the third set of side output features, that is, the output features of the FBIAM module are passed to a receptive field size of 9×9, a void rate of 3, and an output channel number of C. 3 The depth of the separable convolutional layer is then passed through a receptive field size of 3×3, a dilation rate of 12, and an output channel number of 2×C 3 The depth-separable convolutional layer is passed through a receptive field size of 1×1 and the number of output channels is C. 3 The convolutional layer is added to the output features of the FBIAM module to obtain the output features of the second stage. The output features of the second stage are then passed to a receptive field size of 9×9, a void rate of 3, and an output channel number of 2×C 3 The depth of the separable convolutional layer is then passed through a receptive field size of 1×1 and the number of output channels is C 3 The convolution layer is added to the output features of the second stage to obtain the third set of side output features.
[0129] 4) the fourth submodule is the second visual area V4 cortical column;
[0130] First, the output features of the second stage in the first visual area V4 cortical column are passed through a pooling layer and then passed to the first stage module of the fourth area to obtain the output features of the first stage. That is, the output features of the second stage in the third submodule are passed through a pooling layer and downsampled to half of the original resolution. Then, they are passed to a receptive field size of 7×7, a hole rate of 3, and an output channel number of C. 4The depth of the separable convolutional layer is used to obtain the output features of the first stage.
[0131] Among them, C 4 =512.
[0132] Then the output features of the first stage are passed to the second stage module of the fourth zone, and the output is added to the output features of the first stage to obtain the output features of the second stage. Finally, the output features of the second stage are passed to the third stage module of the fourth zone, and the output is added to the output features of the second stage to obtain the fourth set of side output features, that is, the output features of the first stage are passed to a receptive field size of 9×9, a void rate of 3, and an output channel number of C. 4 The depth of the separable convolutional layer is then passed through a receptive field size of 3×3, a dilation rate of 12, and an output channel number of 2×C 4 The depth-separable convolutional layer is passed through a receptive field size of 1×1 and the number of output channels is C. 4 The convolutional layer is added to the output features of the first stage to obtain the output features of the second stage. The output features of the second stage are then passed to a receptive field size of 9×9, a void rate of 3, and an output channel number of 2×C. 4 The depth of the separable convolutional layer is then passed through a receptive field size of 1×1 and the number of output channels is C 4 The convolution layer is added to the output features of the second stage to obtain the fourth set of side output features.
[0133] In this embodiment, step S3 is specifically as follows:
[0134] S31, arranging the four groups of side output features obtained in step S2 from large to small according to feature size;
[0135] S32, based on the sorting result of step S31, passing the first set of features and the second set of features into the first FFIAM to obtain the first set of attention-modulated side output features;
[0136] like Figure 5 As shown, the first set of features and the second set of features are first fused through a feedforward channel attention module to obtain a channel attention result, and then the channel attention result is passed to the multi-scale spatial attention module to obtain the first set of attention-modulated side output features, as follows:
[0137] The first set of features is sequentially subjected to a downsampling, an LN layer, and an output channel C with a receptive field size of 1×1. 5The convolution layer is used to obtain the alignment result of the first set of features, and the alignment result of the first set of features is passed to the spatial average pooling and the spatial maximum pooling in parallel to obtain the output results of the spatial average pooling and the output results of the spatial maximum pooling respectively. Then the two output results are successively passed to an output channel with a receptive field size of 1×1 and the number of channels is k×C 5 The convolutional layer, a GELU, a receptive field size of 1×1, and the number of output channels is C 5 The convolution layer is passed through a sigmoid function to obtain two processed results. The two processed results are then added and passed through the sigmoid function to perform Hadamard product with the alignment result of the first set of features, and the result is added to the second set of features to obtain the channel attention result. The channel attention result is then passed in parallel to the channel average pooling and channel maximum pooling and the two results are concatenated. The concatenated result is passed in parallel to a convolution layer with a receptive field size of 3×7 and an output channel number of 1, a convolution layer with a receptive field size of 5×13 and an output channel number of 1, a convolution layer with a receptive field size of 7×3 and an output channel number of 1, and a convolution layer with a receptive field size of 13×5 and an output channel number of 1. The four results are added and passed through the sigmoid function to perform Hadamard product with the channel attention result. The product result is added to the channel attention result and shuffled through the channel to obtain the multi-scale spatial attention result, that is, the side output features of the first set of attention modulation.
[0138] Among them, C 5 Indicates the larger number of channels of the two inputs, and k is set to 4. In the text recognition model, the three FFIAM module structures are consistent with the processing flow.
[0139] S33, passing the obtained first group of attention-modulated side output features and the third group of features into the second FFIAM to obtain the second group of attention-modulated side output features;
[0140] S34. The second set of attention-modulated side output features and the fourth set of features are passed into the third FFIAM to obtain the third set of attention-modulated side output features.
[0141] In this embodiment, step S5 is specifically as follows:
[0142] The probability output of the character recognition in step S4 is set to P cls , whose size is N×M; the text label is represented by Y cls , whose size is N×M.
[0143] Among them, N represents the number of samples in a batch, and M represents the total number of recognized categories.
[0144] Then use the cross entropy loss function to calculate Pcls With Y cls The recognition loss L cls , the calculation expression is as follows:
[0145]
[0146] Among them, y ij represents the true label of sample i in category j, p ij represents the predicted probability of sample i in category j.
[0147] In this embodiment, step S6 is specifically as follows:
[0148] S61, passing the attention-modulated side output features obtained in step S3 and the multi-scale fusion features obtained by the MFP module in step S2 into the EFM module and integrating them to obtain an edge prediction map;
[0149] First, the three groups of attention-modulated side output features obtained in step S3 and the multi-scale fusion features obtained in step S2 are arranged in ascending order of size, and then the first and second groups of features are passed into the first EFM to obtain the output result of the first EFM, as follows:
[0150] like Figure 6 As shown, the two sets of input features are respectively passed through a receptive field size of 5×5 and the number of output channels is C 6 After a convolutional layer, an LN layer, and a GELU, the smaller feature is upsampled to twice the original resolution and added to the other to get the output of the first EFM.
[0151] Among them, C 6 Represents the smaller number of channels of the two inputs.
[0152] Then the output result of the first EFM and the third set of features are passed into the second EFM to obtain the output result of the second EFM. The specific processing flow is the same as that of the first EFM. The output result of the second EFM and the fourth set of features are passed into the third EFM to obtain the output result of the third EFM, that is, the edge prediction map.
[0153] S62, calculating the edge loss through the edge loss function based on the edge prediction image obtained in step S61 and the edge image pseudo label in step S2;
[0154] Set the edge map pseudo label obtained in step S2 to be Y edge =(y j ,j=1,…,|Y edge |),y j ∈[0,1].
[0155] Among them, yj Represents the pixel value at the jth pixel.
[0156] Then Y edge Classified into positive sample set and negative sample set, where the positive sample set and negative sample set are represented as Y + ={y j ,y j >η} and Y - ={y j ,y j =0}, all other pixel values are ignored, η represents the pixel value threshold, which is set to 0.2;
[0157] Then set the edge prediction map obtained in step S61 to be represented as P edge =(p j ,j=1,…,|P edge |),p j ∈[0,1].
[0158] Among them, p j Represents the value after a sigmoid function is processed at the jth pixel.
[0159] Finally, the edge loss function is used to calculate the edge map pseudo label Y edge and edge prediction graph P edge The edge loss L edge , the calculation expression is as follows:
[0160]
[0161] Among them, α and β represent weight coefficients, which are used to balance positive and negative samples, and λ represents the weight of controlling the size of the coefficient, which is set to 1.
[0162] In this embodiment, in step S7, the weighted sum of the identification loss in step S5 and the edge loss in step S6 is as follows:
[0163] Denote the total loss as L total , the calculation expression is as follows:
[0164] L total =L cls +γ·L edge
[0165] Among them, γ represents the weight coefficient, which controls the loss balance and is set to 0.01.
[0166] This embodiment conducted ablation experiments and comparative experiments on the MNIST dataset, and a large number of experimental results proved that the DNNs text recognition model with feedback interaction and the MFP, FBIAM, FFIAM and EFM modules constructed by the present invention can significantly enhance the grating visual illusion perception ability of the model. Table 1 compares the recognition results of the method of the present invention with other methods, showing the quantitative comparison of the method of the present invention with other methods.
[0167] Table 1
[0168]
[0169] As shown in Table 1, the method of the present invention uses TOP1 recognition accuracy as the evaluation index, that is, the category with the maximum model output probability is taken as the recognition result. Compared with other comparison methods, the method of the present invention is significantly higher than other methods, which is conducive to applying the method of the present invention to a wider range of visual tasks such as scene character recognition and camouflage target detection.
[0170] In summary, the method of the present invention constructs a DNNs text recognition model with grating visual illusion perception based on the relevant neural mechanism of grating visual illusion perception, and constructs MFP, FBIAM, FFIAM, and EFM modules at the same time, and uses the fusion of edge loss and recognition loss to guide the DNNs model to learn global contour perception ability during training, that is, to learn global shape preferences rather than local features, and give the model grating visual illusion perception ability, thereby improving the accuracy and robustness of text recognition, and helping to improve the recognition accuracy of text presented in a grating visual illusion manner on printed materials and billboards in scene character recognition tasks, and improve system reliability. At the same time, the method of the present invention can also be applied to camouflage target detection tasks, using the global contour perception properties of visual illusions to improve the accuracy of camouflage target detection.
[0171] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific variations and combinations that do not deviate from the essence of the present invention based on the technical revelations disclosed by the present invention, and these variations and combinations are still within the protection scope of the present invention.
Claims
1. A grating visual illusion text recognition method, the specific steps are as follows: S1. Construct a DNNs text recognition model with grating visual illusion perception; The DNNs text recognition model includes: Text recognition model, edge detection model; The edge detection model is used to provide edge map pseudo labels; The text recognition model includes: 1 MFP module, 4 cortical column modules, 3 FFIAM modules, 3 EFM modules, and 1 classification layer module; Among them, the four cortical column modules are the primary visual cortex V1 cortical column, the secondary visual cortex V2 cortical column and two visual area 4 V4 cortical columns; The primary visual cortex V1 cortical column includes: 1 primary first-stage module, 1 FBIAM module, 1 primary second-stage module, and 1 primary third-stage module; the secondary visual cortex V2 cortical column includes: 1 pooling layer, 1 secondary first-stage module, 1 FBIAM module, 1 secondary second-stage module, and 1 secondary third-stage module; the first visual area IV V4 cortical column includes: 1 pooling layer, 1 area IV first-stage module, 1 FBIAM module, 1 area IV second-stage module, and 1 area IV third-stage module; the second visual area IV V4 cortical column includes: 1 pooling layer, 1 area IV first-stage module, 1 area IV second-stage module, and 1 area IV third-stage module; S2. Use the MNIST training set to train the DNNs text recognition model. First, input the training image into the edge detection model and the text recognition model respectively. The edge detection model obtains the edge map pseudo label, and the text recognition model obtains four sets of side output features and a set of multi-scale fusion features obtained based on the MFP module. S3, passing the four sets of side output features in step S2 into the FFIAM module to obtain three sets of attention-modulated side output features; S4, passing a group of the smallest feature size among the attention-modulated side output features obtained in step S3 into the classification layer module, and sequentially passing through a global average pooling layer, an LN layer, and a fully connected linear layer to obtain a probability output of text recognition; Among them, the probability output of text recognition is the probability of belonging to each category, and the MNIST data set used has a total of 10 categories; S5, calculating the recognition loss by using the cross entropy loss function to calculate the recognition loss of the text recognition probability output in step S4 and the text label; S6, the attention modulated side output features obtained in step S3 and the multi-scale fusion features obtained by the MFP module in step S2 are passed to the EFM module and integrated to obtain an edge prediction map, and then the edge loss is calculated by the edge loss function with the edge map pseudo label in step S2; S7, performing weighted summation of the recognition loss in step S5 and the edge loss in step S6 to obtain a fusion loss function, and then iterating the training cycle to a preset maximum number of training times to complete the training of the DNNs text recognition model and obtain a trained model; S8. Based on the DNNs text recognition model trained in step S7, the MNIST test set is used for testing, that is, the test image is input into the trained model for testing to obtain the text recognition result, thereby completing the grating visual illusion text recognition.
2. The grating optical illusion text recognition method according to claim 1, characterized in that: The step S2 is specifically as follows: S21, passing the training image into the edge detection model to obtain the edge map pseudo label; S22, passing the training image into the MFP module to obtain a set of multi-scale fusion features; The input image is passed into four convolutional layers with a convolution kernel size of 3×3 and a number of output channels of C1, with dilation rates of 2, 4, 6, and 7 respectively. The four output results are added together, and after an LN layer and channel shuffle, the multi-scale fusion features are output; Where, C1=64; S23, sequentially passing the multi-scale fusion features in step S22 through four sub-modules consisting of depthwise separable convolutions with different receptive field sizes, i.e., four cortical column modules, to obtain four sets of side output features; 1) The first submodule is the primary visual cortex V1 cortical column; First, the multi-scale fusion feature is passed into the primary first-stage module to obtain the output feature of the first stage, that is, the multi-scale fusion feature is first passed into a depth-separable convolutional layer with a receptive field size of 5×5 and an output channel number of C1, and then passes through a depth-separable convolutional layer with a receptive field size of 3×3, a void rate of 2, and an output channel number of 2×C1, and then passes through a convolutional layer with a receptive field size of 1×1 and an output channel number of C1, and the obtained output is added to the multi-scale fusion feature in step S22 to obtain the output feature of the first stage; Then the output features of the first stage are passed into the FBIAM module, and the second set of side output features are fed back to the FBIAM module to obtain the output features of the FBIAM module. That is, the two input features are first fused through a feedback channel attention module to obtain the channel attention result, and then the channel attention result is passed into the multi-scale spatial attention module to obtain the output features of the FBIAM module, as follows: The second set of side output features are first upsampled to twice the original resolution through deconvolution, and then passed into a convolution layer with a receptive field size of 1×1 and an output channel number of C and an LN layer in sequence to obtain the alignment results of the second set of side output features; the output features of the first stage are passed into spatial average pooling and spatial maximum pooling in parallel to obtain the output results of spatial average pooling and the output results of spatial maximum pooling, respectively, and then the two output results are passed into a shared convolution layer with a receptive field size of 1×1 and an output channel number of k×C, a GELU, and a shared convolution layer with a receptive field size of 1×1 and an output channel number of C, and then the two processed results are added and passed through the sigmoid function and then Hadamard product is performed with the output features of the first stage to obtain the processed output features of the first stage The result is added to the alignment result of the second set of side output features to obtain the channel attention result; the channel attention result is then passed in parallel to the channel average pooling and channel maximum pooling and the two results are spliced, and the spliced result is passed in parallel to a convolutional layer with a receptive field size of 3×7 and an output channel number of 1, a convolutional layer with a receptive field size of 5×13 and an output channel number of 1, a convolutional layer with a receptive field size of 7×3 and an output channel number of 1, and a convolutional layer with a receptive field size of 13×5 and an output channel number of 1. The four results are added and passed through the sigmoid function, and then the Hadamard product is performed with the channel attention result. The product result is added to the channel attention result and shuffled through the channel to obtain the multi-scale spatial attention result, that is, the output feature of the FBIAM module; Where C represents the smaller number of channels in the two inputs, and k is set to 4; Finally, the output features of the FBIAM module are passed into the primary second-stage module, and the obtained output is added to the output features of the FBIAM module to obtain the output features of the second stage, and the output features are passed into the primary third-stage module to obtain the first set of side output features, that is, the output features of the FBIAM module are passed into a depth-separable convolutional layer with a receptive field size of 4×4, a hole rate of 2, and an output channel number of 2×C1, and then passed through a convolutional layer with a receptive field size of 1×1 and an output channel number of C1, and the obtained output is added to the output features of the FBIAM module to obtain the output features of the second stage, and then the output features of the second stage are passed into a depth-separable convolutional layer with a receptive field size of 4×4, a hole rate of 2, and an output channel number of C1 to obtain the first set of side output features; 2) The second submodule is the cortical column of the secondary visual cortex V2; First, the output features of the second stage in the primary visual cortex V1 cortical column are passed through a pooling layer and then passed to the secondary first stage module to obtain the output features of the first stage, that is, the output features of the second stage in the first submodule are passed through a pooling layer and downsampled to half of the original resolution, and then passed to a deep separable convolutional layer with a receptive field size of 6×6, a void rate of 2, and an output channel number of C2 to obtain the output features of the first stage; Where, C2=128; Then, the output features of the first stage are passed into the FBIAM module, and the third set of side output features are fed back to the FBIAM module to obtain the output features of the FBIAM module; Among them, the FBIAM module structure and processing flow in the primary visual cortex V1 cortical column, the secondary visual cortex V2 cortical column, and the first fourth visual area V4 cortical column were consistent; Finally, the output features of the FBIAM module are passed to the secondary second-stage module, and the output is added to the output features of the FBIAM module to obtain the output features of the second stage, and this output feature is passed to the secondary third-stage module, and the output is added to the output features of the second stage to obtain the second set of side output features, that is, the output features of the FBIAM module are passed to a depth-separable convolutional layer with a receptive field size of 7×7, a hole rate of 2, and an output channel number of C2, and then passed through a receptive field size of 3×3, a hole rate of 6, and an output channel number of The output features of the second stage are obtained by passing the output features of the second stage into a deep separable convolutional layer with a receptive field size of 7×7, a void rate of 2, and an output channel number of 2×C2, and then passing through a convolutional layer with a receptive field size of 1×1 and an output channel number of C2, and adding the output to the output features of the second stage to obtain the second set of side output features. 3) The third submodule is the first visual area V4 cortical column; First, the output features of the second stage in the cortical column of the secondary visual cortex V2 are passed through a pooling layer and then passed to the first stage module of the fourth area to obtain the output features of the first stage, that is, the output features of the second stage in the second submodule are passed through a pooling layer and downsampled to half of the original resolution, and then passed to a deep separable convolutional layer with a receptive field size of 7×7, a void rate of 3, and an output channel number of C3 to obtain the output features of the first stage; Wherein, C3=256; Then, the output features of the first stage are passed into the FBIAM module, and the fourth set of side output features are fed back to the FBIAM module to obtain the output features of the FBIAM module; Finally, the output features of the FBIAM module are passed to the second-stage module in the fourth zone, and the output is added to the output features of the FBIAM module to obtain the output features of the second stage, and this output feature is passed to the third-stage module in the fourth zone, and the output is added to the output features of the second stage to obtain the third set of side output features, that is, the output features of the FBIAM module are passed to a depthwise separable convolutional layer with a receptive field size of 9×9, a dilation rate of 3, and an output channel number of C3, and then passed through a receptive field size of 3×3, a dilation rate of 12, and an output channel number of C3. The output features of the second stage are obtained by passing the output features of the second stage into a deep separable convolutional layer with a receptive field size of 9×9, a void rate of 3, and an output channel number of 2×C3, and then passing through a convolutional layer with a receptive field size of 1×1 and an output channel number of C3, and adding the output to the output features of the second stage to obtain the output features of the second stage; the output features of the second stage are then passed into a deep separable convolutional layer with a receptive field size of 9×9, a void rate of 3, and an output channel number of 2×C3, and then passing through a convolutional layer with a receptive field size of 1×1 and an output channel number of C3, and adding the output to the output features of the second stage to obtain the third set of side output features; 4) the fourth submodule is the second visual area V4 cortical column; First, the output features of the second stage in the first visual area V4 cortical column are passed through a pooling layer and then passed to the first stage module of the fourth area to obtain the output features of the first stage, that is, the output features of the second stage in the third submodule are passed through a pooling layer and downsampled to half of the original resolution, and then passed to a deep separable convolutional layer with a receptive field size of 7×7, a hole rate of 3, and an output channel number of C4 to obtain the output features of the first stage; Among them, C4=512; Then the output features of the first stage are passed to the second stage module of the fourth zone, and the output is added to the output features of the first stage to obtain the output features of the second stage. Finally, the output features of the second stage are passed to the third stage module of the fourth zone, and the output is added to the output features of the second stage to obtain the fourth set of side output features, that is, the output features of the first stage are passed to a depth-separable convolutional layer with a receptive field size of 9×9, a dilation rate of 3, and an output channel number of C4, and then passed through a receptive field size of 3×3, a dilation rate of 12, and an output channel number of C4. The depthwise separable convolutional layer with the number of channels being 2×C4 passes through a convolutional layer with a receptive field size of 1×1 and an output channel number of C4 again, and the output is added to the output features of the first stage to obtain the output features of the second stage; the output features of the second stage are then passed into a depthwise separable convolutional layer with a receptive field size of 9×9, a void rate of 3, and an output channel number of 2×C4, and then pass through a convolutional layer with a receptive field size of 1×1 and an output channel number of C4, and the output is added to the output features of the second stage to obtain the fourth set of side output features.
3. The grating optical illusion text recognition method according to claim 1, characterized in that: The step S3 is specifically as follows: S31, arranging the four groups of side output features obtained in step S2 from large to small according to feature size; S32, based on the sorting result of step S31, passing the first set of features and the second set of features into the first FFIAM to obtain the first set of attention-modulated side output features; The first set of features and the second set of features are first fused through a feedforward channel attention module to obtain a channel attention result, and then the channel attention result is passed to the multi-scale spatial attention module to obtain the first set of attention-modulated side output features, as follows: The first set of features is sequentially passed through a downsampling, an LN layer, and a convolutional layer with a receptive field size of 1×1 and an output channel number of C5 to obtain the alignment results of the first set of features. The alignment results of the first set of features are passed into the spatial average pooling and the spatial maximum pooling in parallel to obtain the output results of the spatial average pooling and the output results of the spatial maximum pooling, respectively. The two output results are then passed into a convolutional layer with a receptive field size of 1×1 and an output channel number of k×C5, a GELU, and a convolutional layer with a receptive field size of 1×1 and an output channel number of C5 to obtain two processed results. The two processed results are then added and passed through the sigmoid function to perform the Hadamard product with the alignment result of the first set of features, and the result is combined with the alignment result of the second set of features. Add to get the channel attention result; then pass the channel attention result in parallel to the channel average pooling and channel maximum pooling and splice the two results, and pass the spliced result in parallel to a convolutional layer with a receptive field size of 3×7 and an output channel number of 1, a convolutional layer with a receptive field size of 5×13 and an output channel number of 1, a convolutional layer with a receptive field size of 7×3 and an output channel number of 1, and a convolutional layer with a receptive field size of 13×5 and an output channel number of 1. Add the four results and pass them through the sigmoid function and then perform the Hadamard product with the channel attention result. Add the product result to the channel attention result and shuffle the channel to get the multi-scale spatial attention result, that is, the side output features of the first group of attention modulation; Wherein, C5 represents the larger number of channels of the two inputs, and k is set to 4; in the text recognition model, the structure of the three FFIAM modules is consistent with the processing flow; S33, passing the obtained first group of attention-modulated side output features and the third group of features into the second FFIAM to obtain the second group of attention-modulated side output features; S34. The second set of attention-modulated side output features and the fourth set of features are passed into the third FFIAM to obtain the third set of attention-modulated side output features.
4. The method for recognizing grating visual illusion characters according to claim 1, characterized in that: The step S5 is specifically as follows: The probability output of the character recognition in step S4 is set to P cls , whose size is N×M; the text label is represented by Y cls , whose size is N×M; Among them, N represents the number of samples in a batch, and M represents the total number of recognized categories; Then use the cross entropy loss function to calculate P cls With Y cls The recognition loss L cls , the calculation expression is as follows: Among them, y ij represents the true label of sample i in category j, p ij represents the predicted probability of sample i in category j.
5. The method for recognizing grating visual illusion characters according to claim 1, characterized in that: The step S6 is specifically as follows: S61, passing the attention-modulated side output features obtained in step S3 and the multi-scale fusion features obtained by the MFP module in step S2 into the EFM module and integrating them to obtain an edge prediction map; First, the three groups of attention-modulated side output features obtained in step S3 and the multi-scale fusion features obtained in step S2 are arranged in ascending order of size, and then the first and second groups of features are passed into the first EFM to obtain the output result of the first EFM, as follows: The two sets of input features are sequentially passed through a convolutional layer with a receptive field size of 5×5 and an output channel number of C6, an LN layer, and a GELU. The smaller feature is upsampled to twice the original resolution and added to the other to obtain the output result of the first EFM. Where C6 represents the smaller number of channels in the two inputs; Then, the output result of the first EFM and the third set of features are passed into the second EFM to obtain the output result of the second EFM. The specific processing flow is the same as that of the first EFM. The output result of the second EFM and the fourth set of features are passed into the third EFM to obtain the output result of the third EFM, that is, the edge prediction map. S62, calculating the edge loss through the edge loss function based on the edge prediction image obtained in step S61 and the edge image pseudo label in step S2; Set the edge map pseudo label obtained in step S2 to be Y edge =(y j ,j=1,…,|Y edge |),y j ∈[0,1]; Among them, y j represents the pixel value at the jth pixel; Then Y edge Classified into positive sample set and negative sample set, where the positive sample set and negative sample set are represented as Y + ={y j ,y j >η} and Y - ={y j ,y j =0}, all other pixel values are ignored, η represents the pixel value threshold, which is set to 0.2; Then set the edge prediction map obtained in step S61 to be represented as P edge =(p j ,j=1,…,|P edge |),p j ∈[0,1]; Among them, p j Represents the value after a sigmoid function is processed at the jth pixel; Finally, the edge loss function is used to calculate the edge map pseudo label Y edge and edge prediction graph P edge The edge loss L edge , the calculation expression is as follows: Among them, α and β represent weight coefficients, and λ represents the weight of the control coefficient size, which is set to 1.
6. The method for recognizing grating visual illusion characters according to claim 1, characterized in that: In step S7, the weighted sum of the recognition loss in step S5 and the edge loss in step S6 is as follows: Denote the total loss as L total , the calculation expression is as follows: L total =L cls +γ·L edge Among them, γ represents the weight coefficient, which is set to 0.01.
Citation Information
Patent Citations
Construction method of image-text recognition model
CN117333884A
Electrode catalyst layer for polymer electrolyte membrane fuel cell using low-temperature substrate sputtering and manufacturing method thereof
KR1020250012447A
OCR image sample generation method and apparatus, print font verification method and apparatus, and device and medium
WO2021212658A1