Image recognition method, device and equipment based on frequency division feature extraction
Through the bilateral parameterized activation function and frequency-dividing feature extraction method, the high-frequency details and low-frequency structure of the image are explicitly separated, solving the problem of insufficient frequency information utilization in traditional convolutional neural networks, and achieving more efficient text feature extraction and recognition.
Patent Information
- Application Number
- CN202510647717.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-08-08
AI Technical Summary
Traditional convolutional neural networks lack explicit division and utilization of frequency information in feature extraction, resulting in limited performance improvement in complex scenarios and difficult to explain.
The frequency division feature extraction method based on the two-sided parameterized activation function is adopted to explicitly separate the high-frequency details and low-frequency structure of the image through multi-directional translation operation, and a frequency division improvement scheme of channel-by-channel or shared parameters is constructed, high-frequency and low-frequency branches are generated, and they are input to the neural network for text classification and recognition.
It improves the analytical accuracy of text features and interpretability of neural networks, enhances the nonlinear computing ability of feature extraction, breaks through the bottleneck of linear transformation of traditional convolution layers, and improves the accuracy of text recognition in complex scenarios.
Smart Images

Figure CN120449860A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image recognition technology, and in particular to an image recognition method, device and equipment based on frequency division feature extraction. Background Art
[0002] Optical character recognition technology is widely used in the field of artificial intelligence, playing a particularly important role in document recognition, bill recognition, and natural scene text recognition. Text images are characterized by a clear distinction between high-frequency information, such as text edges and textures, and low-frequency information, such as the overall text outline. Therefore, introducing a neural network that explicitly utilizes frequency information is crucial for improving text recognition accuracy.
[0003] However, traditional convolutional neural networks mainly rely on linear transformations for feature extraction, lack explicit division and utilization of frequency information, and their internal operating mechanisms are opaque and difficult to explain, limiting their performance in complex scenarios. Existing improvement schemes attempt to decompose feature maps into multiple components through translation operations using neural networks, and use prediction and update operators to achieve frequency division, which is equivalent to the linear transformation of the convolutional layer. However, existing technologies only process through linear operators, and ultimately simply add the high-frequency and low-frequency branches before outputting them. In essence, it is still a linear operation without introducing nonlinear transformations. The feature extraction capability is limited, and the utilization of frequency information is insufficient, failing to effectively enhance key frequency features. Summary of the Invention
[0004] The present invention provides an image recognition method, apparatus and device based on frequency division feature extraction to divide, process and fuse the frequency information in the image to be recognized, thereby improving the efficiency and ability of feature extraction and enhancing the interpretability of the neural network.
[0005] According to one aspect of the present invention, there is provided an image recognition method based on frequency division feature extraction, the method comprising:
[0006] Acquire an image to be identified, and obtain each translation component according to the image to be identified;
[0007] Construct a bilateral parameterized activation function, and construct various lifting schemes based on the bilateral parameterized activation function and each translation component;
[0008] Determine a target implementation method corresponding to the image to be recognized, and generate high-frequency branches and low-frequency branches by predicting and updating according to each boosting scheme and target implementation method, wherein the target implementation method includes channel-by-channel deep frequency boosting or shared parameter frequency boosting;
[0009] The output features are calculated based on each high-frequency branch and low-frequency branch, and the output features are input into the neural network for text classification and recognition, and the recognition results are output.
[0010] Optionally, obtaining each translation component according to the image to be identified includes: inputting the image to be identified into a neural network, obtaining feature maps of different levels through convolution and pooling operations in the neural network; selecting a feature map of any level as a target feature map; and performing a translation operation on the target feature map to obtain four translation components.
[0011] Optionally, constructing a bilateral parameterized activation function includes: determining a mathematical form of the activation function based on the ReLU function, wherein the mathematical form includes bilateral learnable parameters; using a backpropagation algorithm to calculate the gradient of the loss function with respect to the bilateral learnable parameters, and updating the bilateral learnable parameters based on the gradient and a preset optimization algorithm to generate a bilateral parameterized activation function.
[0012] Optionally, each lifting scheme is constructed based on the bilateral parameterized activation function and each translation component, including: taking one of the translation components as the reference component, and forming each lifting pair with the reference component and the other components; setting a prediction operator and an update operator for each lifting pair, and using the bilateral parameterized activation function as the activation function of the prediction operator and the update operator to generate each lifting scheme.
[0013] Optionally, when the target implementation method is channel-by-channel deep frequency division enhancement, prediction and updating are performed according to each enhancement scheme and target implementation method to generate each high-frequency branch and low-frequency branch, including: configuring the operator parameters of the prediction operator and the update operator for each input channel of the neural network to generate an independent prediction operator and an independent update operator corresponding to each input channel; performing separate operations based on each independent prediction operator and independent update operator on each input channel to generate each high-frequency branch and low-frequency branch.
[0014] Optionally, when the target implementation method is frequency division boosting of shared parameters, prediction and updating are performed according to each boosting scheme and target implementation method to generate each high-frequency branch and low-frequency branch, including: jointly configuring the operator parameters of the prediction operator and the update operator for each input channel of the neural network to generate a shared prediction operator and a shared update operator; and performing prediction and updating based on the shared prediction operator and the shared update operator through each input channel to generate each high-frequency branch and low-frequency branch.
[0015] Optionally, the output features are calculated based on each high-frequency branch and the low-frequency branch, including: summing each high-frequency branch to obtain a high-frequency component; using an enhancement module to enhance the high-frequency component to obtain an enhanced high-frequency component; and multiplying the enhanced high-frequency component with the low-frequency branch to obtain the output features.
[0016] According to another aspect of the present invention, there is provided an image recognition device based on frequency division feature extraction, the device comprising:
[0017] A translation component generation module is used to obtain an image to be identified and obtain various translation components according to the image to be identified;
[0018] A lifting scheme construction module is used to construct a bilateral parameterized activation function, and to construct each lifting scheme based on the bilateral parameterized activation function and each translation component;
[0019] A frequency division feature extraction module is used to determine the target implementation method corresponding to the image to be recognized, and to generate high-frequency branches and low-frequency branches by predicting and updating according to various lifting schemes and target implementation methods, where the target implementation methods include channel-by-channel deep frequency division lifting or shared parameter frequency division lifting;
[0020] The output feature calculation module is used to calculate the output features based on each high-frequency branch and low-frequency branch, input the output features into the neural network for text classification and recognition, and output the recognition results.
[0021] According to another aspect of the present invention, an electronic device is provided, comprising:
[0022] at least one processor;
[0023] and a memory communicatively coupled to the at least one processor;
[0024] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the image recognition method based on frequency division feature extraction described in any embodiment of the present invention.
[0025] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement an image recognition method based on frequency division feature extraction as described in any embodiment of the present invention when executed.
[0026] The technical solution of the embodiment of the present invention constructs an improvement scheme through bilateral parameterized activation functions, which can explicitly separate the high-frequency details and low-frequency structures of the image based on multi-directional translation operations, thereby improving the parsing accuracy of text features. By constructing bilateral parameterized activation functions and integrating the improvement scheme, it is possible to inject nonlinear computing capabilities into the feature extraction process, breaking through the bottleneck that the traditional convolutional layer can only achieve linear transformation, allowing the model to adaptively adjust the efficiency of positive and negative interval signal transmission and enhance the flexibility of feature expression. Through the frequency division and improvement implementation method of channel-by-channel or shared parameters, it is possible to specifically extract the specific frequency features of each channel or cross-channel interaction features.
[0027] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0029] Figure 1 This is a flowchart of an image recognition method based on frequency division feature extraction provided in accordance with the first embodiment of the present invention;
[0030] Figure 2 1 is a schematic diagram of a calculation process of a channel-by-channel deep frequency division boosting method provided in accordance with the first embodiment of the present invention;
[0031] Figure 3 1 is a schematic diagram of a calculation process of a frequency division and boosting method with shared parameters provided in accordance with the first embodiment of the present invention;
[0032] Figure 4 is a flowchart of another image recognition method based on frequency division feature extraction provided according to the second embodiment of the present invention;
[0033] Figure 5 2 is a schematic structural diagram of an image recognition device based on frequency division feature extraction according to a third embodiment of the present invention;
[0034] Figure 6 The present invention is a schematic structural diagram of an electronic device for implementing an image recognition method based on frequency division feature extraction according to an embodiment of the present invention. DETAILED DESCRIPTION
[0035] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0036] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0037] Example 1
[0038] Figure 1 A flowchart of an image recognition method based on frequency division feature extraction is provided for the first embodiment of the present invention. This embodiment is applicable to image optical character recognition scenarios. The method can be executed by an image recognition device based on frequency division feature extraction. The image recognition device based on frequency division feature extraction can be implemented in the form of hardware and / or software. The image recognition device based on frequency division feature extraction can be configured in a computer controller. Figure 1 As shown, the method includes:
[0039] S110 , obtaining an image to be identified, and obtaining each translation component according to the image to be identified.
[0040] The translation component refers to the sub-feature map obtained by translating the feature map of the image to be identified according to the pixel position. In the wavelet lifting framework, the feature map is usually divided into four translation components.
[0041] Optionally, obtaining each translation component according to the image to be identified includes: inputting the image to be identified into a neural network, obtaining feature maps of different levels through convolution and pooling operations in the neural network; selecting a feature map of any level as a target feature map; and performing a translation operation on the target feature map to obtain four translation components.
[0042] Specifically, the image to be recognized can be input into a neural network front-end, such as a ResNet or MobileNet architecture. Features are extracted layer by layer through successive convolutional and pooling layers. For example, the convolutional layer uses a 3×3 convolution kernel with a stride of 1 and padding of 1 to convolve the input image to extract local features. The pooling layer uses a 2×2 max pooling with a stride of 2 to downsample the convolution output, reducing the resolution of the feature map and reducing computational effort. Through multiple layers of convolution, activation, and pooling operations, feature maps at different levels are ultimately generated. The controller then selects the output of a particular layer within the neural network's intermediate layers as the target feature map. The controller refers to the computer controller that performs image recognition. The target feature map must meet size requirements, typically a three-dimensional tensor of H×W×C, where H and W are the height and width of the feature map, and C is the number of channels. Both H and W must be greater than or equal to 2 to facilitate subsequent translation operations. For example, if the input image size is 224×224, after several layers of convolution and pooling, the intermediate layer feature map with a size of 28×28×64 can be selected as the target feature map.
[0043] In a specific embodiment, four translation operations can be performed on the target feature map to generate four components. The original component can be directly amplified using the target feature map without translation. The target feature map is translated to the right by 1 pixel, and the left boundary is processed by filling 0 or copying edge pixels to obtain the right-shifted component. The same is true for the down-shifted component and the right-down-shifted component. Each translation component is expressed as:
[0044] I0=(x0,y0)=(2x,2y)
[0045] I1=(x1,y1)=(2x+1,2y)
[0046] I2=(x2,y2)=(2x,2y+1)
[0047] I3=(x3, y3)=(2x+1, 2y+1);
[0048] Among them, x and y represent the horizontal and vertical coordinates of the target feature map, I0 represents the original component, x0 and y0 represent the horizontal and vertical coordinates of the original component, I1 represents the right shift component, x1 and y1 represent the horizontal and vertical coordinates of the right shift component, I2 is the downward shift component, x2 and y2 represent the horizontal and vertical coordinates of the downward shift component, I3 is the right downward shift component, and x3 and y3 represent the horizontal and vertical coordinates of the right downward shift component.
[0049] S120: Construct a bilateral parameterized activation function, and construct each lifting scheme according to the bilateral parameterized activation function and each translation component.
[0050] Among them, the Bilateral Parametric Rectified Linear Unit (BPReLU) is an improved activation function. By constructing a bilaterally parameterized activation function, the nonlinear expression capability of the neural network can be enhanced, replacing the linear operator in the traditional lifting scheme. By adaptively adjusting the gradient between positive and negative intervals through learnable parameters, the feature extraction efficiency is improved. The lifting scheme refers to the wavelet transform method, which consists of three steps: splitting, prediction, and updating. It can achieve multi-scale analysis of signals without the Fourier transform. The lifting scheme is used to separate the translation component into high-frequency and low-frequency information.
[0051] Optionally, constructing a bilateral parameterized activation function includes: determining a mathematical form of the activation function based on the ReLU function, wherein the mathematical form includes bilateral learnable parameters; using a backpropagation algorithm to calculate the gradient of the loss function with respect to the bilateral learnable parameters, and updating the bilateral learnable parameters based on the gradient and a preset optimization algorithm to generate a bilateral parameterized activation function.
[0052] It can be known that the traditional ReLU function expression is:
[0053]
[0054] Where f(x) is the output value after activation by the ReLU function, and x is the input value. The output of the negative interval in the ReLU function is always 0, which results in the neuron's gradient being 0 in the negative region and the slope being fixed to 1 in the positive interval, lacking flexibility. The bilateral parameterized activation function introduces learnable parameters for both the positive and negative intervals based on the traditional ReLU. The expression is:
[0055]
[0056] Here, α is a learnable parameter in the negative interval that controls the leakage of negative signals. β is a learnable parameter in the positive interval that allows for adjustment of the amplification of positive signals. f(x) is the output value after BPReLU activation, and x is the input value, i.e., the output value of the previous layer in the neural network. While the linear output is retained in the positive interval, a learnable parameter, the positive slope parameter, is introduced to allow the network to automatically adjust the amplification or reduction of positive signals. In traditional ReLU, this parameter is fixed to 1. Furthermore, the output of the negative interval is no longer forced to be 0. Instead, another learnable parameter is introduced to allow a certain proportion of negative signals to leak out. In traditional ReLU, this parameter is set to 0. Ultimately, the mathematical form of the activation function becomes: the positive and negative intervals are each controlled by two independent learnable parameters, forming a bilateral parameterization characteristic.
[0057] It's important to note that during neural network training, learnable parameters must be adjusted based on the loss function. First, forward propagation is performed. When the input signal is positive, the signal is scaled using a positive slope parameter; when the input is negative, the signal is scaled using a negative slope parameter, resulting in the output of the activation function. Loss calculation is then performed, comparing the network's predictions with the true labels to calculate the loss value, for example, using cross-entropy loss. Finally, backpropagation is performed, starting from the loss value and working backward to derive the gradient of each parameter. The gradient refers to the degree to which each parameter affects the loss. The gradient of the positive slope parameter reflects the contribution of all positive input signals to the loss, while the gradient of the negative slope parameter reflects the contribution of all negative input signals to the loss.
[0058] Furthermore, after obtaining the parameter gradient, the controller will also use a preset optimization algorithm, such as stochastic gradient descent, Adam, etc. to update the parameter value. For example, by initializing the positive slope parameter to 1 and the negative slope parameter to a smaller positive number, such as 0.25, excessive signal suppression in the negative area is avoided. Then, the optimization algorithm is used to calculate the parameter update amount based on the gradient size and the preset learning rate. For example, stochastic gradient descent directly uses the gradient multiplied by the learning rate as the update amount; the Adam algorithm combines historical gradient information and adaptively adjusts the update step size. Finally, the current parameter value is subtracted from the update amount to obtain the new parameter value, and the process of forward propagation, loss calculation, backpropagation and parameter update is repeated until the network loss converges or the preset number of training times is reached.
[0059] In summary, by setting bilateral learnable parameters, the network can automatically adapt to different data distributions and optimize the transmission efficiency of positive and negative signals. For example, in text recognition, a negative slope parameter can help the network better capture low-frequency information at the edges of text, while a positive slope parameter can enhance high-frequency details of text structure.
[0060] Optionally, each lifting scheme is constructed based on the bilateral parameterized activation function and each translation component, including: taking one of the translation components as the reference component, and forming each lifting pair with the reference component and the other components; setting a prediction operator and an update operator for each lifting pair, and using the bilateral parameterized activation function as the activation function of the prediction operator and the update operator to generate each lifting scheme.
[0061] Specifically, the controller first selects the original component (i.e., the untranslated feature map) from the four translated components as the reference component. The reference component is then paired with the other three translated components to form three lifted pairs. Each lifted pair is used to extract feature differences in a specific direction: the reference component and the right-shifted component for horizontal relationships, the reference component and the downward-shifted component for vertical relationships, and the reference component and the right-downward-shifted component for diagonal relationships.
[0062] It should be noted that in each boosted pair, the controller uses the non-reference component to predict the value of the reference component. When setting the prediction operator, the pixel values of the non-reference components are linearly combined using learnable weights and biases. The result of this linear combination is then fed into the BPReLU, where the predicted value is nonlinearly adjusted using learnable slope parameters in positive and negative ranges. The slope parameter in the positive range allows the network to automatically optimize the amplification of positive signals. The slope parameter in the negative range allows a certain proportion of negative signals to be retained, avoiding the zero-value truncation problem of traditional ReLU. When setting the update operator, the predicted difference can be used to update the non-reference component to generate a low-frequency component. Specifically, the residual obtained by subtracting the predicted value from the reference component represents high-frequency details. This residual is then adjusted using learnable weights and added to the non-reference component. The BPReLU is then nonlinearly activated to generate the updated low-frequency component.
[0063] Ultimately, the three resulting lifting pairs process features in the horizontal, vertical, and diagonal directions, respectively. This allows for comprehensive capture of common line directions in text images and improves the completeness of feature extraction. This includes a horizontal lifting scheme, which uses the right-shifted component to predict the original component, extracting high-frequency horizontal details, such as the differences between the left and right edges of text. The right-shifted component is then updated with the high-frequency details to generate low-frequency horizontal contours, highlighting the horizontal structure of text. A vertical lifting scheme uses the downward-shifted component to predict the original component, extracting high-frequency vertical details, such as the differences between the top and bottom edges of text. The downward-shifted component is then updated with the high-frequency details to generate low-frequency vertical contours, highlighting the vertical structure of text. A diagonal lifting scheme uses the right-downward-shifted component to predict the original component, extracting high-frequency diagonal details, such as differences at text corners. Finally, the right-downward-shifted component is updated with the high-frequency details to generate low-frequency diagonal contours, highlighting text corners and diagonal structure.
[0064] S130. Determine a target implementation method corresponding to the image to be recognized, and perform prediction and update according to each enhancement scheme and target implementation method to generate each high-frequency branch and low-frequency branch, wherein the target implementation method includes channel-by-channel deep frequency division enhancement or shared parameter frequency division enhancement.
[0065] The DepthWise Frequency-Dividing Lifting Scheme module (FDLS-DW) is a lightweight implementation of FDLS. Each input channel uses a separate set of prediction and update operators, employing per-channel computation. This approach reduces the number of parameters and makes it suitable for lightweight models. The Frequency-Dividing Lifting Scheme module with Shared parameters (FDLS-Share) is a standard implementation of FDLS. Prediction operators are independent per channel, while update operators are shared across channels, similar to a convolutional layer. The update operators in FDLS-Share are computed across channels, reducing the number of parameters while maintaining cross-channel interaction. The high-frequency branch refers to the detailed information extracted by the prediction step of the lifting scheme, such as edges and textures, corresponding to the high-frequency components of the image. The low-frequency branch refers to the overall contour information retained by the update step, corresponding to the low-frequency components of the image. The high-frequency branch is used to enhance detail recognition, while the low-frequency branch is used to preserve structural features. Combining the two improves text recognition accuracy and interpretability.
[0066] Optionally, when the target implementation method is channel-by-channel deep frequency division enhancement, prediction and updating are performed according to each enhancement scheme and target implementation method to generate each high-frequency branch and low-frequency branch, including: configuring the operator parameters of the prediction operator and the update operator for each input channel of the neural network to generate an independent prediction operator and an independent update operator corresponding to each input channel; performing separate operations based on each independent prediction operator and independent update operator on each input channel to generate each high-frequency branch and low-frequency branch.
[0067] When the goal is to achieve channel-by-channel deep frequency division and boosting, the controller can split the input feature map into multiple independent channels. For each channel, a set of prediction operator and update operator parameters are set separately. The prediction operator is used to predict the value of the reference component based on the translation component, and the prediction operator parameters for each channel are learned independently. The update operator is used to update the translation component based on the prediction residual, and the update operator parameters for each channel are also learned independently.
[0068] Figure 2A schematic diagram of the operation process of a channel-by-channel deep frequency division and lifting method is provided for the first embodiment of the present invention. In this embodiment, the input feature map represents the original feature map data of the image to be identified. In FDLS-DW, it is split into multiple single-channel feature maps by channel, and each channel is subsequently processed independently. For each single-channel input feature map, the prediction operator attempts to predict the reference component based on the translation component through linear transformation and activation function operation to obtain an intermediate feature map. The intermediate feature map is obtained after the prediction operator operation and is the intermediate result of the prediction process. Its size is the same as the input feature map, but its content is the features processed by the prediction operator, representing the prediction result of the input feature map based on the translation component. The update operator takes the predicted residual (i.e., the difference between the reference component and the predicted value of the input feature map) as input, performs a weighted adjustment on the translation component, generates a low-frequency component by adding it to the translation component and processing it through the activation function, and obtains the final output feature map after the update operator operation.
[0069] Specifically, when using the channel-by-channel deep frequency division boosting module FDLS-DW, each input channel has a separate set of prediction and update operators. After three prediction steps and three update steps, three high-frequency branches and one low-frequency branch are obtained. The specific method is as follows:
[0070]
[0071]
[0072] in, represents the high-frequency branch, I l represents the low-frequency branch, P i and U i Represents the i-th group of prediction operators and update operators, I0 represents the original component, I i Represents other components besides the original components. Since the prediction and update operators are performed separately on each input channel, the high-frequency branch and frequency division branch I l The number of channels is C in , at this time, you need to add a 1×1 convolution layer to transform the number of channels into output channels C out .
[0073] Optionally, when the target implementation method is frequency division boosting of shared parameters, prediction and updating are performed according to each boosting scheme and target implementation method to generate each high-frequency branch and low-frequency branch, including: jointly configuring the operator parameters of the prediction operator and the update operator for each input channel of the neural network to generate a shared prediction operator and a shared update operator; and performing prediction and updating based on the shared prediction operator and the shared update operator through each input channel to generate each high-frequency branch and low-frequency branch.
[0074] Figure 3 A schematic diagram of the operation process of a frequency division and boosting method with shared parameters is provided for the first embodiment of the present invention. Among them, the input feature map represents the original feature map data of the image to be identified. In FDLS-Share, it will be subsequently subjected to component splitting based on different translation operations and other processing. For each input feature map, or its translation component, the prediction operator attempts to predict the reference component based on the translation component through linear transformation and activation function operation to obtain an intermediate feature map. The intermediate feature map is obtained after the prediction operator is operated, and represents the prediction result of the input feature map based on the translation component. The black mesh structure represents the characteristics of shared parameters. In FDLS-Share, the parameters of the update operator are shared across channels, that is, subsequent operations of different channels will use the same set of update operator parameters. Different from the independent update operator parameters of each channel in the channel-by-channel deep frequency division and boosting, the sharing mechanism can reduce the number of parameters, improve computational efficiency, and also promote information interaction between channels to a certain extent. The update operator takes the predicted residual, that is, the difference between the reference component of the input feature map and the predicted value, as input, performs weighted adjustment on the translation component, generates a low-frequency component by adding it to the translation component and processing it through the activation function, and obtains the final output feature map after the update operator operation.
[0075] Specifically, when using the shared parameter frequency division boosting module FDLS-Share, each input channel has a separate set of prediction operators. After three prediction steps, three channels with the number C are obtained. in The high-frequency branch of:
[0076]
[0077] in, represents the high-frequency branch, P i Represents the i-th group of prediction operators, I0 represents the original component. The update operator is cross-channel, similar to the convolution layer, each output channel corresponds to C in An update operator, C in The update operator acts on C channels in The high-frequency branch Obtained C in The feature maps are added to get the output feature map of a channel. Therefore, C out The output channels contain C out ×C in Update operator. At this time, a 1×1 convolution layer is needed to transform the number of channels of I0 into C out , the prediction steps are as follows:
[0078]
[0079] in, represents the high-frequency branch, I lrepresents the low-frequency branch, P i and U i They represent the i-th group of prediction operators and update operators respectively, and I0 represents the original component.
[0080] S140: Calculate output features based on each high-frequency branch and low-frequency branch, input the output features into a neural network for text classification and recognition, and output the recognition results.
[0081] Among them, the output feature refers to the feature map after the fusion of high-frequency and low-frequency information, which is input into the neural network after adjusting the weight through the SE module. The SE module is an attention mechanism module that can adaptively adjust the channel weights of the feature map. In this embodiment, the SE module can be applied to the high-frequency branch to extract the channel-level attention vector. The attention vector is then multiplied by the low-frequency branch to enhance the weight of the key frequency component and optimize the feature fusion effect. Text classification and recognition refers to the use of a neural network, such as a fully connected layer, to map the output features to the label space to complete the classification and recognition of text content, replacing the serial workflow of traditional OCR and improving the robustness in complex scenarios.
[0082] The technical solution of the embodiment of the present invention constructs an improvement scheme through bilateral parameterized activation functions, which can explicitly separate the high-frequency details and low-frequency structures of the image based on multi-directional translation operations, thereby improving the parsing accuracy of text features. By constructing bilateral parameterized activation functions and integrating the improvement scheme, it is possible to inject nonlinear computing capabilities into the feature extraction process, breaking through the bottleneck that the traditional convolutional layer can only achieve linear transformation, allowing the model to adaptively adjust the efficiency of positive and negative interval signal transmission and enhance the flexibility of feature expression. Through the frequency division and improvement implementation method of channel-by-channel or shared parameters, it is possible to specifically extract the specific frequency features of each channel or cross-channel interaction features.
[0083] Example 2
[0084] Figure 4 This is a flowchart of an image recognition method based on frequency division feature extraction provided by the second embodiment of the present invention. This embodiment adds a specific process of calculating output features based on each high-frequency branch and low-frequency branch on the basis of the above-mentioned first embodiment. Among them, the specific contents of steps S210-S230 are roughly the same as those of steps S110-S130 in the first embodiment, so they will not be repeated in this embodiment. Figure 4 As shown, the method includes:
[0085] S210 , obtaining an image to be identified, and obtaining each translation component according to the image to be identified.
[0086] Optionally, obtaining each translation component according to the image to be identified includes: inputting the image to be identified into a neural network, obtaining feature maps of different levels through convolution and pooling operations in the neural network; selecting a feature map of any level as a target feature map; and performing a translation operation on the target feature map to obtain four translation components.
[0087] S220: Construct a bilateral parameterized activation function, and construct each lifting scheme according to the bilateral parameterized activation function and each translation component.
[0088] Optionally, constructing a bilateral parameterized activation function includes: determining a mathematical form of the activation function based on the ReLU function, wherein the mathematical form includes bilateral learnable parameters; using a backpropagation algorithm to calculate the gradient of the loss function with respect to the bilateral learnable parameters, and updating the bilateral learnable parameters based on the gradient and a preset optimization algorithm to generate a bilateral parameterized activation function.
[0089] Optionally, each lifting scheme is constructed based on the bilateral parameterized activation function and each translation component, including: taking one of the translation components as the reference component, and forming each lifting pair with the reference component and the other components; setting a prediction operator and an update operator for each lifting pair, and using the bilateral parameterized activation function as the activation function of the prediction operator and the update operator to generate each lifting scheme.
[0090] S230. Determine a target implementation method corresponding to the image to be recognized, and perform prediction and update according to each enhancement scheme and target implementation method to generate each high-frequency branch and low-frequency branch, wherein the target implementation method includes channel-by-channel deep frequency division enhancement or shared parameter frequency division enhancement.
[0091] Optionally, when the target implementation method is channel-by-channel deep frequency division enhancement, prediction and updating are performed according to each enhancement scheme and target implementation method to generate each high-frequency branch and low-frequency branch, including: configuring the operator parameters of the prediction operator and the update operator for each input channel of the neural network to generate an independent prediction operator and an independent update operator corresponding to each input channel; performing separate operations based on each independent prediction operator and independent update operator on each input channel to generate each high-frequency branch and low-frequency branch.
[0092] Optionally, when the target implementation method is frequency division boosting of shared parameters, prediction and updating are performed according to each boosting scheme and target implementation method to generate each high-frequency branch and low-frequency branch, including: jointly configuring the operator parameters of the prediction operator and the update operator for each input channel of the neural network to generate a shared prediction operator and a shared update operator; and performing prediction and updating based on the shared prediction operator and the shared update operator through each input channel to generate each high-frequency branch and low-frequency branch.
[0093] S240 , summing up the high-frequency branches to obtain a high-frequency component.
[0094] Specifically, in the channel-by-channel deep crossover boost or shared parameter crossover boost mode, each channel will generate a high-frequency branch. The three high-frequency branches are summed to obtain a high-frequency component that better reflects the two-dimensional characteristics:
[0095]
[0096] Among them, I h represents the high frequency component, and Represents each high-frequency branch. High-frequency branch details may include information such as edges and textures in the image. Different channels may focus on high-frequency features of different types or directions. The summation operation aggregates them together to obtain an overall high-frequency component.
[0097] S250: Use an enhancement module to enhance the high-frequency component to obtain an enhanced high-frequency component.
[0098] Among them, the enhancement module refers to the SE module. For example, the enhancement module will perform a global average pooling operation on the obtained high-frequency components, compress the high-frequency components in the spatial dimension, and obtain a vector containing only the channel number dimension, with the purpose of obtaining the overall information of each channel. Then, the vector is input into the fully connected layer. The fully connected layer processes the input vector through a learnable weight matrix, first reducing the dimension, and then raising the dimension back to the same number of channels as the original high-frequency component through a fully connected layer. Finally, the activation function is used to process the vector after dimensionality increase, controlling the value between 0 and 1 to obtain the channel attention weights, and these weights are multiplied by the original high-frequency components to highlight the key channels and enhance the high-frequency components.
[0099] S260: Multiply the enhanced high-frequency component by the low-frequency branch to obtain output features.
[0100] Specifically, the purpose of multiplying the enhanced high-frequency component with the low-frequency branch is to fuse the enhanced high-frequency detail information and the low-frequency structural information, so that the final output features have both the general structure of the image and rich details, which can be used for calculation or task processing in subsequent network layers.
[0101] S270: Input the output features into a neural network for text classification and recognition, and output the recognition results.
[0102] The technical solution of the embodiment of the present invention, by fusing high-frequency and low-frequency branches to generate output features, can deeply integrate the detailed and structural information of text images. Combined with the classification capabilities of neural networks, it can achieve accurate text recognition in complex scenarios. The high-frequency branch's attention enhancement mechanism further highlights key features, while the low-frequency branch preserves the overall structure. The synergistic effect of the two improves the model's robustness to complex text, such as those with multiple fonts and low contrast.
[0103] Example 3
[0104] Figure 5 This is a structural diagram of an image recognition device based on frequency division feature extraction provided by the third embodiment of the present invention. Figure 5 As shown, the device includes: a translation component generation module 310, which is used to obtain an image to be identified and obtain various translation components according to the image to be identified;
[0105] A lifting scheme construction module 320 is used to construct a bilateral parameterized activation function and construct each lifting scheme according to the bilateral parameterized activation function and each translation component;
[0106] The frequency division feature extraction module 330 is used to determine the target implementation method corresponding to the image to be recognized, and to generate high-frequency branches and low-frequency branches by predicting and updating according to various lifting schemes and target implementation methods, wherein the target implementation methods include channel-by-channel deep frequency division lifting or shared parameter frequency division lifting;
[0107] The output feature calculation module 340 is used to calculate output features based on each high-frequency branch and low-frequency branch, input the output features into the neural network for text classification and recognition, and output the recognition results.
[0108] Optionally, the translation component generation module 310 is specifically used to: input the image to be identified into the neural network, obtain feature maps of different levels through convolution and pooling operations in the neural network; select the feature map of any level as the target feature map; and perform a translation operation on the target feature map to obtain four translation components.
[0109] Optionally, the improvement scheme construction module 320 specifically includes an activation function construction unit, which is used to: determine the mathematical form of the activation function based on the ReLU function, wherein the mathematical form includes bilateral learnable parameters; use the back propagation algorithm to calculate the gradient of the loss function with respect to the bilateral learnable parameters, and update the bilateral learnable parameters based on the gradient and the preset optimization algorithm to generate a bilateral parameterized activation function.
[0110] Optionally, the lifting scheme construction module 320 specifically includes a lifting scheme construction unit, which is used to: take one of the translation components as the reference component, and respectively form each lifting pair with the reference component and other components; set a prediction operator and an update operator for each lifting pair, and use the bilateral parameterized activation function as the activation function of the prediction operator and the update operator to generate each lifting scheme.
[0111] Optionally, the frequency division feature extraction module 330 specifically includes a channel-by-channel deep frequency division enhancement unit, which is used to: configure the operator parameters of the prediction operator and the update operator for each input channel of the neural network to generate an independent prediction operator and an independent update operator corresponding to each input channel; perform separate operations on each input channel based on its own independent prediction operator and independent update operator to generate each high-frequency branch and low-frequency branch.
[0112] Optionally, the frequency division feature extraction module 330 specifically includes a shared parameter frequency division enhancement unit, which is used to: jointly configure the operator parameters of the prediction operator and the update operator for each input channel of the neural network to generate a shared prediction operator and a shared update operator; and perform prediction and update based on the shared prediction operator and the shared update operator through each input channel to generate each high-frequency branch and low-frequency branch.
[0113] Optionally, the output feature calculation module 340 is specifically used to: sum each high-frequency branch to obtain a high-frequency component; use an enhancement module to enhance the high-frequency component to obtain an enhanced high-frequency component; multiply the enhanced high-frequency component with the low-frequency branch to obtain an output feature.
[0114] The technical solution of the embodiment of the present invention constructs an improvement scheme through bilateral parameterized activation functions, which can explicitly separate the high-frequency details and low-frequency structures of the image based on multi-directional translation operations, thereby improving the parsing accuracy of text features. By constructing bilateral parameterized activation functions and integrating the improvement scheme, it is possible to inject nonlinear computing capabilities into the feature extraction process, breaking through the bottleneck that the traditional convolutional layer can only achieve linear transformation, allowing the model to adaptively adjust the efficiency of positive and negative interval signal transmission and enhance the flexibility of feature expression. Through the frequency division and improvement implementation method of channel-by-channel or shared parameters, it is possible to specifically extract the specific frequency features of each channel or cross-channel interaction features.
[0115] An image recognition device based on frequency division feature extraction provided by an embodiment of the present invention can execute an image recognition method based on frequency division feature extraction provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects of the execution method.
[0116] Example 4
[0117] Figure 6A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0118] like Figure 6 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0119] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0120] The processor 11 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as an image recognition method based on frequency division feature extraction.
[0121] In some embodiments, an image recognition method based on frequency division feature extraction may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the image recognition method based on frequency division feature extraction described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to perform an image recognition method based on frequency division feature extraction in any other appropriate manner (for example, by means of firmware).
[0122] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0123] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0124] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0125] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0126] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0127] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.
[0128] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.
[0129] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. An image recognition method based on frequency division feature extraction, characterized in that: include: Acquire an image to be identified, and obtain each translation component according to the image to be identified; Constructing a bilateral parameterized activation function, and constructing each lifting scheme based on the bilateral parameterized activation function and each translation component; Determining a target implementation method corresponding to the image to be recognized, and performing prediction and updating according to each of the boosting schemes and the target implementation method to generate each high-frequency branch and each low-frequency branch, wherein the target implementation method includes channel-by-channel deep frequency boosting or parameter-sharing frequency boosting; Output features are calculated based on the high-frequency branches and the low-frequency branches, the output features are input into a neural network for text classification and recognition, and a recognition result is output.
2. The method according to claim 1, characterized in that Obtaining each translation component according to the image to be identified includes: The image to be recognized is input into the neural network, and feature maps of different levels are obtained through convolution and pooling operations in the neural network; Select the feature map of any level as the target feature map; A translation operation is performed on the target feature map to obtain four translation components.
3. The method according to claim 1, characterized in that The construction of the bilateral parameterized activation function includes: Determine a mathematical form of an activation function based on the ReLU function, wherein the mathematical form includes bilateral learnable parameters; The gradient of the loss function with respect to the bilaterally learnable parameters is calculated using a back-propagation algorithm, and the bilaterally learnable parameters are updated based on the gradient and a preset optimization algorithm to generate a bilaterally parameterized activation function.
4. The method according to claim 3, characterized in that The constructing of each lifting scheme according to the bilateral parameterized activation function and each translation component includes: Taking one of the translation components as a reference component, and combining the reference component with other components to form lifting pairs; A prediction operator and an update operator are set for each of the lifting pairs, and the bilateral parameterized activation function is used as the activation function of the prediction operator and the update operator to generate each lifting scheme.
5. The method according to claim 4, characterized in that When the target implementation method is channel-by-channel deep frequency division boosting, the prediction and updating according to each boosting scheme and the target implementation method to generate each high-frequency branch and low-frequency branch includes: The operator parameters of the prediction operator and the update operator are configured for each input channel of the neural network to generate an independent prediction operator and an independent update operator corresponding to each input channel; Each high-frequency branch and low-frequency branch are generated by performing separate operations on each input channel based on its own independent prediction operator and independent update operator.
6. The method according to claim 4, characterized in that When the target implementation method is frequency division boosting with shared parameters, the prediction and updating according to each boosting scheme and the target implementation method to generate each high-frequency branch and low-frequency branch includes: Commonly configuring operator parameters of a prediction operator and an update operator for each input channel of a neural network to generate a shared prediction operator and a shared update operator; Prediction and update are performed based on the shared prediction operator and the shared update operator respectively through each input channel to generate each high-frequency branch and low-frequency branch.
7. The method according to claim 1, characterized in that The calculating output features according to each of the high-frequency branches and the low-frequency branches includes: Summing the high-frequency branches to obtain a high-frequency component; Using an enhancement module to enhance the high frequency component to obtain an enhanced high frequency component; The enhanced high-frequency component is multiplied by the low-frequency branch to obtain output features.
8. An image recognition device based on frequency division feature extraction, characterized in that: include: A translation component generation module, configured to obtain an image to be identified and obtain each translation component based on the image to be identified; A lifting scheme construction module, configured to construct a bilateral parameterized activation function, and construct each lifting scheme according to the bilateral parameterized activation function and each translation component; A frequency division feature extraction module is configured to determine a target implementation method corresponding to the image to be recognized, and to predict and update the target implementation method according to each of the boosting schemes and the target implementation method to generate each high-frequency branch and low-frequency branch, wherein the target implementation method includes channel-by-channel deep frequency division boosting or shared parameter frequency division boosting; The output feature calculation module is used to calculate output features based on each of the high-frequency branches and the low-frequency branches, input the output features into the neural network for text classification and recognition, and output the recognition results.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively coupled to the at least one processor; The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.
10. A computer storage medium, characterized in that The computer storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the method according to any one of claims 1 to 7 when executed.