SIM card identification method based on LCS and low-rank CP reconstruction attention
By applying LCS and low-rank CP reconstruction attention technology in SIM card recognition, automatic identification of SIM card numbers is achieved, and the problems of inefficient business processing and misoperation in the existing technology are solved, and the accuracy and efficiency of identification are improved.
Patent Information
- Application Number
- CN202510287563.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-24
AI Technical Summary
In the prior art, business processing relies on manual input of card numbers, which is inefficient and prone to misoperation. The text recognition model cannot effectively identify card numbers, resulting in less obvious improvement in efficiency.
The SIM card recognition method based on LCS and low-rank CP reconstruction attention is adopted, and the SIM card image is extracted and recognized through the large-core sparse convolution module and the low-rank CP tensor reconstruction attention module to automatically identify the card number.
It improves the efficiency of business processing, reduces the error rate of manual input, realizes the rapid and accurate identification of SIM card numbers, and solves the problems of inefficiency and misoperation in the existing technology.
Smart Images

Figure CN120198642A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition. Specifically, it relates to a SIM card recognition method, device, readable computer storage medium and service system based on LCS and low-rank CP reconstruction attention. Background Art
[0002] Taking the SIM card as an example, the SIM card number consists of 19 digits and 1 uppercase English letter. When handling relevant services, it is usually necessary for staff to manually input the SIM card number, which increases the processing time of batch service handling and the probability of misoperation.
[0003] In the prior art, there is a way to automatically extract the SIM card number by using a text recognition tool. However, for the text recognition tool, all text information will be recognized during the recognition process, and there is a large amount of irrelevant redundant information in the recognized text information. After further screening, the SIM card number information needs to be obtained, and the improvement of efficiency is not obvious.
[0004] In summary, in the business handling process involving inputting the card number in the prior art, manually inputting the card number by staff has the problems of low efficiency and easy misoperation. And the text recognition model in the prior art cannot recognize the card number, and still requires further sorting by staff, and the efficiency improvement is not obvious. Summary of the Invention
[0005] The main purpose of this application is to provide a SIM card recognition method, device, readable computer storage medium and service system based on LCS and low-rank CP reconstruction attention, so as to at least solve the problem of low efficiency in the prior art that business handling depends on manual input.
[0006] To achieve the above object, according to one aspect of the present application, a SIM card recognition method based on LCS and low-rank CP reconstruction attention is provided, including: recognizing a first target image to obtain label data, where the first target image is an image containing a target object, the target object includes a preset pattern and a unique identifier of the preset pattern, and the label data includes the coordinates of the minimum bounding rectangle of the unique identifier and the unique identifier; detecting the region of the unique identifier in the first target image through an alternative detection model to obtain a third target image, and the alternative detection model at least includes a large kernel sparse convolution module, and the large kernel sparse convolution module performs feature extraction through large kernel convolution, sparse convolution and depthwise separable convolution; recognizing the unique identifier of the third target image through an alternative recognition model to obtain a recognition flag, and the alternative recognition model at least includes a low-rank CP tensor reconstruction attention module and a multi-head low-rank CP tensor reconstruction module, the low-rank CP tensor reconstruction attention module decomposes the self-attention tensor into rank-one tensors, and the multi-head low-rank CP tensor reconstruction module slices the input features according to the rank-one tensors and extracts self-attention features respectively and then fuses them; determining a comprehensive performance parameter according to the third target image, the coordinates, the recognition flag and the unique identifier, and when the comprehensive performance parameter is less than or equal to a first threshold, determining the alternative detection model and the alternative recognition model as a target detection model and a target recognition model; recognizing the unique identifier of the first target image obtained in the business process according to the target recognition model and the target detection model to obtain the unique identifier; batch inputting the unique identifier into a target system, and the target system is used to activate the target object according to the unique identifier.
[0007] Optionally, after obtaining the first target image, the method further includes: randomly adjusting the background color of the first target image to an image with the color three-channel values within a first preset range to obtain a plurality of fourth target images; randomly adjusting the font color of each fourth target image to an image with the color three-channel values within a second preset range to obtain a plurality of fifth target images; processing each fifth target image sequentially through random salt and pepper noise, Gaussian blur and random rotation to obtain a plurality of second target images.
[0008] Optionally, the alternative detection model further includes a feature pyramid fusion module and a DB detection head decoding module. Detecting the region of the unique identifier in the first target image through the alternative detection model to obtain a third target image includes: performing feature extraction on the first target image through the large kernel sparse convolution module to obtain a first global feature; performing upsampling on the first global feature at different scales through the feature pyramid fusion module and fusing it through residual connection to obtain a second global feature; performing differentiable binarization on the second global feature through the DB detection head decoding module to obtain a foreground image and a background image, and determining the foreground image as the third target image.
[0009] Optionally, the first target image is subjected to feature extraction by a large kernel sparse convolution module to obtain a first global feature, including: processing the first target image through a point convolution with a convolution kernel size of 1 and splitting it along the first channel dimension to obtain a first feature map and a second feature map, processing the first feature map through a horizontal convolution and processing the second feature map through a vertical convolution, and obtaining a third global feature through concat splicing; processing the third global feature through a point convolution with a convolution kernel size of 1 and splitting it along the second channel dimension to obtain a third feature map and a fourth feature map; performing feature extraction on the third feature map through a depthwise separable convolution with a dilation rate of 1 and a receptive field of a first preset size and to obtain a fourth global feature, and performing feature extraction on the fourth feature map through a depthwise separable convolution with a dilation rate of 1 and a receptive field of a second preset size and to obtain a fifth global feature; performing concat splicing on the fourth global feature and the fifth global feature to obtain a sixth global feature, and performing a residual connection on the sixth global feature and the first target image processed through a point convolution with a convolution kernel size of 1 to obtain a first global feature.
[0010] Optionally, the alternative recognition model further includes a basis vector generation module and a feature encoding and decoding module. The unique identifier of the third target image is recognized by the alternative recognition model to obtain a recognition flag. The method includes: adjusting the third target image respectively according to the length, width and number of channels by the basis vector generation module to obtain a fifth feature map, a sixth feature map and a seventh feature map, processing the fifth feature map through the linear layer of the basis vector generation module to obtain a query vector, processing the sixth feature map through the linear layer of the basis vector generation module to obtain a key vector, and processing the seventh feature map through the linear layer of the basis vector generation module to obtain a value vector; decomposing and reconstructing the self-attention mask according to the query vector, the key vector and the value vector by the low-rank CP tensor reconstruction attention module to obtain a target attention mask; decomposing the third target image based on the number of channels by the multi-head low-rank CP tensor reconstruction module to obtain a plurality of sub-feature maps, reconstructing each sub-feature map respectively according to the target attention mask to obtain a plurality of seventh global features, and splicing the seventh global features to obtain an eighth global feature; performing feature extraction and feature encoding on the eighth global feature through a plurality of large kernel sparse convolution modules and a plurality of low-rank CP tensor reconstruction attention modules in sequence by the feature encoding and decoding module to obtain an encoded recognition result, and decoding the encoded recognition result by a Transformer model to obtain a recognition flag.
[0011] Optionally, the self-attention mask is decomposed and reconstructed by the low-rank CP tensor reconstruction attention module according to the query vector, key vector, and value vector to obtain the target attention mask, including: calculating the outer product of the query vector and the key vector to obtain the first basis vector; decomposing the self-attention mask into a linear combination of the first basis vectors to obtain the second basis vector; processing the second basis vector through the softmax activation function and calculating the outer product with the value vector to obtain a rank-one tensor, and linearly combining through the rank-one tensor to obtain the target attention mask.
[0012] Optionally, the eighth global feature is subjected to feature extraction and feature encoding through multiple large-kernel sparse convolution modules and multiple low-rank CP tensor reconstruction attention modules by the feature encoding and decoding module to obtain the encoded recognition result, including: extracting features of the eighth global feature through the first large-kernel sparse convolution module to obtain the ninth global feature; extracting features of the ninth global feature through the second large-kernel sparse convolution module to obtain the tenth global feature; encoding the features of the tenth global feature through the first low-rank CP tensor reconstruction attention module to obtain the first encoded result; encoding the features of the first encoded result through the second low-rank CP tensor reconstruction attention module to obtain the second encoded result; encoding the features of the second encoded result through the third low-rank CP tensor reconstruction attention module to obtain the encoded recognition result.
[0013] According to another aspect of the present application, there is provided a SIM card recognition device based on LCS and low-rank CP reconstruction attention. The device includes: a first acquisition unit configured to recognize a first target image to obtain label data. The first target image is an image containing a target object, and the target object includes a preset pattern and a unique identifier of the preset pattern. The label data includes the coordinates of the minimum bounding rectangle of the unique identifier and the unique identifier; a first processing unit configured to detect the region of the unique identifier in the first target image through an alternative detection model to obtain a third target image. The alternative detection model at least includes a large kernel sparse convolution module, and the large kernel sparse convolution module performs feature extraction through large kernel convolution, sparse convolution, and depthwise separable convolution; a second processing unit configured to recognize the unique identifier of the third target image through an alternative recognition model to obtain a recognition flag. The alternative recognition model at least includes a low-rank CP tensor reconstruction attention module and a multi-head low-rank CP tensor reconstruction module. The low-rank CP tensor reconstruction attention module decomposes the self-attention tensor into rank-one tensors, and the multi-head low-rank CP tensor reconstruction module slices the input features according to the rank-one tensors and separately extracts self-attention features and then fuses them; a first determination unit configured to determine a comprehensive performance parameter according to the third target image, coordinates, recognition flag, and unique identifier, and when the comprehensive performance parameter is less than or equal to a first threshold, determine the alternative detection model and the alternative recognition model as a target detection model and a target recognition model; a third processing unit configured to recognize the unique identifier of the first target image obtained during the service process according to the target recognition model and the target detection model to obtain the unique identifier; an input unit configured to batch input the unique identifier into a target system, and the target system is configured to activate the target object according to the unique identifier.
[0014] According to still another aspect of the present application, there is provided a computer-readable storage medium. The computer-readable storage medium includes a stored program, wherein when the program runs, it controls the device where the computer-readable storage medium is located to execute any one of the methods.
[0015] According to yet another aspect of the present application, there is provided a service system, including: one or more processors, a memory, and one or more programs. Wherein, the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include the method for any one of them.
[0016] Applying the technical solution of the present application in the above SIM card recognition method based on LCS and low-rank CP reconstruction attention, first, for the first target image, label data is obtained. The first target image is an image containing a target object, and the target object includes a preset pattern and a unique identifier of the preset pattern. The label data includes the coordinates of the minimum bounding rectangle of the unique identifier and the unique identifier. Then, the region of the unique identifier in the first target image is detected through an alternative detection model to obtain a third target image. The alternative detection model at least includes a large kernel sparse convolution module, and the large kernel sparse convolution module performs feature extraction through large kernel convolution, sparse convolution, and depthwise separable convolution. After that, the unique identifier of the third target image is recognized through an alternative recognition model to obtain a recognition flag. The alternative recognition model at least includes a low-rank CP tensor reconstruction attention module and a multi-head low-rank CP tensor reconstruction module. The low-rank CP tensor reconstruction attention module decomposes the self-attention tensor into rank-one tensors, and the multi-head low-rank CP tensor reconstruction module slices the input features according to the rank-one tensors and separately extracts self-attention features and then fuses them. After that, comprehensive performance parameters are determined based on the third target image, coordinates, recognition flag, and unique identifier, and when the comprehensive performance parameters are less than or equal to the first threshold, the alternative detection model and the alternative recognition model are determined as the target detection model and the target recognition model. After that, the unique identifier of the first target image obtained during the business process is recognized according to the target recognition model and the target detection model to obtain the unique identifier. Finally, the unique identifier is batch-input into the target system, and the target system is used to activate the target object according to the unique identifier. The present application improves the large kernel convolution, introduces dilated convolution for sparse sampling to expand the field of view and improve the global nature of the features. At the same time, through depthwise separable convolution, the convolution process is divided into depth convolution and pointwise convolution, reducing the number of parameters and avoiding the reduction of the operation rate caused by large kernel convolution. At the same time, through the low-rank CP tensor reconstruction attention module, the self-attention mask is decomposed into rank-one tensors, reducing the number of parameters and computational complexity of the model, and then the self-attention mask is reconstructed by the rank-one tensors, improving the accuracy while increasing the operation speed. The multi-head low-rank CP tensor reconstruction module slices the input image based on the channel dimension, calculates and fuses the attention features of different blocks respectively, and improves the model recognition accuracy by extracting global features. The card number recognition is realized through the detection model and the recognition model after training, and the SIM card number is one specific embodiment. This method solves the problem of low efficiency in business handling relying on manual input in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 FIG. 6 shows a hardware structure block diagram of a mobile terminal for a SIM card recognition method based on LCS and low-rank CP reconstruction attention provided in an embodiment of the present application;
[0018] Figure 2It shows a schematic flowchart of a SIM card recognition method based on LCS and low-rank CP reconstruction attention provided according to an embodiment of the present application;
[0019] Figure 3 It shows a structural block diagram of a SIM card recognition device based on LCS and low-rank CP reconstruction attention provided according to an embodiment of the present application.
[0020] Among them, the above-mentioned drawings include the following reference numerals:
[0021] 102, processor; 104, memory; 106, transmission device; 108, input / output device. Detailed implementation manners
[0022] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the drawings and in combination with the embodiments.
[0023] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0024] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so as to describe the embodiments of the present application here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0025] For the convenience of description, some nouns or terms related to the embodiments of the present application are described below:
[0026] CP (Canonical Polyadic decomposition, canonical multilinear tensor decomposition): The CP tensor decomposition theory, where a high-order tensor can be decomposed into a linear combination of multiple rank-one tensors. A rank-one tensor refers to one that can be obtained by the outer product of a group of vectors. Tensor decomposition decomposes a high-dimensional tensor into a low-rank tensor, reducing the computational complexity of the model and improving the computational efficiency of the model.
[0027] LSC (Large kernel Sparse Convolution): The large kernel sparse convolution samples the feature information of the image using a convolution kernel with a larger receptive field. In this application, a sparse convolution with holes is used to expand the receptive field of the convolution kernel and reduce the number of parameters of the model. The large kernel sparse convolution is decomposed into strip-shaped convolutions with the same size but different directions to obtain richer feature information.
[0028] DB (Differentiable Binarization): The differentiable thresholding detection head decodes the feature map to obtain a threshold mask map and a probability map. The threshold mask map filters the probability map to obtain a binary foreground image and a background image. Compared with traditional binarization operations, differentiable binarization is easier to optimize and more adaptable during model training.
[0029] As introduced in the background art, in the existing business handling process involving inputting card numbers, manually inputting card numbers by staff has the problems of low efficiency and easy occurrence of misoperations. Moreover, the existing text recognition models cannot recognize card numbers, and further sorting by staff is still required, resulting in an insignificant improvement in efficiency. To solve the problem of low efficiency in business handling relying on manual input in the prior art, an embodiment of this application provides a SIM card recognition method, device, readable computer storage medium, and business system based on LCS and low-rank CP reconstruction attention.
[0030] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.
[0031] The method embodiments provided in the embodiments of this application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 is a hardware structure block diagram of a mobile terminal of a SIM card recognition method based on LCS and low-rank CP reconstruction attention according to an embodiment of the present invention. As Figure 1 shown, the mobile terminal may include one or more ( Figure 1 only one is shown in Figure 1 a processor 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above mobile terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above mobile terminal. For example, the mobile terminal may further include more or fewer components than Figure 1 shown inFigure 1 The different configurations shown.
[0032] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the SIM card recognition method based on LCS and low-rank CP reconstruction attention in the embodiments of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories may be connected to the mobile terminal through a network. Examples of the above networks include, but are not limited to, the Internet, enterprise intranets, local area networks, mobile communication networks, and combinations thereof. The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of the mobile terminal. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0033] In this embodiment, a SIM card recognition method based on LCS and low-rank CP reconstruction attention running on a mobile terminal, a computer terminal, or a similar computing device is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0034] Figure 2 is a flowchart of the SIM card recognition method based on LCS and low-rank CP reconstruction attention according to an embodiment of the present application. As Figure 2 shown, the method includes the following steps:
[0035] Step S201, recognize a first target image to obtain label data. The first target image is an image containing a target object, and the target object includes a preset graphic and a unique identifier of the preset graphic. The label data includes the coordinates of the minimum bounding rectangle of the unique identifier and the unique identifier;
[0036] Specifically, the above-mentioned first target image is identified and labeled by a labeling tool, the coordinates of the above-mentioned unique identifier and the minimum bounding rectangle of the unique identifier are recorded, and a text information dataset containing the above-mentioned unique identifier and coordinates is constructed.
[0037] Step S202, detecting the area of the unique identifier in the first target image through an alternative detection model to obtain a third target image. The alternative detection model at least includes a large kernel sparse convolution module, and the large kernel sparse convolution module extracts features through large kernel convolution, sparse convolution, and depthwise separable convolution;
[0038] Specifically, the large kernel sparse convolution module LSC is used to extract features. There are two layers of large kernel conventional convolution and large kernel depthwise separable sparse convolution. Then, the feature pyramid fusion module is used to fuse feature maps of different scales to obtain richer feature information. The fused feature map is decoded by the DB detection head to generate a threshold map and a probability map, which are converted into a binary image, thereby determining the position information of the card number and obtaining the above-mentioned third target image. Among them, the above-mentioned large kernel conventional convolution realizes an enlarged receptive field through parallel horizontal and vertical convolutions, and the above-mentioned large kernel depthwise separable sparse convolution reduces the computational complexity through sparse convolution and depthwise separable convolution. This is to avoid the increase in computational complexity and the resulting efficiency decline due to the introduction of large kernel convolution.
[0039] Step S203, identifying the unique identifier of the third target image through an alternative recognition model to obtain an identification flag. The alternative recognition model at least includes a low-rank CP tensor reconstruction attention module and a multi-head low-rank CP tensor reconstruction module. The low-rank CP tensor reconstruction attention module decomposes the self-attention tensor into rank-one tensors, and the multi-head low-rank CP tensor reconstruction module slices the input features according to the rank-one tensors and then fuses the self-attention features extracted separately;
[0040] Specifically, a basis vector generation module is used to generate orthogonal query vectors, key vectors, and value vectors based on the above-mentioned third target image. Then, the low-rank CP tensor reconstruction attention module generates multiple rank-one tensors through outer products to decompose the existing self-attention tensor. After re-linear combination, a self-attention matrix is obtained. The multi-head low-rank CP tensor reconstruction module further enhances the global dependence of features by adopting a multi-head attention mechanism. Then, through the feature encoding and decoding module, after serializing the feature map, it is input into the Transformer decoder to obtain the serialized features of the card number, and finally the card number information is output.
[0041] Step S204, determining the comprehensive performance parameter according to the third target image, coordinates, identification flag, and unique identifier, and when the comprehensive performance parameter is less than or equal to the first threshold, determining the alternative detection model and the alternative recognition model as the target detection model and the target recognition model;
[0042] Specifically, the results identified by the model are compared with the above-mentioned tag data to quantify the comprehensive performance of the model, and the above-mentioned comprehensive performance parameters are obtained. Then, the alternative detection model and the alternative recognition model are evaluated based on the quantified model performance to determine whether they can be put into use. That is, when the comprehensive performance parameter is less than or equal to the first threshold, the alternative detection model and the alternative recognition model are determined as the target detection model and the target recognition model.
[0043] Step S205: Identify the unique identifier of the first target image obtained in the business process according to the target recognition model and the target detection model to obtain the unique identifier.
[0044] Step S206: Batch input the unique identifier into the target system, which is used to activate the target object according to the unique identifier.
[0045] Specifically, after the model training is completed, the above-mentioned target recognition model and the above-mentioned target detection model are used to process the images extracted during the business handling process to extract the above-mentioned unique identifier (which can be the SIM card number in one embodiment). Then, the above-mentioned unique identifier obtained by recognition is batch input into the system for card activation.
[0046] It can be understood that this application integrates two technologies in deep learning, namely large kernel sparse convolution (LSC) and low-rank CP tensor reconstruction attention (CPA). The former captures more global information by expanding the receptive field, and the latter reduces the computational complexity of the self-attention mechanism through tensor decomposition technology to improve the model efficiency. By combining the two, the computational delay increase caused by adjusting the convolutional kernel to a large kernel convolution is reduced, ensuring the model accuracy and computational speed.
[0047] Through this embodiment, first, for the first target image, label data is obtained. The first target image is an image containing a target object, and the target object includes a preset graphic and a unique identifier of the preset graphic. The label data includes the coordinates of the minimum bounding rectangle of the unique identifier and the unique identifier. Then, the region of the unique identifier in the first target image is detected by an alternative detection model to obtain a third target image. The alternative detection model at least includes a large kernel sparse convolution module, and the large kernel sparse convolution module performs feature extraction through large kernel convolution, sparse convolution, and depthwise separable convolution. After that, the unique identifier of the third target image is recognized by an alternative recognition model to obtain a recognition flag. The alternative recognition model at least includes a low-rank CP tensor reconstruction attention module and a multi-head low-rank CP tensor reconstruction module. The low-rank CP tensor reconstruction attention module decomposes the self-attention tensor into rank-one tensors, and the multi-head low-rank CP tensor reconstruction module slices the input features according to the rank-one tensors and extracts self-attention features respectively and then fuses them. After that, comprehensive performance parameters are determined based on the third target image, coordinates, recognition flag, and unique identifier, and when the comprehensive performance parameters are less than or equal to a first threshold, the alternative detection model and the alternative recognition model are determined as the target detection model and the target recognition model. After that, the unique identifier of the first target image obtained during the business process is recognized according to the target recognition model and the target detection model, and the unique identifier is obtained. Finally, the unique identifier is batch-input into the target system, and the target system is used to activate the target object according to the unique identifier. This application improves the large kernel convolution, introduces dilated convolution for sparse sampling to expand the field of view and improve the globality of features. At the same time, through depthwise separable convolution, the convolution process is divided into depth convolution and pointwise convolution, reducing the number of parameters and avoiding the reduction of the operation rate caused by large kernel convolution. At the same time, the low-rank CP tensor reconstruction attention module decomposes the self-attention mask into rank-one tensors, reducing the number of parameters and computational complexity of the model, and then reconstructs the self-attention mask through rank-one tensors, improving the accuracy while increasing the operation speed. The multi-head low-rank CP tensor reconstruction module slices the input image based on the channel dimension, calculates and fuses the attention features of different blocks respectively, and improves the model recognition accuracy by extracting global features. Through the trained detection model and recognition model, the card number recognition is realized, and the SIM card number is one specific embodiment. This method solves the problem of low efficiency in business handling relying on manual input in the prior art.
[0048] In an alternative implementation manner, in order to ensure the stability and accuracy of model training, after obtaining the first target image, the above method further includes:
[0049] Step S301, randomly adjust the background color of the first target image to an image with the values of the three color channels within a first preset range to obtain a plurality of fourth target images;
[0050] Specifically, for the above-mentioned first target image obtained (optionally, it can be the original image of the SIM card), the background color is randomly adjusted to the light color range with the color three channels being (170, 255), and multiple above-mentioned fourth target images are generated. The formula is as follows:
[0051] img background =(R(random(170,255),G(random(170,255),B(random(170,255))
[0052] Where img background is the background image, and R, G, and B are the above-mentioned color three channels.
[0053] It can be understood that the above operations are used to simulate the color changes that occur in the background of the SIM card under different lighting conditions to ensure that the model can adapt to various scenarios of brightness and color temperature.
[0054] Step S302: According to each fourth target image, the font color is randomly adjusted to an image with the value of the color three channels within the second preset range to obtain multiple fifth target images;
[0055] Specifically, the font in the above-mentioned fourth target image is randomly adjusted, that is, the font color is randomly adjusted to the dark color range with the color three channels being (0, 120) to generate the above-mentioned fifth target image. The formula is as follows:
[0056] img foreground =(R(random(0,120),G(random(0,120),B(random(0,120))
[0057] Where img foreground is the foreground image, and then the foreground image and the background image are randomly combined to obtain the above-mentioned fifth target image.
[0058] It can be understood that the adjustment of the font color is based on the same purpose as the background adjustment to cope with color changes caused by low illumination, reflected light, etc., so as to enhance the adaptability of the model to text color changes.
[0059] Step S303: Each fifth target image is processed sequentially by random salt-and-pepper noise, Gaussian blur, and random rotation to obtain multiple second target images.
[0060] In one embodiment, the above-mentioned fifth image is processed by performing random salt-and-pepper noise, Gaussian blur, and random rotation. The formula for Gaussian blur processing is img = Gauss(7,7)(img), where img is the above-mentioned fifth target image, and (7,7) is the sliding window. The formula for random rotation processing is img1 = Rotate(-2,2)(img), where img1 is the above-mentioned second target image, and (-2,2) is the random floating-point number of the text rotation angle.
[0061] It can be understood that in the training of deep learning models, data augmentation techniques are widely used to improve the generalization ability and robustness of the models, enabling them to maintain a high accuracy rate when facing inputs under different environments and conditions. In this application, by randomly adjusting the image background color, font color, and applying operations such as salt-and-pepper noise, Gaussian blur, and rotation, the diversity of training data can be increased, enabling the model to learn a more extensive feature representation, and thus performing more stably and accurately in practical applications.
[0062] In addition, the above-mentioned data augmentation means may also include image scaling, shearing, brightness and contrast adjustment, etc., to adapt to a wider range of application scenarios.
[0063] In order to segment the above-mentioned unique identifier to improve the accuracy of the recognition model, in an optional implementation manner, the above-mentioned step S202 includes:
[0064] Step S2021, extracting features from the first target image through a large kernel sparse convolution module to obtain a first global feature;
[0065] Specifically, a large kernel sparse convolution (LSC) module is used to extract features from the first target image. The LSC module includes two layers of structures: large kernel regular convolution and large kernel depthwise separable sparse convolution. The former expands the receptive field through parallel horizontal and vertical convolutions, and the latter extracts global features while maintaining computational efficiency through sparse convolution and depthwise separable convolution. The above operations enrich the feature information through the double-layer structure.
[0066] Step S2022, performing upsampling on the first global feature at different scales through a feature pyramid fusion module, and fusing it through residual connection to obtain a second global feature;
[0067] Specifically, the above-mentioned first global feature is upsampled at different scales through a feature pyramid fusion module, and fused with shallow information through residual connection to obtain the above-mentioned second global feature. The formula is as follows:
[0068] {L i} = LSC_Block(X), i = 3, 4, 5
[0069] {Si (L i + UpSample(L i-1 ), for i = 3, 4, 5
[0070] Z = Linear(Concat(S3, S4, S5))
[0071] Where X is the input feature of the feature pyramid fusion module, {L i} is the feature information obtained after upsampling on the i-th layer, L i-1 is the high-level global semantic feature information, {S i} is the fusion information obtained at the corresponding scale on the i-th layer, and Z is the encoded information after concatenating the fusion information, that is, the above-mentioned second global technical feature.
[0072] Step S2023: Use the DB detection head decoding module to perform differentiable binarization on the second global feature to obtain a foreground image and a background image, and determine the foreground image as the third target image.
[0073] Specifically, use the DB detection head decoding module to perform differentiable binarization on the above-mentioned second global technical feature, that is, based on the extraction of the threshold map and the probability map, and generate a binarized foreground image and background image based on the threshold map and the probability map. Specifically, the probability map represents the probability that each pixel belongs to the foreground, and the threshold map is the foreground threshold of the pixel. When the probability is greater than the threshold, the pixel is determined to be the foreground. The foreground map generated in this process is the third target image, that is, the precise segmentation area of the unique identifier (such as the SIM card number).
[0074] It can be understood that in this application, large kernel sparse convolution is used for feature extraction, which expands the receptive field and facilitates the extraction of the global features of the first target image. Furthermore, through the multi-scale feature fusion of the feature pyramid, different details of the target can be captured, and precise segmentation is performed through the differentiable binarization of the DB detection head based on the refined features after fusion.
[0075] Through the above embodiments, introducing large kernel sparse convolution significantly improves the comprehensiveness of feature extraction compared with traditional convolution in processing irregular texts. Furthermore, the feature pyramid fusion ensures that the model can achieve precise processing for texts of different sizes, while the differentiable binarization of the DB detection head simplifies the operations in the model training process. The object detection model that fuses the above technologies demonstrates high efficiency and accuracy while reducing the number of model parameters, and achieves a stable detection effect for the dynamic changes of SIM card images during the business handling process.
[0076] To improve the effect and speed of large kernel convolution, in an optional implementation, the above step S2021 includes:
[0077] Step S20211: Process the first target image through point convolution with a kernel size of 1 and split it along the first channel dimension to obtain a first feature map and a second feature map. Process the first feature map through horizontal convolution and the second feature map through vertical convolution, and then obtain the third global feature through concat splicing.
[0078] Specifically, process the above first target image through point convolution with a kernel size of 1 to adjust the channel dimension of the convolution kernel. The input features are converted into feature maps through point convolution and then split into a first feature map and a second feature map along the channel dimension. Apply horizontal convolution and vertical convolution to the first feature map and the second feature map respectively. These two operations help the model capture the features in the horizontal and vertical directions of the image. Subsequently, splice the processed feature maps through concat to form the third global feature. The use of horizontal and vertical convolutions ensures that the model can understand the text information from different angles.
[0079] It can be understood that the above operation decomposes the convolution with a receptive field size of (k,k) into parallel (3,k) and (k,3) convolution kernels. Decompose the conventional convolution into two strip-shaped convolutions, horizontal and vertical, to handle irregular texts in the actual scenario. The horizontal convolution kernel and the vertical convolution kernel extract horizontal and vertical feature information, and the above third global feature is obtained after concat splicing in the channel dimension. The formula is as follows:
[0080] X in =f 1×1 (X)
[0081] X1,X2=Split(X in )
[0082] X mid =f 1×1 (Concat(f 3×k (X1),f k×3 (X2)))
[0083] X mid =LSC k×k (X in )
[0084] Among them, f 1×1 is the above point convolution, X1 and X2 are the above first feature map and the above second feature map respectively, f 3×k and f k×3 are the horizontal convolution kernel and the vertical convolution kernel respectively, LSC k×k represents the large kernel convolution with a receptive field of k, and X mid is the feature map after point convolution processing.
[0085] Step S20212: Process the third global feature through point convolution with a convolution kernel size of 1 and split it along the second channel dimension to obtain a third feature map and a fourth feature map;
[0086] Specifically, perform point convolution processing on the third global feature again and split it along the second channel dimension to form a third feature map and a fourth feature map. The formula is as follows:
[0087] X3,X4 = Split(X mid )
[0088] where X3 and X4 are the above-mentioned third feature map and fourth feature map respectively.
[0089] Step S20213: Extract features from the third feature map through depthwise separable convolution with a dilation rate of 1 and a receptive field of the first preset size, and obtain a fourth global feature. Extract features from the fourth feature map through depthwise separable convolution with a dilation rate of 1 and a receptive field of the second preset size to obtain a fifth global feature;
[0090] Specifically, extract features from the above-mentioned third feature map and fourth feature map respectively through depthwise separable convolution with receptive fields of the first preset size and the second preset size, and their dilation rates are both 1, to obtain the above-mentioned fourth global feature and fifth global feature. The formula is as follows:
[0091]
[0092] where represents depthwise separable convolution with a dilation rate of 1 and a convolution kernel receptive field size of (5,k), and X out is the sixth global feature.
[0093] It can be understood that by introducing depthwise separable convolution, compared with traditional convolution, the number of parameters can be greatly reduced, thereby reducing the computational cost and maintaining effective feature extraction.
[0094] Step S20214: Concatenate the fourth global feature and the fifth global feature to obtain a sixth global feature, and perform a residual connection between the sixth global feature and the first target image processed through point convolution with a convolution kernel size of 1 to obtain a first global feature.
[0095] Specifically, concatenate the fourth global feature and the fifth global feature through concat, and then perform a residual connection with the first target image processed through point convolution to obtain a first global feature. Residual connection helps the model learn more complex features, while avoiding the problem of gradient disappearance and reducing the number of iterations.
[0096] Through the above embodiments, the large kernel sparse convolution can improve the performance of extracting local and global features in images through convolution and splitting operations. The setting of horizontal convolution and vertical convolution can help the model capture the features in the horizontal and vertical directions of the image to cope with the deflection of the image. The residual connection expands the depth of the feature representation learned by the model, avoids the vanishing gradient during the training process, and speeds up the training of the model. In the field of image recognition, the multi-scale feature extraction and residual connection mechanism of the large kernel sparse convolution module is not limited to SIM cards, but can also be applied to the recognition of other objects with specific arrangement features, such as barcodes, text recognition, etc. Through multi-scale feature extraction and residual connection, the accuracy and efficiency of recognition can be significantly improved.
[0097] In a specific embodiment, the LSC-Net network constructed based on the large kernel sparse convolution module is used to detect the position information of the SIM card number. It is assumed that the data set includes 200 images, and the ratio of the training set to the test set is 8:2 respectively, and the image size is (640, 640). The optimizer of the model is Adam, and the learning rate is 0.001. The number of iterations is 300, and the size of each batch is 10. The data volume corresponding to different convolution kernel sizes during its training process is shown in Table 1.
[0098] Table 1
[0099] Large kernel convolution kernel Large kernel sparse convolution F1(%) Number of parameters (M) (5,5) (9,9) 93.9 2.675 (7,7) (13,13) 95.1 4.340 (9,9) (17,17) 96.3 6.670
[0100] It can be seen that when the sizes of the large kernel convolution kernels are 5, 7, and 9 respectively, the corresponding large kernel sparse convolution kernels are 9, 13, and 17. From the experimental results, it can be seen that as the convolution kernel gradually becomes larger, the F1 recognition accuracy of the model improves. The large kernel sparse convolution of (17, 17) reaches a recognition accuracy of 96.3%, but the number of parameters of the model is 6.670M, which is 4.394M more than that of (5, 5). In order to balance the scale and accuracy of the model, this method selects the large kernel sparse convolution of (9, 9) as the base model, and the number of parameters of the model is 2.675M.
[0101] In order to reduce the computational complexity of the unique identifier recognition process, in an alternative embodiment, step S203 includes:
[0102] Step S2031, adjust according to the length, width, and number of channels of the third target image respectively through the basis vector generation module to obtain the fifth feature map, the sixth feature map, and the seventh feature map. Process the fifth feature map through the linear layer of the basis vector generation module to obtain the query vector, process the sixth feature map through the linear layer of the basis vector generation module to obtain the key vector, and process the seventh feature map through the linear layer of the basis vector generation module to obtain the value vector;
[0103] Specifically, according to the length, width, and number of channels of the third target image (i.e., the image after card number positioning), the adjusted feature maps, namely the fifth feature map, the sixth feature map, and the seventh feature map, are respectively obtained through the basis vector generation module. Then, these feature maps are processed through a linear layer to respectively generate a query vector, a key vector, and a value vector.
[0104] In a specific implementation, according to the CP tensor decomposition theory, a high-rank tensor can be decomposed into a linear combination of rank-one tensors. The computational complexity of the self-attention mask is quadratic. This application uses CP tensor decomposition to reconstruct the self-attention mask, reducing the complexity of the model. In this embodiment, it is assumed that X represents the input feature of the self-attention module, and H, W, and C represent the length, width, and number of channels of the feature map. The feature map X is adjusted in shape through Reshape to respectively obtain features of sizes H×(WC), W×(HC), and C×(HW). The features in the three dimensions are input into a linear layer to generate a query vector Q, a key vector K, and a value vector V. The formula is as follows:
[0105] X ∈ R H×W×C
[0106] Q = Linear Q (Reshape(X))
[0107] = (q1, q2, q3,..., q r ), Q ∈ R H×r
[0108] K = Linear K (Reshape(X))
[0109] = (k1, k2, k3,..., k r ), K ∈ R W×r
[0110] V = Linear V (Reshape(X))
[0111] = (v1, v2, v3,..., v r ), V ∈ R C×r
[0112] Among them, q r , k r and v r are the basis vectors of the query vector, the key vector, and the value vector.
[0113] Step S2032, reconstruct the attention module through low-rank CP tensors. Decompose and reconstruct the self-attention mask according to the query vector, the key vector, and the value vector to obtain the target attention mask;
[0114] Specifically, by decomposing the self-attention mask into the outer product of the basis vectors of the query vector and the key vector, the self-attention mask can be decomposed and reconstructed through the above query vector, key vector, and value vector to obtain the target attention mask.
[0115] Step S2033: Decompose the third target image based on the number of channels through the multi-head low-rank CP tensor reconstruction module to obtain multiple sub-feature maps. Reconstruct each sub-feature map according to the target attention mask to obtain multiple seventh global features, and splice the seventh global features to obtain the eighth global feature;
[0116] Specifically, the multi-head low-rank CP tensor reconstruction module decomposes the third target image into multiple sub-feature maps along the channel dimension. Each sub-feature map is reconstructed according to the target attention mask to obtain the seventh global feature, and then these seventh global features are spliced to form the eighth global feature.
[0117] In a specific implementation, the feature map is decomposed into multiple feature sub-maps X from the channel dimension i , and each feature sub-map is respectively subjected to CP tensor self-attention reconstruction to obtain Y i and then re-spliced to obtain the final output Y:
[0118]
[0119] Y i = CP_Attention(λ i ; X i )
[0120] Y = Linear(Concat(Y i ))
[0121] where λ i is the learnable parameter in each head of self-attention.
[0122] Step S2034: Through the feature encoding and decoding module, perform feature extraction and feature encoding on the eighth global feature through multiple large-kernel sparse convolution modules and multiple low-rank CP tensor reconstruction attention modules to obtain an encoded recognition result, and decode the encoded recognition result through a Transformer model to obtain a recognition flag.
[0123] Specifically, through the feature encoding and decoding module, perform hybrid feature encoding on the eighth global feature, including feature extraction by the large-kernel sparse convolution module and feature encoding by the low-rank CP tensor reconstruction attention module, to obtain an encoded recognition result. Subsequently, use the Transformer model to decode the encoded recognition result and output the text recognition result, that is, the recognition flag.
[0124] Through the above embodiments, the low-rank CP tensor reconstruction attention module reduces the computational complexity and improves the computational efficiency of the model by decomposing the self-attention mask into rank-one tensors. The generation of query vectors, key vectors, and value vectors enables the model to focus on key regions in the image and improves the recognition accuracy. The use of the multi-head low-rank CP tensor reconstruction module further enhances the expressive power and generalization ability of the model, enabling the model to perform feature learning and recognition from multiple perspectives and dimensions.
[0125] In an alternative embodiment, in order to reconstruct the self-attention mask, step S2032 includes:
[0126] Step S20321, calculating the outer product of the query vector and the key vector to obtain the first basis vector;
[0127] Step S20322, decomposing the self-attention mask into a linear combination of the first basis vectors to obtain the second basis vector;
[0128] Step S20323, processing the second basis vector through the softmax activation function and calculating the outer product with the value vector to obtain a rank-one tensor,
[0129] Step S20324, obtaining the target attention mask through a linear combination of the rank-one tensors.
[0130] In a specific implementation, the basis vector q of the query vector i and the basis vector k of the key vector i perform an outer product to obtain the basis vector qk i , and the self-attention mask QK T is decomposed into a linear combination of the basis vectors qk i . The basis vector of the value vector performs an outer product with the self-attention mask to obtain r rank-one tensors qkv i . The linear combination of the rank-one tensors obtains the final attention mask CP_Attention (the above target attention mask), that is, the above target attention mask. The formula is as follows:
[0131] QK T =λ q (q1,q2,q3,...,q r )λ k (k1,k2,k3,...,k r ) T
[0132] =λ q λ k (qk1,qk2,qk3,...,qk r )
[0133]
[0134] Among them, λ q and λ k represent trainable parameters. The softmax activation function is used to enhance the feature weights with larger weights in the attention mask and reduce the feature weights with smaller weights, so as to prevent data overflow and improve the stability of model training. To prevent data overflow and improve the stability of model training.
[0135] Through the above embodiments, by calculating the outer product of the query vector and the key vector, the basis vectors can be quickly generated, and then the self-attention mask is decomposed into a linear combination of the basis vectors, reducing the computational complexity. By introducing the softmax activation function to transform the basis vectors into a probability distribution, the calculation process of the attention mechanism is further optimized.
[0136] In order to obtain the above recognition flag, in an optional implementation manner, the above step S2034 includes:
[0137] Step S20341: Extract features from the eighth global feature through the first large-kernel sparse convolution module to obtain the ninth global feature;
[0138] Step S20342: Extract features from the ninth global feature through the second large-kernel sparse convolution module to obtain the tenth global feature;
[0139] Step S20343: Perform feature encoding on the tenth global feature through the first low-rank CP tensor reconstruction attention module to obtain the first encoding result;
[0140] Step S20344: Perform feature encoding on the first encoding result through the second low-rank CP tensor reconstruction attention module to obtain the second encoding result;
[0141] Step S20345: Perform feature encoding on the second encoding result through the third low-rank CP tensor reconstruction attention module to obtain the encoded recognition result.
[0142] Specifically, the module receives the eighth global feature from the previous stage of processing. This feature set has undergone preliminary feature fusion. Through a series of LSC and CPA modules, deep feature extraction and encoding are performed on the eighth global feature, and finally an encoded recognition result is generated for the precise recognition of unique identification. Specifically, feature extraction includes the eighth global feature being subjected to feature extraction through the first large-kernel sparse convolution module to generate the ninth global feature. Among them, a convolution kernel of size (9, 9) is used to further expand the receptive field and capture more detailed information. Then, the ninth global feature is further subjected to deep feature extraction through the second large-kernel sparse convolution module to generate the tenth global feature. In this process, through multiple layers of LSC modules, the model can learn richer local and global feature information, preparing for subsequent feature encoding. The feature encoding part includes: the tenth global feature is input into the first low-rank CP tensor reconstruction attention module for feature encoding to generate the first encoding result. This process utilizes the CPA mechanism. By generating and performing outer product operations on query, key, and value vectors, the computational complexity of the attention mechanism is reduced. At the same time, through the multi-head attention mechanism, the efficiency and accuracy of feature encoding are improved. The first encoding result is further encoded through the second low-rank CP tensor reconstruction attention module and the third low-rank CP tensor reconstruction attention module, and finally an encoded recognition result is generated. The specific formulas for the above operations are as follows:
[0143] L1 = LSC(X)
[0144] L2 = LSC(L1)
[0145] L3 = CP_Attention(LSC(L2))
[0146] L4 = CP_Attention(LSC(L3))
[0147] L5 = CP_Attention(LSC(L4))
[0148] Among them, L1 and L2 are the ninth global feature and the tenth global feature, and L3, L4, and L5 are the first encoding result, the second encoding result, and the encoded recognition result.
[0149] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0150] The embodiment of the present application also provides a SIM card recognition device based on LCS and low-rank CP reconstruction attention. It should be noted that the SIM card recognition device based on LCS and low-rank CP reconstruction attention in the embodiment of the present application can be used to execute the SIM card recognition method based on LCS and low-rank CP reconstruction attention provided by the embodiment of the present application. The device is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that realizes a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0151] The following introduces the SIM card recognition device based on LCS and low-rank CP reconstruction attention provided by the embodiment of the present application.
[0152] Figure 3 It is a structural block diagram of the SIM card recognition device based on LCS and low-rank CP reconstruction attention according to the embodiment of the present application. As Figure 3 shown, the device includes:
[0153] The first acquisition unit 10 is used to recognize the first target image to obtain label data. The first target image is an image containing a target object, and the target object includes a preset graphic and a unique identifier of the preset graphic. The label data includes the coordinates of the minimum bounding rectangle of the unique identifier and the unique identifier;
[0154] The first processing unit 20 is used to detect the area of the unique identifier in the first target image through an alternative detection model to obtain a third target image. The alternative detection model at least includes a large kernel sparse convolution module, and the large kernel sparse convolution module performs feature extraction through large kernel convolution, sparse convolution, and depthwise separable convolution;
[0155] The second processing unit 30 is used to recognize the unique identifier of the third target image through an alternative recognition model to obtain a recognition flag. The alternative recognition model at least includes a low-rank CP tensor reconstruction attention module and a multi-head low-rank CP tensor reconstruction module. The low-rank CP tensor reconstruction attention module decomposes the self-attention tensor into rank-one tensors, and the multi-head low-rank CP tensor reconstruction module slices the input features according to the rank-one tensors and respectively extracts self-attention features and then fuses them;
[0156] The first determination unit 40 is used to determine a comprehensive performance parameter according to the third target image, coordinates, recognition flag, and unique identifier, and determine the alternative detection model and the alternative recognition model as the target detection model and the target recognition model when the comprehensive performance parameter is less than or equal to the first threshold;
[0157] A third processing unit 50, configured to identify a unique identifier of a first target image obtained in a service process according to a target recognition model and a target detection model, so as to obtain the unique identifier;
[0158] An input unit 60, configured to batch input the unique identifier into a target system, and the target system is configured to activate a target object according to the unique identifier.
[0159] Through this embodiment, a first acquisition unit obtains label data for a first target image, where the first target image is an image including a target object, the target object includes a preset graphic and a unique identifier of the preset graphic, and the label data includes the coordinates of the minimum bounding rectangle of the unique identifier and the unique identifier; a first processing unit detects the region of the unique identifier in the first target image through an alternative detection model to obtain a third target image, and the alternative detection model at least includes a large kernel sparse convolution module, and the large kernel sparse convolution module performs feature extraction through large kernel convolution, sparse convolution, and depthwise separable convolution; a second processing unit identifies the unique identifier of the third target image through an alternative recognition model to obtain a recognition flag, and the alternative recognition model at least includes a low-rank CP tensor reconstruction attention module and a multi-head low-rank CP tensor reconstruction module, and the low-rank CP tensor reconstruction attention module decomposes the self-attention tensor into a rank-one tensor, and the multi-head low-rank CP tensor reconstruction module slices the input features according to the rank-one tensor and respectively extracts self-attention features and then fuses them; a first determination unit determines a comprehensive performance parameter according to the third target image, the coordinates, the recognition flag, and the unique identifier, and when the comprehensive performance parameter is less than or equal to a first threshold, determines the alternative detection model and the alternative recognition model as the target detection model and the target recognition model; the third processing unit identifies the unique identifier of the first target image obtained in the service process according to the target recognition model and the target detection model to obtain the unique identifier; the input unit batches the unique identifier into the target system, and the target system is configured to activate the target object according to the unique identifier. In this application, the large kernel convolution is improved, and dilated convolution is introduced for sparse sampling to expand the field of view and improve the globality of features. At the same time, the convolution process is divided into depth convolution and pointwise convolution through depthwise separable convolution to reduce the number of parameters and avoid the reduction of the operation rate caused by large kernel convolution. At the same time, the low-rank CP tensor reconstruction attention module decomposes the self-attention mask into a rank-one tensor to reduce the number of parameters and computational complexity of the model, and then recombines the self-attention mask through the rank-one tensor to improve the accuracy while increasing the operation speed. The multi-head low-rank CP tensor reconstruction module slices the input image based on the channel dimension, calculates and fuses the attention features of different blocks respectively, and improves the model recognition accuracy by extracting global features. The detection model and the recognition model after training are used to implement the card number recognition, where the SIM card number is a specific embodiment. This device solves the problem that service handling in the prior art relies on manual input and has low efficiency.
[0160] To ensure the stability and accuracy of model training, in an alternative implementation, the above device further includes:
[0161] A fourth processing unit, configured to, after obtaining the first target image, randomly adjust the background color of the first target image to an image whose color three-channel values are within a first preset range, to obtain a plurality of fourth target images;
[0162] A fifth processing unit, configured to randomly adjust the font color of each fourth target image to an image whose color three-channel values are within a second preset range, to obtain a plurality of fifth target images;
[0163] A sixth processing unit, configured to process each fifth target image sequentially through random salt-and-pepper noise, Gaussian blur, and random rotation, to obtain a plurality of second target images.
[0164] To segment the above unique identifier to improve the accuracy of the recognition model, in an alternative implementation, the above first processing unit includes:
[0165] A first processing module, configured to perform feature extraction on the first target image through a large kernel sparse convolution module to obtain a first global feature;
[0166] A second processing module, configured to perform upsampling on the first global feature at different scales through a feature pyramid fusion module and perform fusion through residual connection to obtain a second global feature;
[0167] A third processing module, configured to perform differentiable binarization on the second global feature through a DB detection head decoding module to obtain a foreground image and a background image, and determine the foreground image as the third target image.
[0168] To improve the effect and speed of large kernel convolution, in an alternative implementation, the above first processing module includes:
[0169] A first processing sub-module, configured to process the first target image through a point convolution with a convolution kernel size of 1 and perform splitting along the first channel dimension to obtain a first feature map and a second feature map, process the first feature map through a horizontal convolution and process the second feature map through a vertical convolution, and perform concatenation through concat to obtain a third global feature;
[0170] A second processing sub-module, configured to process the third global feature through a point convolution with a convolution kernel size of 1 and perform splitting along the second channel dimension to obtain a third feature map and a fourth feature map;
[0171] The third processing sub-module is used to perform feature extraction on the third feature map through depthwise separable convolution with a dilation rate of 1 and a receptive field of the first preset size, and obtain a fourth global feature. Then, perform feature extraction on the fourth feature map through depthwise separable convolution with a dilation rate of 1 and a receptive field of the second preset size to obtain a fifth global feature;
[0172] The fourth processing sub-module is used to perform concat splicing on the fourth global feature and the fifth global feature to obtain a sixth global feature, and perform a residual connection on the sixth global feature and the first target image processed by point convolution with a kernel size of 1 to obtain a first global feature.
[0173] In order to reduce the computational complexity of the unique identifier recognition process, in an optional implementation manner, the above-mentioned second processing unit includes:
[0174] The fourth processing module is used to adjust the third target image respectively according to its length, width and number of channels through the basis vector generation module to obtain a fifth feature map, a sixth feature map and a seventh feature map. Process the fifth feature map through the linear layer of the basis vector generation module to obtain a query vector, process the sixth feature map through the linear layer of the basis vector generation module to obtain a key vector, and process the seventh feature map through the linear layer of the basis vector generation module to obtain a value vector;
[0175] The fifth processing module is used to decompose and reconstruct the self-attention mask according to the query vector, key vector and value vector through the low-rank CP tensor reconstruction attention module to obtain a target attention mask;
[0176] The sixth processing module is used to decompose the third target image based on the number of channels through the multi-head low-rank CP tensor reconstruction module to obtain multiple sub-feature maps, reconstruct each sub-feature map respectively according to the target attention mask to obtain multiple seventh global features, and splice the seventh global features to obtain an eighth global feature;
[0177] The seventh processing module is used to perform feature extraction and feature encoding on the eighth global feature through multiple large-kernel sparse convolution modules and multiple low-rank CP tensor reconstruction attention modules in sequence through the feature encoding and decoding module to obtain an encoded recognition result, and decode the encoded recognition result through the Transformer model to obtain a recognition flag.
[0178] In order to reconstruct the self-attention mask, in an optional implementation manner, the above-mentioned fifth processing module includes:
[0179] The calculation sub-module is used to calculate the outer product of the query vector and the key vector to obtain a first basis vector;
[0180] The fifth processing sub-module is used to decompose the self-attention mask into a linear combination of first basis vectors to obtain second basis vectors;
[0181] The sixth processing sub-module is used to process the second basis vectors through the softmax activation function and calculate the outer product with the value vectors to obtain a rank-one tensor.
[0182] The seventh processing sub-module is used to perform a linear combination through the rank-one tensor to obtain the target attention mask.
[0183] In order to obtain the above recognition flag, in an optional implementation manner, the above seventh processing module includes:
[0184] The eighth processing sub-module is used to extract features from the eighth global feature through the first large-kernel sparse convolution module to obtain the ninth global feature;
[0185] The ninth processing sub-module is used to extract features from the ninth global feature through the second large-kernel sparse convolution module to obtain the tenth global feature;
[0186] The tenth processing sub-module is used to perform feature encoding on the tenth global feature through the first low-rank CP tensor reconstruction attention module to obtain the first encoding result;
[0187] The eleventh processing sub-module is used to perform feature encoding on the first encoding result through the second low-rank CP tensor reconstruction attention module to obtain the second encoding result;
[0188] The twelfth processing sub-module is used to perform feature encoding on the second encoding result through the third low-rank CP tensor reconstruction attention module to obtain the encoded recognition result.
[0189] The above SIM card recognition device based on LCS and low-rank CP reconstruction attention includes a processor and a memory. The above first acquisition unit, first processing unit, second processing unit, first determination unit, third processing unit, input unit, etc. are all stored in the memory as program units, and the processor executes the above program units stored in the memory to implement corresponding functions. The above modules are all located in the same processor; or, the above modules are respectively located in different processors in any combination form.
[0190] The processor contains a kernel, and the kernel retrieves the corresponding program units from the memory. One or more kernels can be set, and the accuracy and efficiency of SIM card number recognition can be improved by adjusting the kernel parameters.
[0191] The memory may include non-permanent memory in the form of computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.
[0192] Embodiments of the present invention provide a computer-readable storage medium, and the computer-readable storage medium includes a stored program. When the program runs, it controls the device where the computer-readable storage medium is located to execute the SIM card identification method based on LCS and low-rank CP reconstruction attention.
[0193] Embodiments of the present invention provide a processor, and the processor is used to run a program. When the program runs, it executes the SIM card identification method based on LCS and low-rank CP reconstruction attention.
[0194] Embodiments of the present invention provide a service system. The service system includes a processor, a memory, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements at least the steps of the SIM card identification method based on LCS and low-rank CP reconstruction attention.
[0195] This application also provides a computer program product. When executed on a data processing device, it is adapted to execute a program initialized with at least the steps of the SIM card identification method based on LCS and low-rank CP reconstruction attention.
[0196] Obviously, those skilled in the art should understand that the various modules or steps of the present invention described above can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. They can be implemented with program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to implement. In this way, the present invention is not limited to any specific combination of hardware and software.
[0197] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0198] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, as well as the combination of flows and / or blocks in the flowchart and / or block diagram. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0199] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0200] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0201] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0202] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0203] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0204] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0205] From the above description, it can be seen that the above embodiments of the present application achieve the following technical effects:
[0206] 1), The SIM card recognition method based on LCS and low-rank CP reconstruction attention of the present application. First, for the first target image, label data is obtained. The first target image is an image containing a target object, and the target object includes a preset pattern and a unique identifier of the preset pattern. The label data includes the coordinates of the minimum bounding rectangle of the unique identifier and the unique identifier. Then, the region of the unique identifier in the first target image is detected by an alternative detection model to obtain a third target image. The alternative detection model at least includes a large kernel sparse convolution module, and the large kernel sparse convolution module extracts features through large kernel convolution, sparse convolution, and depthwise separable convolution. After that, the unique identifier of the third target image is recognized by an alternative recognition model to obtain a recognition flag. The alternative recognition model at least includes a low-rank CP tensor reconstruction attention module and a multi-head low-rank CP tensor reconstruction module. The low-rank CP tensor reconstruction attention module decomposes the self-attention tensor into rank-one tensors, and the multi-head low-rank CP tensor reconstruction module slices the input features according to the rank-one tensors and extracts self-attention features respectively and then fuses them. After that, a comprehensive performance parameter is determined according to the third target image, coordinates, recognition flag, and unique identifier, and when the comprehensive performance parameter is less than or equal to a first threshold, the alternative detection model and the alternative recognition model are determined as the target detection model and the target recognition model. After that, the unique identifier of the first target image obtained in the business process is recognized according to the target recognition model and the target detection model to obtain the unique identifier. Finally, the unique identifier is batch-input into the target system, and the target system is used to activate the target object according to the unique identifier. The present application improves the large kernel convolution by introducing dilated convolution for sparse sampling to expand the field of view and improve the global nature of the features. At the same time, the convolution process is divided into depth convolution and pointwise convolution by depthwise separable convolution to reduce the number of parameters and avoid the large kernel convolution from reducing the operation rate. At the same time, the low-rank CP tensor reconstruction attention module decomposes the self-attention mask into rank-one tensors to reduce the number of parameters and computational complexity of the model, and then reconstructs the self-attention mask through the rank-one tensors to improve the operation speed and accuracy at the same time. The multi-head low-rank CP tensor reconstruction module slices the input image based on the channel dimension, calculates and fuses the attention features of different blocks respectively, and improves the model recognition accuracy by extracting global features. The card number recognition is realized through the detection model and recognition model after training, and the SIM card number is one specific embodiment. This method solves the problem that business handling in the prior art relies on manual input and has low efficiency.
[0207] 2) The SIM card recognition device based on LCS and low-rank CP reconstruction attention of the present application. The first acquisition unit obtains label data for the first target image, where the first target image is an image containing a target object, the target object includes a preset pattern and a unique identifier of the preset pattern, and the label data includes the coordinates of the minimum bounding rectangle of the unique identifier and the unique identifier. The first processing unit detects the region of the unique identifier in the first target image through an alternative detection model to obtain a third target image. The alternative detection model at least includes a large kernel sparse convolution module, and the large kernel sparse convolution module performs feature extraction through large kernel convolution, sparse convolution, and depthwise separable convolution. The second processing unit identifies the unique identifier of the third target image through an alternative recognition model to obtain an identification flag. The alternative recognition model at least includes a low-rank CP tensor reconstruction attention module and a multi-head low-rank CP tensor reconstruction module. The low-rank CP tensor reconstruction attention module decomposes the self-attention tensor into rank-one tensors, and the multi-head low-rank CP tensor reconstruction module slices the input features according to the rank-one tensors and separately extracts self-attention features and then fuses them. The first determination unit determines a comprehensive performance parameter based on the third target image, coordinates, identification flag, and unique identifier, and when the comprehensive performance parameter is less than or equal to the first threshold, determines the alternative detection model and the alternative recognition model as the target detection model and the target recognition model. The third processing unit identifies the unique identifier of the first target image obtained during the service process according to the target recognition model and the target detection model to obtain the unique identifier. The input unit batches the unique identifier and inputs it into the target system, and the target system is used to activate the target object according to the unique identifier. The present application improves the large kernel convolution, introduces dilated convolution for sparse sampling to expand the field of view and improve the globality of features. At the same time, through depthwise separable convolution, the convolution process is divided into depth convolution and pointwise convolution to reduce the number of parameters and avoid the large kernel convolution from reducing the operation rate. At the same time, through the low-rank CP tensor reconstruction attention module, the self-attention mask is decomposed into rank-one tensors to reduce the number of parameters and computational complexity of the model, and then the self-attention mask is reconstructed by the rank-one tensors to improve the operation speed and accuracy at the same time. The multi-head low-rank CP tensor reconstruction module slices the input image based on the channel dimension, calculates and fuses the attention features of different blocks respectively, and improves the model recognition accuracy by extracting global features. The card number recognition is realized through the detection model and the recognition model after training, where the SIM card number is one specific embodiment. This device solves the problem in the prior art that business handling relies on manual input and has low efficiency.
[0208] The foregoing are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A SIM card recognition method based on LCS and low-rank CP reconstruction attention, characterized in that: include: Recognize a first target image to obtain label data, wherein the first target image is an image containing a target object, the target object includes a preset graphic and a unique identifier of the preset graphic, and the label data includes coordinates of a minimum circumscribed rectangle of the unique identifier and the unique identifier; Detecting the area uniquely identified in the first target image through an alternative detection model to obtain a third target image, wherein the alternative detection model at least includes a large-core sparse convolution module, and the large-core sparse convolution module performs feature extraction through large-core convolution, sparse convolution, and depth-separable convolution; The unique identifier of the third target image is identified by an alternative recognition model to obtain an identification mark, wherein the alternative recognition model at least includes a low-rank CP tensor reconstruction attention module and a multi-head low-rank CP tensor reconstruction module, wherein the low-rank CP tensor reconstruction attention module decomposes the self-attention tensor into a rank tensor, and the multi-head low-rank CP tensor reconstruction module slices the input features according to the rank tensor and extracts the self-attention features respectively and then fuses them; determining a comprehensive performance parameter according to the third target image, the coordinates, the identification mark and the unique identifier, and determining the candidate detection model and the candidate recognition model as the target detection model and the target recognition model when the comprehensive performance parameter is less than or equal to a first threshold; Identifying the unique identifier of the first target image acquired during the business process according to the target recognition model and the target detection model to obtain the unique identifier; The unique identifiers are input into a target system in batches, and the target system is used to activate the target objects according to the unique identifiers.
2. The method according to claim 1, characterized in that: After acquiring the first target image, the method further includes: According to the first target image, the background color is randomly adjusted to an image in which the values of the three color channels are within a first preset range to obtain a plurality of fourth target images; According to each of the fourth target images, the font color is randomly adjusted to an image whose values of the three color channels are within a second preset range to obtain a plurality of fifth target images; Each of the fifth target images is processed in turn by random salt and pepper noise, Gaussian blur and random rotation to obtain a plurality of second target images.
3. The method according to claim 1, characterized in that: The alternative detection model further includes a feature pyramid fusion module and a DB detection head decoding module. The alternative detection model is used to detect the region uniquely identified in the first target image to obtain a third target image, including: Performing feature extraction on the first target image by using the large-kernel sparse convolution module to obtain a first global feature; Upsampling the first global features at different scales through the feature pyramid fusion module, and fusing them through residual connections to obtain second global features; The second global feature is differentiably binarized by the DB detection head decoding module to obtain a foreground image and a background image, and the foreground image is determined as the third target image.
4. The method according to claim 3, characterized in that The first target image is subjected to feature extraction by the large-kernel sparse convolution module to obtain a first global feature, including: The first target image is processed by point convolution with a convolution kernel size of 1 and segmented along a first channel dimension to obtain a first feature map and a second feature map, the first feature map is processed by horizontal convolution and the second feature map is processed by vertical convolution, and concatenated to obtain a third global feature; Processing the third global feature by point convolution with a convolution kernel size of 1 and segmenting it along the second channel dimension to obtain a third feature map and a fourth feature map; Perform feature extraction on the third feature map by performing depth-separable convolution with a dilation rate of 1 and a convolution kernel receptive field of a first preset size and a sum, and obtain a fourth global feature; perform feature extraction on the fourth feature map by performing depth-separable convolution with a dilation rate of 1 and a convolution kernel receptive field of a second preset size and a sum, and obtain a fifth global feature; The fourth global feature and the fifth global feature are concat-joined to obtain a sixth global feature, and the sixth global feature and the first target image processed by point convolution with a convolution kernel size of 1 are residually linked to obtain the first global feature.
5. The method according to claim 1, characterized in that The alternative recognition model further includes a basis vector generation module and a feature encoding and decoding module. The unique identifier of the third target image is identified by the alternative recognition model to obtain an identification mark. The method includes: The basis vector generation module is used to adjust the length, width and number of channels of the third target image to obtain a fifth feature map, a sixth feature map and a seventh feature map, the fifth feature map is processed by the linear layer of the basis vector generation module to obtain a query vector, the sixth feature map is processed by the linear layer of the basis vector generation module to obtain a key vector, and the seventh feature map is processed by the linear layer of the basis vector generation module to obtain a value vector; Reconstructing the attention module through a low-rank CP tensor decomposes and reconstructs the self-attention mask according to the query vector, the key vector, and the value vector to obtain a target attention mask; Decomposing the third target image based on the number of channels by the multi-head low-rank CP tensor reconstruction module to obtain a plurality of sub-feature maps, reconstructing each of the sub-feature maps according to the target attention mask to obtain a plurality of seventh global features, and splicing the seventh global features to obtain an eighth global feature; The feature encoding and decoding module extracts and encodes the eighth global feature in turn through multiple large-kernel sparse convolution modules and multiple low-rank CP tensor reconstruction attention modules to obtain a coded recognition result, and the coded recognition result is decoded through a Transformer model to obtain the recognition mark.
6. The method according to claim 5, characterized in that The self-attention mask is decomposed and reconstructed according to the query vector, the key vector and the value vector by a low-rank CP tensor reconstructing attention module to obtain a target attention mask, including: Calculate the outer product of the query vector and the key vector to obtain a first basis vector; Decomposing the self-attention mask into a linear combination of the first basis vectors to obtain a second basis vector; The second basis vector is processed by a softmax activation function and the outer product with the value vector is calculated to obtain the rank tensor. The target attention mask is obtained by linearly combining the rank tensor.
7. The method according to claim 5, characterized in that The feature encoding and decoding module extracts and encodes the eighth global feature in turn through a plurality of the large-core sparse convolution modules and a plurality of low-rank CP tensor reconstruction attention modules to obtain a coding recognition result, including: Extracting the eighth global feature through a first large-kernel sparse convolution module to obtain a ninth global feature; Extracting the ninth global feature through a second large-kernel sparse convolution module to obtain a tenth global feature; Performing feature encoding on the tenth global feature through a first low-rank CP tensor reconstruction attention module to obtain a first encoding result; Performing feature encoding on the first encoding result through a second low-rank CP tensor reconstruction attention module to obtain a second encoding result; The second encoding result is feature encoded through a third low-rank CP tensor reconstruction attention module to obtain the encoding recognition result.
8. A SIM card recognition device based on LCS and low-rank CP reconstruction attention, characterized in that: The device comprises: A first acquisition unit is used to identify a first target image to obtain label data, wherein the first target image is an image containing a target object, the target object includes a preset graphic and a unique identifier of the preset graphic, and the label data includes coordinates of a minimum circumscribed rectangle of the unique identifier and the unique identifier; A first processing unit, configured to detect the area uniquely identified in the first target image through an alternative detection model to obtain a third target image, wherein the alternative detection model at least includes a large-core sparse convolution module, and the large-core sparse convolution module performs feature extraction through large-core convolution, sparse convolution, and depthwise separable convolution; A second processing unit is used to identify the unique identifier of the third target image through an alternative recognition model to obtain an identification mark, wherein the alternative recognition model at least includes a low-rank CP tensor reconstruction attention module and a multi-head low-rank CP tensor reconstruction module, wherein the low-rank CP tensor reconstruction attention module decomposes the self-attention tensor into a rank tensor, and the multi-head low-rank CP tensor reconstruction module slices the input features according to the rank tensor and extracts the self-attention features respectively and then fuses them; a first determining unit, configured to determine a comprehensive performance parameter according to the third target image, the coordinates, the identification mark and the unique identifier, and determine the candidate detection model and the candidate recognition model as the target detection model and the target recognition model when the comprehensive performance parameter is less than or equal to a first threshold; A third processing unit, configured to identify the unique identifier of the first target image acquired during the business process according to the target recognition model and the target detection model, so as to obtain the unique identifier; The input unit is used to input the unique identifiers in batches into a target system, and the target system is used to activate the target objects according to the unique identifiers.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the method according to any one of claims 1 to 7.
10. A business system, characterized in that: include: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include methods for executing any one of claims 1 to 7.
Citation Information
Cited By
Modulation identification method and system based on high order moment and spectrum thereof
CN120750707A