Lightweight image recognition method, system and device and storage medium
By optimizing the image recognition model through depthwise separable convolution and channel attention mechanisms, and combining the H-Swish activation function and block caching mechanism, the problem of limited computing resources in medical auxiliary diagnosis is solved, and efficient and accurate image recognition is achieved.
Patent Information
- Application Number
- CN202510784846.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-11-07
AI Technical Summary
Image recognition in medical auxiliary diagnosis involves a large amount of computation, making it difficult to apply effectively in primary healthcare institutions with limited computing resources.
A lightweight image recognition method is adopted, which uses depthwise separable convolutional modules and channel attention mechanisms, combined with the H-Swish activation function to optimize the model structure, and improves CPU adaptability through a block caching mechanism.
It reduces the computational load of image recognition, improves computational efficiency and accuracy on the CPU, adapts to different hardware environments, and meets the needs of medical image recognition.
Smart Images

Figure CN120912896A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of image recognition, and particularly relates to a lightweight image recognition method, system, device and storage medium. BACKGROUND
[0002] Image recognition performs well in the field of medical auxiliary diagnosis. However, medical auxiliary diagnosis has a high requirement for accuracy, and therefore requires a large amount of computing power. For example, DenseNet-121 or ResNet-152, these models usually have hundreds of layers, and use dense connections (DenseNet) or residual connections (ResNet) to alleviate the problem of gradient disappearance, ensuring the effective training of deep networks. And whether the parameter quantity is large. Processing an X-ray film of a standard size (such as 224x224 or 512x512), ResNet-152 may require billions of floating-point operations (GFLOPs). DenseNet-121 is relatively less, but it is also computationally intensive.
[0003] However, the computing resources of the computing devices of some primary medical institutions are limited, and therefore how to reduce the computing amount of image recognition for auxiliary diagnosis while ensuring the accuracy requirement is a technical problem to be solved. SUMMARY
[0004] In view of the above shortcomings of the prior art, the present application provides a lightweight image recognition method, system, device and storage medium to solve the above technical problems.
[0005] In a first aspect, the present application provides a lightweight image recognition method, comprising: obtaining an image to be recognized; inputting the image to be recognized into a pre-trained recognition model to obtain a category corresponding to the image to be recognized; The recognition model comprises four depth separable convolution modules, and a channel attention mechanism is embedded after the second depth separable convolution module and the fourth depth separable convolution module, respectively, and the depth separable convolution module adopts an H-Swish activation function.
[0006] In an optional embodiment, the image to be recognized is obtained, comprising: obtaining an original image and pre-processing the original image, wherein the pre-processing comprises: scaling the image to a specified pixel size; applying histogram equalization to a grayscale image, and equalizing the Y channel after converting a color image to YUV space; using a non-local mean denoising technique using image enhancement technology.
[0007] In an optional embodiment, the deep separable convolution module comprises: a deep convolution layer with a kernel size of 3x3 and a number of groups equal to the number of input channels; a point-wise convolution layer with a kernel size of 1x1 for channel fusion and dimension transformation; an H-Swish activation function.
[0008] In an optional embodiment, the channel attention mechanism is multiplied with the input feature map channel by channel to re-label the feature importance, and the channel attention mechanism comprises: a global average pooling layer outputting a feature map of 1x1xC; a first fully connected layer for compressing channels to C / 16; a second fully connected layer for restoring the original number of channels.
[0009] In an optional embodiment, the recognition model further comprises an output layer, which comprises: a 1x1 convolution layer for reducing the channel dimension to 5 through linear transformation to obtain a 5-channel spatial activation map; a global average pooling for converting the 5-channel spatial activation map into 5 scalar values; a Softmax classification function for determining the corresponding class according to the 5 scalar values.
[0010] In an optional embodiment, the method further comprises: dividing the feature map extracted from the to-be-recognized image by the recognition model into feature map blocks according to the cache structure of the CPU, and caching the feature map blocks; reorganizing the input data of the convolution kernel and the calculation strategy of the convolution kernel so that the convolution kernel processes the feature map blocks in turn; merging the processing results of all feature map blocks, and transmitting the merged processing results to the input layer.
[0011] In an optional embodiment, dividing the feature map extracted from the to-be-recognized image by the recognition model into feature map blocks according to the cache structure of the CPU, and caching the feature map blocks, comprises: calculating the optimal block size (T w ,T h ) so that: (T w +2P)(T h +2P)xCxs element ≤αS cache wherein the feature map size is WxHxC, represented as WxHxC; the convolution kernel size is KxK; the cache capacity is S cache ; and the size of each data element is s elementP is the padding size; and a is a safety factor to reserve space for the kernel and intermediate results.
[0012] In a second aspect, the present application also provides a lightweight image recognition system, comprising: An acquisition module is configured to acquire an image to be recognized. An identification module is configured to input the image to be recognized into a pre-trained identification model to obtain a category corresponding to the image to be recognized. The identification model comprises four depth separable convolution modules, and a channel attention mechanism is embedded after the second depth separable convolution module and the fourth depth separable convolution module, respectively, and the depth separable convolution module adopts an H-Swish activation function.
[0013] In a third aspect, a device is provided, comprising: A memory is configured to store a lightweight image recognition program. A processor is configured to implement the steps of the lightweight image recognition method provided in the first aspect when the lightweight image recognition program is executed.
[0014] In a fourth aspect, a computer-readable storage medium is provided, and the storage medium stores a lightweight image recognition program, and the steps of the lightweight image recognition method provided in the first aspect are implemented when the lightweight image recognition program is executed by a processor.
[0015] The lightweight image recognition method, system, device and storage medium provided by the present application have the beneficial effects that the multi-layer feature extraction module is constructed by the depth separable convolution, the parameter amount is reduced, the channel attention mechanism is added at the key level to improve the feature expression ability, the H-Swish is used to replace the ReLU to balance the calculation efficiency and the nonlinear expression ability, thereby reducing the calculation amount of the medical image recognition and reducing the requirement for the computing power of the computing device.
[0016] In addition, the design principle of the present application is reliable, the structure is simple, and it has very wide application prospect. BRIEF DESCRIPTION OF DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced as follows, and obviously, other drawings can also be obtained by those skilled in the art without creative labor.
[0018] Figure 1 is a schematic flowchart of the method of one embodiment of the present application.
[0019] Figure 2 is a flowchart of the training and application of the identification model of the method of one embodiment of the present application.
[0020] Figure 3 is a schematic block diagram of a system according to an embodiment of the present application.
[0021] Figure 4 is a schematic block diagram of a device according to an embodiment of the present application. DETAILED DESCRIPTION
[0022] In order to make the technical personnel in the art better understand the technical solutions in the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor should belong to the scope of protection of the present application.
[0023] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs. The terminology used in the specification of the present application is only for the purpose of describing specific embodiments and is not intended to limit the present application.
[0024] The lightweight image recognition method provided by the embodiments of the present application is executed by a computer device, and correspondingly, the lightweight image recognition system runs in the computer device.
[0025] Figure 1 is a schematic flow chart of a method according to an embodiment of the present application. In which, Figure 1 The execution subject can be a lightweight image recognition system. According to different needs, the order of steps in the flow chart can be changed, and some can be omitted.
[0026] As Figure 1 shown, the method comprises: S1. obtaining an image to be recognized; S2. inputting the image to be recognized into a pre-trained recognition model to obtain a category corresponding to the image to be recognized; The recognition model comprises four depth separable convolution modules, and a channel attention mechanism is embedded after the second depth separable convolution module and after the fourth depth separable convolution module, respectively, and the depth separable convolution module adopts an H-Swish activation function.
[0027] In an embodiment of the present application, the image to be recognized is an original eye image output by a fundus imaging device, which is used for diagnosing cataract.
[0028] After obtaining the original eye image, it is preprocessed: Size normalization: uniformly scale the image to 224x224 pixels; apply histogram equalization for grayscale images, and equalize the Y channel after converting color images to YUV space; apply CLAHE algorithm (Clip Limit=2.0, TileGrid Size=8x8), and use non-local mean denoising (h=10, search window=21x21).
[0029] The pre-processed eye image is input into the recognition model, and the degree of cataract is obtained.
[0030] Specifically, the recognition model includes four depth separable convolution modules, and a channel attention mechanism is embedded after the second depth separable convolution module and the fourth depth separable convolution module, and an H-Swish activation function is used.
[0031] The depth separable convolution module includes: a depth convolution layer with a kernel size of 3x3 and a grouping number equal to the input channel number; a point-wise convolution layer with a kernel size of 1x1 for channel fusion and dimension transformation; and an H-Swish activation function.
[0032] The channel attention mechanism is multiplied with the input feature map channel by channel to re-rate the feature importance, and the channel attention mechanism includes: a global average pooling layer outputting a 1x1xC feature map; a fully connected layer 1 for compressing the channel to C / 16; and a fully connected layer 2 for restoring the original channel number.
[0033] The recognition model further includes an output layer, which includes: a 1x1 convolution layer for reducing the channel dimension to 5 through linear transformation to obtain a 5-channel spatial activation map; a global average pooling for converting the 5-channel spatial activation map into 5 scalar values; and a Softmax classification function for determining the corresponding category according to the 5 scalar values. Replacing the fully connected layer with a 1x1 convolution layer can reduce the computational load.
[0034] Specifically, the model is used for cataract diagnosis, adopts a lightweight depth separable convolution architecture, includes four depth separable convolution modules, and embeds a channel attention mechanism at a key position, and finally realizes the classification and severity evaluation of cataract through an optimized output layer. The overall architecture design of the model fully considers the needs of the ophthalmic clinical application scene, ensures the diagnosis accuracy, reduces the computational resource demand, and is convenient for integration on ophthalmic equipment.
[0035] Depth separable convolution module: Each depth separable convolution module is composed of three parts: Depth convolution layer: using a 3x3 convolution kernel, the grouping number is equal to the input channel number Point-wise convolution layer: using a 1x1 convolution kernel, realizing channel fusion and dimension transformation H-Swish activation function: provides non-linear transformation capability The computational complexity of convolution operation is:
[0036] While the computational complexity of depthwise separable convolution is:
[0037] When the number of output channels is large, the computational complexity of depthwise separable convolution is about 1 / 9 of traditional convolution, significantly reducing the model complexity.
[0038] Application advantages in cataract diagnosis: Reducing parameter quantity: suitable for processing high-resolution ophthalmic images; Multi-scale feature extraction: can capture both microscopic structure and macroscopic morphological features of the lens Prevent overfitting: better generalization ability on limited ophthalmic image dataset.
[0039] (Three) Channel attention mechanism Channel attention mechanism reweights the channels of feature maps, enabling the model to adaptively focus on important channels and suppress unimportant channels. The specific implementation steps are: Global average pooling: compress the input feature map into a 1×1×C feature vector; Channel compression and recovery: realize dimension reduction (to C / 16) and dimension increase (restore to C) through two fully connected layers; Feature reweighting: multiply the obtained attention weights with the original feature map channel by channel.
[0040] Role in cataract diagnosis: Enhance feature channels related to lesions: for example, in nuclear cataract, highlight the feature response of the lens nucleus region; Suppress noise and irrelevant information: reduce the interference of factors such as corneal reflection and instrument artifacts on diagnosis; Improve model interpretability: visualize attention weights to help doctors understand the basis of model decision-making.
[0041] In this model, the channel attention mechanism is embedded after the second and fourth depthwise separable convolution modules. The consideration of this design is: Attention mechanism after the second module: focus on local detailed features of the lens, such as subtle changes in early cortical opacity; Attention mechanism after the fourth module: focus on higher-level semantic features, such as the overall morphology of the lens and the distribution pattern of opacity.
[0042] (Four) H-Swish activation function H-Swish is an approximate implementation of the Swish activation function, which has the following advantages over traditional activation functions: Smooth activation curve: better gradient propagation characteristics on the positive half-axis; Non-zero gradient: alleviates the problem of gradient vanishing, helping to train deeper networks; High computational efficiency: avoids the exponential operation in the Swish function, more suitable for mobile deployment.
[0043] Application advantages in ophthalmic images: Better preservation of lens edge information: particularly important for the diagnosis of posterior subcapsular cataract; Enhanced response to weak lesion characteristics: helpful for early detection of cataract.
[0044] Four, output layer design and cataract diagnosis classification The output layer of this model uses the following structure: 1×1 convolutional layer: reduces the channel dimension to 5, obtaining a 5-channel spatial activation map; Global average pooling: converts the 5-channel spatial activation map into 5 scalar values; Softmax classification function: converts the 5 scalar values into a probability distribution, corresponding to 5 diagnostic categories.
[0045] The advantages of using 1×1 convolution instead of fully connected layer: Reduced computational complexity: for a feature map with input size H×W×C, the parameter amount of 1×1 convolution is C×5, much smaller than that of fully connected layer; Preservation of spatial information: 1×1 convolution can preserve the spatial structure of the feature map, which is crucial for locating the lens lesion area; Translation invariance: makes the model insensitive to lesion location, improving diagnostic stability.
[0046] (Three) Cataract diagnosis category design This model classifies cataract diagnosis into 5 categories: Normal lens, early nuclear cataract, moderate nuclear cataract, advanced nuclear cataract, and other types of cataract (including cortical and posterior subcapsular).
[0047] This classification method is highly consistent with the actual application needs of clinical practice, providing clear diagnostic reference for doctors.
[0048] This model can be integrated into an ophthalmic slit lamp microscope or fundus camera to provide real-time cataract diagnosis suggestions. When examining patients, the system can display automatic diagnosis results and confidence scores simultaneously, helping doctors make more accurate diagnostic decisions.
[0049] Hardware specifications: CPU: Intel(R) Core(TM) i7-8700 CPU @ 3.20GHz, 32GB (RAM). Please refer to [reference needed]. Figure 2 The specific process of model training and inference: Step 1: Local Image Annotation. Image annotation by folder name is supported. A standard dataset in a preset format is obtained.
[0050] The cataract annotation dataset contains 4000 slit-lamp images, covering mild, moderate, and severe cataracts, as well as nuclear and cortical subtypes. Image size normalization was performed: images were uniformly scaled to 224×224 pixels; histogram equalization was applied to grayscale images, and Y-channel equalization was applied to color images after conversion to YUV space; the CLAHE algorithm (Clip Limit=2.0, Tile Grid Size=8×8) was applied, and non-local means denoising was used (h=10, search window=21×21).
[0051] The processed images are then labeled.
[0052] Step 2: Model training. The PPLCNet algorithm is used to build a multi-layer feature extraction module based on depthwise separable convolutions to reduce the number of parameters. Channel attention mechanism is added to key layers to improve feature expression ability. H-Swish is used to replace ReLU to balance computational efficiency and non-linear expression ability. The learning rate decay strategy and batch size are adjusted based on the CPU computing characteristics.
[0053] The PPLCNet model architecture consists of four depthwise separable convolutional modules, each containing three convolutional layers (3×3 kernel size, alternating strides of 1 / 2). Channel attention (compression ratio = 16) is added after the second and fourth modules. H-Swish (β = 1.5, dynamically adjusting the non-linear response) is used. Global average pooling layers and fully connected layers (output nodes = 5, corresponding to normal, mild, moderate, severe, and subtype labels) are employed. Initial learning rate = 0.001, weight decay = 0.05, and cosine annealing decay (T_max = 50, η_min = 0.0001). Focal Loss (α = 0.8, γ = 2.0, addressing class imbalance). Batch Size: 16 (limited by CPU memory). Training duration: 100 epochs.
[0054] Step 3: Visualize the training process (TensorBoard, VisualDL) to view training logs and training reports.
[0055] Step 4: Model accuracy evaluation, displaying the model's accuracy, F1 score, precision, and recall.
[0056] Accuracy, F1 Score (macro average), Precision, Recall, and single image processing time (ms).
[0057] Step five: model inference, used for clinical auxiliary diagnosis to recognize the cataract image.
[0058] On the basis of the above embodiment, in order to further improve the adaptability of the medical image recognition method provided by the above embodiment to CPU, a block caching mechanism is adopted as an implementable manner.
[0059] 1. According to the cache structure of CPU, the feature map extracted from the to-be-recognized image by the recognition model is divided into feature map blocks, and the feature map blocks are cached.
[0060] Calculate the optimal block size (T w ,T h ) so that: (T w +2P)(T h +2P)×C×s element ≤αS cache Where, the feature map size is width × height × channel number, denoted as W × H × C; the convolution kernel size is K × K; the cache capacity is S cache ; each data element size is s element ; P is the padding size; alpha is the safety factor, which reserves space for the kernel and intermediate results.
[0061] Find the solution that makes T w ×T h max by traversal: import math def find_optimal_tile(P, C, s_element, alpha, S_cache, W, H):A_max =(alpha * S_cache) / (C * s_element),max_area = 0, optimal_tile = (0, 0), # Possible width range (considering step alignment) for T_w in range(16, min(W, int(math.sqrt(A_max))), 16):# Maximum possible height, max_h = min(H, int((A_max / (T_w + 2*P)) - 2*P)),for T_h in range(16, max_h, 16):area = (T_w + 2*P) * (T_h + 2*P),actual_area = T_w * T_h,if area<= A_max and actual_area>max_area:max_area =actual_area,optimal_tile = (T_w, T_h),return optimal_tile。
[0062] The pyramid processing order (also known as the tiling traversal strategy) is used to spatially divide the feature map based on the optimal tile size calculated, and a cache-friendly efficient processing is achieved through hierarchical calculation.
[0063] The pyramid processing divides the large feature map into multiple fixed-size tiles through tiling, each tile is calculated independently, and the entire feature map is traversed through sliding windows between tiles. The core goal is: spatial locality: ensure that the data of each tile can be completely placed in the cache, reducing memory access; computational locality: reuse the convolution kernel within the tile to improve data reuse rate; the CPU allocates cache addresses for each tile and writes to the cache.
[0064] 2. Reorganize the input data of the convolution kernel and the calculation strategy of the convolution kernel, so that the convolution kernel processes the feature map tiles in turn.
[0065] Set cache-aware traversal order: / / Optimized tiling processing loop, for (int y = 0; y<height; y += tile_h) {for (int x = 0; x<width; x += tile_w) { / / Process the current tile, process_tile(y, x, tile_h, tile_w);}} / / Compare with traditional row-first: memory jumps cause a lot of cache misses.
[0066] Convert data layout: convert data from NHWC (N=batch, H=height, W=width, C=channel) to HWC (height tile, width tile, channel tile): memory layout: (H,W,C)→(T h ,Tw C b ); wherein C b is a channel block (typically 8 or 16 channels in a group).
[0067] Cache-aware convolution kernel rearrangement: Rearrange the convolution kernel as:
[0068] wherein C g =C / C b .
[0069] Convolution can be converted to matrix multiplication:
[0070] For a block, the block matrix multiplication strategy:
[0071] Form a block order according to the position of each block in the feature map, process the blocks in order, and the processing flow for each block: perform im2col conversion; block matrix multiplication; accumulate partial results.
[0072] 3. Merge the processing results of all feature map blocks and transmit the merged processing results to the input layer.
[0073] The advantages of this method are: controlling data volume through block, ensuring that the working set is within the cache capacity, improving spatial locality through data reorganization, increasing data reuse rate through block matrix multiplication, and adapting to different hardware environments.
[0074] In the medical field, image recognition technology is increasingly widely used, from lesion detection of X-ray films to tumor recognition of nuclear magnetic resonance images, all of which rely on efficient and accurate image recognition methods. However, under different hardware environments, especially when running on CPU, the efficiency and performance of medical image recognition will be significantly affected. In order to further improve the adaptability of medical image recognition methods to CPU, using block cache mechanism is a highly feasible and efficient technical solution.
[0075] I. Core idea of block cache mechanism The core of the block cache mechanism is to reasonably divide the feature map, reorganize the data and calculation strategy, fully utilize the cache structure of the CPU, reduce memory access overhead, and improve data reuse rate, thereby improving the running efficiency of medical image recognition on CPU. It mainly includes three key steps: feature map block and cache, convolution kernel input data and calculation strategy reorganization, processing result merging and transmission. Next, each step will be described in detail.
[0076] II. Feature map block and cache (I) Feature map partitioning based on CPU cache structure CPU cache is a high-speed storage component used to temporarily store data that the CPU may frequently access in the near future, thereby reducing the number of accesses to main memory and improving data read speed. To enable the feature maps extracted by the medical image recognition model to utilize the CPU cache more efficiently, the feature maps need to be divided into multiple feature map blocks according to the CPU cache structure.
[0077] Determining the appropriate block size is crucial when partitioning feature maps. This is achieved by calculating the optimal block size (Tw, Th).
[0078] (II) Calculation and Implementation of Optimal Block Size To find the solution that maximizes TwÃTh, we can implement a traversal calculation by writing code. First, we calculate A_max according to the formula, which represents the maximum theoretical area of the block size under the cache capacity limit. Then, we traverse the possible width T_w and height T_h using two nested loops. During the traversal, we calculate the corresponding maximum height max_h based on the current T_w, and then try different T_h values. Each time, we calculate the actual area of the block (actual_area) and the area including padding (area). When the area satisfies the cache capacity limit and actual_area is greater than the currently recorded maximum area (max_area), we update max_area and the optimal block size (optimal_tile).
[0079] (III) Pyramid-style processing sequence After obtaining the optimal block size, the feature map is spatially segmented using a pyramid-style processing order (also known as a block traversal strategy). The core idea of pyramid-style processing is to divide the large feature map into multiple fixed-size blocks (tiles), each of which is computed independently, and the blocks are traversed through the entire feature map using a sliding window (stride).
[0080] This approach has two key objectives: First, it achieves spatial locality, ensuring that each block of data is fully cached, reducing memory accesses. When the CPU accesses data, reading from the cache is significantly faster than reading from main memory. When data for each block can be found in the cache, the CPU's waiting time for data transfer from main memory is greatly reduced. Second, it achieves computational locality, reusing convolution kernels within blocks to improve data reuse. During convolution calculations on each block, the convolution kernel can be used multiple times within the block, avoiding frequent readings of the same kernel data from memory and further improving computational efficiency.
[0081] In actual operation, the CPU allocates a cache address for each tile and writes tile data into the cache, so that subsequent computing operations can quickly access the required data.
[0082] III. Convolution kernel input data and calculation strategy reorganization (I) Set cache-aware traversal order In traditional image data processing, the row-major traversal method is usually used. However, this method can cause memory jumps, resulting in a large amount of data not being in the CPU cache, thus generating a large number of cache misses and reducing computing efficiency. To solve this problem, a cache-aware traversal order is set. By performing loop traversal with tile height tile_h and tile width tile_w as the step size, it is ensured that when processing each tile, the data in the cache is used as much as possible, reducing unnecessary memory access.
[0083] (II) Convert data layout The conversion of data layout is also one of the key steps to improve the adaptability of the CPU. The data is converted from the common NHWC (N=batch, H=height, W=width, C=channel) format to the HWC (height tile, width tile, channel tile) format, and the layout in memory is converted from (H, W, C) to (T h , T w , C b ). Among them, C b is the channel block, usually 8 or 16 channels are divided into a group. This conversion of data layout makes the storage of data in memory more in line with the needs of tile processing, and can better utilize the CPU cache to improve the locality of data access.
[0084] (III) Cache-aware convolution kernel rearrangement In addition to data layout conversion, the convolution kernel is also rearranged. The convolution kernel is rearranged into a specific form, where C g =C / C b . Through this rearrangement, the convolution operation can be converted into matrix multiplication, which is of great significance to improving computing efficiency. Because there are many mature optimization algorithms and implementations of matrix multiplication in computers, after converting the convolution operation into matrix multiplication, these optimization techniques can be used to speed up the calculation.
[0085] For a tile, the tile matrix multiplication strategy is used. According to the position of each tile in the feature map, a tile order is formed, and the tiles are processed in turn. The processing flow for each tile includes: first, perform im2col conversion to convert the two-dimensional convolution operation into matrix multiplication form; then perform tile matrix multiplication; finally, accumulate the partial results. Through such a processing flow, the computing power and cache characteristics of the CPU are fully utilized, improving the efficiency of convolution calculation.
[0086] IV. Merging and transmission of processing results After completing the processing of all feature map tiles, the processing results of each tile need to be merged. The process of merging the processing results needs to ensure the accuracy and integrity of the data, avoiding data loss or incorrect merging. The merged processing results will be transmitted to the input layer for further processing or output.
[0087] V. Advantages of the tile-based caching mechanism This medical image recognition method based on the tile-based caching mechanism has many significant advantages. First, by controlling the data volume through tiling, the working set is ensured to be within the cache capacity. This means that during computation, most data can be quickly read from the cache, greatly reducing the number of accesses to the main memory and reducing the latency of data transmission, thereby improving overall computing efficiency.
[0088] Second, data reorganization improves spatial locality. Whether it is data layout conversion or convolution kernel rearrangement, it makes the storage and access of data in memory more reasonable, allowing the CPU to more efficiently utilize the data in the cache, further improving computing performance.
[0089] Third, the tiled matrix multiplication increases the data reuse rate. During the tiling process, the convolution kernel and data are reused within the tile, avoiding a large amount of redundant data reading and improving the utilization of computing resources.
[0090] Finally, this method has strong adaptability and can adapt to different hardware environments. Whether it is the size difference of CPU cache capacity or the difference in other hardware characteristics, by reasonably adjusting the tile size and related parameters, the performance of medical image recognition can be optimized to some extent, ensuring efficient operation in various CPU environments.
[0091] In summary, the medical image recognition method based on the tile-based caching mechanism significantly improves the adaptability to CPUs by scientifically and reasonably tiling feature maps, reorganizing data and computing strategies, and merging and transmitting results. It provides strong technical support for the efficient application of medical image recognition technology on the CPU platform, and is expected to play a greater role in medical diagnosis, disease analysis, and other fields, promoting the further development of medical imaging technology.
[0092] In some embodiments, the lightweight image recognition system can include multiple functional modules composed of computer program segments. The computer programs of each program segment in the lightweight image recognition system can be stored in the memory of the computer device and executed by at least one processor to perform the functions of lightweight image recognition (see Figure 1 Description).
[0093] In this embodiment, the lightweight image recognition system can be divided into a plurality of functional modules according to the functions performed by the system, as shown in the figure. The module referred to in the present application refers to a series of computer program segments capable of being executed by at least one processor and capable of completing a fixed function, which are stored in the memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments. Figure 3 The module referred to in the present application refers to a series of computer program segments capable of being executed by at least one processor and capable of completing a fixed function, which are stored in the memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments.
[0094] The acquisition module is configured to acquire an image to be recognized. The recognition module is configured to input the image to be recognized into a pre-trained recognition model to obtain a category corresponding to the image to be recognized. The recognition model includes four depth separable convolution modules, and a channel attention mechanism is embedded after the second depth separable convolution module and the fourth depth separable convolution module, and an H-Swish activation function is used.
[0095] Figure 4 The lightweight image recognition method provided by the embodiments of the present application can be applied to a device. Those skilled in the art can understand that the device structure involved in the embodiments of the present application does not constitute a limitation on the device, and the device can include more or fewer components than the illustration, or combine certain components, or different component arrangements. In the embodiments of the present application, the device includes but is not limited to a laptop computer, a desktop computer, a workstation, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the embodiments of the present application described and / or claimed herein.
[0096] The device 400 can include a processor 410, a memory 420, and a communication unit 430. These components communicate through one or more buses, and those skilled in the art can understand that the structure of the server shown in the figure does not constitute a limitation on the present application, and it can be a bus structure or a star structure, and can include more or fewer components than the illustration, or combine certain components, or different component arrangements.
[0097] The memory 420 can be used to store the execution instructions of the processor 410, and the memory 420 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk. When the execution instructions in the memory 420 are executed by the processor 410, the device 400 can perform part or all of the steps in the following method embodiments.
[0098] The processor 410 is the control center of the storage device, which connects various parts of the entire electronic device through various interfaces and lines, and performs various functions of the electronic device and / or processes data by running or executing software programs and / or modules stored in the memory 420 and calling data stored in the memory. The processor can be composed of an integrated circuit (IC), for example, it can be composed of a single packaged IC, or it can be composed of multiple packaged ICs with the same function or different functions connected. For example, the processor 410 can only include a central processing unit (CPU). In the embodiments of the present application, the CPU can be a single operation core or can include multiple operation cores.
[0099] The communication unit 430 is used to establish a communication channel, so that the storage device can communicate with other devices. Receive user data sent by other devices or send user data to other devices.
[0100] The present application also provides a computer storage medium, wherein the computer storage medium can store a program, and the program can include part or all of the steps in the embodiments provided by the present application when executed. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.
[0101] Those skilled in the art can clearly understand that the technical solutions in the embodiments of the present application can be implemented by means of software plus necessary universal hardware platforms. Based on such an understanding, the technical solutions in the embodiments of the present application can be embodied in the form of a software product, which is stored in a storage medium, such as a USB flash disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and the like, and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a second device, a network device, or the like) to execute all or part of the steps of the methods described in the embodiments of the present application.
[0102] The same or similar parts among the various embodiments in the specification can be referred to each other. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description in the method embodiments.
[0103] In the several embodiments provided in the present application, it should be understood that the disclosed system and method can be implemented in other ways. For example, the system embodiments described above are merely schematic. For example, the division of the modules is merely a logical function division. In actual implementation, another division manner can be used. For example, a plurality of modules or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed modules can be indirect coupling or communication connection through some interfaces. The coupling or communication connection can be electrical, mechanical or in other forms.
[0104] The modules described as separate components can or can not be physically separate, and the components displayed as modules can or can not be physical modules, i.e., can be located in one place or distributed on a plurality of network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the embodiments.
[0105] In addition, each functional module in the various embodiments of the present application can be integrated into a processing module, or each module can exist physically independently, or two or more modules can be integrated into one module.
[0106] Although the present application has been described in detail with reference to the preferred embodiments, it should be understood that the application is not limited to those preferred embodiments. Various modifications and equivalents can be made by those skilled in the art without departing from the spirit and scope of the application. Any and all modifications and equivalents are intended to be included within the scope of the present application.
Claims
1. A lightweight image recognition method, characterized by, The method comprises: obtaining an image to be identified; inputting the image to be identified into a pre-trained identification model to obtain a category corresponding to the image to be identified; the identification model comprises four deep separable convolution modules, and a channel attention mechanism is embedded after the second deep separable convolution module and the fourth deep separable convolution module, and the deep separable convolution module adopts an H-Swish activation function.
2. The method of claim 1, wherein, The method comprises: obtaining an original image and pre-processing the original image, wherein the pre-processing comprises: scaling the image to a specified pixel size; performing histogram equalization on a grayscale image, and performing equalization on a Y channel after converting a color image to a YUV space; performing non-local mean denoising using an image enhancement technique.
3. The method of claim 1, wherein, The deep separable convolution module comprises: a deep convolution layer with a kernel size of 3*3 and a grouping number equal to an input channel number; a point-wise convolution layer with a kernel size of 1*1, which is used for channel fusion and dimension transformation; an H-Swish activation function.
4. The method of claim 1, wherein, The channel attention mechanism is multiplied with an input feature map channel by channel to re-rate the importance of the feature, and the channel attention mechanism comprises: a global average pooling layer, which outputs a feature map with a size of 1*1*C; a first fully connected layer, which is used for compressing channels to C / 16; a second fully connected layer, which is used for restoring the original channel number.
5. The method of claim 1, wherein, The identification model further comprises an output layer, and the output layer comprises: a 1*1 convolution layer, which is used for reducing a channel dimension to 5 through linear transformation to obtain a 5-channel spatial activation map; global average pooling, which is used for converting the 5-channel spatial activation map into 5 scalar values; a Softmax classification function, which is used for determining a corresponding category according to the 5 scalar values.
6. The method of claim 1, wherein, The method further comprises: dividing a feature map extracted from the image to be identified by the identification model into feature map blocks according to a cache structure of a CPU, and caching the feature map blocks; recombining input data of a convolution kernel and a calculation strategy of the convolution kernel to enable the convolution kernel to sequentially process the feature map blocks; merging processing results of all the feature map blocks, and transmitting the merged processing results to an input layer.
7. The method of claim 6, wherein, The method further comprises: Computing the optimal tile size (T w ,T h ) such that: (T w +2P)(T h +2P) x C x s element ≤ αS cache Wherein, the feature map size is width x height x channel number, represented as W x H x C; the convolution kernel size is K x K; the cache capacity is S cache ; each data element size is s element ; P is the padding size; and a is a safety factor, reserving space for the kernel and intermediate results.
8. A lightweight image recognition system, characterized by, dividing a feature map extracted from the image to be identified by the identification model into feature map blocks according to a cache structure of a CPU, and caching the feature map blocks, comprising: The method comprises: an obtaining module, configured to obtain an image to be identified; an identification module, configured to input the image to be identified into a pre-trained identification model to obtain a category corresponding to the image to be identified; 9. An apparatus, comprising: the identification model comprises four deep separable convolution modules, and a channel attention mechanism is embedded after the second deep separable convolution module and the fourth deep separable convolution module, and the deep separable convolution module adopts an H-Swish activation function. The method comprises: a memory, configured to store a lightweight image identification program; 10. A computer readable storage medium storing a computer program, characterized in that, a processor, configured to implement steps of the lightweight image identification method according to any one of claims 1-7 when the processor executes the lightweight image identification program. The readable storage medium stores a lightweight image identification program, and the lightweight image identification program implements steps of the lightweight image identification method according to any one of claims 1-7 when the lightweight image identification program is executed by a processor.