Drug image recognition method based on artificial intelligence

By using orthogonal light source projection grid pattern and point cloud processing technology in the drug recognition system, combined with knowledge distillation, the problems of low accuracy and high computational complexity of stereoscopic drug recognition are solved, and efficient and rapid identification on edge devices are achieved.

CN120496042APending Publication Date: 2025-08-15苏州聚辰源创科技有限公司
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510619177.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

When existing drug identification systems deal with three-dimensional structure complex drugs, especially under high-speed production lines, the recognition accuracy is low and the calculation complexity is high, making it difficult to deploy on resource-constrained edge devices.

Method used

The structured light source projection grid pattern in three orthogonal directions is adopted, combined with projection contour inversion algorithm and point cloud processing, and the three-dimensional feature extraction and rapid identification of drugs is achieved through compressed vision transformers and knowledge distillation technology.

Benefits of technology

It improves the recognition accuracy of three-dimensional structure complex drugs, reduces the computing complexity and resource requirements, is suitable for edge devices and mobile terminals, and has efficient and fast drug recognition capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496042A_ABST
    Figure CN120496042A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medicine image recognition, and discloses a medicine image recognition method based on artificial intelligence, which comprises the following steps: simultaneously projecting grid patterns through structured light sources in three orthogonal directions, and forming three orthogonal projection shadows of a medicine on a single image sensor; reconstructing a visual shell of the medicine from the three orthogonal projection shadows by using a projection contour inversion algorithm; the reconstructed visual shell is converted into simplified point cloud representation, and key geometric and texture features are reserved through point cloud sampling; performing feature extraction on the simplified point cloud by applying a compressed visual converter to generate feature vectors of the medicine; the knowledge distillation technology is combined, and through matching with a pre-established medicine three-dimensional feature library, rapid and accurate identification of the medicine is realized; according to the method, high-precision, high-speed and low-resource-consumption medicine identification is realized, and the method is suitable for high-speed production line and edge equipment deployment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of drug image recognition, and more specifically, to a drug image recognition method based on artificial intelligence. Background Art

[0002] With the increasing automation of pharmaceutical production and increasingly stringent drug regulatory requirements, rapid and accurate automatic drug identification has become a critical requirement in the pharmaceutical industry. Currently, drug identification primarily relies on two-dimensional image analysis technology, but it faces numerous challenges when processing complex-shaped drugs (such as capsules, pills, and irregular tablets).

[0003] Existing drug recognition systems primarily rely on traditional two-dimensional image processing and simple machine learning algorithms to classify and identify drugs by analyzing features such as color, shape, and surface texture. While these methods perform well for standard tablets, recognition accuracy significantly decreases for drugs with complex three-dimensional structures, especially on high-speed production lines (>10 tablets / second). Furthermore, while traditional three-dimensional reconstruction methods can obtain more complete drug structural information, they are computationally complex and time-consuming, making them difficult to meet real-time recognition requirements.

[0004] The existing technology mainly has the following technical problems: it is difficult to fully capture the three-dimensional structural characteristics of drugs by relying on two-dimensional image analysis; for drugs with large differences in appearance under different viewing angles, the recognition accuracy is significantly affected by the viewing angle; the computational complexity is large and the processing delay is high, making it difficult to deploy applications on resource-constrained edge devices.

[0005] Therefore, a drug image recognition method is needed that can efficiently obtain three-dimensional information of drugs, quickly extract key features, and achieve lightweight deployment on edge devices. Summary of the Invention

[0006] The present invention provides an artificial intelligence-based drug image recognition method to solve the technical problems in related technologies that are difficult to quickly and accurately identify drugs with complex three-dimensional structures and cannot be deployed on resource-constrained devices.

[0007] The present invention provides a drug image recognition method based on artificial intelligence, comprising the following steps: A grid pattern is simultaneously projected by structured light sources in three orthogonal directions, forming three orthogonal projected shadows of the drug on a single image sensor; The visual shell of the drug is reconstructed from three orthogonal projected shadows using the projected contour inversion algorithm; Convert the reconstructed visual hull into a compact point cloud representation, preserving key geometric and texture features through point cloud sampling; Applying a compressed visual transformer to the simplified point cloud for feature extraction to generate a feature vector for the drug; Combined with knowledge distillation technology, the key knowledge of drug identification in the large model is compressed into a lightweight model, and the feature vector is processed to generate a three-dimensional feature representation of the drug. By matching it with the pre-established three-dimensional feature library of drugs, fast and accurate identification of drugs can be achieved.

[0008] In a preferred embodiment, in the step of simultaneously projecting a grid pattern through three structured light sources in orthogonal directions, the three structured light sources are respectively located in the X-axis, Y-axis and Z-axis directions, and each light source projects a grid pattern with a specific periodicity.

[0009] In a preferred embodiment, the step of reconstructing the visual shell of the drug from three orthogonal projection shadows using a projection profile inversion algorithm comprises: The collected images are preprocessed and the drug contour projections in three orthogonal directions are extracted using an adaptive threshold segmentation algorithm; Based on the three extracted orthogonal projection contours, the three-dimensional visual shell of the drug is reconstructed using the voxel space projection inversion algorithm; A horseshoe network is used to refine the initially reconstructed rough visual shell to improve the accuracy of restoring surface details.

[0010] In a preferred embodiment, the step of converting the reconstructed visual shell into a compact point cloud representation comprises: The Mach cube algorithm is used to convert the voxel model into an initial point cloud and extract the surface information of the visual shell; Apply the farthest point sampling algorithm to downsample the initial point cloud, reducing the number of points while preserving the geometric information to the greatest extent; The characteristic points on the drug surface are identified through curvature analysis, and a higher sampling density is used for these areas.

[0011] In a preferred embodiment, the step of applying a compressed visual transformer to the simplified point cloud to perform feature extraction comprises: Normalize the point cloud into the unit sphere and divide it into local regions. Use the K-nearest neighbor algorithm to build local connectivity for each point. Introducing a sparse attention mechanism to process point cloud data and reduce computational complexity; A multi-level feature extraction hierarchy is constructed, which gradually abstracts from point-level features, local features to global features, and uses skip connections to fuse features at different levels.

[0012] In a preferred embodiment, the sparse attention mechanism is implemented in the following way: Perform linear transformation on point cloud features to obtain query, key and value matrices; Calculate the similarity between the query matrix and the key matrix and perform normalization; A sparse mask matrix is introduced to retain only the most relevant k attention connections, reducing computational complexity; Multiply the processed attention weights with the value matrix to obtain the weighted feature representation.

[0013] In a preferred embodiment, the step of achieving rapid and accurate drug identification includes: Construct a larger teacher network and a lightweight student network, and transfer knowledge by minimizing the similarity difference between the output feature vectors of the two networks; Applying principal component analysis and quantization techniques to the feature vectors output by the student network further compresses the feature vectors while retaining key identification information; A three-dimensional feature vector library of drugs is constructed, and the cosine similarity calculation is used to query the similarity between drugs and samples in the library to identify the most matching drug category.

[0014] In a preferred embodiment, the step of combining the knowledge distillation technology includes: Use temperature parameters to adjust the hardness or softness of knowledge transfer; The loss function is calculated by calculating the difference between the probability distribution of the output features of the teacher network and the student network; The student network parameters are optimized according to the loss function so that it can learn the key drug recognition knowledge from the teacher network.

[0015] In a preferred embodiment, the further compression of the feature vector includes three steps of feature dimensionality reduction, scalar quantization and entropy coding, which compresses the feature vector from 1MB to 2KB and retains key identification information.

[0016] In a preferred embodiment, an artificial intelligence-based drug image recognition system is used to perform an artificial intelligence-based drug image recognition method, comprising: A multi-projection acquisition device comprising three orthogonally arranged structured light sources and an image sensor; A visual shell reconstruction module is used to reconstruct the three-dimensional visual shell of the drug from multiple projection shadows; A point cloud processing module for converting visual hulls into compact point cloud representations; Feature extraction module, including a compressed visual transformer, for extracting drug features from point clouds; The lightweight inference module uses knowledge distillation technology to compress models and features to achieve fast and accurate identification of drugs.

[0017] The beneficial effects of the present invention are: By simultaneously projecting a grid pattern using structured light sources in three orthogonal directions, multi-angle projection information of the drug can be acquired in a single acquisition. This significantly improves data acquisition efficiency compared to traditional methods and solves the problem of traditional methods requiring multiple acquisitions from multiple angles. An improved projection contour inversion algorithm and horseshoe network are used for visual shell reconstruction and refinement to accurately capture the three-dimensional structural characteristics of drugs and significantly improve the recognition accuracy of drugs with complex three-dimensional structures. The introduction of compressed visual transformers and sparse self-attention mechanisms greatly reduces the computational complexity of point cloud processing, shortens system acquisition and reconstruction time, and improves detection speed, far exceeding traditional 3D reconstruction methods; Combining knowledge distillation and feature compression technology reduces inference time, enables low-resource deployment on edge devices and mobile terminals, and reduces power consumption; The robustness to changes in viewing angle is improved, and a high recognition accuracy rate is maintained for the same medicine that is damaged, partially occluded, or from different production batches, which significantly improves the adaptability and reliability of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 This is a flow chart of a drug image recognition method based on artificial intelligence of the present invention; Figure 2 is a detailed flow chart of forming three orthogonal projected shadows of a drug on a single image sensor according to the present invention; Figure 3 is a detailed flow chart of the present invention for reconstructing the visual shell of a drug from three orthogonal projection shadows; Figure 4 is a detailed flow chart of the present invention for preserving key geometric and texture features through point cloud sampling; Figure 5 is a detailed flow chart of generating a characteristic vector of a drug according to the present invention; Figure 6 It is a detailed flow chart of the present invention for realizing rapid and accurate identification of medicines. DETAILED DESCRIPTION

[0019] The subject matter described herein will now be discussed with reference to example embodiments. It should be understood that these embodiments are discussed solely to enable those skilled in the art to better understand and implement the subject matter described herein, and that the functions and arrangements of the elements discussed may be varied without departing from the scope of this specification. Various examples may omit, substitute, or add various processes or components as needed. Furthermore, features described in some examples may be combined in other examples.

[0020] At least one embodiment of the present invention discloses a drug image recognition method based on artificial intelligence, such as Figures 1 to 6As shown, the following steps are included: Step 1: Using three orthogonal structured light sources to simultaneously project a grid pattern, three orthogonal shadows of the drug are formed on a single image sensor. The specific steps include: Step 1.1, configuring an orthogonal structured light source array; According to an embodiment of the present application, three mutually orthogonal structured light sources are configured, respectively located in the X-axis, Y-axis and Z-axis directions, and each light source projects a grid pattern with a specific periodicity.

[0021] The spatial frequency of the grid pattern is chosen to be 5 to 10 line pairs / mm to avoid diffraction effects while ensuring resolution.

[0022] In some embodiments, the structured light sources can utilize different spectral characteristics. For example, the X-axis light source can utilize a red LED (wavelength 635-650nm), the Y-axis light source can utilize a green LED (wavelength 520-540nm), and the Z-axis light source can utilize a blue LED (wavelength 450-470nm). This spectral differentiation allows simultaneous capture of projection information from three directions within a single frame, eliminating the need for temporal separation.

[0023] Structured light sources can also be distinguished by using the same wavelength but different modulation frequencies. For example, the X-axis light source is modulated at a frequency of 10kHz, the Y-axis light source is modulated at a frequency of 15kHz, and the Z-axis light source is modulated at a frequency of 20kHz. The projection information in each direction is extracted through frequency domain separation.

[0024] Step 1.2, configure the image sensor system; According to an embodiment of the present application, an image acquisition system is formed using an image sensor with high resolution (≥8MP) and high frame rate (≥100fps) in combination with a wide-angle lens (field of view ≥90°).

[0025] The image sensor is placed in a position where it can receive shadows projected from three directions simultaneously, forming an orthogonal projection geometry with the three light sources.

[0026] Step 1.3, multiple shadows are collected synchronously; As a drug passes through the inspection area, three orthogonal light sources flash synchronously, with exposure times controlled to less than 1ms. The image sensor captures the drug's shadows cast in three directions within the same frame. Furthermore, using color filters or time-division multiplexing, the shadow information from these three directions is separated within a single frame.

[0027] Step 2, using the projection contour inversion algorithm, reconstruct the visual shell of the drug from three orthogonal projection shadows; The specific steps include: Step 2.1, projection contour extraction; The collected images are preprocessed (including denoising and contrast enhancement), and then the adaptive threshold segmentation algorithm is used to extract the drug contour projections in three orthogonal directions. .

[0028] The contour extraction adopts the improved Canny edge detection algorithm, and the threshold is adaptively adjusted to adapt to different lighting conditions.

[0029] Step 2.2, projection profile inversion calculation; According to the three extracted orthogonal projection contours, the voxel space projection inversion algorithm is applied to reconstruct the three-dimensional visual shell of the drug, which is mathematically expressed as: ; in, Represents the drug object to be reconstructed, represents the projection operation in the i-th direction, represents the corresponding projection contour, represents the back-projection operation, Indicates pharmaceutical objects The visual shell reconstruction result of Indicates the intersection operation of the back-projection results in three directions. The index of the projection direction (from 1 to 3, corresponding to the three orthogonal directions X, Y, and Z respectively). This formula describes the mathematical process of calculating the 3D visual hull of the drug by inverting the projection contours in three orthogonal directions.

[0030] For high-speed flowing drugs, a motion compensation mechanism can be introduced to estimate the tiny displacement of the drugs during the acquisition process, make corresponding corrections to the projection contour, and then perform inversion calculations to reduce the impact of motion blur on reconstruction accuracy.

[0031] Step 2.3, voxel model refinement; According to an embodiment of the present application, a horseshoe network (Hourglass Network) is used to refine the initially reconstructed rough visual shell to improve the restoration accuracy of surface details.

[0032] The network input is the preliminary reconstructed voxel model and the original projection image, and the output is a refined high-precision voxel model with a resolution of 0.1mm.

[0033] The horseshoe network is an encoder-decoder neural network that includes a downsampling path, an upsampling path, and a skip connection.

[0034] According to an embodiment of the present application, the network consists of 5 downsampling layers and 5 upsampling layers, and each layer uses a 3D convolution operation.

[0035] The convolution kernel size of the downsampling layer is 3×3×3, the stride is 2, and the number of channels is 32, 64, 128, 256, and 512 respectively; The upsampling layer uses deconvolution operation with a convolution kernel size of 4×4×4, a stride of 2, and the number of channels is 512, 256, 128, 64, and 32, respectively.

[0036] Skip connections are set between the upsampling and downsampling paths to directly connect the feature maps of the corresponding layers to preserve high-resolution detail information.

[0037] The network is trained using voxel supervision, and the loss function includes binary cross entropy loss and surface smoothness regularization: ; in, Represents the total loss function, which is the objective function that needs to be minimized during network training; is the binary cross entropy loss, which is used to measure the difference between the predicted voxel model and the true voxel model; is the surface smoothness regularization term, which is used to constrain the smoothness of the reconstructed surface and reduce noise and irregular bumps; It is a trade-off coefficient used to balance the relative importance of the binary cross entropy loss and the surface smoothness regularization term. Its value is 0.1, which means that the influence of the smoothness constraint is one tenth of the binary cross entropy loss.

[0038] Step 3: Convert the reconstructed visual shell into a compact point cloud representation, preserving key geometric and texture features through point cloud sampling; The specific steps include: Step 3.1, conversion from voxel to point cloud; According to an embodiment of the present application, a Marching Cubes algorithm is used to convert a voxel model into an initial point cloud to extract surface information of the visual shell.

[0039] The converted initial point cloud typically contains 300K to 500K points.

[0040] Step 3.2, point cloud downsampling and simplification; According to an embodiment of the present application, a Farthest Point Sampling (FPS) algorithm is applied to downsample the initial point cloud, thereby reducing the number of points while retaining geometric information to the greatest extent.

[0041] Therefore, the sampled point cloud is controlled between 50K and 100K points to ensure that key geometric structures are captured while reducing the computational burden of subsequent processing.

[0042] Step 3.3, key feature preservation enhancement; Curvature analysis is used to identify characteristic points on the surface of the drug (such as edges, concave-convex structures, embossed marks, etc.), and a higher sampling density is used for these areas to ensure that these detailed information that is crucial for identification is retained during the point cloud simplification process.

[0043] Step 4: Apply the compressed visual transformer to the simplified point cloud to extract features and generate a feature vector of the drug; The specific steps include: Step 4.1, point cloud preprocessing and grouping; According to an embodiment of the present application, the point cloud is normalized to a unit sphere and divided into N local regions. In addition, the K-Nearest Neighbors (KNN) algorithm is used to construct a local connection relationship for each point, forming a graph structure for subsequent attention calculations.

[0044] Point cloud normalization consists of two steps: center alignment and scale normalization. First, the geometric center of the point cloud is calculated and moved to the origin, and then the point cloud is scaled to fit within the unit sphere (the distance from the farthest point to the origin is 1).

[0045] The local area division adopts the spherical uniform sampling method to evenly divide the unit sphere into 32 areas. The points in each area are marked according to the distance from the center point for subsequent local feature extraction.

[0046] Step 4.2, sparse self-attention calculation; The method provided in this application introduces a sparse attention algorithm to process point cloud data, reducing computational complexity: ; in, Represents the query matrix, which is obtained by linear transformation of point cloud features and is used to represent the features of the point that currently needs attention; Represents the key matrix, which is also obtained by linear transformation of point cloud features and is used to calculate the correlation with the query matrix; The representation value matrix is obtained by another linear transformation of the point cloud features and contains the information that actually needs to be aggregated; Represents the feature dimension, which is the dimension size of the feature vector in the attention mechanism and is used to normalize the dot product operation; Represents a sparse mask matrix, which is a binary matrix that only contains The most relevant attention connection position is 1, and the rest of the positions are 0 to reduce computational complexity; Indicates the number of attention connections retained for each point, which is much smaller than the total number of points ; Indicates the total number of points in the point cloud; represents the Hadamard product (element-wise multiplication), which is used to apply the mask matrix to the attention score; represents the softmax normalization function, which converts the attention score into a probability distribution; represents the transpose multiplication of the query matrix and the key matrix to calculate the correlation score between the points; represents a scaled dot product operation that stabilizes the gradient by dividing by the square root of the feature dimension.

[0047] According to an embodiment of the present application, the architecture of the compressed visual transformer includes 4 hidden layers, each layer consists of a multi-head sparse self-attention module and a feedforward network. The number of multi-head attention heads is 8, and the feature dimension of each head is 64. Sparse mask matrix The construction is based on the k-nearest neighbor graph, where each point only establishes attention connections with its k nearest neighbors (k=16), which reduces the computational complexity from Reduce to .

[0048] When processing pharmaceutical point cloud data, the sparse attention algorithm pays special attention to characteristic areas on the drug surface, such as tiny embossing and concave-convex textures. By analyzing the local curvature distribution of the point cloud, it assigns higher feature dimension weights to high-curvature areas (such as edges and embossed text), enhancing sensitivity to key recognition features.

[0049] In some implementations, an adaptive sparsity strategy can be used to dynamically adjust the sparsity based on the local feature distribution of the point cloud, retaining more attention connections for feature-rich areas and using higher sparsity for flat areas. Specifically, the adaptive sparsity parameter can be defined as: ; in, Indicates a point The number of adaptive sparse connections determines how many other points the point establishes attention connections with; is the basic sparse connectivity number, which indicates the number of neighbors connected to each point by default; Yes The normalized curvature value of the area is in the range of 0 to 1. The larger the value, the higher the curvature of the area, that is, the richer the surface features; is the adjustment coefficient that controls the degree of adaptation. A larger value indicates a more pronounced effect of curvature on the number of connections. This formula implements a mechanism for automatically adjusting the density of attention connections based on local geometric features.

[0050] In order to further improve computational efficiency, a progressive precision control mechanism can be introduced in attention calculations, using lower-precision floating-point representations (such as FP16) for shallow layers of the network and higher-precision (such as FP32) for deep layers, thereby further reducing computational resource requirements while maintaining recognition accuracy.

[0051] Step 4.3, multi-scale feature fusion; According to the embodiments of this application, a multi-level feature extraction hierarchy is constructed, gradually abstracting from point-level features, local features, and global features. Skip connections are used to fuse features at different levels, ensuring that both the micro-texture and macro-shape information of the drug are captured simultaneously.

[0052] The multi-scale feature fusion structure is designed in a pyramid form, including four levels: point-level features (dimension 64), local area features (dimension 128), medium area features (dimension 256) and global features (dimension 512).

[0053] Features at each level are fused through an adaptive weight mechanism: ; in, Represents the final feature vector after fusion, which is a weighted combination of multi-scale features; Represents the features of the i-th level, corresponding to feature representations at four different scales: point-level features, local area features, medium area features, and global features; is the learned weight coefficient, which determines the importance of the i-th level feature in the final fusion feature. These weights are automatically learned through network training; Indicates that the sum of all weight coefficients is equal to 1, ensuring weight normalization and preventing feature magnitude changes.

[0054] This multi-scale fusion architecture can simultaneously capture both subtle textures (such as embossed characters and micro-indentations) and overall shapes (such as the curvature of capsules and the edge contours of tablets) of a drug, significantly improving the ability to distinguish between drugs with similar appearances. In capsule and tablet classification tasks, multi-scale feature fusion improves recognition accuracy by approximately 12% compared to a single-scale representation.

[0055] Step 5: Combining knowledge distillation technology, the key knowledge for drug identification in the large model is compressed into a lightweight model. The feature vectors are processed to generate a 3D feature representation of the drug. By matching it with a pre-established 3D feature library of drugs, rapid and accurate drug identification is achieved. The specific steps include: Step 5.1, teacher-student knowledge distillation; According to the embodiment of the present application, a larger teacher network (about 100MB) and a lightweight student network (about 1MB) are constructed. The teacher network has a more complex structure and higher expressive power, while the student network adopts a more efficient network architecture.

[0056] Knowledge transfer is performed by minimizing the KL divergence between the output feature vectors of two networks: ; in, represents the knowledge distillation loss function, which is used to measure the effect of the student network learning from the teacher network; and are the output features of the teacher and student networks, respectively, which contain the representation information of the input data of each network; It is the softmax function, which converts the output features into probability distribution; is the temperature parameter, which controls the hardness or softness of knowledge transfer. A higher temperature value will produce a smoother probability distribution, which is conducive to transferring subtle knowledge in the teacher network. Represents the Kullback-Leibler divergence, which is used to measure the difference between two probability distributions. The smaller the value, the better the student network imitates the behavior of the teacher network. It is the square of the temperature parameter, which serves as a scaling factor for the KL divergence and is used to balance the gradient size of the loss function.

[0057] The teacher network uses a complete compressed visual transformer architecture, consisting of eight hidden layers, each with a feature dimension of 512, and a total parameter size of approximately 100MB. The student network uses a lightweight architecture, consisting of three hidden layers with a feature dimension of 128, and introduces depthwise separable convolutions instead of standard convolutions, with a total parameter size of approximately 1MB.

[0058] In the knowledge distillation process, the temperature parameter Set to 3.0 to generate a softer probability distribution, which is conducive to knowledge transfer.

[0059] Knowledge distillation adopts a three-stage training strategy: in the first stage, only the teacher network is trained, in the second stage, the teacher network parameters are fixed for knowledge distillation, and in the third stage, both the teacher and student networks are fine-tuned simultaneously.

[0060] Step 5.2, feature vector compression encoding; According to an embodiment of the present application, principal component analysis (PCA) and quantization techniques are applied to the feature vectors output by the student network, further compressing the feature vectors from 1MB to 2KB while retaining key identification information. Furthermore, the specific steps include feature dimensionality reduction, scalar quantization, and entropy coding.

[0061] Feature dimensionality reduction uses principal component analysis to reduce the original 128-dimensional feature vector to 32 dimensions, preserving approximately 95% of the variance information. Scalar quantization quantizes each 32-bit floating-point eigenvalue into an 8-bit integer, using a nonlinear quantization strategy to employ finer quantization intervals for densely populated areas. Entropy coding employs Huffman coding, assigning codes of varying lengths based on the frequency distribution of eigenvalues, further reducing storage requirements.

[0062] While reducing the size of feature vectors, compression coding learns the optimal quantization parameters and coding dictionary through offline training, thereby retaining the key information required for drug identification to the greatest extent possible.

[0063] In some embodiments, a drug category-aware compression strategy may be employed, using specially optimized compression parameters for different categories of drugs (such as tablets, capsules, and pills), to further improve the recognition accuracy of specific categories of drugs.

[0064] For example, for compressed tablets with complex surface textures, more high-frequency features related to the texture can be retained during the compression process; for capsule-type drugs with significant shape changes, more emphasis is placed on retaining low-frequency features related to the geometric shape.

[0065] The compression process can also use an iterative quantization strategy, which optimizes the quantization parameters through multiple iterations to minimize the quantization error. In each iteration, the residual is calculated based on the result of the previous quantization, and then the residual is requantized. Finally, the multiple quantization results are combined to form a more accurate quantization representation.

[0066] Step 5.3, feature matching and drug identification; The method provided in this application constructs a three-dimensional drug feature vector library, which contains standard feature representations of various types of drugs. It should be noted that cosine similarity is used to calculate the similarity between the query drug and the samples in the library to identify the most matching drug category: ; in, is the feature vector of the drug to be identified, which represents the compressed feature representation extracted from the point cloud of the drug to be identified; It is the feature vector of a drug in the library, which represents the standard drug feature representation pre-stored in the database; Represents the dot product of two feature vectors and calculates the raw similarity between them; Represents a vector The L2 norm (Euclidean norm) of the vector is calculated as the square root of the sum of the squares of the elements of the vector and is used to normalize the length of the vector; Represents a vector The L2 norm of , also used for normalization; Represents the cosine similarity calculation formula. The result range is [-1, 1]. The closer the value is to 1, the more similar the two drug features are. It is used for final drug matching and identification.

[0067] Application examples of this embodiment: According to the examples of this application, this implementation was put to practical use on a tablet quality inspection line at a pharmaceutical manufacturer. This line processes approximately 50,000 tablets per hour, including standard round tablets, irregularly shaped tablets, and capsules. Accurate identification and quality inspection of the tablets are required at a high flow rate (10 to 20 tablets per second).

[0068] Implementation process example: According to an embodiment of the present application, the system implementation includes the following specific processes: Hardware platform construction: Three 60WLED structured light sources (wavelengths of 450nm, 550nm, and 650nm) are used, equipped with a 12MP industrial camera (frame rate 120fps), and the processing platform is an edge computing device equipped with an RTX3060 graphics card.

[0069] Multiple Shadow Projection Capture: As a tablet passes through the inspection area, three LED light sources flash synchronously (exposure time 0.8ms), forming orthogonal shadow projections in red, green, and blue channels on the image sensor. The captured image has a resolution of 4096 × 3072 pixels, with each pixel corresponding to approximately 0.05 mm in real space.

[0070] Visual hull reconstruction: The acquired three-channel images were preprocessed, and contours were extracted using a modified Canny algorithm (low threshold 50, high threshold 150). The initial visual hull was reconstructed using a projective inversion algorithm. A pretrained horseshoe network was then used to refine the voxel model. The refined voxel resolution was 256×256×256, corresponding to a practical spatial resolution of 0.1 mm.

[0071] Point cloud processing and feature extraction: The voxel model is converted to an initial point cloud (approximately 400,000 points) using the Mach Cube algorithm. This is then reduced to 80,000 points using the Farthest Point Sampling algorithm. A higher density of sampling points is retained for areas of high curvature, such as embossed text and edges on the tablet. A compressed visual transformer processes the point cloud to generate a 512-dimensional feature vector.

[0072] Lightweighting and Recognition: Knowledge distillation is used to compress the feature extraction network to 1MB. PCA is then used to reduce the feature vector to 32 dimensions. 8-bit quantization and entropy coding are then used to compress the network to a final size of 2KB. The network is then compared with a pre-established drug feature library, and cosine similarity is used to calculate the match. A threshold of 0.85 is set, and a match above this threshold is considered a successful recognition.

[0073] In some implementations, the system can also integrate quality inspection capabilities, analyzing the reconstructed 3D model to detect defects such as gaps, cracks, and deformations. This is achieved by registering and comparing the reconstructed 3D model of the drug with a standard model, calculating surface deviations, and determining if a quality issue exists when the deviation exceeds a preset threshold.

[0074] The system can also be deployed on mobile devices. By simplifying the light source system (such as using the mobile phone flash and ambient light as light sources in different directions), using the mobile phone camera to capture multi-angle images, and running a lightweight recognition algorithm on the mobile phone processor, it can achieve portable identification of drugs. It is suitable for scenarios such as pharmacies, hospitals, and patients' homes.

[0075] Technical effect verification: The experimental results show that this method has significant advantages over traditional methods, as shown in Table 1 and Table 2: Table 1: Comparison of recognition accuracy;

[0076] Table 2: Comparison of processing performance;

[0077] After three months of continuous testing on an actual production line, this method maintained a stable recognition accuracy rate of over 93%, and its processing speed met the production requirement of 20 tablets per second. Model compression enabled the system to run stably on edge computing devices, significantly reducing deployment and maintenance costs. Furthermore, even for damaged or partially obscured tablets, this method maintained an accuracy rate of over 85%, surpassing the 65% achieved by traditional methods.

[0078] The above describes an embodiment of the present invention, but this embodiment is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Ordinary technicians in this field can also make more forms of equivalent embodiments based on the inspiration of this embodiment, all of which are protected by this embodiment.

Claims

1. A drug image recognition method based on artificial intelligence, characterized in that: The following steps are involved: A grid pattern is simultaneously projected by structured light sources in three orthogonal directions, forming three orthogonal projected shadows of the drug on a single image sensor; The visual shell of the drug is reconstructed from three orthogonal projected shadows using the projected contour inversion algorithm; Convert the reconstructed visual hull into a compact point cloud representation, preserving key geometric and texture features through point cloud sampling; Applying a compressed visual transformer to the simplified point cloud for feature extraction to generate a feature vector for the drug; Combined with knowledge distillation technology, the key knowledge of drug identification in the large model is compressed into a lightweight model, and the feature vector is processed to generate a three-dimensional feature representation of the drug. By matching it with the pre-established three-dimensional feature library of drugs, fast and accurate identification of drugs can be achieved.

2. The method for drug image recognition based on artificial intelligence according to claim 1, characterized in that: In the step of simultaneously projecting a grid pattern through three structured light sources in orthogonal directions, the three structured light sources are respectively located in the X-axis, Y-axis and Z-axis directions, and each light source projects a grid pattern with a specific periodicity.

3. The method for drug image recognition based on artificial intelligence according to claim 1, characterized in that: The step of reconstructing the visual shell of the medicine from three orthogonal projection shadows using a projection contour inversion algorithm comprises: The collected images are preprocessed and the drug contour projections in three orthogonal directions are extracted using an adaptive threshold segmentation algorithm; Based on the three extracted orthogonal projection contours, the three-dimensional visual shell of the drug is reconstructed using the voxel space projection inversion algorithm; A horseshoe network is used to refine the initially reconstructed rough visual shell to improve the accuracy of restoring surface details.

4. The method for drug image recognition based on artificial intelligence according to claim 1, characterized in that: The step of converting the reconstructed visual shell into a compact point cloud representation comprises: The Mach cube algorithm is used to convert the voxel model into an initial point cloud and extract the surface information of the visual shell; Apply the farthest point sampling algorithm to downsample the initial point cloud, reducing the number of points while preserving the geometric information to the greatest extent; The characteristic points on the drug surface are identified through curvature analysis, and a higher sampling density is used for these areas.

5. The method for drug image recognition based on artificial intelligence according to claim 1, characterized in that: The step of applying a compressed visual transformer to the simplified point cloud to perform feature extraction comprises: Normalize the point cloud into the unit sphere and divide it into local regions. Use the K-nearest neighbor algorithm to build local connectivity for each point. Introducing a sparse attention mechanism to process point cloud data and reduce computational complexity; A multi-level feature extraction hierarchy is constructed, which gradually abstracts from point-level features, local features to global features, and uses skip connections to fuse features at different levels.

6. The method for drug image recognition based on artificial intelligence according to claim 5, characterized in that: The sparse attention mechanism is implemented in the following way: Perform linear transformation on point cloud features to obtain query, key and value matrices; Calculate the similarity between the query matrix and the key matrix and perform normalization; A sparse mask matrix is introduced to retain only the most relevant k attention connections, reducing computational complexity; Multiply the processed attention weights with the value matrix to obtain the weighted feature representation.

7. The method for drug image recognition based on artificial intelligence according to claim 1, characterized in that: The steps of achieving rapid and accurate drug identification include: Construct a larger teacher network and a lightweight student network, and transfer knowledge by minimizing the similarity difference between the output feature vectors of the two networks; Applying principal component analysis and quantization techniques to the feature vectors output by the student network further compresses the feature vectors while retaining key identification information; A three-dimensional feature vector library of drugs is constructed, and the cosine similarity calculation is used to query the similarity between drugs and samples in the library to identify the most matching drug category.

8. The method for drug image recognition based on artificial intelligence according to claim 1, characterized in that: The steps of combining the knowledge distillation technology include: Use temperature parameters to adjust the hardness or softness of knowledge transfer; The loss function is calculated by calculating the difference between the probability distribution of the output features of the teacher network and the student network; The student network parameters are optimized according to the loss function so that it can learn the key drug recognition knowledge from the teacher network.

9. The method for drug image recognition based on artificial intelligence according to claim 7, characterized in that: The further compression of the feature vector includes three steps: feature dimensionality reduction, scalar quantization and entropy coding, which compresses the feature vector from 1MB to 2KB and retains key identification information.

10. An artificial intelligence-based drug image recognition system, used to execute the artificial intelligence-based drug image recognition method according to any one of claims 1 to 9, characterized in that: include: A multi-projection acquisition device comprising three orthogonally arranged structured light sources and an image sensor; A visual shell reconstruction module is used to reconstruct the three-dimensional visual shell of the drug from multiple projection shadows; A point cloud processing module for converting visual hulls into compact point cloud representations; Feature extraction module, including a compressed visual transformer, for extracting drug features from point clouds; The lightweight inference module uses knowledge distillation technology to compress models and features to achieve fast and accurate identification of drugs.

Citation Information

Cited By

  • Bionic robot vision collaboration method based on multiple vision modules and robot vision device

    CN120839803A

  • Bionic robot vision coordination method based on multi-vision module and robot vision device

    CN120839803B