Non-destructive Counting Method of Tablets in Medicine Bottles Based on Acoustic Analysis and Machine Learning Model
Through acoustic analysis and machine learning models, combined with principal component analysis and convolutional neural network, the accurate judgment of the number of remaining drugs in the drug bottle is achieved, solving the problem that it is difficult for blind people to accurately judge the drug dosage, and avoiding drug contamination.
Patent Information
- Application Number
- CN202411544286.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-31
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-10-31
AI Technical Summary
It is difficult for blind people to accurately judge the amount of drugs remaining in the bottle, and the existing technology cannot achieve accurate judgment, and manual counting is easy to contaminate drugs.
The lossless counting method of tablets in the medicine bottle based on acoustic analysis and machine learning models is adopted. The sound wave spectrum when the drug is shaken is collected, the characteristics are extracted using spectral subtraction and linear weighting methods, and machine learning training is carried out in combination with principal component analysis and convolutional neural network to accurately judge the number of drugs.
It realizes an accurate judgment of the amount of remaining drugs in the medicine bottle, solves the problem of difficulty in estimating the amount of drugs when taking medicine by blind people, and avoids drug contamination.
Smart Images

Figure CN119397240B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for non-destructively counting tablets in a medicine bottle based on acoustic analysis and a machine learning model, belonging to the fields of acoustic processing and machine learning. Background Art
[0002] For people who need to take medicine for a long time, they usually need to prepare medicine in advance. Normal people can rely on visual effects to estimate the remaining amount of medicine in the medicine bottle. However, for blind people, they need to open the medicine box and pour out the medicine, and then judge the remaining amount of medicine in the medicine box by touching one by one. But this will cause many inconveniences. First, it may contaminate the medicine, resulting in the loss of drug efficacy or bacterial infection; second, it is also difficult for blind people to put the poured-out medicine back into the bottle through the bottle mouth, which is time-consuming and laborious.
[0003] Publication No. CN202740391U discloses a voice medicine box for blind people. By setting a voice chip, the voice chip stores information such as the name of the medicine, applicable symptoms, and taking method. When a blind person takes medicine, they touch the touch sensor on the box body or the box cover with their hands. The touch sensor sends a signal to the voice chip, and the voice chip transmits the voice signal to the speaker. The speaker emits the sound of information such as the name of the medicine, applicable symptoms, and taking method. However, this voice medicine box for blind people does not announce the quantity inside the medicine box.
[0004] Currently, there is no device in the market that can know the remaining amount of medicine in a medicine box / bottle. After research, most blind people can only roughly judge the remaining amount of medicine based on the sound emitted by shaking the medicine bottle at present. However, this method cannot achieve accurate judgment. And for many diseases, continuous and uninterrupted medication is required. Therefore, it is necessary to accurately know the remaining amount of medicine in order to prepare medicine in time. Summary of the Invention
[0005] In view of the above defects of the prior art, the present invention proposes a method and device for non-destructively counting tablets in a medicine box based on acoustic analysis and a machine learning model. By analyzing samples of different categories and different quantities, multiple machine learning models are built for the same sample. The standard deviation between the calculation instance and different models can be used to accurately judge the quantity of the medicine to be measured. The present invention can be used for blind people to judge the remaining amount of medicine in a medicine bottle.
[0006] The method for non-destructively counting tablets in a medicine box based on acoustic analysis and a machine learning model proposed by the present invention includes:
[0007] S1: Collect the sound wave spectrum and noise spectrum of different types and different quantities of medicine when shaken in the medicine box;
[0008] S2: Obtain the effective sound spectrum by using spectral subtraction;
[0009] S3: Combine the time-domain features and frequency-domain features of the voice signal based on the linear weighting method. Among them, the time-domain features include: Short-Time Energy Spectrum (STES), Short-Time Zero-Crossing Rate (STZCR); the frequency-domain features include: Short-Time Fourier Transform (STFT), Mel-Frequency Cepstral Coefficients (MFCC), Linear Prediction Cepstral Coefficients (LPCC).
[0010] The linear weighting formula is as follows, where the coefficients are obtained through training to get the optimal solution.
[0011] MixFeature = α * STES + β * STZCR + γ * STFT + δ * MFCC + ω * LPCC
[0012] S4: Reduce the dimension of the combined feature parameters based on the Principal Component Analysis (PCA) method, and process them into a two-dimensional image format data of 128 * 128 * 3, which is convenient to be fed into the Convolutional Neural Network (CNN) model for processing;
[0013] S5: Use the Convolutional Neural Network (CNN) to perform machine learning training on the input data.
[0014] The data is first preprocessed, and the file name is named with i - j for marking, where i is the number of drugs in the medicine bottle / medicine box, and j is the number of experiments.
[0015] In the CNN model, the data with the same i is grouped together. The input of CNN is a wavelet scale map image with a size of 128 * 128, and it has three color channels (RGB). Each image is represented as a 3D tensor with a shape of (128, 128, 3), where the first two dimensions represent the spatial size of the image, and the third dimension represents the color channel.
[0016] CNN first applies a series of convolutional operations to extract features from the input image. Each convolutional layer consists of a set of learnable filters (or kernels), which are convolved with the input image to generate feature maps. The filter is a small matrix (e.g., 3 * 3), which slides on the input image and calculates the dot product between the filter values and the corresponding image values.
[0017] To enable machine learning to learn more complex models, we introduce an activation function ReLU, and the function is as follows:
[0018] ReLU(x) = max(0, x)
[0019] The ReLU function introduces non-linearity into the model, enabling it to learn more complex patterns. In the data, it ensures that only positive eigenvalue features are passed to the next layer, while setting negative eigenvalue features to zero.
[0020] After convolution and activation operations, a pooling layer is applied to reduce the size of the spatial feature map. In the solution of the present invention, the pooling operation is max pooling, which selects the maximum value within a small window (e.g., 2×2) of the feature map. This downsampling operation reduces the computational complexity and provides translational invariance for small displacements in the input image. The max pooling expression is as follows:
[0021] P(x,y) = max{I(x+i,y+j)|i,j∈[0,p-1]}
[0022] where P(x,y) is the pooled feature map, and p×p is the size of the pooling window. After several convolutional layers and pooling layers, the obtained feature map is flattened into a one-dimensional vector and passed to the fully connected (dense) layer. At this stage, each neuron in the fully connected layer is connected to all neurons in the previous layer, enabling the network to learn the non-linear combinations of the high-level features extracted by the convolutional layers. The last layer in the CNN is the fully connected layer with 11 neurons, corresponding to 11 categories. The softmax activation function is applied to the output of this layer to generate class probabilities:
[0023]
[0024] where zi is the input of the i-th neuron, and K is the total number of categories (11 in this example). The softmax function ensures that the output is a probability distribution, and the sum of all probabilities is equal to 1. The category with the highest probability is selected as the predicted category of the input image. After being optimized by the Adam optimizer, this function can obtain the final model.
[0025] S6: Adjust the initial parameters and repeat S3 - S5 multiple times. We can obtain different corresponding models for the sounds and quantities of different types of drugs when shaken in the medicine bottle;
[0026] S7: In the prediction stage, when we need to explore the remaining quantity of drugs in a medicine bottle, first shake it by hand and collect the sound signal through the microphone of our lossless counting device;
[0027] S8: Use spectral subtraction to obtain pure sound data;
[0028] S9: Input the data we obtained into all the models;
[0029] S10: Calculate the standard deviation among different quantities of the same type of drug, and select the group with the smallest standard deviation. The quantity corresponding to this group is the required value;
[0030] S11: Inform the blind person of the remaining quantity of drugs in the medicine bottle through the speaker of our device.
[0031] The beneficial effects of the present invention are:
[0032] By proposing a non-destructive counting method and device based on acoustic analysis and machine learning models, combining traditional acoustics and convolutional neural networks, it can effectively judge the remaining amount of medicine in the medicine bottle when the remaining amount of medicine is small by using machine learning methods and a pre-prepared model library, and achieve relatively accurate judgment. It solves the problem that it is difficult for blind people to estimate the remaining amount of medicine in the medicine bottle when taking medicine, and at the same time avoids the problem of secondary pollution of drugs caused by touching the drugs for counting by hand. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0034] Figure 1 It is a schematic diagram of the composition of a non-destructive counting device for tablets in a medicine bottle based on acoustic analysis and machine learning models proposed by the present invention.
[0035] Figure 2 It is a flowchart of a non-destructive counting method provided by an embodiment of the present invention.
[0036] Figure 3 It is a schematic diagram of the structure of a convolutional neural network provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0037] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will further describe the embodiments of the present invention in detail with reference to the drawings.
[0038] Embodiment 1:
[0039] This embodiment provides a non-destructive counting method for tablets in a medicine bottle based on acoustic analysis and machine learning models. The method is implemented based on a non-destructive counting device, as Figure 1 shown. The non-destructive technology device includes a sound collection module composed of a microphone for collecting the sound emitted by the collision of drugs in the bottle during shaking, a signal analysis module composed of a calculation chip for combining the sound signal to judge the remaining amount of drugs in the medicine bottle / medicine box, and a result playback module composed of a speaker for notifying the blind of the remaining amount of drugs in the medicine bottle / medicine box after obtaining the result.
[0040] As Figure 2 shown; the non-destructive counting method for tablets in a medicine bottle based on acoustic analysis and machine learning models includes:
[0041] S1: Collect the sound wave spectrum and noise spectrum of the sound emitted when different types and different quantities of drugs are shaken in the medicine bottle / medicine box;
[0042] S2: Obtain the effective sound spectrum using spectral subtraction;
[0043] S3: Based on the linear weighting method, perform parameter combination on the time-domain features and frequency-domain features of the effective sound spectrum to obtain the combined feature (MixFeature).
[0044] Among them, the time-domain features include: short-time energy spectrum (STES), short-time zero-crossing rate (STZCR); the frequency-domain features include: short-time Fourier transform (STFT), Mel-frequency cepstral coefficients (MFCC), linear prediction cepstral coefficients (LPCC).
[0045] The linear weighting formula is as follows, where the coefficients of each item obtain the optimal solution through training.
[0046] MixFeature = α * STES + β * STZCR + γ * STFT + δ * MFCC + ω * LPCC
[0047] S4: Perform dimensionality reduction on the combined feature based on the principal component analysis method (PCA), and process it into a two-dimensional image format data of 128 * 128 * 3, which is convenient to be fed into the convolutional neural network (CNN) model for processing.
[0048] S5: Use the convolutional neural network (CNN) model to perform machine learning training on the input data.
[0049] The data is first preprocessed, and the file name is named with i - j for marking, where i is the number of drugs in the medicine bottle / medicine box, and j is the number of experiments. In the CNN model, the data with the same i is grouped. The input of the CNN is an image of size 128 * 128, with three color channels (RGB). Each image is represented as a 3D tensor with the shape of (128, 128, 3), where the first two dimensions represent the spatial size of the image, and the third dimension represents the color channel.
[0050] The CNN model includes three convolutional pooling layers, a flattening layer, a fully connected layer, and an output layer. As Figure 3 shown, the CNN model first applies a series of convolutional operations to extract features from the input image. Each convolutional layer consists of a set of learnable filters (or kernels), which are convolved with the input image to generate feature maps. The filter is a small matrix (such as 3 * 3), which slides on the input image and calculates the dot product between the filter values and the corresponding image values.
[0051] In order to enable machine learning to learn more complex models, the present invention introduces an activation function ReLU, and the function is as follows:
[0052] ReLU(x) = max(0, x).
[0053] The ReLU function introduces non-linearity into the model, enabling it to learn more complex patterns. In the data, it ensures that only positive eigenvalues are passed to the next layer, while setting negative eigenvalues to zero. After convolutional and activation operations, a pooling layer is applied to reduce the size of the spatial feature map. The pooling operation adopted in the present invention is max pooling, which selects the maximum value within a small window (e.g., 2×2) of the feature map. This downsampling operation reduces the computational complexity and provides translational invariance for small displacements in the input image. The max pooling expression is as follows:
[0054] P(x,y)=max{I(x+i,y+j)|i,j∈[0,p-1]}
[0055] Where P(x,y) is the pooled feature map and p×p is the size of the pooling window. After several convolutional and pooling layers, the resulting feature map is flattened into a one-dimensional vector and passed to the fully connected (dense) layer. At this stage, each neuron in the fully connected layer is connected to all neurons in the previous layer, enabling the network to learn the non-linear combinations of the high-level features extracted by the convolutional layers. The last layer in the CNN is the fully connected layer with 11 neurons, corresponding to 11 neurons.
[0056] The softmax activation function is applied to the output of this layer to generate class probabilities:
[0057]
[0058] Where zi is the input of the i-th neuron and K is the total number of classes (11 in this example). The softmax function ensures that the output is a probability distribution and the sum of all probabilities equals 1. The class with the highest probability is selected as the predicted class of the input acoustic wave spectrum. The function is optimized by the Adam optimizer to obtain the final model.
[0059] S6: Adjust the initial parameters and repeat S3 - S5 multiple times to obtain multiple corresponding models of different sounds and quantities when different types of drugs are shaken in medicine bottles / medicine boxes;
[0060] S7: In the prediction stage, collect sound signals through the microphone of the lossless counting device;
[0061] S8: Use spectral subtraction to obtain pure sound data;
[0062] S9: Input the obtained pure sound data into all models;
[0063] S10: Calculate the standard deviation among different quantities of the same type of drug, and select the group with the smallest standard deviation. The quantity corresponding to this group is the one sought;
[0064] S11: Inform the blind of the remaining amount of medicine in the medicine bottle / medicine box through the speaker of the non-destructive counting device.
[0065] To verify the effectiveness of the method of the present invention, 22 groups of data obtained by hand shaking using the above method were used. The amount of medicine in the bottle was a random value between 1 and 10, generated by a random integer generator. The non-destructive counting device provided by the present invention was used to judge the amount of medicine in the bottle, and then the box was opened for verification:
[0066] Definition: The results are shown in Table 1, where predicted label is the model prediction label, true label is the true label, and custom loss is the custom loss function.
[0067] Table 1 Results of true label, model prediction label and custom loss function
[0068]
[0069] According to Table 1, the counting accuracy of the amount of medicine in the medicine bottle obtained by using the non-destructive counting method provided by the present invention is as high as 95.45%, and the standard deviation of custom loss is as low as 0.0298, which shows that the non-destructive counting device provided by the present invention has good stability.
[0070] In summary, the method provided in this example combines traditional acoustics and convolutional neural networks, and can effectively judge the remaining amount of medicine when the remaining amount of medicine in the medicine bottle is small by using machine learning methods and a pre-prepared model library. Our device has also effectively achieved our goal and realized relatively accurate judgment.
[0071] Some steps in the embodiments of the present invention can be implemented by software, and the corresponding software program can be stored in a readable storage medium, such as an optical disc or a hard disk, etc.
[0072] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A non-destructive counting method for tablets in medicine bottles based on acoustic analysis and machine learning models, characterized in that: The method comprises: Step S1: collecting sound wave spectra and noise spectra of the sounds made when different types and quantities of drugs are shaken in medicine bottles / medicine boxes; Step S2: Obtaining an effective sound spectrum by spectral subtraction; Step S3: performing parameter combination on the time domain features and frequency domain features of the effective sound spectrum based on a linear weighting method; Step S4: Reduce the dimension of the characteristic parameter combination based on the principal component analysis method and process it into a 128*128*3 input data; Step S5: Use the convolutional neural network model to perform machine learning training on the input data to obtain a trained model so that the amount of medicine can be judged based on the sound made when the medicine is shaken in the medicine bottle / medicine box.
2. The method according to claim 1, characterized in that The time domain features in step S2 include: short-time energy spectrum and short-time zero-crossing rate; the frequency domain features include: short-time Fourier transform, Mel cepstral coefficients and linear prediction complex cepstral coefficients.
3. The method according to claim 2, characterized in that The convolutional neural network model includes three layers of convolutional pooling layers, flattening layers, fully connected layers and output layers; in the convolutional pooling layer, a series of convolution operations are first applied to extract features from the input image, and each convolutional layer consists of a group of learnable filters; secondly, the introduced activation function ReLU is used to ensure that only positive eigenvalues are passed to the next layer, while setting the negative eigenvalues to zero; after the convolution and activation operations, the pooling layer is applied to reduce the size of the spatial feature map.
4. The method according to claim 3, characterized in that The flattening layer is used to flatten the obtained feature map into a one-dimensional vector and pass it to the fully connected layer. Each neuron in the fully connected layer is connected to all neurons in the previous layer, so that the network can learn the high-level features extracted by the nonlinear combination convolution layer; the output layer uses the softmax activation function to obtain the generated class probability and selects the class with the highest probability as the output result.
5. The method according to claim 4, characterized in that Before the step S5 uses the convolutional neural network model to perform machine learning training on the input data, the step also includes: The input data is preprocessed and classified, wherein data with the same number are grouped together.
6. A non-destructive counting device for tablets in medicine bottles based on acoustic analysis and machine learning models, characterized in that: The device includes a sound collection module for collecting the sound produced by the collision of medicines in the medicine bottle / medicine box, a signal analysis module composed of a computing chip for judging the amount of remaining medicines in the medicine bottle / medicine box in combination with the sound signal, and a result playback module for notifying the blind person of the amount of remaining medicines in the medicine bottle / medicine box after obtaining the result; the computing chip includes the execution steps of the method described in any one of steps 1-5 above.
7. The non-destructive counting device according to claim 6, characterized in that: The sound collection module is implemented by a microphone.
8. The non-destructive counting device according to claim 6, characterized in that: The result playback module is implemented using a loudspeaker.
Citation Information
Patent Citations
Voice pill case for blind people
CN202740391U
Method for identifying abnormal sound signal based on convolutional neural network
CN109473120A
Speaker counting method and device based on deep learning, equipment and storage medium
CN113903328A