Pharmacy receipt processing method and device, electronic equipment and storage medium
By OCR identification and data mining of pharmacy receipts, the problem of low efficiency in pharmacy inventory management is solved, efficient and accurate data collection and analysis are achieved, and inventory management is optimized.
Patent Information
- Application Number
- CN202510133056.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-02-06
AI Technical Summary
The commodity transactions of traditional Chinese medicine stores rely on manual statistics, resulting in inaccurate or inefficient data statistics, and ineffective use of the data value in receipts, resulting in inefficient inventory management.
By obtaining the ticket image set and OCR recognition, drug sales information is extracted, and data mining is carried out to obtain the mining results required for drug management.
It realizes efficient and automated processing of pharmacy receipts, improves the speed and accuracy of data collection and analysis, optimizes inventory management, and avoids the problems of insufficient or overstock.
Smart Images

Figure CN120047957A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and particularly to a method, device, electronic device, and storage medium for processing pharmacy receipts. Background Art
[0002] With the development of information technology, image recognition technology and big data analysis are increasingly widely used in the field of business intelligence. Especially for physical stores, through the digital processing and analysis of sales receipts, the data processing efficiency of merchants can be effectively improved, and then the performance in aspects such as inventory management and promotion activity design can be optimized. For the commodity transactions in pharmacies, in the related technologies, the sales data of commodities (such as drugs) are mainly counted manually, and the inventory planning of commodities is carried out according to the statistical data relying on the experience of store clerks. It is impossible to discover potential data value from the receipts, which often leads to problems such as inaccurate data statistics or low efficiency, and is likely to cause problems of insufficient inventory or overstock. It can be seen that there are problems of low efficiency in pharmacy management in the related technologies. Summary of the Invention
[0003] To solve the above technical problems, this application provides a method, device, electronic device, and storage medium for processing pharmacy receipts.
[0004] In a first aspect, this application provides a method for processing pharmacy receipts, including: obtaining a receipt image set, where the receipt image set includes multiple receipt images; performing OCR recognition on each receipt image in the receipt image set to obtain a set of receipt data, where each receipt data in the set of receipt data corresponds to a receipt image, and each receipt data includes drug sales information; performing data mining based on the set of receipt data to obtain a mining result, where the mining result is used for drug management.
[0005] By adopting the above technical solution, it is possible to achieve efficient and automated processing of pharmacy receipts, and improve the speed and accuracy of data collection and analysis. Specifically, obtaining the receipt image set and extracting drug sales information through OCR recognition, the automated processing of receipt data can significantly improve the data processing efficiency, reduce the time and cost of manual operations, and avoid the problems of high error rate and low efficiency in traditional manual statistics; further, based on the obtained set of receipt data for data mining, it is possible to more accurately predict drug demand, optimize inventory management, avoid problems of insufficient or excessive inventory, and can reveal valuable information hidden in a large amount of data, such as drug demand trends, customer purchase habits, etc., so as to help pharmacies better manage inventory and design promotion activities, achieving the effect of improving pharmacy management efficiency.
[0006] Optionally, data mining is performed on a set of receipt data to obtain mining results, including: performing association analysis on a set of receipt data to obtain a target frequent item set in the set of receipt data, where the target frequent item set includes a drug combination that co-occurs in the set of receipt data more than or equal to a preset frequency threshold, and the mining results include the target frequent item set.
[0007] By adopting the above technical solution, association analysis is performed on a set of receipt data, and it is possible to accurately identify which drugs are often purchased simultaneously to form a target frequent item set. Based on these frequent item sets, the pharmacy can reasonably arrange the positions of relevant drugs on the shelves, promote associated sales, and enhance the customer shopping experience. It is possible to discover drug combinations that are frequently purchased from a set of receipt data, thereby helping the pharmacy to more scientifically conduct product layout and inventory management.
[0008] Optionally, after obtaining the target frequent item set in the set of receipt data, the above method further includes: adjusting the layout of the target drug combination on the shelf according to the target frequent item set, where the target drug combination is any drug combination included in the target frequent item set.
[0009] By adopting the above technical solution, it is possible to adjust the layout of the target drug combination on the shelf according to the target frequent item set in the receipt data. Specifically, by performing association analysis on a set of receipt data, drug combinations that are often purchased together during the sales process are identified, and then these drug combinations are placed together, thereby enhancing the customer shopping experience, reducing the time for customers to search for the required drugs, and increasing sales. At the same time, this layout adjustment also helps the pharmacy to better manage inventory and avoid inventory backlog or shortage problems caused by unreasonable drug placement.
[0010] Optionally, data mining is performed on a set of receipt data to obtain mining results, including: predicting according to a set of receipt data using a target prediction model to obtain a prediction result, where the prediction result includes the demand for the target drug, the mining results include the prediction result, and the target prediction model is trained using historical sales data.
[0011] By adopting the above technical solution, it is possible to efficiently predict the future demand for the target drug according to a set of receipt data using the target prediction model, which not only improves the scientificity and accuracy of drug management, but also can, to a certain extent, avoid inventory shortages or surpluses caused by human experience judgment errors, thereby enhancing the overall operation efficiency and service quality of the pharmacy.
[0012] Optionally, the target prediction model is trained as follows: Obtain a training sample data set, where each training sample in the training sample data set includes the first sales data of the sample drug in the past preset period and the second sales data after the preset period, and the second sales data is used to represent the actual sales data of the sample drug within a preset duration after the preset period; Use the training sample data set to train the initial prediction model until the loss value between the predicted sample result output by the initial prediction model and the actual sample result meets the preset convergence condition to end the training, and use the initial prediction model at the end of the training as the target prediction model, where the second sales data includes the actual sample result, and in the case where the preset convergence condition is not met, adjust the model parameters in the initial prediction model.
[0013] By adopting the above technical solution, a training sample data set is obtained, and each training sample contains the first sales data in the past preset period and the actual sales data in the subsequent preset duration. Using the training sample data set for training ensures the quality of the basic data for model training. Using these training sample data to train the initial prediction model and continuously adjusting the model parameters minimize the error between the predicted result output by the model and the actual sales data, thereby improving the prediction accuracy. When the model reaches the preset convergence condition, the finally formed prediction model can more accurately predict the future drug demand, help the pharmacy better manage inventory and design promotion activities, reduce the risk of inventory backlog or shortage, and improve the operation efficiency.
[0014] Optionally, perform OCR recognition on each receipt image in the receipt image set to obtain a set of receipt data, including: Establish an OCR recognition model, where the OCR recognition model is obtained as follows: Collect receipt samples from the pharmacy and perform labeling processing to obtain a training receipt sample set; Use the receipt sample set to train the original recognition model through a reinforcement learning algorithm, so that the original recognition model selects a preset processing strategy under different input conditions to obtain the OCR recognition model, where the preset processing strategies include: a strategy for dynamically adjusting the image sharpening degree and a rotation angle compensation strategy; Use the OCR recognition model to perform OCR recognition on each receipt image in the receipt image set to obtain a set of receipt data.
[0015] By adopting the above technical solution, receipt samples of pharmacies are collected and labeled to obtain a trained receipt sample set, ensuring the pertinence of the OCR recognition model and improving the generalization ability of the model; the original recognition model is trained using the reinforcement learning algorithm, enabling the model to select the optimal preset processing strategy under different input conditions, enhancing the adaptability and robustness of the model; the application of the dynamic adjustment strategy for image sharpness degree and the rotation angle compensation strategy effectively solves the problem of inconsistent receipt image quality and further improves the accuracy and speed of OCR recognition; finally, the trained OCR recognition model is used to recognize receipt images, ensuring the high-quality generation of a set of receipt data and laying a solid foundation for subsequent data mining and analysis. This technical solution can significantly improve the recognition accuracy and efficiency of receipt images.
[0016] Optionally, the above method further includes: obtaining a plurality of user feature vectors from a set of receipt data, where each user feature vector among the plurality of user feature vectors includes basic features, behavioral features, and time features; performing clustering analysis on the plurality of user feature vectors using the K-means clustering algorithm to obtain a target clustering result, where the mining result includes the target clustering result; and performing personalized recommendation according to the clustering result.
[0017] By adopting the above technical solution, multi-dimensional feature information of users can be extracted from a set of receipt data, and clustering analysis is performed on these feature vectors using the K-means clustering algorithm to obtain a target clustering result. This method not only improves the accuracy of data analysis but also provides personalized recommendation services for different users according to the target clustering result, further enhancing the user experience and satisfaction. Specifically: through the processing of receipt data, a plurality of user feature vectors including basic features, behavioral features, and time features are obtained, making the user portrait more comprehensive and accurate; the K-means clustering algorithm is used to classify the user feature vectors, and the generated target clustering result can reflect the behavioral patterns and preferences of different types of users; based on the target clustering result, the system can formulate differentiated recommendation strategies for different types of user groups, improving the accuracy of recommendations and user stickiness.
[0018] In the second aspect of the present application, a processing device for pharmacy receipts is further provided, including: an acquisition module for acquiring a receipt image set, where the receipt image set includes a plurality of receipt images; a recognition module for performing OCR recognition on each receipt image in the receipt image set to obtain a set of receipt data, where each receipt data in the set of receipt data corresponds to a receipt image, and each receipt data includes drug sales information; and a processing module for performing data mining on the set of receipt data to obtain a mining result, where the mining result is used for drug management.
[0019] In the third aspect of the present application, an electronic device is further provided, including a memory and a processor. A computer program is stored on the memory, and when the processor executes the program, the method steps of any one of the above are implemented.
[0020] In the fourth aspect of the present application, a computer-readable storage medium is further provided. The computer-readable storage medium stores instructions, and when the instructions are executed, the method steps of any one of the above are executed.
[0021] In summary, one or more technical solutions provided in the present application have at least the following technical effects or advantages: 1. It can achieve efficient and automated processing of pharmacy receipts, which can help pharmacies better manage inventory and design promotional activities, achieving the effect of improving pharmacy management efficiency; 2. It can discover frequently purchased drug combinations from a set of receipt data, thereby helping pharmacies more scientifically arrange product layouts and manage inventory, enhancing the shopping experience of users; 3. It can use the target prediction model to efficiently predict the future demand for target drugs based on a set of receipt data, which not only improves the scientificity and accuracy of drug management, but also can, to a certain extent, avoid inventory shortages or surpluses caused by human experience judgment errors, thereby enhancing the overall operation efficiency and service quality of pharmacies; 4. Using the trained OCR recognition model to recognize receipt images ensures the high-quality generation of a set of receipt data, and can significantly improve the recognition accuracy and efficiency of receipt images; 5. Conducting clustering analysis on a set of receipt data to obtain the target clustering result, and can also provide personalized recommendation services for different users according to the target clustering result, further enhancing the user experience and satisfaction. Description of the Drawings
[0022] Figure 1 is a flowchart of a method for processing pharmacy receipts provided by an embodiment of the present application; Figure 2 is a structural block diagram of a device for processing pharmacy receipts provided by an embodiment of the present application; Figure 3 is a schematic structural diagram of an electronic device disclosed by an embodiment of the present application.
[0023] Description of the reference numerals: 300 - electronic device; 301 - processor; 302 - communication bus; 303 - user interface; 304 - network interface; 305 - memory. Detailed Embodiments
[0024] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments.
[0025] In the description of the embodiments of this application, words such as "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design solution described as "for example" or "for instance" in the embodiments of this application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "for example" or "for instance" is intended to present relevant concepts in a specific manner.
[0026] In the description of the embodiments of this application, the meaning of the term "a plurality" refers to two or more. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. The terms "include", "comprise", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0027] The following will describe the embodiments of this application in conjunction with the attached Figures 1 - 3 for illustration.
[0028] This application provides a method for processing pharmacy receipts. Referring to Figure 1 , Figure 1 is a flowchart of a method for processing pharmacy receipts provided by an embodiment of this application. The method includes: Step S101: Obtain a set of receipt images, where the set of receipt images includes a plurality of receipt images; Step S102: Perform OCR recognition on each receipt image in the set of receipt images to obtain a set of receipt data, where each receipt data in the set of receipt data corresponds to a receipt image, and each receipt data includes drug sales information; Step S103: Perform data mining based on the set of receipt data to obtain a mining result, where the mining result is used for drug management.
[0029] Through the above steps, the efficient and automated processing of pharmacy receipts can be achieved, improving the speed and accuracy of data collection and analysis. Specifically, obtaining the receipt image set and extracting drug sales information through OCR recognition, the automated processing of receipt data can significantly improve the efficiency of data processing, reduce the time and cost of manual operations, and avoid the problems of high error rates and low efficiency in traditional manual statistics; further, data mining based on the obtained set of receipt data can more accurately predict drug demand, optimize inventory management, avoid problems of insufficient or excessive inventory, and can reveal valuable information hidden in a large amount of data, such as drug demand trends, customer purchase habits, etc., thus helping pharmacies better manage inventory and design promotional activities, and improving operational efficiency and service quality.
[0030] In this embodiment, by obtaining a receipt image set and performing OCR (Optical Character Recognition) to convert the text information in the images into processable electronic data, sales information is obtained, and further data mining is performed on this information to support drug management. For example, receipt images generated during the sales process in a pharmacy can be collected through scanning, photographing, or other technical means to form a receipt image set. Using optical character recognition (OCR) technology, each image in the receipt image set is recognized, and the text information in the images is converted into editable and analyzable electronic data, that is, a set of receipt data containing drug sales information is generated. For example, the receipt data may include sales information such as drug name, quantity, price, transaction time, and total transaction price. Data mining is performed on the obtained set of receipt data to extract valuable information such as sales quantity, sales amount, and drug types, and further analysis and decision-making are based on this information. For example, association rule analysis is performed on a set of receipt data to discover the most frequently purchased drug combinations, and then the layout of the shelves can be adjusted to place the drugs in the associated drug combinations together or in nearby locations; or, based on a set of receipt data, predictions are made to predict the demand or sales trend of a certain drug in the future, so as to facilitate the formulation of corresponding inventory plans; or, a set of receipt data is analyzed to classify customers with similar purchase behaviors, so as to accurately formulate marketing strategies, such as designing promotional activities or personalized recommendation services. Through the automated receipt processing process, this embodiment can greatly shorten the data processing time and improve the work efficiency of merchants. Based on the mining results of sales data, merchants can more accurately predict future demand, thus formulating more reasonable inventory plans. With real-time and accurate data support, pharmacies can more effectively make management decisions such as promotional activity design and inventory planning, thereby improving the overall management level. In some actual application scenarios, some pharmacies newly connect to the intelligent marketing system and do not store previous orders. The receipts are only printed with real-time input data and no information storage is performed. For example, in some promotional activities, pharmacies may use independent devices to print receipts, and these devices are not connected to the pharmacy's computer system. Therefore, it is necessary to collect some marketing data through OCR recognition to assist in marketing.
[0031] In an alternative embodiment, data mining is performed on a set of receipt data to obtain a mining result, including: performing association analysis on a set of receipt data to obtain a target frequent item set in the set of receipt data, where the target frequent item set includes drug combinations that co-occur in the set of receipt data more than or equal to a preset frequency threshold, and the mining result includes the target frequent item set.
[0032] In the above embodiments, by performing association analysis on a set of receipt data, it is possible to accurately identify which drugs are often purchased together, forming target frequent item sets. Based on these frequent item sets, pharmacies can reasonably arrange the positions of relevant drugs on the shelves, promote associated sales, and enhance the customer shopping experience. It is possible to discover drug combinations that are frequently purchased from a set of receipt data, thereby helping pharmacies to more scientifically arrange product layouts and manage inventories. At the same time, this data analysis method can also reveal potential market trends, guiding pharmacies to better plan inventories and avoid inventory backlogs or out-of-stock problems.
[0033] In practical applications, after OCR recognition and generation of receipt data containing drug sales information, the data can be preprocessed, such as cleaning and sorting, to ensure the accuracy and consistency of the data. Then, association analysis is performed on a set of preprocessed receipt data. This is an important method in data mining for discovering associations between data items. In this scenario, association analysis aims to find out which drug combinations often appear together during pharmacy sales, that is, target frequent item sets, which refer to those drug combinations that frequently appear simultaneously in sales records. For example, if many users purchase drug A while also purchasing drug B, it indicates that they may be complementary products or customers are accustomed to purchasing them together. In this way, drug A and drug B can be regarded as a drug combination in the target frequent item set. To determine which drug combinations are frequent, a preset occurrence threshold needs to be set. Only those drug combinations whose co-occurrence times are greater than or equal to this threshold will be regarded as target frequent item sets. The preset occurrence threshold can be set according to actual application needs and can also be adjusted as needed; based on the results of the association analysis, a mining result containing target frequent item sets is generated. These target frequent item sets can provide valuable sales information for pharmacies. The methods in the related technologies may only stay at the simple statistical level of receipt data and do not deeply explore the correlations and trends between data. Through association analysis in this embodiment, potential associations between drugs can be discovered, providing deeper market insights for pharmacies. Through association analysis in this embodiment, pharmacies can more deeply understand sales data, discover the correlations and trends between drugs, thereby enhancing data analysis capabilities. Based on the results of the target frequent item sets, pharmacies can formulate more precise promotion strategies, such as bundling sales, combined discounts, etc., to attract customers and increase sales. Additionally, knowing which drug combinations are often purchased together by customers can help pharmacies optimize product displays, place related drugs together, facilitate customer selection, and improve the shopping experience; the target frequent item sets obtained through association analysis can also provide guidance for the inventory management of pharmacies, ensuring an adequate inventory of popular drug combinations and avoiding out-of-stock or overstocked inventory situations.
[0034] The following is an example for illustration. Suppose the sales receipt data of a certain pharmacy is as follows: ① cold medicine, vitamin C; ② cold medicine, vitamin C, mask; ③ cold medicine, mask; ④ vitamin C, mask; ⑤ cold medicine, vitamin C. A preset frequency threshold can be defined first, such as 3 times, that is, only those drug combinations that appear simultaneously in at least 3 transactions will be regarded as the combinations in the target frequent item set. After defining the preset frequency threshold, frequent item sets can be searched for in a set of receipt data. For example, [cold medicine, vitamin C] as mentioned above, which indicates that when a user buys cold medicine, there is a high probability that they will buy vitamin C. In practical applications, the minimum support or confidence threshold can also be set to identify frequent item sets and association rules. For example, first extract the names, quantities and corresponding transaction IDs of all drugs from the receipts, remove irrelevant fields (such as cashier numbers), standardize the drug name format, and use association rule mining algorithms such as Apriori or FP-Growth, set the minimum support and confidence thresholds, and identify frequent item sets and association rules.
[0035] In an optional embodiment, after obtaining the target frequent item set in a set of receipt data, the above method further includes: adjusting the layout of the target drug combination on the shelf according to the target frequent item set, where the target drug combination is any drug combination included in the target frequent item set.
[0036] In the above embodiment, the layout of the target drug combination on the shelf can be adjusted according to the target frequent item set in the receipt data. Specifically, through the association analysis of a set of receipt data, the drug combinations that are often purchased together during the sales process are identified, and then these drug combinations are placed together, thereby enhancing the customer shopping experience, reducing the time for customers to search for the required drugs, and increasing the sales volume. At the same time, this layout adjustment also helps the pharmacy to better manage the inventory and avoid inventory backlog or shortage problems caused by unreasonable drug placement.
[0037] The drug combinations that are often purchased together are identified through data mining techniques, such as the drug combinations in the target frequent item set. The layout of the drugs on the pharmacy shelf is adjusted using this information so that the drug combinations that are often purchased together are placed adjacent or close to each other, thereby improving the shopping convenience of users and possibly promoting sales, achieving the effect of enhancing the operation efficiency; in the traditional pharmacy layout, the placement of drugs may be arranged based on product categories, brands or the habits of the staff, and may not take into account the purchase habits of customers, resulting in customers spending more time searching for relevant drugs; in this embodiment, by placing the drug combinations that are often purchased together in close positions, the time for customers to search for drugs can be reduced, and the shopping experience of users is enhanced.
[0038] In an alternative embodiment, data mining is performed on a set of receipt data to obtain a mining result, including: predicting using a target prediction model based on a set of receipt data to obtain a prediction result, where the prediction result includes the demand for a target drug, the mining result includes the prediction result, and the target prediction model is trained using historical sales data.
[0039] In the above embodiment, the target prediction model can be used to efficiently predict the future demand for the target drug based on a set of receipt data. This not only improves the scientificity and accuracy of drug management but also, to a certain extent, avoids problems such as insufficient or excessive inventory caused by human experience judgment errors, thereby enhancing the overall operation efficiency and service quality of the pharmacy.
[0040] A target prediction model is trained using historical sales data. This model can be a machine learning model, such as linear regression, decision tree, random forest, neural network, etc., or a deep learning model. The goal of the model is to predict the demand for the target drug in the future. Optionally, the current set of receipt data can be input into the trained target prediction model for prediction to obtain a prediction result, or a part of the data of some drugs included in the set of receipt data can be input into the trained target prediction model for prediction to obtain a prediction result. The prediction result includes the demand for the target drug in the future. The target drug can be one or more drugs. For example, predict the demand for the target drug in the next week or month. The prediction result is used as part of the mining result to guide the inventory management, procurement plan, sales strategy, etc. of the pharmacy. The target prediction model of this embodiment is trained based on historical sales data, which means that the model can identify various factors affecting drug sales (such as seasonal changes, epidemic trends, etc.) and make predictions accordingly. In practical applications, the model learns the patterns and trends in historical data to predict future demand changes. For example, the model input values (feature variables) may include time series data, drug categories, price changes, promotional activities, etc., and the model output value (target variable) is the demand for the drug. Traditional pharmacy inventory management may rely on experience or intuition, resulting in inventory backlogs or shortages. Through the target prediction model in this embodiment, the drug demand can be predicted more accurately, thereby optimizing inventory management. Through this embodiment, using the target prediction model, the pharmacy can more accurately predict the drug demand, thereby optimizing inventory management, reducing inventory backlogs and shortages. The procurement plan based on the prediction result can more precisely meet the operation needs of the pharmacy, reduce procurement costs, and improve profitability. By optimizing inventory management and procurement plans and formulating precise sales strategies, the pharmacy can overall enhance operation efficiency, reduce operation costs, and improve market competitiveness.
[0041] In an optional embodiment, the target prediction model is trained as follows: Obtain a training sample data set, where each training sample in the training sample data set includes the first sales data of the sample drug in the past preset period and the second sales data after the preset period, and the second sales data is used to represent the actual sales data of the sample drug within a preset duration after the preset period; Use the training sample data set to train the initial prediction model until the loss value between the predicted sample result output by the initial prediction model and the actual sample result meets the preset convergence condition to end the training, and use the initial prediction model at the end of the training as the target prediction model, where the second sales data includes the actual sample result, and in the case where the preset convergence condition is not met, adjust the model parameters in the initial prediction model.
[0042] In the above embodiment, obtaining the training sample data set, where each training sample contains the first sales data in the past preset period and the actual sales data within the subsequent preset duration, and using the training sample data set for training ensures the quality of the basic data for model training. Using these training sample data to train the initial prediction model and continuously adjusting the model parameters minimize the error between the predicted result output by the model and the actual sales data, thereby improving the accuracy of the prediction. When the model reaches the preset convergence condition, the finally formed prediction model can more accurately predict the future drug demand, help pharmacies better manage inventory and design promotional activities, reduce the risk of inventory backlog or shortage, and improve operational efficiency. This embodiment can effectively train the target prediction model, which is based on historical sales data, and ensures the accuracy and reliability of the model through accurate sample data sets and scientific training methods.
[0043] First, obtain a training sample dataset containing multiple training samples. Each training sample includes the first sales data of the sample drugs within a preset past period and the second sales data within a preset duration after this preset period. Here, the second sales data is actually the actual sales data of the sample drugs after this preset period and is used as the label or target value for model training. Next, use this training sample dataset to train the initial prediction model. During the training process, the model will attempt to predict the second sales data (target value) based on the first sales data (input features) in each training sample. The difference (i.e., the loss value) between the predicted result output by the model and the actual sample result will be calculated and used to evaluate the performance of the model. If the loss value between the predicted result and the actual result does not meet the preset convergence condition, the model parameters in the initial prediction model will be adjusted to improve the prediction performance of the model. This process will be repeated until the loss value meets the preset convergence condition and the training ends. At the end of the training, the initial prediction model at this time is determined as the target prediction model for subsequent receipt data mining and drug demand prediction. Using the above training samples to train the initial prediction model, the goal is to let the model learn to infer the second sales data (i.e., the future sales trend) from the first sales data. During the training process, the model will continuously adjust its own parameters to minimize the difference (loss value) between the predicted value and the true value. When the loss value reaches a certain preset convergence condition, it is considered that the model has fully learned the sales pattern. At this time, the training ends, and the final version of the model is determined as the target prediction model. In this embodiment, by using historical sales data as the training basis, the model can identify various factors affecting sales and make more accurate demand predictions accordingly. Compared with traditional empirical judgments, this method provides higher accuracy and reliability. As new data is continuously added, the model can be retrained or fine-tuned to always keep it up-to-date and be able to quickly respond to market changes and technological advancements. It provides reliable sales prediction information for pharmacy management, helping them make more informed business decisions, such as the best order time, quantity, and evaluating the potential of new product introductions. Based on the prediction results of the target prediction model, pharmacies can formulate inventory management strategies more precisely, reduce inventory backlogs and shortages, and improve operational efficiency.
[0044] The above target prediction model can be constructed based on algorithms such as ARIMA or LSTM, which can capture trends and seasonal variations in time series. Taking the LSTM model as an example, the training sample set is composed of the historical sales data of the pharmacy. The LSTM can handle more types of input features and is good at capturing complex non-linear patterns. For example, the input features include time series data, external factors, lag variables, time windows, etc. The time series data can be the daily, weekly, or monthly sales volume of drug C. The external factors can include weather conditions (such as meteorological data like temperature, humidity, etc.), holiday flags (such as whether it is a holiday), and market activities (such as promotional activities, advertising, etc.). The LSTM model can automatically learn the importance of historical data through internal mechanisms and can choose to include the sales data of the past few days as additional inputs to enhance the model performance. Additionally, to capture short-term trends, a fixed-length time window (such as 7 days, or other time) is usually defined. Each training sample will include all the feature values within this window. The output target can be the sales volume in a future period, that is, the target value to be predicted. For the LSTM model, the output is an estimate of the sales volume of drug C in a future time period (such as tomorrow, next week, or next month) based on historical data. Taking the prediction of the sales volume of the next day as an example, the above first sales data is the daily sales data of a past period (such as January 1st - January 7th, 2 years ago), and the actual sales volume on the day after 7 consecutive days (that is, January 8th) is the actual result. The result predicted by the model is compared with this actual result; the above first sales data is multiple groups of data, and each group of data is the sales data of a consecutive 7 days in the past; during the model training process, grid search or random search can be used to find the optimal hyperparameter configurations such as the learning rate and batch size. The generalization ability of the model can be tested through the cross-validation method, and the performance of the model is evaluated on the test set. Commonly used metrics include the mean squared error (MSE), mean absolute error (MAE), etc. If the model performance is poor, the parameters can be adjusted and returned. Once the model is trained and passed the validation, it can be applied to new time points to predict future sales trends.
[0045] In an optional embodiment, OCR recognition is performed on each receipt image in the receipt image set to obtain a set of receipt data, including: establishing an OCR recognition model, where the OCR recognition model is obtained through the following method: collecting receipt samples of the pharmacy and performing labeling processing to obtain a training receipt sample set; training the original recognition model using the receipt sample set through a reinforcement learning algorithm to enable the original recognition model to select a preset processing strategy under different input conditions to obtain the OCR recognition model, where the preset processing strategies include: a dynamic adjustment strategy for image sharpening degree and a rotation angle compensation strategy; using the OCR recognition model to perform OCR recognition on each receipt image in the receipt image set to obtain a set of receipt data.
[0046] In the above embodiments, receipt samples of pharmacies are collected and labeled to obtain a set of labeled receipt samples, ensuring the pertinence of the OCR recognition model and improving the generalization ability of the model. The original recognition model is trained using a reinforcement learning algorithm, enabling the model to select the optimal preset processing strategy under different input conditions, thus enhancing the adaptability and robustness of the model. The application of the strategies for dynamically adjusting the image sharpening degree and the rotation angle compensation effectively solves the problem of inconsistent receipt image quality, further improving the accuracy and speed of OCR recognition. Finally, the trained OCR recognition model is used to recognize receipt images, ensuring the high-quality generation of a set of receipt data and laying a solid foundation for subsequent data mining and analysis. Through this embodiment, the recognition accuracy and efficiency of receipt images can be significantly improved.
[0047] First, receipt samples of pharmacies are collected and labeled. The labeling process may include annotating information such as text, numbers, barcodes, etc. on the receipts for subsequent training and recognition. Then, using the collected set of receipt samples, the original recognition model is trained through a reinforcement learning algorithm. Reinforcement learning is a method of enabling the model to learn the optimal strategy through interactions with the environment. Here, the model is trained to select preset processing strategies under different input conditions to optimize the recognition effect. The preset processing strategies include the strategy for dynamically adjusting the image sharpening degree and the rotation angle compensation strategy. These strategies are designed to preprocess the image according to the actual situation of the receipt image to improve the accuracy of OCR recognition. Using the trained OCR recognition model, according to the preset processing strategies, OCR recognition is performed on each receipt image in the receipt image set. The recognition process may include steps such as image preprocessing (such as sharpening, rotation compensation, etc.), character segmentation, feature extraction, and character recognition. Finally, the OCR recognition model outputs a set of receipt data, including information such as text, numbers, barcodes, etc. on the receipts. This information can be used for subsequent data analysis and processing. Through the introduction of reinforcement learning and preset processing strategies in this embodiment, the recognition accuracy of the OCR recognition model under different input conditions can be significantly improved, reducing the situations of misrecognition and missed recognition. Through preset processing strategies such as the strategy for dynamically adjusting the image sharpening degree and the rotation angle compensation strategy, flexible preprocessing can be performed according to the actual situation of the receipt image, enhancing the adaptability of the model. By collecting diverse receipt samples of pharmacies and conducting reinforcement learning training, the generalization ability of the OCR recognition model can be improved, enabling it to recognize more types and styles of receipt images. Using the OCR recognition model to automatically recognize receipt images can greatly accelerate the speed of data processing, reduce errors and costs of manual input, and improve the operational efficiency of pharmacies.
[0048] In an optional embodiment, the above method further includes: obtaining a plurality of user feature vectors according to a set of receipt data, where each user feature vector among the plurality of user feature vectors includes a basic feature, a behavior feature, and a time feature; performing clustering analysis on the plurality of user feature vectors by using the K-means clustering algorithm to obtain a target clustering result, where the mining result includes the target clustering result; and performing personalized recommendation according to the clustering result.
[0049] In the above embodiment, multi-dimensional feature information of users can be extracted from a set of receipt data, and clustering analysis is performed on these feature vectors by using the K-means clustering algorithm to obtain a target clustering result. This method not only improves the accuracy of data analysis, but also can provide personalized recommendation services for different users according to the target clustering result, further enhancing the user experience and satisfaction. Specifically: by processing the receipt data, a plurality of user feature vectors including basic features, behavior features, and time features are obtained, making the user portrait more comprehensive and accurate; the K-means clustering algorithm is used to classify the user feature vectors, and the generated target clustering result can reflect the behavior patterns and preferences of different types of users; based on the target clustering result, the system can formulate differentiated recommendation strategies for different types of user groups, improving the accuracy of recommendation and user stickiness.
[0050] Extract the basic features, behavior features, and time features of users from a set of receipt data to construct a plurality of user feature vectors. For example, the basic features may include basic information such as the age and gender of the user, the behavior features may include the purchase frequency, preferred drug types, and single purchase amount, etc., which are extracted from the receipt data, and the time features may include the purchase time, purchase cycle, seasonal changes, etc. Use the K-means clustering algorithm to cluster these user feature vectors, group users with similar features, and perform personalized product recommendations for users in different groups according to the clustering result. For example, the group of chronic disease patients is mostly people over 60 years old, who rely on specific drugs for a long time to treat chronic diseases; those who pay attention to health and beauty are mostly young female customer groups aged 20-35, etc. Use the K-means clustering algorithm to perform clustering analysis on the above constructed plurality of user feature vectors, aiming to identify user groups with similar shopping patterns or preferences. In this way, customers can be divided into different categories, and each category represents a group of users with common characteristics. In the related technology, the lack of in-depth analysis of user behavior and preferences leads to the pharmacy's inability to accurately grasp user needs and affects the effectiveness of sales strategies; in this embodiment, by extracting user feature vectors and performing clustering analysis, the purchase behavior and needs of users can be more accurately identified, so as to provide personalized recommendations. By extracting user feature vectors in multiple dimensions, user data can be more comprehensively utilized, improving the utilization rate of data and the accuracy of recommendations.
[0051] This application provides a marketing data processing and analysis method based on OCR recognition of pharmacy receipts. By adopting the OCR recognition technology optimized by deep learning algorithms, the information recognition accuracy of pharmacy receipts is greatly improved, the need for manual intervention is reduced, and a large amount of time and costs are saved. Moreover, the receipt data recognized by OCR is combined with big data analysis to mine marketing rules and enhance marketing effectiveness.
[0052] Collect a large number of real pharmacy receipt samples, perform labeling processing on them, and establish a training data set; use reinforcement learning algorithms to train the model so that it learns to select the best preprocessing and recognition strategies under different input conditions, such as dynamically adjusting the image sharpening degree, rotation angle compensation, etc. Continuously evaluate the model performance through the test system, and continuously optimize the model structure and parameter settings to obtain a receipt recognition model. Use the receipt recognition model to recognize real receipt images to obtain receipt data (order number, detail order number, membership card number, drug code, drug name, drug specification, number of boxes, order time, manufacturer, chain branch, store number, store name, retail unit price, transaction total price, etc.); perform data analysis on the receipt data, combine it with the time series prediction model, and statistically analyze the pharmacy transaction rules to assist marketing.
[0053] This application also provides a processing device for pharmacy receipts, as Figure 2 shown, Figure 2 which is a structural block diagram of a processing device for pharmacy receipts provided by an embodiment of this application. The device includes: An acquisition module 201, configured to acquire a receipt image set, where the receipt image set includes multiple receipt images; An identification module 202, configured to perform OCR recognition on each receipt image in the receipt image set to obtain a set of receipt data, where each receipt data in the set of receipt data corresponds to a receipt image, and each receipt data includes drug sales information; A processing module 203, configured to perform data mining based on the set of receipt data to obtain a mining result, where the mining result is used for drug management.
[0054] It should be noted that when the device provided in the above embodiment realizes its functions, only the above-mentioned division of each functional module is used for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be elaborated here.
[0055] The present application also provides a computer-readable storage medium, in which instructions are stored, and when the instructions are executed, the method steps described in any one of the above are performed.
[0056] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical discs that can store computer programs.
[0057] The present application also discloses an electronic device. As Figure 3 shown, Figure 3 FIG. 10 is a schematic structural diagram of an electronic device disclosed in an embodiment of the present application. The electronic device 300 may include: at least one processor 301, at least one communication bus 302, a user interface 303, at least one network interface 304, and a memory 305.
[0058] Among them, the communication bus 302 is used to realize the connection and communication between these components.
[0059] Among them, the user interface 303 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 303 may further include a standard wired interface and a wireless interface.
[0060] Among them, the network interface 304 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).
[0061] Among them, the processor 301 may include one or more processing cores. The processor 301 connects various parts within the entire electronic device (such as a server) through various interfaces and lines. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 305, and by calling the data stored in the memory 305, it performs various functions of the server and processes data. Optionally, the processor 301 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 301 may integrate a combination of one or more of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communication. It can be understood that the above-mentioned modem may not be integrated into the processor 301 and may be implemented separately by a single chip.
[0062] Among them, the memory 305 may include random access memory (RAM) and may also include read-only memory. Optionally, the memory 305 includes a non-transitory computer-readable storage medium. The memory 305 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 305 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store the data involved in the above-mentioned various method embodiments. Optionally, the memory 305 may also be at least one storage device located far from the aforementioned processor 301. Refer to Figure 3 , as a computer storage medium, the memory 305 may include an operating system, a network communication module, a user interface module, and an application program for a method of processing pharmacy receipts.
[0063] In Figure 3In the electronic device 300 shown, the user interface 303 is mainly used to provide an interface for the user to input and obtain the data input by the user; while the processor 301 can be used to call an application program of a processing method for pharmacy receipts stored in the memory 305. When executed by one or more processors 301, the electronic device 300 is caused to execute one or more of the methods as described in the above embodiments. It should be noted that, for the foregoing method embodiments, for simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0064] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0065] In several implementation manners provided by the present application, it should be understood that the disclosed device or system can be implemented in other ways. For example, the device or system embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some service interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical or other form.
[0066] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0067] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0068] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a memory and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present application. The aforementioned memory includes various media that can store program codes, such as USB flash drives, mobile hard disks, magnetic disks, or optical discs.
[0069] The above are only exemplary embodiments of the present disclosure, and the scope of the present disclosure cannot be limited thereby. That is, any equivalent changes and modifications made in accordance with the teachings of the present disclosure still fall within the scope covered by the present disclosure. Those skilled in the art will readily think of other embodiments of the present disclosure after considering the disclosure of the specification.
[0070] The present application aims to cover any variations, uses, or adaptive changes of the present disclosure that follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not recorded in the present disclosure.
Claims
1. A method for processing pharmacy receipts, characterized in that: include: Acquire a ticket image set, wherein the ticket image set includes a plurality of ticket images; Performing OCR recognition on each receipt image in the receipt image set to obtain a set of receipt data, wherein each receipt data in the set of receipt data corresponds to a receipt image, and each receipt data includes drug sales information; Data mining is performed based on the set of receipt data to obtain mining results, wherein the mining results are used for drug management.
2. The method according to claim 1, characterized in that Data mining is performed based on the set of receipt data to obtain mining results, including: Perform association analysis on the set of receipt data to obtain a target frequent item set in the set of receipt data, wherein the target frequent item set includes drug combinations that appear together in the set of receipt data a number of times greater than or equal to a preset number threshold, and the mining result includes the target frequent item set.
3. The method according to claim 2, characterized in that After obtaining the target frequent item set in the set of receipt data, the method further includes: The layout of the target drug combination on the shelf is adjusted according to the target frequent item set, wherein the target drug combination is any drug combination included in the target frequent item set.
4. The method according to claim 1, characterized in that: Data mining is performed based on the set of receipt data to obtain mining results, including: Predictions are made using a target prediction model based on the set of receipt data to obtain prediction results, wherein the prediction results include the demand for target drugs, the mining results include the prediction results, and the target prediction model is trained using historical sales data.
5. The method according to claim 4, characterized in that The target prediction model is trained in the following way: Acquire a training sample data set, wherein each training sample in the training sample data set includes first sales data of the sample drug within a preset period in the past and second sales data after the preset period, wherein the second sales data is used to represent actual sales data of the sample drug within a preset time period after the preset period; The initial prediction model is trained using the training sample data set until the loss value between the predicted sample results output by the initial prediction model and the actual sample results meets the preset convergence condition and the training is terminated, and the initial prediction model at the end of the training is used as the target prediction model, wherein the second sales data includes the actual sample results, and when the preset convergence condition is not met, the model parameters in the initial prediction model are adjusted.
6. The method according to claim 1, characterized in that Perform OCR recognition on each receipt image in the receipt image set to obtain a set of receipt data, including: Establish an OCR recognition model, wherein the OCR recognition model is obtained by the following method: Collect pharmacy receipt samples and label them to obtain a training receipt sample set; The original recognition model is trained by using the receipt sample set through a reinforcement learning algorithm, so that the original recognition model selects a preset processing strategy under different input conditions to obtain the OCR recognition model, wherein the preset processing strategy includes: a dynamic adjustment image sharpness strategy and a rotation angle compensation strategy; The OCR recognition model is used to perform OCR recognition on each ticket image in the ticket image set according to the preset processing strategy to obtain the set of ticket data.
7. The method according to claim 1, characterized in that The method further comprises: Obtaining a plurality of user feature vectors according to the set of receipt data, wherein each of the plurality of user feature vectors includes a basic feature, a behavior feature, and a time feature; Performing cluster analysis on the multiple user feature vectors using a K-means clustering algorithm to obtain a target clustering result, wherein the mining result includes the target clustering result; Personalized recommendation is performed according to the clustering result.
8. A device for processing drugstore receipts, characterized in that: include: An acquisition module, used to acquire a receipt image set, wherein the receipt image set includes a plurality of receipt images; A recognition module, used for performing OCR recognition on each receipt image in the receipt image set to obtain a set of receipt data, wherein each receipt data in the set of receipt data corresponds to a receipt image, and each receipt data includes drug sales information; The processing module is used to perform data mining based on the set of receipt data to obtain mining results, wherein the mining results are used for drug management.
9. An electronic device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor executes the program, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, and when the instructions are executed, the method according to any one of claims 1 to 7 is performed.
Citation Information
Patent Citations
Medicine combination recommendation method and device, electronic equipment and storage medium
CN111914163A
Bill information identification method, device and equipment based on OCR (Optical Character Recognition) and storage medium
CN118397642A