A medicine package information intelligent identification method and system
By acquiring image information of drug packaging, extracting various types of character information, and performing contextual logic judgment and Bayesian inference network verification, the problem of low recognition efficiency and misreading in drug packaging recognition systems under diverse materials and dynamic environments is solved, achieving higher accuracy and robustness.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NANTONG MATERNAL & CHILD HEALTH CARE HOSPITAL
- Filing Date
- 2026-04-30
- Publication Date
- 2026-05-29
AI Technical Summary
Existing pharmaceutical packaging identification systems are inefficient and prone to misreading when faced with diverse packaging materials and dynamic production environments, making it difficult to meet stringent industry standards and regulatory requirements.
By acquiring image information of drug packaging, various types of character information are extracted, including character stroke structure, character texture and geometric contour. The weight of character information is dynamically set according to the type and degree of defects, and the recognition results are verified and corrected through Bayesian inference network that combines contextual logic judgment and multi-dimensional information fusion.
It significantly improves the accuracy and robustness of drug packaging information recognition, can adaptively process complex images, reduce the false recognition rate, and meet stringent industry standards and regulatory requirements.
Smart Images

Figure CN122116379A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent identification of pharmaceutical packaging information, and in particular to a method and system for intelligent identification of pharmaceutical packaging information. Background Technology
[0002] Accurate and efficient identification of key information on drug packaging is crucial for ensuring product quality, achieving end-to-end traceability, and effectively preventing counterfeiting at every stage of drug production and distribution. However, in actual automated production lines, the diversity of drug packaging materials and the dynamic changes in the production environment present significant challenges to existing intelligent identification systems in terms of image acquisition and information processing.
[0003] For example, on a modern automated pharmaceutical production line, sensors trigger cameras to take pictures as medicine boxes pass by, and the system then processes the acquired images. Initially, the production line typically handled standardized paper boxes with a matte finish. Under these ideal conditions, the system performed exceptionally well. With a fixed camera angle and preset lighting parameters, the system could obtain clear, high-contrast images of the top surface of the boxes. The image processing unit could reliably locate key information areas such as the printed batch number, expiration date, and regulatory code, and subsequent character recognition and code parsing modules accurately output the recognition results. The system compared these results with reference information in the production plan database. Boxes with incorrect information or those that could not be identified were pushed to the defective product channel by a robotic arm.
[0004] However, with changing market demands, pharmaceutical companies introduced a series of new packaging formats to enhance product image and anti-counterfeiting features. One type was a soft bag with a highly reflective metallic film, and another was a brown glass bottle with a curved shape. When these new packages were put into production, the existing identification system immediately faced severe challenges. First, under fixed top lighting, the highly reflective film would produce specular reflection, creating large overexposed areas—glaring bright spots. These bright spots would randomly cover the production batch number or regulatory code, causing significant loss of image information captured by the camera. Consequently, the character recognition module could not identify complete characters, and the system frequently misclassified qualified products as defective. Similarly, for curved glass bottles, their curved shape will cause light to converge or diverge, not only forming irregular highlight stripes, but also causing perspective distortion of the text printed on the side due to the curvature of the bottle. For example, the originally straight number "1" may appear curved in the image, and the edges of straight characters become blurred, which exceeds the processing capabilities of the original character recognition module.
[0005] Furthermore, the mechanical vibrations of the production line itself also cause interference. The minute vibrations generated by the high-speed conveyor belt and surrounding equipment are transmitted to the pharmaceutical packaging. If the packaging happens to be at the peak or trough of this vibration at the moment the camera exposes, the captured image will exhibit slight motion blur. This blur, superimposed on the previously mentioned issues of highlights, distortion, and shadows, further degrades image quality. For example, a number already partially obscured by reflections, after slight motion blur, may become extremely similar in outline to another number, causing the system to make an incorrect identification. A large number of qualified products are incorrectly rejected, requiring manual secondary inspection and release. This not only fails to improve efficiency but also becomes a new bottleneck in the production process, completely contradicting the original intention of the system deployment. Summary of the Invention
[0006] This application discloses a method and system for intelligent identification of drug packaging information, aiming to solve the technical problems of low efficiency, easy misreading, and difficulty in meeting strict industry standards and regulatory requirements of traditional identification methods in the drug production and distribution process, as well as the challenges faced by existing intelligent identification systems in image acquisition and information processing.
[0007] In a first aspect, this application discloses a method for intelligent identification of drug packaging information, comprising the following steps: The process involves: acquiring image information of drug packaging; extracting various types of character information from the image information, including character stroke structure, character texture, and character geometric contours; determining the types and degrees of defects present in the image information; setting the weight of various types of character information based on the defect types and degrees; integrating the various types of character information with the set weights to obtain a preliminary comprehensive description of the characters; performing contextual logic judgment on the preliminary comprehensive description of the characters to obtain a final comprehensive description of the characters; and using preset format specifications, production plan data, and relationships between characters to verify and correct the recognition results.
[0008] Optionally, a contextual logic judgment is performed on the preliminary comprehensive description of the character to obtain the final comprehensive description of the character; including: The probability of character visual features, the probability of applicability of verification algorithms, and the probability of MES matching are determined. The probability of character visual features is the probability that each character is recognized as a specific character in the image. The probability of applicability of verification algorithms is the probability of applicability of each verification algorithm. The probability of MES matching is the probability of matching each candidate batch number with the MES topology. By using character visual feature probabilities, verification algorithm applicability probabilities, and MES matching probabilities as input nodes, a multi-dimensional information fusion Bayesian inference network is constructed. Generate all logically possible candidate batch number sequences; Based on the Bayesian inference network, evaluate the posterior probability of each candidate batch number sequence; The recognition result with the highest posterior probability is selected as the final output to obtain the final comprehensive description of the character.
[0009] Optionally, determine the applicability probability of the verification algorithm, including: Extract unstructured information related to drug packaging to infer product classification intent; unstructured information includes one or more of the following: product name, production batch, and production line number; Determine the applicability probability of the verification algorithm based on the product classification intent.
[0010] Optionally, determine the MES matching probability, including: When the target expected batch number information is obtained, the target expected batch number information is deconstructed into a batch number topology structure containing known characters, placeholders and their possible value ranges; Narrow down the possible values of the placeholder based on the auxiliary information; The initial comprehensive description of the characters is matched with the batch number topology, and characters with low confidence and preset values are filled with possible values of placeholders. Calculate the matching probability between each candidate batch number and the batch number topology.
[0011] Optionally, the method also includes: Send data requests to multiple preset data sources, including manufacturing execution systems, enterprise resource planning systems, and laboratory information management systems; Configure a data source reliability parameter for each of the multiple data sources. The data source reliability parameter is preset based on the historical data accuracy, update frequency and response speed of the data source. Receive expected batch number information from multiple data sources; Check for any conflicts in the expected batch number information; When there are conflicts between expected batch number information, the conflicting expected batch number information is weighted according to the data source reliability parameters of each data source. Select the batch number information with the highest weighting value as the target expected batch number information.
[0012] Optionally, a contextual logic judgment is performed on the preliminary comprehensive description of the character to obtain the final comprehensive description of the character, including: For each character initially identified, a list containing multiple candidate characters is generated, and a visual similarity score is assigned to each candidate character; Based on the preset format specifications and production plan data, a context compatibility evaluation is performed on each candidate character in a list containing multiple candidate characters; Construct a character association path graph. The nodes of the character association path graph are the candidate characters for each character position, and the edges of the character association path graph represent the strength of the contextual association between characters. The overall confidence score of each character association path is calculated. The overall confidence score combines the visual similarity score, context compatibility score and the strength of association between characters for each character. The character association path with the highest overall confidence level is selected as the final comprehensive description of the character.
[0013] Optionally, when the overall confidence of multiple character association paths is the same, the method further includes: Multiple character association paths are used as candidate recognition results for differential feature extraction; Query historical production data and product quality control records related to the current batch to obtain historical production data and product quality control records; Based on the differentiated characteristics, historical production data, and product quality control records, the candidate identification results are evaluated a second time to obtain the candidate identification results after the second evaluation. The candidate recognition result with the highest priority after secondary evaluation is selected as the final comprehensive description of the character.
[0014] Optionally, determine the type and severity of defects present in the image information, including: Multi-scale decomposition of image information yields image components at different scales; Identify and extract texture features related to the microstructure of materials from image components at different scales; Frequency domain analysis of image information yields periodic patterns related to the material's microstructure. Based on texture features and periodic patterns, a description of the material's microstructure features is constructed; Defect features are extracted from image information to obtain the visual features of potential defects; Compare the visual characteristics of potential defects with the descriptions of the material's microstructure characteristics; When the visual features of a potential defect are highly similar to the description of the material's microstructure features, the feature stripping process is initiated. In the feature stripping process, the image information is subjected to adaptive filtering to suppress the influence of the material's microstructure; Defect regions are segmented from the filtered image information to obtain the actual defect regions. Based on the actual defect area, quantify the type and extent of the defect.
[0015] Optionally, after performing multi-scale decomposition on the image information to obtain image components of different scales, the method further includes: Energy leakage is detected for each scale component; When an energy leak is detected, the leaked energy is compensated in a targeted manner. Independence assessment of each scale component; Based on the independence assessment results, the components were orthogonalized.
[0016] Secondly, this application also discloses a smart identification system for drug packaging information, the system comprising: The image information acquisition module is used to acquire image information of the drug packaging. The character information extraction module is used to extract various types of character information from image information; these various types of character information include character stroke structure, character texture, and character geometric contour. The defect analysis module is used to analyze image information to obtain the types and degrees of defects present in the image information; The weight setting module is used to set the weight of various types of character information according to the defect type and defect severity. The character description integration module is used to integrate various types of character information with set weights to obtain a preliminary comprehensive description of the character; The context logic judgment module is used to perform context logic judgment on the preliminary comprehensive description of the character to obtain the final comprehensive description of the character. The context logic judgment uses preset format specifications, production plan data and the relationship between characters to verify and correct the recognition results.
[0017] Beneficial effects The intelligent identification method for pharmaceutical packaging information disclosed in this application acquires image information of pharmaceutical packaging and extracts various types of character information from the image information, including character stroke structure, character texture, and character geometric contour, achieving comprehensive capture of character features. Simultaneously, this method can determine the type and degree of defects in the image information and dynamically set the proportion of various types of character information according to the defect situation, enabling the identification process to adaptively address image quality issues. Subsequently, the multi-type character information with set proportions is integrated to obtain a preliminary comprehensive description of the character. Through contextual logic judgment, using preset format specifications, production plan data, and inter-character relationships, the identification result is verified and corrected, ultimately obtaining the final comprehensive description of the character. Through the above technical solution, this application effectively solves the challenges of image acquisition and information processing caused by the diversity of pharmaceutical packaging materials and dynamic changes in the production environment in existing technologies. It significantly improves the accuracy, robustness, and intelligence level of pharmaceutical packaging information identification, overcomes the shortcomings of traditional identification methods such as low efficiency and susceptibility to misreading, and meets increasingly stringent industry standards and regulatory requirements. Attached Figure Description
[0018] Figure 1 This is a schematic flowchart of a method for intelligent identification of drug packaging information provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of another intelligent identification method for drug packaging information provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of an intelligent identification system for drug packaging information provided in an embodiment of the present invention. Detailed Implementation
[0019] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments. The components of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0020] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0021] The following specific embodiments will provide a detailed introduction and explanation of the intelligent identification method for drug packaging information provided in this application.
[0022] Reference Figure 1 This invention provides a method for intelligent identification of pharmaceutical packaging information, comprising the following steps: S1, Obtain image information of the drug packaging.
[0023] Specifically, drug packaging can be scanned one by one by a handheld scanning device operated manually, or a fixed industrial camera can be set up on the production line and manually triggered by the operator to take pictures.
[0024] S2. Extract various types of character information from image information.
[0025] The various types of character information include character stroke structure, character texture, and character geometric outline.
[0026] Specifically, existing image processing libraries can be used for preliminary feature extraction. For example, edge detection algorithms can be used to obtain the geometric contours of characters, gray-level co-occurrence matrices can be used to analyze character texture, and morphological operations can be used to extract character stroke structures.
[0027] S3. Determine the type and degree of defects present in the image information.
[0028] Specifically, threshold-based image segmentation methods can be used to identify potential defect areas, and the degree of defect can be initially quantified by calculating the area or pixel intensity of the defect area. For example, a brightness threshold can be set, and areas below the threshold can be marked as stain defects.
[0029] S4. Set the weight of various types of character information according to the defect type and defect severity.
[0030] Specifically, a set of static weighting rules can be pre-defined. For example, when image blur is detected, the weighting of character texture information is reduced, while the weighting of character stroke structure is increased.
[0031] S5. Integrate the various types of character information with set weights to obtain a preliminary comprehensive description of the character.
[0032] Specifically, feature vectors of different types of character information can be concatenated and multiplied by their respective weight coefficients to form a preliminary feature description vector.
[0033] S6. Perform contextual logic judgment on the preliminary comprehensive description of the character to obtain the final comprehensive description of the character.
[0034] Among them, the context logic judgment uses preset format specifications, production plan data, and the relationship between characters to verify and correct the recognition results.
[0035] Specifically, a rule-based expert system can be built to verify the initial identification results based on preset batch number format rules (e.g., batch numbers must contain a specific combination of letters and numbers) and simple production plan queries (e.g., the range of current production batch numbers). Character association judgments can be limited to simple matching of adjacent characters. However, this judgment based on simple rules and limited associations may have limited corrective capabilities when dealing with highly ambiguous or multi-ambiguous characters.
[0036] Compared to existing methods that rely on single-feature recognition or fixed-rule verification, the intelligent identification method for pharmaceutical packaging information proposed in this application represents a significant advancement. Traditional methods often exhibit limitations in recognition accuracy and robustness when facing challenges such as the diversity of pharmaceutical packaging materials, unstable printing quality, and dynamic changes in the production environment. This application significantly enhances the feature extraction capability for complex images by acquiring and integrating various types of character information, including character stroke structure, character texture, and character geometric contours, and adaptively adjusting the weight of these information based on the type and degree of defects present in the image. Furthermore, by introducing a contextual logic judgment mechanism based on preset format specifications, production plan data, and inter-character relationships, this application can intelligently verify and correct the preliminary recognition results, effectively reducing the false recognition rate and improving the overall accuracy and reliability of recognition. Therefore, this application can better adapt to the complex needs of actual automated production lines, providing more accurate and efficient technical support for pharmaceutical quality control and traceability.
[0037] In some of the embodiments described above in this application, when performing contextual logic judgment on the preliminary comprehensive description of characters, relying solely on preset format specifications, production plan data, and simple verification and correction based on the relationships between characters may not adequately handle the inherent uncertainties and ambiguities in the recognition process. Especially when the confidence level of character recognition is low or multiple reasonable interpretations exist, it is difficult to systematically evaluate the reliability of different recognition results, thereby affecting the accuracy and robustness of the final recognition. To address this, this application further proposes a method for contextual logic judgment based on a Bayesian inference network using multi-dimensional information fusion, to more accurately evaluate and select the final recognition result.
[0038] like Figure 2 As shown, in order to perform contextual logic judgments on the preliminary comprehensive description of a character and obtain the final comprehensive description of the character, this application may further include the following steps: S101. Determine the probability of character visual features, the probability of applicability of the verification algorithm, and the probability of MES matching.
[0039] Among them, the character visual feature probability is the probability that each character is recognized as a specific character in the image; for example, a blurry character may be recognized as "B" with a probability of 0.6 and as "8" with a probability of 0.3.
[0040] The applicability probability of a verification algorithm is the applicability probability of each verification algorithm; for example, some batch numbers may contain specific check bits, while others may not, so it is necessary to evaluate the effectiveness of the corresponding verification algorithms.
[0041] The MES matching probability is the probability of matching each candidate batch number with the MES topology.
[0042] S102. Using the character visual feature probability, the applicability probability of the verification algorithm, and the MES matching probability as input nodes, construct a Bayesian inference network that integrates multi-dimensional information.
[0043] Bayesian inference networks can probabilistically fuse information from different sources to handle uncertainties in the recognition process. A Bayesian inference network is a graphical model that uses nodes to represent random variables and directed edges to represent conditional dependencies between variables, thereby calculating the posterior probability of each variable given observed evidence.
[0044] A Bayesian inference network is a directed acyclic graph where nodes represent random variables and directed edges represent conditional dependencies between variables. In this application, the network is constructed to fuse character visual feature probabilities, verification algorithm applicability probabilities, and MES matching probabilities to infer the correctness of candidate batch number sequences.
[0045] 1. Network node definition: Core Inference Node (Hidden Variable): We can define a node representing the "true state of the candidate batch sequence," for example, we call it "true batch state (S)." This node is usually binary, indicating "the candidate batch sequence is correct" or "the candidate batch sequence is incorrect." This is the variable we ultimately want to infer through the network.
[0046] Evidence Nodes (Observational Variables): Input nodes include "Character Visual Feature Probability (V)", "Verification Algorithm Applicability Probability (A)", and "MES Matching Probability (M)". These are observational evidences that we can obtain from image processing and system data. To simplify the definition of the conditional probability table, these probability values are usually discretized into several states, such as "high", "medium", "low" or "pass", "fail", etc.
[0047] 2. Network structure (inter-node connections): A common, simplified network structure assumes that, given the "true batch number state (S)," the three evidence nodes (V, A, M) are conditionally independent of each other. This structure is similar to the Naive Bayes classifier, but its fusion capability remains strong.
[0048] Point from the “Real Batch Number Status (S)” node to the “Character Visual Feature Probability (V)” node.
[0049] The node points from the “True Batch Number Status (S)” node to the “Verification Algorithm Applicability Probability (A)” node.
[0050] Point from the “Real Batch Status (S)” node to the “MES Matching Probability (M)” node.
[0051] This means that the true state of a batch number sequence will affect the probability of the visual features we observe, the applicability of the verification algorithm, and the MES matching probability.
[0052] 3. Definition of Conditional Probability Table (CPT): Prior probability P(S): This is the prior probability of the "true batch status (S)", that is, the probability that a candidate batch sequence is correct or incorrect before any observational evidence is available. For example, P(S=correct) and P(S=incorrect).
[0053] Conditional probability P(V | S): This is the probability of observing a specific "character visual feature probability (V)" given the "true batch number status (S)". For example: P(V=high | S=correct): If the batch number sequence is correct, the probability that its visual feature is high.
[0054] P(V=low|S=correct): The probability that the visual feature probability is low if the batch sequence is correct (e.g., poor image quality but the recognition result is correct).
[0055] P(V=high | S=error): The probability that the visual feature of a batch number sequence is high if the sequence is incorrect (e.g., a false alarm in visual recognition).
[0056] P(V=low | S=error): The probability that the visual feature of a batch sequence is low if the batch number sequence is erroneous.
[0057] Conditional probability P(A | S): Similar to P(V | S), it is defined as the probability of observing a specific "application probability (A)" given a "true batch number status (S)".
[0058] Conditional probability P(M | S): Similar to P(V | S), it is defined as the probability of observing a specific MES matching probability (M) given a "true batch number status (S)".
[0059] Parameter learning method: Parameter learning in Bayesian inference networks primarily involves determining the specific values of the prior probability P(S) and all conditional probability tables P(V|S), P(A|S), and P(M|S). The most commonly used method is supervised learning based on historical data.
[0060] 1. Data collection and preprocessing: We collected a large amount of historical drug packaging image data, ensuring that each image was accompanied by its corresponding authentic and accurate batch number information. This authentic batch number information will serve as "truth value" labels in the learning process.
[0061] For each image, the recognition process is simulated, and the corresponding observed values of "character visual feature probability (V)", "verification algorithm applicability probability (A)" and "MES matching probability (M)" are calculated.
[0062] Discretize these continuous probability values into several predefined states (e.g., "high", "medium", "low") to facilitate the construction of discrete CPT.
[0063] 2. Frequency statistics and parameter estimation: Learning P(S): P(S=correct) and P(S=incorrect) are estimated by counting the frequency of correct and incorrect batch number sequences in all batch number sequences in the statistical data set.
[0064] Learning P(V|S): Statistically analyze the frequency of V in the "high", "medium", and "low" states among all samples of genuine batch number sequences (S=correct) to estimate P(V=high|S=correct), P(V=medium|S=correct), and P(V=low|S=correct).
[0065] In the sample of all erroneous batch number sequences (S=erroneous), the frequency of V being in the "high", "medium", and "low" states is statistically analyzed to estimate P(V=high|S=erroneous), P(V=medium|S=erroneous), and P(V=low|S=erroneous).
[0066] Learning P(A | S) and P(M | S): Using the same method as learning P(V | S), we count the frequency of occurrence of each state of A and M in the two cases of S=correct and S=incorrect.
[0067] Smoothing: To avoid the problem of some combinations never appearing in the dataset, resulting in a probability of zero (e.g., P(V=low | S=correct) is 0 in the training data), techniques such as Laplace smoothing are usually used to assign a small non-zero probability to all possible combinations.
[0068] 3. Expert knowledge and experience: In situations with insufficient data or in certain specific scenarios, CPT can be initialized or adjusted by incorporating the knowledge of domain experts. Experts can directly provide estimates of certain conditional probabilities based on their experience, such as, "If the batch number is correct, then the probability of a visual recognition result is 95%."
[0069] For example: Suppose we have a candidate batch number sequence "XYZ789", and we want to use a Bayesian inference network to determine whether it is a genuine batch number. We have already obtained the following simplified CPT (V, A, M are discretized as "high" and "low") through parameter learning: Prior probability: P(S=correct) = 0.9 (assuming most recognition results are correct); P(S=Error) = 0.1; Conditional probability table P(V|S): P(V=high | S=correct) = 0.95; P(V=low | S=correct) = 0.05; P(V=high | S=error) = 0.20 (error identification may also look visually similar); P(V=low | S=error) = 0.80; Conditional probability table P(A|S): P(A=High | S=Correct) = 0.98; P(A=low | S=correct) = 0.02; P(A=high | S=error) = 0.10; P(A=low | S=error) = 0.90; Conditional probability table P(M | S): P(M=High | S=Correct) = 0.90; P(M=low | S=correct) = 0.10; P(M=high | S=error) = 0.15; P(M=low | S=error) = 0.85; Now, we identify the candidate batch number sequence “XYZ789” and obtain the following observational evidence: Character visual feature probability (V) = High; The applicability probability (A) of the verification algorithm is high; MES matching probability (M) = low (e.g., the batch number information in the MES system has not been fully updated). We need to calculate P(S=correct | V=high, A=high, M=low) and P(S=incorrect | V=high, A=high, M=low).
[0070] Based on Bayes' theorem and the assumption of conditional independence: P(S=correct| V, A, M) = [P(V | S=correct) * P(A | S=correct) * P(M | S=correct)* P(S=correct)] / P(V, A, M); P(S=Error|V, A, M) = [P(V | S=Error) * P(A | S=Error) * P(M | S=Error)* P(S=Error)] / P(V, A, M); First, calculate the molecular part: Molecular correctness = P(V=high | S=correct) * P(A=high | S=correct) * P(M=low | S=correct) * P(S=correct) = 0.95 * 0.98 * 0.10 * 0.9 = 0.08379; Molecular error = P(V=high | S=error) * P(A=high | S=error) * P(M=low | S=error) * P(S=error) = 0.20 * 0.10 * 0.85 * 0.1 = 0.0017; Then calculate the denominator P(V, A, M) (normalization constant): P(V, A, M) = Numerator_Correct + Numerator_Incorrect = 0.08379 + 0.0017 = 0.08549; Finally, the posterior probability was calculated as follows: P(S=correct | V=high, A=high, M=low) = 0.08379 / 0.08549 ≈0.980; P(S=Error|V=High, A=High, M=Low) = 0.0017 / 0.08549 ≈ 0.020; In this example, although the MES matching probability is low, the probability of visual features and the applicability probability of the verification algorithm are both very high. After comprehensive judgment, the Bayesian inference network concludes that the posterior probability of the candidate batch number sequence "XYZ789" being the real batch number is as high as 98.0%. This shows that the network can effectively integrate multi-dimensional information and make more robust and accurate judgments even when there is uncertainty or conflict in a certain information source.
[0071] S103. Generate all logically possible candidate batch number sequences.
[0072] In practical applications, after constructing the Bayesian inference network, all logically possible candidate batch number sequences are generated. These sequences are all potential, logically consistent batch number combinations derived from a preliminary comprehensive description of the characters and predefined format specifications and character association rules.
[0073] S104. Based on the Bayesian inference network, evaluate the posterior probability of each candidate batch number sequence.
[0074] Specifically, Bayesian inference networks can probabilistically fuse information from different sources to handle uncertainties in the recognition process. A Bayesian inference network is a graphical model that uses nodes to represent random variables and directed edges to represent conditional dependencies between variables, thereby calculating the posterior probability of each variable given observed evidence.
[0075] S105. Select the recognition result with the highest posterior probability as the final output to obtain the final comprehensive description of the character.
[0076] Specifically, the posterior probability of each candidate batch number sequence can be evaluated based on the Bayesian inference network described above. The posterior probability reflects the likelihood that the sequence is a genuine batch number after considering all input information (visual features, applicability of the verification algorithm, MES matching). Finally, the recognition result with the highest posterior probability is selected as the final output, thus obtaining the final comprehensive description of the character.
[0077] This application's solution systematically addresses the limitations of traditional contextual logic judgments in handling uncertainty and ambiguity by introducing a Bayesian inference network that integrates multi-dimensional information. Specifically, the character visual feature probability provides the confidence level of the character itself; the applicability probability of the verification algorithm provides additional verification information from the perspective of data integrity and format standardization; and the MES matching probability associates the recognition result with actual production data, ensuring the correctness of the business logic of the recognition result. By using these probabilistic information from different dimensions and sources as input nodes, the Bayesian inference network can comprehensively evaluate all possible candidate batch number sequences within a rigorous probabilistic framework. By calculating the posterior probability of each sequence, the network can quantitatively reflect the authenticity of each candidate sequence, thus enabling it to make the most reasonable judgment even when there is ambiguous or conflicting information. This method avoids misjudgments that may be caused by simple rule-based judgments, improving the accuracy and robustness of recognition.
[0078] Through the above technical solution, this application overcomes the shortcomings of traditional contextual logic judgment in handling complex and uncertain recognition scenarios. By fusing character visual feature probabilities, verification algorithm applicability probabilities, and MES matching probabilities in multiple dimensions, and using a Bayesian inference network for systematic evaluation, the accuracy and reliability of the recognition results are significantly improved. This solution can effectively handle the ambiguity, uncertainty, and multiple potential reasonable interpretations in character recognition, thus enabling it to output highly reliable pharmaceutical packaging information recognition results even when facing challenges such as poor printing quality, character deformation, or partial missing parts. This greatly reduces the need for manual review and improves the efficiency and quality of automated recognition.
[0079] In some preferred embodiments, it is assumed that when identifying a drug batch number, a certain character is initially identified as both "B" and "8" with a high probability.
[0080] First, the system determines the probability of the character's visual features. For example, the visual recognition module determines that the probability of the character being "B" is 0.55, and the probability of it being "8" is 0.45.
[0081] Secondly, the system determines the applicability probability of the verification algorithm. Assuming that, based on the product classification intent, the batch number should follow a format containing a specific check digit, and that this verification algorithm has high applicability in distinguishing between "B" and "8", for example, if the batch number contains "B", the probability of passing the verification is 0.9, and if it contains "8", the probability of passing the verification is 0.1.
[0082] Next, the system determines the MES matching probability. By querying the MES system, it is found that the probability of a batch number sequence containing "B" matching the actual production data in the current production plan is 0.8, while the probability of a batch number sequence containing "8" matching the actual production data is 0.2.
[0083] Subsequently, these probabilities are used as input nodes to construct a Bayesian inference network. This network comprehensively considers these probabilities to generate all logically possible candidate batch number sequences. For example, if the batch number is "ABC123B", one candidate sequence is "ABC123B", and another could be "ABC1238".
[0084] Next, the Bayesian inference network evaluates the posterior probability of each candidate batch number sequence. For example, the network calculates that the posterior probability of the sequence "ABC123B" might be 0.75, while the posterior probability of the sequence "ABC1238" might be 0.20.
[0085] Ultimately, because the sequence "ABC123B" has the highest posterior probability, the system selects it as the final recognition result, thus obtaining the final comprehensive description of the character. In this way, even if the visual recognition of a single character is ambiguous, the system can make a more accurate and reliable judgment by fusing multi-dimensional information.
[0086] In some embodiments described above, this application proposes constructing a Bayesian inference network that fuses multi-dimensional information by determining the probability of character visual features, the applicability probability of verification algorithms, and the MES matching probability as input nodes, in order to perform contextual logic judgment on the preliminary comprehensive description of characters. However, in practical applications, the applicability of verification algorithms is often closely related to factors such as the specific type of drug product, production batch, or production line. If this contextual information is not fully considered when determining the applicability probability of the verification algorithm, the weight allocation of the Bayesian inference network for different verification algorithms may be inaccurate, thus affecting the accuracy and reliability of the final recognition result.
[0087] In this regard, this application further proposes steps for determining the applicability probability of the above-mentioned verification algorithm, including: S201. Extract unstructured information related to drug packaging to infer product classification intent.
[0088] The unstructured information includes one or more of the following: product name, production batch, and production line number.
[0089] Specifically, unstructured information can be identified by optical character recognition (OCR) technology to identify non-critical areas of packaging images, or obtained from images through preset template matching, keyword recognition, and other methods.
[0090] Inferring product classification intent refers to identifying the product category or specific production context corresponding to the current drug packaging based on extracted unstructured information. For example, identifying the product name "Amoxicillin Capsules" can infer that it belongs to the antibiotic category; identifying the production line number "L001" can infer that it may follow the verification rules of a specific production line. Inferring product classification intent can utilize machine learning models, such as text classifiers, to analyze and categorize the extracted unstructured information, or it can be matched using a predefined rule base. The goal is to associate the current identification task with a specific product or production environment in order to select the most appropriate verification strategy.
[0091] S202. Determine the applicability probability of the verification algorithm based on the product classification intent.
[0092] In practical applications, determining the applicability probability of a verification algorithm based on the product classification intent means that after inferring the product classification intent, the system assigns a corresponding applicability probability to different verification algorithms based on a pre-established knowledge base or rule set. For example, for a specific product category (such as vaccines), there may be a strict set of batch number verification rules, in which case the applicability probability of the verification algorithm corresponding to that rule will be significantly increased; while for another product category (such as health supplements), its batch number rules may be relatively lenient, and the applicability probability of the relevant verification algorithm will be adjusted accordingly. This determination method makes the selection of verification algorithms more intelligent and context-based, avoiding a "one-size-fits-all" verification strategy, thereby improving the accuracy and robustness of identification.
[0093] The proposed solution first extracts unstructured information related to the drug packaging, such as product name, production batch number, or production line number, thereby obtaining rich contextual information about the current object to be identified. Based on this unstructured information, the system can infer the specific product classification intent, such as identifying which product category the drug belongs to and which production line it was produced on. It is precisely because of this detailed product classification intent that the system can dynamically and specifically determine the applicability probability of each verification algorithm according to the characteristics of different products or production environments. For example, some products may have strict verification rules for specific characters in the batch number, while others may have special requirements for the production date format. In this way, the determination of the applicability probability of the verification algorithm is no longer static or generalized, but can accurately reflect the actual needs of the current identification task and the scope of application of the verification rules, thus providing more accurate and contextualized input for the subsequent Bayesian inference network, effectively solving the problem of insufficient generalization in determining the applicability probability of the verification algorithm in the basic solution.
[0094] In some preferred embodiments, a specific example is given below. Suppose it is necessary to identify batch number information on a drug package. First, the system acquires the image information of the drug package. While extracting character information, the system also extracts unstructured information related to the drug package. For example, through optical character recognition (OCR), the system identifies the product name on the package as "Cold Relief Granules," the production batch number as "20230101," and the production line number as "Line A."
[0095] Based on this unstructured information, the system infers that the product classification intent is "Traditional Chinese Medicine Cold Remedy, produced by Line A". According to the preset knowledge base, the system understands that the batch numbers of "Traditional Chinese Medicine Cold Remedy" products usually follow the format of "year + month + date + serial number", and that products produced by "Line A" may use a specific checksum algorithm.
[0096] Therefore, the system dynamically adjusts the applicability probabilities of various verification algorithms based on the product classification intent. For example, the applicability probability of an algorithm verifying the format "year + month + date + serial number" is set to 0.95; the applicability probability of an algorithm verifying the specific check code "A line" is set to 0.90; and the applicability probability of other irrelevant verification algorithms (such as the Western medicine batch number verification algorithm) is significantly reduced, for example, set to 0.10. These context-adjusted applicability probabilities of the verification algorithms are then input into the Bayesian inference network and fused with other probability information to more accurately evaluate the posterior probability of the candidate batch number sequence, ultimately obtaining a more reliable recognition result.
[0097] The determination of the MES matching probability mentioned above includes: S301. When the target expected batch number information is obtained, the target expected batch number information is deconstructed into a batch number topology structure containing known characters, placeholders and their possible value ranges.
[0098] The target expected batch number information refers to the expected batch number data about the current production batch, obtained from external systems (such as manufacturing execution systems, enterprise resource planning systems, etc.). After obtaining this information, it needs to be structured, decomposing it into a fixed, known character part and a variable placeholder part. For example, a batch number format might be "production date-batch number", where the production date is a known character and the batch number is a placeholder. The possible value range of the placeholder refers to the set of characters or numerical range that the placeholder might appear in actual production; for example, the batch number might be limited to the digits 0-9 or a specific letter combination.
[0099] S302. Narrow down the possible value range of the placeholder based on auxiliary information.
[0100] Auxiliary information can be understood as additional data related to the current production batch or product, such as production line configuration, product type, production date, shift information, etc. By utilizing this auxiliary information, the value range of placeholders can be further limited, thereby reducing the complexity and error rate of matching. For example, if the auxiliary information indicates that the current production is for a specific product, and its batch number typically contains only specific letters, then the possible values of the placeholders can be narrowed down to these specific letters.
[0101] The following will explain in detail how to narrow down the data based on this auxiliary information, and provide examples: 1. Narrow down the range of placeholders based on production line configuration: Different production lines may have specific encoding rules or character sets for certain parts of the batch number. For example, some production lines may only use specific printing equipment that may only support a limited character set, or may enforce the use of specific identifiers at specific locations in the batch number.
[0102] Specific explanation: If the batch number topology contains a placeholder representing a production line identifier, and the auxiliary information explicitly indicates that the current product was produced on "Production Line A", and if the configuration of "Production Line A" specifies that the production line identifier in its batch number can only be "L1" or "L2", then the possible value range of this placeholder can be narrowed from all possible production line identifiers (e.g., L1, L2, L3, L4, etc.) to only include "L1" and "L2".
[0103] For example: Suppose the batch number format is "YYMMDD-LINE-SEQ", where "LINE" is a production line identifier placeholder. If the auxiliary information shows that the current product comes from "Production line configuration: high-speed filling line", and the identifiers used by this filling line in the batch number are limited to "GZ1" and "GZ2", then the possible value range of the "LINE" placeholder is narrowed down from all possible production line identifiers (such as GZ1, GZ2, PK1, PK2, etc.) to {GZ1, GZ2}.
[0104] 2. Narrow down the range of placeholders based on product type: Different types of products (such as tablets, injections, oral liquids, etc.) often have their own unique batch number coding specifications or specific character requirements. These specifications can help limit the values of certain placeholders in the batch number.
[0105] Specific explanation: If the batch number topology contains a placeholder indicating the product category or dosage form, and the auxiliary information explicitly states that the current product is "tablets". If all tablet products use "P" as the dosage form identifier in a specific position in the batch number, then the possible values of this placeholder can be narrowed down from all possible dosage form identifiers (such as P, Z, K, etc.) to only containing "P".
[0106] For example, suppose the batch number format is "PRODTYPE-YYMMDD-BATCH", where "PRODTYPE" is a product type placeholder. If the auxiliary information shows "Product Type: Injectable", and it is known that all injectable batch numbers begin with "INJ", then the possible values for the "PRODTYPE" placeholder can be narrowed down from all product type codes (such as TAB, CAP, INJ, etc.) to {INJ}.
[0107] 3. Narrowing the range of placeholders based on the production date: Batch numbers usually contain the production date or information derived therefrom (such as year, month, abbreviation of date, Julian day, etc.). When auxiliary information provides the exact production date, the date-related placeholders in the batch number can be precisely limited.
[0108] Specific explanation: If the batch number topology contains a placeholder representing the production date (e.g., "YYMMDD"), and the auxiliary information provides "Production Date: August 15, 2023", then the possible value range of the "YYMMDD" placeholder can be directly narrowed down to "230815". If the batch number contains a placeholder representing the production week, and the production date is known to be August 15, the corresponding production week can be calculated, thus narrowing down the range of that placeholder.
[0109] For example: Suppose the batch number format is “YEAR-WEEK-SEQ”, where “YEAR” and “WEEK” are placeholders for the year and week. If the auxiliary information shows “Production Date: October 26, 2023”, then the “YEAR” placeholder can be reduced to {23}, and the “WEEK” placeholder can be reduced to {43} (week 43 of 2023).
[0110] 4. Narrow down the placeholder range based on shift information. In some production scenarios, batch numbers include shift identifiers to distinguish products produced in different shifts. When auxiliary information provides the current production shift, the value of the shift placeholder in the batch number can be narrowed down.
[0111] Specific explanation: If the batch number topology contains a placeholder representing a shift (e.g., "S"), and the auxiliary information provides "Shift Information: Night Shift", then if it is known that night shifts are usually represented by the character "N", then the possible values of this placeholder can be narrowed down from all shift identifiers (e.g., A, B, N, etc.) to only containing "N".
[0112] For example: Suppose the batch number format is “YYMMDD-SHIFT-BATCH”, where “SHIFT” is a shift placeholder. If the auxiliary information displays “Shift Information: Morning Shift”, and it is known that the batch number sequence number for the morning shift only uses the numbers 0-4, then the possible value range of the “BATCH” placeholder can be narrowed down from 0-9 to 0-4.
[0113] By using the above methods and various auxiliary information to accurately reduce the placeholders in the batch number topology, the search space and uncertainty in the batch number matching process can be significantly reduced, thereby improving the accuracy of completing characters with low confidence and ultimately improving the calculation accuracy of MES matching probability and the overall reliability of intelligent identification of drug packaging information.
[0114] S303. Match the preliminary comprehensive description of the characters with the batch number topology, and fill in the possible values of placeholders for characters with low confidence and preset values.
[0115] Specifically, the initial comprehensive description of the characters is the raw character recognition result output by the image recognition module, which may contain some characters with low confidence. During the matching process, the initially recognized character sequence is compared with the deconstructed batch number topology. For character positions where the recognition confidence is below a preset threshold, the system will attempt to intelligently complete the character using the possible value range of placeholders. For example, if a character is recognized as "B" with a very low confidence, and the placeholder in the batch number topology at that position may have a value of "8" or "B", a more reasonable inference and completion can be made based on the context and the placeholder range.
[0116] S304. Calculate the matching probability between each candidate batch number and the batch number topology.
[0117] After completing character completion and matching, it is necessary to quantify the degree of conformity between each possible candidate batch number sequence and the batch number topology. The calculation of this matching probability can comprehensively consider the visual similarity of characters, the reasonableness of placeholder completion, and the overall sequence's fit with the topology. For example, the final matching probability can be obtained by calculating a weighted average of character-level matching scores and structure-level matching scores.
[0118] This application's solution introduces target expected batch number information and deconstructs it into a batch number topology, providing a clear structured reference for calculating MES matching probabilities. When the target expected batch number information is obtained, it is decomposed into known characters and placeholders, enabling the system to distinguish between the fixed and variable parts of the batch number. Furthermore, by utilizing auxiliary information to narrow down the possible values of the placeholders, uncertainty during the matching process is effectively reduced, improving matching accuracy. When matching the preliminary comprehensive description of characters with the batch number topology, for characters with low recognition confidence, intelligent completion using the possible values of placeholders can be performed, thus compensating for potential deficiencies in image recognition and enhancing the ability to recognize blurred or damaged characters. Finally, by calculating the matching probability of each candidate batch number with the batch number topology, a more accurate and reliable MES matching input is provided to the Bayesian inference network, thereby improving the overall accuracy of contextual logic judgment.
[0119] Through the above technical solution, this application can significantly improve the accuracy and robustness of MES matching probability determination. Specifically, by deconstructing the target expected batch number information and using auxiliary information to narrow down the placeholder range, the system can more effectively correct and complete errors when faced with incomplete or low-confidence character recognition results, thereby reducing the misjudgment rate caused by image recognition errors. Furthermore, this method makes the calculation of MES matching probability more refined and intelligent, providing high-quality input for subsequent Bayesian inference networks, thus improving the overall accuracy and reliability of intelligent recognition of pharmaceutical packaging information. Its advantages are particularly evident when dealing with packaging information with complex batch number formats or local defects.
[0120] In some preferred embodiments, suppose the batch number on a drug package is “20230815-A01”, but the image recognition result is blurry in the “A01” part, and it is initially identified as “20230815-?01”, where the confidence level of the “?” character is low.
[0121] First, the system obtains the target expected batch number information. For example, the expected batch number format for the current production batch obtained from the MES system is "YYYYMMDD-XXX", where "YYYYMMDD" is a known string, and "XXX" is a placeholder representing a three-letter alphanumeric combination. At this point, the target expected batch number information is deconstructed into a batch number topology structure, containing the known string "20230815" and the placeholder "XXX", and the possible values of the placeholder "XXX" are all three-letter alphanumeric combinations.
[0122] Furthermore, based on auxiliary information, such as the product type on the current production line, the system determines that the third character of the product batch number is usually the uppercase letter AZ. Therefore, the possible values for the placeholder "XXX" are narrowed down to "XX[AZ]".
[0123] Next, the system matches the initially identified "20230815-?01" with the batch number topology "20230815-XXX". Since the confidence level of the "?" character is below the preset threshold, the system attempts to complete the sequence using the possible values of the placeholder "XXX". Considering that the auxiliary information restricts the third character to uppercase, the system prioritizes completing the "?" with an uppercase letter.
[0124] Finally, the system calculates the matching probability of each candidate batch number (e.g., "20230815-A01", "20230815-B01", etc.) with the batch number topology. For example, if "20230815-A01" has the highest matching degree with the preliminary recognition result and fits the narrowed range of placeholders, its matching probability will be evaluated as the highest and input into the Bayesian inference network as the final result of the MES matching probability. In this way, even when there is uncertainty in character recognition, intelligent error correction and completion can be performed using structured information and auxiliary information, thereby improving the accuracy of recognition.
[0125] The method also includes: S401, Send data requests to multiple preset data sources.
[0126] These data sources can include Manufacturing Execution Systems (MES), Enterprise Resource Planning (ERP) systems, and Laboratory Information Management (LIM) systems. MES typically contain real-time production batch information and production progress; ERP systems may provide more macro-level production planning and material management information; and LIM systems may store batch-related quality inspection data. By requesting information from these different data sources, more comprehensive and multi-dimensional expected batch number information can be obtained.
[0127] S402. Configure a data source reliability parameter for each of the multiple data sources.
[0128] Among them, the data source reliability parameters are preset based on the historical data accuracy, update frequency and response speed of the data source.
[0129] The reliability parameter of this data source is preset based on metrics such as the accuracy, update frequency, and response speed of the data source's historical data. For example, a data source with high historical data accuracy, fast update frequency, and fast response speed will have its reliability parameter set to a higher value, and vice versa. This parameter aims to quantify the confidence level of the information provided by different data sources.
[0130] S403: Receive expected batch number information from multiple data sources.
[0131] S404. Check for any conflicts in the expected batch number information.
[0132] For example, if the batch number provided by the Manufacturing Execution System is "A123" while the batch number provided by the Enterprise Resource Planning System is "B456", then a conflict is considered to exist.
[0133] S405. When there are conflicts between expected batch number information, the conflicting expected batch number information shall be weighted according to the data source reliability parameters of each data source.
[0134] Specifically, for each conflicting batch number, its weight will be determined by the reliability parameter of its source data source. For example, if the reliability parameter of the Manufacturing Execution System is 0.9, the reliability parameter of the Enterprise Resource Planning System is 0.7, and the reliability parameter of the Laboratory Information Management System is 0.8, when they provide different batch number information, the system will weight the respective batch number information according to these parameters.
[0135] S406. Select the batch number information with the highest weighted value as the target expected batch number information.
[0136] This application's solution effectively addresses the accuracy and reliability issues in determining the target expected batch number information when assessing MES matching probability by introducing multi-data source information acquisition, data source reliability parameter configuration, and a weighted processing mechanism for conflicting expected batch number information. Specifically, by sending data requests to multiple preset data sources such as the Manufacturing Execution System (MES), Enterprise Resource Planning (ERP) system, and Laboratory Information Management System (LIMS), comprehensive expected batch number information related to pharmaceutical packaging can be collected, avoiding the limitations or errors that may exist with a single data source. Furthermore, the data source reliability parameters configured for each data source objectively quantify the confidence levels of different information sources, providing a basis for subsequent conflict resolution. When the received expected batch number information conflicts, the system does not simply select randomly or report an error; instead, it weights the conflicting expected batch number information according to the preset data source reliability parameters, giving higher weight to information from more reliable data sources. Thus, the batch number information with the highest weighted value is ultimately selected as the target expected batch number information, ensuring the accuracy and authority of the acquired batch number information in a complex and ever-changing information environment. This provides a solid foundation for subsequent batch number topology deconstruction, matching, and MES matching probability calculation.
[0137] Through the above technical solution, this application can significantly improve the accuracy and reliability of obtaining the target batch number information. Compared with solutions that rely solely on a single data source or fail to effectively handle information conflicts, this application effectively avoids batch number identification errors caused by the singleness of the data source or information conflicts by integrating multi-source information and introducing a weighted processing mechanism based on the reliability of the data source. This not only improves the calculation accuracy of the MES matching probability but also enhances the robustness of the entire intelligent identification method for drug packaging information. It enables the method to stably and accurately identify drug packaging information even in the face of complex and ever-changing production data environments, thereby effectively reducing the false identification rate and improving the efficiency and safety of drug production quality control.
[0138] In some preferred embodiments, suppose the batch number of a certain batch of medicine needs to be identified. The system first sends data requests to the Manufacturing Execution System (MES), Enterprise Resource Planning (ERP) System, and Laboratory Information Management System (LIMS). The MES returns batch number "20230815A", the ERP System returns batch number "20230815B", and the LIMS also returns batch number "20230815A". At this point, the system detects that the batch numbers provided by the MES and LIMS are consistent, but conflict with the batch number provided by the ERP System. Based on pre-set data source reliability parameters, assuming the reliability parameter of the MES is 0.9, the reliability parameter of the ERP is 0.7, and the reliability parameter of the LIMS is 0.8, the system will perform weighted processing on the conflicting expected batch number information "20230815A" and "20230815B". The weighted value of batch number "20230815A" is (0.9 + 0.8) = 1.7 (because both the Manufacturing Execution System and the Laboratory Information Management System provide this information), while the weighted value of batch number "20230815B" is 0.7. Since "20230815A" has the highest weighted value, the system ultimately selects "20230815A" as the target expected batch number. In this way, even in the event of data conflicts, the system can intelligently judge and select based on the reliability of the data source, ensuring the accuracy of the batch number information.
[0139] In some embodiments described above, this application proposes a scheme to adjust the set of nonlinear relation functions to escape local optima when the optimization process gets stuck in a local state. However, in actual implementation, local states may take various forms, such as saddle point regions or flat regions. If these different types of local states are not distinguished and targeted adjustment strategies are not adopted, the adjustment efficiency may be low, or even unable to effectively escape local optima, thus affecting the efficiency and accuracy of finding the global optimum. Therefore, this application further proposes a more refined method for adjusting the set of nonlinear relation functions, which identifies the specific type of local state and adopts differentiated strategies to optimize the optimization process.
[0140] The above preliminary comprehensive description of the character is subjected to contextual logic judgment to obtain the final comprehensive description of the character, specifically including: S501. For each character initially identified, generate a list containing multiple candidate characters and assign a visual similarity score to each candidate character.
[0141] Specifically, for each initially identified character, a list containing multiple candidate characters is generated. For example, if a character is initially identified as "B", but its visual features are also highly similar to "8", then the list may contain both "B" and "8". Simultaneously, each candidate character in the list is assigned a visual similarity score, which reflects the degree of visual matching between the candidate character and the actual character region in the image. This score can be obtained, for example, through pixel-level comparison or feature vector distance calculation.
[0142] S502. Based on the preset format specifications and production plan data, perform a context compatibility evaluation on each candidate character in a list containing multiple candidate characters.
[0143] The formatting specifications can include rules regarding character type (e.g., numbers, letters, special symbols), length, and arrangement order. Production planning data can provide information such as the expected batch number, production date, and expiration date of the current batch of products, thereby limiting the possible range of character values. This information allows for the evaluation of whether each candidate character conforms to the overall contextual logic in a specific position; for example, letter candidate characters would have a lower compatibility score in the numeric portion of the batch number.
[0144] S503. Construct a character association path graph.
[0145] In this graph, the nodes are the candidate characters for each character position, and the edges represent the strength of the contextual association between characters.
[0146] For example, if a batch number has five character positions, and each character position has several candidate characters, then these candidate characters constitute the nodes of the graph. The edges of this path graph represent the strength of the contextual association between characters, which can be determined, for example, based on the statistical co-occurrence frequency of characters, language models, or predefined grammatical rules. For example, "EXP" is usually followed by a date, so the association strength between "P" and date numbers would be relatively high.
[0147] The character association path graph is essentially a directed acyclic graph (DAG), also known as a "Trellis graph," which visualizes the candidate characters at each position in the initially identified character sequence and their interrelationships. Its construction process can be divided into the following steps: Identify nodes (candidate characters at character positions): For each character position in the initially identified drug packaging information, the system generates a list containing multiple candidate characters. For example, if a character at a certain position is blurry in the image and might be identified as "B" or "8", then "B" and "8" are two candidate characters for that position.
[0148] At the same time, a visual similarity score is assigned to each candidate character. This score reflects the degree of visual matching between the candidate character and the corresponding region in the image, and is usually given by the underlying image recognition algorithm (such as an OCR engine). The higher the score, the better the visual matching.
[0149] In the path graph, each candidate character at each character position is treated as an independent node.
[0150] 2. Determine the edges (contextual relationships between characters): The edges in a path graph connect candidate characters at adjacent character positions. For example, if the candidate characters at the first character position are "A" and "C", and the candidate characters at the second character position are "B" and "D", then there will be edges from "A" to "B", "A" to "D", "C" to "B", and "C" to "D".
[0151] These edges represent the contextual relationships between characters. Each edge is assigned a quantified value of "inter-character association strength," which reflects the plausibility or co-occurrence probability of the candidate characters of the preceding and following characters in the context.
[0152] By following the steps above, a character association path graph is constructed from the start character position to the end character position. In this graph, each complete path from the start node to the end node represents a logically possible character sequence.
[0153] S504. Calculate the overall confidence level of each character association path.
[0154] The overall confidence score integrates the visual similarity score, contextual compatibility score, and inter-character association strength for each character.
[0155] For example, it can be calculated using weighted summation, Bayesian inference, or other multi-source information fusion algorithms. This fusion method ensures that the recognition result is not only visually matching, but also logically and contextually plausible.
[0156] Character association strength is a quantitative indicator that measures the probability of two adjacent candidate characters co-occurring in a specific context. It integrates various information such as pre-defined formatting specifications, production plan data, and statistical correlations between characters. Specific quantitative calculation methods can include the following: 1. Association strength based on preset format specifications: Drug batch numbers, expiration dates, and other information usually follow strict format specifications, such as "first two letters + last four numbers" or "date format is YYYYMMDD".
[0157] If two adjacent candidate character combinations conform to a preset format specification, a higher association strength is assigned. For example, in the "alphanumeric" format, the association strength of an alphanumeric candidate character followed by a letter candidate character is higher than that of an alphanumeric candidate character followed by another letter candidate character.
[0158] Quantification method: Rules can be set, and combinations that meet the rules will receive a high score (e.g., 0.9), while combinations that do not meet the rules will receive a low score (e.g., 0.1).
[0159] 2. Correlation strength based on production planning data: Production planning data provides expected information for the current batch or product line. For example, the current batch number may be limited to a certain range, or the characters at a certain position must be specific letters or numbers.
[0160] A combination of two adjacent candidate characters is assigned a very high association strength if it closely matches a known pattern or expected sequence in the production plan data. For example, if the production plan explicitly states that the third character of the batch number is "X", then any combination with "X" as the third candidate character will receive a higher association strength.
[0161] Quantification method: You can set the highest score (e.g., 0.99) for an exact match, a medium score (e.g., 0.7) for a partial match or within the allowed range, and a low score (e.g., 0.05) for a no match.
[0162] 3. Strength based on character statistical association (N-gram model): By analyzing a large amount of historical and valid drug packaging information data, the co-occurrence frequency or probability of different character pairs (bigrams) or character sequences (trigrams, etc., N-grams) can be statistically determined.
[0163] If two adjacent candidate characters frequently appear together in historical data, they are assigned a higher association strength. For example, in English batch numbers, the co-occurrence frequency of "ST" may be much higher than that of "SX", so the association strength of "S" followed by "T" is higher than that of "S" followed by "X".
[0164] Quantification method: N-gram probability can be used directly as the association strength, or it can be normalized before use.
[0165] 4. Comprehensive quantitative calculation: The association strengths from the different sources mentioned above can be fused to obtain the final "inter-character association strength". Common fusion methods include weighted average, product, or Bayesian fusion.
[0166] For example, a weighted product approach can be used: Association Strength = (Formatting Standards Score) * (Production Planning Score) * (Statistical Association Score). Each score can be weighted according to its importance.
[0167] Suppose we want to identify a drug batch number, whose format is "two letters + four numbers", and the production plan data indicates that the current batch number should start with "AB". The preliminary identification results show candidates "A (0.9), C (0.1)" in the first character position, "B (0.8), D (0.2)" in the second character position, and "1 (0.95), 7 (0.05)" in the third character position.
[0168] 1. Build nodes: Location 1: Node (A, 0.9), Node (C, 0.1); Location 2: Node (B, 0.8), Node (D, 0.2); Location 3: Node (1, 0.95), Node (7, 0.05); 2. Calculate the strength of the association between characters and construct edges: Formatting guidelines: The first two digits of the batch number should be letters.
[0169] Letter-letter combinations (such as AB, AD, CB, CD) that conform to the format will receive a high base score.
[0170] Letter-number combinations (such as B1, B7, D1, D7) that conform to the format will receive a high base score.
[0171] Production planning data: Batch numbers should begin with "AB". The combination of "A" followed by "B" has a very strong correlation. The combination of "C" followed by "B" has a weaker correlation.
[0172] Statistical association: It is assumed that in historical data, "AB" has a high co-occurrence frequency, "AD" has a low frequency, and "CB" and "CD" have even lower frequencies. Among the numbers, "12" is more common than "17".
[0173] Quantification example (simplified): Edge (position 1, A) -> (position 2, B): Formatting guidelines: Letter-letter, high score (0.9); Production plan: Conforms to "AB" format, extremely high score (0.99); Statistical association: High frequency co-occurrence, high score (0.95); Overall correlation strength = 0.9 * 0.99 * 0.95 ≈ 0.85; Edge (position 1, A) -> (position 2, D): Formatting guidelines: Letter-letter, high score (0.9); Production plan: Does not conform to the "AB" prefix, low score (0.1); Statistical association: Low-frequency co-occurrence, low score (0.2); Overall correlation strength = 0.9 * 0.1 * 0.2 ≈ 0.018; Edge (position 2, B) -> (position 3, 1): Formatting guidelines: Letters-numbers, high score (0.9); Production plan: No specific restrictions, medium score (0.8); Statistical correlation: common combinations, high score (0.9); Overall association strength = 0.9 * 0.8 * 0.9 ≈ 0.648; By calculating the overall association strength of all possible adjacent candidate character pairs and using it as the edge weights, a complete character association path graph is constructed. Subsequent steps will search this graph for the path with the highest overall confidence, thus obtaining the final recognition result.
[0174] S505. Select the character association path with the highest overall confidence as the final comprehensive description of the character.
[0175] This application's solution effectively addresses the limitations of traditional methods in handling ambiguous or vague characters by constructing a multi-dimensional evaluation system and path optimization mechanism. First, by generating a candidate list and assigning visual similarity scores to each initially identified character, it ensures that potential correct characters are not overlooked even if there is uncertainty in the initial recognition. Second, by introducing pre-defined format specifications and production plan data for contextual compatibility evaluation, the rationality of each candidate character within the overall context is considered, effectively eliminating illogical misidentifications. Furthermore, by constructing a character association path graph and quantifying the strength of contextual associations between characters, the overall coherence and rationality of the character sequence are included in the evaluation scope, avoiding locally optimal but overall unreasonable recognition results. Thus, the calculation of comprehensive confidence organically integrates visual information, contextual compatibility, and the strength of character associations, ensuring that the final recognition result is the optimal choice under multiple constraints, significantly improving the accuracy and robustness of the recognition.
[0176] Through the above technical solutions, this application can significantly improve the accuracy and reliability of intelligent recognition of pharmaceutical packaging information. Specifically, by generating a candidate character list and performing multi-dimensional evaluation, the system can more effectively handle recognition uncertainties caused by poor image quality, blurred characters, or partial occlusion, reducing the false recognition rate. Furthermore, through contextual compatibility evaluation and the construction of character association path graphs, the recognition results not only have high confidence at the individual character level but also conform to logic and norms at the entire information sequence level, thereby effectively avoiding overall recognition failure due to local errors. This comprehensive judgment mechanism enables the system to provide more stable and accurate recognition results when facing complex and ever-changing real-world production environments, reducing reliance on manual review and improving automation levels and production efficiency.
[0177] In some preferred embodiments, a specific example is given below. Suppose we need to identify a drug batch number "ABC12345", but due to printing quality issues, the character "B" in the image is blurry and may be initially identified as "B" or "8", while the character "4" is also partially missing and may be initially identified as "4" or "A".
[0178] First, for each initially identified character, the system generates a list of candidate characters and assigns a visual similarity score. For example, for the second character, the list might be [B (visual similarity 0.8), 8 (visual similarity 0.7)]; for the sixth character, the list might be [4 (visual similarity 0.75), A (visual similarity 0.6)].
[0179] Secondly, based on the preset format specifications (e.g., the first three digits of the batch number are letters, and the last five are numbers) and production plan data (e.g., the current batch number should begin with "ABC"), these candidate characters are evaluated for context compatibility. In this case, for the second character position, the candidate character "B" will have a much higher context compatibility score than "8" because the batch number is expected to begin with "ABC". For the sixth character position, the candidate character "4" will have a much higher context compatibility score than "A" because that position is expected to be a number.
[0180] Next, the system constructs a character association path graph. Nodes in the graph represent candidate characters for each character position, and edges represent the strength of contextual association between characters. For example, in the sequence "ABC", the association strength between "A" and "B" is high, while the association strength between "A" and "8" is low. Similarly, in a sequence of numbers, the association strength between numbers is usually high.
[0181] The system then calculates the overall confidence score for all logically possible character association paths. For example, the overall confidence score for the path "ABC12345" combines the visual similarity score, contextual compatibility score, and inter-character association strength for each character. The overall confidence score for the path "A8C123A5," however, is significantly lower due to the low contextual compatibility scores of "8" and "A" and the low association strength with adjacent characters.
[0182] Ultimately, the system selects the character association path with the highest overall confidence level, namely "ABC12345," as the final identification result for the drug batch number. Through this multi-dimensional and comprehensive judgment, even when there is uncertainty in the recognition of a single character, it can be effectively corrected through contextual information, thereby obtaining an accurate recognition result.
[0183] When the overall confidence of multiple character association paths is the same, the method also includes: S601. Multiple character association paths are used as candidate recognition results for differential feature extraction.
[0184] Differentiating features refer to subtle but crucial features that can distinguish these paths. For example, they may include local deformation features of characters, small differences at the connection of strokes, subtle changes in character spacing, etc. These features may be ignored or have low weight in the initial visual similarity assessment.
[0185] S602. Query the historical production data and product quality control records related to the current batch to obtain the historical production data and product quality control records.
[0186] Historical production data can include information such as the production date, production line, operators, and equipment status of the current or similar batches of products. This data helps in understanding common deviation patterns in the production process. Product quality control records may contain information such as the types and severity of defects in previous batches of products, as well as the corresponding character recognition accuracy, providing empirical references for assessing the reliability of current recognition results.
[0187] S603. Based on the differentiated characteristics, historical production data, and product quality control records, a secondary evaluation is conducted on the candidate identification results to obtain the candidate identification results after the secondary evaluation.
[0188] Secondary evaluation is a deeper judgment process that comprehensively considers subtle visual differences and practical constraints in the production process. For example, weights can be assigned to different differentiating features, and each candidate recognition result can be re-scored or re-ranked by combining the frequency or error rate of a certain character pattern in historical data with the strict requirements for the recognition of specific characters in quality control records.
[0189] S604. Select the candidate recognition result with the highest priority after the second evaluation as the final comprehensive description of the character.
[0190] The "priority" here refers to the result of a secondary evaluation, reflecting which candidate identification result best matches the actual situation and production requirements after comprehensively considering all available information. In this way, even with the same overall confidence level, the most reliable identification result can be selected by introducing more dimensions of data and a more refined evaluation mechanism.
[0191] This application's solution effectively addresses the recognition ambiguity problem encountered when multiple character association paths share the same overall confidence level during contextual logic judgment by introducing multi-dimensional information for secondary evaluation. Specifically, when the system initially identifies multiple character sequences with the same highest confidence level, traditional single-confidence evaluation mechanisms cannot make further decisions. This solution first extracts differentiated features to delve into the subtle visual differences between these paths, which may reflect variations in actual character form or printing quality. Simultaneously, by querying historical production data and product quality control records, external knowledge and experience related to the production process and product quality are introduced, allowing the evaluation to extend beyond the image itself and incorporate the actual production context. Therefore, using these differentiated features and external data as input for secondary evaluation of candidate recognition results enables verification and correction from a more macroscopic and professional perspective. This multi-dimensional information fusion and secondary evaluation mechanism allows the system to make more accurate and robust judgments in complex or ambiguous recognition scenarios, thereby improving the overall reliability and accuracy of recognition.
[0192] In some preferred embodiments, assuming that when identifying the drug batch number "ABC123", the system initially identifies two character association paths, "ABC123" and "ABG123", and the comprehensive confidence scores of these two paths are exactly the same, then, according to the above scheme, the system will use these two paths as candidate identification results.
[0193] First, differential feature extraction is performed. The system may find that, at the positions of the characters "C" and "G", although the visual similarity scores are high, there are slight deformation differences at the ends or connections of the strokes. For example, the opening of "C" may be slightly rounded, while the lower stroke of "G" may have slight breaks or ink spread.
[0194] Secondly, the system will query historical production data and product quality control records related to the current batch. For example, historical production data may show that during a specific period, the lower strokes of the character "G" on this production line were frequently blurred due to printhead wear. Meanwhile, product quality control records may indicate that before shipment, this batch of products had strict printing quality requirements for the letter "C" in the batch number, while the requirements for "G" were relatively more lenient.
[0195] Next, based on these differentiated characteristics, historical production data, and product quality control records, a secondary evaluation is conducted on "ABC123" and "ABG123". In this secondary evaluation, the system may assign higher weight to the printing quality characteristics of the "C" character, combined with error-prone patterns of the "G" character in historical data. For example, if the ambiguous characteristics of the "G" character in the "ABG123" path highly match common defect patterns in historical data, while the characteristics of the "C" character in the "ABC123" path better meet quality control requirements, then the priority of "ABC123" will be significantly increased.
[0196] Ultimately, the system will select "ABC123," which has the highest priority after secondary evaluation, as the final comprehensive description of the drug batch number. In this way, even when the initial identification confidence level is the same, a more accurate and reliable judgment can be made by introducing deeper visual features and external production knowledge.
[0197] In some embodiments described above, this application proposes methods for determining the type and degree of defects present in image information. However, in practical applications, the microstructure of the pharmaceutical packaging material itself (such as paper fibers, plastic particles, printing dots, etc.) may visually resemble actual defects (such as missing ink, scratches, stains), which may affect the accuracy of defect recognition, thereby impacting the subsequent weighting of character information and the reliability of the final recognition result. If the above problems are not addressed, inherent material features may be misjudged as defects, or real defects may be ignored, thus reducing the overall performance of intelligent recognition of pharmaceutical packaging information. Therefore, this application further proposes a more accurate method for determining the type and degree of defects present in image information, aiming to effectively distinguish between real defects and the material's microstructure.
[0198] In this regard, this application further proposes methods for determining the type and degree of defects present in image information, including: S701. Perform multi-scale decomposition on the image information to obtain image components of different scales.
[0199] Specifically, multi-scale decomposition of image information refers to breaking down the original image into a series of image components represented at different spatial frequencies. For example, methods such as wavelet transform, Gaussian pyramid, or Laplacian pyramid can be used to capture detailed information of the image at different scales.
[0200] S702. Identify and extract texture features related to the microstructure of the material from image components of different scales.
[0201] Texture features can be statistical features (such as gray-level co-occurrence matrix, local binary mode), structural features, or spectral features, used to characterize the inherent fine structure of a material surface.
[0202] S703. Perform frequency domain analysis on the image information to obtain periodic patterns related to the microstructure of the material.
[0203] Frequency domain analysis of image information, such as through Fourier transform, can reveal periodic patterns in the image. These periodic patterns are often related to the manufacturing process or inherent structure of the material, such as woven textures or printing dots, thus yielding periodic patterns related to the material's microstructure.
[0204] S704. Construct a description of the material's microstructure features based on texture characteristics and periodic patterns.
[0205] Based on the extracted texture features and periodic patterns, a comprehensive description of the material's microstructure can be constructed. This description accurately characterizes the inherent visual properties of pharmaceutical packaging materials, rather than actual defects.
[0206] S705. Extract defect features from image information to obtain visual features of potential defects.
[0207] Furthermore, defect feature extraction is performed on the image information to identify all possible abnormal areas in the image and obtain the visual characteristics of potential defects. These potential defects may include scratches, stains, missing ink, blurred printing, etc.
[0208] S706. Compare the visual characteristics of potential defects with the description of the material's microstructure characteristics.
[0209] The purpose of the comparison is to distinguish which visual features are inherent to the material and which are actual defects.
[0210] S707. When the visual characteristics of a potential defect are highly similar to the description of the material's microstructure characteristics, initiate the feature stripping process.
[0211] When the comparison results show a high similarity between the visual features of a potential defect and the description of the material's microstructure features, it indicates that the potential defect is likely not a real defect, but rather a feature of the material itself. At this point, the system will initiate the feature stripping process.
[0212] S708. In the feature stripping process, the image information is subjected to adaptive filtering to suppress the influence of the material's microstructure.
[0213] In the feature stripping process, the image information undergoes adaptive filtering. This filtering is customized based on the material's microstructure characteristics, aiming to selectively suppress or remove components in the image that are similar to the material's microstructure, thereby minimizing their interference with defect identification.
[0214] S709. Perform defect region segmentation on the filtered image information to obtain the real defect region.
[0215] After filtering, the influence of the material's microstructure on the image information is significantly suppressed. Based on this, defect region segmentation of the image information can more accurately identify the true defect areas. For example, methods such as threshold segmentation, edge detection, region growing, or deep learning segmentation models can be used.
[0216] S710. Based on the actual defect area, quantify the type and extent of the defect.
[0217] For example, the types of defects can be scratches, stains, bubbles, etc.; the degree of defects can be measured by area size, severity, contrast, etc.
[0218] like Figure 3 As shown in the figure, this invention also provides a smart identification system for pharmaceutical packaging information. The system includes: The image information acquisition module is used to acquire image information of the drug packaging. The character information extraction module is used to extract various types of character information from image information; these various types of character information include character stroke structure, character texture, and character geometric contour. The defect analysis module is used to analyze image information to obtain the types and degrees of defects present in the image information; The weight setting module is used to set the weight of various types of character information according to the defect type and defect severity. The character description integration module is used to integrate various types of character information with set weights to obtain a preliminary comprehensive description of the character; The context logic judgment module is used to perform context logic judgment on the preliminary comprehensive description of the character to obtain the final comprehensive description of the character. The context logic judgment uses preset format specifications, production plan data and the relationship between characters to verify and correct the recognition results.
[0219] This application also provides a computer-readable storage medium. All or part of the processes in the above method embodiments can be executed by a computer program instructing related hardware. This program can be stored in the computer-readable storage medium, and when executed, it can include the processes of the above method embodiments. The computer-readable storage medium can be an internal storage unit of the task execution device (including a data sending end and / or a data receiving end) of any of the foregoing embodiments, such as the hard disk or memory of the task execution device. The computer-readable storage medium can also be an external storage device of the terminal device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the terminal device. Further, the computer-readable storage medium can include both the internal storage unit of the task execution device and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by the task execution device. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.
[0220] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0221] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, essentially, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.
[0222] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be covered within the scope of protection of this application.
Claims
1. A method for intelligent identification of pharmaceutical packaging information, characterized in that, include: Obtain image information of drug packaging; Extract various types of character information from the image information; The various types of character information include character stroke structure, character texture, and character geometric outline; Determining the type and degree of defects present in the image information includes: performing multi-scale decomposition on the image information to obtain image components at different scales; identifying and extracting texture features related to the material's microstructure from the image components at different scales; performing frequency domain analysis on the image information to obtain periodic patterns related to the material's microstructure; constructing a feature description of the material's microstructure based on the texture features and the periodic patterns; extracting defect features from the image information to obtain visual features of potential defects; comparing the visual features of potential defects with the feature description of the material's microstructure; initiating a feature stripping process when the visual features of potential defects are highly similar to the feature description of the material's microstructure; in the feature stripping process, performing adaptive filtering on the image information to suppress the influence of the material's microstructure; segmenting the filtered image information into defect regions to obtain real defect regions; and quantifying the type and degree of defects based on the real defect regions. The weight of each type of character information is determined based on the defect type and defect severity. By integrating the various types of character information with set weights, a preliminary comprehensive description of the character is obtained; The initial comprehensive description of the character is subjected to contextual logic judgment to obtain the final comprehensive description of the character; the contextual logic judgment uses preset format specifications, production plan data and the relationship between characters to verify and correct the recognition result.
2. The intelligent identification method for pharmaceutical packaging information according to claim 1, characterized in that, The process of performing contextual logic judgments on the preliminary comprehensive description of the character to obtain the final comprehensive description of the character includes: The probability of character visual features, the probability of applicability of verification algorithms, and the probability of MES matching are determined. The probability of character visual features is the probability that each character is identified as a specific character in the image. The probability of applicability of verification algorithms is the probability of applicability of each verification algorithm. The probability of MES matching is the probability of matching each candidate batch number with the MES topology. Using the character visual feature probability, the applicability probability of the verification algorithm, and the MES matching probability as input nodes, a Bayesian inference network with multi-dimensional information fusion is constructed. Generate all logically possible candidate batch number sequences; Based on the Bayesian inference network, evaluate the posterior probability of each candidate batch number sequence; The recognition result with the highest posterior probability is selected as the final output to obtain the final comprehensive description of the character.
3. The intelligent identification method for pharmaceutical packaging information according to claim 2, characterized in that, Determining the applicability probability of the verification algorithm includes: Extract unstructured information related to the drug packaging to infer the product classification intent; the unstructured information includes one or more of the following: product name, production batch, and production line number; Based on the product classification intent, determine the applicability probability of the verification algorithm.
4. The intelligent identification method for pharmaceutical packaging information according to claim 2, characterized in that, Determining the MES matching probability includes: When the target expected batch number information is obtained, the target expected batch number information is deconstructed into a batch number topology structure containing known characters, placeholders and their possible value ranges; Narrow down the possible values of the placeholder based on the auxiliary information; The preliminary comprehensive description of the characters is matched with the batch number topology, and characters with low confidence levels and preset values are filled in with the possible values of the placeholders. Calculate the matching probability of each candidate batch number with the batch number topology.
5. The intelligent identification method for pharmaceutical packaging information according to claim 4, characterized in that, The method further includes: Send data requests to multiple preset data sources, including manufacturing execution systems, enterprise resource planning systems, and laboratory information management systems; Configure a data source reliability parameter for each of the multiple data sources. The data source reliability parameter is preset based on the historical data accuracy, update frequency and response speed of the data source. Receive expected batch number information from the multiple data sources; Check whether there are any conflicts in the expected batch number information; When the expected batch number information conflicts with each other, the conflicting expected batch number information is weighted according to the data source reliability parameter of each data source. The batch number information with the highest weighting value is selected as the target expected batch number information.
6. The intelligent identification method for pharmaceutical packaging information according to claim 1, characterized in that, The process of performing contextual logic judgments on the preliminary comprehensive description of the character to obtain the final comprehensive description of the character includes: For each character initially identified, a list containing multiple candidate characters is generated, and a visual similarity score is assigned to each candidate character; Based on the preset format specifications and production plan data, a context compatibility evaluation is performed on each candidate character in the list containing multiple candidate characters; Construct a character association path graph, where each node in the character association path graph is a candidate character at each character position, and the edges in the character association path graph represent the strength of the contextual association between characters; Calculate the overall confidence score for each character association path, which integrates the visual similarity score, context compatibility score, and inter-character association strength for each character. The character association path with the highest overall confidence level is selected as the final comprehensive description of the character.
7. The intelligent identification method for pharmaceutical packaging information according to claim 6, characterized in that, When the overall confidence level of multiple character association paths is the same, the method further includes: The multiple character association paths are used as candidate recognition results for differential feature extraction. Query historical production data and product quality control records related to the current batch to obtain historical production data and product quality control records; Based on the differentiated features, the historical production data, and the product quality control records, the candidate identification results are evaluated a second time to obtain the candidate identification results after the second evaluation. The candidate recognition result with the highest priority after secondary evaluation is selected as the final comprehensive description of the character.
8. The intelligent identification method for pharmaceutical packaging information according to claim 1, characterized in that, After performing multi-scale decomposition on the image information to obtain image components of different scales, the method further includes: Energy leakage is detected for each scale component; When an energy leak is detected, the leaked energy is compensated in a targeted manner. Independence assessment of each scale component; Based on the independence assessment results, the components are orthogonalized.
9. A smart identification system for pharmaceutical packaging information, characterized in that, For performing the method according to any one of claims 1-9, the system comprises: The image information acquisition module is used to acquire image information of the drug packaging. The character information extraction module is used to extract various types of character information from the image information; the various types of character information include character stroke structure, character texture, and character geometric contour. The defect analysis module is used to analyze the image information to obtain the type and degree of defects present in the image information; The weight setting module is used to set the weight of the various types of character information according to the defect type and defect severity. The character description integration module is used to integrate the various types of character information after setting a certain weight to obtain a preliminary comprehensive description of the character; The context logic judgment module is used to perform context logic judgment on the preliminary comprehensive description of the character to obtain the final comprehensive description of the character; the context logic judgment uses preset format specifications, production plan data and the relationship between characters to verify and correct the recognition result.