A drug whole-process circulation traceability system and method
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-12
- Publication Date
- 2026-08-14
AI Technical Summary
[0004]本发明提供一种药品全流程流通追溯系统及方法,以解决现有的问题,即现有技术中,缺乏一种能够自动适配不同系统编码规则、智能关联多环节数据、并实时输出全链条追溯结果的系统,导致“来源可查、去向可追、责任可究”的药品追溯目标难以实现
[0015]本发明的技术方案的有益效果是:通过对药品的大量药品单元所对应的追溯数据进行大数据分析,确定了药品在全流程流通过程中所表现出来的拓扑结构,从而使得可以基于该拓扑结构,在即使追溯数据的编码规则和数据格式存在差别的情况下,仍能够确定具体药品单元在该拓扑结构中的全流程流通的环节链,从而实现对药品单元的全链条追溯;解决了药品追溯系统中因编码规则不统一、数据分散异构导致的全链条关联准确率低的技术问题;通过构建基于时间切片的多维度关联分析框架,实现了异构数据的自动适配与关联,且能够适应药品流通网络的动态演化特性;通过全流程流通模型的构建,提升了药品全流程追溯的完整性与准确性,实现了异构数据的自动适配与关联。
Smart Images

Figure CN122022850B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of pharmaceutical data management technology, specifically to a pharmaceutical full-process traceability system and method. Background Technology
[0002] Current drug traceability systems primarily rely on "one item, one code" technology. However, different systems at each stage still use different coding rules and data formats, resulting in data being scattered across production, distribution, and usage systems, hindering the effective establishment of a complete supply chain. The production stage often uses NMPA electronic regulatory codes, the distribution stage frequently uses custom batch numbers, while the usage stage employs different traceability code formats. This lack of data format consistency, closed system interfaces, and automated data collection leads to insufficient accuracy in the drug traceability system's end-to-end tracking, seriously impacting drug safety supervision and patient medication safety.
[0003] In the existing technology, there is a lack of a system that can automatically adapt to different system coding rules, intelligently associate data from multiple links, and output full-chain traceability results in real time, making it difficult to achieve the drug traceability goal of "traceable source, traceable destination, and accountable responsibility". Summary of the Invention
[0004] This invention provides a drug full-process circulation traceability system and method to solve the existing problem that the existing technology lacks a system that can automatically adapt to different system coding rules, intelligently associate data from multiple links, and output full-chain traceability results in real time, making it difficult to achieve the drug traceability goal of "traceable source, traceable destination, and accountable responsibility".
[0005] The present invention provides a drug whole-process circulation traceability system and method, which adopts the following technical solution: One embodiment of the present invention provides a method for tracing the entire process of drug distribution, the method comprising the following steps: Obtain multi-source traceability data of several drug units of the target drug, wherein the multi-source traceability data includes traceability coding information, date information, subject information and location information of the corresponding drug unit at each stage of the circulation process; By dividing time into several time slices, and based on the traceability coding information, date information, and location information of all drug units, the correlation between drug units in different links under each information dimension is comprehensively analyzed to obtain the process correlation coefficient between drug units; combined with the process correlation coefficient of drug units in any time slice, a local circulation network in the corresponding time slice is constructed. By utilizing the process correlation coefficient of the drug unit and combining the main information of the drug unit in the circulation process, the local circulation network of all time slices is spliced into a graph structure to obtain the full-process circulation model of the target drug. A full-process circulation model is used to generate a visual diagram of the entire circulation trajectory of each drug unit.
[0006] Optionally, the specific method for obtaining the process correlation coefficient is as follows: Based on the traceability coding information, date information, and location information of all drug units in any stage, clusters are formed to obtain several clusters. Based on the correlation of drug units within clusters in adjacent links of the circulation process in terms of traceability coding information, date information, and location information, and combined with preset weights, the process correlation coefficient between drug units is calculated.
[0007] Optionally, the method for calculating the process correlation coefficient between drug units based on the association of drug units within clusters of adjacent links in the circulation process in terms of traceability coding information, date information, and location information, and in combination with preset weights, includes the following specific methods: For any adjacent drug units, calculate the time correlation between drug units based on the distribution and differences of date information in the cluster to which the drug unit belongs; The coding logic matching degree between drug units is calculated based on the traceability coding information of the drug units; the Euclidean distance between drug units is calculated based on the location information of the drug units, thereby obtaining the spatial correlation degree between drug units; Preset weights for the temporal correlation, coding logic matching degree, and spatial correlation between drug units, and obtain the process correlation coefficient between drug units through weighted fusion.
[0008] Optionally, the specific method for obtaining the temporal correlation between the drug units is as follows: The time evolution coefficients between drug units are calculated based on the discreteness and distribution level of the time information of drug units in adjacent links within their respective clusters. By combining the time evolution coefficients between drug units and the differences in corresponding time information, the time correlation between drug units is obtained.
[0009] Optionally, the specific method for obtaining the coding logic matching degree between the drug units is as follows: The traceability coding information corresponding to the drug units in adjacent links is processed by a preset rule parser to obtain the corresponding coding data. The Jaccard coefficient of the corresponding coding data between drug units is used as the coding logic matching degree between drug units.
[0010] Optionally, the specific method for obtaining the local circulation network is as follows: The effective association threshold is determined based on the statistical distribution of the process association coefficient, and drug unit pairs formed by corresponding drug units when the process association coefficient is greater than or equal to the effective association threshold are selected. The combination of the main body and the links corresponding to each drug unit is used as network nodes, the relationship between drug units and their contained drug units is used as directed edges, and the process association coefficient between drug units and their contained drug units is used as the initial weight of the edges, thereby forming a local flow sub-network between adjacent links in any time slice. All adjacent links within a time slice are spliced together to form the corresponding local circulation subnetwork. Nodes are deduplicated and merged according to their main identifiers, while edges retain their directionality and weight information, thereby constructing the corresponding local circulation network.
[0011] Optionally, the specific method for obtaining the end-to-end circulation model is as follows: The main information mentioned includes at least pharmaceutical companies, distribution companies, hospitals, and pharmacies involved in the distribution process; Based on the distribution of time information of drug units in any time slice and the distribution of the quantity of drugs distributed through distribution companies, the circulation mode of the corresponding time slice is obtained. By combining the circulation modes of each time slice, the local circulation networks corresponding to all time slices are spliced together to form a full-process circulation model.
[0012] Optionally, the specific method for obtaining the flow mode of the time slice is as follows: For any time slice, the average time interval from production to use of the drug unit within the time slice is used as the flow rate characteristic of the corresponding time slice; The standard deviation of the number of drug units distributed by all distribution companies in the distribution process is calculated, and the concentration characteristics of the corresponding time slices are calculated based on the standard deviation, where the standard deviation and concentration characteristics are negatively correlated. Among all distribution companies, the distribution company that obtains more drug units than the average number of drug units distributed by all distribution companies is denoted as the target distribution company. The proportion of the target distribution company among all distribution companies is used as the main activity feature of the corresponding time slice. Using the circulation speed characteristics, concentration characteristics, and subject activity characteristics of time slices as inputs, a pre-trained neural network is used to identify the circulation mode corresponding to the time slice. The circulation mode includes conventional distribution mode, emergency direct distribution mode, and hybrid mode.
[0013] Optionally, the specific method for obtaining the time slice is as follows: Based on the date information of all drug units, a circulation time series is constructed. The circulation time series of the target drug is divided into several sequence segments according to a preset time granularity. The interval formed by the maximum and minimum values of the elements in each sequence segment is used as the corresponding time slice.
[0014] A drug full-process circulation traceability system includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements any of the steps of the drug full-process circulation traceability method described above.
[0015] The beneficial effects of the technical solution of this invention are as follows: By conducting big data analysis on traceability data corresponding to a large number of drug units, the topological structure exhibited by drugs in the entire circulation process is determined. This allows for the identification of the specific drug unit's link in the entire circulation chain within this topological structure, even when there are differences in the coding rules and data formats of the traceability data, thus achieving full-chain traceability of drug units. It solves the technical problem of low accuracy in the entire chain association caused by inconsistent coding rules and dispersed, heterogeneous data in drug traceability systems. By constructing a multi-dimensional association analysis framework based on time slices, automatic adaptation and association of heterogeneous data are achieved, and it can adapt to the dynamic evolution characteristics of the drug circulation network. Through the construction of the full-process circulation model, the completeness and accuracy of drug full-process traceability are improved, and automatic adaptation and association of heterogeneous data are realized. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a flowchart illustrating the steps of a drug whole-process circulation traceability method according to the present invention; Figure 2 This is a structural block diagram of a drug full-process circulation traceability system according to the present invention. Detailed Implementation
[0018] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of a drug whole-process circulation traceability system and method proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.
[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0020] The following description, in conjunction with the accompanying drawings, details the specific scheme of the drug full-process circulation traceability system and method provided by the present invention.
[0021] Please see Figure 1 The diagram illustrates a flowchart of a drug whole-process circulation traceability method according to an embodiment of the present invention, which includes the following steps: Step S001: Obtain multi-source traceability data for several drug units of the target drug.
[0022] It should be noted that, due to the typically large quantities involved in drug production, directly incorporating multi-source traceability data from all drug units into big data analysis to construct a complete distribution topology would require substantial computing resources when defining the entire drug distribution process. Therefore, when acquiring multi-source traceability data for drug units, it is first necessary to screen and sample from the drugs that require full-process distribution traceability, and then further preprocess the traceability data of the sampled drug units.
[0023] Specifically, in order to implement the drug end-to-end traceability method proposed in this embodiment, it is first necessary to obtain multi-source traceability data of several drug units of the target drug. The specific process is as follows: First, any drug is selected as the target drug, and each unit of the target drug is taken as a drug unit. All drug units of the target drug are sampled evenly to obtain several drug units. Multi-source traceability data of each drug unit is then obtained. The multi-source traceability data includes traceability coding information, date information, location information, and main information of the drug unit at different stages. The traceability coding information includes the electronic supervision code of the production stage, the custom batch number of the distribution stage, and the traceability code of the usage stage.
[0024] The date information includes the drug's production date, distribution time, and usage time.
[0025] The main information includes pharmaceutical companies, distribution companies, pharmacies, and hospitals.
[0026] Then, the multi-source traceability data of pharmaceuticals is standardized.
[0027] As an optional embodiment, the standardization processing of multi-source traceability data for pharmaceuticals includes the following specific methods: converting production date, distribution time, and usage time into ISO 8601 format (YYYY-MM-DD HH:MM:SS) to eliminate differences in time formats across different systems; extracting batch identifiers from the electronic supervision code in the production process (e.g., the 12th-15th digits of the NMPA code being the batch number), the rule features of custom batch numbers in the distribution process (e.g., the "YYMMDD-001" format), and the coding structure of the traceability code in the usage process; and mapping the main information of pharmaceutical companies, distribution companies, pharmacies, hospitals, etc., to standardized organization codes (e.g., unified social credit codes).
[0028] Based on this, the feature extraction of the circulation cycle is added: based on the date information of the drug unit, the circulation time series is constructed, and the historical circulation data of the target drug is divided into several time slices according to the preset time granularity (such as monthly or quarterly). The drug units in each time slice form a sub-circulation network for subsequent dynamic evolution analysis.
[0029] Thus, multi-source traceability data for several drug units of the target drug have been obtained through the above methods.
[0030] Step S002: Obtain several time slices by dividing the time. Based on the traceability coding information, date information and location information of all drug units, comprehensively analyze the correlation between drug units in different links under each information dimension to obtain the process correlation coefficient between drug units. Combine the process correlation coefficient of drug units in any time slice to construct a local circulation network in the corresponding time slice.
[0031] It should be noted that, since the drug distribution process typically involves multiple stages—production, sales, and use—and these stages are managed by different systems, drug traceability data cannot be effectively linked across the entire chain. To avoid the impact of multiple codes for a single drug on the effectiveness and accuracy of end-to-end traceability results, this method matches and links multi-source traceability data for drug units to construct a local distribution network within corresponding time slices. This facilitates the subsequent construction of an end-to-end distribution model, enabling the association of drug distribution information across all time periods and stages. This achieves data interoperability throughout the entire distribution process, facilitating further visualized traceability.
[0032] As a preferred embodiment, the specific method for obtaining the process correlation coefficient is as follows: Based on the traceability coding information, date information, and location information of all drug units in any stage, clusters are formed to obtain several clusters. Based on the correlation of drug units within clusters in adjacent links of the circulation process in terms of traceability coding information, date information, and location information, and combined with preset weights, the process correlation coefficient between drug units is calculated.
[0033] As an optional embodiment, the method of clustering all drug units within a link based on the traceability coding information, date information, and location information of all drug units in any link to obtain several clusters includes: taking the array formed by the traceability data of any drug unit in any link as a subarray of that link, using the distance of the subarrays, and using the DBSCAN clustering algorithm to cluster all drug units of the target drug to obtain several clusters under the corresponding link.
[0034] It should be noted that in this step, the standardized timestamps and GPS coordinates are normalized by maximum and minimum values and concatenated into a numerical vector. Euclidean distance is used as the distance metric input for the DBSCAN clustering algorithm. In addition, the clustering analysis is performed independently within each time slice to ensure that drug units in different circulation cycles are not incorrectly clustered into the same cluster, thus maintaining the isolation of the time dimension.
[0035] Furthermore, since the time spent at different stages in the drug production and distribution process typically varies, and even the time spent at the same stage can differ, setting fixed values for correlation analysis of multi-source traceability data of drug units at different stages cannot accurately reflect the correlation in the time dimension. Additionally, the differences in traceability coding information among multi-source traceability data affect the results of logical correlation analysis of batches. Therefore, this method combines traceability coding information from different stages with time information to perform logical correlation analysis of batches. Moreover, drug units occupy different positions at different stages in the distribution process; therefore, location information is used to further determine the spatial correlation of drug units at different stages.
[0036] As an optional embodiment, the method for calculating the process correlation coefficient between drug units based on the association of traceability coding information, date information, and location information within clusters of adjacent links in the circulation process, combined with preset weights, includes the following specific method: For any adjacent drug units, calculate the time correlation between drug units based on the distribution and differences of date information in the cluster to which the drug unit belongs; The coding logic matching degree between drug units is calculated based on the traceability coding information of the drug units; the Euclidean distance between drug units is calculated based on the location information of the drug units, thereby obtaining the spatial correlation degree between drug units; Preset weights for the temporal correlation, coding logic matching degree, and spatial correlation between drug units, and obtain the process correlation coefficient between drug units through weighted fusion.
[0037] As an optional embodiment, the specific method for obtaining the time correlation between drug units includes: for any two adjacent stages, designated as the first stage and the second stage; using standardized date information, calculating the absolute value of the difference between corresponding times between drug units belonging to the first stage and the second stage respectively, as the time interval between the corresponding drug units; recording the set of time intervals between all any two drug units belonging to the first stage and the second stage respectively as the total time interval set of the first stage and the second stage; and obtaining the time correlation between drug units based on the distribution of elements in the total time interval set and the time interval between drug units.
[0038] As an optional embodiment, the specific method for calculating the time correlation between the drug units can be: in, Indicates the temporal correlation between drug units; The function is a maximum-minimum normalization function, used to linearly map data to... interval; Represents the time evolution coefficient; The standard deviation of all time intervals in the total time interval set. It is the mean of all time intervals in the total time interval set; This refers to the time corresponding to the drug unit belonging to the first stage; This refers to the time corresponding to the drug unit belonging to the second stage; Represents the absolute value function; Represents an exponential function with the natural constant as the base; This indicates the time interval between drug units.
[0039] Regarding the calculation method for the time interval between drug units, since the time information corresponding to the drug units is in ISO8601 format, in a specific embodiment of the calculation method for the time interval, the specific calculation method is as follows: Let the time information corresponding to the drug unit in the first stage be a string. The time information corresponding to the drug unit in the second stage is a string. The format is "YYYY-MM-DD HH:MM:SS"; and The inputs are then parsed by the `datetime.strptime` function in the Python `datetime` module to obtain the corresponding time objects. and ; Set the time object and Converted to a unified standard timestamp value, denoted as follows: and to obtain The result is then converted into a value in days, as... .
[0040] It should be noted that time correlation aims to quantify the impact of the dispersion of time distribution between adjacent links on the reliability of the correlation. The more concentrated the time interval distribution between adjacent links, the more regular the flow time between the two links, and the higher the reliability of the time correlation. Conversely, if the time interval distribution is discrete (such as the presence of a large number of abnormally fast or slow flows), even if the time difference between two drug units is small, the reliability of the correlation may be reduced due to the overall chaotic distribution. This avoids the limitations of simply relying on the absolute value of the time difference for correlation, and evaluates local time differences within the context of the overall time distribution, which is more in line with the actual business logic of drug distribution.
[0041] As an optional embodiment, the specific method for obtaining the coding logic matching degree between drug units includes: converting the traceability coding information corresponding to drug units belonging to the first stage and the second stage respectively through a preset rule parser to obtain the coding data of each drug unit in the corresponding stage, wherein the coding data contains several key values; and using the Jaccard coefficient between the coding data of drug units in the corresponding stage as the coding logic matching degree between drug units.
[0042] As an optional embodiment, the method of converting the traceability coding information corresponding to the drug units belonging to the first and second stages respectively through a preset rule parser to obtain the coding data of each drug unit in the corresponding stage includes the following specific method: the rule parser has a built-in multi-level coding parsing rule library, which includes NMPA electronic supervision code parsing rules, GS1 coding parsing rules, and custom batch number parsing rules; for the NMPA electronic supervision code in the production stage, the rule parser identifies its 20-bit coding structure: extracting the first 7 bits as the drug identification code (including enterprise identification code). The system extracts information such as industry information, drug name, dosage form, approval number, and packaging specifications. Digits 8-16 are extracted as the production serial number (uniquely identifying each drug unit), and digits 17-20 are extracted as check digits. The extracted drug identification code is combined with the production serial number to form a key-value pair {drug identification code: production serial number}, which serves as the encoding data for that drug unit in the production process. For custom batch numbers in the distribution process, the rule parser identifies the encoding format using regular expression matching rules: if the batch number format is "YYMMDD-XXX" (where YYMMDD represents year, month, and day, and XXX represents the serial number), then... If the batch number contains a distributor code segment (e.g., "DZ-XXX-YYMMDD"), then the distributor code "DZ" is extracted as the main key. The extracted segments are combined to form a key-value set {time key: batch sequence key, main key: distributor code}, which serves as the coded data for the drug unit in the distribution process. For traceability codes in the usage stage, the rule parser first identifies the coding system type: if it is a medical insurance traceability code, then the drug classification code, region code, and sequence number are extracted from the medical insurance code. The sequence number is used; if it is an internal traceability code of a medical institution, the institution code, the batch number entering the warehouse, and the department code are extracted; the extracted code segments are combined according to the preset hierarchical relationship to form a tree-structured coding data, with the root node being the institution identifier and the child nodes containing batch branches and department branches; before performing parsing, the rule parser first performs validity verification on the input traceability coding information: the validity of the NMPA code is verified by the check bit verification algorithm, and the compliance of the custom batch number is verified by the format template matching; for codes that fail verification, they are marked as abnormal codes and a manual review process is triggered, and they do not participate in subsequent association analysis.
[0043] As an optional embodiment, the specific method for obtaining the spatial correlation between the drug units is as follows: calculate the Euclidean distance between drug units in adjacent links through location information, and then obtain the spatial correlation between drug units based on the Euclidean distance.
[0044] As an optional embodiment, the specific method for calculating the spatial correlation between the drug units can be: ;in Indicates the spatial correlation between drug units; Represents the natural constant; The preset attenuation coefficient; This represents the Euclidean distance between drug units.
[0045] It should be noted that, in a specific embodiment of the present invention, an attenuation coefficient is set. The specific implementation of this invention is not limited to any particular aspect, but can be adjusted according to the actual situation. In addition, the transfer of drugs from production to use requires physical space, and the geographical locations of adjacent links should show reasonable spatial continuity. Therefore, the exponential decay model reflects the intuitive understanding that the closer the distance, the higher the probability of correlation.
[0046] As an optional embodiment, the specific method for calculating the process correlation coefficient between the drug units can be: ;in, Indicates time weight; Indicates batch weight; Indicates spatial weights; Indicates the temporal correlation between drug units; Indicates the degree of logical matching between drug units; This indicates the spatial correlation between drug units.
[0047] It should be noted that the preset time weights... Batch weight Spatial weights In a specific embodiment of the present invention, the corresponding preset method is as follows: since the continuity of time is a key factor in drug distribution, therefore, the following is set... Since batch logic is the core basis for drug traceability, but it needs to be handled with caution due to the influence of coding rules, therefore, a specific setting is implemented. Since spatial distance can help eliminate false cross-regional associations, therefore, it is set... In other embodiments, adjustments may be made according to actual circumstances, and the embodiments of the present invention are not specifically limited.
[0048] In addition, since the time, batches and the entities involved in different stages of the continuous production and distribution of the target drug are different, in order to clarify the circulation association of the drug units of the target drug, this method uses multi-dimensional association analysis to determine the process association coefficient of the drug units, which reflects the association relationship formed by a large number of drug units in the target drug in different dimensions (time dimension, batch logic dimension and spatial location dimension).
[0049] As an optional embodiment, the specific method for obtaining the local distribution network is as follows: An effective association threshold is determined based on the statistical distribution of the process association coefficients; drug unit pairs formed by drug units whose process association coefficients are greater than or equal to the effective association threshold are selected; the combination of the subject and process corresponding to each drug unit is used as a network node; the association relationship formed by the drug units contained in the drug unit pair is used as a directed edge; and the process association coefficient between the drug units contained in the drug unit pair is used as the initial weight of the edge, thereby forming a local distribution sub-network between adjacent processes within any time slice; the local distribution sub-networks corresponding to all adjacent process pairs within the time slice are spliced together, nodes are deduplicated and merged according to the subject identifier, and edges maintain directionality and weight information, thereby constructing the corresponding local distribution network.
[0050] As an optional embodiment, the specific method for obtaining the effective association threshold is as follows: taking any cluster in the first stage as the first cluster and any cluster in the second stage as the second cluster, extracting the process association coefficients of drug unit pairs in the first and second clusters of all adjacent stages, and performing normal distribution fitting, obtaining the mean and standard deviation in the normal distribution fitting results, and determining the effective association threshold based on the mean and standard deviation.
[0051] It should be noted that, in a specific embodiment of the present invention, the following is taken: As an effective association threshold, among which This represents the mean in the normal distribution fitting result; This represents the standard deviation in the fit to the normal distribution; additionally, exceeding... When the correlation strength of a drug unit is high, it represents the actual distribution path; otherwise, it is mostly system noise or false correlation. In addition, if there are multiple correlation pairs of the same drug unit in adjacent links, only the pair with the highest process correlation coefficient is retained to avoid multiple correlations of the same drug unit in the distribution process.
[0052] Thus, the local flow network within each time slice is obtained using the above method.
[0053] Step S003: Using the process correlation coefficient of the drug unit and combining the main information of the drug unit in the circulation process, the local circulation network of all time slices is spliced into a graph structure to obtain the full-process circulation model of the target drug.
[0054] It should be noted that this step aims to integrate the local distribution networks of each time slice to construct a full-process distribution model covering the entire life cycle and distribution links of a drug. This model can reflect the complete distribution topology of a drug from production to use, and even in the case of multiple codes for a single item, the distribution path of the drug unit can be inferred through the correlation.
[0055] As a preferred embodiment, the specific method for obtaining the full-process circulation model is as follows: based on the distribution of time information of drug units in any time slice and the distribution of the quantity of drugs circulated through distribution companies, the circulation mode of the corresponding time slice is obtained; combined with the circulation modes of each time slice, the local circulation networks corresponding to all time slices are spliced together to form the full-process circulation model.
[0056] As an optional embodiment, the specific method for obtaining the circulation mode of the time slice is as follows: For any time slice, the average time interval from production to use of the drug unit within the time slice is used as the circulation speed feature of the corresponding time slice; the standard deviation of the number of drug units circulated by all distribution companies in the distribution process is calculated, and the concentration feature of the corresponding time slice is calculated based on the standard deviation, wherein the standard deviation is negatively correlated with the concentration feature; among all distribution companies, the distribution companies whose number of circulated drug units is greater than the average number of circulated drug units of all distribution companies are identified as target distribution companies, and the proportion of the target distribution companies among all distribution companies is used as the main activity feature of the corresponding time slice; the circulation speed feature, concentration feature, and main activity feature of the time slice are used as inputs, and the circulation mode corresponding to the time slice is identified through a pre-trained neural network, wherein the circulation mode includes conventional distribution mode, emergency direct distribution mode, and hybrid mode.
[0057] As an optional embodiment, the specific calculation method for the concentration feature can be: ,in This indicates the concentration characteristics of the corresponding time slice; This represents the standard deviation of the number of drug units distributed by all distribution companies in the distribution process. This represents an exponential function with the natural constant as its base.
[0058] As an optional embodiment, the training method of the pre-trained neural network is as follows: A circulation modality classification neural network is constructed, which adopts a three-layer fully connected structure: the input layer contains 3 neurons, receiving normalized values of circulation speed features, concentration features, and subject activity features respectively; the hidden layer contains 10 neurons and uses the ReLU activation function; the output layer contains 3 neurons and uses the Softmax activation function, corresponding to the probability outputs of the regular distribution modality, emergency direct distribution modality, and mixed modality respectively; historical drug circulation data is collected as training samples: labeled circulation records covering different time periods, different drug types, and different circulation scales are extracted from the drug traceability database, and the labeling is manually determined by business experts based on the actual circulation path characteristics to belong to the modality category; the training samples are divided into a training set and a validation set; the circulation speed features, concentration features, and subject activity features of each sample are calculated to form a three-dimensional feature set. The model is trained using the cross-entropy loss function and the Adam optimizer, with an initial learning rate of 0.001, a batch size of 32, and 100 training epochs. During training, if the validation set loss does not decrease for 10 consecutive epochs, an early stopping mechanism is triggered, and the optimal model parameters are saved. After training, the model performance is evaluated using a confusion matrix, requiring an accuracy of at least 90% for the regular distribution modality, a recall of at least 85% for the emergency direct distribution modality, and an F1-score of at least 0.88 for the mixed modality. If the evaluation metrics are not met, the training samples are expanded and the number of hidden layer neurons is adjusted (increased to 15-20), and the model is retrained until the performance requirements are met. The trained neural network model is then packaged into a circulation modality recognition service. The service takes three feature values of the time slice to be recognized as input and outputs the probability distribution of the time slice belonging to each circulation modality. The modality corresponding to the highest probability value is taken as the recognition result.
[0059] It should be noted that the parameters set during the training process of the neural network are empirical values set in one embodiment of the present invention. In other embodiments, they can be adjusted according to the actual situation. The embodiments of the present invention do not impose specific limitations.
[0060] As an optional embodiment, the method of combining the flow modes of each time slice and splicing the local flow networks corresponding to all time slices to form a full-process flow model includes: Step A: For the same drug unit, establish a globally unique identifier (GUID) between its production stage nodes and its usage stage nodes. The GUID is generated by hashing the drug identification code and the production serial number. The GUID is used to normalize the identity of nodes appearing in different time slices of the same drug unit. Step B: If the circulation modes of adjacent time slices are both regular distribution modes, a progressive splicing method is used: retain the nodes of overlapping links in the two time slices, use the main identifier of the overlapping node as the anchor point, and perform topological fusion of the downstream link of the previous time slice and the upstream link of the next time slice; if the previous time slice is a regular distribution mode and the next time slice is an emergency direct distribution mode, a bridging splicing method is used: insert a mode conversion node between the two time slices, mark the abrupt change point of the circulation path, and maintain the integrity of the internal network structure of the two time slices; if adjacent time slices are both emergency direct distribution modes, a direct connection splicing method is used: ignore the intermediate distribution links, establish a direct association edge between the production link node and the user link node, and mark the edge with the "emergency direct distribution" attribute; Step C: For nodes representing the same distribution entity (such as the same distribution company) in different time slices, merge them into a single node in the global network. The attributes of the merged node include the set of timestamps of the entity in each time slice. For multiple edges across time slices (i.e., there are multiple related edges generated by time slices between the same pair of nodes), use a weighted average method to merge the edge weights. Step D: Add time-dimensional edges to the global network to connect corresponding nodes of the same drug unit in adjacent time slices, forming a spatiotemporally continuous circulation trajectory; perform topology optimization on the whole process circulation model: delete weakly related edges with weights lower than 30% of the global average weight to eliminate false associations caused by data noise; detect and merge duplicate paths, and for multiple paths with the same starting point and ending point and the same sequence of intermediate nodes, retain the one with the highest weight. Step E: Establish a hash index table with the drug unit GUID as the primary key and the distribution path as the value, supporting fast path retrieval based on drug units; establish an inverted index with the distribution entity as the key and the associated drug unit set as the value, supporting batch traceability queries based on the entity; persist the assembled full-process distribution model as a graph database storage format, with nodes storing entity information and link attributes, and edges storing process association coefficients and timestamp information.
[0061] It should be noted that, for the graph structure splicing process, firstly, a global identity identifier for each drug unit is established using GUIDs to solve the problem of identity determination across time slices; secondly, differentiated splicing strategies are formulated based on distribution modalities to reflect the network evolution patterns under different distribution modes (gradual evolution of routine distribution, abrupt switching of emergency states); then, topology integration is achieved through node alignment and edge weight fusion to ensure the consistency of the global network; finally, temporal enhancement and topology optimization are used to improve the quality and usability of the model. This method forms a complete data flow from the local network to the global model, supporting multi-granularity traceability analysis from individual drug units to the overall distribution topology.
[0062] Thus, the entire process distribution model of the target drug is obtained through the above methods.
[0063] Step S004: Generate a visual diagram of the entire process flow trajectory for each drug unit using the full-process flow model.
[0064] Specifically, firstly, for each drug unit of the target drug, starting from the starting node in the production process, the path is traversed along the edge weight from high to low until the terminal node in the usage process is reached, thus obtaining the circulation path sequence of each drug unit, where each node contains standardized main information.
[0065] During the path sequence acquisition process, the edge with the highest correlation strength is selected first to ensure that the path is based on real circulation evidence; if there are multiple equivalent paths, that is, the correlation strength is the same, the path with the best time continuity is selected, that is, the time slice sequence is the most coherent.
[0066] Then, a full-process flow trajectory visualization is generated. In the full-process flow trajectory visualization, the node labels display standardized subject names, and the node colors are mapped based on the subject type.
[0067] Finally, add a timestamp next to the node, with the timestamp format uniformly set to ISO8601.
[0068] By following the above steps, the entire process of drug distribution traceability can be completed.
[0069] Please see Figure 2 The diagram illustrates a structural block diagram of a drug full-process circulation traceability system according to an embodiment of the present invention. The system includes a memory 202, a processor 201, and a computer program 2021 stored in the memory 202 and executable on the processor 201. When the processor 201 executes the computer program, it implements any of the steps of the drug full-process circulation traceability method described above.
[0070] The memory can be volatile or non-volatile, or may include both. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which serves as an external cache. Many forms of RAM are available by way of example, but not limitation. Examples include Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDRSDRAM), Enhanced Synchronous DRAM (ESDRAM), Sync Link DRAM (SLDRAM), and Direct Rambus RAM (DRRAM).
[0071] The aforementioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. General-purpose processors can be microprocessors or any conventional processor. It is worth noting that the processor can be a processor supporting Advanced Reduced Instruction Set Machines (ARM) architecture.
[0072] This embodiment uses big data analysis on traceability data corresponding to a large number of drug units to determine the topological structure of the drug in the entire circulation process. Based on this topological structure, even if there are differences in the coding rules and data formats of the traceability data, it is still possible to determine the link chain of the entire circulation of a specific drug unit in the topological structure, thereby realizing full-chain traceability of drug units and improving the integrity and accuracy of traceability data.
[0073] It should be noted that the embodiments used in this example The model is only used to represent negative correlations and the results of the constraint model output are in Within this range, in specific implementations, other models with the same purpose can be substituted; this embodiment is merely an example. The description will be based on a model, without making specific limitations on it. This refers to the input of the model.
[0074] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for tracing the entire process of drug distribution, characterized in that, The method includes the following steps: Take any drug as the target drug, take each piece of the target drug as a drug unit, and uniformly sample all drug units of the target drug to obtain several drug units. Obtain multi-source traceability data of several drug units of the target drug. The multi-source traceability data includes traceability code information, date information, subject information and location information of the corresponding drug unit at each link in the circulation process. By dividing time into several time slices, and based on the traceability coding information, date information, and location information of all drug units, the correlation between drug units in different links under each information dimension is comprehensively analyzed to obtain the process correlation coefficient between drug units; combined with the process correlation coefficient of drug units in any time slice, a local circulation network in the corresponding time slice is constructed. By utilizing the process correlation coefficient of the drug unit and combining the main information of the drug unit in the circulation process, the local circulation network of all time slices is spliced into a graph structure to obtain the full-process circulation model of the target drug. A full-process circulation trajectory visualization of each drug unit is generated through a full-process circulation model; The specific method for obtaining the process correlation coefficient is as follows: Based on the traceability coding information, date information, and location information of all drug units in any stage, clusters are formed to obtain several clusters. For any adjacent drug units, calculate the time correlation between drug units based on the distribution and differences of date information in the cluster to which the drug unit belongs; The coding logic matching degree between drug units is calculated based on the traceability coding information of the drug units; the Euclidean distance between drug units is calculated based on the location information of the drug units, thereby obtaining the spatial correlation degree between drug units; Weights are preset for the temporal correlation, coding logic matching degree, and spatial correlation between drug units, and the process correlation coefficient between drug units is obtained by weighted fusion. The specific method for obtaining the local distribution network is as follows: The effective association threshold is determined based on the statistical distribution of the process association coefficient, and drug unit pairs formed by corresponding drug units when the process association coefficient is greater than or equal to the effective association threshold are selected. The combination of the main body and the links corresponding to each drug unit is used as network nodes, the relationship between drug units and their contained drug units is used as directed edges, and the process association coefficient between drug units and their contained drug units is used as the initial weight of the edges, thereby forming a local flow sub-network between adjacent links in any time slice. All adjacent links within a time slice are spliced together to form the corresponding local circulation subnetwork. Nodes are deduplicated and merged according to their main identifiers, while edges retain their directionality and weight information, thereby constructing the corresponding local circulation network.
2. The method for tracing the entire process of drug distribution according to claim 1, characterized in that, The specific method for obtaining the temporal correlation between the drug units is as follows: The time evolution coefficients between drug units are calculated based on the discreteness and distribution level of the time information of drug units in adjacent links within their respective clusters. By combining the time evolution coefficients between drug units and the differences in corresponding time information, the time correlation between drug units is obtained.
3. The method for tracing the entire process of drug distribution according to claim 1, characterized in that, The specific method for obtaining the coding logic matching degree between the drug units is as follows: The traceability coding information corresponding to the drug units in adjacent links is processed by a preset rule parser to obtain the corresponding coding data. The Jaccard coefficient of the corresponding coding data between drug units is used as the coding logic matching degree between drug units.
4. The method for tracing the entire process of drug distribution according to claim 1, characterized in that, The specific method for obtaining the full-process circulation model is as follows: The main information mentioned includes at least pharmaceutical companies, distribution companies, hospitals, and pharmacies involved in the distribution process; Based on the distribution of time information of drug units in any time slice and the distribution of the quantity of drugs distributed through distribution companies, the circulation mode of the corresponding time slice is obtained. By combining the circulation modes of each time slice, the local circulation networks corresponding to all time slices are spliced together to form a full-process circulation model.
5. The method for tracing the entire process of drug distribution according to claim 4, characterized in that, The specific method for obtaining the flow mode of the time slice is as follows: For any time slice, the average time interval from production to use of the drug unit within the time slice is used as the flow rate characteristic of the corresponding time slice; The standard deviation of the number of drug units distributed by all distribution companies in the distribution process is calculated, and the concentration characteristics of the corresponding time slices are calculated based on the standard deviation, where the standard deviation and concentration characteristics are negatively correlated. Among all distribution companies, the distribution company that obtains more drug units than the average number of drug units distributed by all distribution companies is denoted as the target distribution company. The proportion of the target distribution company among all distribution companies is used as the main activity feature of the corresponding time slice. Using the circulation speed characteristics, concentration characteristics, and subject activity characteristics of time slices as inputs, a pre-trained neural network is used to identify the circulation mode corresponding to the time slice. The circulation mode includes conventional distribution mode, emergency direct distribution mode, and hybrid mode.
6. The method for tracing the entire process of drug distribution according to claim 1, characterized in that, The specific method for obtaining the time slice is as follows: Based on the date information of all drug units, a circulation time series is constructed. The circulation time series of the target drug is divided into several sequence segments according to a preset time granularity. The interval formed by the maximum and minimum values of the elements in each sequence segment is used as the corresponding time slice.
7. A drug whole-process circulation traceability system, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the drug whole-process circulation traceability method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Drug identification information intelligent association method and system based on multi-source data fusion
CN120494855A
Drug temperature and humidity digital monitoring and quality tracing system based on full life cycle
CN121616159A