A machine learning-based pet pathogen drug resistance trend prediction system
By constructing a machine learning-based pet pathogen resistance trend prediction system, the problems of low data processing efficiency, delayed results, and insufficient feature extraction in pet pathogen resistance monitoring have been solved. This system enables real-time monitoring of resistance trends and scientific medication guidance, and improves prediction accuracy and data mining capabilities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 广西爱宠生物科技有限公司
- Filing Date
- 2026-03-30
- Publication Date
- 2026-07-03
AI Technical Summary
Existing technologies suffer from low data processing efficiency, delayed monitoring results, insufficient feature extraction capabilities, difficulty in capturing long-term nonlinear evolution trends, and a lack of efficient integration and in-depth mining of multi-source drug resistance data.
A machine learning-based pet pathogen drug resistance trend prediction system is adopted, including a multi-source heterogeneous data integration module, a data standardization preprocessing module, a high-dimensional feature representation learning module, a spatiotemporal evolution modeling module, a drug resistance trend prediction module, and a decision support and early warning feedback module. Data is collected through a distributed crawler protocol, semantic alignment and feature extraction are performed, a potential semantic association model is constructed, the diffusion patterns of drug resistance in time and space are captured, and early warning and medication guidance are provided.
It enables real-time monitoring and efficient prediction of drug resistance in pet pathogens, improves data processing efficiency, accurately captures long-term nonlinear evolution trends, provides scientific medication guidance and risk warning, reduces the overuse of antimicrobial drugs, and enhances data mining value and prediction accuracy.
Smart Images

Figure CN122337686A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of electronic digital data processing technology, and in particular to a machine learning-based system for predicting trends in drug resistance to pet pathogens. Background Technology
[0002] Current technologies for monitoring drug resistance in pet pathogens primarily rely on periodic laboratory pathogen isolation and offline drug susceptibility testing. This traditional method of manual recording and experience-based analysis results in low data processing efficiency and significant lag in monitoring results, making it difficult to reflect the dynamic drift of pathogen resistance in real time. Furthermore, existing statistical analysis models exhibit insufficient feature extraction capabilities and poor generalization performance when processing heterogeneous data from pets of different regions and breeds, failing to effectively identify the complex nonlinear relationships between multidimensional environmental factors and pathogen evolution. On the other hand, due to the lack of efficient integration and in-depth analysis of multi-source drug resistance data, existing prediction methods are often limited to single strains or short-term fluctuations, making it difficult to accurately capture long-term drug resistance evolution trends. This results in a lack of sufficient data support and decision-making basis when responding to sudden drug resistance risks.
[0003] Therefore, there is an urgent need to provide a predictive method that has high data processing efficiency, more timely monitoring results, better feature extraction capabilities, and the ability to capture long-term nonlinear evolution trends in the process of monitoring drug resistance to pet pathogens. Summary of the Invention
[0004] The purpose of this invention is to provide a pet pathogen drug resistance trend prediction system based on machine learning, which solves the problems of low data processing efficiency, delayed monitoring results, insufficient feature extraction ability, and difficulty in capturing long-term nonlinear evolution trends in the existing technology for monitoring pet pathogen drug resistance.
[0005] The technical solution of the present invention: This invention provides a machine learning-based system for predicting trends in drug resistance to pet pathogens, comprising: The multi-source heterogeneous data integration module is used to collect and integrate pet strain-related parameters in real time through a distributed crawler protocol and a standardized application programming interface, and to perform accuracy verification. The data standardization preprocessing module is used to call the built-in veterinary microbiology terminology ontology library to perform semantic alignment on the raw data and to perform estimation on the missing drug susceptibility results, generating a standardized pathogen resistance dataset. The high-dimensional feature representation learning module is used to construct heterogeneous graph structures for pathogens, antimicrobial drugs and host background, extract deep topological features between nodes, perform nonlinear dimensionality reduction, and construct a potential semantic association model. The spatiotemporal evolution modeling module is used to capture the long-range dependence of drug resistance levels within the historical 12- to 36-month period, and to analyze the diffusion topology of pathogen drug resistance in the geospatial dimension to generate evolutionary features. The drug resistance trend prediction module is used to process the evolutionary features under a multi-task learning framework and output the predicted drug resistance rate of a specific pathogen to a specific antimicrobial drug in the next 1 month, 3 months, 6 months and 12 months. The decision support and early warning feedback module is used to automatically generate clinical medication guidance suggestions and generate drug resistance risk early warning reports based on the predicted drug resistance rate value and the preset risk threshold.
[0006] Furthermore, the multi-source heterogeneous data integration module is also used to connect to a third-party cold chain transportation monitoring system through an open Internet of Things protocol to obtain temperature fluctuation data of laboratory samples during transportation; and to obtain socio-economic indicators of the area where the sampling point is located, and input the socio-economic indicators as external constraint variables into the spatiotemporal evolution modeling module.
[0007] Furthermore, the data standardization preprocessing module, when processing pet clinical diagnosis and treatment data, is also used to perform deep text mining on the host's past cases using natural language processing technology, extract the dosage of antibiotics used, the dosing cycle and treatment effect, and convert them into numerical features and input them into the high-dimensional feature representation learning module; and when the temperature fluctuation data exceeds the preset temperature range, the drug sensitivity results of the corresponding samples are marked as low weight.
[0008] Furthermore, the high-dimensional feature representation learning module is also used to perform a feature decoupling process after extracting deep topological features. By introducing an orthogonal constraint loss function, the factors affecting drug resistance are decomposed into three independent feature subspaces: biological genetic factors, environmental transmission factors, and drug-induced factors.
[0009] Furthermore, the spatiotemporal evolution modeling module, when analyzing the diffusion topology in the geospatial dimension, also introduces a centrality index based on complex network theory. By calculating the betweenness centrality and eigenvector centrality of sampling points in the pathogen transmission network, it identifies transportation hub areas that play a key role in the spread of drug resistance.
[0010] Furthermore, the drug resistance trend prediction module is also used to introduce an adversarial training mechanism, which generates an adversarial network to construct virtual drug resistance samples with perturbation characteristics, and uses the confidence interval of the predicted drug resistance rate output by the Bayesian neural network layer. When the confidence level is lower than 0.8, a secondary calibration procedure is automatically triggered to retrieve historical archive data to perform supplementary training.
[0011] Furthermore, the drug resistance trend prediction module also has a built-in synchronization interface to the global pathogen drug resistance gene database, which is used to automatically acquire the sequence characteristics of newly discovered drug resistance genes worldwide and compare them with the pathogen gene map monitored by the system; when similar fragments are found, the risk warning level of the corresponding region is automatically raised.
[0012] Furthermore, the decision support and early warning feedback module is also used to provide customized risk alerts for different breeds of pets, identify the susceptibility of specific breeds to specific drug-resistant strains by analyzing historical big data, and provide an interactive analysis interface to respond to users' custom input of antimicrobial drug dosage adjustment parameters and simulate the decline curve of drug resistance rate in the corresponding region over the next two years.
[0013] Furthermore, it also includes a closed-loop evaluation feedback module, which is used to periodically compare the predicted values output by the drug resistance trend prediction module with the actual values subsequently monitored, and automatically adjust the weight parameters in the spatiotemporal evolution modeling module according to the magnitude of the prediction error index.
[0014] Furthermore, a self-calibration module is also provided, which is used to extract real-time data from the most recent observation period every 24 hours to perform small sample learning and perform incremental correction on the output deviation of the drug resistance trend prediction module.
[0015] Compared with the prior art, the beneficial effects of the present invention include at least the following: (1) By constructing a multi-source heterogeneous data integration module and a standardized preprocessing module, this invention achieves deep integration and unified characterization of multi-dimensional data such as pet pathogens, clinical data, drugs and environment. Compared with the traditional monitoring mode that relies on single experimental data, this invention can explore the related factors that affect the occurrence of drug resistance from a wider range of dimensions, effectively solve the problem of data silos, and improve drug resistance monitoring from scattered laboratory observations to systematic global perception, effectively enhancing the mining value of raw data.
[0016] (2) This invention introduces high-dimensional feature representation learning and spatiotemporal evolution modeling technology based on machine learning, which can accurately capture the subtle drift of pathogen resistance on the time axis and the diffusion pattern in geographic space. This invention overcomes the limitations of traditional statistical models in handling nonlinear and long-term dynamic evolution, can identify complex drug resistance evolution patterns, and capture early warning signals before large-scale outbreaks of drug resistance, greatly improving the foresight and accuracy of prediction.
[0017] (3) The drug resistance trend prediction module in this invention outputs results directly related to the decision support and early warning feedback module, which can provide veterinarians with highly scientific drug use guidance and provide management departments with quantitative risk warnings. Its intelligent auxiliary decision-making mechanism replaces the traditional empirical drug use model, which can effectively curb the abuse of antibacterial drugs and reduce the evolutionary pressure of drug resistance genes from the source. Attached Figure Description
[0018] Figure 1 A schematic diagram of the system provided by the present invention; Figure 2 This is a schematic diagram of the spatiotemporal evolution modeling module in the system provided by the present invention. Detailed Implementation
[0019] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0020] To better understand this invention, it should be noted that, In the fields of animal healthcare and pathogenic biology, with the continuous growth of the pet population and the widespread use of antimicrobial drugs in small animal clinical treatment, the evolution of drug resistance in pet-borne pathogens has become a key factor affecting animal health and even public health security. Therefore, establishing a scientific and efficient system for monitoring and analyzing drug resistance trends is of paramount importance for guiding precise clinical drug use, curbing the cross-species spread of drug-resistant genes, and maintaining biosecurity. The core of research in this field lies in revealing the mechanisms of bacterial drug resistance and its temporal and spatial transmission patterns through the analysis of massive amounts of pathogen data.
[0021] Among them, the prediction of drug resistance trends of pet pathogens aims to use historically monitored pathogen information, drug susceptibility test data and relevant host background information to conduct a prospective assessment of the changes in drug resistance levels of specific species or specific drugs in the future through mathematical modeling. Effective prediction results can provide veterinarians with early warning of drug use, help medical institutions optimize antimicrobial drug management strategies, and assist regulatory authorities in formulating targeted prevention and control measures within the region.
[0022] Example Please refer to the above as well. Figures 1-2 The present invention provides a machine learning-based system for predicting the trend of drug resistance in pet pathogens, the system comprising: The multi-source heterogeneous data integration module is used to collect and integrate pet strain types, strain origin sites, minimum inhibitory concentration values, host species, sampling point geographical coordinates, and local meteorological parameters in real time through distributed crawler protocols and standardized application interfaces. It also uses a data quality assessment mechanism to perform accuracy verification on the incoming data and then remove abnormal records that do not conform to logical relationships. The data standardization preprocessing module is used to call the built-in veterinary microbiology terminology ontology library to perform semantic alignment on the raw data, and to use an iterative interpolation algorithm based on matrix factorization to estimate the missing drug sensitivity results, thereby generating a standardized pathogen resistance dataset with differential privacy technology. The high-dimensional feature representation learning module is used to construct heterogeneous graph structures for pathogens, antimicrobial drugs and host background. It uses graph convolutional neural networks to extract deep topological features between nodes and performs nonlinear dimensionality reduction by introducing an autoencoder structure with a contrastive learning mechanism, thereby constructing a potential semantic association model that reflects the commonalities of drug resistance mechanisms. The spatiotemporal evolution modeling module is used to capture the long-range dependence of drug resistance levels in the historical 12-month to 36-month range using causal convolutional networks with dilated convolutional layers. It combines graph attention model and drug selection pressure simulation operator to analyze the diffusion topology of pathogen drug resistance in the geospatial dimension and generate evolutionary features that characterize the evolutionary potential of pathogens in a specific spatiotemporal context. The drug resistance trend prediction module is used to integrate gradient boosting decision trees and deep residual networks using an ensemble learning strategy. Under the multi-task learning framework, it processes the evolutionary features generated by the spatiotemporal evolution modeling module and outputs the predicted drug resistance rate of a specific pathogen to a specific antimicrobial drug in the next 1 month, 3 months, 6 months and 12 months. The decision support and early warning feedback module is used to automatically generate clinical medication guidance suggestions based on the predicted drug resistance rate value and the preset risk threshold, and to map the regional risk value to a digital map to generate a drug resistance risk early warning report in the form of a heat map.
[0023] In some implementation processes, the role of the multi-source heterogeneous data integration module is to collect and integrate pet clinical diagnosis and treatment data, laboratory pathogen detection data, antimicrobial drug use records, and environmental epidemiological parameters from different terminals in real time. Specifically, in actual operation, the data acquired by the multi-source heterogeneous data integration module also includes multi-dimensional raw data such as host age, past medication history, sampling time, and local temperature and humidity.
[0024] Furthermore, the multi-source heterogeneous data integration module connects to a third-party cold chain transportation monitoring system through an open IoT protocol to acquire temperature fluctuation data of laboratory samples during transportation. This cross-domain data access mechanism enables the system to monitor the quality of samples throughout their entire lifecycle, from sampling to testing. Simultaneously, when collecting environmental epidemiological parameters, the multi-source heterogeneous data integration module also correlates local socioeconomic indicators, such as the per capita number of pets, the density of pet hospitals, and the consumption level of human antimicrobial drugs. These socioeconomic indicators are then used as external constraint variables input into the spatiotemporal evolution modeling module. These macroeconomic indicators, as external constraint variables, help reveal the macroeconomic background of cross-species transmission of drug resistance in the context of zoonotic diseases, thereby improving the comprehensiveness of the data.
[0025] In some implementation processes, the multi-source heterogeneous data integration module also establishes a data quality assessment mechanism, so that every piece of incoming data will be checked for completeness and accuracy. For example, the system will check whether the minimum inhibitory concentration value is within a reasonable biological range and whether the host age is consistent with the physiological state in the medical records. Then, the system will automatically remove abnormal records that do not conform to logical relationships, further ensuring that the data basis for subsequent analysis has high reliability.
[0026] Furthermore, the raw data collected by the multi-source heterogeneous data integration module flows to the data standardization preprocessing module. This module performs semantic alignment, missing value completion, and noise filtering on the raw data, thereby generating a standardized pathogen resistance dataset. Specifically, during semantic alignment, the module uses a built-in veterinary microbiology terminology ontology to convert pathogen and drug names from different institutions and using different naming conventions into international standard codes, thus eliminating ambiguity caused by naming differences. During missing value completion, the module employs an iterative interpolation algorithm based on matrix factorization, combined with statistical characteristics of similar pathogens, to scientifically estimate missing drug susceptibility results.
[0027] In some embodiments, for pet clinical diagnosis and treatment data, the data standardization preprocessing module extracts key information such as the dosage of previously used antibiotics, the dosing cycle, and the treatment effect using natural language processing technology, and converts it into numerical features. If the sample transport environment exceeds a preset temperature range, the data standardization preprocessing module will mark the drug sensitivity result of the sample with low weight to prevent false negative data caused by sample inactivation from interfering with the prediction model. Next, the generated standardized pathogen resistance dataset is processed using differential privacy technology during storage. While ensuring that the statistical characteristics of the population are accurately reflected, sensitive information about individual pets is obfuscated, thereby meeting the information security requirements for data sharing among multiple institutions.
[0028] Furthermore, the data standardization preprocessing module inputs the standardized pathogen resistance dataset into the high-dimensional feature representation learning module. The high-dimensional feature representation learning module utilizes manifold learning and embedding techniques to map the discretized features in the standardized pathogen resistance dataset to a high-dimensional continuous vector space, constructing a latent semantic association model of pathogen-drug interaction. The specific implementation process is as follows: Heterogeneous graph structures were constructed targeting pathogens, antimicrobial drugs, and host background. Graph convolutional neural networks are used to extract deep topological features between nodes; The high-dimensional features are reduced nonlinearly by using an autoencoder structure, retaining more than 95% of the feature information contribution. Generate dense vectors that reflect the commonalities of drug resistance mechanisms.
[0029] In detail, when constructing the latent semantic association model, the high-dimensional feature representation learning module introduces a contrastive learning mechanism. By bringing pathogen-drug pairs with similar resistance mechanisms closer together in the vector space, while simultaneously widening the distance between unrelated pairs, it effectively enhances the system's sensitivity to newly emerging variant strains. Next, after feature extraction, the high-dimensional feature representation learning module executes a feature decoupling process, decomposing the factors affecting drug resistance into three independent subspaces: biological genetic factors, environmental transmission factors, and drug-induced factors. This makes the system's prediction logic more interpretable, allowing technicians to retrospectively analyze the specific dominant causes leading to the increase in drug resistance.
[0030] Furthermore, please refer to the attached document. Figure 2 The dense vectors generated by the high-dimensional feature representation learning module are fed into the spatiotemporal evolution modeling module. This module, based on deep recurrent neural networks and spatial attention mechanisms, analyzes the dynamic drift patterns of pathogen resistance over time and its diffusion topology in the geographic spatial dimension. The spatiotemporal evolution modeling module consists of a time series analysis submodule and a spatial correlation analysis submodule. Specifically, the time series analysis submodule uses a causal convolutional network with dilated convolutional layers to capture the long-range dependencies of resistance levels over a historical 12- to 36-month period. The spatial correlation analysis submodule utilizes a graph attention model to calculate the weighting factors of resistance gene transmission between different regions based on the geographical proximity of sampling points and pet transport frequency, thereby identifying potential high-incidence clusters of resistance.
[0031] In some embodiments, the spatial correlation analysis submodule constructs a spatial correlation matrix by calculating the Euclidean distance between sampling points and the socioeconomic distance based on the referral density of pet hospitals, and dynamically adjusts the contribution between nodes using an attention mechanism. Simultaneously, the spatial correlation analysis submodule also introduces a centrality index based on complex network theory. By calculating the betweenness centrality and eigenvector centrality of sampling points in the pathogen transmission network, it identifies key transportation hubs or large pet distribution centers that play a crucial role in the spread of drug resistance, thereby providing a basis for decision-making to cut off the drug resistance transmission chain.
[0032] In some embodiments, the spatiotemporal evolution modeling module further includes a drug selection pressure simulation operator. This operator quantitatively assesses the impact of antibiotic selection pressure on the rate of drug resistance evolution by acquiring data on the total sales volume and prescription frequency of antimicrobial drugs within a specific region. This impact is integrated into the hidden states of a deep recurrent neural network through a nonlinear transformation function, and the calculated evolutionary state vector represents the potential drug resistance evolution of the pathogen under a specific spatiotemporal context. This process performs modeling through the following mathematical logic: in, This is the spatiotemporal evolution state vector at the current moment, which is derived from the hidden state vector at the previous moment. The input feature vector at the current time step and drug selection stress factor vector Obtained by performing a nonlinear transformation, This is the weight matrix. The term represents the bias term, and the symbol represents the activation function. This mechanism enables the prediction results to reflect fluctuations in drug resistance caused by human intervention.
[0033] Furthermore, the evolutionary features generated by the spatiotemporal evolution modeling module are input into the drug resistance trend prediction module. This module, through a multi-task learning framework, outputs predicted drug resistance rates of specific pathogens to specific antimicrobial drugs over a predetermined period. It employs an ensemble learning strategy, fusing gradient-boosting decision trees and deep residual networks. Specifically, during training, a dynamically weighted loss function is introduced, assigning higher learning weights to rare drug resistance phenotypes with smaller sample sizes, thereby addressing prediction bias caused by data imbalance.
[0034] In some embodiments, the drug resistance trend prediction module can simultaneously predict the drug resistance level change trend over the next 1 month, 3 months, 6 months, and 12 months. The module employs an adversarial training mechanism to improve robustness. Specifically, during the training phase, virtual drug resistance samples with perturbation features are constructed by generating an adversarial network, inducing the model to learn more robust decision boundaries. Simultaneously, within a multi-task learning framework, the drug resistance trend prediction module not only outputs a deterministic predicted value of the drug resistance rate but also outputs the confidence interval of the prediction result through a Bayesian neural network layer.
[0035] In some embodiments, since prediction tasks for the next 1, 3, 6, and 12 months are highly correlated, sharing the underlying representation layer enables the model to learn more robust features. Each task's branch performs personalized parameter fine-tuning based on its respective time span. Within the multi-task learning framework, the drug resistance trend prediction module achieves collaborative optimization of prediction tasks. Specifically, the Bayesian neural network layer provides a variance estimate for each prediction value by introducing the probability distribution of the weights; this variance represents the uncertainty of the prediction. When the confidence interval is too wide, the decision support module will include a warning in the output, indicating to the user that the current prediction risk is high and should be carefully considered in conjunction with clinical experience.
[0036] In detail, the drug resistance trend prediction module employs an improved weighted loss calculation method to enhance the accuracy of capturing key drug resistance evolution nodes. The principle is as follows: in, Let be the total loss function value, and N be the total number of prediction tasks, covering prediction targets across different time spans. The term represents the th... The dynamic weighting coefficients for each task are adjusted in real time based on the reciprocal of the historical prediction error. Y represents the actual observed drug resistance rate, while Y with arrows represents the predicted value. is a regularization term used to prevent overfitting of the model, where represents the regularization strength, and represents the set of model parameters.
[0037] During implementation, when the confidence level is below 0.8, the system will automatically trigger a secondary calibration procedure to retrieve earlier historical archive data for supplementary training until the predicted output reaches the preset stability standard.
[0038] In some embodiments, the drug resistance trend prediction module also incorporates a synchronization interface with a global pathogen drug resistance gene database. When novel drug-resistant genes or highly drug-resistant strains are discovered globally, the system can automatically acquire their sequence characteristics and compare them with the pathogen gene map monitored by the system. If similar fragments are found, the system will immediately raise the risk warning level for that region, thereby achieving early identification of imported drug resistance risks.
[0039] Furthermore, the output of the drug resistance trend prediction module is ultimately processed by the decision support and early warning feedback module. This module matches the output with preset risk thresholds and automatically generates clinical medication guidance and regional drug resistance risk early warning reports. Specifically, the decision support and early warning feedback module includes a dynamically updated clinical decision knowledge base. When the predicted drug resistance rate exceeds 30%, the system automatically marks the drug as requiring cautious use; when the resistance rate exceeds 60%, the system marks it as not recommended for use and simultaneously recommends alternative drug combinations with high sensitivity.
[0040] In some embodiments, the decision support and early warning feedback module has a geographic information system visualization function, which maps regional risk values onto a digital map and intuitively displays the evolution trajectory of drug resistance risk in the form of a heat map. At the same time, the system supports multi-dimensional slice analysis according to administrative divisions, animal hospital levels and pet types, thereby providing accurate data support for epidemic prevention departments to formulate differentiated regulatory policies.
[0041] In detail, users can click on any administrative region on the map to view an animation of drug resistance evolution over the past five years and compare it with the risk levels of surrounding areas. The color depth of the heatmap represents the level of drug resistance, while the direction and thickness of the arrows represent the spread trend and intensity of drug resistance risk. The system also supports segmentation by pet species, such as viewing the distribution of feline pathogen resistance to a specific antibiotic, which is of great significance for developing differentiated clinical protocols.
[0042] In detail, the decision support and early warning feedback module can provide customized risk alerts for specific pet breeds. Since different pet breeds have different physiological metabolisms, the system analyzes historical big data to identify the susceptibility of specific breeds to certain drug-resistant strains, and then achieves precise reach in the early warning report by level and category.
[0043] In some embodiments, the decision support and early warning feedback module also provides an interactive analysis interface. The decision support and early warning feedback module will send the generated early warning reports to the relevant veterinary practitioners and administrators in real time via encrypted email, mobile push and dedicated web management backend. Users can customize prediction parameters, such as simulating the decline curve of the resistance rate in the region in the next 2 years after assuming that the usage of a certain antimicrobial drug decreases by 50%.
[0044] In some embodiments, for veterinarians working in pet hospitals, early warning information is directly pushed to their work phone's mobile application; for managers in animal disease prevention departments, the system generates a detailed regional drug resistance situation analysis report every week and sends it via encrypted email; by dragging a slider on the interface to adjust the expected consumption of antimicrobial drugs, researchers can intuitively see the real-time changes in the drug resistance rate prediction curve, providing them with a powerful simulation tool to assess the potential effects of different intervention policies.
[0045] In some embodiments, the clinical decision knowledge base of the decision support and early warning feedback module adopts a combination of expert systems and deep learning. The expert system pre-sets fixed medical guidelines, such as the prohibition of certain drugs in young pets. The deep learning component dynamically adjusts the recommendation level based on the latest prediction results, thereby enabling the system to generate early warning reports that not only include textual descriptions but also automatically generate prediction curves, spatial risk distribution maps, and medication comparison tables.
[0046] In some embodiments, the system provided by this invention further includes a closed-loop evaluation feedback module. The closed-loop evaluation feedback module periodically compares the predicted value of the drug resistance trend prediction module with the actual value subsequently monitored, calculates the prediction error index, and then, based on the magnitude of the error, the closed-loop evaluation feedback module automatically adjusts the weight parameters in the spatiotemporal evolution modeling module using an online learning algorithm to achieve the self-evolution of the model.
[0047] In some embodiments, the closed-loop evaluation feedback module is automatically triggered whenever actual drug susceptibility testing data is entered into the system. It calculates not only the global mean absolute error but also the errors for different bacterial species and drugs. If the prediction error continues to rise in a specific region or for a specific drug, the system automatically triggers a retraining mechanism. The online learning algorithm allows the model to fine-tune its parameters using newly arriving data streams without interrupting service. This incremental learning mode enables the system to quickly adapt to new situations arising from sudden mutations in pathogens.
[0048] In some embodiments, the system provided by this invention further includes a self-calibration module, which performs a self-calibration once every 24 hours. By extracting real-time data from the most recent observation period to perform small-sample learning, it incrementally corrects the output deviation of the drug resistance trend prediction module, thereby ensuring that the prediction model is always synchronized with the real pathogen evolution environment.
[0049] In some embodiments, the self-calibration module runs automatically every morning at midnight. It assesses the quality of all newly accessed data in the past 24 hours and performs a small gradient descent update. This high-frequency fine-tuning ensures that the model can capture those subtle, nascent drug resistance trends.
[0050] In some embodiments, the system provided by this invention runs in a cloud-based distributed computing environment, achieving elastic scaling of computing power through containerized deployment. It can automatically allocate computing resources based on the current computing load. When faced with sudden high-concurrency queries or large-scale model training tasks, the system can horizontally expand to dozens of computing nodes within seconds, ensuring that the response time remains at an optimal level. Specifically, when processing monitoring data on a scale of millions, the response time for a single trend analysis does not exceed 5 minutes. Standardized pathogen resistance datasets are aggregated to a central database through an encrypted transmission channel. The database's storage structure is optimized using columnar storage, thereby improving the efficiency of performing cross-year retrospective queries on drug resistance trends of specific bacterial species. Specifically, compared to traditional row-based databases that require scanning a large number of irrelevant fields when performing cross-year statistics, columnar storage allows the system to read only key columns such as resistance rate, date, and bacterial species, which increases the generation speed of long-term trend charts by more than 10 times. The encrypted data transmission channel employs an advanced transport layer security protocol, combined with end-to-end encryption algorithms, to ensure that monitoring data is not tampered with or stolen throughout the entire transmission path from the terminal to the cloud.
[0051] In some implementation processes, a multi-threaded concurrent acquisition mechanism is adopted when the multi-source heterogeneous data integration module is in operation. Specifically, for pet clinical diagnosis and treatment data, the information system interface of the cooperating pet hospitals is polled every hour to obtain data items including chief symptoms, physical examination results, preliminary diagnosis, medication prescriptions, and treatment outcomes. This text information undergoes a de-identification processor before flowing into the data standardization preprocessing module. The de-identification processor uses a regular expression-based replacement algorithm to replace the pet owner's name, phone number, address, and other private information with a unique anonymous identifier.
[0052] In some embodiments, the data standardization preprocessing module employs a hybrid matching strategy based on edit distance and bidirectional encoder representation model when processing semantic alignment. Specifically, when encountering a non-standard drug name, the system first performs fuzzy matching in the local veterinary microbiology terminology ontology. If the matching score is lower than a preset threshold, deep semantic analysis is initiated to identify the chemical composition category of the drug and classify it into the closest standard terminology.
[0053] In some embodiments, the data standardization preprocessing module uses an iterative imputation algorithm based on matrix factorization to fill in missing values. Specifically, it constructs a large sparse matrix from the standardized pathogen resistance dataset, where rows represent different case samples and columns represent different antimicrobial drugs. The algorithm learns the low-rank structure of this matrix to uncover potential correlations between different drug resistances. For example, if a pathogen exhibits high-level resistance to cephalosporins, then its missing drug susceptibility values for other lactam drugs will be assigned a higher probability of resistance. This imputation method based on collective intelligence reflects biological logic better than simple mean imputation.
[0054] In some embodiments, a graph convolutional neural network is constructed in the high-dimensional feature representation learning module, which defines each pathogen, each drug, and each host context as a node in the graph, and the edges between nodes represent their interactions. For example, if a drug is frequently used to treat a disease caused by a certain pathogen, there is an edge between them, where the weight of the edge is determined by the frequency of historical treatments. The graph convolutional neural network fuses the neighborhood information of the nodes through multiple layers of aggregation operations. After three layers of convolution operations, each node obtains a feature vector containing global topological information.
[0055] In some embodiments, in the high-dimensional feature representation learning module, the encoder employs a layer-by-layer reduction neuron design, compressing the original features of thousands of dimensions into a 128-dimensional latent vector space; the decoder then attempts to recover the original features from this latent vector. By minimizing the reconstruction error, it is ensured that these 128-dimensional vectors capture the most essential drug resistance mechanism features. The feature decoupling process introduces an orthogonal constraint loss function, forcing the vectors of the three subspaces—biological genetic factors, environmental propagation factors, and drug-induced factors—to remain orthogonal, enabling technicians to independently observe the impact of changes in a particular factor on the final drug resistance trend. For example, by fixing the genetic and environmental vectors and only changing the drug-induced vector, fluctuations in drug resistance rates under different drug intensities can be simulated.
[0056] In some embodiments, when analyzing time series data, the spatiotemporal evolution modeling module utilizes a causal convolutional network with dilated convolutional layers to effectively handle historical monitoring data spanning several years. Traditional recurrent neural networks are prone to the vanishing gradient problem when processing long sequences, while causal convolution, by increasing the receptive field of the convolutional kernel, can easily cover a 36-month time window. Each layer of the dilated convolution increases the dilation rate by a power of 2, allowing the lower layers to focus on short-term monthly fluctuations while the higher layers focus on long-term annual cyclical patterns.
[0057] In some embodiments, the spatial correlation analysis submodule uses a graph attention model to calculate the influence between sampling points in real time, and this influence is not entirely dependent on physical distance. For example, two cities that are far apart but have become closely connected due to frequent pet trade or transshipment will be given a higher spatial correlation weight in the model. The spatial correlation analysis submodule uses pet transshipment frequency data to construct a dynamic flow graph, and the attention mechanism automatically identifies drug-resistant genes that may spread along these flows.
[0058] In some embodiments, the drug selection pressure simulation operator in the spatiotemporal evolution modeling module can enhance the predictive depth of the model. It establishes a nonlinear regression equation to describe the dynamic relationship between antibiotic usage and the incidence of drug resistance mutations. This operator not only considers the current drug usage but also calculates the cumulative exposure. This indicator is converted into an offset vector and injected into the computation process of the deep recurrent neural network, enabling the model to predict the consequences of human intervention.
[0059] In some embodiments, the adversarial training mechanism in the drug resistance trend prediction module generates a generator in the adversarial network that attempts to create falsified drug resistance data sufficient to confuse the drug resistance trend prediction module, for example, by fine-tuning temperature or drug dosage data. As a discriminator, the drug resistance trend prediction module must still provide correct predictions under such interference. This process prompts the module to seek out the essential characteristics that truly determine drug resistance trends, rather than relying on some accidental data coincidences, thereby improving the system's performance in extreme environments.
[0060] The system provided by this invention constructs a complete closed loop from bottom-level data acquisition to high-level decision support through the collaborative work of a multi-source heterogeneous data integration module, a data standardization preprocessing module, a high-dimensional feature representation learning module, a spatiotemporal evolution modeling module, a drug resistance trend prediction module, a decision support and early warning feedback module, a closed-loop evaluation feedback module, and a self-calibration module. The fusion of multi-source data broadens the system's perception scope, the application of deep learning algorithms enhances the depth of trend analysis, and the intelligent early warning feedback mechanism ensures that technical measures can be transformed into practical epidemic prevention results. This system demonstrates extremely high application value in both clinical medication guidance for individual pets and regional public health monitoring. Furthermore, through in-depth mining and accurate prediction of pathogen evolution patterns, this system can effectively alleviate the current situation of antibiotic abuse in the pet industry.
[0061] In detail, this system decomposes the drug resistance rate prediction task into two sub-tasks: classification and regression. The classification task is responsible for determining whether the drug resistance level will cross a specific warning threshold, such as transitioning from sensitive to resistant; the regression task is responsible for predicting the specific percentage of drug resistance. This dual-track prediction model further enhances the reference value of the output results. On the other hand, when processing geospatial data, this system also considers inter-city traffic flow data; by integrating pet logistics information obtained from third-party transportation interfaces, the system can more accurately simulate the risk of drug-resistant strains spreading across regions via transportation.
[0062] In detail, regarding the evolution of core algorithms in the logic processing layer, the high-dimensional feature representation learning module continuously incorporates the latest graphics algorithms; for example, to handle the complex gene mutation relationships of pathogens, the system introduces a dynamic graph neural network. Unlike static graphs, dynamic graphs can update the connection structure between nodes in real time based on the gain or loss of gene fragments, thus more realistically simulating the biological process of drug resistance acquisition. The spatiotemporal evolution modeling module incorporates an individual mobility model into the spatial attention mechanism to predict the interaction frequency of pets between different communities within a city.
[0063] In detail, within the decision support and early warning feedback module, the system also supports deep integration with the electronic prescription system of veterinary hospitals. When a veterinarian's antibiotic prescription does not match the current predicted drug resistance trend, the prescription system will automatically display an early warning prompt, along with a list of recommended alternative drugs and their rationale. This real-time intervention mechanism can significantly improve the coverage of scientific drug use.
[0064] In detail, during the operation of the closed-loop evaluation feedback module, each weight update is stored as a new version in the model library, and the current version is compared with the previous version. The switch is only completed when the new version outperforms the old version on the test set. This robust deployment strategy prevents performance degradation due to interference from abnormal noisy data. Through continuous iteration and self-evolution, the system provided by this invention maintains a high degree of sensitivity and accurate capture capability for pathogen resistance evolution trends in complex pet medical environments.
[0065] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A machine learning-based system for predicting trends in drug resistance to pet pathogens, characterized in that, include: The multi-source heterogeneous data integration module is used to collect and integrate pet strain-related parameters in real time through a distributed crawler protocol and a standardized application programming interface, and to perform accuracy verification. The data standardization preprocessing module is used to call the built-in veterinary microbiology terminology ontology library to perform semantic alignment on the raw data and to perform estimation on the missing drug susceptibility results, generating a standardized pathogen resistance dataset. The high-dimensional feature representation learning module is used to construct heterogeneous graph structures for pathogens, antimicrobial drugs and host background, extract deep topological features between nodes, perform nonlinear dimensionality reduction, and construct a potential semantic association model. The spatiotemporal evolution modeling module is used to capture the long-range dependence of drug resistance levels within the historical 12- to 36-month period, and to analyze the diffusion topology of pathogen drug resistance in the geospatial dimension to generate evolutionary features. The drug resistance trend prediction module is used to process the evolutionary features under a multi-task learning framework and output the predicted drug resistance rate of a specific pathogen to a specific antimicrobial drug in the next 1 month, 3 months, 6 months and 12 months. The decision support and early warning feedback module is used to automatically generate clinical medication guidance suggestions and generate drug resistance risk early warning reports based on the predicted drug resistance rate value and the preset risk threshold.
2. The prediction system according to claim 1, characterized in that, The multi-source heterogeneous data integration module is also used to connect to a third-party cold chain transportation monitoring system through an open Internet of Things protocol to obtain temperature fluctuation data of laboratory samples during transportation; and to obtain socio-economic indicators of the area where the sampling point is located, and input the socio-economic indicators as external constraint variables into the spatiotemporal evolution modeling module.
3. The prediction system according to claim 2, characterized in that, The data standardization preprocessing module, when processing pet clinical diagnosis and treatment data, is also used to perform deep text mining on the host's past cases using natural language processing technology, extract the dosage of antibiotics used, the dosing cycle and treatment effect, and convert them into numerical features and input them into the high-dimensional feature representation learning module; and when the temperature fluctuation data exceeds the preset temperature range, the drug sensitivity results of the corresponding samples are marked as low weight.
4. The prediction system according to claim 1, characterized in that, The high-dimensional feature representation learning module is also used to perform a feature decoupling process after extracting deep topological features. By introducing an orthogonal constraint loss function, the factors affecting drug resistance are decomposed into three independent feature subspaces: biological genetic factors, environmental transmission factors, and drug-induced factors.
5. The prediction system according to claim 1, characterized in that, The spatiotemporal evolution modeling module, when analyzing the diffusion topology in the geographic spatial dimension, also introduces a centrality index based on complex network theory. By calculating the betweenness centrality and eigenvector centrality of sampling points in the pathogen transmission network, it identifies transportation hub areas that play a key role in the spread of drug resistance.
6. The prediction system according to claim 1, characterized in that, The drug resistance trend prediction module is also used to introduce an adversarial training mechanism. It generates an adversarial network to construct virtual drug resistance samples with perturbation characteristics and uses the confidence interval of the predicted drug resistance rate output by the Bayesian neural network layer. When the confidence level is lower than 0.8, a secondary calibration procedure is automatically triggered to retrieve historical archive data for supplementary training.
7. The prediction system according to claim 1, characterized in that, The drug resistance trend prediction module also has a built-in synchronization interface to the global pathogen drug resistance gene database, which is used to automatically acquire the sequence characteristics of newly discovered drug resistance genes worldwide and compare them with the pathogen gene map monitored by the system; when similar fragments are found, the risk warning level of the corresponding region is automatically raised.
8. The prediction system according to claim 1, characterized in that, The decision support and early warning feedback module is also used to provide customized risk alerts for different breeds of pets, identify the susceptibility of specific breeds to specific drug-resistant strains by analyzing historical big data, and provide an interactive analysis interface to respond to users' custom input of antimicrobial drug dosage adjustment parameters and simulate the decline curve of drug resistance rate in the corresponding region over the next two years.
9. The prediction system according to claim 1, characterized in that, It also includes a closed-loop evaluation feedback module, which is used to periodically compare the predicted values output by the drug resistance trend prediction module with the actual values subsequently monitored, and automatically adjust the weight parameters in the spatiotemporal evolution modeling module according to the magnitude of the prediction error index.
10. The prediction system according to claim 1, characterized in that, It also has a self-calibration module, which extracts real-time data from the most recent observation period every 24 hours to perform small-sample learning and incrementally corrects the output bias of the drug resistance trend prediction module.