A machine learning-based helicobacter pylori antibody drug preparation method and a blockchain application system for research and development thereof

By combining machine learning and blockchain technology, the problems of drug resistance and high cost in traditional antibiotic regimens have been solved, enabling the preparation and promotion of personalized Helicobacter pylori antibody drugs, improving treatment efficacy and reducing costs.

CN116052799BActive Publication Date: 2026-01-23CHANGZHOU INST OF DALIAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211676968.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-04-06
Filing Date
2022-12-26
Publication Date
2026-01-23
Estimated Expiration
2042-12-26

AI Technical Summary

Technical Problem

Existing technologies cannot effectively prepare specific Helicobacter pylori antibody drugs to address individual differences among patients, resulting in poor treatment outcomes. Furthermore, traditional antibiotic regimens suffer from drug resistance and high costs.

Method used

By employing a hybrid deep meta-learning approach based on machine learning and combining it with blockchain technology, this study collects and analyzes Helicobacter pylori gene sample data. Using a small-sample differential hybrid deep meta-learning model and an adaptive hierarchical clustering and typing method, it identifies the pathogen subtypes of patients and prepares personalized antibody drugs. The antibody information is recorded and disseminated using blockchain.

Benefits of technology

This enables the preparation of antibody drugs based on patient specificity, improving treatment efficacy, simplifying the preparation process, reducing costs, and encouraging patient participation in the development and application of novel antibody drugs through a blockchain incentive mechanism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116052799B_ABST
    Figure CN116052799B_ABST
Patent Text Reader

Abstract

The present application relates to the field of biological software information technology, and more particularly to a method and system for preparing an H. pylori antibody drug based on machine learning, wherein step one is to analyze the subtype and characteristics of H. pylori genes, step two is to establish an H. pylori subtype gene library, step three is to identify the specific genotype of H. pylori in patients, and step four is to prepare a specific H. pylori antibody treatment drug. The mixed deep regional family feature extraction method proposed in the present application can improve parameter sensitivity and realize model generalization through multiple simulated sampling small sample data sets based on the common knowledge of relevant patient information data, and the difference-based mixed deep meta-learning model of the small sample solves the problems of high analysis difficulty of newly emerging pathogen subtypes of H. pylori and the lack of personalized antibody drugs, thereby simplifying the personalized antibody preparation process, shortening the preparation period, and reducing the preparation cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of bio-software information technology and blockchain technology, specifically to a method for preparing Helicobacter pylori antibody drugs based on machine learning and a blockchain application system for its research and development. Background Technology

[0002] Helicobacter pylori (Hp) is a Gram-negative bacterium commonly found in the gastric mucosa of humans and non-human primates, easily leading to gastritis, gastric ulcers, and gastric cancer. In my country, over 50% of the population is infected, a large number that warrants attention. Hp was first isolated in 1983, and treatment regimens for Hp infection have been continuously improved and refined. Currently, the most widely accepted regimen is the quadruple therapy (PPI + bismuth agent + two antibiotics, 14 days). Although the safety and effectiveness of this regimen have been affirmed, it still has shortcomings in clinical practice that are difficult to overcome. These mainly include: (1) Increased drug resistance: Due to the non-standard use of antibiotics, the problem of Hp drug resistance is becoming increasingly prominent. Studies have reported that the resistance rates of Hp to tetracycline, amoxicillin, clarithromycin, levofloxacin and metronidazole are 0-12.9%, 5.5%-9.5%, 17.8%-42.1%, 27%-41.7% and 29.5%-96.8%, respectively. The multidrug resistance (MDR) rate is 25.2%, and the eradication rate of the first-line regimen has dropped to below 75%; (2) High drug costs: The variety of drugs, large dosages, long treatment courses, and the increase in eradication failure rate and recurrence rate have increased the medical expenses burden on patients; (3) Poor compliance: The medication steps are cumbersome and there are many precautions, resulting in poor patient compliance. Therefore, it is urgent to find better eradication drug alternatives.

[0003] Traditional antibiotic treatments are often ineffective against drug-resistant Helicobacter pylori variants, necessitating antibody therapy. Egg yolk antibody drugs are relatively easy to prepare, and the use of egg yolk antibody IgY as a prevention and treatment method has gained traction in China in recent years, maturing with ongoing research. However, currently available antibody drugs for treating Helicobacter pylori primarily use broad-spectrum drugs prepared by mixing Helicobacter pylori ATCC 11637, ATCC 26695, and CagA gene knockout strains in a ratio of 1–2:1–2:1–2. Due to the numerous drug-resistant subspecies of Helicobacter pylori and the potential for new mutations, broad-spectrum antigen drugs cannot be precisely targeted to each patient, resulting in an efficacy rate of only 20-30% for drug-resistant patients, often falling short of the desired therapeutic effect. Because of individual differences among patients, personalized antibodies need to be customized according to the patient's condition. Therefore, there is an urgent need to develop a drug preparation method for specific Helicobacter pylori antibodies for different patients to solve the problem that the current antibody prevention and treatment effects are not highly correlated with the patient's condition.

[0004] Machine learning is a multidisciplinary field that has emerged in the last 20 years or so. Machine learning theory primarily involves designing and analyzing algorithms that enable computers to "learn" automatically. Machine learning algorithms are a class of algorithms that automatically analyze data to obtain patterns and use these patterns to predict unknown data. In terms of algorithm design, machine learning theory focuses on feasible and effective learning algorithms.

[0005] Blockchain technology is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Blockchain is a chain-like data structure that combines data blocks sequentially according to time order, and is a distributed ledger that is cryptographically guaranteed to be immutable and unforgeable. Due to its decentralized, tamper-proof, and autonomous characteristics, blockchain is receiving increasing attention and application.

[0006] However, how to better apply machine learning and blockchain technology to the preparation and commercial promotion of drugs with specific Helicobacter pylori antibodies for different patients is a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0007] To address the above issues, this invention provides a method for preparing Helicobacter pylori antibody drugs based on machine learning and blockchain technology, employing a hybrid deep meta-learning approach, and a blockchain application system for its development. This effectively applies machine learning and blockchain technology to the research, development, preparation, and promotion of specific Helicobacter pylori antibody drugs for different patients.

[0008] The aforementioned machine learning-based method for preparing Helicobacter pylori antibody drugs includes the following steps: 1. Collecting Helicobacter pylori gene sample data; 2. Using a small-sample differential mixed deep meta-learning model to perform machine learning on the gene sample data, recording the specificity of the sample data; learning the regional family genetic mixed characteristics of the pathogen subtype data in the sample data, and establishing a pathogen subtype gene database; 3. Using an adaptive hierarchical clustering and typing method to determine the pathogen subtype of new patients and whether there are drug-immune gene mutations, and identifying the pathogen subtype of new patients; 4. Matching or preparing specific Helicobacter pylori gene antibody therapeutic drugs based on the identified pathogen subtypes.

[0009] For the technical solution described above, a further preferred embodiment is that step one involves collecting Helicobacter pylori pathogen gene sample data.

[0010] We collected existing Helicobacter pylori gene subtype data and collected different Helicobacter pylori genes from patients with insignificant treatment effects. We obtained the gene data of the samples through gene sequencing and manually analyzed and labeled the pathogen subtypes, disease characteristics, and drug-immune gene mutation sample data of different gene fragments.

[0011] For the technical solution described above, a further preferred embodiment is that step two: establishing a pathogen subtype gene database: under the condition that there is already a small amount of sample data of different pathogen gene fragments, a differential hybrid deep meta-learning model based on small samples is used to perform machine learning on the sample data, and the specificity of the sample data is recorded; the regional family genetic hybrid characteristics of the pathogen subtype data in the sample data are learned, and a pathogen subtype gene database is constructed.

[0012] For the technical solution described above, in a further preferred embodiment, step three: identifying the pathogen subtype of the new patient: using an adaptive hierarchical clustering typing method to determine the pathogen subtype of the patient and whether there is a drug-immune gene mutation. If the pathogen subtype and drug-immune gene mutation of the new patient are successfully identified, the sample data corresponding to the pathogen gene fragment of the new patient are marked into the pathogen subtype gene database established in step two. If the identification fails, the new mutation characteristics of the new patient are recorded and updated into the pathogen subtype gene database established in step two. The sample data corresponding to the pathogen gene fragment of the patient is then passed into the differential hybrid deep meta-learning model to update the model.

[0013] For the technical solution described above, a further preferred embodiment is that step four involves matching or preparing a specific Helicobacter pylori gene antibody therapeutic drug according to the identified pathogen subtype. If the pathogen subtype of the new patient is a type already existing in the pathogen subtype gene database of step two, then an existing corresponding specific IgY drug can be directly matched for treatment. If the pathogen subtype of the new patient is a new mutation type not existing in the pathogen subtype gene database of step two, then a new IgY drug corresponding to the Helicobacter pylori gene subtype of the new patient is prepared.

[0014] For the technical solution described above, in a further preferred embodiment, the differential hybrid deep meta-learning model based on small sample data in step two uses a hybrid recurrent neural network and a convolutional neural network to extract the regional family genetic features of patient information data. Combined with the meta-learning paradigm, a hybrid deep meta-learning model architecture suitable for feature learning of small sample data is designed.

[0015] For the technical solution described above, a further preferred embodiment of the differential hybrid deep meta-learning model based on small samples has a two-step architecture:

[0016] Step 1: Design a model-independent few-sample meta-learning task architecture. First, construct the learning task by sampling data multiple times from the existing patient sample data. Train the model for the learning task to obtain a gradient update, calculated as follows: In the formula, vector θ i ′ is the task T updated using the i-th gradient descent. i The parameters obtained later, α is a hyperparameter. It is model f θ The loss function is used, and then the model parameters θ are trained using different task data to optimize the model. The performance of the meta-objective function is: In the formula, TOP(T) is the function for finding the maximum value. The meta-optimization of the small sample element is performed on the parameters θ of the meta-objective function model, and the design goal of the meta-objective function model is to update task T using the i-th gradient descent. i The parameter θ obtained afterwards i The calculation is performed, and the meta-optimization between the learning tasks is carried out using stochastic gradient descent (SGD) to update the model parameters θ to: Where β represents the element step size;

[0017] Step 2: Use a hybrid deep regional family genetic feature extraction method to learn the local features of the data. The hybrid deep regional family genetic feature extraction method consists of three parts: convolutional neural network (CNN), multilayer perceptron (MLP) and long short-term memory network (LSTM). The convolutional neural network is a one-dimensional convolutional neural network.

[0018] For the technical solution described above, a further preferred embodiment of the second step is as follows:

[0019] (1) Extract slice data with the same time step d from the k sequentially collected sample data X in step one to obtain the original input data X∈R. k×d A one-dimensional convolutional neural network is used to learn the features of the original input data X: C = f(W*X + b), where C represents the learned features, W and b are the parameters of the one-dimensional convolution kernel, * represents the one-dimensional convolution operation, and f is the ReLU non-linear activation function;

[0020] (2) Add a max pooling layer to filter out the salient features that need to be retained in the high-level feature vector. Introduce a multilayer perceptron (MLP) to fit the global spatial features with the local spatial features. The multilayer perceptron mapping is represented as: O = f(W·C + b), where O represents the global spatial features and C represents the local spatial features.

[0021] (3) Introducing a long short-term memory recurrent network (LSTM) to extract regional family genetic characteristics. First, the forget gate unit f controls the amount of information of the previous local feature in the sequence. The input unit is responsible for receiving the local features of the current time step. The output unit generates the regional family genetic characteristics classified and analyzed from the patient's sample data.

[0022] (4) Input the regional family genetic characteristics of the patient’s sample data obtained after processing in steps (1)-(3) into the Softmax layer to obtain the predicted classification results and record the specificity of the patient’s sample data.

[0023] For the technical solution described above, a further preferred embodiment of step three includes adopting an adaptive Helicobacter pylori typing method based on hierarchical clustering, designing an adaptive threshold loss function to learn the optimal hierarchical division, thereby improving typing accuracy while limiting model complexity and achieving accurate classification of pathogen subtypes.

[0024] For the technical solution described above, a further preferred embodiment of the third step, the adaptive hierarchical clustering and typing method, comprises the following steps:

[0025] 1) Construct a hierarchical clustering model:

[0026] ① Normalize and imput specific values ​​in the patient sample dataset. Let the entire patient sample dataset to be analyzed be D = {x1, x2, ..., x...} n The dataset contains n samples, each with m pathogen subtypes, i.e., A = {a1, a2, ..., a...}. m All sample data were normalized, first for each pathogen subtype a. i (i = 1, ..., m), the min-max normalization method is used to map the non-missing values ​​of all sample data to the interval [0, 1]; secondly, the value 0 is used to initialize and fill all missing pathogen subtypes in the patient sample dataset D;

[0027] ② Calculate the distance correlation between pathogen subtypes of any two sample data, and define... Represents sample data x n In a m The values ​​in the dimension are represented by the formula Relation(i,j,k,l) ​​to indicate pathogen subtype a. i With pathogen subtype a j Distance correlation presented in the sample data of row k and row l in For sample data x k Central Asian type a i The value, For sample data xl Pathogen subtype a i The value, For sample data x k Pathogen subtype a j The value, For sample data x l Pathogen subtype a j The value of the correlation result can be either positive or negative. The function Score(i,j,k,l) ​​is defined to represent the pathogen subtype a. i With pathogen subtype a j The correlation score between the sample data in row k and row l is calculated using the following formula:

[0028]

[0029] ③ To calculate the correlation between any two pathogen subtypes, first define the correlation distance variable W between the two pathogen subtypes. ij , used to represent the correlation index between the i-th pathogen subtype and the j-th pathogen subtype, W ij The calculation formula is as follows: Define set P as the correlation between any two pathogen subtypes in the sample dataset D, i.e., P = {p 12 ,p 13 ,...,p 1m ,...,p (m-1)m}, where p ij Indicates pathogen subtype a i With a j The percentage of correlation, p ij The calculation formula is defined as follows:

[0030] Where the denominator represents pathogen subtype a i With pathogen subtype a j The number of comparisons in all sample data, where k represents the k-th sample data, l represents the l-th sample data, and n is the total number of all sample data. If p ij A value of 100% indicates that the pathogen subtype a is the same in any two samples in the sample dataset. i With a j Both show a positive correlation trend;

[0031] ④ Perform bottom-up hierarchical clustering of pathogen subtype set A based on set P;

[0032] 2) Validation Feedback: A feedback network is constructed between the clinical validation results and the hierarchical clustering model. The clinical validation results are fed back to the hierarchical clustering model to learn and optimize the parameter thresholds of the hierarchical clustering model and improve the hierarchical clustering of the pathogen subtypes. Then, the extraction process of the classification features of the pathogen subtypes is adaptively adjusted and optimized through the feedback network. The hierarchical clustering is classification by level, that is, each level has its own classification features. If the classification features do not conform to the current level, they are adjusted to the level where the features conform.

[0033] For the technical solution described above, a further preferred embodiment, step ④, performing bottom-up hierarchical clustering of pathogen subtype set A based on set P, specifically includes the following four steps:

[0034] (a) Based on the above steps ①-③, obtain the correlation set P between any two pathogen subtypes in the sample dataset, and express the calculation process as P = calculate(D,A);

[0035] (b) Sort the set P from largest to smallest, denoted as sort(P), and define each element in the sample dataset as a separate group, with each element having a level of 0;

[0036] (c) Define the first element of the sorted set P as p. max Let p max Represents pathogen subtype a e With pathogen subtype a f The percentage of correlation, where E represents pathogen subtype a e The current group, F represents pathogen subtype a f Given the current group and a correlation percentage threshold K, if p max If the pathogen subtype is greater than or equal to K, then E and F are grouped together, denoted as agg(E,F). e Or pathogen subtype a f If both are already members of a certain group, then group the two groups up the hierarchy; then increment the hierarchy of that group by one, and set p... max The operation of removing a set P is denoted as remove(p max The correlation percentage threshold is represented by K. When the correlation between two pathogen subtypes (or two groups) is greater than K, they are considered to be in the same hierarchical cluster.

[0037] (d) Repeat step (c) until p. max empty or p max When the value is less than K, the algorithm terminates, and the current hierarchical clustering result is the set of pathogen subtypes with the strongest correlation.

[0038] In a further preferred embodiment of the above-described technical solution, step four involves preparing specific anti-Helicobacter pylori duck (chicken) egg yolk antibodies using a high-activity IgY antibody preparation technique based on the pathogen subtype analysis results.

[0039] This invention also provides a blockchain application system for the research and development of Helicobacter pylori antibody drugs, including a data acquisition and analysis module, a gene database module, an antibody information module, and a blockchain module; the gene database module includes a machine learning module, which is a differential hybrid deep meta-learning model module; the system is applied to the execution of any of the above-described methods for preparing Helicobacter pylori antibody drugs.

[0040] A further preferred embodiment of the technical solution described above involves connecting the relevant data from the data acquisition module, gene database module, and antibody information module to the blockchain via a hardware data acquisition device. This data becomes the patient data, gene bank data, and antibody information data in the blockchain module, respectively. The patient's basic information data is recorded into the blockchain module via the data acquisition module and broadcast to other nodes. The gene bank module receives the gene data, uses a small-sample differential hybrid deep meta-learning model to accurately classify the pathogen gene sample data into pathogen subtypes, records this classification into the blockchain module, and broadcasts it to other nodes. The antibody information module receives the antibody information and records the corresponding antibody treatment drug information into the blockchain module and broadcasts it to other nodes. This process completes the comparison between the new patient and existing therapeutic antibodies.

[0041] The beneficial effects of this invention are:

[0042] (1) This invention solves the problems of high difficulty in analyzing newly emerging pathogen subtypes of Helicobacter pylori and lack of personalized antibody preparation, and simplifies the personalized antibody preparation process, shortens the preparation cycle and reduces the preparation cost.

[0043] (2) In the Helicobacter pylori subtype data feature learning of the present invention, the idea of ​​meta-learning is used to construct a model-independent meta-learning architecture, so that the model can accurately and quickly learn model parameters from a very small amount of pathogen subtype data.

[0044] (3) The hybrid deep regional family genetic feature extraction method proposed in this invention can improve parameter sensitivity and achieve model generalization based on the common knowledge of relevant patient information data through multiple simulated small sample datasets, and simultaneously optimize convolutional neural networks and recurrent neural networks to achieve the task of learning regional family genetic hybrid features that can quickly adapt to the differences in Helicobacter pylori pathogen subtype data.

[0045] (4) This study investigates an adaptive hierarchical clustering method for Helicobacter pylori typing based on clinical patient sample data. This addresses the issue that the effectiveness of antibody treatment is highly correlated with the patient's condition. Due to individual differences among patients, personalized antibodies need to be customized according to the patient's situation. However, overemphasizing patient information while neglecting the common characteristics of regional and familial inheritance can lead to overfitting in typing, making it difficult to widely apply the resulting antibodies. The root cause of the above problems lies in the typing granularity being too high or too low, manifesting as overfitting or underfitting of the typing model.

[0046] (5) This method can also encourage new patients to participate in the development of new antibody drugs through the incentive mechanism of blockchain, so that patients can participate in and share the profits of future new antibody drugs. Attached Figure Description

[0047] To better illustrate the technical solution of this invention, the following is a description of the invention with accompanying drawings:

[0048] Figure 1 This is a flowchart of the method of the present invention;

[0049] Figure 2 This is a diagram showing the amplification results of the cagA and hpaA genes in the patients of this invention;

[0050] Figure 3 This is a schematic diagram of the patient clinical sample data hybrid deep feature learning model framework of the present invention;

[0051] Figure 4 This is a structural diagram of the Helicobacter pylori adaptive hierarchical clustering method based on patient sample data according to the present invention;

[0052] Figure 5 This is a schematic diagram of hierarchical clustering based on the correlation between pathogen subtypes in this invention. Detailed Implementation

[0053] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.

[0054] Example 1

[0055] like Figure 1-5 A machine learning-based method for preparing Helicobacter pylori antibody drugs.

[0056] Step 1: Record Helicobacter pylori gene sample data and analyze the subtypes and characteristics of Helicobacter pylori genes:

[0057] We collect existing Helicobacter pylori gene subtype data, and collect different Helicobacter pylori genes from patients with insignificant treatment effects. We obtain gene data from the samples through gene sequencing, and manually analyze and label different gene fragments to identify pathogen subtypes, disease characteristics, drug-immune gene mutations, and other sample data. In other words, we use feasible methods to extract Helicobacter pylori gene sample data from patients, use gene analysis technology to analyze pathogen gene fragments, and record sample data such as pathogen classification, pathogen subtypes, disease characteristics, drug-immune gene mutations, etc., corresponding to different pathogen gene fragments.

[0058] Helicobacter pylori was collected from patients in the stomach, cultured in vitro, and the pathogenic gene CagA [Helicobacter pylori Puno135, Gene ID: 56917366] was amplified and sequenced. Case selection was performed first: This example selected Helicobacter pylori from five patients with gastric Hp-positive stomachs and severe gastric diseases preserved at Peking University First Hospital. All of these patients were resistant to amoxicillin, clarithromycin, and metronidazole.

[0059] The pathogenic bacteria obtained from clinical patients are cultured in vitro. Specifically, Helicobacter pylori solid culture medium supplemented with 0.1 wt% glucose, 0.2 wt% compound amino acids, 5 wt% fetal bovine serum, 3 wt% rabbit serum, and 5 wt% horse serum is used for streaking culture. Single, transparent, pinpoint-sized colonies are picked, passaged and purified on plates to obtain single colonies of Helicobacter pylori and establish strains.

[0060] Drug resistance tests were performed on newly established strains, specifically testing their resistance to amoxicillin, clarithromycin, and metronidazole. Note: Antibody therapy is more meaningful for patients who are resistant to antibiotics. In this example, drug-resistant strains were selected for identification. Helicobacter pylori DNA was extracted using a bacterial DNA extraction kit (TIANGEN). DNA extraction: Helicobacter pylori strains from 1.0 to 5.0E+09 were collected, and DNA was extracted using the bacterial DNA extraction kit (TIANGEN), following all instructions. The strain-specific toxin-associated protein A gene of Helicobacter pylori was selected, and primer sequences were designed (as shown in Table 1) for PCR amplification of the CagA gene.

[0061] The reaction conditions were as follows: cagA: 94℃ for 1 min, 62℃ for 1 min, 72℃ for 1 min, 72℃ for 5 min; hpaA: 94℃ for 1 min, 62℃ for 1 min, 72℃ for 1 min, 72℃ for 5 min. The amplified products were analyzed by 2% agarose gel electrophoresis and UV gel imaging. The target gene was ligated into the PMD-18T vector, transformed into *E. coli* DH5α, identified by enzyme digestion, and then sent to Sangon Biotech for sequencing. The sequences were analyzed using DNAMAN software.

[0062] Primers were designed to amplify Helicobacter pylori-specific toxin-associated proteins and adhesin A genes.

[0063] Table 1. Primer information for CagA, a protein associated with Helicobacter pylori toxin.

[0064]

[0065] The reaction conditions were as follows: cagA: 94℃ for 1 min, 62℃ for 1 min, 72℃ for 1 min, 72℃ for 5 min; hpaA: 94℃ for 1 min, 62℃ for 1 min, 72℃ for 1 min, 72℃ for 5 min. The amplified products were analyzed by 2% agarose gel electrophoresis, observed using a UV gel imaging system, and photographed.

[0066] The results are as follows Figure 2 As shown, patients 1 and 4, who were isolated, had their cagA gene amplified, while patient 5 had their hpaA gene amplified. Patients 2 and 3 were negative for both cagA and hpaA genes.

[0067] Step Two Figure 3 Establish a gene database for Helicobacter pylori pathogen subtypes.

[0068] Given a small amount of diverse pathogen gene fragments as sample data, a small-sample differential hybrid deep meta-learning model is used to perform machine learning on the sample data. The hybrid deep meta-learning method is used to learn the characteristics that can quickly adapt to the differences in Helicobacter pylori subtypes in a small sample, and the specificity of the sample data is recorded. The regional family genetic hybrid characteristics of the pathogen subtype data in the sample data are learned to construct a pathogen subtype gene database.

[0069] In pathogen subtype classification analysis, the extremely small sample size of patient information data, significant regional differences and familial characteristics, and frequent genetic recombination pose significant challenges. Therefore, this invention studies a hybrid deep meta-learning model for Helicobacter pylori subtype differences based on small samples, including a model-independent small-sample meta-learning architecture and a hybrid deep regional familial feature extraction method. The goal is to learn the hybrid regional familial features of patient information data to improve the accuracy of pathogen subtype classification.

[0070] Patient information sample data also suffers from problems such as high gene recombination rates and complex genetic variations, making it difficult for existing methods to meet the requirements of accurate classification analysis under such complex conditions with small samples. Given a limited amount of diverse pathogen gene sample data, a differential hybrid deep meta-learning model based on small samples is used to perform machine learning on the pathogen gene sample data. A hybrid recurrent neural network and convolutional neural network are employed to extract regional and familial genetic features from the patient information data. Combining this with a meta-learning paradigm, a hybrid deep meta-learning model architecture suitable for feature learning on small sample data is designed to record the specificity of the pathogen gene sample data.

[0071] First, a regional feature learning network is constructed to address the regional characteristics of patient information data. Second, a genetic feature learning network is constructed to address the genetic characteristics of patient information data. Finally, considering the small sample size of patient information data, a hybrid deep meta-learning model with feature learning generalization ability is constructed by combining the regional genetic feature learning network, enabling it to learn features that quickly adapt to the differences in pathogen subtypes in small samples.

[0072] Specifically, the differential hybrid deep meta-learning model architecture based on small samples consists of two steps:

[0073] Step 1: Design a model-independent few-shot meta-learning task architecture. Meta-learning has been widely used to solve few-shot problems. Its core idea is to train a set of initial parameters and then adjust the gradients in one or more steps based on the initial parameters to achieve the goal of quickly adapting to new tasks with only a small amount of data. In the pathogen subtype data feature learning of this invention, the idea of ​​meta-learning is adopted to construct a model-independent meta-learning architecture, which enables the model to accurately and quickly learn model parameters from a very small amount of bacterial subtype data.

[0074] First, the learning task is constructed by sampling data multiple times from the existing patient sample data. Then, the model is trained for the learning task, resulting in a gradient update, calculated using the following formula: In the formula, vector θ i ′ is the task T updated using the i-th gradient descent. i The parameters obtained later, α is a hyperparameter. It is model f θ The loss function is used, and then the model parameters θ are trained using different task data to optimize the model. The performance of the meta-objective function is: In the formula, TOP(T) is the function for finding the maximum value. The meta-optimization of the small sample element is performed on the parameters θ of the meta-objective function model, and the design goal of the meta-objective function model is to update task T using the i-th gradient descent. i The parameter θ obtained afterwards iThe calculation is performed, and the meta-optimization between the learning tasks is carried out using stochastic gradient descent (SGD) to update the model parameters θ to: Where β represents the element step size.

[0075] Step 2: Learn local features of the data using a hybrid deep regional family genetic feature extraction method. This method consists of three parts: a convolutional neural network (CNN), a multilayer perceptron (MLP), and a long short-term memory network (LSTM). The convolutional neural network used is a one-dimensional convolutional neural network. Specifically, it involves four steps:

[0076] (1) Extract slice data with the same time step d from the sample data X collected sequentially by k sensors to obtain the original input data X∈R. k×d A one-dimensional convolutional neural network is used to learn the features of the original input data X: C = f(W*X + b), where C represents the learned features, W and b are the parameters of the one-dimensional convolution kernel, * represents the one-dimensional convolution operation, and f is the ReLU non-linear activation function;

[0077] (2) In the convolutional neural network, one-dimensional filters of different lengths are used to fully perceive the spatial features of patient information data. Then, a max pooling layer is added to filter the significant features that need to be retained in the high-level feature vector. A multilayer perceptron (MLP) is introduced to combine local spatial features with global spatial features. The multilayer perceptron mapping is represented as: O = f(W·C + b), where O represents global spatial features and C represents local spatial features.

[0078] (3) A Long Short-Term Memory Recurrent Network (LSTM) is introduced to extract family genetic characteristics. In the LSTM, the family dependence of the sample data is captured through the special structure of the storage unit. First, the forget gate unit f controls the amount of information of the previous local feature in the sequence, the input unit is responsible for receiving the local features at the current time step, and the output unit generates the regional family genetic characteristics classified and analyzed from the patient's sample data.

[0079] To train the model, it's necessary to minimize the model's cross-entropy loss function based on the backpropagation algorithm. Because neural network operations are black-box operations, it's difficult to track and evaluate the results of backpropagation; experiments and parameter adjustments are needed to observe the results. In this embodiment, we assume the training sample x... i , and its corresponding true probability distribution label y i ∈R K Where K represents the number of categories, This means that the sample belongs to the j-th class, otherwise it is zero. Let Labels representing the predicted probability distribution. The cross-entropy loss function with weighted penalty term is:

[0080] Where N is the number of training samples, λ>0 is the hyperparameter for weighting the penalty term, and θ represents the trainable parameter;

[0081] The cross-entropy loss function is calculated via backpropagation through the Softmax layer as shown in the formula:

[0082] Where δ j Representing the residual terms of the output layer, δp, δo,j, δf,jδi,jδc ’ j represents the residual value of each network at the corresponding layer, and δ represents the residual value.

[0083] In Long Short-Term Memory (LSTM) recurrent network learning, the backpropagation calculation formula is:

[0084] Among them W oh W fh W ih W ch This only indicates multiple layers and has no practical meaning. In the Multilayer Perceptron (MLP) part, the loss function is passed through the minimum batch stochastic gradient descent algorithm. The passing process is as follows:

[0085]

[0086] In a convolutional neural network (CNN), the loss is denoted as E. c Then, the backpropagation of the pooling layer l in a convolutional neural network (CNN) is:

[0087] Where a c z represents the result of the activation function. c This indicates a weighted sum, and up(·) represents the upsampling operation of the pooling layer;

[0088] Backpropagation for layer l of a convolutional layer is as follows:

[0089] Where * denotes a convolution-like operation. This means convolution kernel matrix Rotate 180 degrees.

[0090] (4) Finally, the regional family genetic characteristics of the patient’s sample data obtained through the above (1)-(3) steps are input into the Softmax layer to obtain the predicted classification results and record the specificity of the patient’s sample data.

[0091] See step three Figure 4 Identify the pathogen subtype in new patients.

[0092] An adaptive hierarchical clustering-based typing method is adopted, and an adaptive threshold loss function is designed to learn the optimal hierarchical division. This improves typing accuracy while limiting model complexity, thus achieving accurate classification of pathogen subtypes.

[0093] Because antibody efficacy is closely related to patient disease characteristics, when performing pathogen subtyping treatment, it is necessary to incorporate individual information such as the patient's physical condition and clinical manifestations into the core data. However, overemphasizing patient information while neglecting common characteristics such as family and geographical location can easily lead to overfitting. Therefore, how to balance patient individuality and commonality and determine the boundary of pathogen subtyping is the core issue of pathogen subtyping. The process involves determining the patient's pathogen subtype and the presence of drug-immune gene mutations. If a new patient's pathogen subtype and drug-immune gene mutation are successfully identified, the sample data corresponding to the pathogen gene fragment of the new patient is marked into the pathogen subtype gene database established in step two. If identification fails, the patient's new mutation characteristics and disease characteristics are recorded and updated into the pathogen subtype gene database established in step two. The sample data corresponding to the patient's pathogen gene fragment is then fed into the differential hybrid deep meta-learning model to update the model.

[0094] The adaptive hierarchical clustering classification method is described in [reference needed]. Figure 5 ,

[0095] 1) Construct a hierarchical clustering model:

[0096] ① Normalize and imput specific values ​​in the patient sample dataset. Let the entire patient sample dataset to be analyzed be D = {x1, x2, ..., x...} n The dataset contains n samples, each with m pathogen subtypes, i.e., A = {a1, a2, ..., a...}. m All sample data were normalized, first for each pathogen subtype a. i (i = 1, ..., m), the min-max normalization method is used to map the non-missing values ​​of all sample data to the interval [0, 1]; secondly, the value 0 is used to initialize and fill all missing pathogen subtypes in the patient sample dataset D;

[0097] ② Calculate the distance correlation between pathogen subtypes of any two sample data, and define... Represents sample data x n In a m The values ​​in the dimension are represented by the formula Relation(i,j,k,l) ​​to indicate pathogen subtype a. i With pathogen subtype a j Distance correlation presented in the sample data of row k and row l in For sample data xk Central Asian type a i The value, For sample data x l Pathogen subtype a i The value, For sample data x k Pathogen subtype a j The value, For sample data x l Pathogen subtype a j The correlation results can be either positive or negative. Therefore, based on the learned distance correlation between subtypes in any sample data, the function Score(i,j,k,l) ​​is defined to represent the pathogen subtype a. i With pathogen subtype a j The correlation score between the sample data in row k and row l is calculated using the following formula:

[0098]

[0099] ③ To calculate the correlation between any two pathogen subtypes, first define the correlation distance variable W between the two pathogen subtypes. ij , used to represent the correlation index between the i-th pathogen subtype and the j-th pathogen subtype, W ij The calculation formula is as follows:

[0100] To more intuitively represent the correlation between pathogen subtypes, the formula in step ③ is further refined into a percentage form. Let set P be the correlation between any two pathogen subtypes in the sample dataset D, i.e., P = {p 12 ,p 13 ,...,p 1m ,...,p (m-1)m}, where p ij Indicates pathogen subtype a i With a j The percentage of correlation, p ij The calculation formula is defined as follows:

[0101] Where the denominator represents pathogen subtype a i With pathogen subtype a j The number of comparisons in all sample data, where k represents the k-th sample data, l represents the l-th sample data, and n is the total number of all sample data. If p i j A value of 100% indicates that the pathogen subtype a is the same in any two samples in the sample dataset. i With a j Both show a positive correlation trend;

[0102] ④ Perform bottom-up hierarchical clustering of pathogen subtype set A based on set P, including the following 4 steps:

[0103] (a) Based on the above steps ①-③, obtain the correlation set P between any two pathogen subtypes in the sample dataset, and express the calculation process as P = calculate(D,A);

[0104] (b) Sort the set P from largest to smallest, denoted as sort(P), and define each element in the sample dataset as a separate group, with each element having a level of 0;

[0105] (c) Define the first element of the sorted set P as p. max Assume p max Represents pathogen subtype a e With pathogen subtype a f The percentage of correlation, where E represents pathogen subtype a e The current group, F represents the current group of pathogen subtype af, given a correlation percentage threshold K, if p max If the pathogen subtype is greater than or equal to K, then E and F are grouped together, denoted as agg(E,F). e Or pathogen subtype a f If both are already members of a certain group, then group the two groups up the hierarchy; then increment the hierarchy of that group by one, and set p... max The operation of removing a set P is denoted as remove(p max The correlation percentage threshold is represented by K. When the correlation between two pathogen subtypes (or two groups) is greater than K, they are considered to be in the same hierarchical cluster group. Typically, 60% ≤ K ≤ 80%.

[0106] (d) Repeat step (c) until p. max empty or p max When the value is less than K, the algorithm terminates, and the current hierarchical clustering result is the set of pathogen subtypes with the strongest correlation.

[0107] 2) Validation Feedback: A feedback network is constructed between the clinical validation results and the hierarchical clustering model. The clinical validation results are fed back to the hierarchical clustering model to learn and optimize the parameter thresholds of the hierarchical clustering model and improve the hierarchical clustering of pathogen subtypes. Furthermore, the feedback network adaptively adjusts and optimizes the extraction process of classification features for pathogen subtypes, achieving accurate pathogen subtype classification through hierarchical clustering without affecting treatment efficacy. This avoids overfitting and provides an analytical basis for ensuring the reliability of pathogen subtype classification analysis and personalized antibody diagnosis and treatment in clinical practice. Hierarchical clustering involves classifying by levels; each level has its own classification features, and if a classification feature does not conform to the current level, it is adjusted to a level where the feature conforms.

[0108] Step 4: Match or prepare specific Helicobacter pylori gene antibody therapeutic drugs according to the pathogen subtypes identified in Step 3. If the pathogen subtype of the new patient is a type already existing in the pathogen subtype gene database of Step 2, then directly match the existing corresponding specific IgY drug for treatment; if the pathogen subtype of the new patient is a new mutation type not existing in the pathogen subtype gene database of Step 2, then prepare a new IgY drug corresponding to the Helicobacter pylori gene subtype of the new patient. In this embodiment, specific anti-Helicobacter pylori duck (chicken) egg yolk antibodies are prepared using high-activity IgY antibody preparation technology according to the analysis results.

[0109] Duck (chicken) egg yolk antibodies belong to the category of antibody drugs. Traditional broad-spectrum egg yolk antibodies are prepared by mixing multiple Helicobacter pylori pathogen subtypes in a certain proportion to prepare universal immune antibodies. However, this method suffers from insufficient specificity, resulting in poor therapeutic effects and prolonged treatment duration. Identifying the specific bacterial subtype causing the disease in patients and preparing specific antibody drugs can significantly improve the cure rate and shorten the treatment cycle. Duck (chicken) egg yolk antibodies are inexpensive and have significant advantages over antibodies prepared from rabbits, sheep, horses, cattle, etc. Furthermore, the avian immune system can recognize more immune response sites and more antigenic determinants, thus exhibiting higher titers and affinity for antibodies produced by mammalian proteins or biomolecules. Duck (chicken) egg yolk antibodies represent an antibody resource with great development potential. In actual production, the preparation of personalized antibodies requires steps such as bacterial extraction, antigen preparation, obtaining an immune host, and antibody purification. Existing methods are mostly complex processes and immature technologies with low antibody activity. This invention employs a highly active IgY antibody preparation technology to obtain smaller defective IgY molecules, enhancing their adhesion to the gastric mucosa. It also utilizes a unique anti-pathogen solution formula and optimizes the injection cycle to shorten the research and development time and enhance antibody activity.

[0110] This invention utilizes meta-learning combined with small sample sizes and a unique feature extraction model. The experimental results are shown in Table 2. After a comparative experiment on in vivo treatment with egg yolk powder, the results are shown in the table: immunized eggs were made into egg yolk powder for in vivo treatment; each type of bacteria corresponds to each person who has their oral bacteria collected, and each patient takes the egg yolk powder from one egg each time. After 10 days of administration, the effectiveness is over 100%.

[0111] Table 2 shows the changes in Helicobacter pylori C13 levels and symptoms before and after antibody administration based on the present invention.

[0112]

[0113] Example 2

[0114] A blockchain application system for Helicobacter pylori antibody drug development includes a data acquisition and analysis module, a gene database module, an antibody information module, and a blockchain module. The gene database module includes a machine learning module, which is a differential hybrid deep meta-learning model module. The system is used to execute the method described in Example 1. Specifically, the system connects relevant data from the data acquisition module, gene database module, and antibody preparation module to the blockchain via a hardware data acquisition device, becoming patient data, gene bank data, and antibody information data in the blockchain module, respectively; these also form the data base for personalized tags in the blockchain module. Basic patient information data is recorded into the blockchain module through the data acquisition module and broadcast to other nodes. The gene bank module receives the gene data, uses a small-sample-based differential hybrid deep meta-learning model to accurately classify the pathogen gene sample data into pathogen subtypes, records the classification into the blockchain module, and broadcasts it to other nodes. The antibody information module receives the antibody information and records the corresponding antibody treatment drug information into the blockchain module and broadcasts it to other nodes. This process completes the comparison between new patients and existing therapeutic antibodies. If a new patient's sample is found to be a new pathogen subtype, and the patient agrees to prepare antibodies for the new subtype and apply for a patent, the profits generated from this in the future will be shared with the patient. This smart contract is automatically completed by the blockchain module.

[0115] The embodiments described above are merely examples of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention.

Claims

1. A method for preparing Helicobacter pylori antibody drugs based on machine learning, characterized in that, Includes the following steps: Step 1: Collect Helicobacter pylori gene sample data; Step 2: Use a small-sample-based differential hybrid deep meta-learning model to perform machine learning on the gene sample data, recording the specificity of the sample data; learn the regional family genetic hybrid characteristics of the pathogen subtype data in the sample data, and establish a pathogen subtype gene database; Step 3: Use an adaptive hierarchical clustering and typing method to determine the pathogen subtype of new patients and whether there are drug-immune gene mutations, and identify the pathogen subtype of new patients; Step 4: Match or prepare specific Helicobacter pylori gene antibody therapeutic drugs for the identified pathogen subtypes; Step 2 describes the use of a small-sample-based differential hybrid deep meta-learning model that employs a hybrid Long Short-Term Memory Network (LSTM) and Convolutional Neural Network (CNN) to extract regional family genetic features from patient information data. Combined with the meta-learning paradigm, a hybrid deep meta-learning model architecture suitable for feature learning on small-sample data is designed. The architecture of the differential hybrid deep meta-learning model based on small samples includes: (1) Design a model-independent few-sample meta-learning task architecture; (2) The local features of the data are learned using a hybrid deep regional family genetic feature extraction method. The hybrid deep regional family genetic feature extraction method consists of three parts: a sequentially connected convolutional neural network (CNN), a multilayer perceptron (MLP), and a long short-term memory network (LSTM). The convolutional neural network is a one-dimensional convolutional neural network. The adaptive hierarchical clustering classification method described in step three has the following specific steps: 1) Construct a hierarchical clustering model: ① Normalize and imput specific values ​​in the patient sample dataset. Let the entire patient sample dataset to be analyzed be D = {x1, x2, ..., x...} n The dataset contains n samples, each with m pathogen subtypes, i.e., A = {a1, a2, ..., a...}. m }, normalize all sample data, first for each pathogen subtype a i (i=1,...,m), the min-max normalization method is used to map the non-missing values ​​of all sample data to the interval [0,1]; secondly, the value 0 is used to initialize and fill all missing pathogen subtypes in the patient sample dataset D; ② Calculate the distance correlation between pathogen subtypes of any two sample data, and define... Represents sample data x n In a m The values ​​in the dimension are represented by the formula Relation(i,j,k,l) ​​to indicate pathogen subtype a. i With pathogen subtype a j Distance correlation presented in the sample data of row k and row l in For sample data x k Central Asian type a i The value, For sample data x l Pathogen subtype a i The value, For sample data x k Pathogen subtype a j The value, For sample data x l Pathogen subtype a j The value of the correlation result can be either positive or negative. The function Score(i,j,k,l) ​​is defined to represent the pathogen subtype a. i With pathogen subtype a j The correlation score between the sample data in row k and row l is calculated using the following formula: ③ To calculate the correlation between any two pathogen subtypes, first define the correlation distance variable W between the two pathogen subtypes. ij , used to represent the correlation index between the i-th pathogen subtype and the j-th pathogen subtype, W ij The calculation formula is as follows: Define set P as the correlation between any two pathogen subtypes in the sample dataset D, i.e., P = {p 12 ,p 13 ,...,p 1m ,...,p (m-1)m }, where p ij Indicates pathogen subtype a i With a j The percentage of correlation, p ij The calculation formula is defined as follows: Where the denominator represents pathogen subtype a i With pathogen subtype a j The number of comparisons in all sample data, where k represents the k-th sample data, l represents the l-th sample data, and n is the total number of all sample data. If p ij A value of 100% indicates that the pathogen subtype a is the same in any two samples in the sample dataset. i With a j Both show a positive correlation trend; ④ Perform bottom-up hierarchical clustering of pathogen subtype set A based on set P; 2) Validation feedback: A feedback network is constructed between the clinical validation results and the hierarchical clustering model. The clinical validation results are fed back to the hierarchical clustering model to learn and optimize the parameter thresholds of the hierarchical clustering model and improve the hierarchical clustering of the pathogen subtype. Then, the extraction process of the classification features of the pathogen subtype is adaptively adjusted and optimized through the feedback network.

2. The method for preparing Helicobacter pylori antibody drugs based on machine learning according to claim 1, characterized in that: The steps for collecting Helicobacter pylori gene sample data include collecting existing Helicobacter pylori gene subtype data, collecting different Helicobacter pylori genes from patients with insignificant treatment effects, obtaining gene data of the samples through gene sequencing, and manually analyzing and labeling pathogen subtypes, disease characteristics, and drug-immune gene mutation sample data of different gene fragments.

3. The method for preparing Helicobacter pylori antibody drugs based on machine learning according to claim 1, characterized in that: The step three, identifying the pathogen subtype of a new patient, includes: using an adaptive hierarchical clustering typing method to determine the pathogen subtype of the new patient and whether there is a drug-immune gene mutation. If the pathogen subtype and drug-immune gene mutation of the new patient are successfully identified, the sample data corresponding to the pathogen gene fragment of the new patient are marked into the pathogen subtype gene database established in step two. If the identification fails, the new mutation characteristics of the new patient are recorded and updated into the pathogen subtype gene database established in step two. The sample data corresponding to the pathogen gene fragment of the patient is then passed into the differential hybrid deep meta-learning model to update the identification model.

4. The method for preparing Helicobacter pylori antibody drugs based on machine learning according to claim 1, characterized in that: Step four includes: matching or preparing specific Helicobacter pylori gene antibody therapeutic drugs according to the pathogen subtypes identified in step three; if the pathogen subtype of the new patient is a type already existing in the pathogen subtype gene database of step two, then directly matching the existing corresponding specific IgY drug for treatment; if the pathogen subtype of the new patient is a new mutation type not existing in the pathogen subtype gene database of step two, then preparing a new IgY drug corresponding to the Helicobacter pylori gene subtype of the new patient.

5. The method for preparing Helicobacter pylori antibody drugs based on machine learning according to claim 1, characterized in that, The architecture of the differential hybrid deep meta-learning model based on small samples includes: The (1) design of the model-independent few-sample meta-learning task architecture is as follows: First, the learning task is constructed by sampling data multiple times from the existing patient sample data. Then, the model is trained for the learning task to obtain a gradient update. The calculation formula is as follows: In the formula, vector θ i ′ is the task T updated using the i-th gradient descent. i The parameters obtained later, α is a hyperparameter. It is model f θ The loss function is used, and then the model parameters θ are trained using different task data to optimize the model. The performance of the meta-objective function is: Meta-optimization is performed on the parameters θ of the meta-objective function model, which is designed to update task T using the i-th gradient descent. i The parameter θ obtained afterwards i The calculation is performed, and the meta-optimization between the learning tasks is carried out using stochastic gradient descent (SGD) to update the model parameters θ to: Where β represents the element step size.

6. The method for preparing Helicobacter pylori antibody drugs based on machine learning according to claim 1, characterized in that, The specific steps of step two are as follows: (1) Extract slice data with the same time step d from the k sequentially collected sample data X in the first step to obtain the original input data X∈R. k×d A one-dimensional convolutional neural network is used to learn the features of the original input data X: C = f(W*X + b), where C represents the learned features, W and b are the parameters of the one-dimensional convolution kernel, * represents the one-dimensional convolution operation, and f is the ReLU non-linear activation function; (2) Add a max pooling layer to filter out the salient features that need to be retained in the high-level feature vector. Introduce a multilayer perceptron (MLP) to fit the global spatial features with the local spatial features. The multilayer perceptron mapping is represented as: O = f(W·C' + b), where O represents the global spatial features and C' is the local spatial features. (3) Introducing a long short-term memory recurrent network (LSTM) to extract regional family genetic characteristics. First, the forget gate unit f controls the amount of information of the previous local feature in the sequence. The input unit is responsible for receiving the local features of the current time step. The output unit generates the regional family genetic characteristics classified and analyzed from the patient's sample data. (4) Input the regional family genetic characteristics of the patient’s sample data obtained after processing in steps (1)-(3) into the Softmax layer to obtain the predicted classification results and record the specificity of the patient’s sample data.

7. The method for preparing Helicobacter pylori antibody drugs based on machine learning according to claim 1, characterized in that, The step ④, which involves bottom-up hierarchical clustering of pathogen subtype set A based on set P, specifically includes the following four steps: (a) Based on the above steps ①-③, obtain the correlation set P between any two pathogen subtypes in the sample dataset, and express the calculation process as P = calculate(D,A); (b) Sort the set P in descending order, denoted as sort(P), and define each element as a separate group with each element having a level of 0; (c) Define the first element of the sorted set P as p. max Let p max Representative pathogen subtype a e With pathogen subtype a f The percentage of correlation, where E represents pathogen subtype a e The current group, F represents pathogen subtype a f Given the current group and a correlation percentage threshold K, if p max If the pathogen subtype is greater than or equal to K, then E and F are grouped together, denoted as agg(E,F). e Or pathogen subtype a f If both are already members of a certain group, then group the two groups up the hierarchy; then increment the hierarchy of that group by one, and set p... max The operation of removing a set P is denoted as remove(p max The correlation percentage threshold is represented by K. When the correlation between two pathogen subtypes or two groups is greater than K, they are considered to be in the same hierarchical cluster. (d) Repeat step (c) until p. max empty or p max When the value is less than K, the algorithm terminates, and the current hierarchical clustering result is the set of pathogen subtypes with the strongest correlation.

8. The method for preparing Helicobacter pylori antibody drugs based on machine learning according to claim 1, characterized in that, Step four involves preparing specific anti-Helicobacter pylori duck or chicken egg yolk antibodies using IgY antibody technology based on the pathogen subtype analysis results.

9. A blockchain application system for the research and development of Helicobacter pylori antibody drugs, characterized in that, The system includes a data acquisition and analysis module, a gene database module, an antibody information module, and a blockchain module; the gene database module includes a machine learning module, which is a differential hybrid deep meta-learning model module; the system employs the method described in any one of claims 1-8.

10. The system as described in claim 9, characterized in that, The relevant data from the data acquisition and analysis module, gene database module, and antibody information module are uploaded to the blockchain via hardware data acquisition devices, becoming patient data, gene database data, and antibody information data in the blockchain module, respectively. The patient's basic information data is recorded into the blockchain module and broadcast to other nodes through the data acquisition and analysis module. The gene database module receives gene data, uses a small-sample differential hybrid deep meta-learning model to accurately classify pathogen gene sample data into pathogen subtypes, records the data into the blockchain module, and broadcasts it to other nodes. The antibody information module receives the antibody information, records the corresponding antibody therapy drug information into the blockchain module, and broadcasts it to other nodes.

Citation Information

Patent Citations

  • Block chain-based helicobacter pylori egg yolk antibody and preparation method thereof

    CN113845590A

  • Adaptive hierarchical clustering algorithm

    US20140037214A1