Drug delivery tube and drug delivery amount auxiliary decision-making method

By designing a drug delivery tube with airbags and endoscopes, the problem of poor fit between the existing drug delivery tube and minimally invasive wounds is solved, and the CLIP network assists in the decision-making of drug delivery dosage, achieving effective delivery of drug liquid and accurate decision-making of drug delivery dosage.

CN119950968APending Publication Date: 2025-05-09BEIJING FRIENDSHIP HOSPITAL CAPITAL MEDICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510050579.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The existing drug delivery tubes cannot effectively fit with minimally invasive wounds, resulting in leakage of the drug liquid and lack of endoscopic assistance, making it difficult to accurately decide the dose of drug delivery.

Method used

A drug delivery tube is designed, including the main body of the delivery tube and the drug output tube section. An airbag and an endoscope are provided at the end of the main body of the delivery tube. The drug output tube section is equipped with a leaky hole. The airbag is designed as a three-layer structure to better fit the minimally invasive wounds. At the same time, the multimodal fusion layer of the CLIP network is used to assist in the decision-making of the drug delivery dosage through the matching of wound images and drug delivery text information.

Benefits of technology

It realizes an effective fit between the drug delivery tube and the minimally invasive wound, avoids leakage of the drug liquid, and achieves accurate drug delivery decisions through endoscopy, improving the efficiency of wound recovery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119950968A_ABST
    Figure CN119950968A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical instruments, and provides a drug delivery tube and a drug delivery amount auxiliary decision-making method. Comprising a delivery pipe, a medicine output pipe section and a medicine output pipe section, an air bag is arranged at the tail end of the conveying pipe body; a plurality of leaking holes are formed in the medicine output pipe section; an endoscope is further arranged on the outer wall of the medicine output tube section and used for collecting wound images; the air bags comprise a first air bag, a second air bag and a third air bag which are sequentially arranged from top to bottom; after the first airbag, the second airbag and the third airbag are completely expanded, the radius of the second airbag is smaller than that of the first airbag and that of the third airbag, and the radius of the first airbag is equal to that of the third airbag. The device has the advantages that the medicine conveying pipe stretches into the device to obtain a wound, multi-angle wound images are obtained through an endoscope, the radius of the second air bag is smaller than that of the first air bag and that of the third air bag, the whole expanded air bag is high in two ends and low in the middle and better fits a minimally invasive opening, and the limiting effect is better.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical devices, and in particular to a drug delivery tube and a drug delivery dosage auxiliary decision method. Background Art

[0002] Existing drug delivery tubes cannot fit well with micro-wounds and easily leak drug solution; and they do not have an endoscope and cannot assist in drug dosage decisions, which is not conducive to wound recovery.

[0003] In view of this, the present invention is proposed. Summary of the invention

[0004] The purpose of the present invention is to provide a drug delivery tube and a drug delivery amount auxiliary decision method to solve the technical problems existing in the prior art.

[0005] To achieve the above object, the technical solution adopted by the present invention is: a drug delivery tube, comprising: a delivery tube, the delivery tube is composed of a delivery tube body and a drug output tube section, the drug output tube section is an extension of the delivery tube body; an air bag is arranged at the end of the delivery tube body, and the drug output tube section is located above the air bag;

[0006] The drug output tube section is provided with a plurality of leak holes, and the plurality of leak holes are used for extracting the drug solution; an endoscope is also provided on the outer wall of the drug output tube section for collecting wound images;

[0007] The airbags include a first airbag, a second airbag and a third airbag arranged in sequence from top to bottom; after the first airbag, the second airbag and the third airbag are fully expanded, the radius of the second airbag is smaller than the radius of the first airbag and the third airbag, and the radius of the first airbag is equal to the radius of the third airbag.

[0008] In an optional embodiment, one side of the lower portion of the delivery tube body is connected to a liquid inlet pipe; and the three air bags are all connected to the air delivery tube.

[0009] In an optional embodiment, the signal line of the endoscope is attached to the delivery tube and extends to the outside of the delivery tube body through the air delivery tube.

[0010] On the other hand, the present invention also provides a method for assisting decision-making on drug delivery amount, comprising:

[0011] Acquire multiple wound images taken by an endoscope and text information corresponding to each wound image, and each wound image and its corresponding text information constitute a multimodal sample, wherein the text information includes a description of wound symptoms and a drug delivery amount of the wound image;

[0012] The wound image and text information of each multimodal sample are respectively input into the image encoder and text encoder of the CLIP network to obtain the image encoding features and text encoding features of each multimodal sample;

[0013] The image encoding features and text encoding features of each multimodal sample are respectively input into the multimodal fusion layer of the CLIP network to obtain the image fusion features and text fusion features of each multimodal sample;

[0014] The training of the CLIP network is completed by maximizing the similarity between the image fusion features and the text fusion features of the same multimodal sample and minimizing the similarity between the image fusion features and the text fusion features of different multimodal samples.

[0015] The drug delivery amount in each multimodal sample is used as a text label, and a text template is constructed according to the wound symptom description of the new wound image to be decided, and each text label is substituted into the text template to obtain each text information to be matched;

[0016] Input each text information to be matched and the new wound image into the text encoder and image encoder of the trained CLIP network respectively, and finally obtain the similarity of the fusion features of each text information to be matched and the wound image to be decided;

[0017] The text label in the text information to be matched with the greatest similarity is recommended as the medication dosage for the new wound image to assist the doctor in making a medication dosage decision.

[0018] In an optional embodiment, the acquiring of multiple wound images taken by the endoscope and text information corresponding to each wound image includes:

[0019] Dividing the continuous range of drug delivery into a plurality of drug delivery intervals;

[0020] The drug delivery interval corresponding to each wound image is used as the drug delivery amount of each wound image.

[0021] In an optional embodiment, the drug delivery amount in each multimodal sample is used as a text label, and a text template is constructed according to the wound symptom description of the new wound image to be decided, and each text label is substituted into the text template to obtain each text information to be matched, including:

[0022] The wound symptom description of the new wound image to be decided is used as a fixed paragraph of the text template;

[0023] After the fixed paragraph, add the drug delivery amount guide and text label slots in sequence;

[0024] The drug delivery amount in each multimodal sample is used as a text label and added to the text label slot respectively, so as to obtain the text information to be matched respectively.

[0025] In an optional embodiment, the step of inputting the wound image and text information of each multimodal sample into the image encoder and text encoder of the CLIP network to obtain the image encoding features and text encoding features of each multimodal sample, respectively, includes: inputting the wound symptom description in each multimodal sample into the text encoder of the CLIP network to obtain the first text encoding features of each multimodal sample; inputting the complete text information in each multimodal sample into the text encoder of the CLIP network to obtain the second text encoding features of each multimodal sample;

[0026] Accordingly, the training of the CLIP network is completed by maximizing the similarity between the image fusion features and the text fusion features of the same multimodal sample and minimizing the similarity between the image fusion features and the text fusion features of different multimodal samples, including: screening multiple target multimodal samples whose similarity of the first text encoding feature is greater than a set threshold from multiple multimodal samples with the same drug delivery amount; maximizing the similarity between the image fusion features and the text fusion features of the same multimodal sample and minimizing the similarity between the image fusion features and the text fusion features of different multimodal samples by constraining the difference of the second text encoding features between the target multimodal samples to be greater than another threshold, and completing the training of the CLIP network.

[0027] In an optional embodiment, the training of the CLIP network is completed by constraining the difference of the second text encoding features between the target multimodal samples to be greater than another threshold, maximizing the similarity between the image fusion features and the text fusion features of the same multimodal sample, and minimizing the similarity between the image fusion features and the text fusion features of different multimodal samples, including:

[0028] The training of the CLIP network is completed through the following loss function:

[0029]

[0030] Among them, TE p,i and TE q,i The second text encoding features TE representing the two target multimodal samples p and q respectively p and TE q The i-th dimension of , n represents the total number of dimensions; σ represents another threshold, σ>0; T x and I y I represents the text fusion features of multimodal sample x and the image fusion features of multimodal sample y respectively; x represents the image fusion feature of the multimodal sample x; sim(,) represents the vector similarity of two features; α, β and γ represent the weight coefficients respectively.

[0031] The beneficial effects of the present invention are:

[0032] (1) The drug delivery tube in the present invention can be inserted into the wound site, and multi-angle wound images can be obtained through the endoscope. After the acquisition is completed, the drug delivery tube is limited to the minimally invasive surgical opening below the wound through the airbag, and the device is removed after the drug is delivered to the wound through the delivery tube. It should be pointed out that the radius of the second airbag is smaller than the radius of the first airbag and the third airbag, so that the inflated airbag is higher at both ends and lower in the middle, which is more suitable for the minimally invasive opening (the existing technology is low at both ends and high in the middle), and the limiting effect is better.

[0033] (2) The dosage decision-making assistance method of the present invention predicts and recommends the dosage through wound images to assist doctors in making decisions. Specifically, the embodiment of the present invention adopts the CLIP network as the basic structure of the dosage prediction model, and realizes the prediction of the dosage by matching the wound image and the dosage text. In the model training stage, in order to improve the prediction accuracy, the wound symptom description is added to the text information to distinguish different wound conditions under the same dosage; at the same time, in order to avoid the long text of the wound symptom description from drowning the information of the short dosage text, the text encoder is used to encode the separate wound symptom description and the complete text information respectively, and the text encoding features of multimodal samples with similar wound symptom descriptions but different dosages are constrained to have sufficient differences to enhance the proportion of the dosage text information, thereby ensuring the CLIP network's ability to distinguish the dosage. In the model use stage, each dosage is used as a text label, and different text templates are constructed for different wound images to be decided using the wound symptom description to expand the semantic information of the text label, improve the information content of the text encoding features and text fusion features, and ensure the accuracy of the final prediction result. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0035] Figure 1 A schematic diagram of the overall structure of a drug delivery tube provided in an embodiment of the present invention.

[0036] Figure 2 is a flow chart of a drug delivery dosage auxiliary decision method provided by an embodiment of the present invention;

[0037] Figure 3 It is a schematic diagram of a data processing flow of a CLIP network in a training phase provided by an embodiment of the present invention;

[0038] Figure 4It is a schematic diagram of a data processing flow of a CLIP network in a prediction phase provided by an embodiment of the present invention.

[0039] Among them, the accompanying drawings are marked as follows:

[0040] 1-delivery tube, 11-delivery tube body, 111-first closed end, 12-drug output tube section, 121-second closed end, 122-leakage hole, 13-liquid inlet tube; 2-airbag, 21-first airbag, 22-second airbag, 23-third airbag; 3-air delivery tube. DETAILED DESCRIPTION

[0041] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0042] It should be noted that when a component is referred to as being "fixed on" or "disposed on" another component, it may be directly or indirectly located on the other component. When a component is referred to as being "connected to" another component, it may be directly or indirectly connected to the other component. The directions or positions indicated by the terms "upper", "lower", "left", "right", "front", "back", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc. are based on the directions or positions shown in the accompanying drawings and are only for the convenience of description and cannot be understood as limitations on the present technical solution. The terms "first" and "second" are only used for the convenience of description and cannot be understood as indicating or implying relative importance or implicitly specifying the number of technical features. "Multiple" means two or more, and "several" means any number including one, unless otherwise clearly and specifically defined.

[0043] Please see attached Figure 1The purpose of this embodiment is to provide a drug delivery tube, comprising: a delivery tube 1, the delivery tube 1 is composed of a delivery tube body 11 and a drug output tube section 12, the drug output tube section 12 is an extension of the delivery tube body 11; an airbag 2 is arranged at the end of the delivery tube body 11, and the drug output tube section 12 is located above the airbag 2; a plurality of leakage holes 122 are arranged on the drug output tube section 12, and the plurality of leakage holes 122 are used for the output of liquid medicine; an endoscope is also arranged on the outer wall of the drug output tube section 12 for collecting wound images; the airbag 2 includes a first airbag 21, a second airbag 22 and a third airbag 23 arranged in sequence from top to bottom; after the first airbag 21, the second airbag 22 and the third airbag 23 are fully expanded, the radius of the second airbag 22 is smaller than the radius of the first airbag 21 and the third airbag 23, and the radius of the first airbag 21 is equal to the radius of the third airbag 23. The lower side of the delivery tube body 11 is connected to a liquid inlet pipe 13; the three air bags are connected to the air delivery pipe 3. The signal line of the endoscope is attached to the delivery tube 1 and extends to the outside of the delivery tube body 11 through the air delivery pipe 3.

[0044] The drug delivery tube is inserted to obtain the wound, and the wound image from multiple angles is obtained through the endoscope. After the acquisition is completed, the drug delivery tube is limited to the minimally invasive surgical opening below the wound through the airbag, and the device is removed after the drug is delivered to the wound through the delivery tube. It should be pointed out that the radius of the second airbag is smaller than that of the first airbag and the third airbag, so that the inflated airbag is higher at both ends and lower in the middle, which is more suitable for the minimally invasive opening (the existing technology is lower at both ends and higher in the middle), and the limiting effect is better.

[0045] Embodiment 2

[0046] The method in this embodiment can use the drug delivery tube in the first embodiment and use its endoscope to take wound images; or can use an endoscope in the prior art to take wound images and then use other available drug delivery devices to deliver drugs.

[0047] Figure 2 This is a flow chart of a method for assisting decision-making on drug delivery provided by an embodiment of the present invention. The method uses the multimodal medical data accumulated during drug delivery treatment to train a drug delivery prediction model, which can assist in making recommendations on drug delivery based on wound images before drug delivery, and provide a reference for doctors to make drug delivery decisions. Figure 2 As shown, the method specifically comprises the following steps:

[0048] S110, obtaining multiple wound images taken by an endoscope and text information corresponding to each wound image, each wound image and its corresponding text information respectively constitute a multimodal sample, wherein the text information includes a description of wound symptoms and a drug dosage of the wound image.

[0049] Specifically, in the process of drug infusion treatment, this embodiment continuously accumulates historical medical data for the drug infusion amount of the same drug, and the historical medical data includes wound images taken by the endoscope during each treatment, as well as symptom descriptions of the wound images and the final drug infusion amount. This embodiment uses the wound images, symptom descriptions and drug infusion amounts in the same treatment as a multimodal sample, thereby obtaining a large number of multimodal samples as a sample set for training the drug infusion amount prediction model in subsequent operations.

[0050] The symptom description may include wound location, wound morphology, degree of wound deterioration, etc., and the drug delivery amount may be a limited number of optional drug delivery amount intervals. For example, the continuous range of drug delivery amount may be pre-divided into multiple intervals; the interval where the drug delivery amount of each wound image is located is used as the drug delivery amount label in the multimodal sample data.

[0051] S120 , inputting the wound image and text information of each multimodal sample into the image encoder and text encoder of the CLIP network respectively, and obtaining the image coding features and text coding features of each multimodal sample respectively.

[0052] S130, inputting the image coding features and text coding features of each multimodal sample into the multimodal fusion layer of the CLIP network respectively, and obtaining the image fusion features and text fusion features of each multimodal sample respectively.

[0053] S140, completing the training of the CLIP network by constraining the maximization of the similarity between the image fusion features and the text fusion features of the same multimodal sample and the minimization of the similarity between the image fusion features and the text fusion features of different multimodal samples.

[0054] S120-S140 are the basic operations of the drug delivery prediction model training stage. To illustrate these operations, the basic structure of the drug delivery prediction model is first introduced. Specifically, this embodiment uses the CLIP network as the basic structure of the drug delivery prediction model. Figure 3 As shown in the figure, the CLIP network includes a text encoder (Text Encoder), an image encoder (ImageEncoder), a multimodal fusion layer and a similarity calculation layer. Among them, the text encoder is used to encode the text information in the multimodal sample to obtain text encoding features; the image encoder is used to encode the image in the multimodal sample to obtain image encoding features; the multimodal fusion layer is used to convert the text encoding features and the image encoding features into a unified embedding space so that they have the same dimension. The converted text encoding features are called text fusion features, and the converted image encoding features are called image fusion features; the similarity calculation layer is used to calculate the similarity between each pair of text fusion features and image fusion features (I in the figure). i ·T jrepresents the calculation of similarity).

[0055] Based on the above basic structure, in the model training stage, the present embodiment first inputs the wound image of each multimodal sample into the image encoder, and the corresponding text information into the text encoder, and then processes the two encoders, the multimodal fusion layer and the similarity calculation layer in sequence to obtain the similarity between each pair of text fusion features and image fusion features. Then, the training of the CLIP network is completed by constraining the maximization of the similarity between the image fusion features and the text fusion features of the same multimodal sample, and the minimization of the similarity between the image fusion features and the text fusion features of different multimodal samples.

[0056] Optionally, the above constraints can be implemented by constructing a corresponding loss function. After each batch of multimodal samples is processed, the model parameters of the CLIP network are updated according to the loss function. Then, in the trained CLIP network, only the similarity between the image fusion features and the text fusion features from the same multimodal sample is the largest, and the similarities between the remaining image fusion features and the text fusion features from different multimodal samples are very small. Accordingly, in the subsequent prediction stage, the text information corresponding to the wound image can be determined by selecting the largest similarity, including the appropriate amount of medication corresponding to the wound image.

[0057] Compared with the traditional method of using CLIP network for image detection, this embodiment uses the dosage of wound images as the text information of the image for the application scenario of auxiliary decision-making of dosage, so as to achieve the matching between wound images and dosage. At the same time, since different wound conditions will correspond to the same dosage of wounds, and the diversity of wound conditions makes it very difficult to learn the common laws of different wound conditions under the same dosage of wounds, this embodiment also adds the description of wound symptoms of wound images to the text information to distinguish different wound conditions, improve the accuracy of model prediction, and accelerate the convergence speed of the model.

[0058] Furthermore, since the text information of the medication dosage is usually relatively short in actual applications, and the length of the wound symptom description text varies and is usually longer than the medication dosage text, in order to avoid the long text of the wound symptom description from drowning the medication dosage text information, this embodiment can also constrain the output features of the text encoder on the basis of the above-mentioned constraints, so that the difference in text encoding features corresponding to text information with similar wound symptom descriptions but different medication dosages is greater than a set threshold, thereby increasing the proportion of the medication dosage text information in the text encoding features and ensuring the CLIP network's ability to distinguish medication dosages.

[0059] In a specific embodiment, when a text encoder is used to process text information, for each multimodal sample, a separate wound symptom description can be first input into the text encoder to obtain a text encoding feature, and then the complete text information consisting of the wound symptom description and the medication dosage text can be input into the text encoder to obtain another text encoding feature. For the convenience of distinction and description, the one text encoding feature is referred to as the first text encoding feature, and the other text encoding feature is referred to as the second text encoding feature. Then, from multiple multimodal samples with the same medication dosage, multimodal samples whose similarity of the first text encoding feature is greater than a set threshold are screened, which are called target multimodal samples. The loss function is used to constrain the difference in the second text encoding feature between the target multimodal samples to be greater than another threshold, thereby ensuring that each text information with similar wound symptom descriptions but different medication dosages can also correspond to text encoding features with obvious differences.

[0060] Optionally, the CLIP network can be trained using the following loss function:

[0061]

[0062] Among them, TE p,i and TE q,i The second text encoding features TE representing the two target multimodal samples p and q respectively p and TE q The i-th dimension of , n represents the total number of dimensions; σ represents another threshold, σ>0; T x and I y I represents the text fusion features of multimodal sample x and the image fusion features of multimodal sample y respectively; x represents the image fusion feature of the multimodal sample x; sim(,) represents the vector similarity of two features, which can be Euclidean distance, cosine similarity, etc.; α, β and γ represent weight coefficients respectively.

[0063] During training, the above three constraints can be achieved by updating the model parameters by minimizing L. It is used to constrain the second text encoding features of the target multimodal samples to maintain sufficient differences, ∑ x≠y sim(T x ,I y ) is used to constrain the minimization of the similarity between the image fusion features and the text fusion features of different multimodal samples, ∑ x [1-sim(T x ,I x )] is used to constrain the maximization of the similarity between the image fusion feature and the text fusion feature of the same multimodal sample. Of course, the above three loss function terms can also be replaced by other loss function terms with corresponding constraints, and this embodiment does not make specific restrictions.

[0064] S150, taking the drug delivery amount in each multimodal sample as a text label, and constructing a text template according to the wound symptom description of the new wound image to be decided, and the text template and each text label respectively form each text information to be matched.

[0065] After the CLIP network training is completed, this embodiment uses each drug delivery amount as a text label and constructs different text templates for different new wound images to be decided, so as to expand the semantic information of the text label, improve the information content of the text encoding features and text fusion features, and ensure the accuracy of the final prediction results.

[0066] In a specific embodiment, after obtaining a new wound image to be decided, first obtain the wound symptom description of the image, and then use the wound symptom description as a fixed paragraph of the text template; and add the drug delivery amount guide and text label slot in sequence after the fixed paragraph to obtain a complete text template. For example, assuming that the wound symptom description is "the wound is located at..., the wound is red and swollen, and white pus appears", and the drug delivery amount guide is "the corresponding... drug delivery amount is", then the complete text template is: "the wound is located at..., the wound is red and swollen, and white pus appears, and the corresponding... drug delivery amount is {text label}".

[0067] Based on the above template, various drug delivery amounts are added as text labels to the text label slots respectively, and the text information to be matched can be obtained respectively.

[0068] S160, inputting each text information to be matched and the new wound image into the text encoder and image encoder of the trained CLIP network respectively, and finally obtaining the similarity of the fusion features of each text information to be matched and the new wound image.

[0069] Combination Figure 4 After each text information to be matched and the new wound image are input into the trained text encoder and image encoder, they are processed by the two encoders, the multimodal fusion layer and the similarity calculation layer in turn, and finally the similarity sim(Tr,I) between the text fusion feature Tr corresponding to each text information to be matched and the image fusion feature I of the new wound image is obtained, where r is the index of the text label.

[0070] S170: The text label in the text information to be matched with the greatest similarity is suggested as the medication dosage for the new wound image to assist the doctor in making a medication dosage decision.

[0071] Exemplary, combined Figure 4 Finally, the sim(T3,I) corresponding to the third text information to be matched is the largest, then the text label in the text information (i.e., the dosage) is the recommended dosage given by the model, which assists the doctor in making the final decision.

[0072] Furthermore, when another new wound image to be decided is obtained, the operations of S150-S170 will be repeatedly performed for the image, a new text template will be constructed for the image, a new medication dosage will be recommended, and the doctor will be assisted in making the final decision.

[0073] It should be noted that all medical data obtained in this application (including wound images and text information) are authorized by the patient or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.

[0074] In summary, this embodiment provides a method for assisting decision-making on drug dosage, which predicts and recommends drug dosage through wound images to assist doctors in making decisions. Specifically, this embodiment uses the CLIP network as the basic structure of the drug dosage prediction model, and realizes the prediction of drug dosage by matching wound images and drug dosage texts. In the model training stage, in order to improve the prediction accuracy, this embodiment adds the wound symptom description to the text information to distinguish different wound conditions under the same drug dosage; at the same time, in order to avoid the long text of the wound symptom description from drowning the information of the short text of the drug dosage, this embodiment uses a text encoder to process the separate wound symptom description and the complete text information separately, and by constraining the text encoding features of multimodal samples with similar wound symptom descriptions but different drug dosages to have sufficient differences, the proportion of drug dosage text information is strengthened, and the ability of the CLIP network to distinguish drug dosage is ensured. In the model use stage, this embodiment uses each drug dosage as a text label, and uses the wound symptom description to construct different text templates for different new wound images to be decided, so as to expand the semantic information of the text label, improve the information content of the text encoding features and the text fusion features, and ensure the accuracy of the final prediction result.

[0075] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A drug delivery tube, characterized in that: include: A delivery tube (1), the delivery tube (1) comprising a delivery tube body (11) and a drug output tube section (12), the drug output tube section (12) being an extension of the delivery tube body (11); an air bag (2) is arranged at the end of the delivery tube body (11), the drug output tube section (12) being located above the air bag (2); The drug output tube section (12) is provided with a plurality of leakage holes (122), and the plurality of leakage holes (122) are used for extracting the drug solution; an endoscope is also provided on the outer wall of the drug output tube section (12) for collecting wound images; The airbag (2) comprises a first airbag (21), a second airbag (22) and a third airbag (23) which are arranged in sequence from top to bottom; after the first airbag (21), the second airbag (22) and the third airbag (23) are fully expanded, the radius of the second airbag (22) is smaller than the radius of the first airbag (21) and the third airbag (23), and the radius of the first airbag (21) is equal to the radius of the third airbag (23).

2. The drug delivery tube according to claim 1, characterized in that: One side of the lower part of the delivery pipe body (11) is connected to a liquid inlet pipe (13); the three air bags are all connected to the air delivery pipe (3).

3. The drug delivery tube according to claim 1, characterized in that: The signal line of the endoscope is attached to the delivery tube (1) and extends to the outside of the delivery tube body (11) through the gas delivery tube (3).

4. A method for assisting decision making on drug delivery, characterized in that: include: Acquire multiple wound images taken by an endoscope and text information corresponding to each wound image, and each wound image and its corresponding text information constitute a multimodal sample, wherein the text information includes a description of wound symptoms and a drug delivery amount of the wound image; The wound image and text information of each multimodal sample are respectively input into the image encoder and text encoder of the CLIP network to obtain the image encoding features and text encoding features of each multimodal sample; The image encoding features and text encoding features of each multimodal sample are respectively input into the multimodal fusion layer of the CLIP network to obtain the image fusion features and text fusion features of each multimodal sample; The training of the CLIP network is completed by maximizing the similarity between the image fusion features and the text fusion features of the same multimodal sample and minimizing the similarity between the image fusion features and the text fusion features of different multimodal samples. The drug delivery amount in each multimodal sample is used as a text label, and a text template is constructed according to the wound symptom description of the new wound image to be decided, and each text label is substituted into the text template to obtain each text information to be matched; Input each text information to be matched and the new wound image into the text encoder and image encoder of the trained CLIP network respectively, and finally obtain the similarity of the fusion features of each text information to be matched and the wound image to be decided; The text label in the text information to be matched with the greatest similarity is recommended as the medication dosage for the new wound image to assist the doctor in making a medication dosage decision.

5. The method according to claim 4, characterized in that The obtaining of the plurality of wound images taken by the endoscope and the text information corresponding to each wound image includes: Dividing the continuous range of drug delivery into a plurality of drug delivery intervals; The drug delivery interval corresponding to each wound image is used as the drug delivery amount of each wound image.

6. The method according to claim 4, characterized in that The drug delivery amount in each multimodal sample is used as a text label, and a text template is constructed according to the wound symptom description of the new wound image to be decided, and each text label is substituted into the text template to obtain each text information to be matched, including: The wound symptom description of the new wound image to be decided is used as a fixed paragraph of the text template; After the fixed paragraph, add the drug delivery amount guide and text label slots in sequence; The drug delivery amount in each multimodal sample is used as a text label and added to the text label slot respectively, so as to obtain the text information to be matched respectively.

7. The method according to claim 4, characterized in that The wound image and text information of each multimodal sample are respectively input into the image encoder and text encoder of the CLIP network to obtain the image coding features and text coding features of each multimodal sample, including: respectively inputting the wound symptom description in each multimodal sample into the text encoder of the CLIP network to obtain the first text coding features of each multimodal sample; respectively inputting the complete text information in each multimodal sample into the text encoder of the CLIP network to obtain the second text coding features of each multimodal sample; Accordingly, the training of the CLIP network is completed by maximizing the similarity between the image fusion features and the text fusion features of the same multimodal sample and minimizing the similarity between the image fusion features and the text fusion features of different multimodal samples, including: screening multiple target multimodal samples whose similarity of the first text encoding feature is greater than a set threshold from multiple multimodal samples with the same drug delivery amount; maximizing the similarity between the image fusion features and the text fusion features of the same multimodal sample and minimizing the similarity between the image fusion features and the text fusion features of different multimodal samples by constraining the difference of the second text encoding features between the target multimodal samples to be greater than another threshold, and completing the training of the CLIP network.

8. The method according to claim 7, characterized in that The training of the CLIP network is completed by constraining the difference of the second text encoding features between the target multimodal samples to be greater than another threshold, maximizing the similarity between the image fusion features and the text fusion features of the same multimodal sample, and minimizing the similarity between the image fusion features and the text fusion features of different multimodal samples, including: The training of the CLIP network is completed through the following loss function: Among them, TE p,i and TE q,i The second text encoding features TE representing the two target multimodal samples p and q respectively p and TE q The i-th dimension of , n represents the total number of dimensions; σ represents another threshold, σ>0; T x and I y I represents the text fusion features of multimodal sample x and the image fusion features of multimodal sample y respectively; x represents the image fusion feature of the multimodal sample x; sim(,) represents the vector similarity of two features; α, β and γ represent the weight coefficients respectively.