Zebrafish multi-feature fusion-based test sample efficacy prediction method and device
By using a zebrafish multi-feature fusion-based method for predicting drug efficacy, and leveraging multi-dimensional biological characteristic data and a pre-set efficacy prediction model, the method solves the problem of time-consuming traditional drug function evaluation methods and achieves rapid and accurate drug efficacy prediction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- LANXIN JUNYI (GUANGZHOU) MEDICAL TECH CO LTD
- Filing Date
- 2025-10-13
- Publication Date
- 2026-04-21
AI Technical Summary
Traditional zebrafish-based drug function evaluation methods are cumbersome and time-consuming, making it difficult to meet the high-throughput and rapid screening requirements for early drug screening.
A zebrafish-based multi-feature fusion method for predicting the efficacy of test products was adopted. By collecting multi-dimensional biological characteristic data of test products before and after intervention, a fused feature vector was formed, and consensus clustering was performed using a pre-set efficacy prediction model to predict the efficacy of test products.
It shortens the drug efficacy testing cycle and improves the efficiency and accuracy of drug screening.
Smart Images

Figure CN121260299B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of drug efficacy evaluation technology, and in particular to a method and apparatus for predicting the efficacy of test samples based on zebrafish multi-feature fusion. Background Technology
[0002] In the field of drug efficacy evaluation, zebrafish models have become an important tool for drug screening due to their advantages such as short growth cycle, intuitive phenotypic observation, and high conservation of human disease mechanisms. However, traditional zebrafish-based drug function evaluation methods (such as constructing disease models for hyperlipidemia and then performing phenotypic detection) usually involve multiple steps, including model induction, drug intervention, and indicator detection, which are cumbersome and time-consuming.
[0003] Taking the evaluation of the lipid-lowering function of forsythoside A as an example, traditional methods require establishing a high-fat zebrafish model by feeding them a high-fat diet and continuously intervening for 14 days before fat content detection and efficacy verification can be completed, which severely restricts the efficiency of early drug screening. With the increasing demand for high-throughput and rapid screening in innovative drug development, the time-consuming experimental procedures in existing technologies are no longer sufficient to meet the needs of rapid functional prediction of a large number of candidate compounds in the early stages of drug discovery.
[0004] The above content is only used to help understand the technical solution of the present invention and does not represent an admission that the above content is prior art. Summary of the Invention
[0005] The main objective of this invention is to provide a method and device for predicting the efficacy of test samples based on zebrafish multi-feature fusion, aiming to solve the technical problem that traditional zebrafish-based drug function evaluation methods usually require multiple steps such as model induction, drug intervention, and indicator detection, which are cumbersome and time-consuming.
[0006] To achieve the above objectives, the present invention provides a method for predicting the efficacy of test samples based on zebrafish multi-feature fusion, the method comprising:
[0007] Multidimensional biological characteristic data of zebrafish before and after intervention were collected, and the multidimensional biological characteristic data of zebrafish were integrated to form a fusion feature vector of the test sample.
[0008] The fusion feature vector is input into a preset efficacy prediction model, which is used to perform consensus clustering on the fusion feature vector of the test sample and the set of fusion feature vectors of clinical drugs with known efficacy, and outputs the consensus clustering result.
[0009] Based on the consensus clustering results, the test sample is associated with clinical drugs with known efficacy in the cluster to predict the efficacy of the test sample.
[0010] Optionally, the process of constructing the preset efficacy prediction model includes the following steps:
[0011] We acquire multidimensional biological characteristic data of zebrafish before and after intervention with clinical drugs with known efficacy, and integrate the multidimensional biological characteristic data of zebrafish before and after intervention with clinical drugs with known efficacy to form a fusion feature vector of clinical drugs with known efficacy.
[0012] The fusion feature vectors of the clinically known drugs with known efficacy and their corresponding efficacy are used to train a preset deep learning network to obtain the preset efficacy prediction model.
[0013] Optionally, the zebrafish multidimensional biological characteristic data includes brain nerve fluorescence image data, tail wagging movement data, and heart rate physiological data;
[0014] Accordingly, the process of integrating the multi-dimensional biological characteristic data of the zebrafish to form a fused feature vector of the test sample includes:
[0015] The T-score brain activity map is calculated based on the brain nerve fluorescence image data, the tail wag frequency is calculated based on the tail wag motion data, and the heart rate is recorded based on the heart rate physiological data.
[0016] The T-score brain activity map, the tail wag frequency, and the heart rate are integrated to form a fusion feature vector of the test sample.
[0017] Optionally, the step of calculating the T-score brain activity map based on the brain nerve fluorescence image data includes:
[0018] The brain fluorescence image data of zebrafish were aligned with the spatial coordinates of a pre-defined standardized brain template to establish a unified image coordinate system.
[0019] A preset region of interest is defined on the standardized brain template, and the preset region of interest is used to define the analysis area of subsequent neural activity signals;
[0020] For each of the preset regions of interest, the calcium transient counts before and after the intervention of the test sample are obtained, and the numerical difference between the two is calculated to generate the calcium transient difference signal of zebrafish.
[0021] For the calcium transient differential signal of zebrafish, the signal is accumulated along the longitudinal axis in three-dimensional space, and the accumulated calcium transient differential signal is mapped to a two-dimensional plane to generate a brain activity map of zebrafish.
[0022] Based on brain activity map data of the same region of interest for each group of zebrafish, the t-test method was used to calculate the t-score value at the group level, generating a group-level T-score brain activity map.
[0023] Optionally, calculating the tail-wagging frequency based on the tail-wagging motion data includes:
[0024] Collect continuous motion images of zebrafish, and use a labeling tool to select and label the outline of the zebrafish in the continuous motion images to obtain labeled data containing the coordinate information of the tail endpoint and the midpoint of the body, forming a training dataset, namely the tail wagging motion data.
[0025] A keypoint detection model is constructed based on a deep learning framework. The keypoint detection model is trained using the training dataset so that it can output the coordinates of the tail tip and the midpoint of the body of a zebrafish.
[0026] For continuously acquired current zebrafish motion images, the trained key point detection model is used to detect and obtain the coordinates of the zebrafish's tail endpoint and body midpoint in each frame, forming coordinate data containing time series information.
[0027] For coordinate data from two consecutive frames, the vector change of the tail relative to the torso is calculated, and the sequence of angle change of the zebrafish tail is obtained based on the vector change.
[0028] Periodic swaying events are detected in the sequence of angle changes. When the absolute value of the angle change between consecutive frames exceeds a preset swaying angle threshold, it is determined to be a valid tail sway. The number of valid tail sways per unit time is counted to obtain the tail swaying frequency, which represents the number of periodic movements of the zebrafish tail swaying per unit time.
[0029] Optionally, recording heart rate based on the heart rate physiological data includes:
[0030] Images of the heart region of zebrafish are collected, and the location of the heart in the heart region images is labeled using a labeling tool to form a labeled dataset, namely the heart rate physiological data.
[0031] A target organ detection model is constructed based on a deep learning framework. The target organ detection model is trained using the labeled dataset so that it can output the coordinates of the center point of the shadow region of the zebrafish heart.
[0032] The current heart region image of zebrafish is continuously acquired. The center point coordinates of the heart shadow region are detected and obtained frame by frame through a trained target organ detection model. The horizontal axis coordinates of the center point coordinates are extracted to form horizontal axis coordinate time series data corresponding to the acquisition time. The acquisition time is equal to the frame number divided by the video acquisition frame rate.
[0033] The time series data of the horizontal axis is subjected to Fourier transform, and the main frequency components are extracted within a preset frequency range. Based on the extracted main frequency components and the video acquisition frame rate, the number of reciprocating motions of the heart per unit time is calculated to obtain the heart rate of the zebrafish.
[0034] Optionally, after associating the test sample with clinical drugs with known efficacy in the cluster based on the consensus clustering results to predict the efficacy of the test sample, the method further includes:
[0035] The experiment was divided into a normal group, a model group, a positive group, and a test sample group. The normal group served as the basic reference group for the experiment, providing baseline data of physiological indicators under healthy conditions. The model group was used to verify whether the disease model was successfully constructed and to provide baseline indicators under disease conditions. The positive group was used to verify the reliability of the experimental system and to provide a reference for the effects of known effective drugs. The test sample group was used to verify whether the test sample had the expected efficacy.
[0036] The accuracy of the predictive power of the test sample was verified based on the normal group, model group, positive group, and test sample group.
[0037] Furthermore, to achieve the above objectives, the present invention also provides a test sample efficacy prediction device based on zebrafish multi-feature fusion, the test sample efficacy prediction device based on zebrafish multi-feature fusion comprising:
[0038] The data integration module is used to collect multi-dimensional biological characteristic data of zebrafish before and after the intervention of the test sample, and integrate the multi-dimensional biological characteristic data of zebrafish to form a fusion feature vector of the test sample.
[0039] The model clustering module is used to input the fusion feature vector into a preset efficacy prediction model. The preset efficacy prediction model is used to perform consensus clustering on the fusion feature vector of the test sample and the set of fusion feature vectors of clinical drugs with known efficacy, and output the consensus clustering result.
[0040] The efficacy prediction module is used to associate the test sample with clinical drugs with known efficacy in the cluster based on the consensus clustering results, and predict the efficacy of the test sample.
[0041] Furthermore, to achieve the above objectives, the present invention also provides a test sample efficacy prediction device based on zebrafish multi-feature fusion, the device comprising: a memory, a processor, and a test sample efficacy prediction program based on zebrafish multi-feature fusion stored in the memory and executable on the processor, the test sample efficacy prediction program based on zebrafish multi-feature fusion being configured to implement the steps of the test sample efficacy prediction method based on zebrafish multi-feature fusion as described above.
[0042] In addition, to achieve the above objectives, the present invention also provides a storage medium storing a test sample efficacy prediction program based on zebrafish multi-feature fusion, wherein when the test sample efficacy prediction program based on zebrafish multi-feature fusion is executed by a processor, the test sample efficacy prediction program based on zebrafish multi-feature fusion implements the steps of the test sample efficacy prediction method based on zebrafish multi-feature fusion as described in any of the above claims.
[0043] This invention provides a method for predicting the efficacy of a test product based on zebrafish multi-feature fusion. The method includes: collecting multi-dimensional biometric data of zebrafish before and after intervention with the test product; integrating the multi-dimensional biometric data of the zebrafish to form a fusion feature vector of the test product; inputting the fusion feature vector into a preset efficacy prediction model, which performs consensus clustering on the fusion feature vector of the test product with a set of fusion feature vectors of clinical drugs with known efficacy, and outputs a consensus clustering result; and based on the consensus clustering result, associating the test product with clinical drugs with known efficacy in the cluster to predict the efficacy of the test product. Compared with traditional drug function evaluation methods that construct disease models such as hyperlipidemia and then perform phenotypic detection, this invention shortens the drug efficacy detection cycle. Attached Figure Description
[0044] Figure 1 This is a schematic diagram of the structure of a test sample efficacy prediction device based on zebrafish multi-feature fusion in the hardware operating environment of the embodiment of the present invention.
[0045] Figure 2 This is a flowchart illustrating the first embodiment of the method for predicting the efficacy of test samples based on zebrafish multi-feature fusion according to the present invention.
[0046] Figure 3 This is a structural block diagram of the first embodiment of the test sample efficacy prediction device based on zebrafish multi-feature fusion of the present invention.
[0047] Figure 4 This is a schematic diagram of the zebrafish brain neural activity activation atlas of the present invention;
[0048] Figure 5 This is a heatmap showing the differences in the number of signal peaks in different brain regions of zebrafish brains according to the present invention.
[0049] Figure 6 This is a schematic diagram of zebrafish behavior images according to the present invention, where a, b, and c represent different undulation situations;
[0050] Figure 7 Images showing the results of zebrafish light field behavior photography and annotation in this invention: a: before annotation, b: after annotation;
[0051] Figure 8 This is a timing diagram of the zebrafish tail swaying in this invention;
[0052] Figure 9 Example images of zebrafish physiology taken for this invention: a. intestinal peristalsis, b. blood flow velocity estimated by tracking blood cell flow, c. heartbeat;
[0053] Figure 10 This invention presents a functional classification diagram of hypoglycemic drugs, uric acid-lowering drugs, lipid-lowering drugs, antihypertensive drugs, and antiepileptic drugs based on zebrafish brain activity, behavior, and heartbeat. A represents representative images of the zebrafish brain, behavior, and heart; B represents representative T-score BAM, tail wagging, and heart rate diagrams; C is a heatmap illustrating the statistical association between anatomical therapeutic chemical categories and identified categories; and D is a diagram of five phenotypic categories determined through consistent clustering, combining T-score BAM, tail wagging, and heart rate.
[0054] Figure 11 The present invention relates to the effect of forsythoside A on fat in zebrafish, wherein A is a visual diagram of fat deposition in zebrafish and B is a quantitative diagram of fat deposition in zebrafish.
[0055] Table 1 lists the names and uses of the drugs.
[0056] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0057] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0058] Reference Figure 1 , Figure 1 This is a schematic diagram of the structure of a test sample efficacy prediction device based on zebrafish multi-feature fusion, which is part of the hardware operating environment of the embodiment of the present invention.
[0059] like Figure 1As shown, the zebrafish multi-feature fusion-based test sample efficacy prediction device may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen, and optionally, it may also include a standard wired interface or a wireless interface. In this invention, the wired interface of the user interface 1003 may be a USB interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be a high-speed random access memory (RAM) or a non-volatile memory (NVM), such as a disk storage device. The memory 1005 may also optionally be a storage device independent of the aforementioned processor 1001.
[0060] Those skilled in the art will understand that Figure 1 The structure shown does not constitute a limitation on the test sample efficacy prediction device based on zebrafish multi-feature fusion, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0061] like Figure 1 As shown, the memory 1005, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a test sample efficacy prediction program based on zebrafish multi-feature fusion.
[0062] exist Figure 1 In the zebrafish multi-feature fusion-based test sample efficacy prediction device shown, the network interface 1004 is mainly used to connect to the backend server and communicate data with the backend server; the user interface 1003 is mainly used to connect to peripheral devices; the zebrafish multi-feature fusion-based test sample efficacy prediction device calls the zebrafish multi-feature fusion-based test sample efficacy prediction program stored in the memory 1005 through the processor 1001, and executes the zebrafish multi-feature fusion-based test sample efficacy prediction method provided in this embodiment of the invention.
[0063] Based on the above hardware structure, an embodiment of the present invention is proposed for the test sample efficacy prediction method based on zebrafish multi-feature fusion.
[0064] Reference Figure 2 , Figure 2This is a flowchart illustrating the first embodiment of the method for predicting the efficacy of test samples based on zebrafish multi-feature fusion according to the present invention. The first embodiment of the method for predicting the efficacy of test samples based on zebrafish multi-feature fusion according to the present invention is presented.
[0065] In the first embodiment, the method for predicting the efficacy of test samples based on zebrafish multi-feature fusion includes the following steps:
[0066] S10: Collect multi-dimensional biological characteristic data of zebrafish before and after the intervention of the test sample, and integrate the multi-dimensional biological characteristic data of zebrafish to form a fusion feature vector of the test sample.
[0067] It should be noted that the test sample refers to the substance to be tested for efficacy prediction, including but not limited to natural products (such as forsythoside), chemically synthesized drugs, and biological products. These are target samples whose potential pharmacological effects need to be verified using a zebrafish model. Intervention refers to the process of applying the test sample to the zebrafish model, specifically through methods including but not limited to immersion administration, microinjection, and feeding administration. Immersion administration involves placing zebrafish juveniles or embryos in an aqueous solution containing the test sample, allowing absorption through the gills or body surface. Microinjection involves directly injecting the test sample into zebrafish embryos, suitable for substances poorly soluble in water. Feeding administration involves feeding adult fish with feed containing the test sample. The core of intervention is to simulate the drug exposure process in organisms and obtain biological response data after the test sample's action. Zebrafish multidimensional biological characteristic data refers to data collected from multiple biological levels reflecting changes in the zebrafish's state before and after the test sample intervention. Integration refers to the process of preprocessing, normalizing, and fusing features of multidimensional heterogeneous data. A fusion feature vector refers to a numerical vector formed by vectorizing integrated multidimensional data. Each dimension corresponds to a selected biological feature, such as the expression level of a differentially expressed gene, the concentration change of a metabolite, or the rate of change in liver area. It is used to characterize the comprehensive effect of the test sample on zebrafish and serves as input data for subsequent efficacy prediction models.
[0068] In one embodiment, the zebrafish multidimensional biometric data may include brain nerve fluorescence image data, tail wagging motion data, and heart rate physiological data;
[0069] Accordingly, the process of integrating the multi-dimensional biological characteristic data of the zebrafish to form a fused feature vector of the test sample includes:
[0070] The T-score brain activity map is calculated based on the brain nerve fluorescence image data, the tail wag frequency is calculated based on the tail wag motion data, and the heart rate is recorded based on the heart rate physiological data.
[0071] The T-score brain activity map, the tail wag frequency, and the heart rate are integrated to form a fusion feature vector of the test sample.
[0072] It should be noted that brain nerve fluorescence image data is image data that visualizes and records neurons or neural activity in the zebrafish brain using fluorescent labeling technology. Transgenic zebrafish (such as strains expressing fluorescent proteins to label neurons, e.g., GFP-labeled excitatory neurons, RFP-labeled inhibitory neurons) or fluorescent dyes (such as the calcium indicator Fluo-4) can be used to label neural activity, and images of fluorescence signal changes in different brain regions can be acquired under a microscope. It should be understood that neural activity (such as action potentials) causes changes in intracellular calcium ion concentration in zebrafish cells, and the fluorescence intensity of the fluorescent probe changes with the calcium ion concentration. Real-time imaging using confocal microscopy or fluorescence microscopy can capture the dynamic changes in fluorescence intensity during neuronal activation, reflecting functional activity in brain regions. Brain nerve fluorescence image data can be multi-channel, multi-timepoint two-dimensional or three-dimensional images, with each pixel corresponding to the fluorescence intensity value of a specific brain region. T-score brain activity maps refer to heatmaps of significant differences in brain region activity calculated using statistical methods (T-test) based on brain nerve fluorescence image data, used to visualize changes in neural activity in different regions of the zebrafish brain before and after intervention with the test substance.
[0073] Tail wagging data, obtained through video recording and behavioral analysis, reflects the frequency and pattern of tail wagging in zebrafish, indicating their motor ability, neuromuscular coordination, or stress response. Specifically, juvenile zebrafish (e.g., 5 dpf, 5 days post-fertilization) are placed in transparent culture dishes, and side-view videos are captured using a high-speed camera (frame rate ≥30 frames / second) for 1-5 minutes. The tail movement trajectory is tracked using image recognition algorithms (e.g., OpenCV edge detection). A tail wagging event is defined as a reciprocating motion where the tail wagging angle exceeds a threshold (e.g., ±10°), and the number of tail waggings per unit time is counted. Tail wagging frequency can serve as a neurobehavioral indicator; for example, central nervous system stimulants may increase tail wagging frequency, while neurodepressants may decrease it, reflecting the effect of the test substance on motor neuron regulation.
[0074] Heart rate physiological data can be obtained by recording the heartbeat frequency of zebrafish using non-invasive optical detection, reflecting the effects of test products on the cardiovascular system. Specifically, heartbeats can be directly observed under a stereomicroscope, and the number of heartbeats can be counted manually or using image analysis software; alternatively, the zebrafish trunk can be illuminated with an LED light source, and the transmittance changes caused by heartbeats can be received by a photosensitive sensor, with the heart rate extracted via Fourier transform; or the heart rate can be calculated by observing the periodic changes in fluorescence signals from fluorescent proteins specifically expressed in cardiomyocytes. Changes in heart rate directly reflect the effects of drugs on cardiac pacing function, myocardial contractility, or autonomic nervous system regulation, and are key physiological indicators for assessing the cardiovascular toxicity or pharmacological effects of test products.
[0075] Specifically, the calculation of the T-score brain activity map based on the brain nerve fluorescence image data includes:
[0076] The brain fluorescence image data of zebrafish were aligned with the spatial coordinates of a pre-defined standardized brain template to establish a unified image coordinate system.
[0077] A preset region of interest is defined on the standardized brain template, and the preset region of interest is used to define the analysis area of subsequent neural activity signals;
[0078] For each of the preset regions of interest, the calcium transient counts before and after the intervention of the test sample are obtained, and the numerical difference between the two is calculated to generate the calcium transient difference signal of zebrafish.
[0079] For the calcium transient differential signal of zebrafish, the signal is accumulated along the longitudinal axis in three-dimensional space, and the accumulated calcium transient differential signal is mapped to a two-dimensional plane to generate a brain activity map of zebrafish.
[0080] Based on brain activity map data of the same region of interest for each group of zebrafish, the t-test method was used to calculate the t-score value at the group level, generating a group-level T-score brain activity map.
[0081] It should be noted that the standardized brain template is a pre-constructed three-dimensional structural reference map of the zebrafish brain. By averaging a large amount of brain image data from similar zebrafish, standard spatial coordinates for each brain region are defined as the benchmark for subsequent image registration. Spatial coordinate alignment uses an image registration algorithm to map the original fluorescent brain images of each zebrafish onto the coordinate system of the standardized brain template, ensuring that the pixel positions of the same brain region are consistent in different images. Pre-defined regions of interest (ROIs) are specific brain regions manually or automatically divided on the standardized brain template. Each ROI corresponds to a neural nucleus or pathway encoding a specific function, such as the tectum in the midbrain that regulates movement, or the cardiovascular center in the medulla oblongata that regulates heart rate. Calcium transients refer to the instantaneous increase in intracellular calcium ion concentration when neurons are activated, reflected by a sudden increase in the fluorescence intensity of calcium indicators, and are a direct marker of neural activity. A single action potential usually corresponds to one calcium transient, and its count reflects the neuronal firing frequency. For the fluorescence signal time series within each ROI, the number of calcium transients is counted using a thresholding method to obtain the pre-intervention count and post-intervention count. Differential signal calculation refers to calculating the difference in the number of calcium transients before and after intervention for each ROI, reflecting the direction of change in neural activity in that brain region; positive values indicate activation enhancement, and negative values indicate inhibition. Longitudinal axis signal accumulation refers to summing the three-dimensional spatial differential calcium transient signals within each ROI along the Z-axis to obtain the total signal value at each coordinate point on the two-dimensional plane. Two-dimensional brain activity map generation refers to mapping the accumulated total signal value into a pseudo-color image, forming a two-dimensional brain activity map reflecting changes in surface neural activity in the brain region. Group-level T-score calculation refers to performing a two-sample t-test on the brain activity map data of the same ROI for each group of zebrafish to calculate the T-score value.
[0082] In one specific embodiment, by analyzing the calcium imaging fluorescence signal trace data of the zebrafish brain, a brain neural activity activation map is calculated, referring to... Figure 4 , Figure 4 This is a schematic diagram of zebrafish brain neural activity activation atlas according to the present invention, where a: calm state, b: activation signal in the tail, and c: activation signal in the head. Specifically, based on digital image technology, the image segmentation data of the brain neural region of interest is preprocessed, including image data cleaning, image cropping, and image registration; let the reference image be... The image to be registered is Find the optimal transformation :
[0083]
[0084] in, Using the reference image (normalized brain template) in coordinates The pixel value (fluorescence intensity or structural grayscale value) at that location. Single-sample brain nerve fluorescence images to be registered in the original coordinates Pixel value at that location, This is a spatial transformation function that maps the coordinates (x, y) of the image to be registered to new coordinates in the coordinate system of the reference image. The optimal transformation parameters are obtained by minimizing the sum of squared differences in pixel values between the reference image and the registered image, ensuring the best spatial match between the two images.
[0085] The baseline of fluorescence signal trace was calculated by taking the mean value of the original fluorescence signal trace of the first 2.5 minutes of the image frame before drug intervention in the resting state data.
[0086]
[0087] in, The baseline fluorescence intensity represents the mean fluorescence signal at rest before drug intervention. The number of image frames corresponding to the baseline acquisition time can be set to 2.5 minutes before drug intervention; is the raw fluorescence signal of frame t, i.e., the average fluorescence intensity of the corresponding brain region of interest.
[0088] Based on the baseline, the changes in brain neural signal traces before and after drug intervention were calculated using the ΔF / F (Delta F over F) method.
[0089]
[0090] in, The relative rate of change of fluorescence signal in frame t reflects the dynamic changes in neural activity.
[0091] Based on signal processing methods, the number of peaks in the fluorescence signal trace change curves before and after intervention was calculated; pre-intervention signal:
[0092]
[0093] Post-intervention signals:
[0094]
[0095] in, The duration of the signal before intervention, such as recording 5 minutes before drug administration, corresponds to... =900 frames; The total number of recorded frames (before intervention + after intervention, e.g., a total of 10 minutes of recording). =1800 frames).
[0096] Smooth the signal using a Gaussian filter:
[0097]
[0098] in, τ is the time offset (frame interval, τ=−K,…,0,…,K), representing the range of K frames before and after the current frame; K is the half-width of the filter kernel, controlling the smoothing range. The standard deviation of the Gaussian kernel determines the filtering strength;
[0099] Number of signals obtained:
[0100]
[0101] Among them, condition 1: This indicates that the current frame t is a local maximum (peak point), excluding spurious peaks during plateau periods or falling edges; Condition 2: middle This is the amplitude threshold coefficient, used to filter noise: if =1, then the peak value must exceed the signal standard deviation. That is, significantly higher than the noise level; to prevent small fluctuations (such as instrument noise) from being misinterpreted as peaks in neural activity. Condition 3: middle This indicates the threshold for the peak time interval, preventing the oscillation of the same calcium transient from being misjudged as multiple peaks; The standard deviation of the signal reflects the baseline noise level and is used to dynamically adjust the amplitude threshold.
[0102] The difference in the number of peak signals before and after the intervention was calculated to obtain a zebrafish brain neural activation map, which was then used as a reference. Figure 5 , Figure 5 This is a heatmap showing the differences in the number of signal peaks in different brain regions of zebrafish, as presented in this invention.
[0103] In one embodiment, calculating the tail-wagging frequency based on the tail-wagging motion data includes:
[0104] Collect continuous motion images of zebrafish, and use a labeling tool to select and label the outline of the zebrafish in the continuous motion images to obtain labeled data containing the coordinate information of the tail endpoint and the midpoint of the body, forming a training dataset, namely the tail wagging motion data.
[0105] A keypoint detection model is constructed based on a deep learning framework. The keypoint detection model is trained using the training dataset so that it can output the coordinates of the tail tip and the midpoint of the body of a zebrafish.
[0106] For continuously acquired current zebrafish motion images, the trained key point detection model is used to detect and obtain the coordinates of the zebrafish's tail endpoint and body midpoint in each frame, forming coordinate data containing time series information.
[0107] For coordinate data from two consecutive frames, the vector change of the tail relative to the torso is calculated, and the sequence of angle change of the zebrafish tail is obtained based on the vector change.
[0108] Periodic swaying events are detected in the sequence of angle changes. When the absolute value of the angle change between consecutive frames exceeds a preset swaying angle threshold, it is determined to be a valid tail sway. The number of valid tail sways per unit time is counted to obtain the tail swaying frequency, which represents the number of periodic movements of the zebrafish tail swaying per unit time.
[0109] In one specific implementation, continuous images of juvenile zebrafish are captured using behavioral analysis to observe the tail wagging area, wagging angle, and maximum wagging angular velocity, with reference to... Figure 6 , Figure 6 This is a schematic diagram of zebrafish behavior images for this invention. a, b, and c represent different undulation patterns. 1000 images were collected for training and evaluation. The collection and annotation of zebrafish images in bright-field environments are involved. The annotation process involves using a tool to outline the juvenile zebrafish. See [link to data collection and annotation examples] for details. Figure 7 , Figure 7 Images showing the zebrafish bright-field behavioral imaging and annotation results of this invention: a: before annotation, b: after annotation; the Unet model was selected and the TensorFlow deep learning framework was used to build the model. The tail tip and mid-body position of the zebrafish were captured. Based on the tail tip and mid-body position of the zebrafish in consecutive frames, the angular velocity of the tail swing was calculated; the vector of the tail relative to the body was calculated for each frame.
[0110] Frame t:
[0111]
[0112] in, The coordinates of the tail endpoint should be ( , ), The coordinates of the midpoint of the torso are ( , The position vector of the tail relative to the torso should be from the midpoint of the torso to the end point of the tail: Midpoint of the torso Origin, tail endpoint The relative position vector reflects the spatial orientation of the tail in the two-dimensional plane;
[0113] Frame t+1:
[0114]
[0115] Calculate the angle between two vectors (angle change):
[0116]
[0117] in, Let be the polar angle of the vector in frame t. Let be the polar angle of the vector in frame t+1. The change in angle between adjacent frames; from this, the angular velocity is derived. , where Δt is the time interval between frame t and frame t+1. See the diagram for the final result. Figure 8 , Figure 8 This is a timing diagram of the zebrafish tail swaying according to the present invention.
[0118] In one embodiment, recording heart rate based on the heart rate physiological data includes:
[0119] Images of the heart region of zebrafish are collected, and the location of the heart in the heart region images is labeled using a labeling tool to form a labeled dataset, namely the heart rate physiological data.
[0120] A target organ detection model is constructed based on a deep learning framework. The target organ detection model is trained using the labeled dataset so that it can output the coordinates of the center point of the shadow region of the zebrafish heart.
[0121] The current heart region image of zebrafish is continuously acquired. The center point coordinates of the heart shadow region are detected and obtained frame by frame through a trained target organ detection model. The horizontal axis coordinates of the center point coordinates are extracted to form horizontal axis coordinate time series data corresponding to the acquisition time. The acquisition time is equal to the frame number divided by the video acquisition frame rate.
[0122] The time series data of the horizontal axis is subjected to Fourier transform, and the main frequency components are extracted within a preset frequency range. Based on the extracted main frequency components and the video acquisition frame rate, the number of reciprocating motions of the heart per unit time is calculated to obtain the heart rate of the zebrafish.
[0123] It should be noted that 1000 bright-field images of different zebrafish hearts can be collected, and the location of the zebrafish hearts can be marked using annotation tools, for reference. Figure 9 , Figure 9 This is an example image of zebrafish physiology captured by the present invention; the YOLO11 model was selected and the TensorFlow deep learning framework was used to build the model to capture the location of the zebrafish's heart.
[0124] It should be understood that cardiac congestion results in a distinct shadow on the heart. As blood flows between the atria and ventricles, this shadow shifts between them. By capturing the center point (X, Y) of this shadow, the center point will reciprocate along the X-axis during the heart's beating process. Based on the X-axis position of the center point within 30 seconds, a Fourier transform is used to obtain the main frequency of the reciprocating motion. The heart rate can then be calculated based on the time interval between the captured frames. Specifically, let the frame rate be f (fps), the total acquisition time be T = 30 seconds, and the total number of frames be N = f × T.
[0125] The X-coordinate of the center point of the heart shadow in the i-th frame: xi, where i = 0, 1, 2, ..., N−1; the time series: ti = i / f, where i = 0, 1, 2, ..., N−1;
[0126]
[0127] Where f is the sampling frequency and N is the number of sampling points; kmin=[0.5⋅N / f], kma=[5.0⋅N / f].
[0128] It should be understood that, in this embodiment, the zebrafish multidimensional biometric data may also include blood flow velocity data and intestinal peristalsis data. Similarly, the blood flow velocity data and intestinal peristalsis data can be used to form a fusion feature vector of the test sample, and then used for consensus clustering with a set of fusion feature vectors of clinically known drugs, thereby making the efficacy prediction of the test sample more accurate. This embodiment does not limit the type and quantity of zebrafish multidimensional biometric data; data acquisition and clustering analysis can be performed according to actual needs.
[0129] S20: Input the fusion feature vector into a preset efficacy prediction model. The preset efficacy prediction model is used to perform consensus clustering on the fusion feature vector of the test sample and the set of fusion feature vectors of clinical drugs with known efficacy, and output the consensus clustering result.
[0130] S30: Based on the consensus clustering results, associate the test sample with clinical drugs with known efficacy in the cluster to predict the efficacy of the test sample.
[0131] It should be noted that the preset efficacy prediction model is a pre-trained machine learning / deep learning model. Its core function is to predict the potential efficacy of the test product by calculating the feature similarity between the test product and known drugs. Known efficacy clinical drugs are drugs that have been clinically validated and have clearly demonstrated specific therapeutic effects; their feature vectors can constitute the training / reference dataset. In this embodiment, the known efficacy clinical drugs include hypoglycemic drugs, lipid-lowering drugs, uric acid-lowering drugs, antihypertensive drugs, and antiepileptic drugs. Consensus clustering is an ensemble clustering method that integrates multiple clustering results to generate more stable and reliable cluster groups, reducing the randomness or noise impact of a single clustering algorithm. Specifically, the feature vector set is clustered multiple times; the frequency with which any two samples are assigned to the same class in all basic clusters is calculated; secondary clustering is performed based on the consensus matrix to obtain more stable groups. The final cluster grouping output by the model divides the test product and known drugs into several efficacy-related categories. Each sample is assigned a category label; the category label of the test product inherits the main efficacy of the known drugs in that category. Clustering enables efficacy analogy: if a test sample is highly similar in characteristics to a certain class of known drugs, it is inferred that they have similar therapeutic effects, providing candidate directions for drug screening.
[0132] The construction process of the preset efficacy prediction model includes the following steps:
[0133] We acquire multidimensional biological characteristic data of zebrafish before and after intervention with clinical drugs with known efficacy, and integrate the multidimensional biological characteristic data of zebrafish before and after intervention with clinical drugs with known efficacy to form a fusion feature vector of clinical drugs with known efficacy.
[0134] The fusion feature vectors of the clinically known drugs with known efficacy and their corresponding efficacy are used to train a preset deep learning network to obtain the preset efficacy prediction model.
[0135] It should be noted that the preset deep learning network can adopt the MobileNetV3 model architecture, which is convenient for deployment on mobile and embedded devices. Specifically, the preset efficacy prediction model uses consensus clustering for unsupervised machine learning and deep learning methods to build a fully supervised prediction model.
[0136] Taking the functional prediction of forsythoside A in zebrafish as an example, the effects of a group of hypoglycemic drugs, uric acid-lowering drugs, lipid-lowering drugs, antihypertensive drugs, antiepileptic drugs, and forsythoside A on brain activity, behavior, and heartbeat in transgenic zebrafish were first evaluated. The transgenic zebrafish gene encodes a calcium-sensitive fluorescent indicator, referring to... Figure 10 In section A, to quantitatively analyze drug-induced changes in brain activity, behavior, and heart rate, T-score brain activity maps, tail wags, and heart rates of each larva in response to the corresponding drug treatment were generated, referencing... Figure 10In section B, 34 clinical drugs, including hypoglycemic agents, lipid-lowering agents, uric acid-lowering agents, antihypertensive agents, and antiepileptic drugs, were tested (see Table 1). Next, the relationship between T-score BAMs, tail wag, heart rate, and the therapeutic function of the 34 clinical drugs in the library was determined. Using a consistent clustering method, combined with T-score BAMs, tail wag, and heart rate, five phenotypic categories were identified, referring to... Figure 10 D. Further investigation revealed statistically significant associations between multiple categories (e.g., categories 2, 3, and 5) of antiepileptic drugs (N01), lipid-lowering drugs (M02), and hypoglycemic drugs (M01) and the treatment category of the drugs, referring to... Figure 10 The C in the diagram illustrates a heatmap of statistical associations between the anatomical therapeutic chemical category and the identified categories. The heatmap is a consensus matrix illustrating the frequency with which two drugs cluster together in a pair. The colors of the heatmap are proportional to the frequency scores of the consensus matrix, ranging from 0 to 1 (hypergeometric test, P < 0.05). Ultimately, category 3 also includes forsythoside (FP), suggesting that forsythoside may have lipid-lowering effects.
[0137]
[0138] Table 1 List of Drug Names and Their Functions
[0139] Furthermore, after associating the test sample with clinically known drugs in the cluster based on the consensus clustering results to predict the efficacy of the test sample, the method further includes:
[0140] The experiment was divided into a normal group, a model group, a positive group, and a test sample group. The normal group served as the basic reference group for the experiment, providing baseline data of physiological indicators under healthy conditions. The model group was used to verify whether the disease model was successfully constructed and to provide baseline indicators under disease conditions. The positive group was used to verify the reliability of the experimental system and to provide a reference for the effects of known effective drugs. The test sample group was used to verify whether the test sample had the expected efficacy.
[0141] The accuracy of the predictive power of the test sample was verified based on the normal group, model group, positive group, and test sample group.
[0142] It should be noted that, taking the construction of a hyperlipidemia model to verify the lipid-lowering effect of forsythosides as an example, (1) Experimental groups: normal group, model group, positive group, forsythosides group;
[0143] (2) Model construction and test sample intervention
[0144] Juvenile fish selection: Select healthy wild-type AB strain zebrafish that have developed to 5 dpf (days post fertilization) and place them in 6-well cell culture plates, 10 fish / well.
[0145] Model construction: The normal group was given E3 water (5 mM NaCl, 0.17 mM KCl, 0.33 mM CaCl2, and 0.33 mM MgSO4), 5 mL per well, and fed with normal feed; the model group, positive group (atorvastatin calcium), and forsythia suspensa group were given E3 water, 5 mL per well, and fed with high-fat feed; incubated for 7 days, with fresh solution changed daily.
[0146] Intervention with test samples: After the model was constructed, the normal group was given E3 water and fed a normal diet; the model group was given E3 water and fed a high-fat diet; the positive group was given atorvastatin calcium solution (2.5 μg / mL) and fed a high-fat diet; the forsythoside group was given forsythoside (10 μM) and fed a high-fat diet; the intervention lasted for 7 days, with fresh solutions changed daily.
[0147] Oil Red O staining: After the intervention, zebrafish were washed twice with E3 water. Fifteen zebrafish from each group were randomly selected and placed in centrifuge tubes, fixed with 4% paraformaldehyde solution at 4 °C for 24 h. They were then rinsed twice with phosphate-buffered saline (PBS), dehydrated using a gradient of 1,2-propanediol, and then placed in cell culture plates. Oil Red O staining solution was added. After staining, the zebrafish were destained with 1,2-propanediol, observed under a stereomicroscope, and photographed.
[0148] Quantitative statistical analysis of the grayscale values (S) of the oil red O-stained region in zebrafish was performed using ImageJ software. The formula for calculating the relative fat content in zebrafish is as follows:
[0149]
[0150] 6) Data statistics: GraphPad Prism 10 software was used for data processing and graphing. All experimental data are expressed as mean ± SD. One-way ANOVA with Tukey's multiple comparison test was used for comparisons among multiple groups. P < 0.05 was used as the criterion for statistical significance.
[0151] Furthermore, taking the verification of the lipid-lowering effect of forsythoside as an example, based on the above test methods, the effect of forsythoside on the body fat of zebrafish was verified as follows: Figure 11 As shown, A is a visual diagram of fat deposition in zebrafish, and B is a quantitative diagram of fat deposition in zebrafish.
[0152] Depend on Figure 11The results showed that, compared with the normal group, the body fat content of zebrafish in the model group was significantly increased (P<0.001), indicating that the hyperlipidemia model was successfully established in this test. Compared with the model group, the body fat content of zebrafish in the positive group was significantly reduced (P<0.001), indicating that this experiment was effective. Compared with the model group, the body fat content of zebrafish in the forsythoside group was significantly reduced (P<0.001), indicating that forsythoside has a lipid-lowering effect.
[0153] In summary, the model-based experiment to predict the function of forsytholipin in zebrafish only takes 30 minutes, while the traditional method of constructing disease models to evaluate the lipid-lowering effect of forsytholipin takes 14 days. This indicates that the model-based method is more efficient in predicting the function of test samples in zebrafish.
[0154] Furthermore, this embodiment of the invention also proposes a storage medium storing a test sample efficacy prediction program based on zebrafish multi-feature fusion. When the test sample efficacy prediction program based on zebrafish multi-feature fusion is executed by a processor, it implements the steps of the test sample efficacy prediction method based on zebrafish multi-feature fusion as described above.
[0155] In addition, refer to Figure 5 This invention also proposes a test sample efficacy prediction device based on zebrafish multi-feature fusion, the test sample efficacy prediction device based on zebrafish multi-feature fusion includes:
[0156] The data integration module 10 is used to collect multi-dimensional biological characteristic data of zebrafish before and after the intervention of the test sample, and integrate the multi-dimensional biological characteristic data of zebrafish to form a fusion feature vector of the test sample.
[0157] The model clustering module 20 is used to input the fusion feature vector into a preset efficacy prediction model. The preset efficacy prediction model is used to perform consensus clustering on the fusion feature vector of the test sample and the fusion feature vector set of clinical drugs with known efficacy, and output the consensus clustering result.
[0158] The efficacy prediction module 30 is used to associate the test sample with clinical drugs with known efficacy in the cluster based on the consensus clustering results, and predict the efficacy of the test sample.
[0159] Other embodiments or specific implementations of the zebrafish multi-feature fusion-based test sample efficacy prediction device of the present invention can be referred to the above-described method embodiments, and will not be repeated here.
[0160] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0161] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments. In the unit claims listing several devices, several of these devices may be embodied by the same hardware item. The use of the terms first, second, and third, etc., does not indicate any order and can be interpreted as names.
[0162] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as a read-only memory image (ROM) / random access memory (RAM), magnetic disk, optical disk), and includes several instructions to cause a terminal user device (which may be a mobile phone, computer, server, air conditioner, or network user device, etc.) to execute the methods described in the various embodiments of the present invention.
[0163] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.
Claims
1. A method for predicting the efficacy of test samples based on zebrafish multi-feature fusion, characterized in that, The method includes: Multidimensional biological characteristic data of zebrafish before and after intervention were collected, and the multidimensional biological characteristic data of zebrafish were integrated to form a fusion feature vector of the test sample. The fusion feature vector is input into a preset efficacy prediction model, which is used to perform consensus clustering on the fusion feature vector of the test sample and the set of fusion feature vectors of clinical drugs with known efficacy, and outputs the consensus clustering result. Based on the consensus clustering results, the test sample is associated with clinical drugs with known efficacy in the cluster to predict the efficacy of the test sample; wherein, The zebrafish multidimensional biometric data includes brain nerve fluorescence image data, tail wagging motion data, and heart rate physiological data. Accordingly, the process of integrating the multi-dimensional biological characteristic data of the zebrafish to form a fused feature vector of the test sample includes: The T-score brain activity map is calculated based on the brain nerve fluorescence image data, the tail wag frequency is calculated based on the tail wag motion data, and the heart rate is recorded based on the heart rate physiological data. The T-score brain activity map, the tail wag frequency, and the heart rate are integrated to form a fusion feature vector of the test sample; The T-score brain activity map calculated based on the brain nerve fluorescence image data includes: The brain fluorescence image data of zebrafish were aligned with the spatial coordinates of a pre-defined standardized brain template to establish a unified image coordinate system. A preset region of interest is defined on the standardized brain template, and the preset region of interest is used to define the analysis area of subsequent neural activity signals; For each of the preset regions of interest, the calcium transient counts before and after the intervention of the test sample are obtained, and the numerical difference between the two is calculated to generate the calcium transient difference signal of zebrafish. For the calcium transient differential signal of zebrafish, the signal is accumulated along the longitudinal axis in three-dimensional space, and the accumulated calcium transient differential signal is mapped to a two-dimensional plane to generate a brain activity map of zebrafish. Based on brain activity map data of the same region of interest for each group of zebrafish, the t-test method was used to calculate the t-score value at the group level and generate a group-level T-score brain activity map. The calculation of the tail-wagging frequency based on the tail-wagging motion data includes: Collect continuous motion images of zebrafish, and use a labeling tool to select and label the outline of the zebrafish in the continuous motion images to obtain labeled data containing the coordinate information of the tail tip and the midpoint of the body, forming a training dataset, namely the tail wagging motion data. A keypoint detection model is constructed based on a deep learning framework. The keypoint detection model is trained using the training dataset so that it can output the coordinates of the tail tip and the midpoint of the body of a zebrafish. For continuously acquired current zebrafish motion images, the trained key point detection model is used to detect and obtain the coordinates of the tail tip and the midpoint of the body of the zebrafish in each frame, forming coordinate data containing time series information. For coordinate data from two consecutive frames, the vector change of the tail relative to the torso is calculated, and the sequence of angle change of the zebrafish tail is obtained based on the vector change. Periodic swaying events are detected in the sequence of angle changes. When the absolute value of the angle change between consecutive frames exceeds a preset swaying angle threshold, it is determined to be a valid tail sway. The number of valid tail sways per unit time is counted to obtain the tail swaying frequency, which represents the number of periodic movements of the zebrafish tail swaying per unit time.
2. The method for predicting the efficacy of test samples based on zebrafish multi-feature fusion as described in claim 1, characterized in that, The process of constructing the preset efficacy prediction model includes the following steps: We acquire multidimensional biological characteristic data of zebrafish before and after intervention with clinical drugs with known efficacy, and integrate the multidimensional biological characteristic data of zebrafish before and after intervention with clinical drugs with known efficacy to form a fusion feature vector of clinical drugs with known efficacy. The fusion feature vectors of the clinically known drugs with known efficacy and their corresponding efficacy are used to train a preset deep learning network to obtain the preset efficacy prediction model.
3. The method for predicting the efficacy of test samples based on zebrafish multi-feature fusion as described in claim 1, characterized in that, The recording of heart rate based on the aforementioned physiological heart rate data includes: Images of the heart region of zebrafish are collected, and the location of the heart in the heart region images is labeled using a labeling tool to form a labeled dataset, namely the heart rate physiological data. A target organ detection model is constructed based on a deep learning framework. The target organ detection model is trained using the labeled dataset so that it can output the coordinates of the center point of the shadow region of the zebrafish heart. The current heart region image of zebrafish is continuously acquired. The center point coordinates of the heart shadow region are detected and obtained frame by frame through a trained target organ detection model. The horizontal axis coordinates of the center point coordinates are extracted to form horizontal axis coordinate time series data corresponding to the acquisition time. The acquisition time is equal to the frame number divided by the video acquisition frame rate. The time series data of the horizontal axis is subjected to Fourier transform, and the main frequency components are extracted within a preset frequency range. Based on the extracted main frequency components and the video acquisition frame rate, the number of reciprocating motions of the heart per unit time is calculated to obtain the heart rate of the zebrafish.
4. The method for predicting the efficacy of test samples based on zebrafish multi-feature fusion as described in claim 1, characterized in that, After associating the test sample with clinically known drugs in the cluster based on the consensus clustering results to predict the efficacy of the test sample, the method further includes: The experiment was divided into a normal group, a model group, a positive group, and a test sample group. The normal group served as the basic reference group for the experiment, providing baseline data of physiological indicators under healthy conditions. The model group was used to verify whether the disease model was successfully constructed and to provide baseline indicators under disease conditions. The positive group was used to verify the reliability of the experimental system and to provide a reference for the effects of known effective drugs. The test sample group was used to verify whether the test sample had the expected efficacy. The accuracy of the predictive power of the test sample was verified based on the normal group, model group, positive group, and test sample group.
5. A device for predicting the efficacy of test samples based on zebrafish multi-feature fusion, characterized in that, The zebrafish multi-feature fusion-based test sample efficacy prediction device includes: The data integration module is used to collect multi-dimensional biological characteristic data of zebrafish before and after the intervention of the test sample, and integrate the multi-dimensional biological characteristic data of zebrafish to form a fusion feature vector of the test sample. The model clustering module is used to input the fusion feature vector into a preset efficacy prediction model. The preset efficacy prediction model is used to perform consensus clustering on the fusion feature vector of the test sample and the set of fusion feature vectors of clinical drugs with known efficacy, and output the consensus clustering result. The efficacy prediction module is used to associate the test sample with clinical drugs with known efficacy in the cluster based on the consensus clustering results, and predict the efficacy of the test sample. in, The zebrafish multidimensional biometric data includes brain nerve fluorescence image data, tail wagging motion data, and heart rate physiological data. Accordingly, the process of integrating the multi-dimensional biological characteristic data of the zebrafish to form a fused feature vector of the test sample includes: The T-score brain activity map is calculated based on the brain nerve fluorescence image data, the tail wag frequency is calculated based on the tail wag motion data, and the heart rate is recorded based on the heart rate physiological data. The T-score brain activity map, the tail wag frequency, and the heart rate are integrated to form a fusion feature vector of the test sample; The T-score brain activity map calculated based on the brain nerve fluorescence image data includes: The brain fluorescence image data of zebrafish were aligned with the spatial coordinates of a pre-defined standardized brain template to establish a unified image coordinate system. A preset region of interest is defined on the standardized brain template, and the preset region of interest is used to define the analysis area of subsequent neural activity signals; For each of the preset regions of interest, the calcium transient counts before and after the intervention of the test sample are obtained, and the numerical difference between the two is calculated to generate the calcium transient difference signal of zebrafish. For the calcium transient differential signal of zebrafish, the signal is accumulated along the longitudinal axis in three-dimensional space, and the accumulated calcium transient differential signal is mapped to a two-dimensional plane to generate a brain activity map of zebrafish. Based on brain activity map data of the same region of interest for each group of zebrafish, the t-test method was used to calculate the t-score value at the group level and generate a group-level T-score brain activity map. The calculation of the tail-wagging frequency based on the tail-wagging motion data includes: Collect continuous motion images of zebrafish, and use a labeling tool to select and label the outline of the zebrafish in the continuous motion images to obtain labeled data containing the coordinate information of the tail tip and the midpoint of the body, forming a training dataset, namely the tail wagging motion data. A keypoint detection model is constructed based on a deep learning framework. The keypoint detection model is trained using the training dataset so that it can output the coordinates of the tail tip and the midpoint of the body of a zebrafish. For continuously acquired current zebrafish motion images, the trained key point detection model is used to detect and obtain the coordinates of the tail tip and the midpoint of the body of the zebrafish in each frame, forming coordinate data containing time series information. For coordinate data from two consecutive frames, the vector change of the tail relative to the torso is calculated, and the sequence of angle change of the zebrafish tail is obtained based on the vector change. Periodic swaying events are detected in the sequence of angle changes. When the absolute value of the angle change between consecutive frames exceeds a preset swaying angle threshold, it is determined to be a valid tail sway. The number of valid tail sways per unit time is counted to obtain the tail swaying frequency, which represents the number of periodic movements of the zebrafish tail swaying per unit time.
6. A device for predicting the efficacy of test samples based on zebrafish multi-feature fusion, characterized in that, The device includes: a memory, a processor, and a zebrafish multi-feature fusion-based test sample efficacy prediction program stored in the memory and executable on the processor, the zebrafish multi-feature fusion-based test sample efficacy prediction program being configured to implement the steps of the zebrafish multi-feature fusion-based test sample efficacy prediction method as described in any one of claims 1 to 4.
7. A storage medium, characterized in that, The storage medium stores a test sample efficacy prediction program based on zebrafish multi-feature fusion, which, when executed by a processor, implements the steps of the test sample efficacy prediction method based on zebrafish multi-feature fusion as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Method for evaluating efficacy of hair-blacking cosmetics based on zebra fish model
CN118022006A