Drug prediction model training method based on tumor organoid biological sample library data
By establishing homology relationships and dynamic pharmacodynamic labels for organoid drug treatment objects, the problems of culture state differences and drug response confounding in drug prediction models were solved, thereby improving the stability and accuracy of drug prediction models.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU NOGHUI BIOTECHNOLOGY CO LTD
- Filing Date
- 2026-06-11
- Publication Date
- 2026-07-31
AI Technical Summary
In existing technologies, drug prediction models based on tumor organoid biobank data have difficulty effectively distinguishing between drug-induced responses and differences in culture states during the training process. This leads to a decrease in the predictive stability of the model on new batches of organoids or new patient samples, affecting the ranking of candidate drugs and the efficiency of experimental validation.
By establishing the homology of organoid drug treatment objects, generating time-series association records, performing baseline observation standardization characterization, analyzing culture perturbations and retaining true baseline differences, constructing dynamic pharmacodynamic labels and label credibility weights, training neural networks to identify drug-induced response fragments and performing consistency verification, and generating dynamic pharmacodynamic labels and label credibility weights.
It achieves the unification of training sample object boundaries and temporal boundaries, solves the problem of data misuse in organoid biobanks, dynamically filters real drug response fragments, and improves the stability and prediction accuracy of drug prediction models.
Smart Images

Figure CN122490111A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of neural network technology, and in particular to a method for training drug prediction models based on tumor organoid biobank data. Background Technology
[0002] Tumor organoids can retain some structural, proliferative, and drug response characteristics of patient tumor tissue in vitro, and have been increasingly used for anti-tumor drug screening, personalized medicine research, and drug prediction model training. With the construction of organoid biobanks, the biobanks typically accumulate various types of data, including patient origin information, cancer subtype information, lesion origin information, culture process records, passage status records, drug administration records, microscopic imaging data, and endpoint drug efficacy readings. In practical applications, the response results of the same drug on different organoid samples may be affected by the drug's mechanism of action, as well as by factors such as the initial volume of the organoid, proliferation rate, passage number, seeding density, matrix batch, culture time, drug plate position, imaging equipment parameters, and detection time window. For example, the proliferation status of organoids from the same patient may differ in the early post-resuscitation period and after stable passage. The edge well positions of the same drug plate may also cause reading deviations due to evaporation or uneven temperature distribution. These factors will all be incorporated into the subsequent model training data.
[0003] In existing technologies, drug prediction models based on tumor organoid biobank data typically use endpoint survival rate, inhibition rate, IC50, or AUC as pharmacodynamic labels, and combine imaging features, omics features, clinical features, or drug structure features to train deep learning models. To reduce data variability, common methods include data normalization, batch correction, sample screening, data augmentation, and temporal feature extraction. However, when organoid samples come from complex sources, culture process records are not entirely consistent, and detection time windows differ, the above methods mainly focus on unifying data scales or mitigating batch variability. They still lack specific differentiation of the contribution relationship between culture disturbances, baseline growth differences of the sample itself, and drug-induced responses. The endpoint pharmacodynamic label may retain non-pharmacological response features, since the endpoint pharmacodynamic label mainly reflects... When the combined results of preset detection time points lack the correspondence between the baseline trajectory before drug administration and the continuous response trajectory after drug administration, it is difficult to fully distinguish between different processes such as drug-induced inhibition, natural slow growth of organoids, recovery after early response, and delayed drug effects. During the training process, the neural network may fit differences in culture state or baseline growth as drug sensitivity-related features. This may result in the model maintaining good performance in the internal validation of the sample library, but showing decreased predictive stability on new batches of organoids, new center samples, or new patient samples, thereby affecting the ranking of candidate drugs and the efficiency of subsequent experimental validation. Therefore, how to identify non-drug interferences in the pharmacodynamic label before training the drug prediction model and avoid the neural network mislearning culture perturbations and baseline growth differences is a technical problem that still needs to be solved in this field. Summary of the Invention
[0004] This application proposes a method for training drug prediction models based on tumor organoid biobank data to address the problems mentioned in the background art.
[0005] To achieve the above objectives, this application adopts the following technical solution: a drug prediction model training method based on tumor organoid biobank data, comprising:
[0006] S1. Obtain sample library association data to characterize the homology boundary and efficacy observation process of organoid drug treatment objects, perform time-series association construction processing according to sample treatment boundary information and efficacy observation time sequence, and generate time-series association records;
[0007] S2, based on time-series correlation records, performs baseline observation standardization characterization processing on the baseline observation content, and generates the culture perturbation interpretation results and baseline difference retention results according to the dual-boundary analysis processing of culture perturbation interpretation and baseline difference retention.
[0008] S3. Based on the results of culture perturbation interpretation and baseline difference retention, perturbation calibration is performed on the baseline state before drug administration to generate the baseline calibration growth trajectory and the effective range of baseline growth. The response state after drug administration is compared with the baseline calibration growth trajectory and the effective range of baseline growth to screen drug-induced response fragments. Dynamic efficacy labels and label confidence weights are generated according to the consistency verification between the response fragments and the endpoint efficacy.
[0009] S4 inputs the drug prediction input features formed by the information available before drug exposure in the time-series association records into the drug prediction neural network, uses dynamic efficacy labels as the supervision target, and performs model training according to the credible supervision training process limited by the label credibility weights, and outputs a drug prediction model for predicting the sensitivity results or drug response level of candidate drugs.
[0010] Furthermore, the sample library associated data includes biological source data, culture status data, drug treatment data, time-series observation data, and reference verification data. Among them, biological source data is used to define the source of organoid drug treatment objects, culture status data is used to define the culture process and passage status of organoid drug treatment objects, drug treatment data is used to define the conditions for the action of candidate drugs, time-series observation data is used to define the observation process before and after drug exposure, and reference verification data is used to define the basis for untreated reference and endpoint efficacy verification.
[0011] The operation of obtaining sample library association data for characterizing the homology boundary and efficacy observation process of organoid drug treatment objects includes: assigning data roles to the sample library association data according to the data roles corresponding to biological source data, culture status data, drug treatment data, time-series observation data and reference verification data in the sample library recording specifications, drug screening detection procedures and historical verification samples.
[0012] Furthermore, the operation of constructing time-series associations based on sample processing boundary information and drug efficacy observation time sequence includes: establishing homology association keys based on biological source boundaries, culture state boundaries, and drug processing boundaries; and grouping the boundary definition information, drug processing information, and drug efficacy observation information of the same organoid drug treatment object into the same time-series association record according to the homology association keys; screening untreated reference units with homology boundaries to organoid drug treatment objects from the sample library association data with untreated reference roles; when no untreated reference unit that meets the homology reference conditions formed by the sample library verification procedure is found, selecting reference backoff sources in the adjacent passages of the same sample, the same culture batch of the same cancer subtype, and the historical baseline records of the same cancer subtype according to the reference backoff order formed by the sample library verification procedure, and generating reference backoff markers; and performing relative time axis alignment on untreated references, pre-drug observations, post-drug responses, and endpoint drug efficacy records with the drug exposure start time as the time zero point, generating time-series association records containing observation validity markers and reference backoff markers.
[0013] Furthermore, the baseline observation standardization and characterization process includes: extracting the number of organoids, organoid area, and organoid morphological compactness from the baseline observations in the untreated reference and the baseline observations before drug administration, respectively; converting the number of organoids, organoid area, and organoid morphological compactness to a uniform scale according to the distribution boundaries of qualified observation samples determined by the sample library review procedure within the same imaging batch, and performing homogenization processing according to the effective growth state of the organoids; wherein, the organoid morphological compactness results establish a morphological support mapping through historical review samples, mapping the morphological intervals corresponding to the effective growth state in the baseline observations to the first support result, and mapping the morphological intervals corresponding to the fragmented state, incomplete boundary state, abnormal collapse state, and image occlusion state in the baseline observations to the second support result. The first support result is used to retain the corresponding morphological items to participate in the formation of baseline observation characterization, and the second support result is used to mark the corresponding morphological items as morphological anomalies to generate baseline observation characterization for the interpretation of culture perturbation and the preservation of baseline differences.
[0014] Furthermore, the process of generating culture perturbation interpretation results and baseline difference retention results according to the dual-boundary analysis of culture perturbation interpretation and baseline difference retention includes: comparing the baseline observation characteristics corresponding to the untreated reference with the baseline observation characteristics corresponding to the pre-drug-treated baseline observation content at relative time points, and combining the stability of the reference well positions in the drug plate reference record and the common fluctuation trend of untreated observations within the same culture batch to generate culture common offset results; based on the culture common offset results and the offset conversion rules formed by drug screening detection procedures, culture calibration records, or historical verification samples, determining the passage status, matrix batch, inoculation density, etc. The offsets corresponding to the imaging batch and detection time window are calculated, and culture perturbation interpretation results are generated. Historical baseline stability results are generated by retrieving historical baseline records from the same patient source, the same cancer subtype, or the same lesion source, and cross-batch response stability results are generated by retrieving historical response records that have completed label verification. Based on historical baseline stability results, cross-batch response stability results, sample source credibility, and culture perturbation interpretation results, baseline difference retention results are generated. When the number of historical samples does not reach the minimum number specified in the sample library verification procedure, a historical sample insufficiency mark is generated, but the historical sample insufficiency mark is not used as the sole basis for deducting from the baseline difference retention results.
[0015] Furthermore, the perturbation calibration of the pre-drug baseline status includes: determining the set of valid pre-drug time points from the relative time points that are valid before drug exposure; forming a representative pre-drug baseline difference based on the difference between the baseline observation characteristics of the pre-drug baseline in the set of valid pre-drug time points and the baseline observation characteristics of the untreated reference; when the number of valid pre-drug time points reaches the minimum number specified in the sample library review procedure, the representative pre-drug baseline difference is determined using the observation quality-weighted median; when the number of valid pre-drug time points does not reach the minimum number specified in the sample library review procedure, the difference of the most recent valid baseline time point before drug exposure is used to replace the representative pre-drug baseline difference, and a pre-drug baseline point insufficiency marker is generated; and combining the natural growth trajectory provided by the untreated reference, the strength of the true biological difference retention limited by the baseline difference retention result, and the strength of the non-drug-related offset correction limited by the culture perturbation interpretation result to generate a baseline-calibrated growth trajectory.
[0016] Furthermore, the operations for generating the baseline calibration growth trajectory and the effective baseline growth range include: using the baseline calibration growth trajectory as the center of the effective baseline growth range; determining the range width of the effective baseline growth range based on the baseline fluctuation width, culture common offset results, culture perturbation interpretation results, historical baseline stability results, and the set of analytical markers; the baseline fluctuation width is formed from historical qualified untreated baseline observation records; when there are reference rollback markers, drug plate reference anomaly markers, or missing observation time markers, the effective baseline growth range at the corresponding time point is corrected according to the range width correction rules formed by the sample library review procedure, or the subsequent label credibility weights are corrected according to the label credibility weight formation rules; among them, the culture perturbation interpretation results are used to determine the explainable range of non-drug-related changes, and the baseline difference retention results are used to limit the retention range of true biological differences; the two together limit whether the response state after drug administration exceeds the effective baseline growth range.
[0017] Further, the screening of drug-induced response fragments includes: comparing the post-drug response state with the baseline calibrated growth trajectory and the effective baseline growth range; identifying response time points that occur after the start of drug exposure, have valid response observations, exceed the effective baseline growth range, and retain residual responses after processing by culture perturbation interpretation results as drug-induced response time points; when there is a time point in the post-drug response state where the number of organoids is lower than the baseline object number lower limit formed by the sample library verification procedure, if this time point meets the image quality and imaging field of view integrity conditions formed by the drug screening detection procedure, and occurs after the start of drug exposure, then this time point is marked as a low object number response candidate time point and included in the screening of drug-induced response fragments; when there is still a continuous residual response corresponding to the drug exposure time after the post-drug response change is processed by culture perturbation interpretation results, a calibrated residual response marker is generated, and the corresponding response time point is retained as a candidate for drug-induced response fragments; continuous drug-induced response time points with consistent response directions are merged into drug-induced response fragments, and response initiation time, response duration, response intensity, recovery trend, and late-onset markers are generated for the drug-induced response fragments.
[0018] Furthermore, the process of generating dynamic efficacy labels and label credibility weights based on the consistency verification between response fragments and endpoint efficacy includes: generating dynamic response intensity results based on the response initiation time, response duration, response intensity, recovery trend, late-onset markers, and fragment observation quality of the drug-induced response fragment; converting the endpoint efficacy record into a endpoint efficacy normalization result in the same direction as the drug response degree, and comparing the consistency between the dynamic response intensity result and the endpoint efficacy normalization result to generate an endpoint consistency verification result; when the endpoint efficacy record is missing or does not meet the quality review conditions formed by the drug screening detection procedure of the sample library, the weight of the endpoint consistency verification item is reset to zero, and a dynamic efficacy label is formed from the dynamic response intensity result; generating label credibility weights based on the culture perturbation interpretation results, baseline difference retention results, cross-batch response stability results, endpoint consistency verification results, drug-induced response fragment quality, calibrated residual response markers, and parsing marker set, and when the same source of anomaly has already been deducted during the drug-induced response fragment screening process, it will not be deducted again during the label credibility weight formation process of the same response fragment.
[0019] Furthermore, the model training process, performed according to the trusted supervision training process defined by the label's trusted weight, includes: forming drug prediction input features from information available before drug exposure in time-series correlation records; inputting the drug prediction input features into the drug prediction neural network, with dynamic pharmacodynamic labels as the primary supervision target; adjusting the contribution of training samples to model parameter updates based on the label's trusted weight; performing response fragment auxiliary supervision based on drug-induced response fragments; and performing culture perturbation sensitivity constraints based on the interpretation results of culture perturbation and the results of baseline difference retention. It also involves determining cross-batch constraint sample pairs according to the response fragment consistency pairing rules formed by drug action category, response fragment direction, label trusted weight, culture perturbation difference, and baseline difference retention differences, and performing response fragment consistency constraints on these pairs. Training samples not included in the cross-batch constraint sample pairs retain their single-sample supervised training contribution defined by the dynamic pharmacodynamic label and label trusted weight. The culture perturbation sensitivity constraint rules and response fragment consistency pairing rules are formed by review procedures using historical validation sets, verification samples, or sample libraries.
[0020] The beneficial effects of this invention are as follows:
[0021] 1. This invention solves the problem of data from different sources, different culture batches, and different detection time windows in organoid biobanks being mistakenly used as the same training sample by establishing a homology relationship between organoid drug treatment objects before model training and limiting natural growth reference, pre-drug baseline, post-drug response, and endpoint efficacy to the same relative time boundary. It achieves the unification of training sample object boundaries and temporal boundaries, and also achieves data consistency between culture perturbation interpretation results, baseline difference retention results, baseline-calibrated growth trajectory, and dynamic efficacy labels, so that subsequent training supervision signals no longer rely solely on a single endpoint efficacy record.
[0022] 2. This invention first analyzes the non-pharmacological changes that can be explained by culture disturbances, then retains the true baseline differences that have historical stability and cross-batch response support, and constructs a baseline-calibrated growth trajectory and effective baseline growth range accordingly. This solves the problem of natural slow growth, culture state fluctuations and true drug-induced responses being mixed in the endpoint pharmacodynamic label, realizes the dynamic screening of true drug response fragments, and also realizes the retention and judgment of low-number responses, delayed responses and recovery responses.
[0023] 3. This invention divides the work of dynamic pharmacodynamic labels and label confidence weights for neural network training. The dynamic pharmacodynamic labels characterize the intensity of drug-induced response, and the label confidence weights control the training contribution of low-confidence samples. Combined with the culture perturbation sensitivity constraint and the cross-batch response fragment consistency constraint, this invention solves the problem that drug prediction neural networks mislearn culture state differences or low-confidence endpoint records as drug sensitivity features. It achieves controlled utilization of highly perturbation training samples and stable output of candidate drug sensitivity prediction results across culture batches. Attached Figure Description
[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort:
[0025] Figure 1 This is a flowchart of the method of the present invention;
[0026] Figure 2 This is the S3 flowchart of the present invention. Detailed Implementation
[0027] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0028] like Figure 1 and Figure 2 As shown, this invention discloses a method for training a drug prediction model based on tumor organoid biobank data, including the following specific steps:
[0029] In this embodiment, S1 acquires sample library association data used to characterize the homology boundary of organoid drug treatment objects and the drug efficacy observation process. The temporal association construction process is performed according to the sample treatment boundary information and the drug efficacy observation time sequence to generate temporal association records. The temporal association records are used to provide a unified data foundation for subsequent culture perturbation interpretation results, baseline difference retention results, baseline calibration growth trajectory, dynamic drug efficacy labels and label confidence weights.
[0030] The sample library associated data includes biological source data, culture status data, drug treatment data, time-series observation data, and reference verification data. Among them, biological source data is used to define the origin of organoid drug treatment objects; culture status data is used to define the culture process and passage status of organoid drug treatment objects; drug treatment data is used to define the conditions for the action of candidate drugs; time-series observation data is used to define the observation process before and after drug exposure; and reference verification data is used to define the untreated reference and the basis for endpoint efficacy verification. As a preferred implementation, the reference verification data includes drug plate reference records and endpoint efficacy records. The endpoint efficacy records include survival rate, inhibition rate, IC50 or AUC results. In S1, the endpoint efficacy records are only used as the data source for subsequent endpoint efficacy consistency verification and are not directly used as supervision labels for the drug prediction neural network.
[0031] After acquiring the associated data of the sample library, data role assignment is performed on the associated data of the sample library according to the boundary rules formed by the sample library recording specifications, drug screening detection procedures and historical verification samples. The data is respectively classified into biological source boundary, culture state boundary, drug treatment boundary, time series observation boundary and reference verification boundary. After the data role assignment is completed, when S2 generates the culture disturbance interpretation results, it extracts the basis from the culture state boundary and reference verification boundary. When S3 generates the dynamic pharmacodynamic label, it extracts the basis from the time series observation boundary and reference verification boundary. This avoids the mixed use of source attribution, culture process, drug treatment, response observation and endpoint verification data without defining roles.
[0032] For each organoid drug treatment object, a homology association key is established based on the biological source boundary, culture state boundary, and drug treatment boundary. The homology association key is determined by at least the organoid sample identifier, culture batch identifier, passage state identifier, and drug treatment identifier. Among them, the drug treatment identifier is used to define the drug identifier, drug concentration, and dosing start time. According to the homology association key, the boundary definition information, drug treatment information, and efficacy observation information of the same organoid drug treatment object are grouped into the same temporal association record. The output of this process is the temporal association record corresponding to a single organoid drug treatment object, which is used to avoid data from different patient sources, different culture batches, different passage states, or different drug treatment conditions being merged into the same training sample.
[0033] Untreated reference units are selected from the sample library association data with untreated reference roles. Untreated reference units are used to characterize natural growth references that have homologous boundaries with organoid drug treatment objects. During the selection, priority is given to untreated observation units that have the same organoid sample identifier, the same culture batch identifier, and the same passage status identifier as organoid drug treatment objects, and that have not been configured with drug exposure events. Untreated reference units do not point to the same physical organoid object as drug treatment objects. Their role is to provide natural growth references under the same or traceable culture conditions.
[0034] When multiple candidate undosed reference units exist, the baseline reference matching degree is determined based on batch consistency, proximity of initial observation time, distance between drug plate wells, consistency of imaging equipment parameters, and image quality. Specifically, batch consistency, proximity of initial observation time, distance between drug plate wells, consistency of imaging equipment parameters, and image quality are converted into matching sub-results in the range of 0 to 1, and then the baseline reference matching degree is obtained by combining the matching weights formed by historical verification samples. The baseline reference matching degree is a dimensionless result in the range of 0 to 1, which is used to characterize the reference fit between the candidate undosed reference units and the organoid drug treatment objects.
[0035] In a preferred embodiment, the baseline reference matching threshold is set to 0.70 to 0.85. This threshold is dimensionless and is determined by the deviation distribution between homologous untreated references and the pre-drug baseline in the historical verification samples. It is used to eliminate candidate untreated reference units with insufficient reference fit, so that the untreated reference units entering the baseline calibration have a traceable homologous basis. When the number of historical verification samples does not reach the minimum number specified in the sample library verification procedure, the minimum number is preferably 3 to 5 historical reference records of the same source or subtype. When the minimum number is not reached, an initial matching threshold is set according to the sample library verification procedure, and a baseline reference sample shortage mark is generated at the same time. The baseline reference sample shortage mark is used for subsequent baseline growth effective interval correction or label confidence weight correction, and is not used as a basis for judging the absence of drug response.
[0036] If no untreated reference unit that meets the homologous reference conditions formed by the sample library verification procedure is found, the corresponding organoid drug treatment object is not deleted. The reference fallback source is selected according to the reference fallback order formed by the sample library verification procedure. The preferred reference fallback order is: first, select the untreated reference source of adjacent passages of the same sample; then, select the untreated reference source of the same cancer subtype and the same culture batch; finally, select the historical baseline record of the same cancer subtype as the reference source. Each time a reference fallback occurs, a reference fallback mark is generated and written to the time-series association record. The reference fallback mark is passed to S2 and S3 for subsequent baseline growth effective interval correction or label confidence weight correction. It is not used as a basis for judging the absence of drug response. This process is used to retain organoid drug treatment objects with limited sample size, limited number of cultured survivors, or rare cancer types.
[0037] After screening the undosed reference units, the dosing start time in the drug exposure record is used as the time zero point. Relative time axis alignment is performed on the undosed reference, pre-dosing observation, post-dosing response, and endpoint efficacy records. For pre-dosing observation and post-dosing response, the actual dosing start time of the corresponding organoid drug treatment object is used as the time zero point. For the undosed reference, since no drug exposure event occurred, the dosing start time of the matching organoid drug treatment object is used as the virtual time zero point, so that the undosed reference can be compared with the pre-dosing observation and post-dosing response on the same relative time axis. After relative time axis alignment, each observation point is assigned to a preset detection time point. When the time difference between the observation point and the preset detection time point exceeds the time tolerance range, a missing marker is generated at the corresponding time point, and no arbitrary compensation is performed.
[0038] The time tolerance range is determined by the imaging acquisition error records of the same drug screening experiment, and is preferably 4% to 8% of the preset imaging interval. When the preset imaging interval is 24 hours, the time tolerance range is preferably 1 hour to 2 hours. This time tolerance range is a time quantity used to unify the observation time boundary between the untreated reference, the pre-drug observation, and the post-drug response, and to avoid mistaking the acquisition delay as a change in the response start time. Observation points exceeding the time tolerance range do not participate in the formation of the response fragment at the corresponding detection time point, and are transmitted to S2 and S3 through missing markers. When the historical imaging acquisition error records are insufficient, the time tolerance range is determined by the imaging time deviation boundary specified in the drug screening detection procedure.
[0039] Under a unified relative time axis, observation units are formed for the untreated reference, pre-drug observation, and post-drug response. Each observation unit includes observation time, number of organoids, organoid area result, organoid morphological compactness result, image quality result, and well location information. The number of organoids, organoid area result, organoid morphological compactness result, and image quality result can be generated using existing image analysis methods. Their output is only used as a data source for subsequent baseline observation characterization, observation validity markers, and culture perturbation interpretation. Well location information is used to reflect the position of the corresponding observation unit within the drug plate and is only used as a data source for drug plate reference and culture perturbation interpretation, and is not directly used as a basis for judging efficacy.
[0040] For untreated references and pre-drug observations, baseline observation valid markers are generated based on image quality results and the number of organoids. When the image quality results meet the quality conditions formed by the sample library review procedure, and the number of organoids reaches the lower limit of the baseline object number, the corresponding observation unit is marked as a valid baseline observation. The image quality conditions are formed by the distribution of sharpness, exposure integrity, and occlusion ratio in historical review samples. The lower limit of the baseline object number is determined by the drug screening test procedure or historical review samples. As a preferred implementation, when counting by single well, the lower limit of the baseline object number is 5 to 20 organoids; when counting by fixed field of view, the lower limit of the baseline object number is 3 to 10 organoids. The lower limit of the baseline object number is a count to ensure that the baseline status is statistically representative. When the historical review samples are insufficient, the lower limit of the baseline object number is determined by the minimum evaluable sample size specified by the drug screening test procedure. The lower limit of the baseline object number is only used for the formation of baseline observation valid markers for untreated references and pre-drug observations, and is not used to directly exclude low object number responses after drug administration.
[0041] For post-drug administration responses, valid markers for response observations are generated based on image quality results, imaging field of view integrity, and segmentation traceability. If the number of organoids in the post-drug administration response is lower than the baseline lower limit for the number of organs, and this decrease occurs after the onset of drug exposure, it is not directly marked as an invalid observation point. Instead, a low-number response candidate marker is generated and passed to S3 to participate in the identification of drug-induced response fragments. The generation of low-number response candidate markers also needs to meet the following conditions: the image quality results of the corresponding observation point meet the quality conditions formed by the sample library review procedure; the imaging field of view integrity meets the field of view integrity conditions formed by the drug screening detection procedure; the segmentation process has traceable records; and the decrease in the number of organoids is not caused by missing field of view, focal plane drift, or segmentation failure. This processing is used to avoid the misjudgment of actual drug killing, organoid disintegration, or drug-induced disappearance as invalid observations.
[0042] Data that has completed data role assignment, homology association, untreated reference screening, reference rollback, relative time axis alignment, and observation validity marking are written into the time-series association record. The time-series association record includes homology association key, boundary constraint information, drug processing information, untreated reference, pre-dose observation, post-dose response, endpoint efficacy record, observation validity marking, reference rollback marking, low-object-count response candidate marking, and missing marking. The endpoint efficacy record in the time-series association record does not participate in drug sensitivity judgment in S1, but only serves as the data source for endpoint efficacy consistency verification in S3.
[0043] Through the above processing, S1 establishes the homology boundary, reference boundary, and temporal boundary of the organoid drug treatment object. This temporal correlation record provides S2 with the unified input required for interpreting culture perturbations and preserving baseline differences, and provides S3 with the temporal basis required for baseline calibration and post-drug response comparison. It can reduce training sample mismatch caused by different sample sources, different culture batches, different passage states, or different detection time windows.
[0044] In this embodiment, S2 reads the undosed reference unit, pre-dosed baseline observation content, drug plate reference record, biological origin boundary, culture state boundary, reference rollback marker, observation validity marker, and missing marker based on the time-series correlation record output by S1. It performs baseline observation standardization characterization processing on the baseline observation content and generates culture perturbation interpretation results and baseline difference retention results according to the dual-boundary resolution processing of culture perturbation interpretation and baseline difference retention. The culture perturbation interpretation results are used to characterize the non-drug-related changes that can be explained by the culture process, passage status, drug plate reference anomaly, or detection time offset. The baseline difference retention results are used to limit the true biological differences that need to be retained in the subsequent perturbation calibration process.
[0045] First, baseline observation standardization characterization processing was performed on the baseline observation content in the untreated reference unit and the baseline observation content before drug administration. Specifically, the number of organoids, organoid area, and organoid morphological compactness were extracted from the untreated reference unit and the baseline observation content before drug administration, respectively. The three types of observations were then converted to a unified scale. Since the number of organoids is a count, the organoid area is an area measure, and the organoid morphological compactness is a morphological description measure, before the three are included in the baseline observation characterization, they are converted to a unified scale according to the distribution boundary of qualified observation samples determined by the sample library review procedure within the same imaging batch, so that each observation is converted into a dimensionless result in the range of 0 to 1.
[0046] As a preferred implementation, the distribution boundary of qualified observation samples is formed by the 5th to 95th percentiles of qualified observation samples within the same imaging batch. When the observation is below the 5th percentile, it is converted according to the lower boundary; when the observation is above the 95th percentile, it is converted according to the upper boundary. This boundary is used to reduce the impact of extreme segmentation errors, local occlusion, and abnormal large clumps on baseline characterization, and to retain the main effective observation distribution within the same imaging batch. When the number of qualified observation samples within the same imaging batch does not reach the minimum number specified in the sample library review procedure, the distribution boundary of qualified observation samples from adjacent imaging batches within the same culture batch is adopted, or the default distribution boundary formed by historical review samples is adopted, and the insufficient sample situation is recorded in the analytical label set.
[0047] After completing the unified scale conversion, the results of organoid number, organoid area, and organoid morphological compactness are homogenized to ensure that the converted numerical directions are consistent with the effective growth state of the organoids. The results of organoid number and organoid area are homogenized based on the value ranges corresponding to the effective growth state in the historical review samples. The organoid morphological compactness results are established with morphological support mapping through historical review samples. The morphological intervals corresponding to the effective growth state in the baseline observation content are mapped as the first support result, and the morphological intervals corresponding to the fragmented state, incomplete boundary state, abnormal collapse state, and image occlusion state in the baseline observation content are mapped as the second support result. The first support result is used to retain the corresponding morphological items to participate in the formation of baseline observation characterization, and the second support result is used to mark the corresponding morphological items as morphological anomalies. The baseline observation characterization generated in this way serves as the common input for the culture common migration result, the culture perturbation interpretation result, and the baseline difference retention result.
[0048] After completing the baseline observation standardization and characterization process, the baseline observation characterization quantity corresponding to the untreated reference unit is compared with the baseline observation characterization quantity corresponding to the pre-drug baseline observation content under a unified relative time axis to generate the culture common offset result. The culture common offset result is used to characterize the baseline offset that already existed before drug exposure. Its inputs include the baseline observation characterization quantity corresponding to the untreated reference unit, the baseline observation characterization quantity corresponding to the pre-drug baseline observation content, the drug plate reference record, the observation valid marker and the missing marker. During processing, the baseline difference between the untreated reference unit and the pre-drug baseline observation content is first compared at the same relative time point. Then, combined with the stability of the reference well position in the drug plate reference record and the common fluctuation trend of untreated observations within the same culture batch, it is determined whether the baseline difference belongs to the common offset of the same batch.
[0049] Drug plate reference records are used to form reference stability judgments. Their output serves as the basis for forming culture common offset results and culture disturbance interpretation results. They do not directly participate in the formation of dynamic pharmacodynamic label master values. In a preferred implementation, reference well stability is formed by the degree of variation of reference well readings within the same drug plate. When the degree of variation of reference well readings falls within the fluctuation boundary of qualified drug plates in historical verification samples, the corresponding reference well is marked as reference stable. When the degree of variation of reference well readings exceeds the fluctuation boundary, a drug plate reference anomaly mark is generated. The fluctuation boundary is formed by the fluctuation distribution of reference wells of historical qualified drug plates. Preferably, the 75th to 90th percentiles of the fluctuation distribution are taken as the anomaly judgment boundary. When the number of historical qualified drug plate records does not reach the minimum number specified in the sample library verification procedure, the initial fluctuation boundary is formed according to the reference well stability boundary in the drug screening detection procedure.
[0050] Furthermore, based on the common offset results of culture and the offset conversion rules formed by drug screening detection procedures, culture calibration records, or historical verification samples, the offsets corresponding to passage status, matrix batch, inoculation density, imaging batch, and detection time window are determined, and culture disturbance interpretation results are generated. Each offset is a dimensionless result in the range of 0 to 1, used to represent the degree of explanation of the baseline observation offset by the corresponding culture or detection factor. The passage status offset is formed by comparing the current passage number, the stable culture time after recovery, and the stable passage boundary in the sample bank verification procedure; the matrix batch offset is formed by the fluctuation distribution of the current matrix batch in the historical untreated baseline observation; the inoculation density offset is formed by the deviation between the current inoculation density and the median value of the inoculation density of qualified samples in the same culture batch; the imaging batch offset is formed by the proportion of qualified image quality, sharpness fluctuation, and exposure integrity of the current imaging batch; and the detection time window offset is formed by the time difference between the actual observation time and the preset detection time point.
[0051] The culture perturbation interpretation results are formed by combining the culture common offset results and each offset according to the interpretation weights. The interpretation weights are determined by the explanatory power of the corresponding factors in the historical review samples for the fluctuation of the untreated baseline. When the number of historical review samples does not reach the minimum number specified in the sample library review procedure, the initial weights formed by the drug screening test procedure or culture calibration records are used, and the insufficient weight samples are recorded in the analytical label set. The culture perturbation interpretation results are dimensionless results in the range of 0 to 1. When the combined results exceed the range of 0 to 1, the range is limited according to the boundary value. The culture perturbation interpretation results are used to determine the explanatory range of non-drug-related changes, the correction basis for the effective range of baseline growth, and the correction basis for the label credibility weights. They are not output as drug sensitivity results.
[0052] While generating interpretation results of culture perturbation, baseline difference retention results are generated based on biological source boundaries, pre-drug baseline observations, historical baseline records, and historical response records. Specifically, historical baseline records are retrieved in the order of priority for the same patient source, supplemented by the same cancer subtype or the same lesion source, and records with untraceable sources, missing effective observation markers, or failure to meet the quality review conditions formed by the sample bank review procedure are removed to form a historical baseline set. Based on the stability of the baseline observation characteristics in the historical baseline set in different culture batches or different passage stages, historical baseline stability results are generated. Historical baseline stability results are dimensionless results in the range of 0 to 1, used to indicate the degree to which the current baseline difference is stably reproduced in historical baseline records.
[0053] In a preferred implementation, the historical baseline set includes at least three historical baseline records covering at least two culture batches, participating in the formation of historical baseline stability results. This value is used to form the minimum verifiable baseline basis in organoid sample bank scenarios with limited sample size, and to avoid single batch fluctuations being mistakenly identified as stable biological differences. When the number of historical baseline records does not reach this minimum number, a historical sample insufficiency marker is generated. The historical sample insufficiency marker is not used as the sole basis for deducting from the baseline difference retention result. If the same patient source, the same cancer subtype, or the same lesion source consistently shows the same direction of baseline difference in different batches, this difference is used as the basis for retaining the true biological difference; if the historical baseline records show that the difference cannot be stably reproduced, then the difference is not used as the basis for retaining the true biological difference.
[0054] Further, historical response records that have completed label verification are retrieved to form cross-batch response stability results. Historical response records include historical drug-induced response fragments, historical dynamic pharmacodynamic labels, or historical endpoint pharmacodynamic verification conclusions. S2 does not use dynamic pharmacodynamic labels or drug-induced response fragments that have not yet been formed for the current organoid drug treatment object to form cross-batch response stability results. Cross-batch response stability results are formed from historical response records from the same drug, the same cancer subtype, or the same patient source. During processing, it is determined whether the direction of historical response change is consistent in different culture batches, and its reliability is determined by combining the observation quality and verification status of the historical records. Cross-batch response stability results are dimensionless results in the range of 0 to 1, used to indicate the degree of support of historical response records for the current baseline difference that should be retained.
[0055] As a preferred implementation, when there are no fewer than 3 historical response records covering no fewer than 2 culture batches, a stable cross-batch response result is formed. This value is used to form the minimum response consistency basis among different culture batches, avoiding the influence of occasional responses in a single batch on the baseline difference retention judgment. When the number of records does not meet this condition, an insufficient cross-batch response mark is generated, and this item is not used as a separate basis for the baseline difference retention result.
[0056] After obtaining historical baseline stability results and cross-batch response stability results, baseline difference retention results are generated based on these results, sample source credibility, and culture perturbation interpretation results. Sample source credibility is determined by the completeness of patient source identifiers, cancer subtypes, lesion sources, and sampling time within the biological source boundary. Baseline difference retention results are dimensionless results ranging from 0 to 1, used to limit the retention strength of true biological differences during perturbation calibration. When historical baseline stability results and cross-batch response stability results support that the current baseline difference belongs to true biological differences, and the sample source credibility meets the sample library verification procedures, the baseline difference retention results are corrected towards the direction of true biological difference retention. When the culture perturbation interpretation results show that the current baseline difference is mainly explained by the culture state or detection state, and the historical baseline record cannot stably reproduce this difference, the baseline difference retention results are corrected towards the direction of culture perturbation interpretation. If there are historical sample insufficiency markers, cross-batch response insufficiency markers, or reference rollback markers, these markers are added to the parsing marker set and used in S3 for baseline growth effective interval correction or label credibility weight correction, and are not directly used as the basis for judging the absence of drug response.
[0057] The set of analytical markers generated by S2 includes markers for insufficient historical samples, insufficient cross-batch response, drug plate reference anomalies, missing observation time, insufficient baseline reference samples, and reference backtracking. The set of analytical markers is used to convey the processing boundaries of insufficient data sources, reference anomalies, and missing time to S3. If a certain source of anomaly has already been included in the interpretation during the formation of culture common offset results or culture perturbation interpretation results, it will not be used as the basis for correction for the same change in the same stage. If the source of anomaly still retains a residual response corresponding to the drug exposure time after subsequent baseline calibration, it will be identified by S3 as a candidate for drug-induced response fragment.
[0058] S2 outputs time-series correlation records, culture perturbation interpretation results, baseline difference retention results, culture common offset results, historical baseline stability results, cross-batch response stability results, and a set of analytical labels. Among them, the culture perturbation interpretation results are used by S3 to determine the interpretable range of non-pharmacological changes, the baseline difference retention results are used by S3 to limit the retention range of true biological differences, and the set of analytical labels is used by S3 to correct the effective range of baseline growth or the confidence weight of the label. Through the above processing, S2 separates the changes that can be explained by culture and detection factors from the true biological differences that need to be retained within the same object boundary, providing dual boundary constraints for the subsequent construction of baseline calibration growth trajectory and dynamic pharmacodynamic labels.
[0059] In this embodiment, S3 performs perturbation calibration on the pre-drug baseline state based on the time-series correlation records, culture perturbation interpretation results, baseline difference retention results, culture common offset results, historical baseline stability results, cross-batch response stability results, and parsing label set output by S2, generating a baseline-calibrated growth trajectory and a baseline growth effective interval; the post-drug response state is compared with the baseline-calibrated growth trajectory and the baseline growth effective interval to screen drug-induced response fragments; then, according to the consistency verification process between the response fragments and the endpoint efficacy, dynamic efficacy labels and label credibility weights are generated. This step is used to convert the endpoint efficacy record from a single detection result into a training supervision signal with time process constraints, culture perturbation deduction basis, and credibility control basis.
[0060] First, the set of effective time points before drug administration is determined based on the time-series correlation records. The effective time points before drug administration are located before the start of drug exposure, and both the untreated reference unit and the baseline observation content before drug administration have effective observation markers at this time point. Based on the set of effective time points before drug administration, the difference between the baseline observation characterization quantity before drug administration and the baseline observation characterization quantity corresponding to the untreated reference unit is compared to form the representative baseline difference before drug administration.
[0061] As a preferred implementation, when the number of effective time points before drug administration reaches the minimum number specified in the sample library verification procedure, the representative baseline difference before drug administration is determined by the observation quality-weighted median. The minimum number is preferably 2 to 4 effective time points before drug administration, determined by the observation frequency before drug administration in the drug screening detection procedure. This range is applicable to short-cycle continuous observation scenarios before drug administration in organoid drug screening, and can avoid the occasional fluctuation of a single baseline point directly determining the baseline calibration result. When the number of effective time points before drug administration does not reach the minimum number, the difference between the most recent effective baseline time point before the start of drug exposure is used to replace the representative baseline difference before drug administration, and a baseline point insufficiency marker is generated before drug administration. This marker participates in the formation of label credibility weights and is not used as a basis for judging the absence of drug response.
[0062] Based on the natural growth trajectory provided by the untreated reference unit, the interpretation results of culture perturbation, the baseline difference retention results, and the representative baseline difference before drug administration, a baseline-calibrated growth trajectory is generated:
[0063] ;
[0064] in, The index of organoid drug processing objects is determined by the homologous association key in the time-series association record; This represents a relative time point index, determined by the relative time axis in S1; Indicates the first The organoid drug treatment subjects were in the first Baseline calibration growth trajectory values at each relative time point; Indicates the first The untreated reference unit corresponding to the organoid drug treatment object is in the first... The baseline observation representation at each relative time point is derived from the baseline observation normalization representation processing in S2. Indicates the first The baseline difference preservation results for each organoid drug treatment object are derived from the double boundary analysis processing in S2; Indicates the first The interpretation results of culture perturbation of each organoid drug treatment object are derived from the double boundary analysis in S2; Indicates the first The representative baseline differences of each organoid drug treatment subject before administration are formed by the difference between the baseline observation characteristics of the pre-administration baseline within the set of effective time points before administration and the baseline observation characteristics of the unadministered reference unit. This function represents a range constraint. It outputs 0 when the input result is less than 0, 1 when the input result is greater than 1, and the input result itself when the input result is within the range of 0 to 1.
[0065] The baseline observations are dimensionless results within the range of 0 to 1. The representative baseline difference before drug administration is the difference between dimensionless results. Both the baseline difference retention results and the culture perturbation interpretation results are dimensionless results within the range of 0 to 1. Therefore, the product and summation terms in this formula have the same dimensions. The retention correction factor used to form the representative baseline difference before administration uses the baseline difference retention result as the main basis for retaining the true biological difference, and ensures that the culture perturbation interpretation result only corrects the part not supported by the baseline difference retention result. When the baseline difference retention result is close to 1, the representative baseline difference before administration is still retained even if the culture perturbation interpretation result is high. When the baseline difference retention result is low and the culture perturbation interpretation result is high, the representative baseline difference before administration is corrected. This treatment is used to avoid the true biological difference being excessively weakened by the culture perturbation interpretation result.
[0066] This formula uses the natural growth trajectory of the untreated reference unit as the basis for baseline time variation. The representative baseline difference before drug administration represents the initial baseline shift between the organoid drug-treated object and the untreated reference unit. The degree of retention of this initial baseline shift in the baseline-calibrated growth trajectory is jointly controlled by the baseline difference retention result and the culture perturbation interpretation result. The range constraint function is used to ensure that the output baseline-calibrated growth trajectory value is still within the uniform scale range of the baseline observation representation. Reference backtracking, insufficient historical samples, or missing observation time are not repeatedly deducted in this formula, but are further processed through the effective baseline growth interval and label confidence weight.
[0067] The effective baseline growth interval is generated based on the baseline-calibrated growth trajectory. The effective baseline growth interval is centered on the baseline-calibrated growth trajectory. Its width is determined by the baseline fluctuation width, culture common offset results, culture perturbation interpretation results, historical baseline stability results, and analytical label set. The baseline fluctuation width is formed by historical qualified undoped baseline observation records. Preferably, the 75th to 90th percentile of the natural fluctuation half-width in the historical qualified undoped baseline observation records is taken as the baseline boundary. The baseline fluctuation width is the dimensionless half-width result in the range of 0 to 1, preferably 0.05 to 0.20. This range is applicable to the natural growth fluctuation boundary of organoids at a uniform scale of 0 to 1, which can cover slight imaging errors, segmentation fluctuations, and normal proliferation differences. At the same time, it avoids including the obvious response after dosing as a whole in the natural baseline fluctuation. When there are insufficient historical qualified undoped baseline observation records, the baseline fluctuation boundary in the sample library verification procedure is used, and the situation of insufficient historical samples is recorded in the analytical label set.
[0068] When the results of culture common offset or culture perturbation interpretation indicate that there is an interpretable non-pharmacological offset in the current baseline state, the effective baseline growth interval for the corresponding time point is corrected according to the interval width correction rule formed by the sample library review procedure. When there are reference rollback markers, drug plate reference anomaly markers, missing observation time markers, or insufficient baseline points before drug administration markers, the time points corresponding to the markers are written into the parsing marker set, and the subsequent label confidence weights are corrected according to the label confidence weight formation rule. The effective baseline growth interval is used to determine whether the response state after drug administration exceeds the natural baseline fluctuation range, and does not directly output drug sensitivity results.
[0069] The post-drug response state was converted into a post-drug response observation characterization quantity and compared with the baseline calibrated growth trajectory and the effective range of baseline growth. The conversion of the post-drug response state adopted a unified scale and homogenization direction consistent with the baseline observation standardization characterization processing in S2, so that the post-drug response observation characterization quantity could be compared with the baseline calibrated growth trajectory under the same evaluation scale. If the post-drug response observation characterization quantity is within the effective range of baseline growth, the change is classified as an explanation of natural baseline fluctuation or culture disturbance and is not included in the screening of drug-induced response fragments. If the post-drug response observation characterization quantity occurs after the start of drug exposure, has a valid response observation marker, exceeds the effective range of baseline growth, and still retains a residual response after processing the culture disturbance explanation results, then this time point is determined as the drug-induced response time point.
[0070] For time points in the post-drug response state where the number of organoids is lower than the baseline lower limit, if the time point meets the image quality and imaging field of view conditions formed by the drug screening detection procedure, and occurs after the start of drug exposure, then the time point is marked as a candidate time point for low-object-number response and included in the screening of drug-induced response fragments. This rule is used to avoid misjudging drug-induced killing, organoid disintegration, or drug-induced disappearance as invalid observations. If the candidate time point for low-object-number response also has records of field of view loss, focal plane drift, or segmentation failure, then the time point is not used as a drug-induced response time point, but is entered into the analytical label set for label confidence weight correction.
[0071] When the post-drug response change can be covered by the culture perturbation explanation results, it is further determined whether the change has the same temporal relationship and the same direction of synchronous change in the undoped reference unit. If there is a synchronous change in the undoped reference unit, the change is included in the culture perturbation explanation range. If there is still a continuous residual response corresponding to the drug exposure time after baseline calibration, a calibrated residual response label is generated, and the corresponding response time point is retained as a candidate for drug-induced response fragment. The calibrated residual response label is used to characterize that the corresponding response change is still related to the drug exposure time after deducting the culture perturbation, and will not be repeatedly corrected as an interference term of the same response fragment.
[0072] Continuous drug-induced response time points with consistent response directions are merged to generate drug-induced response fragments. During merging, these fragments must meet the minimum response duration or minimum number of effective points conditions defined by the drug screening detection procedure. As a preferred implementation, the minimum response duration condition is at least one detection interval, and the minimum number of effective points condition is at least two consecutive effective time points. When the drug screening experiment only sets a single post-dose detection time point, single-point responses occurring after the start of drug exposure, exceeding the baseline growth effective range, and retaining residual responses are retained as candidate drug-induced response fragments. Subsequent endpoint consistency verification results that do not meet the consistency conditions defined by the sample library verification procedure are considered as drug-induced response fragments. At that time, the supervision contribution of the training sample corresponding to the single-point response is corrected by the label confidence weight. When there is a single missing time point in the segment, the response time points before and after the missing time point are bridged only if the missing time point is caused by the missing observation time marker and the response direction before and after the missing time point is consistent. Each drug-induced response segment includes response initiation time, response duration, response intensity, recovery trend and late marker. Response initiation time is used to characterize the time position of drug effect, response duration is used to characterize the continuity of drug effect, response intensity is used to characterize the degree of deviation from the baseline growth effective range, recovery trend is used to identify the situation of early response followed by decline, and late marker is used to identify drug effect that appears later.
[0073] Dynamic response intensity results are generated based on drug-induced response fragments. These results are dimensionless and range from 0 to 1. Specifically, the response intensity, response duration, recovery trend, fragment observation quality, and fragment weight of each drug-induced response fragment are converted into sub-results within the range of 0 to 1. These sub-results are then combined according to fragment weights determined by historical review samples to form the dynamic response intensity results. The fragment weights are determined by the correspondence between the response fragment and the drug exposure time, the number of valid time points within the fragment, and the fragment observation quality. When multiple drug-induced response fragments exist, fragments whose response intensity, response duration, and fragment observation quality meet the drug screening detection procedures are selected as the drug-induced response fragments used to form the dynamic response intensity results. If there are no valid drug-induced response fragments and no observational gaps or quality anomalies affecting the judgment, the dynamic response intensity result is set to 0. If there are no valid drug-induced response fragments but there are observational gaps, quality anomalies, or reference rollback markers, a marker indicating no valid drug-induced response fragment is generated and passed to the label credibility weight formation process. This marker is not directly used as a conclusion of drug ineffectiveness.
[0074] The endpoint efficacy records are converted into endpoint efficacy normalized results, which are then used as the source of consistency verification for dynamic efficacy labels. The endpoint efficacy normalized results are dimensionless results within the range of 0 to 1, and the numerical direction is consistent with the degree of drug response. For survival rate-type endpoint records, the direction of low survival rate corresponding to high drug response is used for normalization; for inhibition rate-type endpoint records, the direction of high inhibition rate corresponding to high drug response is used for normalization; for IC50-type endpoint records, the direction of low IC50 corresponding to high drug response is used for normalization; for AUC-type endpoint records, the direction of normalization is based on the correspondence between AUC and drug response degree in the drug screening test procedure. The endpoint efficacy normalized results and dynamic response intensity results are compared under the same 0 to 1 evaluation scale to form the endpoint consistency verification results. If the endpoint efficacy record is missing or does not meet the quality review conditions formed by the drug screening test procedure of the sample library, the endpoint consistency verification results are not calculated, and an endpoint verification missing marker is generated.
[0075] Based on the dynamic response intensity results, endpoint efficacy normalization results, and endpoint consistency verification results, a dynamic efficacy label is generated:
[0076]
[0077] in, Indicates the first The dynamic pharmacodynamic labels of each organoid drug treatment object are the output of this formula; This represents the index of the organoid drug treatment object, and its meaning is consistent with the index in the baseline calibration growth trajectory formula; Indicates the first The dynamic response intensity results of each organoid drug treatment object are derived from the response intensity, response duration, recovery trend, fragment observation quality, and fragment weight of the drug-induced response fragment; Indicates the first The endpoint efficacy normalization result for each organoid drug treatment object is formed by converting the endpoint efficacy record through homogenization and unified scaling. Indicates the first The endpoint consistency verification results for each organoid drug treatment object are formed by the degree of consistency between the dynamic response intensity results and the endpoint efficacy normalization results. The weights representing the dynamic response intensity results; The weight representing the endpoint efficacy normalization result; This represents the range constraint function, which has the same meaning as the range constraint function in the baseline calibration growth trajectory formula.
[0078] The dynamic response intensity result, endpoint efficacy normalization result, and endpoint consistency verification result are all dimensionless results within the range of 0 to 1. The weights of the dynamic response intensity result and the endpoint efficacy normalization result are both non-negative weights, and their sum is 1. In a preferred embodiment, the weight of the dynamic response intensity result is 0.60 to 0.80, and the weight of the endpoint efficacy normalization result is 1 minus the weight of the dynamic response intensity result. This weight range is determined by the consistency between the dynamic response intensity result, the endpoint efficacy normalization result, and the reviewed drug response conclusion in the historical review samples. Since the weight of the dynamic response intensity result is positive, the denominator is not 0. Since the input quantities in the formula are all within the range of 0 to 1, and the outer layer of the formula sets a range limiting function, the dynamic efficacy label is within the range of 0 to 1.
[0079] This formula uses the dynamic response intensity result as the primary basis for forming the dynamic pharmacodynamic label, and the endpoint pharmacodynamic normalization result as a supplementary basis after endpoint consistency verification. The closer the endpoint consistency verification result is to 1, the more fully the endpoint pharmacodynamic normalization result participates in the formation of the dynamic pharmacodynamic label; the closer the endpoint consistency verification result is to 0, the more the dynamic pharmacodynamic label is formed by the dynamic response intensity result. This normalization gating fusion method is used to avoid directly processing unreliable endpoint records as low drug response intensity, so that the dynamic pharmacodynamic label mainly represents the drug-induced response intensity, and endpoint inconsistency, missing data, or quality anomalies are handled by the label credibility weight. When the endpoint pharmacodynamic record is missing or does not meet the quality verification conditions, the endpoint consistency verification result is not calculated and is treated as 0 in this formula; at this time, the dynamic pharmacodynamic label is formed by the dynamic response intensity result, and the dynamic pharmacodynamic label serves as the supervision target of the drug prediction neural network in S4, while the label credibility weight serves as the adjustment basis for the contribution of training samples in S4.
[0080] It should be noted that the dynamic response intensity results, endpoint efficacy normalization results, and endpoint consistency verification results are used for the construction of supervisory labels for completed drug screening records, and only serve as the basis for the formation of dynamic efficacy labels during the training phase of the drug prediction neural network. During the prediction phase, the input to the drug prediction model is limited to information available before drug exposure and candidate drug characteristics; it is not necessary to obtain the post-dose response state or endpoint efficacy records of the sample to be predicted. Therefore, dynamic efficacy labels are used for constructing supervisory signals for training, while the drug prediction model is used for pre-prediction of candidate drugs; the two are distinct in their usage phases.
[0081] Based on the results of culture perturbation interpretation, baseline difference retention, cross-batch response stability, endpoint consistency verification, drug-induced response fragment quality, post-calibrated residual response markers, and the set of analytical markers, a label confidence weight is generated. The label confidence weight is a dimensionless result ranging from 0 to 1, used to adjust the contribution of training samples in S4 to the neural network loss function; it does not directly represent the strength of drug sensitivity. The culture perturbation interpretation result is used to mark the impact of non-drug-related changes in the current label; the baseline difference retention result is used to mark whether true biological differences are preserved; the cross-batch response stability result is used to mark whether historical responses have cross-batch support; the endpoint consistency verification result is used to mark whether the dynamic response is consistent with the endpoint efficacy; the drug-induced response fragment quality is used to mark whether the response fragment has temporal continuity and observational reliability; the post-calibrated residual response marker is used to mark the residual response basis that still corresponds to the drug exposure time after deducting culture perturbation; the set of analytical markers is used to record the impact of reference backtracking, insufficient samples, drug plate reference anomalies, missing observation time, and missing endpoint verification on the label confidence weight. If the same anomaly source has already been deducted during the drug-induced response fragment screening process, it will not be deducted again during the label confidence weight formation process for the same response fragment.
[0082] S3 outputs time-series correlation records, culture perturbation interpretation results, baseline difference retention results, baseline-calibrated growth trajectory, effective baseline growth interval, drug-induced response fragment set, dynamic response intensity results, dynamic pharmacodynamic labels, label credibility weights, and label formation marker set. The label formation marker set consists of markers from the parsed marker set that participate in the formation of dynamic pharmacodynamic labels and label credibility weights, including markers for insufficient baseline points before dosing, low-object-number response candidate markers, residual response markers after calibration, markers for no effective drug-induced response fragments, and markers for missing endpoint verification. S4 uses dynamic pharmacodynamic labels as the monitoring target, label credibility weights as the basis for sample contribution control, and drug-induced response fragments as auxiliary monitoring information.
[0083] In this embodiment, S4, based on the dynamic pharmacodynamic label, label confidence weight, drug-induced response fragment set, baseline calibration growth trajectory, effective baseline growth range, and label formation mark set output by S3, inputs the drug prediction input features formed by information available before drug exposure in the time-series associated records into the drug prediction neural network. The dynamic pharmacodynamic label is used as the supervision target, and the model is trained according to the confidence supervision training process limited by the label confidence weight. The output is a drug prediction model used to predict the sensitivity results of candidate drugs or the drug response level.
[0084] Drug prediction input features are formed from information available before drug exposure in time-series correlation records. These features include sample source features, pre-exposure baseline features, and candidate drug features. Sample source features are formed by patient source identifier, cancer subtype, lesion source, sampling time, and passage status within the biological source boundary. Pre-exposure baseline features are formed by pre-drug baseline observation content, baseline calibration growth trajectory, effective baseline growth range, and pre-exposure observation completeness. Candidate drug features are formed from drug processing data or sample library drug records, preferably including drug identifier, drug concentration, exposure duration, and drug action category. Drug prediction input features do not include post-drug response status, drug-induced response fragments, dynamic pharmacodynamic labels, and label confidence weights, so that the drug prediction model can output prediction results based on pre-exposure information and candidate drug information during the prediction phase.
[0085] The drug prediction neural network includes a sample baseline encoder, a drug encoder, a drug response interaction layer, and a dynamic efficacy prediction head. The sample baseline encoder encodes the sample source features and pre-exposure baseline features into a sample baseline representation. The drug encoder encodes the candidate drug features into a candidate drug representation. The drug response interaction layer fuses the interaction between the sample baseline representation and the candidate drug representation. The dynamic efficacy prediction head outputs the predicted dynamic efficacy results.
[0086] As a preferred implementation, the drug prediction neural network is implemented using a feedforward neural network or an attention network, with 2 to 5 hidden layers and 64 to 512 hidden units per layer; when the sample library drug records or sample similarity relationship records form a graph structure input, the drug prediction neural network can also be implemented using a graph neural network.
[0087] In the trusted supervised training process, the dynamic pharmacodynamic label serves as the primary supervised target. The label trust weight is used to adjust the contribution of training samples to model parameter updates. The label trust weight is a dimensionless result within the range of 0 to 1, derived from the label trust weight formation process in S3. Training samples whose label trust weight reaches the trust threshold participate in primary supervised training as high-trust supervised samples. Training samples whose label trust weight does not reach the trust threshold but are not deemed unusable by the sample library review procedure retain their single-sample supervised training contribution, and the sample contribution is adjusted according to the label trust weight. Training samples with a label trust weight of 0 do not participate in the primary supervised update of the dynamic pharmacodynamic label. Their pre-exposure baseline features are retained for sample distribution statistics and batch coverage statistics. The trust threshold is a dimensionless threshold within the range of 0 to 1, preferably 0.60 to 0.80, determined by the consistency distribution of the dynamic pharmacodynamic label and the reviewed drug response conclusion in the historical validation set. When the number of historical validation sets is insufficient, the initial trust threshold in the sample library review procedure is used, and a trust threshold insufficient sample marker is generated.
[0088] In addition to the main supervised training, auxiliary supervision is performed based on drug-induced response fragments. Specifically, response initiation time, response duration, response intensity, recovery trend, and late-onset label are extracted from the set of drug-induced response fragments to form auxiliary supervision information for response fragments. When multiple drug-induced response fragments exist, drug-induced response fragments for auxiliary supervision are selected based on response intensity, response duration, fragment observation quality, and label credibility weight. When there are no effective drug-induced response fragments but there are no effective drug-induced response fragment labels, missing observation time labels, or reference backoff labels, the corresponding training samples are not used as the basis for non-response auxiliary supervision. Their training contribution limited by dynamic pharmacodynamic labels and label credibility weights is retained. This processing enables the drug prediction neural network to obtain constraints on response initiation time, duration, recovery state, and late-onset response process while learning dynamic pharmacodynamic labels, avoiding static results formed by only fitting the endpoint pharmacodynamic records.
[0089] Based on the interpretation results of culture perturbation and the retention results of baseline differences, culture perturbation sensitivity constraints are applied. These constraints limit the model from mislearning culture state differences as drug sensitivity features and retain baseline differences supported by real biological differences. As a preferred implementation, when the interpretation results of culture perturbation reach the perturbation constraint threshold and the retention results of baseline differences do not reach the retention confirmation threshold, a first sensitivity constraint is applied to the feature channels corresponding to the culture state boundary in the sample baseline representation. When the interpretation results of culture perturbation reach the perturbation constraint threshold and the retention results of baseline differences reach the retention confirmation threshold, a second sensitivity constraint is applied to the feature channels corresponding to the culture state boundary in the sample baseline representation. The constraint strength of the first sensitivity constraint is greater than that of the second sensitivity constraint. Both the perturbation constraint threshold and the retention confirmation threshold are dimensionless thresholds in the range of 0 to 1, preferably 0.60 to 0.75, determined by the consistency between the interpretation results of culture perturbation in the historical validation set, the retention results of baseline differences, and the drug response review conclusions. When the historical validation set is insufficient, the initial threshold formed by the sample library review procedure is used and updated after adding new review samples.
[0090] The implementation of culture perturbation sensitivity constraints includes: masking, replacing, or perturbing the feature channels corresponding to the culture state boundary, and comparing the predicted dynamic efficacy results before and after the treatment; when the change in the predicted dynamic efficacy results exceeds the sensitivity boundary formed by the historical validation set, a culture perturbation sensitivity constraint record is generated, and sensitivity constraints are applied to the model parameter update process. The sensitivity boundary is a dimensionless difference boundary in the range of 0 to 1, preferably 0.05 to 0.15, formed by the difference distribution of the predicted dynamic efficacy results before and after culture state perturbation in the historical validation set; when the historical validation set is insufficient, the initial sensitivity boundary in the sample library verification procedure is used. This treatment is used to limit the model's dependence on culture state features when the culture perturbation interpretation results are dominant, and to retain the contribution of the corresponding baseline difference to the prediction results when the baseline difference retention results support the real biological differences.
[0091] Based on the response fragment consistency pairing rules formed by drug action category, response fragment direction, label confidence weight, culture perturbation difference, and baseline difference retention difference, cross-batch constraint sample pairs are determined. Training sample pairs must simultaneously meet the following conditions to be included in the cross-batch constraint sample pair: candidate drugs have the same drug identifier or the same drug action category; the response direction of the drug-induced response fragments is consistent; the label confidence weights of both training samples reach the confidence threshold; the culture batches of the two training samples are different; the difference in the culture perturbation interpretation results between the two training samples does not exceed the perturbation difference threshold; and the difference in the baseline difference retention results between the two training samples does not exceed the retention difference threshold. Both the perturbation difference threshold and the retention difference threshold are dimensionless thresholds in the range of 0 to 1, preferably 0.10 to 0.25, determined by the consistency distribution between cross-batch dynamic pharmacodynamic labels and the reviewed drug response conclusions in historical review samples. When historical review samples are insufficient, the initial difference threshold specified in the sample library review procedure is used, and a pairing rule sample shortage marker is generated. Training samples that do not enter the cross-batch constraint sample pair are not removed and still retain the single-sample supervised training contribution limited by the dynamic pharmacodynamic label and label confidence weight.
[0092] For cross-batch constrained sample pairs, a response fragment consistency constraint is applied. During the application, the predicted dynamic efficacy results, response fragment auxiliary supervision results, and drug-induced response fragment direction of the cross-batch constrained sample pairs are compared. If the cross-batch constrained sample pairs meet the response fragment consistency pairing rules in terms of dynamic efficacy label, response fragment direction, and label confidence weight, then the two are used as consistency constraint sample pairs for training. If the two exceed the difference threshold in terms of culture perturbation interpretation results or baseline difference retention results, then no response fragment consistency constraint is applied, and only their respective single-sample supervised training contributions are retained. This treatment is used to avoid the forced consistency of differences in real patient origin or cancer subtype differences and to improve the stability of drug prediction results across culture batches.
[0093] During training, the parameters of the drug prediction neural network are updated by integrating primary supervised training, response fragment auxiliary supervision, culture perturbation sensitivity constraints, and response fragment consistency constraints. After each training round, the dynamic drug efficacy label prediction error, response fragment auxiliary supervision error, culture perturbation sensitivity constraint results, and cross-batch response fragment consistency results are calculated based on the validation set. The dynamic drug efficacy label prediction error is represented by a normalized error in the range of 0 to 1. The upper limit of the allowable dynamic drug efficacy label prediction error in the validation set is preferably 0.15 to 0.25, which is determined by the consistency distribution between the dynamic drug efficacy label and the verified drug response conclusion in the historical validation set. If the difference between the predicted dynamic drug efficacy results corresponding to the culture perturbation sensitivity constraint records in the validation set still exceeds the sensitivity boundary, the corresponding training round is not used as the final model parameter. If the response fragment auxiliary supervision error does not meet the stability conditions formed by the sample library verification procedure in multiple consecutive training rounds, it regresses to the previous stable model parameters. The training round, learning rate, batch size, and stopping condition are determined by the sample library verification procedure or the training validation set.
[0094] The same source of culture disturbance is used for different purposes in S3 and S4. S3 is used to determine whether changes after drug administration form drug-induced response fragments, while S4 is used to constrain the model's fit to culture state characteristics. Changes that have been included in the interpretable range of culture disturbances in S3 and excluded from the drug-induced response fragment screening are no longer penalized as auxiliary supervision objects of response fragments in S4. Residual responses that are retained as candidates for drug-induced response fragments after calibration in S3 are not removed again in S4 due to the same source of culture disturbance. The influence of culture disturbance is only recorded once in the sample-level label confidence weight. This processing is used to avoid repeated deduction of the same abnormal source during fragment screening and model training.
[0095] After training, the drug prediction model is output. For candidate drugs to be predicted, the sample source characteristics, pre-exposure baseline characteristics, and candidate drug characteristics of the organoid samples to be predicted are obtained and input into the trained drug prediction model. The model outputs the predicted dynamic efficacy results and, based on the predicted dynamic efficacy results, outputs the candidate drug sensitivity results or drug response levels. When using drug response levels for output, the level boundaries are determined by the distribution of dynamic efficacy labels in historical verification samples or by drug screening procedures. As a preferred implementation, the predicted dynamic efficacy results are divided into 3 to 5 levels of drug response levels, with the number of levels being the number of discrete levels. This is used to distinguish the response states of candidate drugs with no response, low response, medium response, and high response. When historical verification samples are insufficient, the level boundaries are formed by the response level boundaries in the drug screening procedures and are updated after new verification samples are added. The drug prediction model simultaneously outputs prediction credibility hints, which are formed by the completeness of the pre-exposure baseline observation of the sample to be predicted, the width of the effective baseline growth interval, the completeness of candidate drug characteristics, and the similarity to the distribution of the training samples. The candidate drug sensitivity results or drug response levels output by the drug prediction model are used for organoid drug screening and model evaluation, and are not directly used for clinical dosing decisions.
[0096] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for training a drug prediction model based on tumor organoid biobank data, characterized in that, include: Obtain sample library association data to characterize the homology boundary and efficacy observation process of organoid drug treatment objects, perform time-series association construction processing according to sample treatment boundary information and efficacy observation time sequence, and generate time-series association records; Based on time-series correlation records, baseline observation standardization characterization processing is performed on the baseline observation content, and the culture perturbation interpretation results and baseline difference retention results are generated according to the dual-boundary analysis processing of culture perturbation interpretation and baseline difference retention. Based on the interpretation results of culture perturbation and the results of baseline difference retention, perturbation calibration is performed on the baseline state before drug administration to generate the baseline calibrated growth trajectory and the effective range of baseline growth. The response state after drug administration is compared with the baseline calibrated growth trajectory and the effective range of baseline growth to screen drug-induced response fragments. Dynamic efficacy labels and label confidence weights are generated according to the consistency verification between the response fragments and the endpoint efficacy. The drug prediction input features, formed by information available before drug exposure in time-series association records, are input into the drug prediction neural network. Dynamic efficacy labels are used as the supervision target, and the model is trained according to the credible supervision training process limited by the credible weight of the labels. The output is a drug prediction model used to predict the sensitivity results or drug response level of candidate drugs.
2. The method for training a drug prediction model based on tumor organoid biobank data according to claim 1, characterized in that, The sample library associated data includes biological source data, culture status data, drug treatment data, time-series observation data, and reference verification data. Among them, biological source data is used to define the source of organoid drug treatment objects, culture status data is used to define the culture process and passage status of organoid drug treatment objects, drug treatment data is used to define the conditions for the action of candidate drugs, time-series observation data is used to define the observation process before and after drug exposure, and reference verification data is used to define the untreated reference and the basis for endpoint efficacy verification. The operation of obtaining sample library association data for characterizing the homology boundary and efficacy observation process of organoid drug treatment objects includes: assigning data roles to the sample library association data according to the data roles corresponding to biological source data, culture status data, drug treatment data, time-series observation data and reference verification data in the sample library recording specifications, drug screening detection procedures and historical verification samples.
3. The method for training a drug prediction model based on tumor organoid biobank data according to claim 2, characterized in that, The operations for constructing time-series associations based on sample processing boundary information and efficacy observation time sequence include: establishing homology association keys based on biological source boundaries, culture state boundaries, and drug processing boundaries; and grouping boundary definition information, drug processing information, and efficacy observation information of the same organoid drug-treated object into the same time-series association record according to the homology association keys; screening untreated reference units with homology boundaries to organoid drug-treated objects from the sample library association data with untreated reference roles; when no untreated reference unit that meets the homology reference conditions formed by the sample library verification procedure is found, selecting reference backoff sources in the adjacent passages of the same sample, the same culture batch of the same cancer subtype, and the historical baseline records of the same cancer subtype according to the reference backoff order formed by the sample library verification procedure, and generating reference backoff markers; and performing relative time axis alignment on untreated references, pre-drug observations, post-drug responses, and endpoint efficacy records with the drug exposure start time as the time zero point, generating time-series association records containing observation validity markers and reference backoff markers.
4. The method for training a drug prediction model based on tumor organoid biobank data according to claim 3, characterized in that, The baseline observation standardization and characterization process includes: extracting the number of organoids, organoid area, and organoid morphological compactness from the baseline observations in the untreated reference and the baseline observations before drug administration, respectively; converting the number of organoids, organoid area, and organoid morphological compactness to a uniform scale according to the distribution boundaries of qualified observation samples determined by the sample library review procedure within the same imaging batch, and performing homogenization processing according to the effective growth state of the organoids; wherein, the organoid morphological compactness results establish a morphological support mapping through historical review samples, mapping the morphological intervals corresponding to the effective growth state in the baseline observations to the first support result, and mapping the morphological intervals corresponding to the fragmented state, incomplete boundary state, abnormal collapse state, and image occlusion state in the baseline observations to the second support result. The first support result is used to retain the corresponding morphological items to participate in the formation of baseline observation characterization, and the second support result is used to mark the corresponding morphological items as morphological anomalies to generate baseline observation characterization for culture perturbation interpretation and baseline difference preservation.
5. The method for training a drug prediction model based on tumor organoid biobank data according to claim 4, characterized in that, The process of generating culture perturbation interpretation and baseline difference retention results according to the dual-boundary analysis of culture perturbation interpretation and baseline difference retention includes: comparing the baseline observation characteristics corresponding to the untreated reference with the baseline observation characteristics corresponding to the pre-drug-treated baseline observations at relative time points, and combining the stability of the reference well positions in the drug plate reference record and the common fluctuation trend of untreated observations within the same culture batch to generate culture common offset results; based on the culture common offset results and the offset conversion rules formed by drug screening detection procedures, culture calibration records, or historical verification samples, determining the passage status, matrix batch, inoculation density, and imaging batch. The offset corresponding to the detection time window is calculated, and the culture perturbation interpretation result is generated. Historical baseline stability results are generated by retrieving historical baseline records from the same patient source, the same cancer subtype, or the same lesion source, and historical response records that have completed label verification are generated to form cross-batch response stability results. Based on the historical baseline stability results, cross-batch response stability results, sample source credibility, and culture perturbation interpretation results, baseline difference retention results are generated. When the number of historical samples does not reach the minimum number specified in the sample library verification procedure, a historical sample insufficiency mark is generated. The historical sample insufficiency mark is not used as the sole basis for deducting from the baseline difference retention results.
6. The method for training a drug prediction model based on tumor organoid biobank data according to claim 5, characterized in that, The perturbation calibration of the pre-drug baseline status includes: determining the set of valid pre-drug time points from the relative time points observed before drug exposure begins; forming a representative pre-drug baseline difference based on the difference between the pre-drug baseline observation characteristics and the baseline observation characteristics corresponding to the untreated reference in the set of valid pre-drug time points; when the number of valid pre-drug time points reaches the minimum number specified in the sample library review procedure, the representative pre-drug baseline difference is determined using the observation quality-weighted median; when the number of valid pre-drug time points does not reach the minimum number specified in the sample library review procedure, the difference of the most recent valid baseline time point before drug exposure begins is used to replace the representative pre-drug baseline difference, and a pre-drug baseline point insufficiency marker is generated; and combining the natural growth trajectory provided by the untreated reference, the strength of true biological difference retention limited by the baseline difference retention results, and the strength of non-drug-related offset correction limited by the culture perturbation interpretation results to generate a baseline-calibrated growth trajectory.
7. The method for training a drug prediction model based on tumor organoid biobank data according to claim 6, characterized in that, The operations for generating baseline-calibrated growth trajectories and effective baseline growth intervals include: using the baseline-calibrated growth trajectory as the center of the effective baseline growth interval; determining the interval width of the effective baseline growth interval based on the baseline fluctuation width, culture common offset results, culture perturbation interpretation results, historical baseline stability results, and the set of analytical markers; the baseline fluctuation width is formed from historical qualified untreated baseline observation records; when there are reference rollback markers, drug plate reference anomaly markers, or missing observation time markers, the effective baseline growth interval at the corresponding time point is corrected according to the interval width correction rules formed by the sample library review procedure, or the subsequent label credibility weights are corrected according to the label credibility weight formation rules; among them, the culture perturbation interpretation results are used to determine the explainable range of non-drug-related changes, and the baseline difference retention results are used to limit the retention range of true biological differences. Together, they limit whether the response state after drug administration exceeds the effective baseline growth interval.
8. The method for training a drug prediction model based on tumor organoid biobank data according to claim 7, characterized in that, The process for screening drug-induced response fragments includes: comparing the post-drug response state with the baseline calibrated growth trajectory and the effective baseline growth range; identifying response time points that occur after drug exposure begins, have valid response observations, exceed the effective baseline growth range, and retain residual responses after processing with culture perturbation interpretation results as drug-induced response time points; when there are time points in the post-drug response state where the number of organoids is lower than the baseline object number limit formed by the sample library verification procedure, if this time point meets the image quality and imaging field of view conditions formed by the drug screening detection procedure and occurs after drug exposure begins, then this time point is marked as a low-object-number response candidate time point and included in the screening of drug-induced response fragments; when there are still continuous residual responses corresponding to the drug exposure time after processing with culture perturbation interpretation results for post-drug response changes, a calibrated residual response marker is generated, and the corresponding response time point is retained as a candidate for drug-induced response fragments; continuous drug-induced response time points with consistent response directions are merged into drug-induced response fragments, and response initiation time, response duration, response intensity, recovery trend, and late-onset markers are generated for each drug-induced response fragment.
9. The method for training a drug prediction model based on tumor organoid biobank data according to claim 8, characterized in that, The process of generating dynamic efficacy labels and label credibility weights based on the consistency verification between response fragments and endpoint efficacy includes: generating dynamic response intensity results based on the response initiation time, response duration, response intensity, recovery trend, late-onset markers, and fragment observation quality of the drug-induced response fragment; converting the endpoint efficacy record into a normalized endpoint efficacy result in the same direction as the drug response degree, and comparing the consistency between the dynamic response intensity result and the normalized endpoint efficacy result to generate an endpoint consistency verification result; when the endpoint efficacy record is missing or does not meet the quality review conditions formed by the drug screening test procedure of the sample library, the weight of the endpoint consistency verification item is reset to zero, and a dynamic efficacy label is formed from the dynamic response intensity result; generating label credibility weights based on the culture perturbation interpretation results, baseline difference retention results, cross-batch response stability results, endpoint consistency verification results, drug-induced response fragment quality, calibrated residual response markers, and parsing marker set, and when the same source of anomaly has already been deducted during the drug-induced response fragment screening process, it will not be deducted again during the label credibility weight formation process of the same response fragment.
10. The method for training a drug prediction model based on tumor organoid biobank data according to claim 1, characterized in that, The model training process, which follows a trusted supervision training approach defined by label trust weights, includes the following steps: First, drug prediction input features are formed from information available before drug exposure in time-series correlation records. These features are then input into the drug prediction neural network, with dynamic pharmacodynamic labels serving as the primary supervision target. Second, the contribution of training samples to model parameter updates is adjusted based on label trust weights. Third, response fragment auxiliary supervision is performed based on drug-induced response fragments. Fourth, culture perturbation sensitivity constraints are applied based on the interpretation results of culture perturbation and the baseline difference retention results. Fifth, cross-batch constraint sample pairs are determined according to the response fragment consistency pairing rules formed by drug action category, response fragment direction, label trust weights, culture perturbation differences, and baseline difference retention differences. Response fragment consistency constraints are then applied to these cross-batch constraint sample pairs. Training samples not included in the cross-batch constraint sample pairs retain their single-sample supervised training contribution defined by dynamic pharmacodynamic labels and label trust weights. The culture perturbation sensitivity constraint rules and response fragment consistency pairing rules are formed from historical validation sets, verification samples, or sample library verification procedures.