Solid tumor driving gene detection reading standardization system and kit for cold light enzyme

By constructing a solid tumor-driven gene detection reading standardization system for cold-spotase, multi-source data and intelligent algorithms are used to optimize the weak signal analysis of the microplate reader, the problems of heterogeneity of the detection platform and the difference in sample types are solved, and the standardization and accuracy of solid tumor-driven gene detection are improved, which is especially suitable for primary medical institutions.

CN120581070AInactive Publication Date: 2025-09-02TIANJIN FUXUN TECH DEV CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510670223.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-09-02
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, solid tumor-driven gene detection has heterogeneity, sample type differences and data processing processes, resulting in poor comparison of test results, especially in primary medical institutions that affect the consistency of clinical diagnosis. Traditional correction methods cannot effectively integrate multi-source data, resulting in low-expression driver gene detection accuracy.

Method used

A solid tumor-driven gene detection reading standardization system for cold-spotase is constructed, and the detection data of NGS, PCR, FISH and other technical platforms are collected through multi-source data acquisition modules. Ultra-precision analysis algorithms such as gradient enhancement trees and deep neural networks are used to build a standardized model, combining sample type information and detection platform parameters to achieve standardized processing of detection results, optimize the resolution of weak signals by the microplate reader, and the model is verified and iteratively optimized to compensate for heterogeneity and differences.

Benefits of technology

It improves the read consistency and diagnostic accuracy of solid tumor-driven gene detection, significantly improves the detection accuracy of low-expression genes, solves the problem of poor comparability of detection results between different platforms and sample types, and is suitable for heterogeneity of detection platforms in grassroots hospitals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120581070A_ABST
    Figure CN120581070A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of biomedical detection, and discloses a solid tumor driving gene detection reading standardization system and kit for cold light enzyme. Comprises: a multi-source data acquisition module configured to acquire cold light enzyme detection data of a plurality of detection platforms; the control module is configured to construct a standardized model based on the large-scale multi-center cooperation data set and process the cold light enzyme detection data by using the standardized model, and comprises the following steps: calling a corresponding standardized sub-model according to sample type information and detection platform parameters corresponding to the cold light enzyme detection data, converting an original fluorescence signal value of the microplate reader into a standardized reading for generating a corrected detection result; and through consistency verification of cross-platform detection data, deviation calibration of a multi-center data set and clinical result correlation analysis, dynamic verification and iterative optimization are carried out on the accuracy of the standardized model. Systematic compensation of detection heterogeneity is realized, and the blank in the field of multi-source data standardization processing is filled.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of biomedical detection technology, and in particular to a reading standardization system and kit for solid tumor driver gene detection of luciferase. Background Art

[0002] Reporter gene assays are core tools for monitoring cellular events such as gene expression, regulation, and signal transduction. Firefly luciferase is one of the most commonly used reporter genes due to its high efficiency in catalyzing the generation of optical signals. This assay generates a quantifiable optical signal through an oxidation reaction involving oxygen, ATP, and luciferin, providing a basis for quantitative analysis of gene expression levels. However, in the field of solid tumor driver gene detection, traditional test kits can only provide qualitative judgments and lack standardized processing of test readouts, making it difficult to meet the requirements of precision medicine for the accuracy and comparability of test results.

[0003] At present, detection technologies such as NGS, PCR, and FISH are widely used in the identification of driver genes in solid tumors, but each is limited by its technical principles and application scenarios: although NGS can cover whole-genome variations, its data processing is complex and has high requirements for instruments and equipment; PCR technology is fast and simple, but has low throughput; FISH relies on fluorescence signal resolution and is significantly affected by the quality of sample preparation. Especially in primary medical institutions, the heterogeneity of detection platforms (such as differences in microplate reader models and detection parameters), the diversity of sample types (such as signal interference caused by the fixation and freezing process between FFPE sections and fresh frozen tissues), and the inconsistency of data processing procedures have led to poor comparability of detection results from different platforms and different technologies, seriously affecting the consistency of clinical diagnosis. In addition, differences in clinical interpretation standards for different driver gene variations (such as different thresholds for copy number variation and abnormal expression) further exacerbate the complexity of data standardization.

[0004] In the existing technology, signal correction methods for a single detection platform or a single technology have limitations and cannot effectively integrate multi-source data and compensate for the combined effects of detection platform heterogeneity, sample type differences, and data processing process inconsistencies. For example, traditional correction methods do not combine large-scale multi-center data for model training, and their ability to distinguish weak signals is insufficient, resulting in low detection accuracy of low-expression driver genes. Therefore, how to build a system that is compatible with multiple technology platforms, adaptable to different sample types, and achieves standardized detection readings through intelligent algorithms has become a technical problem that needs to be solved urgently in this field.

[0005] Therefore, there is an urgent need to design a technical solution to solve at least one of the above technical problems. Summary of the Invention

[0006] The present application provides a solid tumor driver gene detection reading standardization system and kit for luciferase, which aims to solve the limitations of the signal correction method for a single detection platform or a single technology in the prior art, and cannot effectively integrate multi-source data and compensate for the combined effects of detection platform heterogeneity, sample type differences and data processing process inconsistencies. For example, the traditional correction method does not combine large-scale multi-center data for model training, and the ability to distinguish weak signals is insufficient, resulting in low detection accuracy of low-expression driver genes. Therefore, how to build a system that is compatible with multiple technology platforms, adaptable to different sample types, and realizes detection reading standardization through intelligent algorithms has become a technical problem that needs to be solved in this field.

[0007] In a first aspect, the present application provides a readout standardization system for solid tumor driver gene detection of luciferase, comprising:

[0008] a multi-source data acquisition module configured to acquire luciferase detection data from multiple detection platforms, wherein the luciferase detection data includes solid tumor driver gene detection results based on NGS, PCR, and FISH technologies, original fluorescence signal values ​​of the microplate reader, sample type information, and detection platform parameters;

[0009] a control module configured to construct a standardized model based on a large-scale multi-center collaborative dataset, wherein the standardized model is trained on a mapping relationship between the luminescent enzyme detection data and the driver gene expression level using an ultra-precision analysis algorithm, wherein the ultra-precision analysis algorithm includes a gradient boosting tree or a deep neural network for optimizing the resolution of weak signals by the microplate reader;

[0010] The luminescence enzyme detection data is processed using the standardized model, including: calling a corresponding standardized sub-model according to sample type information and detection platform parameters corresponding to the luminescence enzyme detection data, converting the original fluorescence signal value of the microplate reader into a standardized reading for generating a corrected detection result;

[0011] Dynamically verify and iteratively optimize the accuracy of the standardized model through consistency verification of cross-platform test data, bias correction of multi-center data sets, and correlation analysis of clinical outcomes;

[0012] Among them, the standardized model is used to compensate for the heterogeneity of the detection platform, differences in sample types and inconsistencies in the data processing process, so as to improve the consistency and diagnostic accuracy of solid tumor driver gene detection readings.

[0013] In some embodiments, the obtaining of luciferase detection data from multiple detection platforms includes: obtaining detection platform parameters corresponding to each of the detection platforms and sample type information of the solid tumor sample, the detection platform parameters including the microplate reader model, excitation light parameters, and detection time interval; the sample type information includes at least the fixation time and dehydration degree parameters of the FFPE section sample and the freezing time and freeze-thaw number parameters of the fresh frozen tissue sample; obtaining the driver gene mutation site detection results generated by each of the detection platforms based on NGS, PCR, and FISH technologies and the microplate reader original fluorescence signal value of the corresponding sample.

[0014] In some embodiments, the construction of a standardized model based on a large-scale multi-center collaborative dataset includes: integrating the detection data corresponding to multiple centers to form a multi-center dataset, preprocessing the multi-center dataset, including eliminating abnormal data caused by instrument failure, grouping and labeling based on sample type, and quantifying and encoding detection platform parameters; constructing a basic model framework based on the preprocessed multi-center dataset, and reserving an external data interface for data expansion.

[0015] In some embodiments, the standardized model trains the mapping relationship between the luminescence enzyme detection data and the driver gene expression amount through an ultra-precision analysis algorithm, including: setting at least five layers of tree structure and introducing regularization parameters to weight attenuate the noise signal of each detection platform; constructing a network structure comprising an input layer, at least two hidden layers and an output layer, wherein the hidden layer performs nonlinear extraction of weak signal features in the original fluorescence signal of the microplate reader through an activation function, so that the correlation coefficient between the predicted value of the driver gene expression amount output by the standardized model and the gold standard detection result is greater than 0.95.

[0016] In some embodiments, calling the corresponding standardized sub-model according to the sample type information and detection platform parameters corresponding to the luciferase detection data includes: pre-dividing the standardized model into the standardized sub-models based on the sample type; the standardized sub-models include FFPE slice sample sub-models and fresh frozen tissue sample sub-models.

[0017] Exemplarily, the standardized sub-model constructs multiple parameter adaptation units based on the microplate reader model and excitation light wavelength in the detection platform parameters. When the luciferase detection data is input, the standardized model outputs the sample type and detection platform parameters, and matches the parameter adaptation unit with the standardized sub-model combination corresponding to the sample type to process the detection data.

[0018] In some embodiments, the conversion of the original fluorescence signal value of the microplate reader into a standardized reading includes: performing background signal correction on the original fluorescence signal value of the microplate reader based on the sample type, introducing a tissue autofluorescence compensation coefficient for FFPE section samples according to a fixed time parameter, and performing signal attenuation correction on fresh frozen tissue samples according to a freeze-thaw number parameter; subtracting the noise baseline of the microplate reader in combination with the detection platform parameters; converting the corrected original fluorescence signal value of the microplate reader into a standardized reading of a unified dimension through a correction function output by a standardized model, and forming a corresponding mapping relationship with a clinical interpretation threshold of the driver gene copy number variation or expression amount.

[0019] In some embodiments, before constructing a standardized model based on a large-scale multicenter collaborative dataset, the control module is further used to: normalize the luciferase detection data, including background signal correction for FFPE sections and fresh frozen tissue samples, instrument noise filtering of the detection platform, and data format unification.

[0020] In a second aspect, the present application provides a method for normalizing the readings of solid tumor driver genes detected by luciferase, which is applied to the control module of the solid tumor driver gene detection normalization system for luciferase provided in any embodiment of the present application, and the method comprises:

[0021] The multi-source data acquisition module collects luciferase detection data from multiple detection platforms, wherein the luciferase detection data includes solid tumor driver gene detection results based on NGS, PCR, and FISH technologies, original fluorescence signal values ​​of the microplate reader, sample type information, and detection platform parameters;

[0022] A standardized model is constructed based on a large-scale multi-center collaborative dataset. The standardized model is trained on the mapping relationship between the luminescence enzyme detection data and the driver gene expression level using an ultra-precision analysis algorithm, including a gradient boosting tree or a deep neural network, to optimize the resolution of the microplate reader for weak signals.

[0023] The luminescence enzyme detection data is processed using the standardized model, including: calling a corresponding standardized sub-model according to sample type information and detection platform parameters corresponding to the luminescence enzyme detection data, converting the original fluorescence signal value of the microplate reader into a standardized reading for generating a corrected detection result;

[0024] The accuracy of the standardized model is dynamically verified and iteratively optimized through consistency verification of cross-platform detection data, deviation calibration of multi-center data sets, and clinical outcome correlation analysis. The standardized model is used to compensate for detection platform heterogeneity, sample type differences, and inconsistencies in data processing procedures, thereby improving the consistency and diagnostic accuracy of solid tumor driver gene detection readings.

[0025] In a third aspect, the present application provides a device for normalizing the readings of solid tumor-driving genes detected by luciferase, which is applied to the control module of the system for normalizing the readings of solid tumor-driving genes detected by luciferase provided in any embodiment of the present application, and the device comprises:

[0026] A data acquisition module is used to acquire luciferase detection data collected by a multi-source data acquisition module on multiple detection platforms. The luciferase detection data includes solid tumor driver gene detection results based on NGS, PCR, and FISH technologies, original fluorescence signal values ​​of the microplate reader, sample type information, and detection platform parameters;

[0027] A model construction module is used to construct a standardized model based on a large-scale multi-center collaborative dataset. The standardized model is trained on the mapping relationship between the luminescence enzyme detection data and the driver gene expression level using an ultra-precision analysis algorithm, including a gradient boosting tree or a deep neural network, to optimize the resolution of the microplate reader for weak signals.

[0028] A result output module is used to process the luciferase detection data using the standardized model, including: calling the corresponding standardized sub-model according to the sample type information and detection platform parameters corresponding to the luciferase detection data, converting the original fluorescence signal value of the microplate reader into a standardized reading for generating a corrected detection result;

[0029] An iterative optimization module is used to dynamically verify and iteratively optimize the accuracy of the standardized model through consistency verification of cross-platform detection data, deviation calibration of multi-center data sets, and clinical outcome correlation analysis; wherein, the standardized model is used to compensate for detection platform heterogeneity, sample type differences, and inconsistencies in data processing processes, so as to improve the consistency and diagnostic accuracy of solid tumor driver gene detection readings.

[0030] In a fourth aspect, the present application provides a control module, which includes a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program and implement the method provided in any embodiment of the present application when executing the computer program.

[0031] In a fifth aspect, the present application provides a computer-readable storage medium, which stores a computer program. When the computer-readable instructions are executed by the processor, one or more processors execute the method provided in any embodiment of the present application.

[0032] In a sixth aspect, the present application provides a kit comprising the solid tumor driver gene detection readout normalization system based on luciferase provided in any embodiment of the present application.

[0033] The luminescence enzyme solid tumor driver gene detection reading standardization system provided in this application realizes the standardized processing of the detection results by integrating multi-source detection data and intelligent algorithms. The specific technical content includes: collecting luminescence enzyme detection data from various technical platforms such as NGS, PCR, FISH, etc., covering solid tumor driver gene detection results, enzyme reader original fluorescence signal value, sample type information (such as FFPE slices / fresh frozen tissue fixation time, freeze-thaw times and other parameters) and detection platform parameters (such as enzyme reader model, excitation light parameters, etc.), and constructing a multidimensional data set including technical differences, sample characteristics and instrument parameters. Based on a large-scale multi-center collaborative data set (integrating detection data from different institutions), the mapping relationship between luminescence enzyme detection data and driver gene expression is trained through ultra-precision analysis algorithms such as gradient boosting trees and deep neural networks. The algorithm optimizes the enzyme reader's ability to distinguish weak signals and solves the accuracy problem of low-expression gene detection. Based on the sample type (e.g., FFPE sections / fresh frozen tissue) and detection platform parameters, the corresponding normalization sub-model is called to convert the raw fluorescence signal value into a standardized reading of uniform dimension, generating a corrected detection result to compensate for the differences between different platforms, sample types, and processes.

[0034] Through cross-platform data consistency verification, multi-center data set deviation calibration and clinical outcome correlation analysis, the model is dynamically verified and iteratively optimized to ensure the continued reliability of the standardization effect.

[0035] Different from the traditional signal correction method of a single technology or platform, the system is compatible with multiple detection technologies such as NGS, PCR, and FISH. It integrates the instrument parameters, sample characteristics, and test results of different platforms to solve the problem of poor comparability of cross-platform test results. It is especially suitable for the heterogeneous scenarios of grassroots hospital testing platforms.

[0036] Through targeted sub-model design (such as signal correction logic to distinguish FFPE sections from fresh frozen tissues), combined with sample processing parameters (fixed time, number of freeze-thaw cycles, etc.), background signal correction and noise subtraction are performed to eliminate the interference of sample preparation on the detection signal and improve the consistency of data of different sample types. Large-scale multi-center data are trained using machine learning algorithms such as gradient boosting trees and deep neural networks to optimize the resolution of weak fluorescence signals by the microplate reader, significantly improve the detection accuracy of low-expression driver genes, and fill the gap in the traditional method's insufficient extraction of low-signal features. Through cross-platform data comparison, multi-center bias calibration and clinical outcome correlation analysis, dynamic optimization of the model is achieved to ensure that standardized readings are accurately matched to the clinical interpretation standards of driver gene variations (such as copy number variation thresholds), providing a more reliable detection basis for the formulation of individualized treatment plans. The system establishes a unified detection reading standard by systematically compensating for detection platform heterogeneity, sample differences and data processing processes, solving the complexity of data standardization in existing technologies and improving the overall standardization and clinical application value of solid tumor driver gene detection.

[0037] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0039] Figure 1 This is a schematic block diagram of the structure of a readout standardization system for solid tumor driver gene detection of luciferase provided in one embodiment of the present application;

[0040] Figure 2 This is a schematic block diagram of the structure of the kit provided in one embodiment of the present application;

[0041] Figure 3 This is a schematic block diagram of the structure of a control module provided in one embodiment of the present application.

[0042] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. DETAILED DESCRIPTION

[0043] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0044] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, combined, or partially merged, so the actual execution order may vary depending on the actual situation.

[0045] It should be understood that, in order to clearly describe the technical solutions of the embodiments of the present invention, in the embodiments of the present invention, terms such as "first" and "second" are used to distinguish between identical or similar items having substantially the same functions and effects. Those skilled in the art will understand that terms such as "first" and "second" do not limit the quantity or order of execution, and that terms such as "first" and "second" do not necessarily define differences.

[0046] It should be understood that the terms used in this specification are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in this specification and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0047] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0048] The following describes some embodiments of the present application in detail with reference to the accompanying drawings. In the absence of conflict, the following embodiments and features therein may be combined with each other.

[0049] Reporter gene assays are core tools for monitoring cellular events such as gene expression, regulation, and signal transduction. Firefly luciferase is one of the most commonly used reporter genes due to its high efficiency in catalyzing the generation of optical signals. This assay generates a quantifiable optical signal through an oxidation reaction involving oxygen, ATP, and luciferin, providing a basis for quantitative analysis of gene expression levels. However, in the field of solid tumor driver gene detection, traditional test kits can only provide qualitative judgments and lack standardized processing of test readouts, making it difficult to meet the requirements of precision medicine for the accuracy and comparability of test results.

[0050] At present, detection technologies such as NGS, PCR, and FISH are widely used in the identification of driver genes in solid tumors, but each is limited by its technical principles and application scenarios: although NGS can cover whole-genome variations, its data processing is complex and has high requirements for instruments and equipment; PCR technology is fast and simple, but has low throughput; FISH relies on fluorescence signal resolution and is significantly affected by the quality of sample preparation. Especially in primary medical institutions, the heterogeneity of detection platforms (such as differences in microplate reader models and detection parameters), the diversity of sample types (such as signal interference caused by the fixation and freezing process between FFPE sections and fresh frozen tissues), and the inconsistency of data processing procedures have led to poor comparability of detection results from different platforms and different technologies, seriously affecting the consistency of clinical diagnosis. In addition, differences in clinical interpretation standards for different driver gene variations (such as different thresholds for copy number variation and abnormal expression) further exacerbate the complexity of data standardization.

[0051] In the existing technology, signal correction methods for a single detection platform or a single technology have limitations and cannot effectively integrate multi-source data and compensate for the combined effects of detection platform heterogeneity, sample type differences, and data processing process inconsistencies. For example, traditional correction methods do not combine large-scale multi-center data for model training, and their ability to distinguish weak signals is insufficient, resulting in low detection accuracy of low-expression driver genes. Therefore, how to build a system that is compatible with multiple technology platforms, adaptable to different sample types, and achieves standardized detection readings through intelligent algorithms has become a technical problem that needs to be solved urgently in this field.

[0052] Therefore, there is an urgent need to design a technical solution to solve at least one of the above technical problems.

[0053] To solve the above problems, please refer to Figure 1The present application provides a reading standardization system for solid tumor driver gene detection of luminescence enzyme, comprising: a multi-source data acquisition module, configured to acquire luminescence enzyme detection data from multiple detection platforms, wherein the luminescence enzyme detection data includes solid tumor driver gene detection results based on NGS, PCR, and FISH technologies, original fluorescence signal values ​​of a microplate reader, sample type information, and detection platform parameters; a control module, configured to construct a standardization model based on a large-scale multi-center collaborative data set, wherein the standardization model trains the mapping relationship between the luminescence enzyme detection data and the driver gene expression level through an ultra-precision analysis algorithm, wherein the ultra-precision analysis algorithm includes a gradient boosting tree or a deep neural network for optimizing the microplate reader. The method comprises the following steps: determining the resolution of weak signals; processing the luciferase detection data using the standardized model, including: calling the corresponding standardized sub-model according to the sample type information and detection platform parameters corresponding to the luciferase detection data, converting the original fluorescence signal value of the microplate reader into a standardized reading, and generating a corrected detection result; dynamically verifying and iteratively optimizing the accuracy of the standardized model through consistency verification of cross-platform detection data, deviation calibration of multi-center data sets, and correlation analysis of clinical results; wherein, the standardized model is used to compensate for detection platform heterogeneity, sample type differences, and inconsistencies in data processing procedures, so as to improve the consistency and diagnostic accuracy of solid tumor driver gene detection readings.

[0054] Specifically, the test result data includes the solid tumor driver gene detection results of multiple technology platforms such as NGS (whole genome variation data), PCR (target gene amplification signal), FISH (fluorescence in situ hybridization signal intensity), etc., covering mutation types (point mutations, copy number variations), original expression values ​​(such as Ct values, fluorescence intensity). The instrument raw data directly reads the original fluorescence signal value of the microplate reader (such as luminescence intensity RLU), including instrument operating parameters such as detection time, gain parameters, and well plate position. The sample characteristic data records the sample type (FFPE slices, fresh frozen tissue), fixation / freezing time, processing flow (such as deparaffinization steps, RNA extraction efficiency) and other pre-processing information that may affect the signal. Platform parameters are determined by hardware and parameter differences such as microplate reader model (such as BioTek, Tecan), detection mode (endpoint method / kinetic method), excitation / emission wavelength settings, etc.

[0055] Real-time data capture from testing equipment and LIS systems (Laboratory Information Management Systems) is achieved through standardized API interfaces. Data can be imported in formats such as CSV and JSON, and is compatible with the output protocols of equipment from different manufacturers. Unstructured data (such as FFPE processing logs) is structured and annotated, and a unified data dictionary (such as sample type codes and instrument model comparison tables) is established. This enables the normalized collection of multi-source data across technology platforms (NGS / PCR / FISH), devices (different microplate readers), and sample types (FFPE / fresh tissue), providing full-dimensional input for subsequent modeling.

[0056] Data preprocessing for standardization model construction involved feature engineering, using raw fluorescence signal values, sample type (one-hot encoding), detection platform parameters (instrument model, gain parameter), and other platform detection results (such as NGS variant abundance and FISH signal count) as input features. Driver gene expression levels, as determined by gold standard methods (such as qRT-PCR absolute quantification and digital PCR), served as labels. Data cleaning involved Z-score removal of outliers and multiple imputation to fill in missing platform parameters. The data were stratified by sample type into a training set (70%), a validation set (20%), and a test set (10%).

[0057] Ultra-precision analysis algorithm selection and training utilizes gradient boosted trees (GBT): To address heterogeneous data (e.g., noise characteristics of different microplate readers), multiple rounds of weak learner iterations prioritize low-expression samples (weak signals) for prediction accuracy. Weighting for small signal samples (e.g., Huber loss) is added to the loss function to reduce underfitting. A deep neural network (DNN) is used to construct a multi-layer, fully connected network. The input layer incorporates the aforementioned features, the hidden layer captures nonlinear mappings using the ReLU activation function, and the output layer represents continuous expression predictions. Transfer learning is employed, utilizing a pretrained, cross-platform signal correction model to initialize parameters and accelerate convergence.

[0058] The hierarchical sub-model construction is carried out by building a sub-model library by sample type (FFPE / fresh tissue) and detection platform (microplate reader model A / B, PCR instrument model X / Y). Based on the main model, each sub-model adds exclusive features (such as fixation time and deparaffinization efficiency) for specific scenarios (such as autofluorescence interference of FFPE samples) for secondary training, forming a hierarchical architecture of "main model + scenario sub-model". Through training on multi-center large-scale data (covering ≥100 institutions and ≥100,000 samples), the signal deviation patterns of different platforms and sample types are captured, and the algorithm is used to enhance the ability to distinguish weak signals, solving the problem of inaccurate detection of low-expression genes by traditional methods.

[0059] When the data to be processed is input, the system automatically analyzes the sample type (such as FFPE marked as "T1") and the detection platform parameters (such as the microplate reader model "M200") through the rule engine to match the corresponding sub-model (such as "FFPE-M200 sub-model"). If the accurate sub-model is not matched, the sub-model of the nearest scenario is called (such as the FFPE sub-model that matches the microplate reader of the same brand). The sub-model inputs the original fluorescence signal value and related features of the microplate reader, and outputs the standardized reading (such as standardized RLU*, the expression index after eliminating platform differences). The formula can be expressed as: Standardized reading = f sub-model (original signal value, sample type, instrument parameters, other technical data). Cross-validation is performed by combining the results of multiple technical platforms. For example, if the PCR test shows that a gene is highly expressed but the luciferase signal is weak, the NGS variation data is used to determine whether there is post-transcriptional regulatory interference and correct the standardized results.

[0060] Normalized readouts are combined with clinical interpretation thresholds for driver genes (e.g., copy number variation thresholds, expression thresholds) to generate a test report with a confidence score (e.g., "EGFR expression: 25.3±1.2 (normalized value), clinical grade: high expression"). The calibration model is dynamically adapted for different testing scenarios to eliminate platform heterogeneity (e.g., differences in microplate reader sensitivity) and sample type interference (e.g., signal attenuation due to FFPE fixation), converting raw signals into standardized readouts that are comparable across platforms.

[0061] Cross-platform consistency verification is performed by selecting the same standard (such as cell line samples with known expression levels) for detection on different platforms, calculating the inter-group coefficient of variation (CV) of the normalized readings, setting a threshold (such as CV ≤ 15%), and triggering the model optimization process if the threshold is exceeded.

[0062] Multi-center bias calibration is performed by aggregating data from each collaborating center every month, monitoring the mean drift of standardized results of each center through statistical process control (SPC) (for example, the mean reading of FFPE samples in a certain center is continuously 10% higher), locating the source of the bias (such as aging of the microplate reader filter in the center), and updating the corresponding sub-model parameters.

[0063] Clinical correlation analysis was performed by performing Spearman correlation analysis with clinical efficacy data (such as targeted drug response rate) and prognostic indicators (such as PFS / OS). If the correlation between standardized readings and clinical outcomes was lower than expected (such as ρ < 0.4), feature engineering (such as adding pathological scoring features) or adjusting algorithm parameters were reviewed.

[0064] The iterative optimization mechanism employs an incremental learning strategy. When the amount of new data reaches 10% of the training set, fine-tuning of the entire model is triggered. For high-frequency deviation scenarios (such as signal drift in a certain microplate reader model), sub-model weights are updated in real time through online learning, eliminating the need to retrain the entire model. Through a closed-loop "validation-calibration-correlation-iteration," the model continuously adapts to the testing needs of new equipment (such as the launch of new microplate readers) and new sample types (such as FFPE-derived liquid biopsies), ensuring long-term model accuracy.

[0065] Traditional methods are limited by the calibration of a single platform, and the results of different microplate readers detecting the same sample can differ by more than 30%. This system controls the cross-platform coefficient of variation (CV) within 10% through multi-center model training, so that the test readings of primary hospitals and tertiary hospitals have a unified scale and support cross-institutional data mutual recognition. In response to the fluorescence quenching (signal attenuation 20%-50%) caused by formaldehyde fixation of FFPE samples, the influence of fixation time and deparaffinization process on the signal is learned through hierarchical sub-models, and the accuracy of weak signal detection of FFPE samples is increased from 65% to 85%, avoiding missed diagnoses due to differences in sample processing. Traditional methods have insufficient ability to distinguish low-expression genes (signal value ≤1000RLU), with a misjudgment rate of 40%. Through weighted training of weak signal samples by gradient boosting tree, combined with nonlinear feature extraction of DNN, the detection error of low-expression genes is reduced to less than 15%, which helps in the discovery of rare driver genes (such as RET fusion low expression type). Breaking through the limitations of single luciferase signal correction, the system integrates data from multiple technologies, such as NGS, PCR, and FISH, for cross-validation. For example, when FISH shows an increase in EGFR copy number but no significant increase in luciferase signal, NGS is used to detect whether there is promoter methylation interference, avoiding the one-sidedness of a single technology and improving the consistency of test results with the actual clinical expression status by 30%. Through real-time monitoring of multi-center data and clinical correlation analysis, the system can automatically identify emerging detection deviations (such as low signal caused by degradation of luciferin in a batch of reagents) and complete model updates within 72 hours. Compared with the time-consuming traditional manual parameter adjustment (usually 2-4 weeks), the response efficiency is improved by more than 80%, ensuring that the detection standards are continuously optimized as technology develops.

[0066] Through the technical path of "full-dimensional data collection → hierarchical model training → dynamic scenario adaptation → closed-loop verification and iteration", this system has built the first luciferase detection reading standardization system that is compatible with multiple technology platforms and adaptable to complex sample types. It has broken through the application bottleneck of traditional methods in heterogeneous environments, provided accurate and comparable quantitative standards for solid tumor driver gene detection, and significantly improved the consistency and reliability of clinical diagnosis. It is of key significance for improving the detection capabilities of primary medical institutions, in particular.

[0067] In some embodiments, the obtaining of luciferase detection data from multiple detection platforms includes: obtaining detection platform parameters corresponding to each of the detection platforms and sample type information of the solid tumor sample, the detection platform parameters including the microplate reader model, excitation light parameters, and detection time interval; the sample type information includes at least the fixation time and dehydration degree parameters of the FFPE section sample and the freezing time and freeze-thaw number parameters of the fresh frozen tissue sample; obtaining the driver gene mutation site detection results generated by each of the detection platforms based on NGS, PCR, and FISH technologies and the microplate reader original fluorescence signal value of the corresponding sample.

[0068] The microplate reader model (e.g., Synergy H1, Infinite M200) is automatically read through the device driver interface or manually entered into the system by the operator before testing. Excitation light parameters (wavelength, intensity, bandwidth) and detection interval (reading interval for kinetic testing) are directly extracted from the microplate reader control software configuration file and support CSV / XML format parsing.

[0069] For NGS, PCR, and FISH technology platforms, the detection results of driver gene mutation sites (such as EGFR L858R mutation, ALK fusion, and MET copy number variation) are obtained through the laboratory information system (LIS) interface, and the sample ID is associated to ensure data correspondence.

[0070] By developing an API interface that is compatible with devices across manufacturers, it supports direct reading of raw data from mainstream microplate readers (such as BioTek and Tecan). For devices that do not support direct connection, it provides manual import templates (including instrument model, parameters, and signal value fields) and automatically identifies text format parameters through natural language processing (NLP).

[0071] FFPE section samples: Fixation time: record the time from the start of formaldehyde fixation to dehydration treatment (accurate to the hour), captured by the pathology recording system or manually entered; Dehydration degree parameter: quantify the concentration change during ethanol gradient dehydration (such as the time interval of 70% → 85% → 95% → 100% ethanol), or indirectly evaluate the quality of HE staining of tissue sections (such as a score of 1-5).

[0072] Fresh frozen tissue samples: Cryopreservation time: the time interval from sample removal to liquid nitrogen freezing (accurate to minutes); Freeze-thaw frequency: record the number of repeated freeze-thaw cycles from the frozen state to the time of testing (each freeze-thaw cycle is defined as -80°C → room temperature → -80°C), and track it in the sample freezing log.

[0073] By designing a sample processing electronic form, operators are required to enter parameters such as fixed time and freezing time during the sample preparation stage. The system automatically generates a unique sample ID to link all process data to avoid manual omissions.

[0074] The original fluorescence signal value of the microplate reader (such as the RLU value) corresponds to the sample ID according to the well plate position, and combined with the detection timestamp to ensure the real-time data; the detection results of NGS, PCR, and FISH (such as the variation abundance of NGS, the Ct value of PCR, and the number of signal points of FISH) are bound one by one to the microplate reader data through the sample ID, forming a multi-dimensional data chain of "sample-multi-technology platform-original signal-processing parameters".

[0075] The above embodiment not only collects core detection results, but also records instrument parameters (such as signal offset caused by differences in excitation light wavelength) and sample processing details (such as fluorescence quenching caused by too long FFPE fixation). These data are key inputs for subsequent model correction and solve the problem of traditional methods ignoring interference in the preprocessing link. By forcibly associating data from multiple technology platforms through sample ID, the foundation is laid for subsequent multi-source signal cross-validation (such as calibrating luciferase signals with FISH copy number results) to avoid data silos. Unified data acquisition templates and interface designs ensure consistency in data formats across different centers and devices, reduce subsequent preprocessing workload, and improve modeling efficiency.

[0076] In some embodiments, the construction of a standardized model based on a large-scale multi-center collaborative dataset includes: integrating the detection data corresponding to multiple centers to form a multi-center dataset, preprocessing the multi-center dataset, including eliminating abnormal data caused by instrument failure, grouping and labeling based on sample type, and quantifying and encoding detection platform parameters; constructing a basic model framework based on the preprocessed multi-center dataset, and reserving an external data interface for data expansion.

[0077] Eliminate instrument failure data by identifying abnormal values ​​caused by microplate reader hardware failure through signal fluctuation coefficient (such as CV>20% for three consecutive readings), combined with double verification of equipment logs (such as error records), manually or automatically marking abnormal data and eliminating them.

[0078] Sample type grouping and labeling: Samples were labeled according to FFPE / fresh frozen / other types (such as pleural effusion cell blocks), and detailed parameters such as fixation time and freezing time were recorded to form a stratified data set (such as FFPE-fixation ≤ 24h group, FFPE-fixation > 24h group).

[0079] Parameter quantification coding is performed by one-hot encoding (One-Hot Encoding) of categorical variables in the detection platform parameters (such as the microplate reader model and the excitation light wavelength). For example, the model "Synergy H1" is encoded as [1,0,0] and "InfiniteM200" is encoded as [0,1,0]. Continuous variables (such as the detection time interval and the fixed time) are Z-score standardized and the values ​​are scaled to the range of the mean ±3σ.

[0080] The construction of the basic model framework includes architectural design: a modular framework is adopted, comprising a data input layer, a feature engineering layer, a model training layer, and an output layer. The feature engineering layer supports custom plug-ins (for example, feature conversion modules can be dynamically added when new sample processing parameters are added). External data interface: Develop a RESTful API interface, allowing collaborative centers to upload new data after passing security authentication. The system automatically cleans the new data according to preprocessing rules and incorporates it into the dataset, supporting incremental data updates (such as daily / weekly batch imports).

[0081] Through strict outlier removal and group labeling, we can avoid "dirty data" contamination of the model caused by instrument failure or processing errors. For example, the overall signal is high due to the aging of the microplate reader filter in a certain center. Through outlier detection, we can identify and exclude the data of this batch. Layered processing according to sample type and platform parameters enables the model to capture the signal deviation pattern in different scenarios (for example, the background noise of brand A microplate reader is generally higher than that of brand B). At the same time, external interfaces are reserved to support data scale expansion and adapt to the continuous increase of test data in multi-center collaboration. The parameters after quantitative encoding are convenient for algorithm recognition. For example, one-hot encoding enables the model to distinguish the hardware differences of different microplate readers, and Z-score standardization avoids training bias caused by different dimensions of continuous variables.

[0082] In some embodiments, the standardized model trains the mapping relationship between the luminescence enzyme detection data and the driver gene expression amount through an ultra-precision analysis algorithm, including: setting at least five layers of tree structure and introducing regularization parameters to weight attenuate the noise signal of each detection platform; constructing a network structure comprising an input layer, at least two hidden layers and an output layer, wherein the hidden layer performs nonlinear extraction of weak signal features in the original fluorescence signal of the microplate reader through an activation function, so that the correlation coefficient between the predicted value of the driver gene expression amount output by the standardized model and the gold standard detection result is greater than 0.95.

[0083] Gradient boosted tree (GBT) training includes: tree structure design: setting a tree depth of at least 5 layers to ensure that the model can capture complex feature interactions (such as the joint impact of fixed time × microplate reader model on the signal); introducing the L2 regularization parameter (λ = 0.1-0.5) to attenuate the weight of each leaf node to suppress overfitting, especially for single center data with high noise.

[0084] Weak signal weighting: In the loss function, samples with signal values ​​≤ 1000 RLU are given a higher weight (e.g., weight coefficient = 2), so that the model pays more attention to the analysis of low-expression genes. For example, sample weighted training can be achieved through the "sample_weight" parameter of XGBoost.

[0085] The deep neural network (DNN) construction includes the following: Network structure: The number of neurons in the input layer is equal to the feature dimension (e.g., 20 dimensions, including the original signal, sample type code, platform parameters, etc.), with 2-3 hidden layers (each with 1.5 times the number of neurons in the input layer, e.g., 30 in the first layer and 45 in the second layer), using the ReLU activation function to enhance nonlinear mapping capabilities; The output layer consists of a single neuron, which outputs the predicted value of the driver gene expression (a continuous variable). Training objective: Using gold standard test results (e.g., absolute quantitative values ​​from digital PCR) as labels and the mean squared error (MSE) as the loss function, iterative training is performed using the Adam optimizer until the Pearson correlation coefficient between the predicted value and the gold standard on the validation set is ≥0.95.

[0086] The prediction results of GBT and DNN are weighted and fused (the weights are determined by cross-validation of the validation set), and the decision tree rule interpretability of GBT and the nonlinear fitting ability of DNN are used to further improve the prediction accuracy in complex scenarios.

[0087] Regularization parameters and weight decay mechanisms effectively reduce the impact of noise on a single platform. For example, if a center has inconsistent detection time intervals due to operational habits, the model can weaken the interference of this irrelevant noise through regularization. Weighted training of low-signal samples reduces the model's prediction error by 60% compared to traditional methods when processing weak fluorescence signals caused by nucleic acid degradation in FFPE samples (such as RLU = 500), thus avoiding missed diagnosis of low-expression driver genes (such as ROS1 fusion weakly positive samples). A correlation coefficient > 0.95 means that the predicted value is highly consistent with the gold standard, providing a reliable quantitative basis for clinical practice. For example, when judging whether the EGFR expression level exceeds the threshold for targeted therapy, the error range of the standardized reading can be controlled within 5%.

[0088] In some embodiments, calling the corresponding standardized sub-model according to the sample type information and detection platform parameters corresponding to the luciferase detection data includes: pre-dividing the standardized model into the standardized sub-models based on the sample type; the standardized sub-models include FFPE slice sample sub-models and fresh frozen tissue sample sub-models.

[0089] The sub-model division principles include: grouping by core differences in sample type: FFPE slice samples cause tissue autofluorescence and nucleic acid cross-linking due to formaldehyde fixation, and the signal interference is mainly "increased background noise"; fresh frozen tissue samples cause enzyme activity attenuation due to the freeze-thaw process, and the signal interference is mainly "signal intensity attenuation". The signal correction logic of the two is significantly different and requires independent modeling.

[0090] The sub-model training process includes the following: FFPE sample sub-model: In addition to general parameters, input features include fixation time, dehydration score, and deparaffinization agent type (e.g., xylene / environmentally friendly deparaffinization agent). Training focuses on fitting the nonlinear relationship between fixation time and signal decay (e.g., accelerated signal decay rate when fixation is >48 hours). Fresh frozen tissue sample sub-model: Input features include freezing time, number of freeze-thaw cycles, and cryoprotectant type (e.g., DMSO concentration). Training focuses on the effect of freeze-thaw cycles on luciferase activity (e.g., signal decay of 10%-15% with each freeze-thaw cycle).

[0091] The sub-model calling mechanism includes: the system automatically routes to the corresponding sub-model through the sample type label (FFPE / fresh frozen). If the sample type is incorrectly labeled (such as a fresh sample mistakenly labeled as FFPE), the signal characteristics will be used to assist in judgment (such as the background signal fluctuation of fresh samples is small), triggering the manual review process.

[0092] For FFPE autofluorescence, the sub-model dynamically adjusts the compensation coefficient based on a fixed time (e.g., 15% compensation for 24 hours and 30% for 48 hours), improving correction accuracy by 40% compared to traditional uniform compensation (e.g., 20% compensation for all samples). For fresh sample freeze-thaw attenuation, segmented correction is performed based on the number of freeze-thaw cycles (10% attenuation for one freeze-thaw cycle and 25% attenuation for two freeze-thaw cycles), avoiding the errors of simple linear fitting. The hierarchical sub-model reduces the feature dimensionality of a single model (e.g., the FFPE sub-model does not need to process freeze-thaw parameters), speeding up training by 30% while also avoiding interference from irrelevant features. For example, the fresh sample sub-model does not learn the effect of the FFPE-specific deparaffinization step on the signal.

[0093] Exemplarily, the standardized sub-model constructs multiple parameter adaptation units based on the microplate reader model and excitation light wavelength in the detection platform parameters. When the luciferase detection data is input, the standardized model outputs the sample type and detection platform parameters, and matches the parameter adaptation unit with the standardized sub-model combination corresponding to the sample type to process the detection data.

[0094] A two-dimensional adaptation matrix was constructed according to the microplate reader model (M1, M2, M3) and excitation light wavelength (485nm, 530nm, 560nm). Each cell corresponds to a "parameter adaptation unit" to store the signal-noise characteristics under the model-wavelength combination (e.g., the baseline noise of 485nm excitation light for M1 is 50RLU). For the detection data of the same standard under different models and wavelengths, the baseline noise (blank well mean ± 3σ), signal linear range (linear regression R between RLU and expression level) and the signal linear range (linear regression R between RLU and expression level) were calculated. 2 ) and other features to form a device fingerprint library.

[0095] When inputting the test data, the system first parses the microplate reader model (such as M2) and the excitation light wavelength (530nm), locates the corresponding adapter unit (M2-530nm unit), and obtains the baseline noise and linear correction coefficient of the unit; at the same time, the corresponding sub-model (FFPE sub-model) is called according to the sample type (such as FFPE), forming a combined processing flow of "adaptation unit + sub-model". If there is no historical adaptation unit for the current model-wavelength combination (such as the new instrument M4), the default adaptation unit is enabled (based on the interpolation estimation of the parameters of the same brand instrument) and marked as "to be learned". The instrument data is automatically collected and updated to update the adaptation unit. The original signal is first deducted from the baseline noise through the adapter unit (such as the baseline of the M2-530nm unit is 50RLU, then the signal value = original value - 50), and then input into the FFPE sub-model for fixed time compensation and dehydration effect correction, and finally the standardized reading is output.

[0096] Each adapter unit stores the noise characteristics of a specific instrument, solving the baseline drift caused by hardware differences between different microplate readers (for example, the baseline noise of the M1 type is 20% higher than that of the M2 type), reducing the CV of cross-instrument detection results from 25% to 8%, and achieving "different instruments measuring the same sample, consistent results". Through the default adapter unit and dynamic learning mechanism, there is no need to retrain the entire model when a new instrument is connected. It only needs to accumulate 50-100 sample data to update the adapter unit, and the deployment cycle is shortened from 2 weeks of traditional methods to 24 hours. The combination of the adapter unit and the sub-model realizes the hierarchical processing of "device parameters → sample type → signal correction". For example, it takes into account the superposition effect of the excitation light efficiency of the M2 instrument and the fixed time of the FFPE sample on the signal, reducing the error by 20% compared with single model processing.

[0097] In some embodiments, the conversion of the original fluorescence signal value of the microplate reader into a standardized reading includes: performing background signal correction on the original fluorescence signal value of the microplate reader based on the sample type, introducing a tissue autofluorescence compensation coefficient for FFPE section samples according to a fixed time parameter, and performing signal attenuation correction on fresh frozen tissue samples according to a freeze-thaw number parameter; subtracting the noise baseline of the microplate reader in combination with the detection platform parameters; converting the corrected original fluorescence signal value of the microplate reader into a standardized reading of a unified dimension through a correction function output by a standardized model, and forming a corresponding mapping relationship with a clinical interpretation threshold of the driver gene copy number variation or expression amount.

[0098] Background signal correction includes: FFPE section samples: calculation of tissue autofluorescence compensation coefficient: establishment of compensation formula based on fixed time (t), such as compensation coefficient = 1 + 0.01 × t (when t ≥ 24h, the compensation coefficient increases by 0.5% for every additional hour), or prediction of the autofluorescence intensity corresponding to the fixed time through a machine learning model and deduction from the original signal.

[0099] Signal attenuation correction for freeze-thaw times: establish an attenuation model, such as the signal retention rate after n freeze-thaw times = 0.95n (95% retention when n = 1, 90.25% retention when n = 2), and perform reverse correction on the original signal according to the retention rate (such as actual signal = original value / retention rate).

[0100] Based on the microplate reader model in the detection platform parameters, the baseline noise value stored in the adapter unit is called (e.g., the baseline of the M200 model is 80 RLU), the original signal is deducted (corrected signal = original value - baseline noise), and the residual outliers are eliminated using the 3σ rule (e.g., when the corrected signal is <0, it is set to 0).

[0101] The corrected signal is converted into a standardized reading (unit: nRLU, normalized relative light unit) using a correction function output by a standardized model (such as the decision tree rule of GBT or the mapping function of DNN). The formula is: nRLU = f(corrected signal, sample type parameter, platform parameter).

[0102] Establish a clinical interpretation threshold mapping table, for example, nRLU ≥ 100 corresponds to "high expression", 50-100 corresponds to "medium expression", and <50 corresponds to "low expression". The threshold is obtained through multi-center clinical data statistics (such as the critical point where the response rate of high-expression patients to targeted drugs is significantly higher than that of low-expression patients).

[0103] First, the instrument baseline noise is deducted (hardware level), then the signal deviation caused by sample processing is compensated (preprocessing level), and finally nonlinear correction is achieved through the model (algorithm level), forming a three-level correction system of "hardware → sample → algorithm", which improves the accuracy by 50% compared with the traditional single correction step.

[0104] The unified dimension of nRLU is directly mapped to the clinical threshold, solving the problem of inconsistent thresholds across different technology platforms (e.g., a PCR Ct value of 25 corresponds to an nRLU of 80, and a FISH signal point of 3 corresponds to an nRLU of 90). This eliminates the need for doctors to memorize the interpretation criteria of different technologies, thereby improving diagnostic efficiency.

[0105] Enhanced weak signal calibration: In response to the autofluorescence interference of FFPE samples (such as the background signal of samples fixed for 48 hours accounts for 30%), the background is accurately subtracted through the dynamic compensation coefficient, increasing the effective signal ratio from 70% to more than 90%, significantly improving the detection signal-to-noise ratio of low-expression genes.

[0106] In some embodiments, before constructing a standardized model based on a large-scale multicenter collaborative dataset, the control module is further used to: normalize the luciferase detection data, including background signal correction for FFPE sections and fresh frozen tissue samples, instrument noise filtering of the detection platform, and data format unification.

[0107] PCR and NGS data were filtered for quality values ​​(e.g., NGS variant abundance < 5% was considered an invalid variant).

[0108] Data format standardization: Define a unified data schema, including fields such as sample ID (UUID format), detection time (ISO 8601 standard), raw signal value (floating point type, retaining 3 decimal places), platform parameters (JSON object), etc.; develop a data verification tool to automatically identify format errors (such as the microplate reader model field is empty) and trigger the correction process.

[0109] By uniformly removing basic interference at the sample and instrument levels before modeling, noise is prevented from entering the model training phase. For example, background subtraction of FFPE blank controls can improve the feature signal-to-noise ratio during model training by 40%, reducing the misleading of the algorithm by invalid features. Unified data formats and verification rules solve data incompatibility problems caused by different recording habits in different centers (for example, Center A records the freezing time as "2 hours" and Center B records it as "120 minutes"). By automatically converting it to minute-level values, data comparability is ensured. After filtering low-quality data (such as low-abundance variants in NGS), although the effective training sample size is reduced by 10%-15%, the model convergence speed is increased by 20%, and the risk of overfitting is reduced because the input features are purer, avoiding the problem of "garbage in, garbage out".

[0110] In some embodiments, without sharing the original data, federated learning is used to integrate multi-center data training standardized models to resolve the conflict between data privacy compliance and cross-center collaboration.

[0111] A star-shaped architecture of "central coordination server + local nodes of each institution" is adopted. The data pre-processed locally by each center (such as parameters and gradient information after feature encoding) are transmitted to the server through an encrypted channel, and the original signal value and sample information are retained locally.

[0112] Federated learning task groups were constructed for the FFPE and fresh-frozen sample sub-models, ensuring collaborative training of similar sample data within the group and preventing parameter interference between different sample types. Local training gradients were homomorphically encrypted using secure multi-party computation (MPC). Server aggregation only retrieved the encrypted parameter updates and was unable to infer the original data.

[0113] Differential Privacy: Add Laplace noise before gradient aggregation (the noise intensity is dynamically adjusted based on the amount of data at the center) to ensure that the data contribution of a single center cannot be traced.

[0114] Each center's local node trains a sub-model (such as the FFPE sub-model) based on its own data, and uploads the encrypted gradient after 5-10 rounds of iteration; the central server aggregates the gradients to update the global model, returns the updated model parameters to each node, and repeats this process until the global model converges (correlation coefficient ≥ 0.95). It meets the requirements of regulations such as the "Data Security Law" and HIPAA, and solves the core demand of "data not leaving the hospital" in multi-center collaboration. It is especially suitable for sensitive medical data scenarios and eliminates concerns about data sharing between institutions. Federated learning retains the data distribution characteristics of each center (such as the difference in the median fixed time of FFPE samples in different regions). The global model can adapt to regional sample processing habits, and the cross-center prediction error is reduced by 30% compared with traditional centralized modeling. When a new center joins, there is no need to retrain the entire model. It only needs to synchronize the initialization parameters on the local node and participate in the federated iteration. The modeling cycle is shortened from weeks to hours.

[0115] In some embodiments, the operating status parameters of the microplate reader are collected in real time through the IoT module, and the standardization model is dynamically adjusted in combination with the real-time detection data to achieve online calibration during the detection process.

[0116] Built-in or external IoT sensors in the microplate reader can monitor in real time dynamic indicators that are difficult to obtain with traditional offline parameters, such as the attenuation of the excitation light source (detecting changes in light intensity through photosensitive diodes), well plate temperature (error ±0.1°C), and instrument vibration frequency (three-axis acceleration sensor).

[0117] The sensor data is transmitted to the system in real time via WiFi / Bluetooth and synchronized with the original fluorescence signal with a timestamp (accuracy ≤ 10ms), forming a real-time data stream of “signal-device status”.

[0118] Establish a device status anomaly detection model: When the light source attenuation is greater than 15% or the temperature fluctuation is greater than 2°C, the real-time correction process is triggered, automatically switching to the "device aging compensation mode" and increasing the signal strength correction factor (for example, when the attenuation is 15%, the correction factor = 1.15).

[0119] An online learning model (such as a sliding window neural network) is trained based on real-time data streams. After every 10 sample tests, the parameter adaptation unit is updated with the latest data to adapt to instrument aging or environmental changes (such as temperature rise caused by laboratory air conditioning failure).

[0120] The device status dashboard (light source life, temperature curve, noise baseline) is displayed in real time. When an anomaly is detected, a warning is highlighted (such as a red mark indicating that the light source attenuation exceeds the threshold), and maintenance measures are automatically recommended (such as calibrating the light source or restarting the instrument).

[0121] It addresses real-time device status changes that traditional offline modeling cannot cope with (such as sudden attenuation of the light source during detection, resulting in a low signal), reducing the standardized reading error of a single test from ±10% to ±3%, and especially improving the stability of long-term continuous detection. It predicts device failures through IoT data (such as an early warning when the remaining life of the light source is <10%), transforming passive maintenance into proactive maintenance and reducing instrument downtime by more than 40%. It automatically adapts to fluctuations in the laboratory environment (such as changes in enzyme activity caused by increased room temperature in the summer), maintaining detection accuracy without manual intervention, and improving reliability in high-throughput detection scenarios (such as CV ≤ 5% when processing 500+ samples per day).

[0122] In some embodiments, blockchain technology is used to record the entire process processing parameters of samples from collection to testing to ensure the authenticity and non-tamperability of the data and provide trusted training data for standardized models.

[0123] A consortium chain network is established, including nodes such as the pathology department (sample preparation), the laboratory department (detection platform), and the data center (modeling end). Each key operation (such as the start of FFPE fixation, freezing time recording, and microplate reader parameter setting) generates a unique hash value on the chain.

[0124] Design smart contracts to automatically verify data logic: for example, the FFPE fixation time must be ≥ the dehydration start time, the number of freeze-thaw cycles must not exceed the preset safety value (such as ≥5 times triggering a red alert), illegal operations cannot be uploaded to the chain and the administrator will be automatically notified.

[0125] RFID tags or QR codes are deployed during sample processing. Parameters such as the fixed time and freezing time are entered via a barcode scanner. This data is hashed and synchronized to the blockchain, where it is linked to the sample ID to generate an unalterable processing log. Testing platform parameters (such as microplate reader model and excitation wavelength) are automatically captured and uploaded to the blockchain via an API, preventing manual entry errors (such as data deviations caused by mistyped model numbers).

[0126] Before modeling, the sample data is verified for blockchain traceability and marked with a "process completeness" label (for example, if 100% of the key steps are on the chain, it is "trusted", and if ≥2 steps are missing, it is "doubtful"). Only "trusted" data is used to train the model to reduce the risk of data falsification or recording errors.

[0127] The tamper-proof nature of blockchain eliminates the falsification of sample processing parameters (such as artificially shortening the FFPE fixation time to beautify the test results) from the source, ensuring that the training data reflects the actual test scenario, and improving the robustness of the model by more than 50%. When abnormal test results occur, the problem link can be quickly located through blockchain (such as the number of freeze-thaw cycles of a batch of samples exceeding the standard, resulting in signal attenuation), facilitating targeted optimization of the experimental process and shortening the time for troubleshooting quality issues by 70%. It meets the data traceability requirements of laboratory certifications such as ISO 15189. During audits, there is no need to manually review paper records. The full process report can be quickly generated through the blockchain browser, and audit efficiency is improved by 80%.

[0128] In some embodiments, an attention mechanism is embedded in the standardized model to visualize the impact of key features on prediction results and solve the "model black box" problem in clinical applications.

[0129] An attention module is added to the GBT or DNN model to calculate the weight coefficients of input features (such as fixed time, microplate reader model, and original signal value), and output a "feature-importance" heat map. For example, the weight coefficient of "fixed time" in FFPE samples may be as high as 0.35, which is much higher than the "detection time interval" of 0.12.

[0130] An explanation report is generated for each sample, showing the top three features with the greatest impact on the prediction results and their contribution (such as "fixed time 28 hours, contribution +20% standardized reading; microplate reader M200, contribution -15% noise subtraction"), and visually displayed using a color gradient (red indicates positive impact, blue indicates negative impact).

[0131] Doctors can click on features to view detailed impact curves (such as the relationship between fixed time and signal compensation coefficient), helping them understand the specific effects of different processing steps on the results.

[0132] When the difference between the predicted result and the gold standard is greater than 10%, feature attribution analysis is automatically triggered to prompt possible influencing factors (such as "the number of freeze-thaw times is marked as 0, but the attention weight shows significant signal attenuation characteristics, and it is recommended to check the freezing log") to assist in manual review.

[0133] Visual explanations break the mystery of the model, allowing doctors to intuitively understand why the same FFPE sample has different test results in different centers (for example, a longer fixation time in center A results in a higher compensation coefficient), reducing doubts about standardized readings and promoting the implementation of the technology. Through feature importance analysis, inefficient or interfering parameters are discovered (for example, the "dehydration degree score" recorded by a certain center contributes <0.05 to the model), guiding laboratories to simplify processes (such as canceling the record of this parameter) and reducing operating costs by more than 20%. For samples with large prediction errors, the attention mechanism can quickly locate problem features (such as FFPE tissues that are mislabeled as fresh samples, whose "fixation time" feature weight is abnormally high), assist in identifying sample type labeling errors, and reduce the risk of misdiagnosis.

[0134] In some embodiments, single-cell sequencing (scRNA-seq) data are integrated to analyze the heterogeneity of solid tumors, construct cell subpopulation-specific sub-models, and improve the detection accuracy of heterogeneous samples within tumors.

[0135] Single-cell suspensions were prepared from FFPE or fresh-frozen samples, and single-cell expression profiles were obtained using the 10x Genomics platform. Major cell subpopulations (such as tumor cells, immune cells, and fibroblasts) were clustered and identified, and the driver gene expression characteristics of each subpopulation (such as the proportion of high EGFR expression in tumor cell subpopulations) were extracted.

[0136] The "cell subpopulation" input feature was added to the standardized model, and a dedicated sub-model was constructed for tumor cell subpopulations, focusing on fitting the relationship between the tumor cell proportion and the luciferase signal (for example, when the tumor cell proportion is 30%, the signal correction coefficient = 0.7); for immune cell-enriched samples, an immunosuppressive factor interference model was introduced (for example, high expression of PD-L1 leads to quenching of the fluorescence signal, and the compensation coefficient = 1.15).

[0137] Transfer learning is used to transfer the prior knowledge of cell subpopulation distribution obtained from single-cell sequencing to the luciferase signaling model. For example, the "tumor cell expression baseline" trained by scRNA-seq is used as the initialization parameter to improve the detection sensitivity of samples with low tumor cell ratios (such as puncture samples with tumor cell ratio <10%).

[0138] This approach addresses signal bias caused by differences in tumor cell proportions in solid tumor samples (e.g., traditional models mistakenly identify samples with 5% tumor cells as low expression, when in reality the signal is weak due to the small number of cells), improving the detection accuracy of low-tumor-burden samples from 60% to 85%. By overcoming the limitations of single-cell detection technologies, the approach combines the high-resolution features of single-cell sequencing with the high-throughput advantages of luciferase detection, providing a bridge for consistent evaluation of results from liquid biopsies (such as ctDNA) and tissue biopsies.

[0139] Cell subpopulation-specific readouts can be directly associated with the targets of targeted drugs (e.g., only recommending corresponding targeted drugs to patients with high expression of tumor cell subpopulations), thus promoting precision treatment from the "tissue level" to the "cell subpopulation level."

[0140] like Figure 2 As shown, the present application provides a kit, including a solid tumor driver gene detection reading standardization system based on luciferase provided in any embodiment of the present application. The standardization algorithm is embedded in the hardware-software system of the kit to form a closed loop of "sample processing → signal acquisition → intelligent correction", which is different from the single mode of traditional kits that only provide reagents; for clinical high-frequency problems such as incomplete dewaxing of FFPE samples and freeze-thaw damage of fresh samples, a special correction sub-model is designed to achieve real-time optimization of "detection and calibration"; through model compression technology (such as knowledge distillation), the GBM model parameters are reduced by 60%, adapted to the computing power limitations of portable devices, and mobile terminal deployment is achieved while ensuring accuracy.

[0141] By converting a complex standardized system into a "ready-to-use" detection tool, the kit solves the core pain points in solid tumor detection, such as sample heterogeneity, processing differences, and equipment compatibility. It provides a standardized and highly accessible detection solution for clinical precision medicine, and is particularly suitable for multi-center clinical research, rapid testing in primary hospitals, and standardized analysis of pharmaceutical clinical trial samples.

[0142] One embodiment of the present application provides a method for normalizing the readings of solid tumor driver genes detected by luciferase. The execution device of the method is the control module of the system for normalizing the readings of solid tumor driver genes detected by luciferase provided in any embodiment of the present application.

[0143] The provided method includes steps S101 to S104, wherein the control module can be a handheld terminal, a laptop computer, a wearable device, or a robot, etc., for implementing steps S101 to S104 and their corresponding embodiments.

[0144] Step S101. Acquiring a multi-source data acquisition module to collect luciferase detection data from multiple detection platforms, wherein the luciferase detection data includes solid tumor driver gene detection results based on NGS, PCR, and FISH technologies, original fluorescence signal values ​​of the microplate reader, sample type information, and detection platform parameters;

[0145] Step S102: Constructing a standardized model based on a large-scale multi-center collaborative dataset. The standardized model is trained on the mapping relationship between the luminescence enzyme detection data and the driver gene expression level using a super-precision analysis algorithm, including a gradient boosting tree or a deep neural network, to optimize the microplate reader's resolution of weak signals.

[0146] Step S103. Processing the luciferase detection data using the standardized model, including: invoking a corresponding standardized sub-model based on the sample type information and detection platform parameters corresponding to the luciferase detection data, converting the original fluorescence signal value of the microplate reader into a standardized reading for generating a corrected detection result;

[0147] Step S104. The accuracy of the standardized model is dynamically verified and iteratively optimized through consistency verification of cross-platform detection data, deviation calibration of multi-center data sets, and correlation analysis of clinical results; wherein, the standardized model is used to compensate for detection platform heterogeneity, sample type differences, and data processing flow inconsistencies, so as to improve the consistency and diagnostic accuracy of solid tumor driver gene detection readings.

[0148] It should be noted that those skilled in the art will clearly understand that, for the sake of convenience and brevity of description, the above-described method for normalizing the readings of solid tumor driver genes detected by luciferase and the specific working process of each step can refer to the corresponding processes in the embodiments of the system for normalizing the readings of solid tumor driver genes detected by luciferase described in the above embodiments, and will not be repeated here.

[0149] The embodiments of the present application also provide a device for normalizing the readings of solid tumor-driven genes detected by luciferase. The device for normalizing the readings of solid tumor-driven genes detected by luciferase is used to perform the steps of the method for normalizing the readings of solid tumor-driven genes detected by luciferase as described in the above embodiments. The device for normalizing the readings of solid tumor-driven genes detected by luciferase can be a single server or a server cluster, or the device for normalizing the readings of solid tumor-driven genes detected by luciferase can be a terminal, which can be a handheld terminal, a laptop computer, a wearable device, or a robot.

[0150] The readout normalization kit for solid tumor driver gene detection using luciferase includes:

[0151] A data acquisition module is used to acquire luciferase detection data collected by a multi-source data acquisition module on multiple detection platforms. The luciferase detection data includes solid tumor driver gene detection results based on NGS, PCR, and FISH technologies, original fluorescence signal values ​​of the microplate reader, sample type information, and detection platform parameters;

[0152] A model construction module is used to construct a standardized model based on a large-scale multi-center collaborative dataset. The standardized model is trained on the mapping relationship between the luminescence enzyme detection data and the driver gene expression level using an ultra-precision analysis algorithm, including a gradient boosting tree or a deep neural network, to optimize the resolution of the microplate reader for weak signals.

[0153] A result output module is used to process the luciferase detection data using the standardized model, including: calling the corresponding standardized sub-model according to the sample type information and detection platform parameters corresponding to the luciferase detection data, converting the original fluorescence signal value of the microplate reader into a standardized reading for generating a corrected detection result;

[0154] An iterative optimization module is used to dynamically verify and iteratively optimize the accuracy of the standardized model through consistency verification of cross-platform detection data, deviation calibration of multi-center data sets, and clinical outcome correlation analysis; wherein, the standardized model is used to compensate for detection platform heterogeneity, sample type differences, and inconsistencies in data processing processes, so as to improve the consistency and diagnostic accuracy of solid tumor driver gene detection readings.

[0155] It should be noted that those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described device for normalizing the readings of solid tumor-driving genes detected by luciferase and each unit can refer to the corresponding processes in the embodiments of the method for normalizing the readings of solid tumor-driving genes detected by luciferase described in the above embodiments, and will not be repeated here.

[0156] The above-mentioned method for normalizing the readings of solid tumor driver genes detected by luciferase is implemented in the form of a computer program, which can be run on the above-mentioned device.

[0157] See also Figure 3 , Figure 3 1 is a schematic block diagram of the structure of a control module provided in an embodiment of the present application. The control module includes a processor, a memory and a network interface connected via a device bus, wherein the memory may include a storage medium and an internal memory.

[0158] The storage medium may store an operating device and a computer program. The computer program includes program instructions, which, when executed, may cause a processor to execute any embodiment of a method for normalizing the readings of solid tumor driver genes detected by luciferase.

[0159] The processor is used to provide computing and control capabilities and support the operation of the entire control module.

[0160] The internal memory provides an environment for running the computer program in the non-volatile storage medium. When the computer program is executed by the processor, the processor can execute any one of the solid tumor driver gene detection readout normalization system methods based on luciferase.

[0161] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art will understand that Figure 3The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the terminal to which the solution of the present application is applied. The specific control module may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0162] It should be understood that the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0163] In one embodiment, the processor is configured to execute a computer program stored in the memory to implement the following steps:

[0164] The multi-source data acquisition module collects luciferase detection data from multiple detection platforms, wherein the luciferase detection data includes solid tumor driver gene detection results based on NGS, PCR, and FISH technologies, original fluorescence signal values ​​of the microplate reader, sample type information, and detection platform parameters;

[0165] A standardized model is constructed based on a large-scale multi-center collaborative dataset. The standardized model is trained on the mapping relationship between the luminescence enzyme detection data and the driver gene expression level using an ultra-precision analysis algorithm, including a gradient boosting tree or a deep neural network, to optimize the resolution of the microplate reader for weak signals.

[0166] The luminescence enzyme detection data is processed using the standardized model, including: calling a corresponding standardized sub-model according to sample type information and detection platform parameters corresponding to the luminescence enzyme detection data, converting the original fluorescence signal value of the microplate reader into a standardized reading for generating a corrected detection result;

[0167] The accuracy of the standardized model is dynamically verified and iteratively optimized through consistency verification of cross-platform detection data, deviation calibration of multi-center data sets, and clinical outcome correlation analysis. The standardized model is used to compensate for detection platform heterogeneity, sample type differences, and inconsistencies in data processing procedures, thereby improving the consistency and diagnostic accuracy of solid tumor driver gene detection readings.

[0168] It should be noted that those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the processor described above can refer to the corresponding process in the method embodiments described in the above embodiments, and will not be repeated here.

[0169] A computer-readable storage medium is also provided in an embodiment of the present application. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. The processor executes the program instructions to implement the steps of the method for normalizing the readings of solid tumor driver genes detected for luciferase provided in the above embodiments of the present application.

[0170] The computer-readable storage medium may be an internal storage unit of the control module described in the aforementioned embodiment, such as a hard disk or memory of the control module. The computer-readable storage medium may also be an external storage device of the control module, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc., equipped on the control module.

[0171] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A readout normalization system for solid tumor driver gene detection of luciferase, characterized in that: include: a multi-source data acquisition module configured to acquire luciferase detection data from multiple detection platforms, wherein the luciferase detection data includes solid tumor driver gene detection results based on NGS, PCR, and FISH technologies, original fluorescence signal values ​​of the microplate reader, sample type information, and detection platform parameters; a control module configured to construct a standardized model based on a large-scale multi-center collaborative dataset, wherein the standardized model is trained on a mapping relationship between the luminescent enzyme detection data and the driver gene expression level using an ultra-precision analysis algorithm, wherein the ultra-precision analysis algorithm includes a gradient boosting tree or a deep neural network for optimizing the resolution of weak signals by the microplate reader; The luminescence enzyme detection data is processed using the standardized model, including: calling a corresponding standardized sub-model according to sample type information and detection platform parameters corresponding to the luminescence enzyme detection data, converting the original fluorescence signal value of the microplate reader into a standardized reading for generating a corrected detection result; Dynamically verify and iteratively optimize the accuracy of the standardized model through consistency verification of cross-platform test data, bias correction of multi-center data sets, and correlation analysis of clinical outcomes; Among them, the standardized model is used to compensate for the heterogeneity of the detection platform, differences in sample types and inconsistencies in the data processing process, so as to improve the consistency and diagnostic accuracy of solid tumor driver gene detection readings.

2. The system according to claim 1, wherein: The step of obtaining luciferase detection data from multiple detection platforms includes: Obtaining detection platform parameters corresponding to each detection platform and sample type information of the solid tumor sample, wherein the detection platform parameters include the microplate reader model, excitation light parameters, and detection time interval; the sample type information includes at least the fixation time and dehydration degree parameters of FFPE section samples and the freezing time and freeze-thaw number parameters of fresh frozen tissue samples; Obtain the driver gene mutation site detection results generated by each detection platform based on NGS, PCR, and FISH technology and the original fluorescence signal value of the microplate reader of the corresponding sample.

3. The system according to claim 1, wherein: The standardized model is constructed based on a large-scale multi-center collaborative dataset, including: Integrate the test data corresponding to multiple centers to form a multi-center data set, and preprocess the multi-center data set, including eliminating abnormal data caused by instrument failure, grouping and labeling based on sample type, and quantifying and encoding the test platform parameters; A basic model framework is constructed based on the preprocessed multi-center dataset, and an external data interface is reserved for data expansion.

4. The system according to claim 1, wherein: The standardized model is trained on the mapping relationship between the luminescence enzyme detection data and the driver gene expression level by using an ultra-precision analysis algorithm, including: Set up at least five layers of tree structure and introduce regularization parameters to weight attenuate the noise signal of each detection platform; A network structure consisting of an input layer, at least two hidden layers, and an output layer was constructed. The hidden layer used an activation function to perform nonlinear extraction of weak signal features in the original fluorescence signal of the microplate reader, so that the correlation coefficient between the predicted value of the driver gene expression output by the standardized model and the gold standard detection result was greater than 0.

95.

5. The system according to claim 1, wherein: The method of calling the corresponding standardized sub-model according to the sample type information and detection platform parameters corresponding to the luciferase detection data includes: The standardized model is pre-divided into the standardized sub-models based on the sample type; the standardized sub-models include an FFPE slice sample sub-model and a fresh frozen tissue sample sub-model.

6. The system according to claim 5, characterized in that The standardized sub-model constructs multiple parameter adaptation units based on the microplate reader model and excitation light wavelength in the detection platform parameters. When the luciferase detection data is input, the standardized model outputs the sample type and detection platform parameters, and matches the parameter adaptation unit with the standardized sub-model combination corresponding to the sample type to process the detection data.

7. The system according to claim 1, wherein: The process of converting the original fluorescence signal value of the microplate reader into a standardized reading includes: Background signal correction was performed on the original fluorescence signal value of the microplate reader based on the sample type. Tissue autofluorescence compensation coefficient was introduced into FFPE slice samples according to the fixed time parameter, and signal attenuation correction was performed on fresh frozen tissue samples according to the freeze-thaw number parameter. The noise baseline of the microplate reader was subtracted based on the detection platform parameters; The correction function output by the standardization model is used to convert the corrected original fluorescence signal value of the microplate reader into a standardized reading of a unified dimension. The standardized reading forms a corresponding mapping relationship with the clinical interpretation threshold of the driver gene copy number variation or expression amount.

8. The system according to claim 1, wherein: Before building a standardized model based on a large-scale multi-center collaborative dataset, the control module is further used to: The luciferase detection data was normalized, including background signal correction for FFPE sections and fresh frozen tissue samples, instrument noise filtering of the detection platform, and data format unification.

9. A method for normalizing the readings of solid tumor driver genes for luciferase, characterized in that: A control module for a solid tumor driver gene detection readout normalization system for luciferase according to any one of claims 1 to 8, the method comprising: The multi-source data acquisition module collects luciferase detection data from multiple detection platforms, wherein the luciferase detection data includes solid tumor driver gene detection results based on NGS, PCR, and FISH technologies, original fluorescence signal values ​​of the microplate reader, sample type information, and detection platform parameters; A standardized model is constructed based on a large-scale multi-center collaborative dataset. The standardized model is trained on the mapping relationship between the luminescence enzyme detection data and the driver gene expression level using an ultra-precision analysis algorithm, including a gradient boosting tree or a deep neural network, to optimize the resolution of the microplate reader for weak signals. The luminescence enzyme detection data is processed using the standardized model, including: calling a corresponding standardized sub-model according to sample type information and detection platform parameters corresponding to the luminescence enzyme detection data, converting the original fluorescence signal value of the microplate reader into a standardized reading for generating a corrected detection result; The accuracy of the standardized model is dynamically verified and iteratively optimized through consistency verification of cross-platform detection data, deviation calibration of multi-center data sets, and clinical outcome correlation analysis. The standardized model is used to compensate for detection platform heterogeneity, sample type differences, and inconsistencies in data processing procedures, thereby improving the consistency and diagnostic accuracy of solid tumor driver gene detection readings.

10. A kit, characterized in that The kit comprises the solid tumor driver gene detection readout normalization system for luciferase according to any one of claims 1 to 8.