Tumor phenotyping method and system based on pathogenic database and machine learning assistance

Through a multimodal feature fusion network based on pathogen database and machine learning, clinical, genetic, and imaging feature extraction and cross-modal dependency modeling of tumor samples are performed, which solves the problem of insufficient cross-modal feature correlation in existing technologies, improves the accuracy and efficiency of tumor phenotyping analysis, and supports precise diagnosis and treatment.

CN120561615BActive Publication Date: 2025-09-30BEIJING KEPTON PHARM TECH DEV CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511044773.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-09-30
Estimated Expiration
2045-07-29

AI Technical Summary

Technical Problem

Existing technologies mostly rely on single-modality data or simple feature splicing, which makes it difficult to capture the deep correlation between cross-modal features, resulting in insufficient accuracy in tumor phenotype judgment.

Method used

Through a method based on pathogenic database and machine learning, the multimodal feature vectors of tumor samples are obtained, and feature extraction is performed using a multimodal feature fusion network. Long-distance dependency modeling of clinical and genetic features is performed, and a second dependency feature map is generated by combining the feature association strength feature map. The target tumor phenotype feature vector is determined and matched with the tumor phenotype database, and the phenotype with the highest confidence is selected as the target result.

Benefits of technology

It improves the accuracy and efficiency of tumor phenotyping analysis, providing support for precise clinical diagnosis and treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120561615B_ABST
    Figure CN120561615B_ABST
Patent Text Reader

Abstract

The present invention discloses a tumor phenotyping method and system based on a pathogenic database and assisted by machine learning, which relates to the field of artificial intelligence. The method comprises: obtaining clinical, genetic, and imaging multimodal feature vectors of a tumor sample, inputting them into a multimodal feature fusion network to extract the path vectors of each modality; performing long-distance dependency modeling on the clinical and genetic features to obtain a first dependency feature graph, and generating a second dependency feature graph in combination with the feature association strength feature graph; determining the target tumor phenotypic feature vector based on the second dependency feature graph and the imaging feature path vector, and finally calculating the matching confidence with the candidate phenotypic feature vectors in the tumor phenotype database, and selecting the phenotype corresponding to the highest confidence as the target result. The present invention improves the accuracy and efficiency of tumor phenotyping analysis by matching the cross-modal feature dependency modeling with the pathogenic database, providing support for clinical precision diagnosis and treatment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and in particular to a method and system for tumor phenotype analysis based on a pathogenicity database and assisted by machine learning. Background Art

[0002] Tumor phenotyping is a key component of precision diagnosis and treatment. Its core lies in integrating multi-source, heterogeneous medical data to accurately classify tumor subtypes. Existing technologies often rely on single-modality data or simple feature splicing, making it difficult to capture the deep connections between cross-modal features, resulting in inaccurate phenotyping. Summary of the Invention

[0003] The purpose of the present invention is to provide a tumor phenotype analysis method and system based on a pathogen database and assisted by machine learning.

[0004] In a first aspect, an embodiment of the present invention provides a method for tumor phenotype analysis based on a pathogenicity database and assisted by machine learning, comprising:

[0005] Obtain tumor sample feature vector;

[0006] Inputting the tumor sample feature vector into the clinical feature path, gene feature path, and imaging feature path in the multimodal feature fusion network of the tumor phenotyping model to obtain a clinical feature path vector, a gene feature path vector, and an imaging feature path vector, respectively;

[0007] Performing long-distance dependency modeling on the clinical feature pathway vector and the gene feature pathway vector to obtain a first dependency feature graph; wherein the first dependency feature graph is used to indicate the strength of association between each first feature in the clinical feature pathway vector and each second feature in the gene feature pathway vector;

[0008] Obtaining a feature association strength feature graph, wherein the feature association strength feature graph is used to indicate a feature association distance between each first feature in the clinical feature pathway vector and each second feature in the gene feature pathway vector, and each bias feature in the feature association strength feature graph is positively correlated with each dependent feature in the first dependent feature graph;

[0009] Determining a second dependency feature graph based on the first dependency feature graph and the feature association strength feature graph, and determining a target tumor phenotype feature vector based on the image feature path vector and the second dependency feature graph;

[0010] Acquiring a tumor phenotype database, wherein the tumor phenotype database includes a plurality of candidate tumor phenotypes and a candidate tumor phenotype feature vector of each candidate tumor phenotype;

[0011] For each of the candidate tumor phenotypes, a phenotype matching confidence between the candidate tumor phenotype feature vector of the candidate tumor phenotype and the target tumor phenotype feature vector is calculated, and the target tumor phenotype is determined based on the highest phenotype matching confidence.

[0012] In a second aspect, an embodiment of the present invention provides a server system, including a server, wherein the server is configured to execute the method described in the first aspect.

[0013] Compared with the existing technology, the beneficial effects provided by the present invention include: using a tumor phenotyping method and system based on a pathogenic database and assisted by machine learning disclosed by the present invention, by obtaining the clinical, genetic, and imaging multimodal feature vectors of the tumor sample, inputting them into the multimodal feature fusion network to extract the path vectors of each modality; performing long-distance dependency modeling on the clinical and genetic features to obtain a first dependency feature graph, and generating a second dependency feature graph in combination with the feature association strength feature graph; determining the target tumor phenotypic feature vector based on the second dependency feature graph and the imaging feature path vector, and finally calculating the matching confidence with the candidate phenotypic feature vector in the tumor phenotype database, and selecting the phenotype corresponding to the highest confidence as the target result. The present invention improves the accuracy and efficiency of tumor phenotyping analysis by matching the cross-modal feature dependency modeling with the pathogenic database, providing support for clinical precision diagnosis and treatment. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly describes the drawings required for use in the embodiments. It should be understood that the following drawings illustrate only certain embodiments of the present invention and should not be construed as limiting the scope of the present invention. Those skilled in the art can, without inventive effort, derive other relevant drawings from these drawings.

[0015] Figure 1 A schematic diagram of the steps of a tumor phenotype analysis method based on a pathogenicity database and assisted by machine learning provided in an embodiment of the present invention;

[0016] Figure 2 A schematic block diagram of the structure of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more apparent, the technical solutions of the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings of the embodiments of the present invention. It should be understood that the described embodiments are only a portion of the embodiments of the present invention, not all of them. Generally, the components of the embodiments of the present invention described and illustrated in the drawings herein may be arranged and designed in a variety of different configurations.

[0018] The specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0019] In order to solve the technical problems in the above background technology, Figure 1 This is a flow chart of a tumor phenotype analysis method based on a pathogenicity database and assisted by machine learning provided in an embodiment of the present disclosure. The following is a detailed introduction to the tumor phenotype analysis method based on a pathogenicity database and assisted by machine learning.

[0020] Step S201, obtaining a tumor sample feature vector;

[0021] Step S202: Inputting the tumor sample feature vector into the clinical feature path, gene feature path, and imaging feature path in the multimodal feature fusion network of the tumor phenotype analysis model to obtain a clinical feature path vector, a gene feature path vector, and an imaging feature path vector, respectively;

[0022] Step S203, performing long-distance dependency modeling on the clinical feature pathway vector and the gene feature pathway vector to obtain a first dependency feature graph; wherein the first dependency feature graph is used to indicate the strength of association between each first feature in the clinical feature pathway vector and each second feature in the gene feature pathway vector;

[0023] Step S204: Obtain a feature association strength characteristic graph, wherein the feature association strength characteristic graph is used to indicate the feature association distance between each first feature in the clinical feature pathway vector and each second feature in the gene feature pathway vector, and each bias feature in the feature association strength characteristic graph is positively correlated with each dependent feature in the first dependent feature graph;

[0024] Step S205: determining a second dependency feature graph based on the first dependency feature graph and the feature association strength feature graph, and determining a target tumor phenotype feature vector based on the image feature path vector and the second dependency feature graph;

[0025] Step S206, obtaining a tumor phenotype database, wherein the tumor phenotype database includes a plurality of candidate tumor phenotypes and a candidate tumor phenotype feature vector of each candidate tumor phenotype;

[0026] Step S207 , for each of the candidate tumor phenotypes, calculating the phenotype matching confidence between the candidate tumor phenotype feature vector of the candidate tumor phenotype and the target tumor phenotype feature vector, and determining the target tumor phenotype based on the highest phenotype matching confidence.

[0027] In an embodiment of the present invention, illustratively, the server first receives multimodal raw data of tumor samples uploaded by the hospital, including clinical data, genetic data, and imaging data, and integrates these heterogeneous data into a unified tumor sample feature vector through feature extraction and encoding processing.

[0028] In terms of clinical data, the patient information corresponding to the sample includes age 62, male sex, 25-year smoking history, clinical tumor marker test results (CEA 4.8 ng / mL, CYFRA21-1 3.5 ng / mL), and no history of hypertension or diabetes. The server uses a structured data encoding module to convert these 12 clinical indicators into a 128-dimensional clinical feature subvector. Each dimension of this subvector corresponds to a quantitative representation of a clinical feature. For example, "25-year smoking history" corresponds to the third dimension of the subvector after normalization, with a value of 0.25; "CEA level 4.8 ng / mL" corresponds to the fifth dimension, with a value of 0.48.

[0029] Regarding genetic data, the sample's genetic testing results showed positive EGFR L858R mutation, negative ALK fusion, and negative KRASG12C mutation. It also included expression data for 200 key driver genes (e.g., EGFR gene expression log2 (TPM) was 7.2, ALK was 3.1, and TP53 was 5.8). The server then processed this data using a gene feature extraction network: a convolutional layer captured local gene interaction features (e.g., co-expression patterns between EGFR and downstream signaling pathway genes), followed by a self-attention layer to mine global gene dependencies, ultimately outputting a 256-dimensional gene feature subvector. The key information, "EGFR mutation positive," was encoded as the 10th dimension of the subvector, with a value of 0.88; "abnormal TP53 expression" corresponded to the 50th dimension, with a value of 0.52.

[0030] Regarding imaging data, the sample chest contrast-enhanced CT 3D volumetric data includes 300 slices (1mm thickness), covering multiple window levels, including lung and mediastinal windows. The server uses a pre-trained 3DResNet model to extract features from the CT data, focusing on imaging features such as tumor nodule size, edge spiculation, and pleural retraction. For example, the sign of "irregular nodules in the right upper lobe" is encoded as the 256th dimension of the image feature subvector, with a value of 0.75; "pleural indentation" corresponds to the 384th dimension, with a value of 0.62. After global average pooling, the image data is converted into a 512-dimensional image feature subvector.

[0031] Finally, the server concatenates the 128-dimensional clinical feature sub-vector, the 256-dimensional gene feature sub-vector, and the 512-dimensional imaging feature sub-vector in sequence to form a tumor sample feature vector with a total dimension of 128+256+512=896, which serves as input data for subsequent model analysis.

[0032] The server inputs the tumor sample feature vector into the multimodal feature fusion network of the tumor phenotyping model. This network consists of three parallel branches: clinical feature pathway, genetic feature pathway, and imaging feature pathway. These branches perform deep feature extraction on the clinical, genetic, and imaging sub-vectors within the feature vector, respectively, to generate more abstract intra-modal features.

[0033] The clinical feature path consists of two fully connected layers and a ReLU activation function. The server inputs the first 128-dimensional clinical subvector of the tumor sample feature vector into this path. The first fully connected layer maps the 128-dimensional features to 256 dimensions, capturing the initial interactions between clinical features (such as the combined effect of "age" and "smoking history"). The second fully connected layer further compresses the features to 128 dimensions and uses the ReLU activation function to enhance nonlinear expression capabilities, ultimately outputting a 128-dimensional clinical feature path vector. For example, the third dimension of this vector has a value of 0.68, corresponding to the cross-feature abstraction of "smoking history and CEA level"; the eighth dimension has a value of 0.72, corresponding to the joint feature abstraction of "age and gender."

[0034] The gene feature path consists of a convolutional layer and a multi-head self-attention layer. The server inputs gene sub-vectors (129-384 dimensions) from the feature vector (a total of 256 dimensions) into this path. The convolutional layer (with a kernel size of 3×3) first extracts local gene features (such as the expression trends of adjacent genes) and outputs 256-dimensional intermediate features. The multi-head self-attention layer (8 attention heads) then performs global dependency modeling on these intermediate features, focusing on capturing the associations between key driver genes (such as the synergistic effect between EGFR mutations and abnormal TP53 expression). The final output is a 256-dimensional gene feature path vector. The 10th dimension has a value of 0.91, corresponding to the association feature between EGFR mutations and downstream signaling pathway activation; the 50th dimension has a value of 0.65, corresponding to the abstract feature of abnormal TP53 expression.

[0035] The image feature path uses a pretrained VisionTransformer (ViT) model fine-tuned on a lung CT image dataset. The server inputs the image subvectors (dimensions 385-896) from the feature vector (a total of 512 dimensions) into this path. ViT's 12-layer Transformer encoder uses a multi-head attention mechanism (12 attention heads) to capture global image features, such as the spatial relationship between the tumor nodule and surrounding blood vessels and pleura, ultimately outputting a 512-dimensional image feature path vector. The 256th dimension has a value of 0.82, corresponding to the enhanced feature of the "spicule sign on the edge of the right upper lung nodule"; the 384th dimension has a value of 0.71, corresponding to the abstract feature of the "pleural traction sign."

[0036] To capture cross-modal associations between clinical features and genetic features (such as the potential biological relationship between "smoking history" and "EGFR mutation"), the server performs long-distance dependency modeling on clinical feature path vectors and genetic feature path vectors, and outputs the first dependency feature graph to quantify the strength of the association between the two.

[0037] Specifically, the server implements this modeling process using a cross-attention mechanism: the clinical feature pathway vector serves as the "query vector," and the gene feature pathway vector serves as both the "key vector" and the "value vector." First, a learnable parameter matrix is ​​used to map the clinical feature pathway vector (128 dimensions) to a "query matrix," and the gene feature pathway vector (256 dimensions) to a "key matrix" and a "value matrix," respectively. The similarity between the query matrix and the key matrix (i.e., the attention score) is calculated, representing the degree of association between the clinical features and the gene features. Finally, the attention score is normalized using a softmax function and then weighted summed with the value matrix to obtain a preliminary dependency feature. To enhance modeling capabilities, the server employs eight attention heads for parallel computation. The outputs of each head are concatenated and integrated through a linear transformation, ultimately resulting in a 128×256 first dependency feature map.

[0038] Each element of this feature map represents the strength of the association between a feature in the clinical feature pathway vector and a feature in the gene feature pathway vector (values ​​range from 0 to 1, with higher values ​​indicating stronger associations). For example, the strength of the association between the third feature in the clinical feature pathway vector (corresponding to "smoking history") and the tenth feature in the gene feature pathway vector (corresponding to "EGFR mutation") is 0.85, indicating a strong association between the two. However, the strength of the association between the eighth feature in the clinical feature pathway vector (corresponding to the "age-sex joint feature") and the 50th feature in the gene feature pathway vector (corresponding to "TP53 expression") is 0.32, indicating a weaker association.

[0039] To enhance the association weights of "neighboring features" between clinical and genetic features (for example, clinical and genetic features at the same position in the original tumor sample feature vector typically have a stronger prior correlation), the server needs to obtain a feature association strength feature map. The elements of this feature map are positively correlated with the elements of the first dependent feature map. In other words, the closer the "association distance" between features, the higher the element value of the feature association strength feature map.

[0040] First, the server defines the "feature index" rule: in the tumor sample feature vector, the i-th dimension of the clinical feature subvector (corresponding to the i-th feature of the clinical feature path vector) and the i-th dimension of the gene feature subvector (corresponding to the i-th feature of the gene feature path vector) have the same "feature index" i (i = 1 to 128; since the gene feature path vector has 256 dimensions, the index of the 129-256-dimensional features is recorded as i+128).

[0041] Based on this, the server calculates a "primary association strength feature map" based on the distance of the feature index: for the i-th feature of the clinical feature pathway vector and the j-th feature of the gene feature pathway vector, the index distance between the two is |ij|, and the element values ​​of the primary association strength feature map decay as the index distance increases (for example, the strength is highest when the distance is 0, and the strength decreases significantly when the distance is 20). For example, when i=j=10 (clinical feature index 10 and gene feature index 10), the index distance is 0 and the primary association strength value is 1.0; when i=10 and j=15 (index distance 5), the strength value is approximately 0.61; when i=10 and j=30 (index distance 20), the strength value is approximately 0.14.

[0042] The server then weights the initial correlation strength feature map using a learnable scaling factor (optimized during training, currently set to 1.2) to produce the final feature correlation strength feature map. The element values ​​of this feature map then form a positive synergy with the element values ​​of the first dependent feature map. Feature pairs with close index distances (e.g., features where i=j) have high strength values, which, together with elements with high correlation strength in the first dependent feature map (e.g., "smoking history" and "EGFR mutation"), enhance the relevance of cross-modal features.

[0043] The server obtains the second dependency feature map by fusing the first dependency feature map and the feature association strength feature map, and then combines it with the image feature path vector to finally determine the target tumor phenotype feature vector through iterative optimization.

[0044] The fusion process uses an element-by-element multiplication approach: each element in the second dependency feature map is equal to the product of the corresponding element in the first dependency feature map and the corresponding element in the feature association strength feature map. This operation not only preserves the original association strengths between clinical and genetic features, but also enhances the weights of neighboring features by biasing feature index distances, filtering out distant noise. For example, the first dependency strength between the 10th feature (index 10) of the clinical feature pathway vector and the 10th feature (index 10) of the genetic feature pathway vector is 0.92, and the feature association strength is 1.2 × 1.0 = 1.2. After fusion, the second dependency strength is 0.92 × 1.2 = 1.104 (significantly increased weight). In contrast, the first dependency strength between the third feature (index 3) of the clinical feature pathway vector and the 10th feature (index 10, distance 7) of the genetic feature pathway vector is 0.85, and the feature association strength is 1.2 × 0.496 ≈ 0.596. After fusion, the second dependency strength is 0.85 × 0.596 ≈ 0.507 (modestly reduced weight).

[0045] Next, the server uses the second dependency feature map to perform a weighted fusion of the image feature path vectors, resulting in a long-range dependency modeling result. Using the second dependency feature map (128×256) as the attention weight, the image feature path vectors (512 dimensions) are globally weighted to capture the modulatory effects of clinical-genetic associations on image features (e.g., the enhancement of the "spicule sign" feature by "EGFR mutation plus smoking history"). This result is then "skip-connected" (element-by-element addition) with the original tumor sample feature vector, preserving the original feature information while incorporating cross-modal correlation features to obtain a preliminary tumor sample feature vector.

[0046] To further enhance the abstract expressiveness of features, the server performs a nonlinear transformation on the primary tumor sample feature vector using a multi-layer perceptron (with two fully connected layers and a Reluctant Unit (ReLU) activation function) to generate a high-dimensional abstract feature vector. This high-dimensional abstract feature vector is then skip-added to the primary tumor sample feature vector to generate an optimized phenotypic feature vector, with the number of optimization rounds set to 1. The server uses this optimized vector as the new tumor sample feature vector and repeats the multimodal feature path extraction, dependency modeling, and fusion optimization process until the number of optimization rounds reaches the preset value (three), ultimately outputting the target tumor phenotypic feature vector (896 dimensions).

[0047] The server calls the tumor phenotype database and determines the target tumor phenotype of the sample by calculating the matching confidence between the target tumor phenotype feature vector and the phenotype feature vector to be selected in the database.

[0048] The tumor phenotype database stores standard data for four common lung cancer phenotypes: "lung adenocarcinoma," "lung squamous cell carcinoma," "small cell lung cancer," and "large cell lung cancer." Each phenotype corresponds to an 896-dimensional candidate tumor phenotype feature vector. These vectors are extracted by the tumor phenotype analysis model from a large number of standard samples (confirmed by pathology) and contain typical clinical, genetic, and imaging feature combinations for each phenotype. For example, in the candidate feature vector for "lung adenocarcinoma," the value of the 10th dimension is 0.90 (corresponding to the standard feature of "EGFR mutation"), and the value of the 256th dimension is 0.85 (corresponding to the standard imaging feature of "spicule sign of right upper pulmonary nodule").

[0049] For each candidate phenotype, the server calculates the spatial proximity (i.e., cosine similarity) between the target tumor phenotype feature vector and the candidate phenotype feature vector. Values ​​closer to 1 indicate greater feature similarity. This spatial proximity is then converted into phenotypic match confidence using a pre-trained confidence adjustment factor (0.88 in this scenario). The calculation results are as follows: the match confidence for "lung adenocarcinoma" is 0.920, for "squamous cell lung cancer" is 0.659, for "small cell lung cancer" is 0.477, and for "large cell lung cancer" is 0.579.

[0050] Finally, the server selected "lung adenocarcinoma" corresponding to the highest matching confidence (0.920) as the target tumor phenotype, and returned the analysis results (including the phenotype name and key associated features such as "EGFR mutation + smoking history + spicule sign imaging") to the hospital terminal system to provide data support for clinical diagnosis.

[0051] Through the above process, the server realizes automated analysis from multimodal tumor sample data to precise tumor phenotypes. It not only integrates multi-dimensional information of clinical, genetic, and imaging, but also mines deep correlations between features through machine learning models, thereby improving the accuracy and efficiency of phenotypic analysis.

[0052] In the embodiment of the present invention, determining the target tumor phenotype feature vector based on the image feature path vector and the second dependency feature graph can be implemented through the following example.

[0053] Determine a phenotypic feature optimization vector according to the image feature path vector and the second dependent feature graph, and increase the number of optimization rounds by one, where the number of optimization rounds starts from the starting round;

[0054] The phenotypic feature optimization vector is used as the tumor sample feature vector, and the step of inputting the tumor sample feature vector into the clinical feature pathway, gene feature pathway, and imaging feature pathway in the multimodal feature fusion network of the tumor phenotyping analysis model is repeatedly performed until the optimization rounds reach a preset number of rounds, and the phenotypic feature optimization vector is used as the target tumor phenotypic feature vector.

[0055] In the embodiment of the present invention, for example, the preset optimization rounds are 3 rounds (the starting round is recorded as 0), and the specific process is as follows: First round optimization (starting round → round 1): the server first converts the second dependency feature map (128×256) as the attention weight, the image feature path vector I (512 dimensions) is weighted fused—— The "clinical-gene strong correlation feature pair" (such as the correlation strength of "smoking history-EGFR mutation" is 0.507) will enhance the weight of the corresponding imaging feature (such as the feature dimension value of "spicule sign imaging" is increased from 0.82 to 0.89), and the long-distance dependency modeling result M (128 dimensions) is obtained. Subsequently, M is jump-connected (element-by-element addition) with the original tumor sample feature vector X (896 dimensions), retaining the original feature information while incorporating cross-modal associations to obtain the primary tumor sample feature vector Then, a multi-layer perceptron (MLP) is used to Perform nonlinear transformation (such as strengthening the weak correlation feature of "age-gene expression") and output a high-dimensional abstract feature vector Z; then Z is combined with Skip connection addition to obtain the phenotypic feature optimization vector , the optimization round is updated to 1. Second round of optimization (round 1 → round 2): the server will As a new tumor sample feature vector, it is re-input into the multimodal feature fusion network: the clinical feature path extracts a more accurate "smoking history-CEA level" cross feature (the dimension value increases from 0.68 to 0.75), the gene feature path strengthens the "EGFR mutation-TP53 expression" collaborative feature (the dimension value increases from 0.65 to 0.72), and the imaging feature path further focuses on the details of the "pleural traction sign" (the dimension value increases from 0.71 to 0.78). Repeat the dependency modeling and fusion process to obtain the optimized vector , the round is updated to 2. The third round of optimization (round 2 → round 3): continue with The network uses iterative learning to further suppress noise features (such as the weak correlation strength of "sex-irrelevant genes" is reduced from 0.32 to 0.21), enhance core phenotypic features (such as the correlation strength of "EGFR mutation-spicule sign imaging" is increased from 0.89 to 0.93), and finally output At this point, the optimization rounds have reached the preset 3 rounds, and the server will The target tumor phenotype feature vector is determined and used for subsequent tumor phenotype database matching.

[0056] In the embodiment of the present invention, determining the phenotypic feature optimization vector based on the image feature path vector and the second dependency feature graph can be implemented through the following examples.

[0057] Obtaining a long-distance dependency modeling result according to the image feature path vector and the second dependency feature graph;

[0058] determining a primary tumor sample feature vector according to the long-distance dependency modeling result and the skip connection output of the tumor sample feature vector;

[0059] Performing a nonlinear transformation on the primary tumor sample feature vector according to the multi-layer perceptron of the tumor phenotype analysis model to obtain a high-dimensional abstract feature vector;

[0060] The phenotypic feature optimization vector is obtained according to the sum of the skip connection outputs of the high-dimensional abstract feature vector and the primary tumor sample feature vector.

[0061] In the embodiment of the present invention, for example, the current optimization round is 1, and the specific process is as follows: Step 1: Generate the long-distance dependency modeling results. The server will (128×256) is used as the attention weight to perform weighted fusion on the image feature path vector I (512 dimensions). The correlation strength of "the third dimension of clinical characteristic path vector (smoking history) - the tenth dimension of gene characteristic path vector (EGFR mutation)" is 0.507, which corresponds to the 256th dimension "spicule sign image" feature in the enhanced I (original value 0.82). The correlation strength between the 10th dimension of the clinical feature path vector (age-sex joint feature) and the 10th dimension of the genetic feature path vector (EGFR mutation) in I is 1.104, further enhancing the 384th dimension of the "pleural traction sign imaging" feature in I (original value 0.71). After global weighting and pooling, a 128-dimensional long-distance dependency modeling result M is obtained, with the 3rd dimension value being 0.12 (corresponding to the cross-modal correlation between "smoking history-EGFR mutation-spiculation sign") and the 10th dimension value being 0.18 (corresponding to the cross-modal correlation between "age-sex-EGFR mutation-pleural traction sign"). Step 2: Determine the primary tumor sample feature vector. The server performs a skip connection between M (128 dimensions) and the tumor sample feature vector X (896 dimensions): Each of the 128-dimensional features of M is element-wise added to the corresponding dimensions of the first 128-dimensional clinical subvectors in X, preserving the original features while incorporating cross-modal correlations. For example, the original value of the third dimension "smoking history" in X is 0.25, and after adding it to the third dimension of M 0.12, it is updated to 0.37; the original value of the tenth dimension "age-sex joint feature" in X is 0.72, and after adding it to the tenth dimension of M 0.18, it is updated to 0.90. After skip connection, the primary tumor sample feature vector of 896 dimensions is obtained. Step 3: Generate high-dimensional abstract feature vectors. The server calls the multi-layer perceptron (MLP) of the tumor phenotype analysis model to Perform nonlinear transformation: MLP contains two fully connected layers (896→1024→896) and ReLU activation function, focusing on strengthening Features that are weakly associated but biologically significant. For example, The original value of the 50th dimension "TP53 expression" was 0.52. After MLP mining of the potential association with "pleural traction sign", the output value increased to 0.68; the original value of the 256th dimension "spicule sign image" was 0.82 and increased to 0.91 after enhancement. Finally, an 896-dimensional high-dimensional abstract feature vector Z was obtained. Step 4: Determination of the phenotypic feature optimization vector. The server compares Z with Perform skip connection: Z's 896-dimensional features and The corresponding dimensions are added element by element to fuse the abstract features with the primary features. For example, the 50th dimension of Z is 0.68 and After adding the 50th dimension 0.52, it is updated to 1.20; Z 256th dimension 0.91 and After adding the 256th dimension value of 0.82, it is updated to 1.73. This operation results in an 896-dimensional phenotypic feature optimization vector X_opt, which is used to enter the next round of optimization iteration.

[0062] In this embodiment of the present invention, for any feature dimension in the tumor sample feature vector, the feature index of the first feature matched by the feature dimension in the clinical feature pathway vector is the same as the feature index of the second feature matched by the feature dimension in the gene feature pathway vector, wherein the feature index is used to indicate the position information of the feature dimension in the tumor sample feature vector;

[0063] The feature association strength feature map is obtained by the following process and can be implemented through the following example.

[0064] Determining a quantitative feature value of the bias feature of the primary-order association strength feature graph according to an index distance between a feature index of each first feature in the clinical feature pathway vector and a feature index of each second feature in the gene feature pathway vector;

[0065] The feature association strength feature map is determined according to a weighted multiplication result of a learnable scaling coefficient and the primary-order association strength feature map.

[0066] In this embodiment of the present invention, the tumor sample feature vector is 896-dimensional, composed of a clinical sub-vector (dimensions 1-128), a genetic sub-vector (dimensions 129-384), and an imaging sub-vector (dimensions 385-896). The specific process is as follows: The server defines a "feature index" as the position of a feature dimension in the tumor sample feature vector. For any feature dimension in the tumor sample feature vector, the feature index of the first feature matched in the clinical feature path vector is the same as the feature index of the second feature matched in the genetic feature path vector. For example, the 10th dimension of the clinical subvector (tumor sample feature vector index 10) corresponds to the first feature of the clinical feature pathway vector, with a feature index of 10. The 10th dimension of the gene subvector (tumor sample feature vector index 128 + 10 = 138) corresponds to the second feature of the gene feature pathway vector, also with a feature index of 10. The clinical pathway vector has 128 dimensions (feature indices 1-128), and the gene pathway vector has 256 dimensions (feature indices 1-256, with 1-128 corresponding to the first 128 dimensions of the gene subvector and 129-256 corresponding to the last 128 dimensions of the gene subvector). Both feature indices correspond one-to-one with the corresponding modal subvectors in the original tumor sample feature vector. The server calculates the quantized eigenvalue of the bias feature in the primary association strength feature graph based on the "index distance" between the feature index of the first feature in the clinical feature pathway vector (denoted as i, 1≤i≤128) and the feature index of the second feature in the gene feature pathway vector (denoted as j, 1≤j≤256). The index distance is defined as |ij|. The quantitative eigenvalue decays exponentially with increasing distance (attenuation coefficient 0.1, empirical value), using the formula: Quantized eigenvalue = exp(-0.1 × index distance). Specific examples are as follows: i = 10, j = 10 (clinical pathway index 10 and gene pathway index 10, corresponding to tumor sample feature vector indices 10 and 138): Index distance = |10-10| = 0, Quantized eigenvalue = exp(-0.1×0) = 1.0 (maximum strength, indicating the strongest a priori association with the index feature); i = 10, j = 15 (clinical pathway index 10 and gene pathway index 15, corresponding to tumor sample feature vector indices 10 and 143): Index distance = |10-15| = 5, Quantized eigenvalue = (medium intensity, indicating strong correlation between neighboring index features); i=10, j=30 (clinical pathway index 10 and gene pathway index 30, corresponding to tumor sample feature vector indices 10 and 158): index distance = |10-30| = 20, quantitative eigenvalue = exp(-0.1×20) = exp(-2) ≈ 0.1353 (low intensity, indicating weak correlation between distant index features); i=10, j=130 (clinical pathway index 10 and gene pathway index 130, corresponding to the second dimension of the 128-dimensional gene subvector): index distance = |10-130| = 120, quantitative eigenvalue = (The intensity is extremely low, indicating that there is almost no prior correlation between the features of the cross-modal sub-vector partitions.) Through the above calculations, the server generates a 128×256 primary-order association intensity feature map, whose element values ​​reflect the index distance bias intensity of the clinical-gene feature pair.

[0067] The server calls the learnable scaling factor γ (optimized during model training, γ=1.2 in the current scenario) and performs weighted multiplication on the primary correlation strength feature map to obtain the final feature correlation strength feature map. Specific examples are as follows: the element value of i=10, j=10 in the primary correlation strength feature map is 1.0, and the corresponding element value of the weighted feature correlation strength feature map is 1.0×1.2=1.2; the primary value of i=10, j=15 is 0.6065, and the weighted value is 0.6065×1.2≈0.7278; the primary value of i=10, j=30 is 0.1353, and the weighted value is 0.1353×1.2≈0.1624; the primary value of i=10, j=130 is , after weighting ≈ In the final generated feature association strength feature map, the element value is negatively correlated with the index distance of the clinical-gene feature pair, and the overall association strength is enhanced by the scaling coefficient, providing a distance bias basis for subsequent fusion with the first dependent feature map.

[0068] In an embodiment of the present invention, the tumor phenotype analysis model is obtained by fine-tuning the multimodal basic model in the following manner, which can be implemented through the following examples.

[0069] For each training reference tumor sample, construct a target phenotype instance and a non-target phenotype instance, wherein the reference tumor phenotype data in the target phenotype instance matches the training reference tumor sample, and the reference tumor phenotype data in the non-target phenotype instance does not match the training reference tumor sample;

[0070] For each of the training reference tumor samples, loading the target phenotype instance that matches the training reference tumor sample into the multimodal base model, obtaining a sample tumor phenotype feature vector of the training reference tumor sample and a first candidate tumor phenotype feature vector of the reference tumor phenotype data in the target phenotype instance, and calculating a first phenotype matching confidence based on the sample tumor phenotype feature vector and the first candidate tumor phenotype feature vector;

[0071] loading the non-target phenotype instance that matches the training reference tumor sample into the multimodal base model, obtaining a sample tumor phenotype feature vector of the training reference tumor sample and a second candidate tumor phenotype feature vector of the reference tumor phenotype data in the non-target phenotype instance, and calculating a second phenotype matching confidence based on the sample tumor phenotype feature vector and the second candidate tumor phenotype feature vector;

[0072] The network parameters of the multimodal basic model are adjusted according to the first phenotype matching confidence and the second phenotype matching confidence of the plurality of training reference tumor samples to obtain the tumor phenotype analysis model.

[0073] In this embodiment of the present invention, for example, a server fine-tunes a multimodal base model (pre-trained using masked feature reconstruction). The training dataset contains 1,000 pathologically confirmed lung cancer samples (including clinical, genetic, and imaging data, as well as true phenotype labels). The specific process is as follows: The server first constructs a "sample-phenotype mapping table" based on the training dataset, recording the matching relationship between each training reference tumor sample and its true phenotype (e.g., sample "TR-LC-001" corresponds to the phenotype "lung adenocarcinoma," "TR-LC-002" corresponds to "lung squamous cell carcinoma," etc.). For each training reference tumor sample, the server constructs two types of instances: Target phenotype instances, which consist of the original data (clinical, genetic, and imaging feature vectors) of the training reference tumor sample and the matching reference tumor phenotype data (the standard feature vector of the true phenotype). For example, the original data for the training reference tumor sample "TR-LC-001" is an 896-dimensional feature vector for a CT image of a 62-year-old male with a 25-year smoking history, positive for EGFR mutations, and a spiculated sign on the right upper lung. Its true phenotype is "lung adenocarcinoma." Therefore, the target phenotype instance includes the original data for this sample and the standard feature vector for "lung adenocarcinoma" (896 dimensions, generated from the average feature vector of 100 confirmed lung adenocarcinoma samples, with a dimension value of 0.92 for "EGFR mutation" and 0.88 for "spiculated sign"). A non-target phenotype instance consists of the original data for the training reference tumor sample and the non-matched reference tumor phenotype data. The server randomly selects three phenotypic data from the "sample-phenotype mapping table" that differ from the current sample's true phenotype (for example, if the true phenotype of "TR-LC-001" is "lung adenocarcinoma," the non-target phenotype data is the standard feature vectors for "squamous lung carcinoma," "small cell lung cancer," and "large cell lung cancer"). The server constructs three non-target phenotype instances, each containing the original data for "TR-LC-001" and a standard feature vector for the non-matching phenotype (for example, the standard vector for "squamous lung carcinoma" has a dimension value of 0.15 for "EGFR mutation" and a dimension value of 0.85 for "central mass imaging"). For each training reference tumor sample, the server loads its target phenotype instance into the multimodal base model. The model consists of two parallel feature extraction branches: one branch processes the original sample data and outputs a sample tumor phenotype feature vector; the other branch processes the reference tumor phenotype data and outputs a first candidate tumor phenotype feature vector. Taking "TR-LC-001" as an example: Sample tumor phenotypic feature vector extraction: The model performs multimodal fusion (clinical-gene-imaging feature interaction) on the original feature vector of "TR-LC-001" (896 dimensions) and outputs an 896-dimensional sample tumor phenotypic feature vector S, where the "EGFR mutation-spicule sign imaging" association dimension value is 0.85, and the "smoking history-age" joint dimension value is 0.72.Extracting the first candidate tumor phenotype feature vector: The model performs the same feature extraction process on the standard feature vector for "lung adenocarcinoma" in the target phenotype instance, outputting an 896-dimensional first candidate tumor phenotype feature vector, T1. The correlation dimension for "EGFR mutation-spiculation imaging" is 0.90, and the combined dimension for "smoking history-age" is 0.75. Calculating the confidence of the first phenotype match: The server calculates the cosine similarity (spatial proximity) between S and T1, yielding the result. =0.88 (the closer the value is to 1, the more similar the features are); then the confidence adjustment coefficient β=0.85 (controlling the confidence scale) of the pre-training is used to calculate the first phenotype matching confidence = =0.88 / 0.85≈1.035 (the goal is to make this value as close to 1 as possible, indicating that the sample is highly similar to the matching phenotype). The server loads the non-target phenotype instances of the training reference tumor samples into the multimodal basic model in sequence, repeats the above feature extraction and confidence calculation process, and obtains the second phenotype matching confidence. Taking the "squamous lung carcinoma" non-target instance of "TR-LC-001" as an example: Second candidate tumor phenotype feature vector extraction: The model extracts features from the "squamous lung carcinoma" standard feature vector and outputs an 896-dimensional second candidate tumor phenotype feature vector , where the correlation dimension value of “EGFR mutation-spicule sign imaging” is 0.20 (EGFR mutation rate is low in squamous cell carcinoma), and the dimension value of “central mass imaging” is 0.85. Second phenotypic matching confidence calculation: Calculate the sample tumor phenotypic feature vector S and The cosine similarity of =0.42; the confidence of the second phenotype match is calculated by the confidence adjustment coefficient β=0.85 =0.42 / 0.85≈0.494 (the goal is to make this value as low as possible, indicating that the sample has a significant phenotypic difference from the non-matching one). Similarly, the server calculates the second phenotypic matching confidence of the non-target instances of "small cell lung cancer" and "large cell lung cancer" of "TR-LC-001", which are 0.382 and 0.451 respectively. The server collects the first phenotypic matching confidence of all training reference tumor samples (such as "TR-LC-001"). =1.035, "TR-LC-002" = 0.987 and the second phenotype matching confidence (3 non-target instances per sample, 3000 in total) The model error is calculated through the loss function, and the network parameters of the multimodal basic model (such as the weight matrix of the clinical / gene / imaging pathway, the parameters of the attention head, etc.) are adjusted through backpropagation. The specific loss calculation logic is: for each sample, all its second phenotype matching confidences are regarded as "negative samples" and the first phenotype matching confidences are regarded as "positive samples". Through the comparison loss function (such as InfoNCE loss) optimization, the confidence of the positive sample is significantly higher than that of the negative sample. For example, the loss component of "TR-LC-001" is = ≈ ≈0.258. A smaller value indicates a stronger model's ability to distinguish between matching and non-matching phenotypes. The server averages the loss components of 1000 samples to obtain a total loss value. Using the Adam optimizer (learning rate 1e-5), the model parameters are iteratively adjusted until the total loss converges (for example, the loss drops from 1.2 to 0.35 after 50 iterations). This results in a fine-tuned tumor phenotyping model. Through this process, the server enables the model to distinguish the degree of match between tumor samples and different phenotypes, improving the accuracy of phenotypic predictions for unknown samples.

[0074] In the embodiment of the present invention, constructing target phenotype instances and non-target phenotype instances for each training reference tumor sample can be implemented through the following examples.

[0075] Obtaining a sample-phenotype mapping table, wherein the sample-phenotype mapping table includes a plurality of reference tumor samples and reference tumor phenotype data that are matched one-to-one with the plurality of reference tumor samples;

[0076] For each training reference tumor sample in the sample-phenotype mapping table, the target phenotype instance is constructed based on the training reference tumor sample and the reference tumor phenotype data matching the training reference tumor sample, and the non-target phenotype instance is constructed based on the training reference tumor sample and the reference tumor phenotype data matching other reference tumor samples other than the training reference tumor sample.

[0077] In an embodiment of the present invention, for example, a server processes a lung cancer training dataset, which contains 1,000 pathologically confirmed reference tumor samples (numbered "TR-LC-001" to "TR-LC-1000"). The specific process is as follows: the server first retrieves the organized "sample-phenotype mapping table" from the hospital pathology database. This table is a structured table, and each row contains a unique identifier (sample ID) of a reference tumor sample, original multimodal data (clinical, genetic, and imaging feature vectors, 896 dimensions), and reference tumor phenotypic data (standard phenotypic feature vectors, 896 dimensions) that matches it one by one. For example, for sample ID "TR-LC-001," the original data consists of an 896-dimensional feature vector for a CT image of a 62-year-old male with a 25-year smoking history, positive for the EGFR L858R mutation, and a spiculated sign on the right upper lung. The matched reference tumor phenotype data is a standard feature vector for "lung adenocarcinoma" (generated by averaging the feature vectors of 100 confirmed lung adenocarcinoma samples, with a "EGFR mutation association" dimension value of 0.92 and a "spiculated sign" dimension value of 0.88). For sample ID "TR-LC-002," the original data consists of an 896-dimensional feature vector for a CT image of a 58-year-old male with a 30-year smoking history, positive for the KRAS G12C mutation, and a central mass. The matched reference tumor phenotype data is a standard feature vector for "lung squamous cell carcinoma" (with a "KRAS mutation association" dimension value of 0.85 and a "central mass" dimension value of 0.90). Sample ID "TR-LC-003": The original data consists of an 896-dimensional feature vector of a CT image of a 70-year-old female with no smoking history, positive for ALK fusion, and diffuse nodules in both lungs. The matched reference tumor phenotype data is a standard feature vector for "small cell lung cancer" (with a dimension value of 0.88 for "ALK fusion association" and 0.91 for "diffuse nodule imaging"). For each training reference tumor sample in the sample-phenotype mapping table, the server extracts its original data and the matched reference tumor phenotype data, combining them into a target phenotype instance. Take the training reference tumor sample "TR-LC-001" as an example: obtain the original multimodal feature vector of "TR-LC-001" from the mapping table (896 dimensions, including clinical features such as "25-year smoking history", gene "EGFR mutation positive", and imaging features such as "spicule sign in the right upper lung"); extract the reference tumor phenotypic data matching the sample - the "lung adenocarcinoma" standard feature vector (896 dimensions, including standard phenotypic patterns such as "EGFR mutation-spicule sign" association features and "smoking history-age" joint features); package the above original data with the reference tumor phenotypic data to generate the target phenotypic instance of "TR-LC-001", which is recorded as (sample data: TR-LC-001 original vector, phenotypic data: lung adenocarcinoma standard vector).The server selects other reference tumor samples with different phenotypes from the current training reference tumor sample "TR-LC-001" from the sample-phenotype mapping table, extracts the matching reference tumor phenotype data, and combines it with the original data of "TR-LC-001" to form a non-target phenotype instance. The specific operation is as follows: "TR-LC-001" and its matching "lung adenocarcinoma" phenotype data are excluded from the mapping table, and reference tumor samples with the phenotypes of "lung squamous cell carcinoma," "small cell lung cancer," and "large cell lung cancer" (such as "TR-LC-002," "TR-LC-003," and "TR-LC-005") are selected; the standard feature vector of "lung squamous cell lung cancer" matching "TR-LC-002" (including the "KRAS mutation-central mass" association feature) and the standard feature vector of "small cell lung cancer" matching "TR-LC-003" (including the "ALK fusion-diffuse nodule" association feature) are extracted. , the standard feature vector for "large cell lung cancer" matching "TR-LC-005" (including the "no clear driver mutation-necrosis imaging" association feature); the original data of "TR-LC-001" is packaged with the three non-matching phenotypic data above to generate three non-target phenotypic instances, respectively denoted as (sample data: TR-LC-001 original vector, phenotypic data: lung squamous cell lung cancer standard vector), (sample data: TR-LC-001 original vector, phenotypic data: small cell lung cancer standard vector), and (sample data: TR-LC-001 original vector, phenotypic data: large cell lung cancer standard vector). Through this process, the server constructs one target phenotypic instance (matching the true phenotype) and three non-target phenotypic instances (non-matching phenotypes) for each training reference tumor sample, providing positive and negative sample pairs for subsequent model fine-tuning.

[0078] In an embodiment of the present invention, the adjusting of the network parameters of the multimodal basic model according to the first phenotype matching confidence and the second phenotype matching confidence of the plurality of training reference tumor samples can be implemented through the following examples.

[0079] For each of the training reference tumor samples, determining a first sample error component according to the first phenotype matching confidence and a first accumulated value of a plurality of second phenotype matching confidences;

[0080] taking an average loss component of the plurality of first sample error components as a first sample error value;

[0081] The network parameters of the multimodal basic model are adjusted according to the first sample error value.

[0082] In the embodiment of the present invention, for example, taking the server processing a lung cancer training data set (containing 1000 training reference tumor samples) as an example, each sample corresponds to a first phenotype matching confidence ( ) and 3 second phenotype matching confidences ( 、 、 ), the specific process is as follows: For each training reference tumor sample, the server first sums up its multiple second phenotype matching confidences to obtain a first cumulative value, and then calculates the first sample error component based on the first phenotype matching confidence. Taking the training reference tumor sample "TR-LC-001" (phenotype "lung adenocarcinoma") as an example: the first phenotype matching confidence is: =1.035 (matching degree between the sample and the target phenotype of “lung adenocarcinoma”, the higher the value, the better the match); second phenotype matching confidence: the matching degrees of the three non-target instances are =0.494 (lung squamous cell carcinoma), =0.382 (small cell lung cancer), =0.451 (large cell lung cancer) (the lower the value, the greater the difference from the non-target phenotype); First cumulative value: sum the three second reliabilities to get 1.327 (0.494 + 0.382 + 0.451); First sample error component: calculated using the contrast loss formula, the logic is "the higher the first reliability, the lower the first cumulative value, the smaller the error", the formula is: Error component = Substituting the data into the equation, we get: = ≈ ≈0.823 (a smaller value indicates a stronger model's ability to distinguish between samples and target / non-target phenotypes). The server calculates the first-sample error component for each of the 1000 training reference tumor samples and takes the average as the first-sample error value. For example, in addition to the error component of 0.823 for "TR-LC-001," the error component of "TR-LC-002" (a lung squamous cell carcinoma sample) is 0.791, and that of "TR-LC-003" (a small cell lung cancer sample) is 0.856. After traversal calculation, the sum of the error components for all 1000 samples is 827.5. The first-sample error value = total error component / number of samples = 827.5 / 1000 = 0.8275 (this value reflects the model's current overall discriminatory ability, and the goal is to gradually reduce it through parameter adjustment). The server inputs the first-sample error value into the optimizer (Adam optimizer with a learning rate of 1e-5), which adjusts the network parameters of the multimodal base model through backpropagation. For example, if the error value of 0.8275 exceeds the preset convergence threshold (0.35), the optimizer calculates the gradient of the error with respect to each network parameter (such as the weight of the fully connected layer of the clinical pathway, the attention head parameters of the gene pathway, and the convolution kernel parameters of the image pathway). The optimizer fine-tunes parameters that contribute significantly to the error in TR-LC-001 (such as the weight of the "EGFR mutation-spicule sign imaging" association feature) by increasing the attention weight corresponding to this feature from 0.85 to 0.89 to enhance the expression of the target phenotype association feature. The optimizer also reduces the weight of the non-target association feature "smoking history-central lung squamous cell carcinoma mass" from 0.32 to 0.28. After completing one round of parameter adjustment, the server recalculates the first-sample error value for all samples. If the error drops to 0.791 (a decrease of 0.0365 from the previous round), the optimization process continues. The above process is repeated until the first-sample error value converges to below 0.35, ultimately resulting in the optimized tumor phenotypic analysis model. Through the above process, the server continuously optimizes model parameters based on error feedback of sample matching confidence, thereby improving the ability to distinguish and match tumor phenotypes.

[0083] In the embodiment of the present invention, for each of the training reference tumor samples, determining the first sample error component according to the first phenotype matching confidence and a first accumulated value of a plurality of second phenotype matching confidences can be implemented through the following example.

[0084] using a second phenotype matching confidence of a non-target phenotype instance for each of the training reference tumor samples as a non-target phenotype index component of an exponential transformation of an Euler number, and summing a plurality of the non-target phenotype index components to obtain a first accumulated value;

[0085] Summing a target phenotypic index component of an exponential transformation using the first phenotypic matching confidence as an Euler number and the first accumulated value to obtain a second accumulated value;

[0086] An inverse logarithm is taken for a sample error ratio having a target phenotypic index component of an exponential transformation with the second accumulated value as the dividend and the first phenotype matching confidence as the Euler number as the divisor to obtain a first sample error component of the training reference tumor sample.

[0087] In the embodiment of the present invention, for example, the server processes the training reference tumor sample "TR-LC-001" (phenotype is "lung adenocarcinoma") as an example. The sample corresponds to a first phenotype matching confidence ( =1.035) and 3 second phenotype matching confidences (lung squamous cell carcinoma =0.494, small cell lung cancer =0.382, large cell lung cancer =0.451). The specific process is as follows: For each non-target phenotype instance of "TR-LC-001", the server uses its second phenotype matching confidence as the exponent of the Euler number (natural constant e, approximately 2.718), calculates the non-target phenotype index component, and then sums multiple components to obtain the first cumulative value. Non-target lung squamous cell carcinoma instance: Second phenotype matching confidence =0.494, non-target phenotype index component = exp(0.494)≈ ≈1.639 (exp is the Euler number exponential transformation function); small cell lung cancer non-target example: =0.382, non-target phenotype index component = exp(0.382)≈ ≈1.465; non-target example for large cell lung cancer: =0.451, non-target phenotype index component = exp(0.451)≈ ≈1.569; First cumulative value: sum the three non-target phenotype index components to get 1.639+1.465+1.569=4.673, which reflects the comprehensive "interference intensity" of all non-target phenotypes. The server will match the first phenotype with confidence. Perform Euler number exponential transformation to obtain the target phenotypic index component, and then sum it with the first accumulated value to obtain the second accumulated value. Target phenotypic index component: =1.035, the target phenotype index component = exp(1.035) ≈ 2.718^1.035 ≈ 2.815, which reflects the "true match strength" of the target phenotype. The second cumulative value: summing the target phenotype index component with the first cumulative value, yields 2.815 + 4.673 = 7.488, which combines the "total strength" of the target phenotype and all non-target phenotypes. The server uses the second cumulative value as the dividend and the target phenotype index component as the divisor to calculate the sample error ratio. The inverse of the natural logarithm (inverse logarithm) of this ratio is then taken to obtain the first sample error component. Sample error ratio: Second cumulative value / target phenotype index component = 7.488 / 2.815≈2.66. The closer this ratio is to 1, the stronger the target phenotype is relative to the non-target phenotype, and the better the model's ability to distinguish. First sample error component: Take the inverse logarithm of the sample error ratio, that is, -ln(2.66)≈-1.016 (ln is the natural logarithm function). The smaller this value is, the stronger the model's ability to distinguish the target / non-target phenotype of the current sample is (for example, when the ratio = 1, the error component = -ln(1) = 0, which is the ideal state). Through the above process, the server calculates the first sample error component ≈-1.016 for "TR-LC-001". Subsequently, the average loss will be calculated based on this component and the error components of other samples to guide model parameter adjustment.

[0088] In the embodiment of the present invention, the calculation of the first phenotype matching confidence according to the sample tumor phenotype feature vector and the first candidate tumor phenotype feature vector can be implemented through the following example.

[0089] Determining the spatial proximity between the sample tumor phenotype feature vector and the first candidate tumor phenotype feature vector;

[0090] The first phenotype matching confidence is determined based on the ratio of the spatial proximity and the confidence adjustment coefficient.

[0091] In the embodiment of the present invention, for example, the server processes the training reference tumor sample "TR-LC-001" (phenotype is "lung adenocarcinoma") as an example. The sample tumor phenotype feature vector is S (896 dimensions, including the "EGFR mutation-spicule sign image" associated dimension value of 0.85 and the "smoking history-age" joint dimension value of 0.72). The first candidate tumor phenotype feature vector is (Standard vector of lung adenocarcinoma, 896 dimensions, including the correlation dimension value of "EGFR mutation-spicule sign image" 0.90 and the joint dimension value of "smoking history-age" 0.75). The specific process is as follows: Step 1: Determine spatial proximity. The server calculates S and The cosine similarity of is used as the spatial proximity: first calculate the dot product of the two vectors (0.85×0.90+0.72×0.75+...+the product of other 894-dimensional elements), the result is 682.3; then calculate the modulus of S (√( +...+squares of other elements))≈28.5, The modulus of S is ≈ 29.1; spatial proximity = dot product / (S modulus × Modulus) = 682.3 / (28.5 × 29.1) ≈ 0.83 (the closer the value is to 1, the more similar the two vector features are). Step 2: Calculate the confidence level of the first phenotype match. The server uses the pre-trained confidence adjustment coefficient β = 0.85 (to control the confidence scale and avoid numerical overflow) and compares the spatial proximity with β: the first phenotype match confidence = 0.83 / 0.85 ≈ 0.976 (this value reflects the confidence level of the match between the sample and the target phenotype; the closer it is to 1, the higher the match).

[0092] In an embodiment of the present invention, the multimodal basic model is trained through the following steps, which can be implemented through the following examples.

[0093] Obtain multiple multimodal sample instances;

[0094] Dividing the multimodal sample instance into dimensions, selecting a target feature dimension from the candidate feature dimensions obtained by the division, and performing feature masking on the target feature dimension;

[0095] Loading the multimodal sample instance after feature masking into a masked feature reconstruction network to obtain a masked dimension reconstruction value;

[0096] Calculating a second sample error value according to the target feature dimension and the masking dimension reconstruction value of the plurality of multimodal sample instances;

[0097] The masked feature reconstruction network is trained according to the second sample error value to obtain the multimodal basic model.

[0098] In this embodiment of the present invention, for example, a server trains a multimodal basic model for lung cancer. The training data consists of 10,000 multimodal sample instances (each an 896-dimensional feature vector, comprising 128 clinical dimensions, 256 genetic dimensions, and 512 imaging dimensions). The specific process is as follows: The server retrieves preprocessed multimodal sample instances from a hospital database. Each instance is an 896-dimensional feature vector integrating clinical, genetic, and imaging data. For example, instance "MM-0001" contains features such as clinical dimension 3 (25-year smoking history, value 0.25), genetic dimension 138 (EGFR mutation positive, value 1.0), and imaging dimension 641 (right upper lung spiculation, value 0.82). The server then divides the 896-dimensional vector into three candidate feature dimensions based on modality: clinical dimension (1-128), genetic dimension (129-384), and imaging dimension (385-896). A masking priority weight is calculated for each candidate dimension (key features have higher weights, such as 0.9 for the EGFR mutation dimension in the gene dimension, 0.7 for the clinical smoking history dimension, and 0.8 for the imaging spiculation dimension). The two dimensions with the highest priority are selected as target feature dimensions, such as gene dimension 138 (EGFR mutation) and imaging dimension 641 (spiculation) for "MM-0001," and their values ​​are assigned the standard masking identifier -1 (indicating that the dimension is masked). The server loads the masked "MM-0001" instance into the masked feature reconstruction network (consisting of three modality branches and one fusion layer). The network infers the true value of the masked dimension through the unmasked features of other dimensions (such as clinical age 62 years, gene TP53 expression 0.58, and imaging pleural traction sign 0.71), and outputs the reconstructed value of the masked dimension: the reconstructed feature value of gene dimension 138 is 0.88 (true value 1.0), and the reconstruction confidence is 0.92; the reconstructed feature value of image dimension 641 is 0.75 (true value 0.82), and the reconstruction confidence is 0.88. The server calculates the error between the true value and the reconstructed value of the target feature dimension for each sample. Taking "MM-0001" as an example: the gene dimension error = =0.0144, image dimension error = = 0.0049, and the sample error component = 0.0144 + 0.0049 = 0.0193. The error components of 10,000 samples were averaged to obtain a second sample error value of 0.021 (initial value). The server backpropagated the second sample error value using the Adam optimizer (learning rate 5e-5) and adjusted network parameters (such as the gene branch attention weight and the image branch convolution kernel value). For example, increasing the prediction weight of "TP53 expression" for "EGFR mutation" (from 0.3 to 0.45) optimized the reconstruction value of gene dimension 138 from 0.88 to 0.95. After 50 rounds of iterative training, the second sample error value dropped to 0.008 (convergence), and the masked feature reconstruction network was then used as the multimodal base model.

[0099] In an embodiment of the present invention, the masking dimension reconstruction value includes a reconstruction eigenvalue and a reconstruction confidence of the reconstruction eigenvalue;

[0100] The calculation of the second sample error value based on the target feature dimension and the masking dimension reconstruction value of the plurality of multimodal sample instances may be implemented through the following example.

[0101] For each of the multimodal sample instances, determining a target feature mask whose reconstructed feature value is consistent with the target feature dimension among the multiple feature masks of the multimodal sample instance;

[0102] Obtaining a second sample error component of the multimodal sample instance according to the inverse logarithm of the reconstruction confidence of the reconstructed feature value of each target feature mask match;

[0103] The second sample error value is calculated according to the second sample error components of a plurality of the multimodal sample instances.

[0104] In an embodiment of the present invention, for example, taking the server processing the multimodal sample instance "MM-0001" (896-dimensional feature vector, including clinical, genetic, and imaging modalities) as an example, the target feature dimensions of this instance are gene dimension 138 (EGFR mutation status, true value 1.0, classification feature) and imaging dimension 641 (burr sign intensity, true value 0.82, numerical feature). The specific process is as follows: After the server performs feature masking on the target feature dimension of "MM-0001", it is loaded into the masked feature reconstruction network to obtain the masked dimension reconstruction value (including the reconstructed feature value and reconstruction confidence). Gene dimension 138 (EGFR mutation): true value 1.0 (positive), reconstructed feature value 0.92 (close to 1.0, correctly classified, considered "consistent"), reconstruction confidence 0.95 (model confidence that the reconstruction is correct); image dimension 641 (spicule intensity): true value 0.82 (numerical feature, preset error threshold ±0.1), reconstructed feature value 0.85 (within the range of 0.72-0.92, considered "consistent"), reconstruction confidence 0.90 (model confidence that the reconstruction is correct). The server marks the two feature masks with consistent reconstructions as "target feature masks." The server takes the inverse logarithm (inverse logarithm) of the reconstruction confidence of each target feature mask and sums them to obtain the second sample error component of "MM-0001." The inverse logarithm of the gene dimension 138 is: -ln(0.95)≈-(-0.051)=0.051 (higher confidence means smaller inverse logarithm values ​​and lower error); the inverse logarithm of the image dimension 641 is: -ln(0.90)≈-(-0.105)=0.105; the second sample error component is: 0.051+0.105=0.156 (smaller values ​​indicate better reconstruction of the masked dimension of the sample). The server iterates over 10,000 multimodal sample instances and repeats the above steps to calculate the second sample error component for each sample. For example, the error component for "MM-0002" (target features: clinical smoking history and KRAS mutation) is 0.172; the error component for "MM-0003" (target features: radiographic pleural traction sign and ALK fusion) is 0.148. The total error component for all 10,000 samples is 1580.3. The server calculates the second sample error value = total error component / number of samples = 1580.3 / 10000 ≈ 0.158 (initial value). Subsequently, through training, this value is reduced to below 0.08 (convergence), completing the training of the multimodal basic model.

[0105] In the embodiment of the present invention, the selecting of the target feature dimension from the candidate feature dimensions obtained by division and the feature masking of the target feature dimension can be implemented through the following examples.

[0106] For each of the candidate feature dimensions obtained by division, determining the masking priority weight of the candidate feature dimension;

[0107] According to the masking priority weight, a predetermined number of the candidate feature dimensions are selected from the candidate feature dimensions as the target feature dimensions;

[0108] The target feature dimension is assigned a standard masking identifier to perform feature masking on the target feature dimension.

[0109] In an embodiment of the present invention, for example, taking the server processing the multimodal sample instance "MM-0001" (896-dimensional feature vector, including 128 clinical dimensions, 256 genetic dimensions, and 512 imaging dimensions) as an example, two target feature dimensions are predetermined to be masked. The specific process is as follows: the server divides the 896-dimensional vector into three candidate feature dimensions according to the modality: clinical dimension (1-128), genetic dimension (129-384), and imaging dimension (385-896), and calculates the masking priority weight for each candidate dimension (the higher the value, the more critical the dimension is to phenotypic analysis, and the more priority masking training is needed). Clinical dimension: Dimension 3 (25-year smoking history, normalized value 0.25) is strongly correlated with lung cancer and is weighted 0.8; Dimension 5 (CEA level 0.48) is weighted 0.6; weights of the remaining dimensions are ≤ 0.5. Genetic dimension: Dimension 138 (EGFR mutation-positive, value 1.0, driver mutation) is weighted 0.95; Dimension 150 (KRAS mutation-negative, value 0.0) is weighted 0.3; weights of the remaining dimensions are ≤ 0.7. Imaging dimension: Dimension 641 (right upper lung spiculation sign, value 0.82, typical imaging sign) is weighted 0.9; Dimension 720 (pleural traction sign, value 0.71) is weighted 0.85; weights of the remaining dimensions are ≤ 0.6. The server sorts all candidate dimensions by masking priority from high to low and selects the top two as target feature dimensions. The ranking results are: gene dimension 138 (0.95) > imaging dimension 641 (0.9) > imaging dimension 720 (0.85) > clinical dimension 3 (0.8), etc. Therefore, the target feature dimensions are determined to be gene dimension 138 (EGFR mutation) and imaging dimension 641 (spicule sign). The server assigns the original value of the target feature dimension to the standard masking identifier -1 (which is uniformly used to indicate masked dimensions). For example, in "MM-0001": gene dimension 138 original value 1.0 → masked value -1; imaging dimension 641 original value 0.82 → masked value -1; the values ​​of the remaining dimensions remain unchanged (for example, clinical dimension 3 remains 0.25, and imaging dimension 720 remains 0.71). The masked examples are used for subsequent training of the masked feature reconstruction network.

[0110] In the embodiment of the present invention, the multimodal sample instance is obtained through the following process, which can be implemented through the following example.

[0111] Access to multiple heterogeneous medical datasets for phenotyping tasks;

[0112] For each of the heterogeneous medical data sets, performing multimodal data segmentation on the heterogeneous medical data set to obtain a plurality of multimodal data units;

[0113] For each of the multimodal data units, a start marker is embedded in the starting point identifier of the multimodal data unit, and an end marker is embedded in the ending point identifier of the multimodal data unit to obtain the multimodal sample instance.

[0114] In an embodiment of the present invention, for example, a server constructs a multimodal sample instance for lung cancer phenotyping analysis. The specific process is as follows: the server retrieves three types of heterogeneous medical data sets for lung cancer phenotyping analysis from the hospital information system (HIS), the genetic testing platform, and the picture archiving system (PACS): a clinical data set: a structured table containing 10,000 patient records, each record containing 12 clinical indicators such as patient ID (such as "PT-0001"), age, gender, smoking history, and tumor markers (CEA, CYFRA21-1, etc.), stored in CSV format; a gene data set: gene sequencing data in FASTA format and expression profile data in Excel format, containing 10,000 patient records, each record containing patient ID, 200 driver gene mutation status (such as EGFR, KRAS) and corresponding expression level ( Value); Image dataset: DICOM-formatted chest CT image data, containing 3D volumetric data from 10,000 patients (1mm slice thickness, 300 slices / patient). Each image is linked to clinical and genetic data using the patient ID. The server uses "patient ID" as the association key to perform multimodal data segmentation on the three heterogeneous datasets, integrating the clinical, genetic, and imaging data of the same patient into a single multimodal data unit. For example, for patient ID "PT-0001": clinical indicators (age 62, smoking history 25 years, CEA = 4.8 ng / mL, etc.) are extracted from the clinical dataset and encoded into a 128-dimensional clinical sub-vector; genetic data (EGFR L858R mutation positive, TP53 expression log2(TPM) = 5.8, etc.) are extracted from the genetic dataset and encoded into a 256-dimensional genetic sub-vector; chest CT data are extracted from the imaging dataset and feature extracted using 3DResNet into a 512-dimensional imaging sub-vector; these three sub-vectors are concatenated in the order of "clinical-genetic-imaging" to obtain an 896-dimensional multimodal data unit (containing the complete multimodal features of patient ID "PT-0001"). The server embeds preset markers at the start and end points of each multimodal data unit to clearly define the data unit boundaries. For example, for the 896-dimensional data unit of "PT-0001": the start marker "[CLS]" (encoding value 0.0) is embedded in the first dimension (starting point identifier) ​​of the data unit to identify the beginning of the data unit; the end marker "[SEP]" (encoding value 1.0) is embedded in the 896th dimension (end point identifier) ​​of the data unit to identify the end of the data unit.

[0115] The original multimodal feature values ​​for the intermediate 894 dimensions are retained (e.g., the second dimension is the normalized value of 0.62 for clinical age, and the 130th dimension is the encoding value of 1.0 for EGFR mutations). The resulting 896-dimensional vector is the multimodal sample instance, denoted as "MM-0001." Through this process, the server integrates heterogeneous medical data into a structured multimodal sample instance for subsequent training of the multimodal foundation model.

[0116] In the embodiment of the present invention, the tumor sample feature vector is obtained through the following process, which can be implemented through the following examples.

[0117] Acquire multimodal clinical input data;

[0118] The multimodal clinical input data is encoded in terms of feature dimensions to obtain the tumor sample feature vector.

[0119] In an embodiment of the present invention, for example, taking the server processing a suspected lung cancer sample as an example, the specific process is as follows: the server retrieves the multimodal clinical input data of the sample from the hospital's heterogeneous system: Clinical data: Obtain patient structured information from the HIS system, including 12 indicators such as age 62 years, male gender, 25-year smoking history, CEA=4.8ng / mL, CYFRA21-1=3.5ng / mL, and no history of hypertension / diabetes; Genetic data: Obtain sequencing reports from the genetic testing platform, including EGFR L858R mutation-positive, ALK fusion-negative, KRASG12C mutation-negative, and expression data of 200 driver genes such as EGFR (log2(TPM)=7.2) and TP53 (log2(TPM)=5.8); Imaging data: Obtain chest enhanced CT three-dimensional volume data (1mm layer thickness, 300 layers) from the PACS system, including multi-window images of lung windows and mediastinal windows. The server encodes and concatenates the multimodal data: Clinical data encoding: 12 indicators are quantized into 128-dimensional vectors. For example, age 62 is normalized to 0.62 (first dimension), male sex is one-hot encoded as [1,0] (second-third dimensions), and a smoking history of 25 years is normalized to 0.25 (fourth dimension). Genetic data encoding: Mutation status (positive = 1, negative = 0) and expression level (normalized to 0-1) are encoded into 256-dimensional vectors. For example, a positive EGFR mutation corresponds to a value of 1.0 in the 10th dimension, and an EGFR expression level of 7.2 is normalized to 0.72 (the 11th dimension). Imaging data encoding: CT image features are extracted using a pre-trained 3DResNet, outputting a 512-dimensional vector. For example, "spicule sign in the right upper lung" corresponds to a value of 0.75 in the 256th dimension. Concatenation: The 128-dimensional clinical, 256-dimensional genetic, and 512-dimensional imaging vectors are concatenated in sequence to obtain an 896-dimensional tumor sample feature vector. Furthermore, all medical data involved in the acquisition and subsequent analysis of tumor sample feature vectors has undergone strict privacy protection. Before providing data to the server, the hospital information system (HIS), genetic testing platform, and picture archiving system (PACS) completely removes patient personal identifying information (such as name, ID number, contact information, and date of birth). Only an anonymous sample number (such as "PT-LC-2023-112") is retained as a unique identifier, which is not mapped to the patient's actual identity. Privacy-sensitive fields in clinical data (such as home address and medical record number) have been completely removed, retaining only medical indicators relevant to phenotypic analysis (such as age, smoking history, and tumor marker levels). Genetic data only provides mutation status (positive / negative) and normalized expression (without raw sequencing reads). In the DICOM file metadata for imaging data, privacy-sensitive fields such as patient name and hospital ID have been replaced with "Anonymous" to ensure that patient identities cannot be reversed.All data transmission uses an encrypted protocol (HTTPS+AES-256). Data access requires multi-factor authentication (such as permission approval + dynamic password) and is only used for the training and verification of tumor phenotyping analysis models. It will not be disclosed to any third party or used for other purposes.

[0120] In an embodiment of the present invention, the multilayer perceptron includes a feature space mapping component, a nonlinear operator, and a feature space inverse mapping component;

[0121] The multi-layer perceptron according to the tumor phenotype analysis model performs nonlinear transformation on the primary tumor sample feature vector to obtain a high-dimensional abstract feature vector, which can be implemented through the following example.

[0122] Performing a first feature space mapping on the primary tumor sample feature vector according to the feature space mapping component to obtain an intermediate tumor phenotype feature vector;

[0123] performing a nonlinear transformation operation on the intermediate tumor phenotype feature vector according to the nonlinear operator to obtain a tumor phenotype feature vector to be processed;

[0124] A second feature space mapping is performed on the tumor phenotype feature vector to be processed according to the feature space inverse mapping component to obtain the high-dimensional abstract feature vector, wherein the high-dimensional abstract feature vector and the primary tumor sample feature vector meet a feature consistency constraint.

[0125] In the embodiment of the present invention, for example, the server processes the primary tumor sample feature vector of the lung cancer sample numbered "PT-LC-2023-112". (896 dimensions, including clinical-gene-imaging cross-modal correlation features) as an example, the multi-layer perceptron (MLP) includes a feature space mapping component, a nonlinear operator, and a feature space inverse mapping component. The specific process is as follows: the server calls the feature space mapping component of the MLP (implemented by a fully connected layer, the weight matrix , bias vector ), the 896-dimensional primary tumor sample feature vector Mapping to a higher-dimensional space, we obtain the intermediate tumor phenotype feature vector M (1024 dimensions). Key feature dimensions: 10th dimension ("EGFR mutation-age" association feature, value 0.90), 256th dimension ("spicule sign image-smoking history" association feature, value 0.82), 50th dimension ("TP53 expression-pleural traction sign" association feature, value 0.52); Mapping process: through linear transformation , expand low-dimensional features to high-dimensional space to capture more complex feature interactions. For example, 0.90 Warp of the 10th Dimension After being mapped with a weight of 0.85, the corresponding feature vector in M ​​contributes to the 120th dimension (value 0.90 × 0.85 + contributions from other dimensions ≈ 1.12). The feature vector 0.82 in the 256th dimension, after being mapped with a weight of 0.92, contributes to the 320th dimension of M (value 0.82 × 0.92 + contributions from other dimensions ≈ 1.05). The output is a 1024-dimensional M, where the features in the higher-dimensional space (e.g., the 120th and 320th dimensions of M) more finely capture the interaction patterns of the original low-dimensional features. The server performs a nonlinear transformation on the intermediate tumor phenotypic feature vector M using a nonlinear operator (LeakyReLU activation function with a negative slope of 0.01) to obtain the processed tumor phenotypic feature vector T (1024 dimensions). Nonlinear enhancement: LeakyReLU retains positive features in M ​​(e.g., M 120th dimension 1.12 → T 120th dimension 1.12), and assigns a small slope to negative features (e.g., M 500th dimension -0.3 → T 500th dimension -0.3×0.01=-0.003), avoiding neuron "death" and enhancing feature nonlinear expression; Noise suppression: Weakly correlated noise features in M ​​(e.g., 0.15 in the 700th dimension, corresponding to the redundant interaction of "sex-irrelevant genes") still maintain a low amplitude (T 700th dimension 0.15) after nonlinear transformation, while the high amplitude of key features (e.g., 120th and 320th dimensions) is retained, achieving implicit screening of feature importance; Output: 1024-dimensional T, the feature distribution is more consistent with the nonlinear correlation law of tumor phenotypes. The server calls the feature space inverse mapping component (implemented by a 1-layer fully connected layer, weight matrix , bias vector ), the tumor phenotype feature vector T (1024 dimensions) to be processed is mapped back to the original feature space to obtain a high-dimensional abstract feature vector Z of 896 dimensions, and Z is consistent with the primary tumor sample feature vector Satisfy feature consistency constraints (dimensional consistency, key feature trend preservation). Inverse mapping process: through linear transformation , compressing the high-dimensional nonlinear features back to 896 dimensions. For example, the 120th dimension (EGFR-related high-dimensional feature 1.12) is compressed by After the corresponding weight (0.88) is mapped, it contributes to the 10th dimension of Z (value 1.12×0.88+contribution of other dimensions ≈ 0.95); after the 320th dimension of T (burr-related high-dimensional feature 1.05) is mapped with a weight (0.84), it contributes to the 256th dimension of Z (value 1.05×0.84+contribution of other dimensions ≈ 0.88); feature consistency constraint: Z and The dimensions are all 896, and the changing trends of key features are consistent. 0.90 in the 10th dimension → 0.95 in the 10th dimension (augmented), 0.82 in the 256th dimension → 0.88 in the 256th dimension (augmented), 0.52 in the 50th dimension → 0.65 in the 50th dimension (enhancement), ensuring that the abstracted features still retain the core phenotypic information of the original sample; Output: 896-dimensional high-dimensional abstract feature vector Z, whose feature expression is more abstract and the key association is more significant, which is used for subsequent Through the above process, the server uses MLP to achieve nonlinear abstraction of the primary feature vector, while retaining the core information and improving the discriminative ability of the features.

[0126] The embodiment of the present invention provides a computer device 100, which includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device 100 executes the aforementioned tumor phenotype analysis method based on the pathogenicity database and machine learning assistance. Figure 2 As shown, Figure 2 This is a block diagram of the structure of a computer device 100 provided in an embodiment of the present invention. Computer device 100 includes a memory 111, a processor 112, and a communication unit 113. To enable data transmission or exchange, memory 111, processor 112, and communication unit 113 are electrically connected to each other, directly or indirectly. For example, these components can be electrically connected via one or more communication buses or signal lines.

[0127] For illustrative purposes, the foregoing description has been made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the present disclosure to the precise forms disclosed. Numerous modifications and variations are possible in light of the above teachings. These embodiments have been selected and described in order to best illustrate the principles of the present disclosure and its practical application, thereby enabling those skilled in the art to best utilize the present disclosure and to utilize various embodiments with various modifications as appropriate for the specific application contemplated.

Claims

1. A tumor phenotyping method based on a pathogenic database and assisted by machine learning, characterized in that: include: Obtain tumor sample feature vector; Inputting the tumor sample feature vector into the clinical feature path, gene feature path, and imaging feature path in the multimodal feature fusion network of the tumor phenotyping model to obtain a clinical feature path vector, a gene feature path vector, and an imaging feature path vector, respectively; Performing long-distance dependency modeling on the clinical feature pathway vector and the gene feature pathway vector to obtain a first dependency feature graph; wherein the first dependency feature graph is used to indicate the strength of association between each first feature in the clinical feature pathway vector and each second feature in the gene feature pathway vector; Obtaining a feature association strength feature graph, wherein the feature association strength feature graph is used to indicate a feature association distance between each first feature in the clinical feature pathway vector and each second feature in the gene feature pathway vector, and each bias feature in the feature association strength feature graph is positively correlated with each dependent feature in the first dependent feature graph; Determining a second dependency feature graph based on the first dependency feature graph and the feature association strength feature graph, and determining a target tumor phenotype feature vector based on the image feature path vector and the second dependency feature graph; Acquiring a tumor phenotype database, wherein the tumor phenotype database includes a plurality of candidate tumor phenotypes and a candidate tumor phenotype feature vector of each candidate tumor phenotype; For each of the candidate tumor phenotypes, calculating a phenotype matching confidence between the candidate tumor phenotype feature vector of the candidate tumor phenotype and the target tumor phenotype feature vector, and determining the target tumor phenotype based on the highest phenotype matching confidence; For any feature dimension in the tumor sample feature vector, the feature index of the first feature matched by the feature dimension in the clinical feature pathway vector is the same as the feature index of the second feature matched by the feature dimension in the gene feature pathway vector, wherein the feature index is used to indicate the position information of the feature dimension in the tumor sample feature vector; The feature association strength feature map is obtained by the following process, including: Determining a quantitative feature value of the bias feature of the primary-order association strength feature graph according to an index distance between a feature index of each first feature in the clinical feature pathway vector and a feature index of each second feature in the gene feature pathway vector; The feature association strength feature map is determined according to a weighted multiplication result of a learnable scaling coefficient and the primary-order association strength feature map.

2. The method according to claim 1, characterized in that Determining the target tumor phenotype feature vector according to the image feature path vector and the second dependency feature graph includes: Obtaining a long-distance dependency modeling result according to the image feature path vector and the second dependency feature graph; determining a primary tumor sample feature vector according to the long-distance dependency modeling result and the skip connection output of the tumor sample feature vector; Performing a nonlinear transformation on the primary tumor sample feature vector according to the multi-layer perceptron of the tumor phenotype analysis model to obtain a high-dimensional abstract feature vector; Obtaining a phenotypic feature optimization vector based on the sum of the skip connection outputs of the high-dimensional abstract feature vector and the primary tumor sample feature vector, and increasing the number of optimization rounds by one, with the optimization rounds counting starting from the starting round; The phenotypic feature optimization vector is used as the tumor sample feature vector, and the step of inputting the tumor sample feature vector into the clinical feature pathway, gene feature pathway, and imaging feature pathway in the multimodal feature fusion network of the tumor phenotyping analysis model is repeatedly performed until the optimization rounds reach a preset number of rounds, and the phenotypic feature optimization vector is used as the target tumor phenotypic feature vector.

3. The method according to claim 1, characterized in that The tumor phenotyping model is obtained by fine-tuning the multimodal base model in the following ways, including: For each training reference tumor sample, construct a target phenotype instance and a non-target phenotype instance, wherein the reference tumor phenotype data in the target phenotype instance matches the training reference tumor sample, and the reference tumor phenotype data in the non-target phenotype instance does not match the training reference tumor sample; For each of the training reference tumor samples, loading the target phenotype instance that matches the training reference tumor sample into the multimodal base model, obtaining a sample tumor phenotype feature vector of the training reference tumor sample and a first candidate tumor phenotype feature vector of the reference tumor phenotype data in the target phenotype instance, and calculating a first phenotype matching confidence based on the sample tumor phenotype feature vector and the first candidate tumor phenotype feature vector; loading the non-target phenotype instance that matches the training reference tumor sample into the multimodal base model, obtaining a sample tumor phenotype feature vector of the training reference tumor sample and a second candidate tumor phenotype feature vector of the reference tumor phenotype data in the non-target phenotype instance, and calculating a second phenotype matching confidence based on the sample tumor phenotype feature vector and the second candidate tumor phenotype feature vector; The network parameters of the multimodal basic model are adjusted according to the first phenotype matching confidence and the second phenotype matching confidence of the plurality of training reference tumor samples to obtain the tumor phenotype analysis model.

4. The method according to claim 3, characterized in that The constructing of target phenotype instances and non-target phenotype instances for each training reference tumor sample includes: Obtaining a sample-phenotype mapping table, wherein the sample-phenotype mapping table includes a plurality of reference tumor samples and reference tumor phenotype data that are matched one-to-one with the plurality of reference tumor samples; For each training reference tumor sample in the sample-phenotype mapping table, the target phenotype instance is constructed based on the training reference tumor sample and the reference tumor phenotype data matching the training reference tumor sample, and the non-target phenotype instance is constructed based on the training reference tumor sample and the reference tumor phenotype data matching other reference tumor samples other than the training reference tumor sample.

5. The method according to claim 3, characterized in that The adjusting of the network parameters of the multimodal basic model according to the first phenotype matching confidence and the second phenotype matching confidence of the plurality of training reference tumor samples includes: using a second phenotype matching confidence of a non-target phenotype instance for each of the training reference tumor samples as a non-target phenotype index component of an exponential transformation of an Euler number, and summing a plurality of the non-target phenotype index components to obtain a first accumulated value; Summing a target phenotypic index component of an exponential transformation using the first phenotypic matching confidence as an Euler number and the first accumulated value to obtain a second accumulated value; Taking the inverse logarithm of a sample error ratio of a target phenotypic index component of an exponential transformation using the second accumulated value as the dividend and the first phenotype match confidence as the Euler number as the divisor, to obtain a first sample error component of the training reference tumor sample; taking an average loss component of the plurality of first sample error components as a first sample error value; The network parameters of the multimodal basic model are adjusted according to the first sample error value.

6. The method according to claim 3, characterized in that The calculating a first phenotype matching confidence level according to the sample tumor phenotype feature vector and the first candidate tumor phenotype feature vector includes: Determining the spatial proximity between the sample tumor phenotype feature vector and the first candidate tumor phenotype feature vector; The first phenotype matching confidence is determined based on the ratio of the spatial proximity and the confidence adjustment coefficient.

7. The method according to claim 3, characterized in that The multimodal basic model is trained by the following steps, including: Obtain multiple multimodal sample instances; Dividing the multimodal sample instance into dimensions, and determining the masking priority weight of the candidate feature dimension obtained by each division; According to the masking priority weight, a predetermined number of the candidate feature dimensions are selected as target feature dimensions from the candidate feature dimensions; Assigning the target feature dimension as a standard masking identifier to perform feature masking on the target feature dimension; Loading the multimodal sample instance after feature masking into a masked feature reconstruction network to obtain a masked dimension reconstruction value; the masked dimension reconstruction value includes a reconstructed feature value and a reconstruction confidence of the reconstructed feature value; For each of the multimodal sample instances, determining a target feature mask whose reconstructed feature value is consistent with the target feature dimension among the multiple feature masks of the multimodal sample instance; Obtaining a second sample error component of the multimodal sample instance according to the inverse logarithm of the reconstruction confidence of the reconstructed feature value of each target feature mask match; Calculating a second sample error value based on the second sample error components of the plurality of multimodal sample instances; The masked feature reconstruction network is trained according to the second sample error value to obtain the multimodal basic model.

8. The method according to claim 7, characterized in that The multimodal sample instance is obtained through the following process, including: Access to multiple heterogeneous medical datasets for phenotyping tasks; For each of the heterogeneous medical data sets, performing multimodal data segmentation on the heterogeneous medical data set to obtain a plurality of multimodal data units; For each of the multimodal data units, a start marker is embedded in the starting point identifier of the multimodal data unit, and an end marker is embedded in the ending point identifier of the multimodal data unit to obtain the multimodal sample instance.

9. A server system, characterized in that: The method comprises a server, wherein the server is configured to execute the method according to any one of claims 1 to 8.