Tumor identification method based on multi-modal feature fusion

By employing a tumor identification method based on multimodal feature fusion, the problem of feature distribution shift caused by heterogeneous multicenter devices is solved, achieving tumor identification with high accuracy and clinical interpretability. This method is adapted to the deployment of edge medical devices and improves the detection rate of early-stage small tumors.

CN121765644APending Publication Date: 2026-03-31SINONEEDLE INTELLIGENCE TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-31
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing tumor identification technologies suffer from problems such as lack of multimodal feature calibration, insufficient semantic association, poor generalization, and lack of clinical interpretability, especially in the case of heterogeneous multicenter equipment, feature distribution shift and sharp drop in recognition performance.

Method used

An event-driven sensor array is used to synchronously acquire visual structure and electrical impedance data. The acquisition parameters are adjusted through a biomimetic spatiotemporal perception algorithm. The structured and characteristic feature vectors are extracted by a heterogeneous twin coding engine, and dynamic weighted fusion and semantic association are performed. Multi-scale dynamic perception cell array, global topological association modeling and cross-modal calibration are used to generate comprehensive feature vectors. Semantic association is enhanced through attention mechanism and InfoNCE contrastive loss, which is suitable for deployment of edge medical devices.

Benefits of technology

It achieves a unified dimension and distribution of multimodal features, improves the accuracy and stability of tumor identification, reduces computing power consumption, meets the interpretability and generalization requirements of clinical diagnosis, and in particular improves the detection rate of early-stage small tumors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121765644A_ABST
    Figure CN121765644A_ABST
Patent Text Reader

Abstract

The invention provides a tumor recognition method based on multi-modal feature fusion, relates to the technical field of tumor recognition, and is used for solving the problems of missing multi-modal feature calibration, insufficient semantic association, poor generalization and insufficient clinical interpretability in the prior art. The method comprises the following steps: firstly, synchronously acquiring visual structure and electrical impedance bimodal data through an event-driven sensor array, and completing preprocessing through pulse coding, STDP rule optimization and quality verification; then, respectively extracting structured feature vectors of a tumor visual structure class and characteristic feature vectors of a numerical attribute class by a heterogeneous twin coding engine, and unifying dimensions and distribution through processes such as multi-scale dynamic perception and cross-modal calibration; then based on an attention mechanism, InfoNCE contrast loss and minority class weight gain, dynamic weighted fusion and semantic association enhancement are performed on the bimodal features, and a comprehensive feature vector is generated; and finally, through clinical logic adaptation and multi-center deviation correction, outputting a tumor benign and malignant identification result through a full-connection classifier, and synchronously generating a clinical interpretable report containing key features and weights. According to the method, the complementary advantages of bimodal information are effectively integrated, the problems of heterogeneous multi-center equipment, unbalanced samples and the like are solved, the tumor recognition accuracy and the early-stage tiny tumor detection rate are improved, the computing power consumption is reduced, edge medical equipment deployment is adapted, and the requirements for low misjudgment and traceability of clinical diagnosis are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of tumor identification technology, and in particular to a tumor identification method based on multimodal feature fusion. Background Technology

[0002] The core requirement for tumor identification is to capture lesion characteristics and distinguish between normal tissue and lesions. However, due to the deficiencies in information dimensions and fusion logic, the existing technology system has always been unable to meet the clinical diagnostic criteria of low misjudgment.

[0003] Current tumor identification technologies mainly rely on two types of single-modal information, both of which suffer from significant information loss. The visual structure modality-dominated method captures visual features such as the shape, boundary, and spatial distribution of tumors through edge detection, texture extraction, and other algorithms, but cannot correlate them with the biochemical characteristics of tumors. On the other hand, the numerical characteristic modality-dominated method achieves identification by analyzing outliers in quantitative indicators such as tumor markers and tissue impedance, but lacks spatial localization and structural support for lesions.

[0004] Among existing multimodal fusion technologies, patent CN120852783A discloses "a method and system for identifying and locating neural tumors based on image recognition." This method generates results through morphological features to achieve the identification and location of neural tumors; however, this method mainly relies on a single modality and does not fully utilize other types of features. In addition, existing multimodal fusion technologies also have many shortcomings, such as feature distribution shifts caused by the heterogeneity of multi-center devices, and a sharp drop in recognition performance when modalities are missing; in particular, the fusion process is a black box operation, lacking clinical interpretability; sample imbalance leads to a high rate of missed diagnoses for a minority of samples, such as early-stage small tumors; and the contradiction between high-dimensional feature redundancy and computational consumption is prominent, making it difficult to adapt to the deployment needs of edge medical devices.

[0005] Therefore, there is an urgent need to design a tumor identification method based on multimodal feature fusion, which can solve the problems of the existing technology while retaining the complementary advantages of multimodal information. Summary of the Invention

[0006] The main objective of this invention is to provide a tumor identification method based on multimodal feature fusion, which solves the problems of missing multimodal feature calibration, insufficient semantic association, poor generalization, and lack of clinical interpretability in the prior art.

[0007] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is: a tumor identification method based on multimodal feature fusion, the method comprising the following steps: S1. Collect visual structural data and electrical impedance structural data of the tumor to construct bimodal data and perform preprocessing; S2. Design a heterogeneous twin coding engine to extract structured feature vectors. With characteristic feature vectors And calibrate the uniform dimension and distribution; S3, Based on Structured Feature Vectors With characteristic feature vectors By dynamically weighted fusion and semantic association enhancement, a comprehensive feature vector is generated. ; S4. For the comprehensive feature vector The tumor is classified and processed, and the tumor identification results and clinically interpretable reports are output.

[0008] In the preferred embodiment, in step S1: The dual-modal data is acquired using a biomimetic spatiotemporal perception and heterogeneous data stream modeling acquisition system. This system consists of an event-driven sensor array, which includes a visual structure data acquisition unit and a resistive structure data acquisition unit. The two units achieve time alignment of data acquisition through a synchronous triggering mechanism. The visual structure data acquisition unit uses an image sensor, which is pre-set to a sampling frequency and outputs grayscale image data of a specified resolution to capture visual information such as the shape, boundary, spatial distribution and local texture of the tumor. The electrical impedance structure data acquisition unit uses an impedance sensor, with pre-set sampling accuracy and sampling frequency consistent with the visual acquisition unit, to synchronously acquire the electrical impedance values ​​of tumor tissue, which are used to quantify the physical properties and biochemical characteristics of the tumor. During the data acquisition process, the dynamic changes of the acquisition environment are perceived in real time through a biomimetic spatiotemporal perception algorithm. When slight patient movement or fluctuations in ambient light are detected, the sensor gain parameters and sampling trigger timing are adjusted to maintain the stability of synchronous data acquisition.

[0009] In the preferred embodiment, the preprocessing of the dual-modal data in step S1 further includes the following steps: S11. Perform pulse coding processing on the acquired continuous data to convert the continuous grayscale values ​​and impedance values ​​into a pulse sequence of 0 or 1, wherein: For visual structural data, calculate the change in grayscale value of each pixel over time. When the change exceeds a preset threshold, output pulse 1; otherwise, output pulse 0. For electrical impedance structural data, the same triggering logic as for visual structural data is used to calculate the time series change of the electrical impedance value. When the change exceeds the threshold, pulse 1 is output, otherwise pulse 0 is output. S12. The pulse timing-dependent plasticity learning rule (STDP) is introduced to adaptively optimize the pulse sequence. The synaptic weights are dynamically adjusted according to the timing relationship of pulse firing, preserving key features in the data and filtering redundant noise. The synaptic weight update formula is as follows: When the presynaptic neuron fires a pulse earlier than the postsynaptic neuron... (1); When the presynaptic neuron fires a pulse later than the postsynaptic neuron... (2); in, This refers to the update amount of synaptic weights; A + and A - τ represents the amplitude coefficients for synaptic weight enhancement and weakening, respectively; Δt is the time difference between the firing of the presynaptic neuron and the postsynaptic neuron; τ is the amplitude coefficients for synaptic weight enhancement and weakening, respectively. + and τ - These are the time constants for weight enhancement and weight reduction, respectively; This learning rule is used to strengthen pulse firing patterns that are strongly correlated with tumor features in pulse sequences, suppress irrelevant noise pulses, and generate a pulse sequence dataset. S13. Perform quality checks on the dual-mode pulse sequence, including data integrity, time synchronization, and signal-to-noise ratio, where: Data integrity requires that the length of the valid pulse sequence be no less than a preset percentage of the acquisition duration; if data is lost, re-acquisition will be triggered. Time synchronization is achieved by calculating the timestamp difference between the dual-mode pulse sequences and comparing it with a preset time difference, and then performing error correction. The signal-to-noise ratio (SNR) is determined by calculating the ratio of signal pulses to noise pulses. If the SNR is lower than the target decibel, the pulse coding threshold and STDP learning rule parameters are readjusted. After passing the verification, the preprocessed bimodal raw dataset is output as the input data for the heterogeneous twin coding engine in step S2.

[0010] In the preferred embodiment, step S2 further includes the following steps: S21. A multi-scale dynamic perceptual cell array is used for feature extraction, global topological association modeling, semantic flow field alignment, and dimensional anchoring calibration to process visual structural information such as tumor morphology and distribution, and output a fixed-dimensional structured feature vector. ; S22. Numerical structural information such as the quantification of tumor physical properties is processed using redundant index screening, semantic embedding generation, and cross-modal calibration adaptation procedures, outputting a structured feature vector. Characteristic feature vectors with consistent dimensions and distribution .

[0011] In the preferred embodiment, step S21 further includes: S211, a multi-scale dynamic sensing cell array is used to capture fine-grained details of tumor edges and local textures in the information, and to realize parallel sensing and fusion of multi-scale information; S212. Global topological association modeling constructs a topological association network between local features, rather than a simple spatial association mapping, to achieve structured dimensional integration from local features to global structural features. S213. By generating a dynamic semantic flow field, the data collected by different devices are mapped to a unified feature space for alignment, and the global structured feature set is aligned. S214. Dimensional Anchoring Calibration: By using a preset dimensional anchor point matrix and combining it with cell adaptive transformation, global topological features are accurately anchored to a preset dimension. Simultaneously, anchor point constraints eliminate scale differences and distribution offsets, outputting a structured feature vector with fixed dimensions and a regular distribution. .

[0012] In the preferred embodiment, step S211 specifically includes: S2111, the multi-scale dynamic sensing cell array is equipped with a multi-scale cell library containing cells of different scales. Each cell has independent scale adjustment capabilities and is dynamically activated based on the local features of the input visual data. The adaptive calculation of cell scale is based on the tumor edge gradient. With texture density gradient The calculation formula is: (3); in, The scale of the currently activated cells; This serves as the initial reference scale; This is the gradient decay coefficient, used to adjust the sensitivity of the scale to gradient changes; This is set as the minimum cell size threshold to avoid a surge in computational load due to excessively small cell sizes. Tumor margin gradient The Sobel operator is used to perform horizontal and vertical edge detection on the visual grayscale image, and then the edge gradient value of each pixel is calculated by taking the square root of the sum of squares; texture density gradient is obtained. By calculating the gray-level co-occurrence matrix, the distribution pattern of gray-level values ​​of adjacent pixels in the image is statistically analyzed, thereby obtaining the gradient change of texture density; for each local region in the image, its edge gradient is calculated. With texture density gradient The average value is substituted into formula (3) to obtain the cell scale that fits the region, and cells with matching scale in the cell library are dynamically activated. The sliding step size of the cell array adopts a gradient-following step size, that is, the step size is automatically densified in key areas with high edge gradients, such as tumor edges, while the step size is adaptively sparse in uniform areas with low edge gradients. The step size calculation formula is: (4); in, This is the current sliding step size; Used as the reference step size; This represents the edge gradient value of the current region. The maximum edge gradient value of the entire image is obtained by statistically analyzing the edge gradient distribution of the training set images. Increase the sampling density of the cell array in key areas such as the tumor margin to further capture fine-grained details; S2112. For each activated cell, an edge-enhancing attention mask is embedded by assigning weights to each pixel within the cell's coverage area, with pixels having higher edge gradient values ​​receiving greater weights. This enhances the response intensity of tumor edge features within the cell's coverage area and suppresses background interference from normal tissue. Based on this mask, the edge-sensitive feature response value of each cell is calculated. The formula is: (5); in, The area covered by the current cell, that is, the set of pixels contained in the cell; Coordinates within the coverage area The pixel grayscale value of the location; coordinates Location-based attention mask weights; coordinates Edge gradient values ​​at the location; This is an exponential enhancement term for the edge gradient, used to further amplify the response of edge features; This formula enhances the tumor edge features within the cell coverage area and suppresses the response of normal tissue areas, enabling each activated cell to output a corresponding edge-sensitive feature response value. ; S2113. By using the cell feature self-organizing fusion rule, the response values ​​of all activated cells are collaboratively integrated to generate a multi-scale fused local feature map. The fusion formula is: (6); in, The total number of activated cells; For the first The contribution weight of each cell is determined by the cell scale and the gradient matching degree. The higher the cell scale and the gradient matching degree of the current region, the greater the weight. The sum of the weights of all cells is 1. For the first Edge-sensitive feature response value of each cell; For feature enhancement operations, To adapt to the first A cell-scale convolutional kernel, with the kernel size consistent with the cell scale, is used to extract local feature patterns in the cell response values. Multi-scale fused local feature map obtained after fusion Each pixel carries cellular-scale information and edge-sensitive information, which can fully characterize the fine-grained details of the tumor's edge, texture, etc.

[0013] In the preferred embodiment, step S212 specifically includes: S2121. Calculate the semantic similarity of each local feature using the cosine similarity formula, and group local features with semantic similarity higher than a preset threshold into the same category; then, perform stratification based on the cell scale corresponding to the features, further clustering local features with similar scales within the same semantic category to form a feature cluster set. The set expression is ,in The number of feature clusters; S2122, Introducing the topological correlation matrix of feature clusters The spatial topological relationships and semantic dependencies among the feature clusters are quantitatively analyzed. The formula for calculating the matrix elements is as follows: (7); in, For feature clusters With feature clusters The strength of the topological association between them, with a value ranging from 0 to 1. The larger the value, the stronger the association between the two. For feature clusters and The semantic similarity is obtained by calculating the cosine similarity of the average feature vectors of the two clusters of features; The spatial center distance between two clusters of features is obtained by calculating the Euclidean distance between the geometric centers of the two clusters of features in the feature map. This is the sum of the topological association strengths of all feature clusters, used to normalize the association strengths. Based on topological correlation matrix Construct a global topological association network, where nodes in the network are feature clusters and the weights of edges represent the corresponding topological association strengths; S2123, Based on Topological Independence Matrix The fused features were purified using feature distillation technology to generate a globally structured feature set that can fully characterize the overall topology of the tumor. The distillation formula is: (8); in, Topological Incidence Matrix The Row vectors containing the association weights of all feature clusters; For feature clusters The global average pooling result; Global structured feature set for The weighted summation of the feature clusters, with each element carrying topological association information between the feature clusters, fully characterizes the overall structural features of the tumor.

[0014] In the preferred embodiment, step S213 specifically includes: S2131. A semantic similarity matrix is ​​obtained by calculating the cosine similarity between each feature vector in the global structured feature set and the preset standard feature vector; a spatial distance matrix is ​​obtained by calculating the Euclidean distance between each feature vector in the feature space; and then a dynamic semantic flow field is generated based on these two matrices. Each vector in the flow field represents the mapping direction and magnitude of the feature point from the original space to the target space. S2132, Global structured feature set Each feature point is substituted into the semantic flow field, and spatial mapping transformation is performed to align the heterogeneous features of multi-center devices. The alignment error is then evaluated by calculating the average Euclidean distance between the mapped features and the standard features. If the standard is not met, the generation parameters of the semantic flow field are readjusted until the accuracy requirements are met.

[0015] In the preferred embodiment, step S214 specifically includes: S2141. Set the dimension anchor matrix The matrix dimension is ,in To preset the output dimension, To determine the feature dimensions of the aligned global structured feature set, a cell-adaptive linear and nonlinear hybrid transformation is used to accurately map the global structured features to the dimension space defined by the anchor points. The mapping formula is as follows: (9); in, This is for the initial mapping of feature vectors; The dimension anchor matrix is ​​obtained through pre-training on a standard dataset; This is the aligned global structured feature set; The anchor point bias vector; It is a non-linear activation function used to introduce non-linear feature transformations and enhance the expressive power of features; It is an adaptive weight matrix, which is adaptively learned through the training process and used to adjust the distribution of the mapped features; S2142, Initial mapping of feature vectors Double normalization, involving both anchor point constraint and distribution alignment, is performed to confine the eigenvalues ​​to a uniform interval. The formula for calculating the anchor point constraint is as follows: (10); in, The feature vector after anchor point constraint; Anchor point matrix of dimensions The minimum value in; Anchor point matrix of dimensions The maximum value in; This formula maps each element value of the feature vector to the range of 0 to 1, eliminating scale differences between different samples. Batch normalization is used for distribution alignment to eliminate distribution shifts between different samples. The calculation formula is as follows: (11); in, The feature vectors after distribution alignment; This represents the mean of the features of the current batch of samples; The variance of the features of the current batch of samples; To prevent constraint values ​​with a denominator of 0; S2143, Aligning the feature vectors after distribution Outputting in one-dimensional vector form yields structured feature vectors. The eigenvalues ​​are distributed in the interval between 0 and 1.

[0016] In the preferred embodiment, step S22 specifically includes: S221. Redundancy index screening is based on the correlation decay model to retain core quantitative indicators that are highly correlated with tumors and independent, and to construct a dynamic decay correlation and redundancy cluster purification screening layer to filter redundant data. S222, Semantic Embedding Generation: Through index semantic embedding and high-dimensional feature generation layer, the filtered discrete indexes are transformed into high-dimensional feature vectors carrying semantic information. S223. Cross-modal calibration adaptation involves constructing an anchor-driven dimension alignment layer and a cross-modal calibration module, outputting structured feature vectors. Dimensionality and distribution consistent characteristic feature vectors This ensures the homology fusion of multimodal features.

[0017] In the preferred embodiment, step S221 specifically includes: S2211. Construct an assessment model for the correlation between indicators and tumor dynamic attenuation. Based on tumor type and stage, dynamically adjust the attenuation coefficient to quantify the true correlation between each indicator and the tumor diagnosis result. For each quantified indicator... The numerical sequence and the tumor determination result sequence are dynamically attenuated and weighted to calculate the correlation assessment value. Its attenuation coefficient is dynamically adjusted according to the time distance between the indicator and the tumor. The greater the time distance, the smaller the attenuation coefficient, thus increasing the weight of recent indicators. S2212. Based on the index correlation distribution entropy of the input samples, dynamically adjust the correlation screening threshold; wherein the correlation distribution entropy is obtained by calculating the information entropy of the correlation evaluation values ​​of all indicators. The larger the entropy value, the more dispersed the correlation distribution of the indicators, and the lower the screening threshold accordingly; the smaller the entropy value, the more concentrated the distribution, and the higher the screening threshold accordingly; further, retain indicators with correlation evaluation values ​​higher than the dynamic threshold, and remove redundant indicators with correlation values ​​lower than the threshold, and obtain a candidate indicator set after preliminary screening; S2213. The correlation between indicators is calculated using the Pearson correlation coefficient. Indicators with correlation coefficients higher than a preset threshold are clustered into a redundant cluster. For each redundant cluster, the indicator with the highest correlation evaluation value and the largest information entropy is extracted as the core indicator. The information entropy is obtained by calculating the probability distribution of the indicator values. The highly correlated indicators that still exist in the candidate indicator set are purified into redundant clusters to make the core indicators have strong distinguishing ability and information richness. Finally, a set of core quantitative indicators that are strongly correlated with tumor identification and have low redundancy is obtained.

[0018] In the preferred embodiment, step S222 specifically includes: S2221. The core quantitative indicators are semantically prioritized according to a descending order of tumor core biomarkers, key imaging quantitative parameters, and core pathological indicators. Core tumor markers include carcinoembryonic antigen (CEA) and alpha-fetoprotein (AFP), which directly reflect the presence of tumors. Key quantitative parameters for images include quantitative parameters related to image performance, such as mean electrical impedance and variance. Core pathological indicators include quantitative indicators related to tumor pathological types; After sorting according to this priority, a one-dimensional index sequence carrying semantic priority is formed. The sequence expression is ,in The number of core quantitative indicators; S2222, Design Indicator Semantic Embedding Matrix The matrix dimension is ,in To predefine high-dimensional feature dimensions; a semantically driven dynamic weight matrix mapping function is used to transform the one-dimensional index sequence... Embedded into the semantic space, and then mapped to the high-dimensional feature space, a high-dimensional feature representation carrying semantic information is generated. The mapping formula is: (12); in, It represents high-dimensional features; The semantic embedding matrix for indicators is initialized using a pre-trained medical data semantic model. Each element in the matrix represents the correlation strength between the corresponding indicator and the semantic dimension. It is a one-dimensional index sequence; For semantic embedding bias; It is a dynamic weight matrix that is adaptively adjusted based on the semantic type of the indicators; Discrete quantification indicators are transformed into continuous high-dimensional feature vectors carrying semantic priority information of the indicators.

[0019] In the preferred embodiment, step S223 specifically includes: S2231. Obtain the dimension anchor point matrix of the dimension anchoring calibration module in step S21 in real time. This serves as a reference standard for dimension alignment; an anchor-driven dimension alignment layer is set after the feature generation layer, outputting the number of nodes and the anchor matrix. Preset output dimensions Consistent with the anchor point constraint rules embedded in step S21, the high-dimensional feature representation is... The input alignment layer is mapped to the target dimension through an anchor-guided adaptive linear transformation and semantic error correction mechanism. To obtain the initial aligned feature vector ; S2232. Using the same anchor point constraint and distribution alignment dual normalization strategy as in step S21, the feature vectors are initially aligned. Standardization processing is performed to distribute and structure feature vectors. Consistent; the anchor point constraint formula and the distribution alignment formula are the same as the corresponding formulas (10) and (11) in step S21, respectively; S2233. A lightweight multilayer perceptron (MLP) modal adapter is introduced to perform cross-modal calibration on the normalized feature vectors and dynamically adjust the distribution of impedance features; the activation function is ReLU, and the number of neurons in the output layer is related to the target dimension. Consistent; during calibration, parameters are optimized by minimizing the distribution difference loss function of the bimodal features, ultimately outputting a characteristic feature vector. Distribution and structured feature vectors Totally consistent.

[0020] In the preferred embodiment, step S3 specifically includes the following steps: S31. Structured feature vectors With characteristic feature vectors Perform dimensional consistency checks to unify the dimensions as follows: If dimensionality deviation exists, return to step S2 for recalibration; after successful verification, use a sparse topological matrix to filter the core feature clusters of the bimodal feature vectors, where: Construct a topological correlation matrix for bimodal features, where matrix elements represent In the mid-feature cluster and For feature clusters, the association strength is obtained by calculating the semantic similarity between the two feature clusters; based on this matrix, core feature cluster pairs are selected, feature cluster pairs with association strength higher than a preset threshold are retained, and redundant feature cluster pairs with low association strength are removed to reduce the redundancy of dual-modal features and reduce the amount of computation and feature interference in the fusion process. S32. Perform a double evaluation on the selected core feature clusters, calculate their respective task importance scores, and obtain information entropy to enhance the rationality of weight allocation; The importance mapping relationship between features and the tumor identification task is learned through an attention mechanism to evaluate the contribution of features. (Attention weight matrix) The dimension is 1×D, with a preset attention bias term. ; Structured feature vectors With characteristic feature vectors The calculation logic for the contribution evaluation value is a linear combination of the attention weight matrix and the feature vector, plus an attention bias term; Information entropy is obtained by calculating the probability distribution of the feature vector elements. To measure the information richness of a feature vector, its probability distribution is represented by normalized feature element values, while a maximum information entropy threshold is introduced. The information entropy is normalized. By combining feature contribution and information entropy, the task importance score of the bimodal features is obtained. and , respectively corresponding to structured feature vectors With characteristic feature vectors The calculation formula is: (13); (14); in, This is the attention weight matrix; This refers to attention bias. and These are the information entropies of the two feature vectors, respectively. The maximum information entropy threshold; S33. Introduce the InfoNCE contrastive loss function to improve the semantic consistency of fused features and strengthen the cross-modal semantic association of bimodal features by bringing the bimodal feature distance of the same tumor sample closer and pushing the feature distance of different samples further apart. First, construct positive and negative sample pairs, where: Positive sample pairs are structured feature vectors of the same tumor sample. With characteristic feature vectors ; Negative sample pairs are the current sample's Compared with other tumor samples or normal tissue samples ; Each sample corresponds to 1 positive sample pair and K negative sample pairs; The formula for calculating the InfoNCE contrast loss is as follows: (15); in, To compare the loss values; This is the cosine similarity calculation function, used to measure the semantic similarity between two feature vectors; These are the positive samples of the current structured feature vector; These are the positive samples of the current characteristic feature vector; For the first Characterized feature vectors of negative samples; This is a temperature coefficient used to adjust the smoothness of the similarity distribution; By minimizing the contrast loss L, the distribution of bimodal features is dynamically adjusted, making the bimodal features of the same tumor sample more tightly clustered in the feature space, while the features of different samples are more dispersed, thus strengthening cross-modal semantic association and improving the discriminative ability of fused features. S34. To address the issue of low proportion and weak feature response of minority samples such as early-stage micro-tumors in clinical data, a category entropy evaluation mechanism is embedded to assign adaptive weight gain to the features of minority samples. First, the distribution entropy of each category is obtained by statistically analyzing the proportion of samples of each category in the training set. Then, the weight gain coefficient is calculated based on the distribution entropy. The distribution entropy of minority class samples is larger, and its formula is: (16); in, This represents the total number of samples in the training set. This represents the current number of minority class samples; The lower the proportion of minority class samples, the larger the weight gain coefficient, thus strengthening the response intensity of minority class sample features; The adjusted score is obtained by multiplying the weight gain coefficient k by the task importance scores e1 and e2 of the bimodal features, respectively. and This increases the weight of minority class features in the fusion process, thereby improving the detection rate of early-stage micro-tumors. S35. Use the Softmax function to score the adjusted task importance. and Normalization is performed to obtain the dynamic weights of the bimodal features. The calculation formula is: (17); (18); in, Structured feature vectors Dynamic weights; For characteristic feature vectors Dynamic weights; For different samples, the weights are automatically adjusted based on the task importance of the features, information entropy, and sample category. Finally, based on dynamic weights , For structured feature vectors With characteristic feature vectors By performing a weighted summation, the final comprehensive feature vector is obtained. .

[0021] In the preferred embodiment, step S4 specifically includes the following steps: S41. Regarding the comprehensive feature vector Clinical logic adaptation and multicenter bias correction are performed, and the feature vectors are adjusted according to semantic priority ranking results. The feature weights of the corresponding dimensions are adjusted to align with the diagnostic logic that prioritizes core clinical biomarkers. Then, the feature mean of the center of the input sample is statistically analyzed. With variance For the comprehensive feature vector Standardization corrections are performed, and the correction formula is as follows: (19); in, The corrected feature vector; To prevent constraint values ​​with a denominator of 0; To offset the systemic biases caused by the heterogeneity of multi-center equipment; S42, A fully connected classifier is constructed to map the corrected feature vectors to the probabilities of benign and malignant tumors. The classifier consists of hidden layers and an output layer, with ReLU as the activation function. The output layer has two neurons, corresponding to benign and malignant tumor categories respectively, and Softmax as the activation function. The corrected feature vectors are then mapped to the tumor malignancy probabilities. Inputting the above fully connected classifier, the probability vector of benign or malignant tumors is obtained through linear transformation and nonlinear activation operation, as shown in the formula: (20); in, This is a probability vector representing benign or malignant conditions. This is the classification layer weight matrix; For classification layer bias terms; The weight matrix and bias terms are adaptively learned through the feature training logic of steps S2 and S3, so that the classification boundary conforms to the distribution pattern of clinical data. Based on classification probability Determine the final identification results and generate interpretable reports simultaneously to meet the traceability requirements of clinical diagnosis.

[0022] This invention provides a tumor identification method based on multimodal feature fusion. It synchronously acquires visual structural and electrical impedance dual-modal data through an event-driven sensor array, and dynamically adjusts acquisition parameters using a biomimetic spatiotemporal perception algorithm to ensure data synchronization stability and integrity. In the preprocessing stage, pulse coding and STDP learning rules effectively filter redundant noise while retaining key features. A heterogeneous twin coding engine accurately extracts structured feature vectors through multi-scale dynamic sensing cell arrays and global topological association modeling. After redundant index screening, semantic embedding, and cross-modal calibration, it generates characteristic feature vectors, achieving a unified dual-modal feature dimension and distribution, thus solving the feature distribution problem caused by heterogeneous multi-center equipment. To address the bias issue, dynamic weighted fusion combines attention mechanisms, information entropy assessment, and InfoNCE contrastive loss to enhance the semantic association and discriminative ability of bimodal features. The category entropy assessment mechanism assigns adaptive weight gain to minority class samples, reducing the missed diagnosis rate of minority class samples such as early-stage small tumors. The fusion process is fully monitorable and traceable. Through clinical logic adaptation and multi-center bias correction, the model's generalization ability is improved. The final tumor identification results have both high accuracy and clinical interpretability, meeting the traceability requirements of clinical diagnosis. At the same time, the entire process reduces computational consumption through optimizations such as core feature cluster screening, adapts to the deployment of edge medical devices, and improves the accuracy, stability, and clinical applicability of tumor identification. Attached Figure Description

[0023] The present invention will be further described below with reference to the accompanying drawings and embodiments: Figure 1 This is a flowchart of the tumor identification method based on multimodal feature fusion according to the present invention. Figure 2 This is a flowchart of step S21 of the tumor identification method based on multimodal feature fusion according to the present invention; Figure 3 This is a flowchart of step S22 of the tumor identification method based on multimodal feature fusion according to the present invention; Detailed Implementation With the clinical scenario of lung tumor identification as the application background, and considering the morphological diversity of lung tumors, this invention is implemented in a specific manner based on its technical solution.

[0024] like Figure 1-3As shown, a tumor identification method based on multimodal feature fusion is proposed, which includes the following steps: S1. Collect dual-modal data and perform preprocessing; S2, the heterogeneous twin coding engine extracts structured feature vectors and characteristic feature vectors, and calibrates the unified dimension and distribution; S3, Based on Structured Feature Vectors With characteristic feature vectors By dynamically weighted fusion and semantic association enhancement, a comprehensive feature vector is generated. ; S4. For the comprehensive feature vector The tumor is classified and processed, and the tumor identification results and clinically interpretable reports are output.

[0025] In the preferred embodiment, in step S1: The data acquisition system consists of an event-driven sensor array, including a visual structure data acquisition unit using a CMOS image sensor IMX290 and a resistive structure data acquisition unit using an impedance sensor AD5933. The two units are time-aligned via an FPGA synchronization trigger module, with a synchronization error ≤1μs. The visual structure data acquisition unit is set to a sampling frequency of 30Hz to avoid data redundancy and outputs grayscale image data with a resolution of 512×512 and a pixel depth of 8 bits, which is used to capture the morphological contour, boundary clarity, spatial distribution within the lung lobe, and local texture details of lung tumors. The electrical impedance structure data acquisition unit is set with a sampling precision of 16 bits and a sampling frequency of 30Hz, consistent with the visual acquisition unit. The sampling range covers the electrical impedance value range of normal lung tissue and tumor tissue as needed, and the electrical impedance value of tumor tissue is collected simultaneously to quantify the physical properties of tumor tissue such as conductivity and dielectric constant. The biomimetic spatiotemporal perception algorithm sets the threshold for detecting ambient light fluctuations, while the threshold for detecting patient movement is calculated based on pixel matching of consecutive frame images. When the threshold is triggered, the sensor gain parameter is adjusted linearly according to the fluctuation amplitude, and the sampling trigger timing is delayed by 0-2ms to maintain the synchronization and stability of data acquisition. In pulse code processing: For visual structural data, calculate the change in gray value ΔG of each pixel over time. Based on the statistics of gray distribution in lung images, take into account both noise filtering and effective feature preservation. A preset threshold is used to output pulse 1 when ΔG is greater than the threshold, otherwise pulse 0 is output. For electrical impedance structural data, the triggering logic consistent with that for visual structural data is used to calculate the time series change ΔZ of the electrical impedance value. A preset threshold is set based on the range of changes in clinical lung tissue electrical impedance. When ΔZ is greater than the threshold, pulse 1 is output, and otherwise pulse 0 is output. Subsequently, the pulse sequence was adaptively optimized using the pulse timing-dependent plasticity learning rule (STDP), and the synaptic weight enhancement amplitude coefficient A was set. + =0.01, weight reduction coefficient A - =0.008, to avoid excessive suppression of effective features; weight enhancement time constant τ + =20ms, weight reduction time constant τ - =30ms, the pulse timing interval matching the sampling frequency; During the optimization process, the synaptic weights are dynamically adjusted using formulas (1) and (2) to strengthen the pulse firing mode that is strongly correlated with tumor characteristics and filter redundant noise such as respiratory motion and environmental electromagnetic interference. If the data loss rate is greater than 5% in a certain period, the sensor array is triggered to re-acquire data for that period. Then, the timestamp difference of the dual-modal pulse sequence is calculated. When the difference is greater than 5ms, the timestamp compensation algorithm of the FPGA synchronization module is used to correct the error and align the time of the two-modal data. Finally, the signal-to-noise ratio (SNR) is obtained by calculating the power spectral density of the pulse sequence. If the SNR is less than 40 dB, the pulse coding threshold and STDP parameters are readjusted until the verification is successful, and the preprocessed bimodal original dataset is output.

[0026] In the preferred embodiment, in step S2: Visual structure data is processed through a multi-scale dynamic perceptual cell array, and impedance data is processed through redundancy filtering, semantic embedding, and cross-modal calibration, outputting structured feature vectors with consistent dimensionality and distribution. With characteristic feature vectors ,in: Structured feature vectors In the extraction and calibration, First, multi-scale dynamic sensing cell array feature extraction is performed. The multi-scale cell library configuration includes cells of three scales: 2×2, 4×4, and 8×8, covering the size range of lung tumors from small lesions (diameter <5mm) to intermediate and late-stage lesions (diameter >20mm). Cell scale adaptive calculation sets the initial baseline scale. Pixel, smallest cell scale To avoid a surge in computation due to excessively small cell size, the gradient decay coefficient... Sensitivity and stability of balancing scale adjustment; calculation of tumor margin gradient using the 3×3 Sobel operator. The texture density gradient is calculated using the gray-level co-occurrence matrix. Statistical analysis of three texture parameters: energy, entropy, and contrast; sliding step size sets the baseline step size. Pixel, the maximum edge gradient of the entire image Based on the statistical results of the lung image training set, the sliding step size is calculated using formula (4) to increase the sampling density of key areas; the edge enhancement attention mask M(x,y) is calculated according to the edge gradient. Allocation, when The larger The larger the value, the more the edge-sensitive feature response value R of each cell is calculated using formula (5), thus enhancing the tumor edge feature response; in the cell feature self-organization fusion, the number of activated cells is adaptively determined according to the size of the lung image, and the cell contribution weight is... The weights of all cells are calculated by weighting the cell scale and gradient matching degree, and the sum of the weights of all cells is 1; convolution kernel The size is consistent with the cell scale, and a multi-scale fused local feature map is generated using formula (6). It characterizes fine-grained features such as tumor margins and textures; Subsequently, global topological association modeling is performed, and feature clustering is conducted. A preset threshold for semantic similarity is calculated based on cosine similarity. Local features with similarity values ​​higher than this threshold are grouped into the same category. The clusters are divided into three scale layers: 2×2, 4×4, and 8×8, based on the cell scale, forming a set of feature clusters. Where k is determined statistically based on lung tumor feature dimensions; topological correlation matrix The correlation strength between feature clusters is quantified by formula (7). The spatial center distance Dist is obtained by considering the feature clusters in... Geometric center coordinates in Calculate the Euclidean distance and normalize it to the 0-1 interval to indicate a correlation; use a 2×2 global average pooling operation. A global structured feature set is generated using formula (8). This allows each element to carry topological association information between feature clusters, thus fully representing the overall structure of lung tumors; Next, the semantic flow field is aligned, where the standard feature vector is trained based on labeled standard lung tumor samples, and then calculated... The semantic similarity matrix is ​​obtained by calculating the cosine similarity between each feature vector and the standard feature vector, and the spatial distance matrix is ​​obtained by calculating the Euclidean distance. These are then fused to generate a dynamic semantic flow field. Each feature point is substituted into the flow field for spatial mapping transformation to align the heterogeneous features of the multi-center device; if the average Euclidean distance between the mapped feature and the standard feature is greater than 1, the semantic flow field generation parameters are adjusted, and the flow field is regenerated until the accuracy requirements are met. Finally, anchoring calibration is performed, and the anchor point matrix is ​​calculated. Its dimensions are obtained through pre-training with standard samples. To balance feature representation capability and computational power consumption, M needs to address this. Feature dimensions; anchor bias vector An all-zero vector, with an adaptive weight matrix. Initialize as a random normal distribution N(0,0.01); the anchor point constraint and distribution alignment are obtained by formula (9) to obtain the initial mapped feature vector. , in formula (10) ,Will Mapped to the 0-1 interval, the distribution shift between different samples is eliminated by normalization using formulas (10) and (11), and finally a one-dimensional structured feature vector is output. The eigenvalues ​​are distributed in the interval 0-1; Characterized feature vectors In the extraction and calibration, First, redundant indicators were screened. The dynamic attenuation correlation assessment model had an initial attenuation coefficient of 0.95, which was dynamically adjusted based on the time distance between the indicator and the tumor. For quantitative indicators such as electrical impedance values, impedance change rate, and impedance spectrum peak values... The numerical sequence of its sampling points is dynamically attenuated and weighted, and the correlation evaluation value is calculated by summing the results with the tumor determination result sequence. By calculating all indicators Information entropy yields the correlation distribution entropy. When the entropy value is too high, the filtering threshold is lowered; when the entropy value is too low, the filtering threshold is raised; retaining... For indicators exceeding a threshold, redundant indicators are removed. A threshold is set using the Pearson correlation coefficient; indicators with correlation coefficients greater than the threshold are clustered into redundant clusters. For each cluster, [the following is extracted / extracted]. The highest-ranking indicator with an information entropy ≥ 2.0 is selected as the core indicator, resulting in the final set of core quantitative indicators. Next, the impedance-related indicators were sorted in descending order by core tumor markers, key imaging quantitative parameters, and core pathological indicators to form a one-dimensional indicator sequence. ,in The average electrical impedance. This is the main peak frequency of the impedance spectrum. Set the impedance phase angle; define the index semantic embedding matrix. The dimension is 256×3, and Dimensionality is consistent, initialized using a pre-trained semantic model of medical data; semantic embedding bias. A vector consisting entirely of zeros, with a dynamic weight matrix. Adjust according to the indicator type, and generate a high-dimensional feature representation using formula (12). This transforms discrete indices into continuous vectors carrying semantic information. Finally, align the anchor-driven dimensions, obtain the dimension anchor matrix A in real time, set the output dimension of the anchor-driven dimension alignment layer to 256 dimensions, and embed the same anchor constraint rules as before. Mapped to the initial alignment feature vector of the target dimension Standardization is performed using formulas (10) and (11) to make... and The distribution is consistent; finally, cross-modal calibration is performed. Its lightweight MLP contains two hidden layers with ReLU activation function and 256 neurons in the output layer; the MSE loss function is used to minimize the distribution difference of the bimodal features, and the MLP parameters are optimized through gradient descent. The final output dimension and distribution are consistent with... The completely consistent characteristic feature vector f2.

[0027] In the preferred embodiment, in step S3: A comprehensive feature vector is generated through core feature cluster selection, task importance assessment, semantic association enhancement, minority class weight gain, and dynamic weighted fusion. ,in: In the core feature cluster selection, the topological correlation matrix of the bimodal features is first constructed, and the matrix dimension is based on... Number of feature clusters and The core indicator is the number of derived feature clusters. The matrix elements obtain the association strength by calculating the cosine similarity between two clusters. Based on the preset association strength threshold of the lesion, feature cluster pairs with association strength greater than the threshold are retained, and redundant cluster pairs with association strength less than the threshold are removed to reduce feature redundancy and computational load. In task importance assessment, attention weight matrix The dimension is set to 1×256, initialized as a random normal distribution N(0,0.001), and the attention bias term is defined. ;pass Linear combination with eigenvectors Calculate separately and The feature contribution is then calculated; subsequently, the information entropy of the feature vector is obtained by calculating the probability distribution of the feature elements. and the maximum information entropy threshold Finally, the task importance score with the maximum information entropy is calculated using formula (13). Formula (14) is used to calculate Task importance score The importance of features in lung tumor identification is assessed by combining feature contribution and information richness. In semantic association enhancement, positive sample pairs are those of the same lung tumor sample. and Negative sample pairs are the current sample's... Compared with K other samples including tumor samples and normal lung tissue samples from different patients In the InfoNCE contrast loss optimization, a reasonable temperature coefficient is set. To enhance similarity discriminative power, the contrast loss is calculated using formula (15). Minimize using the Adam optimizer Multiple iterations allow the bimodal features of the same tumor sample to cluster in the feature space, while the features of different samples are dispersed, thus strengthening cross-modal semantic associations. In the minority class weight gain and dynamic weighted fusion, taking lung tumors as an example, the total number of samples N in the training set should be greater than 5000, including a small number of early-stage small lung tumors with a diameter <5mm and other minority class samples. Example: Calculate the weight gain coefficient using formula (16). This coefficient is related to Multiply each product separately to obtain the adjusted score. This increases the weight of minority class features in the fusion process; In dynamic weighted fusion, the Softmax function of formulas (17) and (18) is used to... Normalization is performed to obtain dynamic weights. and ,and By weight and To each and We perform a weighted summation to obtain the comprehensive feature vector. The eigenvalues ​​are distributed in the interval between 0 and 1.

[0028] In the preferred embodiment, in step S4: First, sort the results according to semantic priority. The feature weights of the corresponding dimensions are adjusted to align the feature weight distribution with the priority logic of clinical biomarkers, thereby correcting the feature vector; subsequently, the historical sample feature mean of the input sample's medical center is statistically analyzed. With variance Using formula (19) Standardization corrections are performed to offset systematic biases introduced by different devices, resulting in corrected feature vectors. ; The fully connected classifier contains two hidden layers with ReLU activation; the output layer has two neurons with Softmax activation; and the classification layer weight matrix... Initialized to N(0, 0.001), the classification layer bias term. Initialized as a vector of all zeros; learning rate 0.0001; optimized using cross-entropy loss function. and The process continues iteratively until the classification boundaries align with the distribution patterns of clinical data on lung tumors. Then, Input the classifier and obtain the benign / malignant probability vector using formula (20). To determine whether it is a benign, malignant tumor or normal tissue.

[0029] Finally, a clinically interpretable report is generated, which includes the top 5 features that have the greatest impact on the classification results, their weights, and clinical reference suggestions.

[0030] The above embodiments are merely preferred technical solutions of the present invention and should not be considered as limitations on the present invention. The scope of protection of the present invention should be limited to the technical solutions described in the claims, including equivalent substitutions of the technical features described in the claims. That is, equivalent substitutions and improvements within this scope are also within the scope of protection of the present invention.

Claims

1. A tumor identification method based on multimodal feature fusion, characterized in that, Includes the following steps: S1. Collect visual structural data and electrical impedance structural data of the tumor to construct bimodal data and perform preprocessing; S2. Design a heterogeneous twin coding engine to extract structured feature vectors. With characteristic feature vectors And calibrate the uniform dimension and distribution; S3, Based on Structured Feature Vectors With characteristic feature vectors By dynamically weighted fusion and semantic association enhancement, a comprehensive feature vector is generated. ; S4. For the comprehensive feature vector The tumor is classified and processed, and the tumor identification results and clinically interpretable reports are output.

2. The tumor identification method based on multimodal feature fusion according to claim 1, characterized in that, Step S1 specifically includes: S11. Perform pulse coding processing on the collected continuous data to convert the continuous gray values ​​and impedance values ​​into a pulse sequence. S12. The pulse timing-dependent plasticity learning rule (STDP) is introduced to adaptively optimize the pulse sequence. The synaptic weights are dynamically adjusted according to the timing relationship of pulse firing, so as to retain the key features in the data and filter out redundant noise. S13. Perform quality checks on the dual-modal pulse sequence, including data integrity, time synchronization, and signal-to-noise ratio. If the quality fails to meet the standards, readjust the pulse coding threshold and STDP learning rule parameters.

3. The tumor identification method based on multimodal feature fusion according to claim 1, characterized in that, Step S2 specifically includes: S21. Employing a multi-scale dynamic perceptual cell array for feature extraction, global topological association modeling, semantic flow field alignment, and dimensional anchoring calibration, this process handles the visual structural information of tumors and outputs structured feature vectors. ; S22. Employing a redundant index screening, semantic embedding generation, and cross-modal calibration and adaptation process, the numerical structural information of tumors is processed, and characteristic feature vectors are output. .

4. The tumor identification method based on multimodal feature fusion according to claim 3, characterized in that, Step S21 specifically includes: S211, a multi-scale dynamic sensing cell array is used to capture fine-grained details of tumor edges and local textures in the information, and to realize parallel sensing and fusion of multi-scale information; S212. Global topological association modeling constructs a topological association network between local features, realizing the structured dimensional integration of local features into global structural features; S213. By generating a dynamic semantic flow field, the data collected by different devices are mapped to a unified feature space for alignment, and the global structured feature set is aligned. S214. Dimensional Anchoring Calibration: By using a preset dimensional anchor point matrix and combining it with cell adaptive transformation, global topological features are anchored to a preset dimension. Simultaneously, anchor point constraints eliminate scale differences and distribution offsets, outputting a structured feature vector with fixed dimensions and a regular distribution. .

5. The tumor identification method based on multimodal feature fusion according to claim 4, characterized in that, Step S211 specifically includes: S2111, the multi-scale dynamic sensing cell array is equipped with a multi-scale cell library based on tumor edge gradients. With texture density gradient The cell scale is adaptively calculated by adjusting the sampling density using a gradient-following step size. The cell scale and step size are calculated using formulas (3) and (4), respectively. The formulas are: (3); in, The scale of the currently activated cells; This serves as the initial reference scale; This is the gradient decay coefficient, used to adjust the sensitivity of the scale to gradient changes; The minimum cell scale threshold is used to avoid excessive computation due to excessively small cell scales. (4); in, This is the current sliding step size; Used as the reference step size; This represents the edge gradient value of the current region. The maximum edge gradient value of the entire image is obtained by statistically analyzing the edge gradient distribution of the training set images. S2112. Embedded edge-enhanced attention mask to strengthen tumor edge feature response, suppress background interference, and calculate the edge-sensitive feature response value of cells. Formula (5) is used; the formula is as follows: (5); in, The area covered by the current cell, that is, the set of pixels contained in the cell; Coordinates within the coverage area The pixel grayscale value of the location; coordinates Location-based attention mask weights; coordinates Edge gradient values ​​at the location; This is an exponential enhancement term for the edge gradient, used to further amplify the response of edge features; S2113. By integrating the response values ​​of all activated cells through the cell feature self-organization fusion rule, a multi-scale fused local feature map is generated. The fusion is performed using formula (6); the formula is as follows: (6); in, The total number of activated cells; For the first The contribution weight of each cell is determined by the cell scale and the gradient matching degree. The higher the cell scale and the gradient matching degree of the current region, the greater the weight. The sum of the weights of all cells is 1. For the first Edge-sensitive feature response value of each cell; For feature enhancement operations, To adapt to the first A cell-scale convolution kernel, with the kernel size consistent with the cell scale, is used to extract local feature patterns in the cell response values.

6. The tumor identification method based on multimodal feature fusion according to claim 4, characterized in that, Step S212 specifically includes: S2121. Form a set of feature clusters through semantic similarity clustering and scale hierarchical analysis. ; S2122, Introducing the topological correlation matrix of feature clusters To quantify the spatial topological relationships and semantic dependencies among feature clusters, matrix elements are calculated using formula (7); the formula is as follows: (7); in, For feature clusters With feature clusters The strength of the topological association between them, with a value ranging from 0 to 1. The larger the value, the stronger the association between the two. For feature clusters and The semantic similarity is obtained by calculating the cosine similarity of the average feature vectors of the two clusters of features; The spatial center distance between two clusters of features is obtained by calculating the Euclidean distance between the geometric centers of the two clusters of features in the feature map. This is the sum of the topological association strengths of all feature clusters, used to normalize the association strengths. S2123, Based on Topological Independence Matrix Features are purified using characteristic distillation techniques to generate a globally structured feature set. The distillation process uses formula (8); the formula is as follows: (8); in, Topological Incidence Matrix The Row vectors containing the association weights of all feature clusters; For feature clusters The result of global average pooling.

7. The tumor identification method based on multimodal feature fusion according to claim 4, characterized in that, Steps S213 and S214 specifically include: S2131. Generating dynamic semantic flow fields based on semantic similarity matrices and spatial distance matrices; S2132. Substitute the feature points in the global structured feature set into the semantic flow field for spatial mapping transformation, align the heterogeneous features of multi-center devices, and verify the alignment error. If the standard is not met, adjust the flow field generation parameters. S2141, via dimensional anchor matrix The cell-adaptive linear and nonlinear hybrid transformation maps the global structured features to the dimensional space defined by the anchor points. The mapping process uses formula (9); the formula is: (9); in, This is for the initial mapping of feature vectors; The dimension anchor matrix is ​​obtained through pre-training on a standard dataset; This is the aligned global structured feature set; The anchor point bias vector; It is a non-linear activation function used to introduce non-linear feature transformations and enhance the expressive power of features; It is an adaptive weight matrix, which is adaptively learned through the training process and used to adjust the distribution of the mapped features; S2142, Initial mapping of feature vectors The anchor point constraint and distribution alignment are subjected to dual normalization processing, with the constraint and alignment using formula (10) and formula (11) respectively; the formulas are: (10); in, The feature vector after anchor point constraint; Anchor point matrix of dimensions The minimum value in; Anchor point matrix of dimensions The maximum value in; (11); in, The feature vectors after distribution alignment; This represents the mean of the features of the current batch of samples; The variance of the features of the current batch of samples; To prevent constraint values ​​with a denominator of 0; S2143, normalize the eigenvectors Output in one-dimensional form to obtain structured feature vectors. .

8. The tumor identification method based on multimodal feature fusion according to claim 3, characterized in that, Step S22 specifically includes: S221. Select core quantitative indicators and filter redundant data based on the correlation decay model; S222. By embedding the semantics of the indicators and generating high-dimensional features, discrete indicators are transformed into high-dimensional feature vectors carrying semantic information. S223. Construct an anchor-driven dimension alignment layer and a cross-modal calibration module, outputting structured feature vectors. Dimensionality and distribution consistent characteristic feature vectors .

9. The tumor identification method based on multimodal feature fusion according to claim 8, characterized in that, Step S221 specifically includes: S2211. Construct a correlation assessment model between indicators and tumor dynamic decay, and calculate the correlation assessment value. Quantify the correlation between each indicator and the tumor diagnosis result; S2212. Based on the distribution entropy of indicator correlation, dynamically adjust the screening threshold, retain highly correlated indicators, and obtain a set of candidate indicators. S2213. Redundant clusters are formed by calculating the correlation between indicators and clustering them. The core indicators of each redundant cluster are extracted to obtain a set of low-redundancy core quantitative indicators.

10. The tumor identification method based on multimodal feature fusion according to claim 8, characterized in that, Steps S222 and S223 specifically include: S2221. Sort the core quantitative indicators according to the preset semantic priority rules to form a one-dimensional indicator sequence. ; S2222, Using the indicator semantic embedding matrix Mapping with a dynamic weight matrix to a one-dimensional index sequence Embedded into the semantic space and mapped to a high-dimensional feature space, a high-dimensional feature representation carrying semantic information is generated. The mapping process uses formula (12); the formula is: (12); in, It represents high-dimensional features; The semantic embedding matrix for indicators is initialized using a pre-trained medical data semantic model. Each element in the matrix represents the correlation strength between the corresponding indicator and the semantic dimension. It is a one-dimensional index sequence; For semantic embedding bias; It is a dynamic weight matrix that is adaptively adjusted based on the semantic type of the indicators. S2231, using dimensional anchor matrix For reference, an anchor-driven dimension alignment layer is set to represent high-dimensional features. Mapped to target dimension Preliminary alignment feature vector ; S2232. Employing a dual normalization strategy of anchor point constraint and distribution alignment, the feature vectors are initially aligned. With structured feature vectors The distribution is consistent, and the constraints and alignment are respectively expressed by formula (10) and formula (11). S2233. Perform cross-modal calibration using a lightweight multilayer perceptron (MLP) modal adapter to output a characteristic feature vector. .

Citation Information

Cited By

  • Self-adaptive identification system for medical capsule endoscope streaming image

    CN122066708A

  • Medical data granularity alignment method and system based on multi-modal fusion

    CN122091257A