Multi-parameter joint detection method and system for early screening of oral cancer
Through a multi-parameter joint detection method and the use of an edge-cloud collaborative computing architecture, oral images, salivary miRNA, and multi-frequency impedance data are integrated to solve the problems of low early lesion detection rate and high misdiagnosis rate in existing technologies, and achieve efficient oral cancer detection.
Patent Information
- Application Number
- CN202511094112.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-08-06
AI Technical Summary
Existing oral cancer early screening technology relies on a single detection modality, resulting in a low detection rate of early lesions, a high misdiagnosis rate, and a lack of multi-band dynamic feature analysis capabilities, which cannot meet clinical real-time needs.
A multi-parameter joint detection method is adopted to obtain oral images, salivary miRNA and multi-frequency impedance data through edge devices, and a lightweight model is used for initial screening. The cloud-based Transformer model is combined to perform deep fusion and hierarchical classification of multimodal features, integrate image, miRNA and impedance data, and dynamically adjust the modal weights.
It significantly improves the detection rate of early cancer and reduces the misdiagnosis rate. It is suitable for large-scale screening in primary medical institutions and realizes efficient oral cancer detection.
Smart Images

Figure CN120600341A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data analysis, and in particular to a multi-parameter combined detection method and system for early screening of oral cancer. Background Art
[0002] Existing oral cancer early screening technology relies on a single detection modality (such as visual inspection or a single biomarker alone), resulting in a low detection rate for early lesions; over-reliance on physician experience leads to a high misdiagnosis rate; and oral cancer detection equipment (such as electronic tongue impedance meters) can only provide static parameters and lack the ability to analyze multi-band dynamic features; secondly, established oral cancer AI mostly uses a single CNN model, which has low accuracy in identifying oral cancer and cannot meet clinical real-time requirements. Summary of the Invention
[0003] In order to solve the above technical problems, a multi-parameter combined detection method and system for early screening of oral cancer are provided. This technical solution solves the above problems.
[0004] In order to achieve the above objects, the technical solution adopted by the present invention is: A multi-parameter combined detection method for early screening of oral cancer, comprising: S1. Obtaining oral multi-dimensional data of a target user based on a collection terminal; wherein the oral multi-dimensional data of the target user includes: oral image data, oral saliva miRNA data, and oral multi-frequency impedance data; S2: Based on edge devices, a lightweight model for classifying suspected oral cancer is pre-set to pre-process the target user's multi-dimensional oral data to generate the target user's oral multimodal feature vector and preliminary oral status classification label; Among them, the lightweight model for classifying suspected oral cancer uses the MobileViTv3 architecture as the image branch, the 1DCNN architecture as the salivary miRNA branch, and the TinyViT architecture as the oral impedance branch; S3. Deploy the Transformer multimodal oral cancer status classification model based on cloud nodes. Perform a joint analysis of the target user's oral multimodal feature vectors and the initial oral status classification labels to generate a classification of the target user's oral cancer status. Among them, the Transformer multimodal oral cancer status hierarchical classification model uses SwinTransformer as the image encoder, ResNet as the miRNA encoder, and TimeSformer as the impedance encoder to form a multimodal oral cancer status hierarchical classification model.
[0005] Preferably, step S2 includes the following contents: Based on the target user's oral image data, the non-local mean filtering algorithm is used to calculate the weighted average of the similarity of adjacent pixel blocks in the image to eliminate noise; Using U-Net image segmentation, a binary mask of the oral image of the target user is generated for the oral image data, and the lesion area of the oral image of the target user is marked to obtain an image of the target oral lesion area; Normalizing the target oral lesion area image to obtain a normalized image tensor of the target oral lesion area; Based on the oral saliva miRNA data of the target user, the oral saliva miRNA data is divided into several concentration intervals and standardized to obtain several oral saliva miRNA multi-dimensional vectors of the target user; Based on the target user's oral multi-frequency impedance data, the impedance amplitude and phase angle in each frequency band in the oral multi-frequency impedance data are marked and substituted into the Cole-Cole curve to extract the target user's oral high-frequency impedance, low-frequency impedance, cell membrane capacitance and distribution coefficient of the impedance spectrum to form the target user's oral multi-frequency impedance vector.
[0006] Preferably, step 2 further comprises: Based on the MobileViTv3 architecture, an image branch feature extraction model for target users is established; Substitute the target oral lesion area image into the target user's image branch feature extraction model, divide the target oral lesion area image into several pixel blocks of equal size, input them into the fully connected layer for linear projection, and generate the target oral lesion area image pixel block embedding vector; The position code of the target oral lesion area image is added and fused with the embedding vector of the target oral lesion area image pixel block. Then, a lightweight multi-head attention mechanism is used to calculate the attention weights of the query, key, and value of each head in the embedding vector of the spatial information dimension of the target oral lesion area image pixel block. The weights are substituted into the fusion output of the linear transformation projection layer to obtain the multi-head attention linear projection of the target oral lesion area image. According to layer normalization, the multi-head attention linear projection of the target oral lesion area image is normalized according to the feature dimension according to self-attention, substituted into the feedforward network and residual connection, and the local-global correlation feature vector of the target oral lesion area is output; Through global average pooling, the local-global correlation feature vector of the target oral lesion area is globally pooled to generate the image feature vector of the target oral lesion area; Based on 1DCNN, a salivary miRNA branch feature extraction model for target users was established; Substitute the target user's multiple oral saliva miRNA multi-dimensional vectors into the target user's saliva miRNA branch feature extraction model, convert the target user's multiple oral saliva miRNA multi-dimensional vectors into three-dimensional tensors, input them into the convolution layer, and perform convolution operations on the target user's three-dimensional tensors of each oral saliva miRNA dimension to generate the target user's multiple oral saliva miRNA dimension three-dimensional vector feature maps as follows: ; in, is the three-dimensional vector feature map of the target user’s i-th oral saliva miRNA dimension, is the activation function, For the The weight of the convolution kernel, for Convolution inputs the target user’s oral saliva miRNA value, is the bias term; Input several three-dimensional vector feature maps of the target user's oral saliva miRNA dimension into the pooling layer, take the maximum value within the pooling window for each three-dimensional vector feature map of the target user's oral saliva miRNA dimension, and output the three-dimensional vector pooling feature map of each oral saliva miRNA dimension of the target user; Based on the three-dimensional vector pooling feature map of each oral saliva miRNA dimension of the target user, the fully connected compression layer is input to flatten the three-dimensional vector pooling feature map of each oral saliva miRNA dimension of the target user into a one-dimensional vector to obtain the oral saliva miRNA feature vector of the target user; Based on the TinyViT architecture, a branch feature extraction model for oral impedance of target users is established; Based on the target user's oral multi-frequency impedance vector, each oral frequency impedance vector of the target user is independently standardized to obtain the target user's oral multi-frequency impedance standardized vector, which is substituted into the target user's oral impedance branch feature extraction model for linear projection to generate the target user's multi-dimensional oral impedance standardized vector, which is input into the target user's oral impedance feature extraction model; According to the single-head attention mechanism, the attention weights of the query, key, and value corresponding to the multi-dimensional oral impedance standardized vector of the target user are calculated and input into the feedforward network layer. The feedforward network layer uses the first fully connected layer to map the attention weights of the query, key, and value corresponding to the multi-dimensional oral impedance standardized vector of the target user. After mapping the multi-dimensional oral impedance standardized vector of the target user to a high-dimensional space, the second fully connected layer is input to compress the high-dimensional vector of the multi-dimensional oral impedance standardized vector of the target user back to the low-dimensional space to obtain the multi-dimensional oral impedance compressed feature vector of the target user; Based on the pooling layer, the target user's multi-dimensional oral resistance compression feature vector is used as input, and the target user's oral resistance compression feature vector is used as output.
[0007] Preferably, step 2 further comprises: The target oral lesion image feature vector, the target user's oral saliva miRNA feature vector, and the target user's oral pressure feature vector are substituted into the meta-learning process. Each feature vector is individually input into a multilayer perceptron with two hidden layers of size 16. The importance of each feature vector in the corresponding oral cancer lesion task is verified, and a scalar weight score is assigned to the target oral lesion image feature vector, the target user's oral saliva miRNA feature vector, and the target user's oral pressure feature vector. Based on the target oral lesion area image feature vector weight score, the target user's oral saliva miRNA feature vector weight score, and the target user's oral obstruction pressure feature vector scalar weight score, each feature vector score is normalized so that the sum of the weights of all feature vectors is 1, and the dynamic weight of the target oral lesion area image feature vector, the dynamic weight of the target user's oral saliva miRNA feature vector, and the dynamic weight of the target user's oral obstruction pressure feature vector are obtained; Based on the dynamic weight of the target oral lesion area image feature vector, the dynamic weight of the target user's oral saliva miRNA feature vector, and the dynamic weight of the target user's oral obstruction pressure feature vector, the target oral lesion area image feature vector, the target user's oral saliva miRNA feature vector, and the target user's oral obstruction pressure feature vector are concatenated in series according to the weighted fusion method to obtain the target user's oral multimodal feature vector as follows: ; in, is the oral multimodal feature vector of the target user, is the dynamic weight of the target oral lesion area image feature vector, is the dynamic weight of the target user's oral saliva miRNA feature vector, is the dynamic weight of the oral pressure feature vector of the target user, is the target oral lesion area image feature vector, is the target user's oral saliva miRNA feature vector, is the oral resistance feature vector of the target user; A deep neural network is pre-trained using known early-stage oral cancer samples. The deep neural network consists of a three-layer fully connected network for normal oral cavity, early-stage oral cancer, and stage-stage oral cancer. The deep neural network takes the target user's oral multimodal feature vector as input and outputs the probability distribution of oral cancer lesions corresponding to the target user's oral multimodal feature vector. Based on the oral cancer lesion probability distribution corresponding to the target user's oral multimodal feature vector, the target user is divided according to the risk threshold interval corresponding to the oral cancer lesion probability to obtain the initial screening oral status classification label.
[0008] Preferably, step 3 includes the following: Based on the Transformer architecture, a multimodal oral cancer status classification model was constructed using SwinTransformer as the image encoder, ResNet as the miRNA encoder, and TimeSformer as the impedance encoder. Based on One-Hot encoding, the target user's initial screening oral status classification label is converted into the target user's initial screening oral status classification coding label and concatenated with the target user's oral multimodal feature vector to obtain the target user's oral multi-dimensional feature-label enhancement vector; Based on the multimodal oral cancer status hierarchical classification model, linear projection and position encoding are performed on the target user's oral multi-dimensional feature-label enhancement vector, and the oral multi-dimensional feature-label enhancement vector is used as the encoder input. The oral cancer probability distribution of the target user's oral multi-dimensional feature-label enhancement vector of each encoder is used.
[0009] Preferably, step 3 further includes: Based on the entropy weight method formula, the uncertainty entropy of the oral cancer probability distribution of the oral multi-dimensional features-label enhancement vector of the target user of each encoder is calculated, and the oral cancer probability distribution weight of the oral multi-dimensional features-label enhancement vector of the target user of each encoder is assigned as follows: ; in, is the oral cancer probability distribution weight of the i-th oral multi-dimensional feature-label enhancement vector of the target user of the j-th encoder, is the uncertainty entropy of the oral cancer probability distribution of the i-th oral multi-dimensional feature-label enhanced vector of the target user of the j-th encoder, is the probability distribution of oral cancer of the i-th oral multi-dimensional feature-label enhanced vector of the target user of the j-th encoder; Based on the target user's oral multi-dimensional features and label enhancement vectors of each encoder, a weighted fusion method is used to obtain the target user's oral multi-dimensional features and label enhancement vectors of oral cancer fusion probability distribution. The method is as follows: ; in, Oral cancer fusion probability distribution of the target user's oral multi-dimensional feature-label enhancement vector; Based on the oral cancer fusion probability distribution of the target user's oral multi-dimensional features and label enhancement vector, the target user's oral status is classified according to the risk threshold interval corresponding to the probability of oral cancer lesions.
[0010] Furthermore, a multi-parameter combined detection system for early screening of oral cancer is used to implement the multi-parameter combined detection method for early screening of oral cancer as described above, comprising: The edge data acquisition module is used to obtain the target user's oral multi-dimensional data based on the acquisition terminal; the target user's oral multi-dimensional data includes: oral image data, oral saliva miRNA data and oral multi-frequency impedance data; The edge node initial screening module is electrically connected to the edge data acquisition module. The edge node initial screening module is used to pre-set a lightweight model for classifying suspicious oral cancer based on the edge device, pre-process the target user's oral multi-dimensional data, and generate the target user's oral multimodal feature vector and initial screening oral status classification label; The cloud node screening module is wirelessly connected to the edge node initial screening module. The cloud node screening module is used to deploy the Transformer multimodal oral cancer status grading and classification model based on the cloud node, and jointly analyze the target user's oral multimodal feature vector and the initial screening oral status classification label to generate the target user's oral cancer status grading and classification.
[0011] Compared with the prior art, the present invention has the following beneficial effects: The present invention proposes a multi-parameter joint detection scheme for early screening of oral cancer. Through the edge-cloud collaborative computing architecture, it integrates oral images, salivary miRNA and multi-frequency impedance multimodal data. The edge device first performs a lightweight model for rapid initial screening, and then uses the cloud-based Transformer model to perform deep fusion and hierarchical classification of multimodal features. The beneficial effects are: the use of multi-parameter joint detection significantly improves the detection rate of early cancer; 2 edge-cloud layered processing realizes efficient screening; the dynamic feature fusion mechanism can adaptively adjust the weight of each modality and reduce the impact of single parameter errors. The present invention effectively solves the problems of insufficient accuracy and reliance on expert experience of traditional screening methods, and is suitable for large-scale screening scenarios in primary medical institutions. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 This is a flow chart of a multi-parameter combined detection method for early screening of oral cancer; Figure 2 This is a multi-parameter combined detection method and system framework diagram for early screening of oral cancer. DETAILED DESCRIPTION
[0013] The following description is intended to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments described below are merely examples, and those skilled in the art may conceive of other obvious variations.
[0014] Reference Figure 1 As shown, a multi-parameter combined detection method for early screening of oral cancer includes: S1. Obtaining oral multi-dimensional data of a target user based on a collection terminal; wherein the oral multi-dimensional data of the target user includes: oral image data, oral saliva miRNA data, and oral multi-frequency impedance data; S2: Based on edge devices, a lightweight model for classifying suspected oral cancer is pre-set to pre-process the target user's multi-dimensional oral data to generate the target user's oral multimodal feature vector and preliminary oral status classification label; Among them, the lightweight model for classifying suspected oral cancer uses the MobileViTv3 architecture as the image branch, the 1DCNN architecture as the salivary miRNA branch, and the TinyViT architecture as the oral impedance branch; Step S2 includes the following: Based on the target user's oral image data, the non-local mean filtering algorithm is used to calculate the weighted average of the similarity of adjacent pixel blocks in the image to eliminate noise; Using U-Net image segmentation, a binary mask of the oral image of the target user is generated for the oral image data, and the lesion area of the oral image of the target user is marked to obtain an image of the target oral lesion area; Normalizing the target oral lesion area image to obtain a normalized image tensor of the target oral lesion area; Based on the oral saliva miRNA data of the target user, the oral saliva miRNA data is divided into several concentration intervals and standardized to obtain several oral saliva miRNA multi-dimensional vectors of the target user; Based on the target user's oral multi-frequency impedance data, the impedance amplitude and phase angle at each frequency band in the oral multi-frequency impedance data are marked and substituted into the Cole-Cole curve to extract the target user's oral high-frequency impedance, low-frequency impedance, cell membrane capacitance, and distribution coefficient of the impedance spectrum, and construct the target user's oral multi-frequency impedance vector; Step 2 also includes: Based on the MobileViTv3 architecture, an image branch feature extraction model for target users is established; Substitute the target oral lesion area image into the target user's image branch feature extraction model, divide the target oral lesion area image into several pixel blocks of equal size, input them into the fully connected layer for linear projection, and generate the target oral lesion area image pixel block embedding vector; The position code of the target oral lesion area image is added and fused with the embedding vector of the target oral lesion area image pixel block. Then, a lightweight multi-head attention mechanism is used to calculate the attention weights of the query, key, and value of each head in the embedding vector of the spatial information dimension of the target oral lesion area image pixel block. The weights are substituted into the fusion output of the linear transformation projection layer to obtain the multi-head attention linear projection of the target oral lesion area image. According to layer normalization, the multi-head attention linear projection of the target oral lesion area image is normalized according to the feature dimension according to self-attention, substituted into the feedforward network and residual connection, and the local-global correlation feature vector of the target oral lesion area is output; Through global average pooling, the local-global correlation feature vector of the target oral lesion area is globally pooled to generate the image feature vector of the target oral lesion area; Based on 1DCNN, a salivary miRNA branch feature extraction model for target users was established; Substitute the target user's multiple oral saliva miRNA multi-dimensional vectors into the target user's saliva miRNA branch feature extraction model, convert the target user's multiple oral saliva miRNA multi-dimensional vectors into three-dimensional tensors, input them into the convolution layer, and perform convolution operations on the target user's three-dimensional tensors of each oral saliva miRNA dimension to generate the target user's multiple oral saliva miRNA dimension three-dimensional vector feature maps as follows: ; in, is the three-dimensional vector feature map of the target user’s i-th oral saliva miRNA dimension, is the activation function, For the The weight of the convolution kernel, for Convolution inputs the target user’s oral saliva miRNA value, is the bias term; Input several three-dimensional vector feature maps of the target user's oral saliva miRNA dimension into the pooling layer, take the maximum value within the pooling window for each three-dimensional vector feature map of the target user's oral saliva miRNA dimension, and output the three-dimensional vector pooling feature map of each oral saliva miRNA dimension of the target user; As a further example, the configuration of the target user's salivary miRNA branch feature extraction model is as follows: Convolution kernel configuration: the number of convolutions is 16, the size of each kernel is 3, the stride is 1, the padding method is to make the output length the same as the input, and the activation function is ReLU; pooling configuration: the pooling window size is 2, the stride is 2, and the padding is not filled; Secondly, since the target user's oral saliva miRNA dimension three-dimensional vector pooling feature map has an odd number, but the pooling window and pooling step size are 2, it will lead to a lack of data in the final pooling window. Therefore, in the initial stage of pooling, it is possible to determine whether the total number of the target user's oral saliva miRNA dimension three-dimensional vector pooling feature map is an odd number. If so, the final pooling window is called to discard the pooling data of the last step. If not, the process is executed normally. Based on the three-dimensional vector pooling feature map of each oral saliva miRNA dimension of the target user, the fully connected compression layer is input to flatten the three-dimensional vector pooling feature map of each oral saliva miRNA dimension of the target user into a one-dimensional vector to obtain the oral saliva miRNA feature vector of the target user; Based on the TinyViT architecture, a branch feature extraction model for oral impedance of target users is established; Based on the target user's oral multi-frequency impedance vector, each oral frequency impedance vector of the target user is independently standardized to obtain the target user's oral multi-frequency impedance standardized vector, which is substituted into the target user's oral impedance branch feature extraction model for linear projection to generate the target user's multi-dimensional oral impedance standardized vector, which is input into the target user's oral impedance feature extraction model; According to the single-head attention mechanism, the attention weights of the query, key, and value corresponding to the multi-dimensional oral impedance standardized vector of the target user are calculated and input into the feedforward network layer. The feedforward network layer uses the first fully connected layer to map the attention weights of the query, key, and value corresponding to the multi-dimensional oral impedance standardized vector of the target user. After mapping the multi-dimensional oral impedance standardized vector of the target user to a high-dimensional space, the second fully connected layer is input to compress the high-dimensional vector of the multi-dimensional oral impedance standardized vector of the target user back to the low-dimensional space to obtain the multi-dimensional oral impedance compressed feature vector of the target user; Based on the pooling layer, the target user's multi-dimensional oral resistance compression feature vector is used as input and the target user's oral resistance compression feature vector is used as output; As a further example, the configuration of the oral impedance feature extraction model for the target user is as follows: Single-head self-attention has a head dimension of 16, scaled dot product attention, a 2-layer MLP (16→32→16) feedforward network, GELU activation function, and mean pooling. Step 2 also includes: The target oral lesion image feature vector, the target user's oral saliva miRNA feature vector, and the target user's oral pressure feature vector are substituted into the meta-learning process. Each feature vector is individually input into a multilayer perceptron with two hidden layers of size 16. The importance of each feature vector in the corresponding oral cancer lesion task is verified, and a scalar weight score is assigned to the target oral lesion image feature vector, the target user's oral saliva miRNA feature vector, and the target user's oral pressure feature vector. As a further step, the importance of each feature vector in the corresponding oral cancer lesion task is verified by updating the MLP weights through gradient backpropagation to determine the contribution of each feature vector to the corresponding oral cancer lesion task. This is well known to those skilled in the art and will not be described in detail here. Based on the target oral lesion area image feature vector weight score, the target user's oral saliva miRNA feature vector weight score, and the target user's oral obstruction pressure feature vector scalar weight score, each feature vector score is normalized so that the sum of the weights of all feature vectors is 1, and the dynamic weight of the target oral lesion area image feature vector, the dynamic weight of the target user's oral saliva miRNA feature vector, and the dynamic weight of the target user's oral obstruction pressure feature vector are obtained; Based on the dynamic weight of the target oral lesion area image feature vector, the dynamic weight of the target user's oral saliva miRNA feature vector, and the dynamic weight of the target user's oral obstruction pressure feature vector, the target oral lesion area image feature vector, the target user's oral saliva miRNA feature vector, and the target user's oral obstruction pressure feature vector are concatenated in series according to the weighted fusion method to obtain the target user's oral multimodal feature vector as follows: ; in, is the oral multimodal feature vector of the target user, is the dynamic weight of the target oral lesion area image feature vector, is the dynamic weight of the target user's oral saliva miRNA feature vector, is the dynamic weight of the oral pressure feature vector of the target user, is the target oral lesion area image feature vector, is the target user's oral saliva miRNA feature vector, is the oral resistance feature vector of the target user; A deep neural network is pre-trained using known early-stage oral cancer samples. The deep neural network consists of a three-layer fully connected network for normal oral cavity, early-stage oral cancer, and stage-stage oral cancer. The deep neural network takes the target user's oral multimodal feature vector as input and outputs the probability distribution of oral cancer lesions corresponding to the target user's oral multimodal feature vector. Based on the oral cancer probability distribution corresponding to the target user's oral multimodal feature vector, the target user is divided according to the risk threshold interval corresponding to the oral cancer probability to obtain the target user's initial screening oral status classification label; When using, combine the above steps: As a further development, a lightweight model of branching factors that focus on oral cancer lesions in target users is preset at the edge nodes to initialize and extract the target user's oral multimodal feature vectors. This reduces the problem of inaccurate classification of oral cancer lesions due to insufficient granularity in uploaded data at the cloud nodes. Meta-learning is used to dynamically assign weights to each feature vector of the target user to avoid modal bias caused by fixed weights (such as over-reliance on images to miss molecular marker cases). Furthermore, the edge node deployment solution of the MobileViTv3 architecture + 1D CNN + TinyViT architecture has low requirements for localized computing resources.
[0015] S3. Deploy the Transformer multimodal oral cancer status classification model based on cloud nodes. Perform a joint analysis of the target user's oral multimodal feature vectors and the initial oral status classification labels to generate a classification of the target user's oral cancer status. Among them, the Transformer multimodal oral cancer status classification model uses SwinTransformer as the image encoder, ResNet as the miRNA encoder, and TimeSformer as the impedance encoder to build a multimodal oral cancer status classification model; Step 3 includes the following: Based on the Transformer architecture, a multimodal oral cancer status classification model was constructed using SwinTransformer as the image encoder, ResNet as the miRNA encoder, and TimeSformer as the impedance encoder. Based on One-Hot encoding, the target user's initial screening oral status classification label is converted into the target user's initial screening oral status classification coding label and concatenated with the target user's oral multimodal feature vector to obtain the target user's oral multi-dimensional feature-label enhancement vector; Based on the multimodal oral cancer status classification model, linear projection and position encoding are performed on the target user's oral multi-dimensional feature-label enhancement vector. The oral multi-dimensional feature-label enhancement vector is used as the encoder input, and the oral cancer probability distribution of the target user's oral multi-dimensional feature-label enhancement vector of each encoder is used; Step 3 also includes: Based on the entropy weight method formula, the uncertainty entropy of the oral cancer probability distribution of the oral multi-dimensional features-label enhancement vector of the target user of each encoder is calculated, and the oral cancer probability distribution weight of the oral multi-dimensional features-label enhancement vector of the target user of each encoder is assigned as follows: ; in, is the oral cancer probability distribution weight of the i-th oral multi-dimensional feature-label enhancement vector of the target user of the j-th encoder, is the uncertainty entropy of the oral cancer probability distribution of the i-th oral multi-dimensional feature-label enhanced vector of the target user of the j-th encoder, is the probability distribution of oral cancer of the i-th oral multi-dimensional feature-label enhanced vector of the target user of the j-th encoder; Based on the target user's oral multi-dimensional features and label enhancement vectors of each encoder, a weighted fusion method is used to obtain the target user's oral multi-dimensional features and label enhancement vectors of oral cancer fusion probability distribution. The method is as follows: ; in, Oral cancer fusion probability distribution of the target user's oral multi-dimensional feature-label enhancement vector; Based on the oral cancer fusion probability distribution of the target user's oral multi-dimensional features and label enhancement vector, the target user's oral status is classified according to the risk threshold interval corresponding to the probability of oral cancer lesions.
[0016] When using, combine the above steps: As a further point, the existing oral cancer multimodal diagnosis system uses static weight fusion, resulting in low-quality modal interference results, black-box decision-making lacks interpretability, and lacks an automated resolution mechanism when modal conflicts occur, resulting in insufficient diagnostic accuracy and low response efficiency for oral lesions. This approach utilizes a multimodal Transformer architecture (SwinTransformer, ResNet, and TimeSformer for image, miRNA, and impedance data, respectively) to concatenate initial screening labels and multimodal features into a feature-label augmented vector. This vector, after linear projection and positional encoding, is fed into an encoder to generate a cancer probability distribution for each modality. An entropy-weighted approach dynamically calculates the uncertainty of each modality's prediction, assigns weights, and performs a weighted fusion. Ultimately, risk levels (low, medium, and high) are assigned based on the fused probability distribution. This approach combines morphological (image), molecular marker (miRNA), and functional (impedance) data to improve early cancer detection rates. Entropy-weighted approaches automatically downweight low-quality modalities to reduce misdiagnosis (lowering the error rate of fixed-weight approaches).
[0017] Reference Figure 2 As shown, a multi-parameter combined detection system for early screening of oral cancer includes: The edge data acquisition module is used to obtain the target user's oral multi-dimensional data based on the acquisition terminal; the target user's oral multi-dimensional data includes: oral image data, oral saliva miRNA data and oral multi-frequency impedance data; The edge node initial screening module is electrically connected to the edge data acquisition module. The edge node initial screening module is used to pre-set a lightweight model for classifying suspicious oral cancer based on the edge device, pre-process the target user's oral multi-dimensional data, and generate the target user's oral multimodal feature vector and initial screening oral status classification label; The cloud node screening module is wirelessly connected to the edge node initial screening module. The cloud node screening module is used to deploy the Transformer multimodal oral cancer status grading and classification model based on the cloud node, and jointly analyze the target user's oral multimodal feature vector and the initial screening oral status classification label to generate the target user's oral cancer status grading and classification.
[0018] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions merely illustrate the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A multi-parameter combined detection method for early screening of oral cancer, characterized in that: include: S1. Based on the acquisition terminal, obtain the target user's oral multi-dimensional data; The target user's oral multi-dimensional data includes: oral image data, oral saliva miRNA data and oral multi-frequency impedance data; S2: Based on edge devices, a lightweight model for classifying suspected oral cancer is pre-set to pre-process the target user's multi-dimensional oral data to generate the target user's oral multimodal feature vector and preliminary oral status classification label; Among them, the lightweight model for classifying suspected oral cancer uses the MobileViTv3 architecture as the image branch, the 1DCNN architecture as the salivary miRNA branch, and the TinyViT architecture as the oral impedance branch; S3. Deploy the Transformer multimodal oral cancer status classification model based on cloud nodes. Perform a joint analysis of the target user's oral multimodal feature vectors and the initial oral status classification labels to generate a classification of the target user's oral cancer status. Among them, the Transformer multimodal oral cancer status hierarchical classification model uses SwinTransformer as the image encoder, ResNet as the miRNA encoder, and TimeSformer as the impedance encoder to form a multimodal oral cancer status hierarchical classification model.
2. A multi-parameter combined detection method for early screening of oral cancer according to claim 1, characterized in that: The step S2 includes the following contents: Based on the target user's oral image data, the non-local mean filtering algorithm is used to calculate the weighted average of the similarity of adjacent pixel blocks in the image to eliminate noise; Using U-Net image segmentation, a binary mask of the oral image of the target user is generated for the oral image data, and the lesion area of the oral image of the target user is marked to obtain an image of the target oral lesion area; Normalizing the target oral lesion area image to obtain a normalized image tensor of the target oral lesion area; Based on the oral saliva miRNA data of the target user, the oral saliva miRNA data is divided into several concentration intervals and standardized to obtain several oral saliva miRNA multi-dimensional vectors of the target user; Based on the target user's oral multi-frequency impedance data, the impedance amplitude and phase angle in each frequency band in the oral multi-frequency impedance data are marked and substituted into the Cole-Cole curve to extract the target user's oral high-frequency impedance, low-frequency impedance, cell membrane capacitance and distribution coefficient of the impedance spectrum to form the target user's oral multi-frequency impedance vector.
3. A multi-parameter combined detection method for early screening of oral cancer according to claim 2, characterized in that: Step 2 also includes: Based on the MobileViTv3 architecture, an image branch feature extraction model for target users is established; Substitute the target oral lesion area image into the target user's image branch feature extraction model, divide the target oral lesion area image into several pixel blocks of equal size, input them into the fully connected layer for linear projection, and generate the target oral lesion area image pixel block embedding vector; The position code of the target oral lesion area image is added and fused with the embedding vector of the target oral lesion area image pixel block. Then, a lightweight multi-head attention mechanism is used to calculate the attention weights of the query, key, and value of each head in the embedding vector of the spatial information dimension of the target oral lesion area image pixel block. The weights are substituted into the fusion output of the linear transformation projection layer to obtain the multi-head attention linear projection of the target oral lesion area image. According to layer normalization, the multi-head attention linear projection of the target oral lesion area image is normalized according to the feature dimension according to self-attention, substituted into the feedforward network and residual connection, and the local-global correlation feature vector of the target oral lesion area is output; Through global average pooling, the local-global correlation feature vector of the target oral lesion area is globally pooled to generate the image feature vector of the target oral lesion area; Based on 1DCNN, a salivary miRNA branch feature extraction model for target users was established; Substitute the target user's multiple oral saliva miRNA multi-dimensional vectors into the target user's saliva miRNA branch feature extraction model, convert the target user's multiple oral saliva miRNA multi-dimensional vectors into three-dimensional tensors, input them into the convolution layer, and perform convolution operations on the target user's three-dimensional tensors of each oral saliva miRNA dimension to generate the target user's multiple oral saliva miRNA dimension three-dimensional vector feature maps as follows: ; in, is the three-dimensional vector feature map of the target user’s i-th oral saliva miRNA dimension, is the activation function, For the The weight of the convolution kernel, for Convolution inputs the target user’s oral saliva miRNA value, is the bias term; Input several three-dimensional vector feature maps of the target user's oral saliva miRNA dimension into the pooling layer, take the maximum value within the pooling window for each three-dimensional vector feature map of the target user's oral saliva miRNA dimension, and output the three-dimensional vector pooling feature map of each oral saliva miRNA dimension of the target user; Based on the three-dimensional vector pooling feature map of each oral saliva miRNA dimension of the target user, the fully connected compression layer is input to flatten the three-dimensional vector pooling feature map of each oral saliva miRNA dimension of the target user into a one-dimensional vector to obtain the oral saliva miRNA feature vector of the target user; Based on the TinyViT architecture, a branch feature extraction model for oral impedance of target users is established; Based on the target user's oral multi-frequency impedance vector, each oral frequency impedance vector of the target user is independently standardized to obtain the target user's oral multi-frequency impedance standardized vector, which is substituted into the target user's oral impedance branch feature extraction model for linear projection to generate the target user's multi-dimensional oral impedance standardized vector, which is input into the target user's oral impedance feature extraction model; According to the single-head attention mechanism, the attention weights of the query, key, and value corresponding to the multi-dimensional oral impedance standardized vector of the target user are calculated and input into the feedforward network layer. The feedforward network layer uses the first fully connected layer to map the attention weights of the query, key, and value corresponding to the multi-dimensional oral impedance standardized vector of the target user. After mapping the multi-dimensional oral impedance standardized vector of the target user to a high-dimensional space, the second fully connected layer is input to compress the high-dimensional vector of the multi-dimensional oral impedance standardized vector of the target user back to the low-dimensional space to obtain the multi-dimensional oral impedance compressed feature vector of the target user; Based on the pooling layer, the target user's multi-dimensional oral resistance compression feature vector is used as input, and the target user's oral resistance compression feature vector is used as output.
4. A multi-parameter combined detection method for early screening of oral cancer according to claim 3, characterized in that: Step 2 also includes: The target oral lesion image feature vector, the target user's oral saliva miRNA feature vector, and the target user's oral pressure feature vector are substituted into the meta-learning process. Each feature vector is individually input into a multilayer perceptron with two hidden layers of size 16. The importance of each feature vector in the corresponding oral cancer lesion task is verified, and a scalar weight score is assigned to the target oral lesion image feature vector, the target user's oral saliva miRNA feature vector, and the target user's oral pressure feature vector. Based on the target oral lesion area image feature vector weight score, the target user's oral saliva miRNA feature vector weight score, and the target user's oral obstruction pressure feature vector scalar weight score, each feature vector score is normalized so that the sum of the weights of all feature vectors is 1, and the dynamic weight of the target oral lesion area image feature vector, the dynamic weight of the target user's oral saliva miRNA feature vector, and the dynamic weight of the target user's oral obstruction pressure feature vector are obtained; Based on the dynamic weight of the target oral lesion area image feature vector, the dynamic weight of the target user's oral saliva miRNA feature vector, and the dynamic weight of the target user's oral obstruction pressure feature vector, the target oral lesion area image feature vector, the target user's oral saliva miRNA feature vector, and the target user's oral obstruction pressure feature vector are concatenated in series according to the weighted fusion method to obtain the target user's oral multimodal feature vector as follows: ; in, is the oral multimodal feature vector of the target user, is the dynamic weight of the target oral lesion area image feature vector, is the dynamic weight of the target user's oral saliva miRNA feature vector, is the dynamic weight of the oral pressure feature vector of the target user, is the target oral lesion area image feature vector, is the target user's oral saliva miRNA feature vector, is the oral resistance feature vector of the target user; A deep neural network is pre-trained using known early-stage oral cancer samples. The deep neural network consists of a three-layer fully connected network for normal oral cavity, early-stage oral cancer, and stage-stage oral cancer. The deep neural network takes the target user's oral multimodal feature vector as input and outputs the probability distribution of oral cancer lesions corresponding to the target user's oral multimodal feature vector. Based on the oral cancer lesion probability distribution corresponding to the target user's oral multimodal feature vector, the target user is divided according to the risk threshold interval corresponding to the oral cancer lesion probability to obtain the initial screening oral status classification label.
5. A multi-parameter combined detection method for early screening of oral cancer according to claim 4, characterized in that: The step 3 includes the following: Based on the Transformer architecture, a multimodal oral cancer status classification model was constructed using SwinTransformer as the image encoder, ResNet as the miRNA encoder, and TimeSformer as the impedance encoder. Based on One-Hot encoding, the target user's initial screening oral status classification label is converted into the target user's initial screening oral status classification coding label and concatenated with the target user's oral multimodal feature vector to obtain the target user's oral multi-dimensional feature-label enhancement vector; Based on the multimodal oral cancer status hierarchical classification model, linear projection and position encoding are performed on the target user's oral multi-dimensional feature-label enhancement vector, and the oral multi-dimensional feature-label enhancement vector is used as the encoder input. The oral cancer probability distribution of the target user's oral multi-dimensional feature-label enhancement vector of each encoder is used.
6. A multi-parameter combined detection method for early screening of oral cancer according to claim 5, characterized in that: Step 3 also includes: Based on the entropy weight method formula, the uncertainty entropy of the oral cancer probability distribution of the oral multi-dimensional features-label enhancement vector of the target user of each encoder is calculated, and the oral cancer probability distribution weight of the oral multi-dimensional features-label enhancement vector of the target user of each encoder is assigned as follows: ; in, is the oral cancer probability distribution weight of the i-th oral multi-dimensional feature-label enhancement vector of the target user of the j-th encoder, is the uncertainty entropy of the oral cancer probability distribution of the i-th oral multi-dimensional feature-label enhanced vector of the target user of the j-th encoder, is the probability distribution of oral cancer of the i-th oral multi-dimensional feature-label enhanced vector of the target user of the j-th encoder; Based on the target user's oral multi-dimensional features and label enhancement vectors of each encoder, a weighted fusion method is used to obtain the target user's oral multi-dimensional features and label enhancement vectors of oral cancer fusion probability distribution. The method is as follows: ; in, Oral cancer fusion probability distribution of the target user's oral multi-dimensional feature-label enhancement vector; Based on the oral cancer fusion probability distribution of the target user's oral multi-dimensional features and label enhancement vector, the target user's oral status is classified according to the risk threshold interval corresponding to the probability of oral cancer lesions.
7. A multi-parameter combined detection system for early screening of oral cancer, characterized in that: A method for implementing a multi-parameter combined detection method for early screening of oral cancer according to any one of claims 1 to 6, comprising: The edge data acquisition module is used to obtain the target user's oral multi-dimensional data based on the acquisition terminal; the target user's oral multi-dimensional data includes: oral image data, oral saliva miRNA data and oral multi-frequency impedance data; The edge node initial screening module is electrically connected to the edge data acquisition module. The edge node initial screening module is used to pre-set a lightweight model for classifying suspicious oral cancer based on the edge device, pre-process the target user's oral multi-dimensional data, and generate the target user's oral multimodal feature vector and initial screening oral status classification label; The cloud node screening module is wirelessly connected to the edge node initial screening module. The cloud node screening module is used to deploy the Transformer multimodal oral cancer status grading and classification model based on the cloud node, and jointly analyze the target user's oral multimodal feature vector and the initial screening oral status classification label to generate the target user's oral cancer status grading and classification.
Citation Information
Patent Citations
Intellectualized multifunction diagnostic device for oral diseases
CN1041523A
Oral cancer tumor staging method based on pathology and CT (Computed Tomography) multi-modal model
CN117830227A
Lightweight multi-mode in-vitro examination information fusion analysis method and system
CN118468212A
Health status assessment system applied to disease follow-up visit
CN119626556A
Power transmission line bird damage related bird classification method based on image recognition
CN119851311A
Cited By
Multi-sign fusion lung respiration monitoring system and method based on LSTM neural network
CN120859475A
Multi-mode oral cancer early screening and diagnosis and treatment integrated system
CN122291085A