Cipn information fusion method and system based on multi-modal information

By employing multimodal information fusion methods and artificial intelligence algorithms, the subjectivity and sensitivity issues in CIPN assessment have been resolved, enabling objective and quantitative assessment of CIPN, improving diagnostic accuracy and treatment outcomes, and promoting the development of personalized medicine.

CN119405277BActive Publication Date: 2025-12-09FUDAN UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411609221.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-12
Publication Date
2025-12-09
Estimated Expiration
2044-11-12

AI Technical Summary

Technical Problem

Existing technologies for assessing chemotherapy-induced peripheral neurotoxicity (CIPN) suffer from high subjectivity, lack of objectivity and sensitivity, making it difficult to provide a comprehensive and quantitative assessment, which affects treatment outcomes and patients' quality of life.

Method used

A multimodal information fusion approach is adopted, which includes assessing the test subjects' subjective questionnaires, walking videos, fine motor skills of fingers, neural conduction velocity, brain function and brain structure. This information is then integrated through artificial intelligence algorithms to form a comprehensive scoring system.

Benefits of technology

It enables objective and quantitative assessment of CIPN, improves diagnostic accuracy and treatment efficacy, allows for early identification and intervention, reduces the risk of treatment interruption, and improves quality of life and treatment outcomes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119405277B_ABST
    Figure CN119405277B_ABST
Patent Text Reader

Abstract

The application provides a CIPN information fusion method and system based on multi-modal information, comprising: evaluating a subjective questionnaire of a to-be-tested person; evaluating walking speed and gait according to a walking video of the to-be-tested person; evaluating finger fine motor skills of the to-be-tested person; evaluating nerve conduction velocity of the to-be-tested person; evaluating brain function of the to-be-tested person; evaluating brain structure of the to-be-tested person; and constructing information fusion according to the subjective questionnaire, the walking speed and gait, the finger fine motor skills, the nerve conduction velocity, the brain function and the brain structure. Compared with a traditional NCI-CTCAE grading system and an EORTC QLQ-CIPN20 questionnaire, the application overcomes subjectivity and repeatability problems, and improves diagnostic accuracy and treatment effect.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of medical assessment tools, in particular, to a CIPN information fusion method and system based on multi-modal information. BACKGROUND

[0002] Peripheral neurotoxicity caused by chemotherapy drugs (CIPN) is one of the common dose-limiting non-hematological toxicities, which presents as symmetrical sensory and motor dysfunction in bilateral distal extremities. Common symptoms include glove and sock numbness, pain, tingling, burning sensation, and sensory dullness. Sensory nerve dysfunction is the most common, which may be accompanied by decreased motor ability and autonomic nervous dysfunction. Traditional chemotherapy drugs such as platinum, taxanes, vinca alkaloids and protease inhibitors, as well as new ADC type antitumor drugs, can cause significant peripheral neurotoxicity. Severe CIPN not only affects the quality of life, but also may lead to treatment interruption, thereby affecting disease control.

[0003] CIPN has dose accumulation, and timely dose reduction or drug withdrawal can alleviate symptoms, but severe toxicity may persist for a long time. Studies have shown that about 20% of patients still have neurotoxicity after using taxanes for 18 months. Long-term toxicity not only reduces the quality of life, but also limits the choice of subsequent drugs, affecting treatment effect. Therefore, timely assessment and treatment of CIPN is crucial for malignant tumor treatment, especially for the goal of prolonging survival and improving quality of life.

[0004] Currently, CIPN assessment mainly relies on the NCI-CTCAE grading system, which assesses the severity of toxicity according to patient symptom description and clinician judgment, but has strong subjectivity and lacks objectivity. Although the EORTC QLQ-CIPN20 questionnaire solves the quantification problem, it still relies on patient self-reporting and cannot overcome the subjectivity and repeatability problems. Electromyography and nerve conduction velocity have objectivity and quantifiable characteristics, but only provide local information, have insufficient sensitivity, and are difficult to be used independently as a grading tool. Therefore, the present application designs an information collection system that can objectively and quantitatively assess the peripheral nerve and muscle function status, fine motor, gait and stride of patients, and develops an artificial intelligence algorithm to integrate these information to form a comprehensive scoring system. This system will help clinicians objectively and accurately assess the occurrence and development of CIPN, and provide an important reference for intervention measures.

[0005] Patent document CN117373671A discloses a patient report assessment tool for evaluating oxaliplatin-induced peripheral neuropathy, which comprises a main scale, a sub-scale and a scoring formula. The main scale evaluates the body organ symptoms, which are divided into grades according to the subjective feeling of the patient on the frequency of the occurrence of the neuropathy, and each grade corresponds to a relative score. The sub-scale includes the occurrence site of the most severe symptom, the duration of the most severe symptom and the impact of the most severe symptom on life. The invention can effectively improve the accuracy and sensitivity of the evaluation results of peripheral neuropathy by evaluating peripheral neuropathy and integrating the characteristics of multiple verified evaluation scales, and can provide a reference for the efficacy evaluation of clinical intervention strategies. However, the invention does not consider the correlation between multiple physiological indicators, has strong subjectivity, and cannot meet the demand for objective evaluation of neurotoxicity. SUMMARY

[0006] In view of the defects in the prior art, the purpose of the present application is to provide a CIPN information fusion method and system based on multi-modal information.

[0007] According to the CIPN information fusion method based on multi-modal information provided by the present application, the following steps are included:

[0008] Step S1: evaluating the subjective questionnaire of the testee;

[0009] Step S2: evaluating the walking speed and gait according to the walking video of the testee;

[0010] Step S3: evaluating the fine motor skills of the testee;

[0011] Step S4: evaluating the nerve conduction velocity of the testee;

[0012] Step S5: evaluating the brain function of the testee;

[0013] Step S6: evaluating the brain structure of the testee;

[0014] Step S7: information fusion according to the subjective questionnaire, walking speed and gait, fine motor skills, nerve conduction velocity, brain function and brain structure.

[0015] Preferably, in the step S1:

[0016] Step S1.1: designing a self-reporting scale, which includes the investigation of sensory symptoms, motor symptoms, autonomic nervous symptoms and the impact on life quality;

[0017] Step S1.2: data collection and preprocessing:

[0018] Collecting the self-reporting scale questionnaire data and normalizing all the data;

[0019]

[0020] wherein x i is the ith data in the self-report questionnaire data of the person to be tested, a is the minimum value of the questionnaire data, and b is the maximum value of the questionnaire data;

[0021] Step S1.3: Setting the index weight:

[0022] According to the opinions of clinical experts, the importance weight of each question in the self-report questionnaire of the person to be tested is determined:

[0023]

[0024] wherein (w i ) j is the importance weight set by the jth clinical expert for the ith index, the range is 1-10, N is the number of experts, and w i is the importance weight of the ith index;

[0025] Step S1.3: Comprehensive score calculation:

[0026] The comprehensive CIPN score is calculated by using the weighted average method, and the formula is as follows:

[0027]

[0028] wherein w i is the importance weight of the ith index, x i is the ith data in the CIPN20 questionnaire data, and s1 is the score of the person to be tested in this scale; M is the total number of indexes.

[0029] Preferably, in the step S2:

[0030] Step S2.1: Video preprocessing

[0031] K frames of images {T1,...,T K} are extracted from the video taken by the person to be tested;

[0032] Step S2.2: Gait feature extraction

[0033] The correlation between consecutive frames is used for background subtraction, the static background is removed, and the moving human body part is retained, specifically:

[0034] Step S2.2.1: Reading the current frame image T t and the previous frame image T t-1 , and calculating the difference image Diff t between the two frames:

[0035] Diff t = |T t - T t-1 |

[0036] Step S2.2.2: using hard threshold to binarize the difference image Diff t , pixels larger than threshold I are foreground:

[0037]

[0038] where F t is the t-th frame foreground image after removing background;

[0039] Using the emplaceAndPop function in the trained pose estimation model OpenPose to detect the toe position {(A1, B1),..., (AK, BK)} of each frame in the video on the foreground image; Where (AK, BK) are the coordinates of the left and right toe positions in the K-th frame, respectively. K K K K

[0040] Step S2.3: gait cycle recognition:

[0041] Calculate the distance between the toe positions in two consecutive frames:

[0042]

[0043] where D m is the distance between the toe positions in the m-th frame and the m+1-th frame.

[0044] Calculate the walking speed of the subject in the video:

[0045]

[0046] where s2 is the walking speed of the person to be tested, T is the number of seconds between the 1st frame and the Kth frame, and V is the walking speed.

[0047] Preferably, in the step S3:

[0048] By randomly appearing click points on the screen of the mobile phone and gradually shortening the time interval between the points, the reaction time of the user is recorded to evaluate the fine motor skills of the user, and the specific steps are as follows:

[0049] Step S3.1: define the screen as an area with a width of C and a height of D;

[0050] Step S3.2: at the end of each time interval t n , randomly select a screen position to appear a click point;​​​​

[0051] Step S3.3: reaction time calculation:

[0052] (Δt) n = t click - t start

[0053] wherein t start is the click point appearance time, t click is the click point click time, (Δt) n is the nth reaction speed;

[0054] Step S3.4: after each click, shorten the time interval of the point appearance:

[0055] t n+1 = a x t n

[0056] wherein a is a shortening coefficient, which is greater than 0 and less than 1; t n is the nth time interval;

[0057] Step S3.5: calculate the average reaction time of the user, record the sequence of multiple reaction times

[0058]

[0059] wherein s3 is the reaction time of the person to be tested, and n1 is the total number of time intervals.

[0060] Preferably, in the step S4:

[0061] The calculation method of the nerve conduction velocity is:

[0062]

[0063] wherein s4 is the nerve conduction velocity of the person to be tested, D is the distance between the stimulation points, ΔL = L1-L2 is the time of detecting the electrical signal between the distal end and the proximal end, L2 is the time point of detecting the electrical signal at the distal end, and L1 is the time point of detecting the electrical signal at the proximal end.

[0064] Preferably, in the step S5:

[0065] Step S5.1: data acquisition:

[0066] An N5-channel NIRS device is used to collect NIRS data of the prefrontal lobe function of N persons to be tested; each channel records the changes of the concentrations of oxyhemoglobin HbO and deoxyhemoglobin HbR, and the number of data recorded by each channel of the person to be tested is N6, and the signal in the jth channel of the ith person to be tested is:

[0067]

[0068] wherein HbO(i) j is the oxygenated hemoglobin signal in the jth channel of the ith person under test, is the N6th oxygenated hemoglobin signal in the jth channel of the ith person under test, HbR(i) j is the deoxygenated hemoglobin signal in the jth channel of the ith person under test, is the N6th deoxygenated hemoglobin signal in the jth channel of the ith person under test;

[0069] Step S5.2: Data denoising:

[0070] A band-pass filter is used to remove high-frequency noise and low-frequency drift in the signal, and a denoised signal is obtained, i.e.

[0071] HbO_filtered(i) j = BF(HbO(i) j , f low , f high )

[0072] HbR_filtered(i) j = BF(HbR(i) j , f low , f high )

[0073] wherein HbO_filtered(i) j is the filtered oxygenated hemoglobin signal in the jth channel of the ith person under test, HbR_filtered(i) i is the filtered deoxygenated hemoglobin signal in the jth channel of the ith person under test, BF is a band-pass filter, f low , f high are the cutoff frequencies;

[0074] Step S5.3: Brain function feature extraction

[0075] Extract the average oxygenated hemoglobin concentration change of all channels of the ith person under test:

[0076]

[0077] wherein N5 is the number of data recorded by the channel;

[0078] Extract the average deoxygenated hemoglobin concentration change of all channels of the ith person under test:

[0079]

[0080] extracting the total blood volume change s5(i) of the i-th person to be tested:

[0081] s5 = AHb0(i) + AHbR(i)

[0082] wherein mean(Hb0_filtered(i) j ) is the average value of the Hb0_filtered(i) j signal of different channels, and mean(HbR_filtered(i) j ) is the average value of the HbR_filtered(i) j signal of different channels.

[0083] Preferably, in the step S6:

[0084] Step S6.1: data collection:

[0085] performing magnetic resonance imaging (MRI) on N persons to be tested, setting scanning parameters;

[0086] Step S6.2: structural feature extraction:

[0087] segmenting the brain tissue into gray matter, white matter and cerebrospinal fluid using an automated tool FreeSurfer; calculating the gray matter volume (GMV) of the prefrontal lobe of the i-th person to be tested using the FreeSurfer toolkit, calculating the average cortical thickness (CT) of the prefrontal lobe region, and calculating the cortical surface area (CSA) of the prefrontal lobe region. i i i

[0088] Step S6.3: calculation of brain structural indicators:

[0089] s6 = GMV i + CT i + CSA i .

[0090] Preferably, in the step S7:

[0091] Step S7.1: collecting the evaluation indicators of N persons to be tested to form a data set P and the CIPN severity level Y of the person to be tested:

[0092] P = {p1,...,p l …,p N},

[0093] Y = {y1,...,y l …,y N},

[0094] wherein p​​​l = {s l1 , s l2 , s l3 , s l4 , s l5 , s l6} T is the evaluation index of the lth person to be tested, Y is the CIPN severity level data matrix of N persons to be tested, the size is 5xN, y l is the CIPN severity level of the lth person to be tested;

[0095] Step S7.2: data standardization, the specific construction method is as follows:

[0096] Step S7.2.1: divide the data of the above N persons to be tested into training set, validation set and test set according to the preset proportion;

[0097] Step S7.2.2: normalize all the above data, the formula is as follows:

[0098]

[0099] Wherein, s ij is the jth feature of the ith person to be tested, i∈{1, 2, 3,..., N}, j∈{1, 2, 3, 4}, sn ij is the jth normalized feature of the ith person to be tested; s minj is the minimum value of the jth feature, s maxj is the maximum value of the jth feature;

[0100] Step S7.3: use Mamba model to construct multi-classification model:

[0101] Step S7.3.1: for the feature vector p i = {sn i1 , sn i2 , sn i3 , sn i4 , sn i5 , sn i6} T of each person to be tested i in the data set P, carry out embedding processing and convert it into high-dimensional feature representation h i ;

[0102] - use linear transformation to map the original normalized feature to high-dimensional embedding space:

[0103] h i = W e ·p i +b e

[0104] wherein h i is the embedded feature vector with size d x 1; W e is the weight matrix of the embedding layer with size d x 6; b e is the bias vector of the embedding layer with size d x 1; and d is the size of the embedding dimension.

[0105] The multi-head attention mechanism is applied to capture the relationship between each normalized feature s ij to enhance the feature representation.

[0106] The embedded feature vector h i is calculated based on the query Q, the key K and the value V:

[0107] Q = W q · h i , K = W k · h i , and V = W v · h i

[0108] wherein Q is the query vector with size d x 1; K is the key vector with size d x 1; V is the value vector with size d x 1; W q , W k and W v are the weight matrices of the query, the key and the value respectively, each with size d x d.

[0109] The self-attention score Attention(Q, K, V) is calculated.

[0110]

[0111] wherein d1 is the dimension of the key vector for scaling the dot product result.

[0112] For the multi-head attention mechanism, the above operation is extended to multiple heads, and multiple attention scores are calculated in parallel.

[0113] MultiHead(Q, K, V) = Concat(head1, …, head h )· W o

[0114] wherein head i represents the calculation result of the i-th attention head; W o is the weight matrix of the output with size d x d.

[0115] After the multi-head attention mechanism, NBA is introduced to extract local and global features from the feature vector h through a set of trainable atomic mapping layers a mNonlinear feature mapping is performed:

[0116]

[0117] where a m is the atomic-level feature representation with size d2x1; W a and b a are the weight matrix and bias vector of the mth atomic mapping layer, respectively; σ is the ReLU nonlinear activation function; and d2 is the dimension size of the atomic-level feature.

[0118] All atomic-level features are integrated into a global feature representation through a pooling operation:

[0119]

[0120] where is the global feature representation extracted by the neural atomic bag with size d2x1; k is the number of atomic mapping layers; and MaxPooling is the max pooling layer.

[0121] Step S7.3.3: The feature vector extracted through the multi-head attention mechanism and the neural atomic bag: The prediction of the CIPN severity level is performed through the classification head.

[0122] The classification head is usually one or more fully connected layers, and the Softmax function is used to output the category probability:

[0123]

[0124] where W c is the weight matrix of the classification head with size Cxd2; b c is the bias vector of the classification head with size Cx1; and y is the CIPN severity level prediction result of the ith person to be tested with size Cx1; and C is the number of CIPN severity levels.

[0125] Step S7.3.4: Model training and updating:

[0126] The model is trained using the cross-entropy loss function, and the model parameters are updated through the backpropagation algorithm. The formula of the cross-entropy loss function is:

[0127]

[0128] where y ic is the true label of the ith person to be tested in category c; and y is the predicted probability of the ith person to be tested in category c; and C is the number of CIPN severity levels.

[0129] Step S7.3.5: Repeat training until convergence

[0130] Steps S7.3.1 to S7.3.4 are repeatedly performed to gradually optimize the model parameters until the performance on the validation set improves inefficiently than a preset standard.

[0131] According to the CIPN information fusion system based on multi-modal information provided by the application, comprising:

[0132] Module M1: evaluating the subjective questionnaire of the testee;

[0133] Module M2: evaluating the walking speed and gait according to the walking video of the testee;

[0134] Module M3: evaluating the fine motor skills of the testee;

[0135] Module M4: evaluating the nerve conduction velocity of the testee;

[0136] Module M5: evaluating the brain function of the testee;

[0137] Module M6: evaluating the brain structure of the testee;

[0138] Module M7: information fusion according to the subjective questionnaire, walking speed and gait, fine motor skills, nerve conduction velocity, brain function and brain structure.

[0139] Preferably, in the module M1:

[0140] Module M1.1: designing a self-reporting scale, containing surveys on sensory symptoms, motor symptoms, autonomic nervous symptoms and the impact on quality of life;

[0141] Module M1.2: data collection and preprocessing:

[0142] Collecting self-reporting scale questionnaire data and normalizing all data;

[0143]

[0144] Wherein x i is the ith data in the self-reporting scale questionnaire data of the testee, a is the minimum value of the questionnaire data, and b is the maximum value of the questionnaire data;

[0145] Module M1.3: index weight setting:

[0146] According to the opinions of clinical experts, the importance weight of each testee's self-reporting scale questionnaire problem is determined:

[0147]

[0148] wherein, w i ) j is the importance weight of the ith index set by the jth clinical expert, and the range is 1-10, N is the number of experts, w i is the importance weight of the ith index;

[0149] Module M1.3: Comprehensive score calculation:

[0150] The comprehensive CIPN score is calculated by using the weighted average method, and the formula is as follows:

[0151]

[0152] wherein, w i is the importance weight of the ith index, x i is the ith data in the CIPN 20 questionnaire data, s1 is the score of the person being tested in this scale; M is the total number of indexes;

[0153] In the module M2:

[0154] Module M2.1: Video preprocessing

[0155] K frames of images {T1,...,T K} are extracted from the video taken by the person being tested;

[0156] Module M2.2: Gait feature extraction

[0157] The correlation between consecutive frames is used for background subtraction to remove the static background and retain the moving human body part, specifically:

[0158] Module M2.2.1: Read the current frame image T t and the previous frame image T t-1 , and calculate the difference image Diff t between the two frames:

[0159] Diff t = |T t -T t-1 |

[0160] Module M2.2.2: Binaryzation of the difference image Diff t using a hard threshold, and the pixels greater than the threshold I are foreground:

[0161]

[0162] wherein, F t is the tth frame of foreground image after removing the background;

[0163] The emplaceAndPop function in the trained pose estimation model OpenPose is used to detect the toe position of each frame in the foreground image video {(A1, B1),..., (A K , B K )}; wherein (A K , B K ) are the coordinates of the left and right toe positions in the Kth frame, respectively;

[0164] Module M2.3: Gait cycle recognition:

[0165] Calculate the distance between the toe positions in two consecutive frames:

[0166]

[0167] where D m is the distance between the toes of the mth frame and the m+1th frame;

[0168] Calculate the walking speed of the subject in the video:

[0169]

[0170] where s2 is the walking speed of the person to be tested, T is the number of seconds between the 1st frame and the Kth frame, and V is the walking speed;

[0171] In the module M3:

[0172] By randomly appearing click points on the phone screen and gradually shortening the time interval between points, the user's reaction time is recorded to evaluate the user's fine motor skills, the specific steps are as follows:

[0173] Module M3.1: Define the screen as an area with a width of C and a height of D;

[0174] Module M3.2: At the end of each time interval t n , randomly select a screen position to appear a click point;

[0175] Module M3.3: Reaction time calculation:

[0176] (Δt) n = t click -t start

[0177] where t start is the appearance time of the click point, t click is the click time of the click point, and (Δt) n is the nth reaction speed;

[0178] Module M3.4: After each click, shorten the time interval between points:

[0179] t n+1 = a x t n

[0180] wherein a is a shortening coefficient, greater than 0 and less than 1; t n is the nth time interval;

[0181] Module M3.5: calculating the average reaction time of the user, recording the sequence of multiple reaction times

[0182]

[0183] wherein s3 is the reaction time of the person to be tested, and n1 is the total number of time intervals;

[0184] In the module M4:

[0185] The calculation method of the nerve conduction velocity is:

[0186]

[0187] wherein s4 is the nerve conduction velocity of the person to be tested, D is the distance between the stimulation points, ΔL = L1-L2 is the time of detecting the electrical signal between the distal end and the proximal end, L2 is the time point of detecting the electrical signal at the distal end, and L1 is the time point of detecting the electrical signal at the proximal end;

[0188] In the module M5:

[0189] Module M5.1: data acquisition:

[0190] An N5-channel NIRS device is used to collect NIRS data of the prefrontal lobe function of N persons to be tested; each channel records the changes in the concentrations of oxyhemoglobin HbO and deoxyhemoglobin HbR, and the number of data recorded by each channel of the person to be tested is N6, and the signal in the jth channel of the ith person to be tested is:

[0191]

[0192] wherein HbO(i) j is the oxyhemoglobin signal in the jth channel of the ith person to be tested, is the N6th oxyhemoglobin signal in the jth channel of the ith person to be tested, HbR(i) j is the deoxyhemoglobin signal in the jth channel of the ith person to be tested, is the N6th deoxyhemoglobin signal in the jth channel of the ith person to be tested;

[0193] Module M5.2: data denoising:

[0194] Using a band-pass filter to remove high-frequency noise and low-frequency drift in the signal, obtaining the de-noised signal, i.e.

[0195] HbO_filtered(i) j =BF(HbO(i) j ,f low ,f high )

[0196] HbR_filtered(i) j =BF(HbR(i) j ,f low ,f high )

[0197] Wherein, HbO_filtered(i) j is the filtered oxygenated hemoglobin signal of the ith person to be tested in the jth channel, HbR_filtered(i) j is the filtered deoxygenated hemoglobin signal of the ith person to be tested in the jth channel, BF is a band-pass filter, f low ,f high is the cutoff frequency;

[0198] Module M5.3: brain function feature extraction

[0199] Extract the average oxygenated hemoglobin concentration change of all channels of the ith person to be tested:

[0200]

[0201] Wherein, N5 is the number of data recorded by the channel;

[0202] Extract the average deoxygenated hemoglobin concentration change of all channels of the ith person to be tested:

[0203]

[0204] Extract the total blood volume change s5(i) of the ith person to be tested:

[0205] s5=ΔHbO(i)+ΔHbR(i)

[0206] Wherein, mean(HbO_filtered(i) j ) is the average value of different channels of HbO_filtered(i) j signal, mean(HbR_filtered(i) j ) is the average value of different channels of HbR_filtered(i) j signal;

[0207] In the module M6:

[0208] Module M6.1: Data acquisition:

[0209] Perform magnetic resonance imaging (MRI) on N subjects to be tested, and set the scanning parameters;

[0210] Module M6.2: Structural feature extraction:

[0211] Segment the brain tissue into gray matter, white matter, and cerebrospinal fluid using the automated tool FreeSurfer; use the FreeSurfer toolkit to calculate the gray matter volume (GMV) of the prefrontal lobe of the ith subject to be tested i Calculate the average cortical thickness (CT) of the prefrontal region i Calculate the cortical surface area (CSA) of the prefrontal region i ;

[0212] Module M6.3: Calculation of brain structural indicators:

[0213] s6 = CMV i + CT i + CSA i

[0214] In the module M7:

[0215] Module M7.1: Collect the assessment indicator set data set P and the CIPN severity level Y of N subjects to be tested:

[0216] P = {p1,...,p l …,p N},

[0217] Y = {y1,...,y l …,y N},

[0218] where p l = {s l1 ,s l2 ,s l3 ,s l4 ,s l5 ,s l6} T is each assessment indicator of the ith subject to be tested, Y is the CIPN severity level data matrix of N subjects to be tested, which has a size of 5 × N, y l is the CIPN severity level of the ith subject to be tested;

[0219] Module M7.2: Data standardization, the specific construction method is as follows:

[0220] Module M7.2.1: Divide the data of the above N persons to be tested into training set, validation set and test set according to the preset proportion;

[0221] Module M7.2.2: Normalize all the above data, the formula is as follows:

[0222]

[0223] Wherein, s ij is the jth feature of the ith person to be tested, i∈{1,2,3,...,N}, j∈{1,2,3,4}, sn ij is the jth normalized feature of the ith person to be tested; s minj is the minimum value of the jth feature, s maxj is the maximum value of the jth feature;

[0224] Module M7.3: Use Mamba model to construct multi-classification model:

[0225] Module M7.3.1: Embedding processing is performed on the feature vector p i of each person to be tested i in the data set P i1 ={sn i2 ,sn i3 ,sn i4 ,sn i5 ,sn i6} T , and convert it into high-dimensional feature representation h i ;

[0226] - Use linear transformation to map the original normalized features to high-dimensional embedding space:

[0227] h i =W e ·p i +b e

[0228] Wherein, h i is the embedded feature vector, with the size of d×1; W e is the weight matrix of the embedding layer, with the size of d×6; b e is the bias vector of the embedding layer, with the size of d×1; d is the size of embedding dimension;

[0229] Apply multi-head attention mechanism to capture the relationship between each normalized feature s ij , and enhance the feature representation;

[0230] Calculate the query Q, key K and value V for the embedded feature vector h i :

[0231] Q=Wq ·h i ,K=W k ·h i ,V=W v ·h i

[0232] where Q is the query vector with size d x 1; K is the key vector with size d x 1; V is the value vector with size d x 1; W q , W k and W v are the weight matrices of query, key and value respectively, all with size d x d;

[0233] Compute self-attention score Attention(Q, K, V)

[0234]

[0235] where d1 is the dimension of the key vector, used to scale the dot product result;

[0236] For multi-head attention mechanism, the above operation is extended to multi-head, and multiple attention scores are computed in parallel:

[0237] MultiHead(Q, K, V) = Concat(head1, …, head h )·W o

[0238] where head i represents the computation result of the i-th attention head; W o is the weight matrix of the output, with size d x d;

[0239] Module M7.3.2: Neural Atom Bag (NBA) feature extraction

[0240] After the multi-head attention mechanism, NBA is introduced to extract local and global features. The feature vector after multi-head attention processing is mapped through a set of trainable atom mapping layers a m to perform nonlinear feature mapping:

[0241]

[0242] where a m is the atom-level feature representation with size d2 x 1; W a and b a are the weight matrix and bias vector of the m-th atom mapping layer respectively; σ is the ReLU nonlinear activation function; d2 is the dimension size of the atom-level feature;

[0243] All atom-level features are integrated into global feature representation through pooling operation:

[0244]

[0245] wherein, is the global feature representation extracted by the neural atomic bag, with the size of d2x1; k is the number of atomic mapping layers; MaxPooling is the maximum pooling layer;

[0246] Module M7.3.3: the feature vector after the multi-head attention mechanism and the neural atomic bag extraction: prediction of the CIPN severity level through the classification head;

[0247] The classification head is usually one or more fully connected layers, and outputs the category probability using the Softmax function:

[0248]

[0249] wherein, W c is the weight matrix of the classification head, with the size of Cxd2; b c is the bias vector of the classification head, with the size of Cx1; is the CIPN severity level prediction result of the ith person to be tested, with the size of Cx1; C is the number of CIPN severity levels;

[0250] Module M7.3.4: model training and updating:

[0251] The model is trained using the cross-entropy loss function, and the model parameters are updated through the back propagation algorithm, and the formula of the cross-entropy loss function is:

[0252]

[0253] wherein, y ic is the true label of the ith person to be tested in category c; is the prediction probability of the ith person to be tested in category c; C is the number of CIPN severity levels;

[0254] Module M7.3.5: repeat training until convergence

[0255] Modules M7.3.1 to M7.3.4 are repeatedly executed, and the model parameters are gradually optimized until the performance improvement efficiency on the validation set is lower than the preset standard.

[0256] Compared with the prior art, the present application has the following beneficial effects:

[0257] 1、The application provides a comprehensive, quantitative and objective CIPN evaluation method by using artificial intelligence algorithms and precise grading systems to objectively and quantitatively evaluate the peripheral nerve and muscle function status, fine motor skills, gait and stride of patients. Compared with the traditional NCI-CTCAE grading system and EORTC QLQ-CIPN20 questionnaire, the application overcomes the subjectivity and repeatability problems, improves the diagnostic accuracy and treatment effect.

[0258] 2、The application has high sensitivity and specificity, can early identify and evaluate the occurrence of CIPN, help clinicians take timely intervention measures, reduce patient pain and treatment interruption risk, and improve overall treatment effect. Through continuous monitoring and evaluation, the system can be personalized according to the individual situation of the patient, improve the quality of life and survival time in the treatment of advanced tumors.

[0259] 3、The application provides a scientific basis for drug research and development and clinical practice. The system can be used for drug safety and effectiveness evaluation, optimize drug dosage and treatment plan, reduce neurotoxicity risk, and promote the research and application of new antitumor drugs. The system has high automation level, reduces manual operation and subjective judgment error, and improves clinical work efficiency.

[0260] 4、The application is not only suitable for CIPN evaluation, but also can be applied to the evaluation of other nervous system diseases and injuries. Through optimization and expansion of algorithms and evaluation indexes, the system provides precise diagnosis and treatment support for more patients, promotes the development of personalized medicine. The system can obtain and analyze patient health data in real time, provide high-quality data support for clinical research, and promote the progress of medical research.

[0261] 5、The application provides a powerful clinical tool to help doctors better manage and evaluate peripheral neurotoxicity in patients during treatment, optimize treatment plans, reduce patient pain and improve quality of life. By building a comprehensive, objective and repeatable evaluation system, the overall treatment effect is improved, and the progress of tumor treatment is promoted. BRIEF DESCRIPTION OF DRAWINGS

[0262] Other features, objects and advantages of the application will become more apparent through reading the following detailed description of non-limiting embodiments, made with reference to the following drawings:

[0263] Figure 1 The flowchart of the application is shown. DETAILED DESCRIPTION

[0264] The application will be described in detail below with specific examples. The following examples will help those skilled in the art to further understand the application, but in no way limit the application. It should be noted that for those skilled in the art, without departing from the concept of the application, a number of changes and improvements can be made. These are within the scope of the application.

[0265] Example 1:

[0266] According to the CIPN information fusion method based on multi-modal information provided by the application, as shown in the formula (1), the method comprises the following steps: Figure 1

[0267] Step S1: evaluating the subjective questionnaire of the testee;

[0268] Specifically, in the step S1:

[0269] Step S1.1: designing a self-reporting scale, including investigation of sensory symptoms, motor symptoms, autonomic nervous symptoms and impact on quality of life;

[0270] Step S1.2: data collection and preprocessing:

[0271] Collecting self-reporting scale questionnaire data, and normalizing all data;

[0272]

[0273] Wherein x i is the i-th data in the self-reporting scale questionnaire data of the testee, a is the minimum value of the questionnaire data, and b is the maximum value of the questionnaire data;

[0274] Step S1.3: index weight setting:

[0275] According to the opinions of clinical experts, the importance weight of each testee's self-reporting scale questionnaire problem is determined:

[0276]

[0277] Wherein, (w i ) j is the importance weight set by the j-th clinical expert for the i-th index, ranging from 1 to 10, N is the number of experts, and w i is the importance weight of the i-th index;

[0278] Step S1.3: comprehensive score calculation:

[0279] The comprehensive CIPN score is calculated by using the weighted average method, and the formula is as follows:

[0280]

[0281] wherein w i is the importance weight of the ith index, x i is the ith data in the CIPN20 questionnaire data, s1 is the score of the to-be-tested person on this scale; M is the total number of indexes.

[0282] Step S2: evaluating the walking speed and gait according to the walking video of the to-be-tested person;

[0283] Specifically, in the step S2:

[0284] Step S2.1: video preprocessing

[0285] K frames of images {T1,...,T K} are extracted from the video taken of the to-be-tested person;

[0286] Step S2.2: gait feature extraction

[0287] The background is removed using the correlation between consecutive frames, and the moving human body part is retained, specifically:

[0288] Step S2.2.1: reading the current frame image T t and the previous frame image T t-1 , and calculating the difference image Diff t between the two frames:

[0289] Diff t = |T t -T t-1 |

[0290] Step S2.2.2: using a hard threshold to binarize the difference image Diff t , and the pixels greater than the threshold I are foreground:

[0291]

[0292] wherein F t is the tth frame of foreground image after removing the background;

[0293] The emplaceAndPop function in the trained posture estimation model OpenPose is used to detect the toe position {(A1,B1),...,(A K ,B K )} of each frame in the video from the foreground image; wherein (A K ,B K ) are the coordinates of the left and right toe positions in the Kth frame, respectively.

[0294] Step S2.3: gait cycle recognition:

[0295] Calculate the distance between the toe positions in two consecutive frames:

[0296]

[0297] where Dm+1is the distance between the toe positions in the mth frame and the m+1th frame; m

[0298] Calculate the walking speed of the subject in the video:

[0299]

[0300] where s2is the walking speed of the subject, T is the time in seconds between the 1st frame and the Kth frame, and V is the walking speed.

[0301] Step S3: Evaluate the fine motor skills of the subject;

[0302] Specifically, in the step S3:

[0303] By randomly appearing click points on the screen of the mobile phone and gradually shortening the time interval between the points, the reaction time of the user is recorded to evaluate the fine motor skills of the user, and the specific steps are as follows:

[0304] Step S3.1: Define the screen as an area with a width of C and a height of D;

[0305] Step S3.2: At each time interval t n , randomly select a screen position to appear a click point;

[0306] Step S3.3: Reaction time calculation:

[0307] (Δt) n = t click -t start

[0308] where t start is the appearance time of the click point, t click is the click time of the click point, and (Δt) n is the nth reaction speed;

[0309] Step S3.4: After each click, shorten the time interval between the points:

[0310] t n+1 = α × t n

[0311] where α is a shortening coefficient, greater than 0 and less than 1; t n is the nth time interval;

[0312] ​Step S3.5: Calculate the average reaction time of the user, record the sequence of multiple reaction times {(At)1,..., (At) n1};

[0313]

[0314] Wherein, s3 is the reaction time of the person to be tested, n1 is the total number of time intervals.

[0315] Step S4: Evaluate the nerve conduction velocity of the person to be tested;

[0316] Specifically, in the step S4:

[0317] The calculation method of the nerve conduction velocity is:

[0318]

[0319] Wherein, s4 is the nerve conduction velocity of the person to be tested, D is the distance between the stimulation points, AL=L1-L2 is the time of detecting the electrical signal between the distal end and the proximal end, L2 is the time point of detecting the electrical signal by the distal end, and L1 is the time point of detecting the electrical signal by the proximal end.

[0320] Step S5: Evaluate the brain function of the person to be tested;

[0321] Specifically, in the step S5:

[0322] Step S5.1: Data acquisition:

[0323] Use the N5-channel NIRS device to collect NIRS data of the prefrontal lobe function of N persons to be tested; record the changes of the concentrations of oxyhemoglobin HbO and deoxyhemoglobin HbR in each channel, and the number of data recorded by each channel of the person to be tested is N6, and the signal in the jth channel of the ith person to be tested is:

[0324]

[0325] Wherein, HbO(i) j is the oxyhemoglobin signal in the jth channel of the ith person to be tested, is the N6th oxyhemoglobin signal in the jth channel of the ith person to be tested, HbR(i) j is the deoxyhemoglobin signal in the jth channel of the ith person to be tested, is the N6th deoxyhemoglobin signal in the jth channel of the ith person to be tested;

[0326] Step S5.2: Data denoising:

[0327] The band-pass filter is used to remove the high-frequency noise and low-frequency drift in the signal, and the de-noised signal is obtained, that is

[0328] HbO_filtered(i) j =BF(HbO(i) j ,f low ,f high )

[0329] HbR_filtered(i) j =BF(HbR(i) j ,f low ,f high )

[0330] Wherein, HbO_filtered(i) j is the filtered oxygenated hemoglobin signal of the ith person to be tested in the jth channel, HbR_filtered(i) j is the filtered deoxygenated hemoglobin signal of the ith person to be tested in the jth channel, BF is a band-pass filter, f low ,f high is the cutoff frequency;

[0331] Step S5.3: brain function feature extraction

[0332] Extract the average oxygenated hemoglobin concentration change of all channels of the ith person to be tested:

[0333]

[0334] Wherein, N5 is the number of data recorded by the channel;

[0335] Extract the average deoxygenated hemoglobin concentration change of all channels of the ith person to be tested:

[0336]

[0337] Extract the total blood volume change s5(i) of the ith person to be tested:

[0338] s5=ΔHbO(i)+ΔHbR(i)

[0339] Wherein, mean(HbO_filtered(i) j ) is the average value of different channels of HbO_filtered(i) j signal, mean(HbR_filtered(i) j ) is the average value of different channels of HbR_filtered(i) j signal.

[0340] Step S6: evaluating the brain structure of the to-be-tested person;

[0341] Specifically, in the step S6:

[0342] Step S6.1: data collection:

[0343] Performing magnetic resonance imaging (MRI) on N to-be-tested persons, and setting scanning parameters;

[0344] Step S6.2: structural feature extraction:

[0345] Segmenting the brain tissue into gray matter, white matter and cerebrospinal fluid using an automated tool FreeSurfer; calculating the gray matter volume (GMV) of the prefrontal lobe of the ith to-be-tested person using the FreeSurfer toolkit i , calculating the average cortical thickness (CT) of the prefrontal lobe region i , and calculating the cortical surface area (CSA) of the prefrontal lobe region i ;

[0346] Step S6.3: calculating the brain structural index:

[0347] s6=GMV i +CT i +CSA i .

[0348] Step S7: constructing a precise grading system according to subjective questionnaires, walking speed and gait, finger fine motor skills, nerve conduction velocity, brain function and brain structure;

[0349] Specifically, in the step S7:

[0350] Step S7.1: collecting the evaluation index of N to-be-tested persons to form a data set P and the CIPN severity level Y of the to-be-tested person:

[0351] P={p1,...,p l …,p N},

[0352] Y={y1,...,y l …,y N},

[0353] wherein p l ={s l1 ,s l2 ,s l3 ,s l4 ,s l5 ,s l6} T is each evaluation index of the lth to-be-tested person, and Y is a CIPN severity level data matrix of N to-be-tested persons, with a size of 5xN, yl This is the severity level of CIPN for the lth person being tested;

[0354] Step S7.2: Data standardization, the specific construction method is as follows:

[0355] Step S7.2.1: Divide the data of the above N test subjects into training set, validation set and test set according to a preset ratio;

[0356] Step S7.2.2: Normalize all the above data using the following formula:

[0357]

[0358] Among them, s ij Let sn be the j-th feature of the i-th test subject, i∈{1,2,3,...,N}, j∈{1,2,3,4}. ij s is the j-th normalized feature of the i-th test subject; minj Let s be the minimum value of the j-th feature. maxj The maximum value of the j-th feature;

[0359] Step S7.3: Construct a multi-class classification model using the Mamba model:

[0360] Step S7.3.1: For each test subject i in the dataset P, the feature vector p i ={sn i1 ,sn i2 ,sn i3 ,sn i4 ,sn i5 ,sn i6} T Embedding is performed to convert it into a high-dimensional feature representation h. i ;

[0361] - Use a linear transformation to map the original normalized features to a high-dimensional embedding space:

[0362] h i =W e ·p i +b e

[0363] Among them, h i It is the embedded feature vector with a size of d×1; W e This is the weight matrix of the embedding layer, with a size of d×6; b e is the bias vector of the embedding layer, with a size of d×1; d is the size of the embedding dimension.

[0364] Multi-head attention mechanism is applied to capture each normalized feature s ijrelationship between them, enhancing the feature representation;

[0365] After embedding, the feature vector h i Calculate the query Q, key K and value V:

[0366] Q = W q ·h i ,K = W k ·h i ,V = W v ·h i

[0367] Where Q is the query vector, with size d x 1; K is the key vector, with size d x 1; V is the value vector, with size d x 1; W q , W k and W v are the weight matrices of the query, key and value respectively, all with size d x d;

[0368] Calculate the self-attention score Attention(Q, K, V)

[0369]

[0370] Where d1 is the dimension of the key vector, used to scale the dot product result;

[0371] For multi-head attention mechanism, the above operation is extended to multi-head, and multiple attention scores are calculated in parallel:

[0372] MultiHead(Q, K, V) = Concat(head1, …, head h )·W o

[0373] Where head i represents the calculation result of the i-th attention head; W o is the weight matrix of the output, with size d x d;

[0374] Step S7.3.2: Neural Atom Bag (NBA) feature extraction

[0375] After the multi-head attention mechanism, NBA is introduced to extract local and global features, and the feature vector is processed by a set of trainable atom mapping layers a m Nonlinear feature mapping:

[0376]

[0377] Where a m is the atomic-level feature representation, with size d2 x 1; W a and b aWm and Bm are the weight matrix and bias vector of the m-th atom mapping layer, respectively; σ is the ReLU nonlinear activation function; d2 is the dimension size of the atomic-level feature;

[0378] All atomic-level features are integrated into global feature representation through pooling operation:

[0379]

[0380] wherein, is the global feature representation extracted by the neural atom bag, with a size of d2x1; k is the number of atom mapping layers; MaxPooling is the maximum pooling layer;

[0381] Step S7.3.3: The feature vector extracted through the multi-head attention mechanism and the neural atom bag: The prediction of the CIPN severity level is performed through the classification head;

[0382] The classification head is usually one or more fully connected layers, and the Softmax function is used to output the category probability:

[0383]

[0384] wherein, W c is the weight matrix of the classification head, with a size of Cxd2; B c is the bias vector of the classification head, with a size of Cx1; is the CIPN severity level prediction result of the i-th person to be tested, with a size of Cx1; C is the number of CIPN severity levels;

[0385] Step S7.3.4: Model training and updating:

[0386] The model is trained using the cross-entropy loss function, and the model parameters are updated through the backpropagation algorithm, and the formula of the cross-entropy loss function is:

[0387]

[0388] wherein, y ic is the true label of the i-th person to be tested in category c; is the predicted probability of the i-th person to be tested in category c; C is the number of CIPN severity levels;

[0389] Step S7.3.5: Repeat the training until convergence

[0390] Repeat steps S7.3.1 to S7.3.4 to gradually optimize the model parameters until the performance improvement efficiency on the validation set is lower than the preset standard, and the trained MANBA model is used to predict the CIPN severity level of new persons to be tested.

[0391] Example 2:

[0392] Example 2 is a preferred example of Example 1, to more specifically illustrate the present application.

[0393] The present application also provides a CIPN information fusion system based on multi-modal information, which can be realized by executing the process steps of the CIPN information fusion method based on multi-modal information, that is, the CIPN information fusion method based on multi-modal information can be understood by those skilled in the art as a preferred embodiment of the CIPN information fusion system based on multi-modal information.

[0394] According to the present application, a CIPN information fusion system based on multi-modal information is provided, comprising:

[0395] Module M1: evaluating the subjective questionnaire of the testee;

[0396] Specifically, in the module M1:

[0397] Module M1.1: designing a self-reporting scale, including investigation of sensory symptoms, motor symptoms, autonomic nervous symptoms and impact on quality of life;

[0398] Module M1.2: data collection and preprocessing:

[0399] Collecting self-reporting scale questionnaire data, and normalizing all data;

[0400]

[0401] wherein x i is the ith data in the self-reporting scale questionnaire data of the testee, a is the minimum value of the questionnaire data, and b is the maximum value of the questionnaire data;

[0402] Module M1.3: index weight setting:

[0403] According to the opinions of clinical experts, the importance weight of each testee's self-reporting scale questionnaire problem is determined:

[0404]

[0405] wherein, (w i ) j is the importance weight set by the jth clinical expert for the ith index, ranging from 1 to 10, N is the number of experts, and w i is the importance weight of the ith index;

[0406] Module M1.3: comprehensive score calculation:

[0407] The comprehensive CIPN score is calculated by using the weighted average method, and the formula is as follows:

[0408]

[0409] Wherein, w i is the importance weight of the i-th index, x i is the i-th data in the CIPN20 questionnaire data, s1 is the score of the to-be-tested personnel on this scale; M is the total number of indexes;

[0410] Module M2: evaluating the walking speed and gait according to the walking video of the to-be-tested personnel;

[0411] In the module M2:

[0412] Module M2.1: video preprocessing

[0413] K frames of images {T1,...,T K} are extracted from the video taken of the to-be-tested personnel;

[0414] Module M2.2: gait feature extraction

[0415] The correlation between consecutive frames is used for background subtraction, removing the static background and retaining the moving human body part, specifically:

[0416] Module M2.2.1: reading the current frame image T t and the previous frame image T t-1 , and calculating the difference image Diff t between the two frames:

[0417] Diff t = |T t -T t-1 |

[0418] Module M2.2.2: using a hard threshold to binarize the difference image Diff t , and the pixels greater than the threshold I are foreground:

[0419]

[0420] Wherein, F t is the t-th frame of foreground image after removing the background;

[0421] The emplaceAndPop function in the trained posture estimation model OpenPose is used to detect the toe position {(A1,B1),...,(A K ,B K )} of each frame in the video of the foreground image; wherein, (A K ,B K) are the coordinates of the left and right toe positions in the Kth frame, respectively;

[0422] Module M2.3: Gait cycle recognition:

[0423] Calculate the distance between the toe positions in two consecutive frames:

[0424]

[0425] where Dm is the distance between the toe positions in the mth frame and the m+1th frame; m

[0426] Calculate the walking speed of the subject in the video:

[0427]

[0428] where s2 is the walking speed of the subject, T is the time in seconds between the 1st frame and the Kth frame, and V is the walking speed;

[0429] Module M3: Assess the fine motor skills of the subject;

[0430] In the module M3:

[0431] By randomly appearing click points on the phone screen and gradually shortening the time interval between points, record the user's reaction time, and thus assess the user's fine motor skills, the specific steps are as follows:

[0432] Module M3.1: Define the screen as an area with a width of C and a height of D;

[0433] Module M3.2: At the end of each time interval t n , randomly select a screen position to appear a click point;

[0434] Module M3.3: Reaction time calculation:

[0435] (Δt) n = t click - t start

[0436] where t start is the appearance time of the click point, t click is the click time of the click point, and (Δt) n is the nth reaction speed; Module M3.4: After each click, shorten the time interval between points:

[0437] t n+1 = α × t n

[0438] where α is a shortening coefficient, greater than 0 and less than 1; t n ​the nth time interval;

[0439] Module M3.5: Calculate the average reaction time of the user, record the sequence of multiple reaction times

[0440]

[0441] wherein s3 is the reaction time of the person to be tested, and n1 is the total number of time intervals;

[0442] Module M4: Evaluate the nerve conduction velocity of the person to be tested;

[0443] In the module M4:

[0444] The calculation method of the nerve conduction velocity is:

[0445]

[0446] wherein s4 is the nerve conduction velocity of the person to be tested, D is the distance between the stimulation points, ΔL=L1-L2 is the time of detecting the electrical signal between the distal end and the proximal end, L2 is the time point of detecting the electrical signal by the distal end, and L1 is the time point of detecting the electrical signal by the proximal end;

[0447] Module M5: Evaluate the brain function of the person to be tested;

[0448] In the module M5:

[0449] Module M5.1: Data acquisition:

[0450] Use an N5-channel NIRS device to collect NIRS data of the prefrontal lobe function of N persons to be tested; record the changes of the concentrations of oxyhemoglobin HbO and deoxyhemoglobin HbR for each channel, and the number of data recorded by each channel of the person to be tested is N6, and the signal in the jth channel of the ith person to be tested is:

[0451]

[0452] wherein HbO(i) j is the oxyhemoglobin signal in the jth channel of the ith person to be tested, is the N6th oxyhemoglobin signal in the jth channel of the ith person to be tested, HbR(i) j is the deoxyhemoglobin signal in the jth channel of the ith person to be tested, is the N6th deoxyhemoglobin signal in the jth channel of the ith person to be tested;

[0453] Module M5.2: Data denoising:

[0454] The band-pass filter is used to remove the high-frequency noise and low-frequency drift in the signal, and the de-noised signal is obtained, that is

[0455] HbO_filtered(i) j = BF(HbO(i) j , f low , f high )

[0456] HbR_filtered(i) j = BF(HbR(i) j , f low , f high )

[0457] Wherein, HbO_filtered(i) j is the filtered oxygenated hemoglobin signal of the ith person to be tested in the jth channel, HbR_filtered(i) j is the filtered deoxygenated hemoglobin signal of the ith person to be tested in the jth channel, BF is a band-pass filter, f low , f high is the cutoff frequency;

[0458] Module M5.3: brain function feature extraction

[0459] Extract the average oxygenated hemoglobin concentration change of all channels of the ith person to be tested:

[0460]

[0461] Wherein, N5 is the number of data recorded by the channel;

[0462] Extract the average deoxygenated hemoglobin concentration change of all channels of the ith person to be tested:

[0463]

[0464] Extract the total blood volume change s5(i) of the ith person to be tested:

[0465] s5 = ΔHbO(i) + ΔHbR(i)

[0466] Wherein, mean(HbO_filtered(i) j ) is the average value of different channels of HbO_filtered(i) j signal, mean(HbR_filtered(i) j ) is the average value of different channels of HbR_filtered(i) j signal;

[0467] Module M6: assessing the brain structure of the person under test;

[0468] In said module M6:

[0469] Module M6.1: data acquisition:

[0470] Performing magnetic resonance imaging (MRI) on N persons under test, setting the scanning parameters;

[0471] Module M6.2: structural features extraction:

[0472] Segmenting the brain tissue into gray matter, white matter and cerebrospinal fluid using the automated tool FreeSurfer; using the FreeSurfer toolkit, calculating the gray matter volume (GMV) of the prefrontal cortex of the ith person under test i , calculating the mean cortical thickness (CT) of the prefrontal region i , calculating the cortical surface area (CSA) of the prefrontal region i ;

[0473] Module M6.3: calculation of the brain structural indices:

[0474] s6 = GMV i + CT i + CSA i

[0475] Module M7: information fusion based on the subjective questionnaire, walking speed and gait, fine motor skills of the fingers, nerve conduction velocity, brain function and brain structure.

[0476] In said module M7:

[0477] Module M7.1: collecting the assessment indices of N persons under test to form a data set P and the CIPN severity level Y of the persons under test:

[0478] P = {p1,..., p l …, p N},

[0479] Y = {y1,..., y l …, y N},

[0480] wherein p l = {s l1 , s l2 , s l3 , s l4 , s l5 , s l6} T are the assessment indices of the ith person under test, and Y is the CIPN severity level data matrix of N persons under test, with a size of 5 x N, y lThis is the severity level of CIPN for the lth person being tested;

[0481] Module M7.2: Data standardization, the specific construction method is as follows:

[0482] Module M7.2.1: Divide the data of the above N test subjects into training set, validation set and test set according to a preset ratio;

[0483] Module M7.2.2: Normalizes all the above data using the following formula:

[0484]

[0485] Among them, s ij Let sn be the j-th feature of the i-th test subject, i∈{1,2,3,...,N}, j∈{1,2,3,4}. ij s is the j-th normalized feature of the i-th test subject; minj Let s be the minimum value of the j-th feature. maxj The maximum value of the j-th feature;

[0486] Module M7.3: Building multi-class classification models using Mamba models:

[0487] Module M7.3.1: For each test subject i in the dataset P, the feature vector p i ={Q i1 ,sn t2 ,sn i3 ,sn i4 ,sn i5 ,sn i6} T Embedding is performed to convert it into a high-dimensional feature representation h. i ;

[0488] - Use a linear transformation to map the original normalized features to a high-dimensional embedding space:

[0489] h i =W e ·p i +b e

[0490] Among them, h i It is the embedded feature vector with a size of d×1; W e This is the weight matrix of the embedding layer, with a size of d×6; b e is the bias vector of the embedding layer, with a size of d×1; d is the size of the embedding dimension.

[0491] Multi-head attention mechanism is applied to capture each normalized feature s ij The relationships between them enhance feature representation;

[0492] The embedded feature vector h i Compute query Q, key K and value V:

[0493] Q = W q · h i K = W k · h i V = W v · h i

[0494] where Q is the query vector with size d x 1; K is the key vector with size d x 1; V is the value vector with size d x 1; W q , W k and W v are the weight matrices of query, key and value respectively, all with size d x d;

[0495] Compute self-attention score Attention(Q, K, V)

[0496]

[0497] where d1 is the dimension of the key vector, used to scale the dot product result;

[0498] For multi-head attention mechanism, the above operation is extended to multi-head, and multiple attention scores are computed in parallel:

[0499] MultiHead(Q, K, V) = Concat(head1, …, head h )· W o

[0500] where head i represents the calculation result of the i-th attention head; W o is the weight matrix of the output, with size d x d;

[0501] Module M7.3.2: Neural Atom Bag (NBA) feature extraction

[0502] After the multi-head attention mechanism, NBA is introduced to extract local and global features, and the feature vector is processed by a set of trainable atom mapping layers a m to perform nonlinear feature mapping:

[0503]

[0504] where a m is the atomic-level feature representation with size d2 x 1; W a and b aWm and Bm are the weight matrix and bias vector of the m-th atom mapping layer, respectively; σ is the ReLU nonlinear activation function; d2 is the dimension size of the atomic-level feature;

[0505] All atomic-level features are integrated into global feature representation through pooling operation:

[0506]

[0507] wherein, is the global feature representation extracted by the neural atom bag, with a size of D2×1; k is the number of atom mapping layers; MaxPooling is the maximum pooling layer;

[0508] Module M7.3.3: The feature vector extracted through the multi-head attention mechanism and the neural atom bag: The prediction of CIPN severity level is made through the classification head;

[0509] The classification head is usually one or more fully connected layers, and the Softmax function is used to output the category probability:

[0510]

[0511] wherein, W c is the weight matrix of the classification head, with a size of C×d2; B c is the bias vector of the classification head, with a size of C×1; is the CIPN severity level prediction result of the i-th person to be tested, with a size of C×1; C is the number of CIPN severity levels;

[0512] Module M7.3.4: Model training and updating:

[0513] The model is trained using the cross-entropy loss function, and the model parameters are updated through the backpropagation algorithm, and the formula of the cross-entropy loss function is:

[0514]

[0515] wherein, y ic is the true label of the i-th person to be tested in category c; is the predicted probability of the i-th person to be tested in category c; C is the number of CIPN severity levels;

[0516] Module M7.3.5: Repeat training until convergence

[0517] Repeat modules M7.3.1 to M7.3.4 to gradually optimize the model parameters until the performance improvement efficiency on the validation set is lower than the preset standard, and the trained MANBA model is used to predict the CIPN severity level of new persons to be tested.

[0518] Embodiment 3:

[0519] Embodiment 3 is a preferred example of Embodiment 1, to more specifically illustrate the present application.

[0520] In view of the defects in the prior art, the purpose of the present application is to provide a CIPN information fusion system based on multi-modal information.

[0521] According to the present application, a precise grading system based on medical multi-modal information is provided, comprising:

[0522] Step S1: evaluating the subjective questionnaire of the person to be tested.

[0523] Step S2: evaluating the walking speed and gait according to the walking video of the person to be tested.

[0524] Step S3: evaluating the fine motor skills of the fingers of the person to be tested.

[0525] Step S4: evaluating the nerve conduction velocity of the person to be tested.

[0526] Step S5: evaluating the brain function of the person to be tested.

[0527] Step S6: evaluating the brain structure of the person to be tested.

[0528] Step S7: constructing a precise grading system using the evaluation values of S1-S6 above.

[0529] Step S8: using the precise grading system constructed in S7 to precisely grade the severity of CIPN of the patient.

[0530] Preferably, in the step S1:

[0531] S1.1 Patient Self-Report Scale Design

[0532] This scale contains surveys on multiple aspects of the patient, specifically designed as follows:

[0533] Part I: Sensory Symptoms

[0534] Please choose the option that best fits your feelings according to the following questions:

[0535] · 0 points: None

[0536] · 1 point: Mild

[0537] · 2 points: Moderate

[0538] · 3 points: Severe

[0539] · 4 points: Very severe

[0540] Do you feel a tingling sensation in your fingers or toes?

[0541] Do you feel a numbness in your fingers or toes?

[0542] Do you feel a burning sensation in your palms or soles?

[0543] Do you feel pain without any provocation?

[0544] Section Two: Motor Symptoms

[0545] Please choose the option that best describes your experience for the following questions:

[0546] • 0: None

[0547] • 1: Mild

[0548] • 2: Moderate

[0549] • 3: Severe

[0550] • 4: Very severe

[0551] Do you feel weakness in your hands or feet?

[0552] Do you find it difficult to lift objects or grip things?

[0553] Do you feel unsteady when walking?

[0554] Section Three: Autonomic Symptoms

[0555] Please choose the option that best describes your experience for the following questions:

[0556] • 0: None

[0557] • 1: Mild

[0558] • 2: Moderate

[0559] • 3: Severe

[0560] • 4: Very severe

[0561] Do you feel lightheaded or dizzy when standing?

[0562] Do you have digestive problems such as constipation or diarrhea?

[0563] Section Four: Impact on Quality of Life

[0564] Please choose the option that best describes your experience for the following questions:

[0565] • 0: Not at all

[0566] • 1 minute: a little

[0567] • 2 minutes: some

[0568] • 3 minutes: quite a lot

[0569] • 4 minutes: very much

[0570] Does your neuropathy symptom affect your daily activities (such as dressing, bathing)?

[0571] Does your neuropathy symptom affect your social activities (such as meeting with friends and family)?

[0572] Does your neuropathy symptom affect your work or study?

[0573] Does your neuropathy symptom affect your sleep?

[0574] S1.2 Data collection and preprocessing

[0575] Collect all the patient self-report scale questionnaire data completed by patients. Standardize all data;

[0576]

[0577] where x i is the ith data in the patient self-report scale questionnaire data, a is the minimum value of the questionnaire data, and b is the maximum value of the questionnaire data.

[0578] S1.3 Setting of index weight

[0579] According to the opinions of clinical experts, determine the importance weight of each patient self-report scale questionnaire question.

[0580]

[0581] where (w i ) j is the importance weight set by the jth clinical expert for the ith index, the range is 1-10, N is the number of experts, and w i is the importance weight of the ith index.

[0582] S1.3 Comprehensive score calculation

[0583] The weighted average method is used to calculate the comprehensive CIPN score of each patient. The formula is as follows:

[0584]

[0585] where w i is the importance weight of the ith index, and x iFor the i-th data in the CIPN20 questionnaire data, s1 is the score of the person being tested on this scale.

[0586] Preferably, in the step S2:

[0587] S2.1 Video preprocessing

[0588] Extract K frames of images from the video of the person being tested

[0589] {T1,...,T K}

[0590] S2.2 Gait feature extraction

[0591] First, background subtraction is performed using the correlation between consecutive frames to remove the static background and retain the moving human body part, specifically:

[0592] Step S2.2.1: Read the current frame image T t and the previous frame image T t-1 , and calculate the difference image between the two frames

[0593] Diff t = |T t -T t-1 |

[0594] Step S2.2.2: Binary the difference image Diff t using a hard threshold, and the pixels greater than the threshold I are the foreground:

[0595]

[0596] where F t is the t-th frame of foreground image after removing the background.

[0597] Then, use the emplaceAndPop function in the trained pose estimation model OpenPose to detect the toe position of each frame in the video from the above foreground image:

[0598] {(A1,B1),...,(A K ,B K )}

[0599] where (A K ,B K ) are the coordinates of the left and right toe positions in the K-th frame, respectively.

[0600] S2.3 Gait cycle recognition

[0601] Calculate the distance between the toe positions in two consecutive frames,

[0602]

[0603] wherein D m is the distance between the toe of the mth frame and the m+1th frame

[0604] Calculate the walking speed of the subject in the video:

[0605]

[0606] wherein s2 is the walking speed of the subject, T is the time in seconds between the 1st frame and the Kth frame, and V is the walking speed.

[0607] Preferably, in said step S3:

[0608] The user's reaction time is recorded by randomly appearing click points on the mobile phone screen, and gradually shortening the time interval between points, so as to evaluate the user's fine motor skills, the specific steps are as follows.

[0609] S3.1 Define the screen as an area with a width of C and a height of D.

[0610] S3.2 At each time interval t n , randomly select a screen position to appear a click point.

[0611] S3.3 Reaction time calculation

[0612] (Δt) n =t click -t start

[0613] t start is the time of the click point, t click is the time of the click point, and Δt is the reaction speed of the nth time

[0614] S3.4 After each click, shorten the time interval of the point.

[0615] t n+1 =α×t n

[0616] α is a shortening coefficient, which is greater than 0 and less than 1

[0617] S3.5 Calculate the average reaction time of the user. Record a sequence of multiple reaction times {(Δt)1,...,(Δt) n},

[0618]

[0619] wherein s3 is the reaction time of the subject.

[0620] Preferably, in said step S4:

[0621] The nerve conduction velocity is calculated as follows:

[0622]

[0623] where s4 is the nerve conduction velocity of the person to be tested, D is the distance between the stimulation points, AL = L1-L2 is the time difference between the detection of the electrical signal at the distal and proximal ends, L2 is the time point at which the electrical signal is detected at the distal end, and L1 is the time point at which the electrical signal is detected at the proximal end.

[0624] Preferably, in the step S5:

[0625] S5.1 Data acquisition

[0626] The prefrontal lobe function of N persons to be tested is acquired using an N5-channel near-infrared spectroscopy (NIRS) device. Each channel records the change in the concentration of oxygenated hemoglobin (HbO) and deoxygenated hemoglobin (HbR). The number of data recorded by each channel of the person to be tested is N6, and the signal in the jth channel of the ith person to be tested is:

[0627]

[0628] S5.2 Data denoising

[0629] The high-frequency noise and low-frequency drift in the above signals are removed using a band-pass filter to obtain the denoised signal, i.e.

[0630] HbO_filtered(i) j =BF(HbO(i) j ,f low ,f high )

[0631] HbR_filtered(i) j =BF(HbR(i) j ,f low ,f high )

[0632] where BF is a band-pass filter, f low ,f high are the cutoff frequencies.

[0633] S5.3 Brain function feature extraction

[0634] The average oxygenated hemoglobin concentration change of all channels of the ith person to be tested is extracted as follows:

[0635]

[0636] The average deoxygenated hemoglobin concentration change of all channels of the ith person to be tested is extracted as follows:

[0637]

[0638] extracting the total blood volume change s5(i) of the i-th person to be tested:

[0639] s5 = AHb0(i) + AHbR(i)

[0640] wherein mean(Hb0_filtered(i) j ) is the mean value of the Hb0_filtered(i) j signal over the different channels and mean(HbR_filtered(i) j ) is the mean value of the HbR_filtered(i) j signal over the different channels.

[0641] Preferably, in said step S6:

[0642] S6.1 data collection

[0643] The N persons to be tested are subjected to a magnetic resonance imaging (MRI) with the following parameters: scan mode: T1-weighted image; scan plane: axial plane; slice thickness: 1 mm; matrix size: 256 x 256; field of view (FOV): 240 mm.

[0644] S6.2 structural features extraction

[0645] The brain tissue is segmented into grey matter, white matter and cerebrospinal fluid using an automated tool (FreeSurfer); using the FreeSurfer toolkit, the grey matter volume GMV i of the prefrontal cortex of the i-th person to be tested is calculated, the mean cortical thickness CT i of the prefrontal cortex region is calculated, the cortical surface area CSA i of the prefrontal cortex region is calculated.

[0646] S6.3 calculation of the cerebral structural indexes

[0647] s6 = GMV i + CT i + CSA i

[0648] Preferably, in said step S7:

[0649] S7.1 collecting the set of evaluation indexes S1-S6 described above for the N persons to be tested, forming a data set P and the CIPN severity level Y of the person to be tested:

[0650] P = {p1,..., p l …, p N},

[0651] Y = {y1,...,y l ...,y N},

[0652] where p l = {s l1 , s l2 , s l3 , s l4 , s l5 , s l6} T is the evaluation index of the lth person to be tested, Y is the CIPN severity level data matrix of N persons to be tested, the size is 5xN, y l is the CIPN severity level of the lth person to be tested.

[0653] S7.2 data standardization, the specific construction method is as follows:

[0654] Step S7.2.1: Divide the data of the above N persons to be tested into training set, validation set and test set according to a certain proportion;

[0655] Step S7.2.2: Normalize all the above data, the formula is as follows:

[0656]

[0657] where s Ij is the jth feature of the ith person to be tested, i∈{1,2,3,...,N}, j∈{1,2,3,4}, sn ij is the jth normalized feature of the ith person to be tested. s minj is the minimum value of the jth feature, s maxj is the maximum value of the jth feature.

[0658] S7.3 use Mamba model to construct multi-classification model

[0659] S7.3.1: Perform embedding processing on the feature vector p i = {sn i1 , sn i2 , sn i3 , sn i4 , sn i5 , sn i6} T of each person to be tested in the data set P, and convert it to a high-dimensional feature representation h i .

[0660] - Use linear transformation to map the original normalized features to a high-dimensional embedding space:

[0661] h i =W e ·p i +b e

[0662] where h i is the embedded feature vector with size d x 1; W e is the weight matrix of the embedding layer with size d x 6; b e is the bias vector of the embedding layer with size d x 1; and d is the size of the embedding dimension.

[0663] The multi-head attention mechanism is applied to capture the relationship between each normalized feature s ij and enhance the feature representation.

[0664] The embedded feature vector h i is used to calculate the query Q, the key K, and the value V:

[0665] Q = W q · h i , K = W k · h i , and V = W v · h i

[0666] where Q is the query vector with size d x 1; K is the key vector with size d x 1; V is the value vector with size d x 1; W q , W k , and W v are the weight matrices of the query, the key, and the value, respectively, all with size d x d.

[0667] The self-attention score Attention(Q, K, V) is calculated as follows:

[0668]

[0669] where d1 is the dimension of the key vector used to scale the dot product result.

[0670] For the multi-head attention mechanism, the above operation is extended to multiple heads, and multiple attention scores are calculated in parallel:

[0671] MultiHead(Q, K, V) = Concat(head1, …, head h ) · W o

[0672] where head i represents the calculation result of the i-th attention head; and W o is the weight matrix of the output with size d x d.

[0673] S7.3.2 Neural Bag of Atoms (NBA) feature extraction

[0674] After the multi-head attention mechanism, NBA is introduced to extract local and global features to enhance the model's ability to capture complex patterns.

[0675] Feature vector after multi-head attention processing Through a set of trainable "atom" mapping layers a m Nonlinear feature mapping is performed:

[0676]

[0677] where a m is the atomic-level feature representation with size d2x1; M a and b a are the weight matrix and bias vector of the mth atomic mapping layer, respectively; σ is the ReLU nonlinear activation function; d2 is the dimension size of the atomic-level feature.

[0678] All atomic-level features are integrated into global feature representation through pooling operation:

[0679]

[0680] where, is the global feature representation extracted by neural bag of atoms with size d2x1; k is the number of atomic mapping layers.

[0681] S7.3.3 Feature vector after multi-head attention mechanism and neural bag of atoms extraction The prediction of CIPN severity level is performed through the classification head.

[0682] The classification head is usually one or more fully connected layers, and uses the Softmax function to output the category probability:

[0683]

[0684] where W c is the weight matrix of the classification head with size Cxd2; b c is the bias vector of the classification head with size Cx1; is the CIPN severity level prediction result of the ith person under test with size Cx1; C is the number of CIPN severity levels.

[0685] S7.3.4 Model training and updating

[0686] The model is trained using the cross-entropy loss function, and the model parameters are updated through the backpropagation algorithm. The cross-entropy loss function formula is:

[0687]

[0688] where y ic is the true label of the ith person to be tested on category c; is the predicted probability of the ith person to be tested on category c.

[0689] S7.3.5 Repeat training until convergence

[0690] Steps S7.3.1 to S7.3.4 are repeated to optimize the model parameters step by step until the performance on the validation set no longer improves significantly. Finally, the trained MANBA model can be used to predict the CIPN severity level of new persons to be tested.

[0691] Those skilled in the art know that, in addition to implementing the system provided by the present application and each device, module, unit thereof in the form of pure computer readable program code, the system provided by the present application and each device, module, unit thereof can also be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers and embedded microcontrollers, etc. to achieve the same functions by logically programming the method steps. Therefore, the system provided by the present application and each device, module, unit thereof can be considered as a hardware component, and the devices, modules, units included therein for implementing various functions can also be considered as structures within the hardware component; the devices, modules, units for implementing various functions can also be considered as both software modules implementing methods and structures within hardware components.

[0692] The specific embodiments of the present application are described above. It needs to be understood that the present application is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essential content of the present application. The embodiments of the present application and the features in the embodiments can be arbitrarily combined with each other without conflict.

Claims

1. A CIPN information fusion method based on multi-modal information, characterized in that, Comprise: Step S1: evaluating the subjective questionnaire of the testee; Step S2: evaluating the walking speed and gait according to the walking video of the testee; Step S3: evaluating the fine motor skills of the testee; Step S4: evaluating the nerve conduction velocity of the testee; Step S5: evaluating the brain function of the testee; Step S6: evaluating the brain structure of the testee; Step S7: information fusion according to the subjective questionnaire, walking speed and gait, fine motor skills, nerve conduction velocity, brain function and brain structure; In the step S7: Step S7.1: collecting the evaluation indexes of N testees to form a data set P and the CIPN severity level Y of the testee: P = {p1,..., p l …,p N}, Y = {y1,...,y l ...,y N}, wherein p l = {s l1 ,s l2 ,s l3 ,s l4 ,s l5 ,s l6} T is the evaluation index of the lth person to be tested, Y is the CIPN severity level data matrix of N persons to be tested, the size of which is 5xN, y l is the CIPN severity level of the lth person to be tested. Step S7.2: data standardization, the specific construction method is as follows: Step S7.2.1: dividing the data of the above N testees into training set, validation set and test set according to the preset proportion; Step S7.2.2: normalizing all the data, the formula is as follows: wherein s ij is the jth feature of the ith person to be tested, i ∈ {1, 2, 3,..., N}, j ∈ {1, 2, 3, 4}, sn ij is the jth normalized feature of the ith person to be tested; s minj is the minimum value of the jth feature, s maxj is the maximum value of the jth feature; Step S7.3: using Mamba model to construct multi-classification model: Step S7.3.1: Perform embedding processing on the feature vector p of each person i in the data set P i = {sn i1 ,sn i2 ,sn i3 ,sn i4 ,sn i5 ,sn i6} T to convert it into a high-dimensional feature representation h i ; Using linear transformation to map the original normalized features to high-dimensional embedding space: h i = W e · p i + b e where h i is the embedded feature vector with size d x 1; W e is the weight matrix of the embedding layer with size d x 6; b e is the bias vector of the embedding layer with size d x 1; and d is the size of the embedding dimension. application of multi-head attention mechanism to capture the relationship between each normalized feature s ij enhance the feature representation; The feature vector h after embedding i Compute query Q, key K, and value V: Q = W q h i K = W k h i V = W v h i where Q is a query vector with size d x 1, K is a key vector with size d x 1, V is a value vector with size d x 1, W q , W k and W v are weight matrices for query, key and value respectively, all with size d x d. where Q is a query vector with size d x 1, K is a key vector with size d x 1, V is a value vector with size d x 1, W q , W k and W v are weight matrices for query, key and value respectively, all with size d x d. Calculate self-attention score Attention(Q, K, V) Where: d1 is the dimension of the key vector, used to scale the dot product result; For multi-head attention mechanism, the above operation is extended to multi-head, and multiple attention scores are calculated in parallel: MultiHead(Q, K, V) = Concat(head1,..., head h ) · W o wherein head i represents the calculation result of the i-th attention head; W o is the weight matrix of the output, with a size of dxd; Step S7.3.2: neural atomic bag: NBA feature extraction After the multi-head attention mechanism, NBA is introduced to extract local and global features, and the feature vector processed by the multi-head attention through a set of trainable atomic mapping layers a m non-linear feature mapping: where a m is an atomic-level feature representation with size d2x1; W a and b a are the weight matrix and bias vector of the m-th atomic mapping layer, respectively; σ is a ReLU nonlinear activation function; and d2 is the dimension size of the atomic-level feature. Integrate all atomic-level features into global feature representation through pooling operation: wherein, is a global feature representation extracted by the neural atom bag with size d2x1; k is the number of atom mapping layers; MaxPooling is a max pooling layer; Step S7.3.3: Feature vector after multi-head attention mechanism and neural atom bag extraction: Prediction of CIPN severity grade by classification head; The classification head is usually one or more fully connected layers, and uses the Softmax function to output the category probability: wherein W c is a weight matrix of the classification head, with size Cxd2; b c is a bias vector of the classification head, with size Cx1; is the CIPN severity grade prediction result of the ith person to be tested, with size Cx1; C is the number of CIPN severity grades; Step S7.3.4: model training and updating: Using cross-entropy loss function to train the model, updating the model parameters through back propagation algorithm, the formula of cross-entropy loss function is: wherein y ic is the true label of the ith person under test on category c; is the predicted probability of the ith person under test on category c; C is the number of CIPN severity levels. Step S7.3.5: repeat training until convergence Repeat steps S7.3.1 to S7.3.4 to gradually optimize the model parameters until the performance improvement efficiency on the validation set is lower than the preset standard.

2. The CIPN information fusion method based on multi-modal information according to claim 1, characterized in that, In the step S1: Step S1.1: design self-report scale, including investigation of sensory symptoms, motor symptoms, autonomic nervous symptoms and influence on quality of life; Step S1.2: data collection and preprocessing: Collect self-report questionnaire data and regularize all data; where x i is the ith data in the self-reported scale questionnaire data of the person to be tested, a is the minimum value of the questionnaire data, and b is the maximum value of the questionnaire data. Step S1.3: index weight setting: According to the opinions of clinical experts, determine the importance weight of each testee's self-report questionnaire problem: wherein (w i ) j is the importance weight set by the jth clinical expert for the ith index, ranging from 1 to 10, N is the number of experts, w i is the importance weight of the ith index; Step S1.3: comprehensive score calculation: Calculate the comprehensive CIPN score by weighted average method, the formula is as follows: wherein w i is the importance weight of the i-th indicator, x i is the i-th data in the CIPN20 questionnaire data, s1 is the score of the to-be-tested person on this scale; M is the total number of indicators.

3. The CIPN information fusion method based on multi-modal information according to claim 1, characterized in that, In the step S2: Step S2.1: video preprocessing extract K frames of images {T1,...,T K} from a video taken of the person under test; Step S2.2: gait feature extraction Using the correlation between consecutive frames for background subtraction, removing the static background and retaining the moving human body part, specifically: Step S2.2.1 : reading the current frame image T t with the previous frame image T t-1 and computing the difference image Diff t : Diff t = |T t -T t-1 | Step S2.2.2: Apply hard thresholding to the difference image Diff t binarize, pixels greater than threshold I are foreground: Ft= Ft- Fb t is the t-th frame foreground image after background removal; The emplaceAndPop function in the pre-trained pose estimation model OpenPose is used to detect the toe position {(A1,B1),..., ... K B K )};wherein, (A K B K These are the coordinates of the left and right toes in the Kth frame, respectively. Step S2.3: gait cycle recognition: Calculate the distance between the toe positions in two consecutive frames: where D m is the distance between the toe of the mth frame and the m+1th frame; Calculate the walking speed of the testee in the video: Where s2 is the walking speed of the testee, T is the number of seconds between the first frame and the Kth frame, and V is the walking speed.

4. The CIPN information fusion method based on multi-modal information according to claim 1, characterized in that, In the step S3: The reaction time of the user is recorded by randomly appearing click points on the mobile phone screen and gradually shortening the time interval between the points, thereby evaluating the fine motor skills of the user, and the specific steps are as follows: Step S3.1: define the screen as an area with a width of C and a height of D; Step S3.2: At each time interval t n At the end, a click point appears at a randomly chosen screen position; Step S3.3: reaction time calculation: (Δt) n = t click - t start where t start is the time of appearance of the click point, t click is the time of click of the click point, (Δt) n is the reaction speed of the nth time; Step S3.4: after each click, shorten the time interval between the points: t n+1 = α x t n wherein a is a shortening factor, greater than 0 and less than 1 ; t n is the nth time interval; Step S3.5: Calculate the average reaction time of the user, record the sequence of multiple reaction times Where s3 is the reaction time of the person to be tested, and n1 is the total number of time intervals.

5. The CIPN information fusion method based on multi-modal information according to claim 1, characterized in that, In the step S4: The calculation method of nerve conduction velocity is: Where s4 is the nerve conduction velocity of the person to be tested, D is the distance between the stimulation points, ΔL=L1-L2 is the time of detecting the electrical signal between the distal and proximal ends, L2 is the time point of detecting the electrical signal at the distal end, and L1 is the time point of detecting the electrical signal at the proximal end.

6. The CIPN information fusion method based on multi-modal information according to claim 1, characterized in that, In the step S5: Step S5.1: data collection: Use the N5 channel NIRS device to collect NIRS data of the frontal lobe function of N persons to be tested; record the changes of oxygenated hemoglobin HbO and deoxygenated hemoglobin HbR concentration for each channel, and the number of data recorded by each channel of the person to be tested is N6, and the signal in the jth channel of the ith person to be tested is: Wherein, HbO(i) j This represents the oxyhemoglobin signal in the j-th channel of the i-th person being tested. HbR(i) represents the N6th oxyhemoglobin signal in the j-th channel of the i-th test subject. j This represents the deoxyhemoglobin signal in the j-th channel of the i-th person being tested. This represents the N6th deoxyhemoglobin signal in the j-th channel of the i-th test subject; Step S5.2: data denoising: Use a band-pass filter to remove high-frequency noise and low-frequency drift in the signal to obtain the denoised signal, that is HbO_filtered(i) j = BF(HbO(i) j ,f low ,f high ) HbR_filtered(i) j = BF(HbR(i) j ,f low ,f high ) HbO_filtered(i) = BF (HbO(i), f j j HbR_filtered(i) = BF (HbR(i), f j low f high cut-off frequency Step S5.3: brain function feature extraction Extract the average oxygenated hemoglobin concentration change of all channels of the ith person to be tested: Where N5 is the number of data recorded by the channel; Extract the average deoxygenated hemoglobin concentration change of all channels of the ith person to be tested: Extract the total blood volume change s5(i) of the ith person to be tested: s5=ΔHbO(i)+ΔHbR(i) where mean(HbO_filtered(i) is the average of the HbO_filtered(i) signal across different channels, mean(HbR_filtered(i) is the average of the HbR_filtered(i) signal across different channels. j j j j ​​​​ 7. The CIPN information fusion method based on multi-modal information according to claim 1, characterized in that, In the step S6: Step S6.1: data collection: Perform magnetic resonance imaging (MRI) on N persons to be tested and set the scanning parameters; Step S6.2: structural feature extraction: segmenting the brain tissue into grey matter, white matter and cerebrospinal fluid using the automated tool FreeSurfer; calculating the grey matter volume GMV of the prefrontal cortex of the ith person under test using the FreeSurfer toolkit i ; calculating the average cortical thickness CT of the prefrontal region i ; calculating the cortical surface area CSA of the prefrontal region i ; Step S6.3: calculation of brain structural indicators: s6 = GMV i + CT i + CSA i .

8. A CIPN information fusion system based on multi-modal information, characterized in that, Including: Module M1: evaluate the subjective questionnaire of the person to be tested; Module M2: evaluate the walking speed and gait according to the walking video of the person to be tested; Module M3: evaluate the fine motor skills of the fingers of the person to be tested; Module M4: evaluate the nerve conduction velocity of the person to be tested; Module M5: evaluate the brain function of the person to be tested; Module M6: evaluate the brain structure of the person to be tested; Module M7: information fusion according to the subjective questionnaire, walking speed and gait, fine motor skills of fingers, nerve conduction velocity, brain function and brain structure; In the module M7: Module M7.1: collect the evaluation indicators of N persons to be tested to form a data set P and the CIPN severity level Y of the person to be tested: P = {p1,..., p l ..., p N}, Y = {y1,...,y l ...,y N}, wherein p l = {s l1 ,s l2 ,s l3 ,s l4 ,s l5 ,s l6} T is the evaluation index of the lth person to be tested, Y is the CIPN severity grade data matrix of N persons to be tested, the size of which is 5xN, y l is the CIPN severity grade of the lth person to be tested. Module M7.2: data standardization, and the specific construction method is as follows: Module M7.2.1: divide the data of the above N persons to be tested into a training set, a validation set and a test set according to a predetermined ratio; Module M7.2.2: normalize all the data, and the formula is as follows: wherein s ij is the jth feature of the ith person to be tested, i ∈ {1, 2, 3,..., N}, j ∈ {1, 2, 3, 4}, sn ij is the jth normalized feature of the ith person to be tested; s minj is the minimum value of the jth feature, s maxj is the maximum value of the jth feature; Module M7.3: use the Mamba model to construct a multi-classification model: Module M7.3.1 : Feature vector p for each person i in the data set P under test i = {sn i1 ,sn i2 ,sn i3 ,sn i4 ,sn i5 ,sn i6} T Perform embedding processing to convert it into a high-dimensional feature representation h i ; - Use linear transformation to map original normalized features to high-dimensional embedding space: h i = W e · p i + b e where h i is the embedded feature vector with size d x 1; W e is the weight matrix of the embedding layer with size d x 6; b e is the bias vector of the embedding layer with size d x 1; and d is the size of the embedding dimension. application of multi-head attention mechanism to capture the relationship between each normalized feature s ij enhance the feature representation; The feature vector h after embedding i Compute query Q, key K, and value V: Q = W q h i K = W k h i V = W v h i where Q is a query vector with size d x 1, K is a key vector with size d x 1, V is a value vector with size d x 1, and W q , W k , and W v are weight matrices for query, key, and value respectively, all with size d x d. Compute self-attention scores Attention(Q, K, V) Where: d1 is the dimension of the key vector, used to scale the dot product result; For multi-head attention mechanism, extend the above operation to multiple heads, and compute multiple attention scores in parallel: MultiHead(Q, K, V) = Concat(head1,..., head h ) · W o wherein head i represents the calculation result of the i-th attention head; W o is the weight matrix of the output, with a size of dxd; Module M7.3.2: Neural Atom Bag: NBA feature extraction After the multi-head attention mechanism, NBA is introduced to extract local and global features, and the feature vector processed by the multi-head attention through a set of trainable atomic mapping layers a m nonlinear feature mapping: where a m is an atomic-level feature representation with size d2x1; W a and b a are the weight matrix and bias vector of the m-th atomic mapping layer, respectively; σ is a ReLU nonlinear activation function; and d2 is the dimension size of the atomic-level feature. Integrate all atomic-level features into global feature representation through pooling operation: wherein, is a global feature representation extracted by the neural atom bag with size d2x1; k is the number of atom mapping layers; MaxPooling is a max pooling layer; Module M7.3.3: Feature vector after passing through multi-head attention mechanism and neural atom bag extraction: Prediction of CIPN severity grade by classification head; The classification head is usually one or more fully connected layers, and uses the Softmax function to output class probabilities: wherein W c is a weight matrix of the classification head, with size Cxd2; b c is a bias vector of the classification head, with size Cx1; is the CIPN severity grade prediction result of the ith person to be tested, with size Cx1; C is the number of CIPN severity grades; Module M7.3.4: Model training and updating: Train the model using the cross-entropy loss function, and update the model parameters through the backpropagation algorithm. The cross-entropy loss function formula is: wherein y ic is the true label of the ith person under test on category c; is the predicted probability of the ith person under test on category c; C is the number of CIPN severity levels. Module M7.3.5: Repeat training until convergence Repeat modules M7.3.1 to M7.3.4 to gradually optimize model parameters until the performance on the validation set improves less than the preset standard.

9. The CIPN information fusion system based on multi-modal information according to claim 8, characterized in that: In the module M1: Module M1.1: Design a self-reporting scale, including surveys of sensory symptoms, motor symptoms, autonomic nervous symptoms, and quality of life impact; Module M1.2: Data collection and preprocessing: Collect self-reporting scale questionnaire data and normalize all data; where x i is the ith data in the self-reported scale questionnaire data of the person to be tested, a is the minimum value of the questionnaire data, and b is the maximum value of the questionnaire data. Module M1.3: Index weight setting: Determine the importance weight of each self-reporting scale questionnaire question for each person being tested according to the opinions of clinical experts: wherein (w i ) j is the importance weight set by the jth clinical expert for the ith index, ranging from 1 to 10, N is the number of experts, w i is the importance weight of the ith index; Module M1.3: Comprehensive score calculation: Calculate the comprehensive CIPN score using the weighted average method, with the formula as follows: wherein w i is the importance weight of the i-th indicator, x i is the i-th data in the CIPN20 questionnaire data, s1 is the score of the to-be-tested person on this scale; M is the total number of indicators; In the module M2: Module M2.1: Video preprocessing extract K frames of images {T1,..., Tk} from a video taken of the person under test K} Module M2.2: Gait feature extraction Use the correlation between consecutive frames for background subtraction to remove static backgrounds and retain moving human body parts, specifically: Module M2.2.1 : reading the current frame image T t with the previous frame image T t-1 and computing the difference image Diff t : Diff t = |T t -T t-1 | Module M2.2.2: using hard thresholding to difference image Diff t binarized, pixels greater than threshold I are foreground: Ft= Ft- Fb t is the t-th frame foreground image after background removal; The emplaceAndPop function in the trained pose estimation model OpenPose is used to detect the toe position {(A1, B1),..., (A K ,B K )} of each frame in the foreground image of the video; wherein (A K ,B K ) are the coordinates of the left toe and right toe positions in the Kth frame, respectively. Module M2.3: Gait cycle recognition: Calculate the distance between the toe positions in two consecutive frames: where D m is the distance between the toe of the mth frame and the m+1th frame; Calculate the walking speed of the subject in the video: Where s2 is the walking speed of the person being tested, T is the number of seconds between the 1st frame and the Kth frame, and V is the walking speed; In the module M3: Evaluate the user's fine motor skills by randomly appearing click points on the phone screen and gradually shortening the time interval between points to record the user's reaction time, with the following specific steps: Module M3.1: Define the screen as an area with a width of C and a height of D; Module M3.2: At each time interval t n At the end, a click point appears at a randomly selected screen location. Module M3.3: Reaction time calculation: (Δt) n = t click -t start where t start is the time of appearance of the click point, t click is the time of click of the click point, (Δt) n is the reaction speed of the nth time; Module M3.4: After each click, shorten the time interval between points: t n+1 = a x t n wherein a is a shortening factor, greater than 0 and less than 1 ; t n is the nth time interval; Module M3.5: Calculate the average reaction time of the user, record a sequence of multiple reaction times Where s3 is the reaction time of the person being tested, and n1 is the total number of time intervals; In the module M4: The calculation method of nerve conduction velocity is: Where s4 is the nerve conduction velocity of the person being tested, D is the distance between the stimulation points, ΔL = L1 - L2 is the time difference between the detection of electrical signals at the distal and proximal ends, L2 is the time point at which the distal end detects the electrical signal, and L1 is the time point at which the proximal end detects the electrical signal; In the module M5: Module M5.1: Data acquisition: Using N5 channel NIRS equipment, the prefrontal function of N persons to be tested is collected; each channel records the change of oxyhemoglobin HbO and deoxyhemoglobin HbR concentration, the number of data recorded by each channel of the person to be tested is N6, and the signal in the jth channel of the ith person to be tested is: Wherein, HbO(i) j This represents the oxyhemoglobin signal in the j-th channel of the i-th person being tested. HbR(i) represents the N6th oxyhemoglobin signal in the j-th channel of the i-th test subject. j This represents the deoxyhemoglobin signal in the j-th channel of the i-th person being tested. This represents the N6th deoxyhemoglobin signal in the j-th channel of the i-th test subject; Module M5.2: data denoising Using a band-pass filter to remove high-frequency noise and low-frequency drift in the signal, the denoised signal is obtained, that is HbO_filtered(i) j = BF(HbO(i) j ,f low ,f high ) HbR_filtered(i) j = BF(HbR(i) j ,f low ,f high ) HbO_filtered(i) = BF (HbO(i), f j j HbR_filtered(i) = BF (HbR(i), f low high f​​ Module M5.3: brain function feature extraction Extract the average oxyhemoglobin concentration change of all channels of the ith person to be tested: Wherein, N5 is the number of data recorded by the channel; Extract the average deoxyhemoglobin concentration change of all channels of the ith person to be tested: Extract the total blood volume change s5(i) of the ith person to be tested: s5 = ΔHbO(i) + ΔHbR(i) wherein mean(HbO_filtered(i) j ) is the mean value of the HbO_filtered(i) j signal over the different channels; and mean(HbR_filtered(i) j ) is the mean value of the HbR_filtered(i) j signal over the different channels. In the module M6: Module M6.1: data acquisition: N persons to be tested are subjected to magnetic resonance imaging (MRI), and the scanning parameters are set; Module M6.2: structure feature extraction: segmenting the brain tissue into grey matter, white matter and cerebrospinal fluid using the automated tool FreeSurfer; calculating the grey matter volume GMV of the prefrontal cortex of the ith person under test using the FreeSurfer toolkit i calculating the average cortical thickness CT of the prefrontal region i calculating the cortical surface area CSA of the prefrontal region i ; Module M6.3: calculation of brain structural index: s6 = GMV i + CT i + CSA i .

Citation Information

Patent Citations

  • Patient report evaluation tool special for evaluating oxaliplatin-induced peripheral neuropathy (OIPN)

    CN117373671A

  • Brain dysfunction auxiliary evaluation method based on multi-modal data fusion

    CN115553752A

  • Multi-dimensional psychological state assessment method based on multi-modal fusion

    CN117796810A