Colorectal cancer early diagnosis model construction method based on Raman spectrum detection
By collecting Raman spectra of colorectal organoids and serum samples, metabolic compensation rate curves and multi-peak joint intensity distribution spectra are generated. Combined with a dual-path neural network, the problem of the inability to effectively capture the dynamic characteristics of metabolic compensation processes and the lack of multimodal feature synergy in existing technologies is solved, thereby improving the sensitivity and accuracy of early diagnosis of colorectal cancer.
Patent Information
- Application Number
- CN202511177751.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2025-11-21
AI Technical Summary
Existing methods for constructing early diagnostic models for colorectal cancer cannot effectively capture the dynamic characteristics of metabolic compensation processes or quantify the cascade effects of metabolic pathways. As a result, the compensatory activation phenomenon unique to early carcinogenesis is not effectively captured. Furthermore, the processing of multi-peak combined intensity distribution lacks a spatial synergistic quantification mechanism for the cascade effects of metabolic pathways, making it impossible to construct a unified evaluation system for metabolic plasticity.
By collecting dynamic and static Raman spectra of colorectal organoids and serum samples, metabolic compensation rate curves and multi-peak joint intensity distribution spectra are generated. A metabolic entropy change model is constructed and input into a dual-path neural network to generate temporal and spatial feature codes. Cross-modal interactive fusion is performed to generate a three-dimensional metabolic stress map, and finally, an early diagnosis model for colorectal cancer is constructed.
It achieves precise capture of the compensatory activation phenomenon unique to early-stage cancer, overcoming the limitation that static spectroscopy cannot quantify metabolic turnover dynamics, and improving the sensitivity and staging accuracy of early colorectal cancer diagnosis.
Smart Images

Figure CN120998470A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of artificial intelligence assisted diagnosis, and in particular to a colorectal cancer early diagnosis model construction method based on Raman spectrum detection. BACKGROUND
[0002] Colorectal cancer early diagnosis technology mainly relies on molecular marker detection and imaging screening. In recent years, non-invasive diagnosis methods based on biological fluids have gradually developed, among which serum tumor marker (such as CEA, CA19-9) detection has become a common clinical method due to its simple operation. With the progress of spectral analysis technology, Raman spectrum has been introduced into the field of cancer diagnosis due to its molecular fingerprint recognition ability, and statistical classification models (such as support vector machine, random forest) are constructed by analyzing the intensity changes of metabolite characteristic peaks in serum to realize the preliminary screening of colorectal cancer. At the same time, the application of organoid model provides a new way for the study of cancerous metabolic dynamics, which simulates the tumor microenvironment by culturing patient-derived organoids in vitro, and captures the metabolic response process by combining time series spectrum.
[0003] The existing colorectal cancer early diagnosis model construction method has significant limitations. Single-point detection relying on serum static spectrum cannot quantify the dynamic characteristics of metabolic compensation process, which leads to the fact that the compensatory activation phenomenon unique to early cancer is not effectively captured, significantly limiting the sensitivity of early detection. In addition, the processing of multi-peak joint intensity distribution lacks a spatial coordination mechanism for the cascade effect of metabolic pathways, and the static entropy change index and dynamic response curve are difficult to align due to the difference in space-time scale, which cannot construct a unified metabolic plasticity evaluation system, resulting in the model failing to capture the overall disorder characteristics of cancer metabolic network. SUMMARY
[0004] In view of the above existing problems, the present application is proposed.
[0005] Therefore, the present application provides a colorectal cancer early diagnosis model construction method based on Raman spectrum detection to solve the problems of insufficient capture of dynamic metabolic characteristics and lack of multi-modal feature coordination.
[0006] To solve the above technical problems, the present application provides the following technical solutions: The application provides a colorectal cancer early diagnosis model construction method based on Raman spectrum detection, which comprises the following steps: collecting organoid dynamic Raman spectrum and serum static Raman spectrum of colorectal organoids and serum samples, generating a metabolic compensation rate curve according to the time sequence change data of the intensity of adenine characteristic peaks in the organoid dynamic Raman spectrum, identifying the intensity distribution of the characteristic peaks of adenine, phenylalanine and phospholipid in the serum static Raman spectrum, generating a multi-peak joint intensity distribution spectrum, constructing a metabolic entropy change model and inputting the multi-peak joint intensity distribution spectrum to generate a patient entropy index, performing spatiotemporal alignment and feature-level fusion on the metabolic compensation rate curve and the patient entropy index to generate a metabolic plasticity feature matrix, constructing a double-path neural network and inputting the metabolic plasticity feature matrix into the double-path neural network to generate time sequence feature encoding and spatial feature encoding. The time sequence feature encoding and the spatial feature encoding are cross-modal interactive fused to generate fused features, the fused features are mapped to the activity gradient plane of the glutamine and fatty acid metabolic pathway through thermal color mapping to generate a three-dimensional metabolic pressure map, a metabolic pathway activity vector is generated according to the three-dimensional metabolic pressure map, the double-path neural network is structurally expanded through the metabolic pathway activity vector, and the double-path neural network after structural expansion is trained and verified to output a colorectal cancer early diagnosis model.
[0007] As a preferred scheme of the colorectal cancer early diagnosis model construction method based on Raman spectrum detection, the steps of collecting the organoid dynamic Raman spectrum and the serum static Raman spectrum of the colorectal organoids and the serum samples are as follows. The patient's colorectal tissue is obtained and the crypt stem cells are separated to form a colorectal organoid, and the patient's venous blood is extracted and centrifuged to form a serum sample. The glutamine pathway antagonist is injected into the colorectal organoid for continuous stimulation, and the organoid dynamic Raman spectrum is collected in real time. The serum sample is added to the surface of the silver nanoparticle substrate for single spectrum scanning to collect the serum static Raman spectrum.
[0008] As a preferred scheme of the colorectal cancer early diagnosis model construction method based on Raman spectrum detection, the steps of generating the metabolic compensation rate curve are as follows. The wavenumber position of the adenine characteristic peak is located from the organoid dynamic Raman spectrum. The time sequence data of the peak intensity of the adenine characteristic peak is extracted, the metabolic compensation rate of adjacent data points is calculated, and a metabolic compensation rate curve is generated.
[0009] As a preferred scheme of the colorectal cancer early diagnosis model construction method based on Raman spectrum detection, the steps of generating the multi-peak joint intensity distribution spectrum are as follows. The serum static Raman spectrum is subjected to denoising, baseline correction and intensity normalization, and the wave number positions of adenine characteristic peaks, phenylalanine characteristic peaks and phospholipid characteristic peaks are located; The intensity integral values of the adenine characteristic peaks, phenylalanine characteristic peaks and phospholipid characteristic peaks are extracted by a peak area integration method to form a multi-peak joint intensity distribution spectrum.
[0010] As a preferred scheme of the early diagnosis model construction method of colorectal cancer based on Raman spectrum detection, the specific steps of generating the patient entropy change index are as follows, A data layer is constructed based on a healthy control group, a modeling layer is constructed based on kernel density estimation, a calculation layer is constructed based on a metabolic disorder quantification mechanism, and a metabolic entropy change model is constructed according to the data layer, the modeling layer and the calculation layer; The multi-peak joint intensity distribution spectrum is input into the metabolic entropy change model, and the patient entropy change index is generated through probability density mapping and negative logarithmic transformation.
[0011] As a preferred scheme of the early diagnosis model construction method of colorectal cancer based on Raman spectrum detection, the specific steps of generating the metabolic plasticity feature matrix are as follows, The metabolic compensation rate curve is subjected to time axis alignment to generate a standardized metabolic compensation rate sequence; The patient entropy change index is extended into a spatiotemporal sequence synchronized with the metabolic compensation rate sequence to generate a patient entropy change index sequence; The metabolic compensation rate sequence and the patient entropy change index sequence are fused to generate a joint feature sequence, and the joint feature sequence is converted into a two-dimensional matrix to output the metabolic plasticity feature matrix.
[0012] As a preferred scheme of the early diagnosis model construction method of colorectal cancer based on Raman spectrum detection, the specific steps of constructing the dual-path neural network are as follows, A time series analysis branch is constructed based on the dynamic fluctuation characteristics of the metabolic plasticity feature matrix, a spatial correlation branch is constructed based on the spatial coordination characteristics of the multi-peak joint intensity distribution spectrum, and an output layer is constructed according to an interface connection mechanism; The dual-path neural network is constructed according to the time series analysis branch, the spatial correlation branch and the output layer.
[0013] As a preferred scheme of the early diagnosis model construction method of colorectal cancer based on Raman spectrum detection, the specific steps of generating the time series feature encoding and the spatial feature encoding are as follows, The metabolic plasticity matrix is input into the dual-path neural network, the metabolic response features are extracted through the time series analysis branch to generate the time series feature encoding; The metabolic spatial correlation features are extracted through the spatial correlation branch to generate the spatial feature encoding.
[0014] As a preferred scheme of the early diagnosis model construction method of colorectal cancer based on Raman spectrum detection, wherein: the three-dimensional metabolic stress atlas is generated, and the specific steps are as follows, The time sequence feature encoding and the space feature encoding are intelligently fused through the attention interaction method to generate the fusion feature. The glutamine metabolic pathway activity axis and the fatty acid metabolic pathway activity axis are defined, and a two-dimensional activity gradient plane is constructed. The metabolic stress intensity is calculated according to the fusion feature, the fusion feature is projected to the activity gradient plane through thermal color mapping, and the three-dimensional metabolic stress atlas is generated by superimposing the metabolic stress intensity.
[0015] As a preferred scheme of the early diagnosis model construction method of colorectal cancer based on Raman spectrum detection, wherein: the three-dimensional metabolic stress atlas is generated, and the specific steps are as follows, The three-dimensional metabolic stress atlas is subjected to spatial grid sampling, the feature point coordinates on the activity gradient plane and the corresponding metabolic stress intensity values are extracted, and a metabolic pathway activity vector is generated. A residual expansion layer is added in front of the output layer of the metabolic pathway activity vector in the double-path neural network to construct a double-path neural network with expanded structure. The double-path neural network with expanded structure is subjected to hierarchical sampling dataset division, cross-entropy loss optimization and diagnosis efficiency verification to output the early diagnosis model of colorectal cancer.
[0016] The present application has the following advantages: the metabolic compensation rate curve is generated by the time sequence change of the organoid dynamic stimulation response under adenine peak, the compensatory activation phenomenon characteristic of early canceration is accurately captured, the limitation that the static spectrum cannot quantify the dynamic of metabolic turnover is broken through; the cross-modal fusion of the time sequence fluctuation and the space collaborative feature is realized by the double-path neural network, the collaborative quantification of the 724-1001-1452 wave number metabolic pathway cascade effect is realized, the technical bottleneck of fragmented analysis of multi-peak characteristics is overcome, and finally the breakthrough improvement of the sensitivity of early diagnosis of colorectal cancer and the accuracy of staging is achieved. BRIEF DESCRIPTION OF DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0018] Figure 1 The flowchart of the early diagnosis model construction method of colorectal cancer based on Raman spectrum detection.
[0019] Figure 2 Flowchart for generating metabolic compensation rate curve.
[0020] Figure 3 Flowchart for generating multi-peak joint intensity distribution spectrum.
[0021] Figure 4 Flowchart for generating feature encoding.
[0022] Figure 5 Flowchart for outputting colorectal cancer early diagnosis model.
[0023] Figure 6 Colorectal region marking image.
[0024] Figure 7 Abdominal abnormal lesion marking image. DETAILED DESCRIPTION
[0025] In order to make the above objectives, features and advantages of the present application more apparent, specific embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0026] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present application. However, it will be apparent to one skilled in the art that the present application can be practiced without the specific details given herein, that the present application can be practiced with other than the described embodiments, and that the present application can be practiced with different or additional components, elements, acts, or steps. Thus, the present application is not limited to the embodiments disclosed herein but instead has wide applicability and scope.
[0027] Secondly, the term "one embodiment" or "an embodiment" as used herein means that a particular implementation can include a particular feature, structure, or characteristic, but every embodiment can not necessarily include the particular feature, structure, or characteristic. Furthermore, the following claims can refer to "one embodiment" or "an embodiment" in the sense of claiming a particular feature, structure, or characteristic of more than one embodiment.
[0028] REFERENCE Figures 1-7 For one embodiment of the present application, the embodiment provides a colorectal cancer early diagnosis model construction method based on Raman spectrum detection, comprising the following steps: S1, collecting organoid dynamic Raman spectrum and serum static Raman spectrum of colorectal organoid and serum sample, and generating metabolic compensation rate curve according to time sequence change data of intensity of adenine characteristic peak in the organoid dynamic Raman spectrum; S1.1, obtaining colorectal tissue of a patient and separating crypt stem cells to form a colorectal organoid; extracting venous blood of the patient and performing centrifugal separation to form a serum sample; It should be noted that after the patient's colorectal biopsy tissue is obtained clinically, it is immediately placed in a pre-cooled HBSS buffer solution to remove mucus and blood residues, a type IV collagenase solution is used to digest the tissue under constant temperature oscillation conditions at 37°C for 30 minutes, after the digestion is terminated, the crypt structure is separated by filtering through a 100 μm cell sieve, the crypt stem cells are collected, mixed with Matrigel, inoculated into a 24-well culture plate, and placed in a growth culture medium for 5-7 days to form a colorectal organoid; 5 mL of fasting venous blood is extracted from the patient and injected into a vacuum blood collection tube, which is centrifuged at 3000 rpm for 15 minutes after standing at room temperature for 30 minutes, and the supernatant serum is collected and divided into cryogenic tubes for storage at -80°C for standby, forming a serum sample.
[0029] S1.2, inject glutamine pathway antagonists into the colorectal organoid for continuous stimulation, and simultaneously collect dynamic Raman spectra of the organoid in real time through the intelligent spectrum sensing terminal; It should be noted that the colorectal organoid is transferred to a special culture dish of the intelligent spectrum sensing terminal, a glutamine pathway antagonist DON solution with a concentration of 10 μM is prepared, and the microenvironment of the organoid is continuously injected at a rate of 0.5 μL / min through microinjection; the intelligent spectrum sensing terminal is started to emit a 633 nm laser light source to continuously scan the colorectal organoid, and Raman spectra in the range of 400-1800 cm -1 with a sampling interval of 5 seconds (defined based on the metabolic response kinetic characteristics) are collected for 60 minutes to generate dynamic Raman spectra of the organoid.
[0030] S1.3, drop the serum sample onto the surface of the silver nanoparticle substrate, and perform a single spectrum scan through the intelligent spectrum sensing terminal to collect the static Raman spectrum of the serum; It should be noted that the serum sample is thawed to room temperature, 0.5 μL is taken and dropped onto the surface of the silver nanoparticle substrate, and a uniform adsorption layer is formed after standing for 10 minutes; the silver nanoparticle substrate is placed on the stage of the intelligent spectrum sensing terminal, the 633 nm laser light source is focused on the center point of the adsorption layer, and a single spectrum scan mode is set, and Raman spectra in the range of 400-1800 cm -1 with a 4-second integration time (defined based on the peak window of Raman scattering effect) are collected to output the static Raman spectrum of the serum.
[0031] S1.4, locate the wave number position of the adenine characteristic peak from the dynamic Raman spectrum of the organoid, extract the peak intensity time series data, calculate the metabolic compensation rate of adjacent data points, and generate a metabolic compensation rate curve.
[0032] It should be noted that wavenumber calibration was performed on the dynamic Raman spectra of organoids to locate the 724 wavenumber position, which is the characteristic peak position of adenine (based on the definition of the specific vibrational mode of the chemical bond in the adenine molecule, representing the CH bond bending vibration mode). The peak intensity value at each sampling time point was extracted to form a time-series data sequence: The discrete data point position corresponding to the 724 wavenumber position was locked in the wavenumber-calibrated dynamic Raman spectra of organoids. The peak intensity value corresponding to the 724 wavenumber position in the dynamic Raman spectra of organoids was extracted at each sampling time point. Sampling was performed continuously for 60 minutes to form a time-series intensity sequence that changes over time. The extraction operation relied entirely on the real-time recording of the physical signal intensity by the intelligent spectral sensing terminal; the peak intensity values came directly from the raw output signal of the intelligent spectral sensing terminal. The metabolic compensation rate at adjacent time points was calculated using the following expression: ; in, Indicates time as The metabolic compensation rate at that time; Indicates the first Peak intensity values corresponding to each sampling time point; Indicates the first Peak intensity values corresponding to each sampling time point; This indicates the interval between adjacent sampling time points, which is 5 seconds. This represents a specific sampling time point; This indicates a conversion from seconds to minutes. The metabolic compensation rate at adjacent sampling time points is smoothed using a moving average with a sliding window. The sliding window size is 7 data points (defined based on the periodic characteristics of the metabolic response signal and the signal-to-noise ratio optimization requirements, covering a 35-second time span to fully cover a single metabolic oscillation cycle and avoid signal truncation distortion). A moving average smoothing method is used to filter high-frequency jitter. The smoothed metabolic compensation rate within the sliding window is calculated point-by-point, and the expression is: ; in, Indicates the first The center point of the sliding window at each sampling time point The corresponding smooth metabolic compensation rate; This represents the index of the data point in the sliding window, with a value range of [ , ]; Indicates the center point; Indicates the center point of the current sliding window. The sampling time point; Indicates the first The sampling time points corresponding to each data point The metabolic compensation rate; By replacing the metabolic compensation rate with a smooth metabolic compensation rate, noise suppression and curve smoothing are achieved, and a metabolic compensation rate response curve is generated, with the horizontal axis representing stimulation time (0-60 minutes) and the vertical axis representing the smooth metabolic compensation rate.
[0033] S2, identify the characteristic peak intensity distribution of adenine, phenylalanine and phospholipid in serum static Raman spectrum, generate multi-peak joint intensity distribution spectrum; construct a metabolic entropy model and input the multi-peak joint intensity distribution spectrum to generate a patient entropy index; S2.1, perform denoising, baseline correction and intensity normalization processing on the serum static Raman spectrum, and locate the wave number positions of the adenine characteristic peak, the phenylalanine characteristic peak and the phospholipid characteristic peak; It should be noted that based on the serum static Raman spectrum, nine-point moving average filtering is used to remove high-frequency noise: the wave number axis of the serum static Raman spectrum is scanned point by point, and the peak intensity values of the four adjacent points (a total of nine points) before and after each target wave number position are taken to calculate the arithmetic mean value, and the arithmetic mean value is used to replace the original target point peak intensity value. By traversing the entire serum static Raman spectrum (400-1800 wave number) through a sliding window, the jitter of the sharp peaks caused by instrument random error is eliminated. By baseline correction, the fluorescence background interference is eliminated: the filtered serum Raman spectrum is continuously identified, and a smooth curve enveloping the lowest point of the filtered serum static Raman spectrum is automatically fitted as the fluorescence background baseline. The fluorescence background baseline is subtracted from the filtered serum static Raman spectrum point by point to eliminate the wide-band background interference caused by sample autofluorescence; After maximum normalization processing, the spectrum intensity is standardized: the maximum peak intensity value in the whole wave number range is identified by globally scanning the baseline-corrected serum static Raman spectrum, and the peak intensity value at each wave number position in the serum static Raman spectrum is compared with the maximum peak intensity value to make all peak intensity values scaled to the [0-1] interval, eliminating the absolute intensity deviation caused by concentration difference or detection condition fluctuation. According to the bending vibration mode of the carbon-hydrogen bond in the adenine molecule, the 724 wave number position is accurately located as the adenine characteristic peak position; based on the symmetric stretching vibration mode of the benzene ring structure of the phenylalanine molecule, the 1001 wave number position is accurately located as the phenylalanine characteristic peak position; based on the bending vibration mode of the methylene group in the phospholipid molecule, the 1452 wave number position is accurately located as the phospholipid characteristic peak position, and the wave number position of the adenine characteristic peak, the phenylalanine characteristic peak and the phospholipid characteristic peak is located.
[0034] S2.2, the intensity integral value of the adenine characteristic peak, the phenylalanine characteristic peak and the phospholipid characteristic peak is extracted by peak area integration method, and a multi-peak joint intensity distribution spectrum is formed; It should be noted that the wave number position ± 5 is set as the wave number integral interval (defined based on the inherent peak width of the biomarker and the instrument resolution, the wave number integral interval of wave number position ± 5 can cover the peak width center area, and ensure that the wave number integral interval of the adenine characteristic peak is 719-729 wave numbers) centered on the target peak wave number position; the peak intensity value of the adenine characteristic peak, the phenylalanine characteristic peak and the phospholipid characteristic peak in each wave number integral interval is extracted from the baseline corrected and maximum normalized serum static Raman spectrum, and the intensity integral value in the wave number integral interval is calculated point by point by using the trapezoidal rule, and the expression is: ; wherein, represents the intensity integral value of the wave number position ; represents the characteristic peak position index, which is 724, 1001 and 1452; represents the wave number integral interval index; represents the lower limit of the wave number integral interval, which is the target peak wave number position-5; represents the upper limit of the wave number integral interval, which is the target peak wave number position+5; represents the peak intensity value of the target peak corresponding to the position in the wave number integral interval ; represents the peak intensity value of the target peak corresponding to the position in the wave number integral interval ; represents the wave number step, that is, the integral infinitesimal width, which is fixed at 1; The adenine characteristic peak intensity integral value, the phenylalanine characteristic peak intensity integral value and the phospholipid characteristic peak intensity integral value are combined into a three-dimensional vector in ascending order of wave number position to generate a multi-peak joint intensity distribution spectrum.
[0035] S2.3, a data layer is constructed based on a healthy control group, a modeling layer is constructed based on kernel density estimation, a calculation layer is constructed based on metabolic disorder quantification mechanism, and a metabolic entropy change model is constructed according to the data layer, the modeling layer and the calculation layer; It should be noted that fresh tumor tissues and paired normal tissues adjacent to cancer of colorectal cancer patients are obtained through clinical surgery and endoscopy center, and histological grading and TNM staging confirmation are performed by pathologists. The patient's preoperative fasting venous blood is collected simultaneously, and the serum is collected using a vacuum blood collection tube, and is divided into cryogenic tubes labeled with patient number, sampling date and clinical staging information, and is immediately stored in a ultra-low temperature freezer at minus 80 degrees Celsius. The serum samples of the healthy control group are collected and processed in the same process and labeled as healthy state. All samples are associated with complete electronic medical record data, which are independently input into the encrypted database by full-time data administrators, and finally form a colorectal cancer special database; Serum samples of 35 healthy controls without digestive system diseases, without malignant tumor history and with normal gastrointestinal endoscopy results in the past 1 year were called from the colorectal cancer special database; the serum samples of the healthy control group were added to the silver nanoparticle substrate surface, and single spectrum scanning was performed by an intelligent spectrum sensing terminal to collect the static Raman spectrum of the healthy control group, and the multi-peak joint intensity distribution spectrum of the healthy control was generated by peak area integration method; The multi-peak joint intensity distribution spectrum of the healthy control group of 35 subjects was collected, and the adenine characteristic peak intensity integral value, phenylalanine characteristic peak intensity integral value and phospholipid characteristic peak intensity integral value of each subject were integrated into a three-dimensional matrix data set (row corresponding to subject number, column corresponding to three characteristic peak intensity integral values) according to the row-column structure, forming a benchmark database representing healthy metabolic homeostasis, and the data layer was constructed; Based on the three-dimensional matrix data set of the healthy control group, the joint distribution of the three groups of intensity integral values of all subjects in the three-dimensional matrix data set was density estimated by using the Gaussian kernel function, the best bandwidth parameter of each intensity integral value dimension was calculated by the Silverman bandwidth optimization criterion, and the joint probability density function describing the healthy metabolic state was constructed according to the Gaussian kernel function and the best bandwidth parameter, and the modeling layer was constructed. The expression of the joint probability density function is: ; The expression of the Gaussian kernel function is: ;
[0036] The expression of the Silverman bandwidth optimization criterion is: ; Wherein, represents the healthy joint probability density; represents the three-dimensional matrix data set; represents the number of healthy control subjects, which is 35; represents the subject index, which is 1- ; represents the characteristic peak dimension, , adenine dimension, , phenylalanine dimension, , phospholipid dimension; represents the best bandwidth parameter of the th dimension; represents the Gaussian kernel function; represents the standardized offset; represents the intensity integral value of the th dimension; represents the intensity integral value of the th The integral value of the intensity of the dimension; The standard deviation, the data dispersion of a dimension of the healthy control group; The interquartile range, the distribution range of the middle 50% data in the healthy control group; The normal distribution coefficient, a statistical constant; The health entropy change index is calculated by the health joint probability density through the negative logarithmic transformation formula, the calculation layer is constructed, and the expression is: ; Wherein, The health entropy change index; According to the three-level series connection of the data layer, the modeling layer and the calculation layer, the metabolic entropy change model is constructed; the metabolic entropy change model is a statistical model directly established by the health control group data, and is not a machine learning model requiring parameter optimization, so the metabolic entropy change model does not need iterative training, and the whole construction process only depends on the data distribution characteristics of the health control group, without back propagation mechanism; the metabolic entropy change model outputs a fixed joint probability density function And the negative logarithmic transformation formula; the fixed joint probability density function And the negative logarithmic transformation formula are directly called for forward calculation when diagnosing a patient, without updating any parameters, so there is no training behavior such as parameter iterative optimization.
[0037] S2.4, input the multi-peak joint intensity distribution spectrum into the metabolic entropy change model, map the probability density and perform negative logarithmic transformation to generate the patient entropy change index.
[0038] It should be noted that the multi-peak joint intensity distribution spectrum generated by the patient serum sample is input into the metabolic entropy change model, the joint probability density function is called to calculate the patient joint probability density, and then the patient entropy change index is generated according to the patient joint probability density through the negative logarithmic transformation formula; It should also be noted that since the health joint probability density indicates that the metabolic state of healthy people is concentrated in the high probability density area (such as the mean of the health joint probability density of the health control group is 0.40), the higher the patient joint probability density, the closer the patient's metabolic characteristics to the health benchmark; when the patient's metabolism is disordered (such as abnormal proliferation of colorectal cancer cells), the multi-peak joint intensity distribution spectrum deviates significantly from the core area of the health control group distribution, resulting in a sharp decrease in the patient joint probability density; and the patient entropy change index is calculated and generated according to the patient joint probability density, and the negative logarithmic operation makes the patient entropy change index increase inversely, therefore, the larger the patient entropy change index value, the greater the deviation of the patient's metabolic state from the health benchmark, that is, the less healthy the patient is.
[0039] S3, time-space alignment and feature-level fusion of the metabolic compensation rate curve and the patient entropy change index are performed to generate a metabolic plasticity feature matrix; S3.1, perform time axis alignment on metabolic compensation rate curve, generate normalized metabolic compensation rate sequence; It should be noted that the unified time range of the metabolic compensation rate curve is set to 0 seconds to 3600 seconds (60 minutes) after the start of glutamine antagonist stimulation, and the time axis is divided into 3600 equidistant points (time stamp sequence is 0, 1, 2,..., 3600 seconds) at a fixed interval of 1 second (based on metabolic response dynamic monitoring standard definition), and finally a standard time reference axis is generated; The time period not covered in the metabolic compensation rate curve (such as a patient's metabolic compensation rate curve for only 50 minutes) is filled with missing points by linear interpolation: for each organoid dynamic Raman spectrum metabolic compensation rate curve, locate the missing time stamp on the standard time reference axis (such as a cancer sample without data at 3000 seconds), find the nearest two time points before and after the missing time stamp and the corresponding metabolic compensation rate, generate the metabolic compensation rate corresponding to the missing time stamp by linear interpolation formula, traverse all missing time stamps to perform linear interpolation, ensure that there is a rate value every second, and generate a continuous normalized metabolic compensation rate curve; Realign the normalized metabolic compensation rate curve to the standard time reference axis to generate a normalized metabolic compensation rate sequence (length 3600 points), the time stamp is strictly aligned with the 0-3600 second integer sequence, and the expression of the linear difference formula is: ; Wherein, represents the metabolic compensation rate corresponding to the time t; represents the time that needs linear interpolation; represents the nearest time point before the linear interpolation time point t; represents the nearest time point after the linear interpolation time point t; represents the metabolic compensation rate corresponding to the time t; represents the metabolic compensation rate corresponding to the time t. represents the metabolic compensation rate corresponding to the time t. represents the metabolic compensation rate corresponding to the time t. represents the metabolic compensation rate corresponding to the time t. represents the metabolic compensation rate corresponding to the time t. represents the metabolic compensation rate corresponding to the time t.
[0040] S3.2, extend the patient entropy index to a spatiotemporal sequence synchronized with the metabolic compensation rate sequence, generate a patient entropy index sequence; It should be noted that a time axis sequence with a length of 3600 (time stamp 0, 1, 2,..., 3600 seconds) is created, the same patient entropy index is filled at each time point, a patient entropy index sequence is generated, and the static metabolic disorder quantification value and the dynamic metabolic response sequence are spatiotemporally aligned.
[0041] S3.3, fuse the metabolic compensation rate sequence and the patient entropy index sequence to generate a joint feature sequence, and convert the joint feature sequence into a two-dimensional matrix to output a metabolic plasticity feature matrix.
[0042] It should be noted that the corresponding elements of the metabolic compensation rate sequence and the patient entropy index sequence are alternately spliced in chronological order by second, each second containing a metabolic compensation rate and a patient entropy index, to generate a joint feature sequence; Based on the metabolic response stage division and the feature fusion requirement, the two-dimensional matrix layout requirement is defined as 60 rows and 120 columns: 60 rows correspond to 3600 seconds (60 minutes) of stimulation throughout, and each row represents a 1-minute time block; 120 columns correspond to each second containing two features (metabolic compensation rate and patient entropy index); the metabolic compensation rate and the patient entropy index are filled into the two-dimensional matrix in chronological order, and finally a metabolic plasticity feature matrix is output.
[0043] S4, construct a double-path neural network, and input the metabolic plasticity feature matrix into the double-path neural network to generate time sequence feature encoding and spatial feature encoding; S4.1, construct a time sequence analysis branch based on the dynamic fluctuation characteristics of the metabolic plasticity feature matrix, construct a spatial correlation branch based on the spatial coordination characteristics of the multi-peak joint intensity distribution spectrum, and construct an output layer according to an interface connection mechanism; It should be noted that the time sequence analysis branch is constructed, a four-level series connection hierarchical structure is defined, the first layer is a one-dimensional convolution operation layer, 60 seconds of metabolic compensation rate and patient entropy index are scanned through a 5-time-point width convolution kernel, the product of the values of 5 consecutive time points in the convolution kernel and the preset weight coefficient (based on the metabolic response period characteristics) is accumulated and summed to generate a local window feature response value; the local window feature response value is traversed by a sliding window with a 1-second step to strengthen the short-time fluctuation mode and weaken the low-frequency noise; a short-time feature map sequence is output which retains the short-time fluctuation characteristics.
[0044] The second layer is a dilated convolution operation layer, which selects non-continuous time point data with an interval of 2 for weighted accumulation, expands the receptive field to 7 seconds to capture cross-minute metabolic oscillation, and traverses the short-time feature map sequence by a sliding window to strengthen the stage transition characteristics, and outputs a long-time feature map sequence which retains the cross-minute oscillation characteristics; The third layer is a max-pooling operation layer, which scans the long-time feature map sequence with an 8-time-point window, outputs the maximum value in the time point window each time, and compresses the time dimension to 7.5 seconds resolution after traversal, retains the fluctuation peak value, and outputs a 450-dimensional time sequence feature map; The fourth layer is a full connection mapping layer, 450-dimensional time sequence feature maps are processed by 32 independent structural units, each unit weights and adds all input feature values and activates output to generate compressed time sequence feature encoding (32-dimensional); the hierarchical connection rule is sequential transmission (convolution→hole→pooling→mapping), and construction of the time sequence analysis branch is completed; The spatial correlation branch is constructed, and a four-level serial hierarchical structure is defined. The first layer is a feature recombination operation layer, which combines the metabolic compensation rate per second and the patient entropy index into a two-dimensional vector. The data per second is merged from two columns into one column to generate a 60x60 spatial feature matrix; The second layer is a cooperative weighting operation layer, which analyzes the intensity change direction of 724 wave number position and 1001 wave number position per second. The same rising and falling mode is assigned a high weight to strengthen the cooperativity (such as cancer group cooperative rising), and the reverse change is assigned a low weight to suppress noise; The third layer is a pathway enhancement operation layer, which identifies the same direction change mode of the intensity of 1001 wave number position and 1452 wave number position per second, and assigns a high enhancement coefficient to the metabolic pathway cascade effect (phenylalanine drives phospholipid synthesis) to amplify the associated features; The fourth layer is a topological aggregation operation layer, which combines adjacent second feature vectors and fuses 724-1001-1452 wave number cooperative characteristics to generate a local spatial topological feature segment, which is compressed into a spatial feature encoding (32-dimensional) by dimension reduction. The hierarchical connection rule is sequential transmission (recombination→weighting→enhancement→aggregation), and the construction of the spatial correlation branch is completed; The time sequence feature encoding of the time sequence analysis branch and the spatial feature encoding of the spatial correlation branch are directly received, transmitted to the downstream through independent channels, and the output layer does not perform any feature transformation or fusion operation, only as an output interface of the double-path terminal encoding, and the construction of the output layer is completed.
[0045] S4.2, constructing a double-path neural network according to the time sequence analysis branch, the spatial correlation branch and the output layer; It should be noted that the time sequence analysis branch and the spatial correlation branch transmit the generated time sequence feature encoding and spatial feature encoding to the output layer through parallel independent processing, and construct a double-path neural network through parallel independent processing and interface connection; The historical sample data of 112 cancer patients and 28 healthy subjects in the colorectal cancer special database is called. 80% of the historical samples (89 cancer and 22 healthy) are divided into the training set, and 20% (23 cancer and 6 healthy) are divided into the validation set; the metabolic plasticity feature matrix of the training set is input into the double-path neural network to perform forward calculation, the time series analysis branch extracts dynamic fluctuation features to generate time series feature encoding, the spatial correlation branch quantifies the multi-peak synergistic characteristics to output spatial feature encoding, and the output layer receives and outputs time series feature encoding and spatial feature encoding; the mean of the time series feature encoding and the mean of the spatial feature encoding of 35 samples of the healthy control group are used as the health reference feature encoding, respectively, to compare the difference between the training set time series feature encoding and the spatial feature encoding and the health reference feature encoding, and guide the training set time series feature encoding to the same stage time series clustering center (such as I stage cancer features close to each other), and the spatial feature encoding to the same stage spatial clustering center (such as I stage cancer spatial encoding close to each other); according to the difference and the aggregation effect, the network weight is adjusted in the reverse direction, the capture ability of the time series analysis branch to the metabolic slope mutation is optimized, and the sensitivity of the spatial correlation branch to the synergistic break is optimized; after each round of training, observe the distribution diagram of the time series feature encoding and the spatial feature encoding of the validation set, when the separation degree of the healthy control group and the training set feature cluster group has no visible improvement for 10 consecutive rounds, terminate the training, and complete the training of the double-path neural network.
[0046] S4.3, input the metabolic plasticity matrix into the double-path neural network, extract the metabolic response features through the time series analysis branch, and generate time series feature encoding; extract the metabolic spatial correlation features through the spatial correlation branch, and generate spatial feature encoding.
[0047] It should be noted that the metabolic plasticity feature matrix is input into the double-path neural network, the time series analysis branch generates time series feature encoding through four-level serial operations (one-dimensional convolution → hollow convolution → maximum pooling → full connection mapping); the synchronous spatial correlation branch generates spatial feature encoding through four-level serial operations (feature reorganization → synergistic weighting → pathway strengthening → topology aggregation); the output layer receives and transmits double-path time series feature encoding and spatial feature encoding through independent interfaces.
[0048] S5, cross-modal interactive fusion of time series feature encoding and spatial feature encoding to generate fusion features; mapping the fusion features to the activity gradient plane of glutamine and fatty acid metabolic pathways through thermal color mapping to generate three-dimensional metabolic stress atlas; S5.1, intelligent fusion of time series feature encoding and spatial feature encoding through attention interaction method to generate fusion features; It should be pointed out that the time sequence feature code is used as the query vector, the space feature code is used as the key vector and the value vector, the correlation weight of time sequence to space is calculated by dot product attention mechanism, and the weighted fusion vector of space feature in time sequence dimension is generated; at the same time, the space feature code is used as the query vector, the time sequence feature code is used as the key vector and the value vector, the correlation weight of space to time sequence is calculated, and the weighted fusion vector of time sequence feature in space dimension is generated; the two weighted fusion vectors are spliced into 64-dimensional joint features, and are compressed to fusion features (32-dimensional) through full connection transformation (64 neurons); the dot product attention interaction formula expression is: ; wherein, represents the weighted fusion vector; represents a normalization function; represents a query vector; represents a key vector; represents a value vector; represents the key vector dimension, which is fixed at 32; represents a matrix transposition operator, which is used to interchange the rows and columns of a matrix; represents a query-key dot product, which is used to calculate the similarity of time sequence-space features; represents a scaled dot product, which prevents gradient disappearance.
[0049] S5.2, define the glutamine metabolic pathway activity axis and the fatty acid metabolic pathway activity axis, and construct a two-dimensional activity gradient plane; It should be pointed out that the glutamine metabolic pathway activity axis is defined as the change direction of the intensity integral value of the adenine characteristic peak at 724 wavenumber position, the positive direction represents the enhancement of glutamine metabolic activity (such as accelerated synthesis of adenine nucleotide in the cancer group), and the negative direction represents metabolic inhibition (such as stable synthesis in the healthy group); the fatty acid metabolic pathway activity axis is defined as the change direction of the intensity integral value of the phospholipid characteristic peak at 1452 wavenumber position, the positive direction represents the increase of fatty acid metabolic activity (such as hyperlipid synthesis in the cancer group), and the negative direction represents the weakening of metabolism (such as balanced synthesis in the healthy group); Construction of two-dimensional activity gradient plane: taking the glutamine metabolic pathway activity axis as the horizontal axis (x-axis) and the fatty acid metabolic pathway activity axis as the vertical axis (y-axis), the two axes intersect orthogonally at the origin, the intensity integral value of the adenine characteristic peak at 724 wavenumber position (glutamine activity) and the intensity integral value of the phospholipid characteristic peak at 1452 wavenumber position (fatty acid activity) in each serum static Raman spectrum are mapped to the plane coordinate point (x, y), forming an activity distribution scatter plot; the activity distribution scatter plot is divided into four quadrants (such as the first quadrant, the double high activity area is the advanced stage of cancer, and the third quadrant, the double low activity area is the healthy group), and the two-dimensional activity gradient plane of quantitative metabolic pathway cooperative activity is generated by coloring the serum static Raman spectrum density gradient (such as the red high-density area indicating high risk of cancer).
[0050] S5.3, Calculate metabolic stress intensity according to fusion features, project fusion features to active gradient plane by heat color mapping, and superimpose metabolic stress intensity to generate three-dimensional metabolic stress atlas.
[0051] It should be noted that the linear output value is generated by weighted accumulation of all dimension feature values of the fusion features, and then compressed to the interval of 0-1 by Sigmoid activation function to generate metabolic stress intensity (the greater the metabolic stress intensity value, the more serious the metabolic disorder); Take the x-axis and y-axis of the two-dimensional active gradient plane as the base, take the two-dimensional active gradient plane coordinate point (x, y) as the spatial position anchor point, create a circular influence area centered on the spatial position anchor point, calculate the distance between all fusion feature points in the circular influence area and the spatial position anchor point, and calculate the density contribution value by Gaussian weight distribution function (the center point contributes the most, and the edge decays by distance), the expression is: ; Wherein, represents the density contribution value; represents the total number of fusion feature points in the circular influence area; represents the fusion feature point index, taking value 1- ; represents the fusion feature value of the th fusion feature point; represents the distance from the th fusion feature point to the spatial position anchor point; represents the Gaussian weight coefficient of the distance ; Traverse the entire two-dimensional active gradient plane, accumulate the density contribution value of the surrounding (for example, within a radius of 0.5) for each circular influence area, generate the heat density value of the current circular influence area, render all heat density values according to the red-blue gradient color system (high-density area red, low-density area blue), form a heat color mapping map covering the two-dimensional active gradient plane; Take the two-dimensional active gradient plane as the base, take each metabolic stress intensity value (z) as the height value, and perform inverse distance weighted interpolation operation on all sample points (x, y, z) in the two-dimensional active gradient plane: take the target sample point as the center, calculate the distance of all sample points within a radius of 1 unit, assign weights according to the distance, and weightedly accumulate the z value of the target sample point, traverse all sample points to generate a continuous height surface; the heat color mapping map is used as the color attribute of the surface, and the red-blue gradient color rendering three-dimensional metabolic stress atlas is output, wherein the x-axis quantifies glutamine metabolic activity, the y-axis quantifies fatty acid metabolic activity, and the z-axis height and color represent metabolic stress intensity; It should be noted that the prior art realizes intestinal adenocarcinoma screening through peripheral blood static Raman spectrum detection combined with a deep learning model, but still has the limitations that static detection cannot quantify metabolic dynamic response and lacks a multi-peak feature collaborative analysis mechanism. The present scheme has breakthrough advantages in dynamic metabolic compensation rate curve generation of organoids, cross-modal fusion of double-path neural networks, and three-dimensional metabolic stress atlas construction, significantly improving early cancer detection rate and staging accuracy.
[0052] S6. Generating a metabolic pathway activity vector based on the three-dimensional metabolic stress atlas, performing structural expansion on the double-path neural network through the metabolic pathway activity vector, and performing training and verification on the structurally expanded double-path neural network to output a colorectal cancer early diagnosis model.
[0053] S6.1. Performing spatial grid sampling on the three-dimensional metabolic stress atlas, extracting feature point coordinates and corresponding metabolic stress intensity values on the activity gradient plane, and generating a metabolic pathway activity vector; It should be noted that an equidistant grid of 0.1x0.1 activity units is created based on the three-dimensional metabolic stress atlas as a base plane, each grid point is traversed to extract the metabolic stress intensity value at the corresponding position, the activity coordinates and metabolic stress intensity value of each grid point are combined into a three-dimensional feature point, and all grid points are arranged in priority order as a metabolic pathway activity vector.
[0054] S6.2. Adding a residual expansion layer in front of the output layer of the double-path neural network according to the metabolic pathway activity vector to construct a structurally expanded double-path neural network; It should be noted that a three-level series of sublayers are defined, the first sublayer is a metabolic activity compression layer, which compresses the metabolic pathway activity vector to a 256-dimensional metabolic feature vector through a weighted merging operation; the second sublayer is an original code expansion layer, which combines the time series feature code and the spatial feature code into a 64-dimensional original code, and fills 192-dimensional zero values to expand to a 256-dimensional alignment vector (the first 64 dimensions are original codes, and the last 192 dimensions are zeros); the third sublayer is a residual addition layer, which adds the 256-dimensional metabolic feature vector and the 256-dimensional alignment vector point by point according to the same position to generate a 256-dimensional residual feature vector; the metabolic activity compression layer and the original code expansion layer receive input in parallel, the output of the metabolic activity compression layer and the original code expansion layer is jointly input into the residual addition layer to perform element-wise addition operation, and the output of the residual addition layer is directly connected to the input interface of the output layer to construct the structurally expanded double-path neural network.
[0055] S6.3. Performing hierarchical sampling dataset division, cross-entropy loss optimization, and diagnosis performance verification on the structurally expanded double-path neural network to output a colorectal cancer early diagnosis model.
[0056] It should be noted that based on the 140 samples (112 cancer + 28 healthy) of the colorectal cancer special database, the cancer stage (I-IV) and the healthy group are stratified sampling, 70% of the samples (78 cancer + 20 healthy) are divided into the training set, 15% (17 cancer + 4 healthy) are the validation set, and 15% (17 cancer + 4 healthy) are the test set, to ensure that the proportion of cancer and healthy groups in the training set, validation set and test set is homogeneously distributed; The training set is input into the structure expanded double-path neural network, and the forward output is a 256-dimensional residual feature vector. The product of each value of the 256-dimensional residual feature vector and the preset weight coefficient is added and summed, and the bias constant is added to generate an initial prediction value. The initial prediction value is compressed to the 0-1 interval by the Sigmoid function, and the cancer probability value is output. The deviation of the cancer probability value and the true label is quantified by the cross-entropy loss function, and the Adam optimizer (learning rate 0.001) is used to update the network weight (time convolution kernel, spatial attention weight and residual expansion layer parameter) by back propagation. Each batch input is 32 samples, and the iteration training is performed until the loss converges. The weight coefficient and the bias constant are not preset fixed values, but are dynamically optimized and generated through the training process: a small value is randomly assigned during initial training (such as weight coefficient-0.1 to 0.1 uniformly distributed, bias constant 0), the training set sample is generated through forward propagation to generate a prediction probability value, the deviation of the prediction probability value and the true label is calculated, the gradient of the deviation on the weight coefficient and the bias constant is backtracked layer by layer according to the chain rule, and the weight coefficient and the bias value are adjusted according to the gradient direction and the learning rate (such as increasing the weight when the gradient is negative). Repeat iteration until the loss value is minimized (such as the prediction probability of the cancer group tends to 1 and the healthy group tends to 0), and fix the optimal weight coefficient and bias constant. By adjusting the cancer determination threshold (defined based on the ROC curve best balance point and clinical diagnosis requirements), the proportion of correctly identifying cancer samples (sensitivity) and the proportion of misjudging healthy samples as cancer (false positive rate) of the structure expanded double-path neural network under the cancer determination threshold are recorded, and the sensitivity-false positive rate relationship curve is drawn. The area under the sensitivity-false positive rate relationship curve (AUC) is used to quantify the overall discrimination performance. When AUC does not improve for 10 consecutive rounds or reaches AUC threshold (such as 0.95), the training is terminated, and the early diagnosis model of colorectal cancer is output.
[0057] The embodiment also provides a computer device suitable for the early diagnosis model construction method of colorectal cancer based on Raman spectrum detection, which comprises a memory and a processor; the memory is used to store computer executable instructions, and the processor is used to execute the computer executable instructions to realize the early diagnosis model construction method of colorectal cancer based on Raman spectrum detection proposed in the above embodiment.
[0058] The computer device can be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen and an input device connected by a system bus. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is configured to perform wired or wireless communication with an external terminal. The wireless communication can be achieved by WIFI, an operator network, NFC (Near Field Communication) or other technologies. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer overlaid on the display screen, or a key, trackball or touchpad arranged on the shell of the computer device, or an external keyboard, touchpad or mouse, etc.
[0059] The embodiment also provides a storage medium having a computer program stored thereon, the program being executed by a processor to implement the method for constructing an early diagnosis model of colorectal cancer based on Raman spectrum detection as described in the above embodiment. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic storage, a flash memory, a magnetic disk or an optical disk.
[0060] In summary, the present application breaks through the limitation of static spectrum that cannot quantify the dynamic of metabolic turnover by generating metabolic compensation rate curve based on the timing change of adenine peak under dynamic stimulation response of organoids, accurately capturing the compensatory activation phenomenon unique to early cancer, and combining the cross-modal fusion of timing fluctuation and spatial coordination characteristics by double-path neural network to achieve the synergistic quantification of 724-1001-1452 wave number metabolic pathway cascade effect, overcoming the technical bottleneck of fragmented analysis of multi-peak characteristics, and finally achieving a breakthrough in sensitivity and staging accuracy of early diagnosis of colorectal cancer.
[0061] It should be noted that the above examples are only used to illustrate the technical solutions of the present application but not limit the present application. Although the present application is described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or equivalently replaced, without departing from the spirit and scope of the technical solutions of the present application, which should be covered in the scope of the claims of the present application.
Claims
1. A method for constructing an early diagnostic model for colorectal cancer based on Raman spectroscopy, characterized in that: comprising, Collecting organoid dynamic Raman spectrum and serum static Raman spectrum of colorectal organoids and serum samples, generating metabolic compensation rate curve according to time series change data of intensity of adenine characteristic peak in organoid dynamic Raman spectrum; Identifying intensity distribution of characteristic peaks of adenine, phenylalanine and phospholipid in serum static Raman spectrum, generating multi-peak joint intensity distribution spectrum; constructing metabolic entropy change model and inputting multi-peak joint intensity distribution spectrum to generate patient entropy change index; Spatiotemporal alignment and feature level fusion of metabolic compensation rate curve and patient entropy change index to generate metabolic plasticity feature matrix; Constructing double-path neural network and inputting metabolic plasticity feature matrix into double-path neural network to generate time series feature encoding and spatial feature encoding; Cross-modal interactive fusion of time series feature encoding and spatial feature encoding to generate fusion feature; mapping fusion feature to activity gradient plane of glutamine and fatty acid metabolic pathway through thermal color mapping to generate three-dimensional metabolic pressure map; Generating metabolic pathway activity vector according to three-dimensional metabolic pressure map, structure expansion of double-path neural network through metabolic pathway activity vector, and training and verification of structure expanded double-path neural network to output early colorectal cancer diagnosis model. 2.The method for constructing a colorectal cancer early diagnosis model based on Raman spectrum detection according to claim 1, wherein: The specific steps of collecting organoid dynamic Raman spectrum and serum static Raman spectrum of colorectal organoids and serum samples are as follows, Obtaining patient colorectal tissue and separating crypt stem cells to form colorectal organoids; extracting patient venous blood and centrifuging to form serum samples; Injecting glutamine pathway antagonist into colorectal organoids for continuous stimulation, and collecting organoid dynamic Raman spectrum in real time; Drop serum samples on the surface of silver nanoparticle substrate for single spectrum scanning to collect serum static Raman spectrum. 3.The method for constructing a colorectal cancer early diagnosis model based on Raman spectrum detection according to claim 2, wherein: The specific steps of generating metabolic compensation rate curve are as follows, Locating the wave number position of the adenine characteristic peak from the organoid dynamic Raman spectrum; Extracting time series data of peak intensity of adenine characteristic peak, calculating metabolic compensation rate of adjacent data points, and generating metabolic compensation rate curve. 4.The method for constructing a colorectal cancer early diagnosis model based on Raman spectrum detection according to claim 3, characterized in that: The specific steps of generating multi-peak joint intensity distribution spectrum are as follows, Performing denoising, baseline correction and intensity normalization on serum static Raman spectrum, locating the wave number position of adenine characteristic peak, phenylalanine characteristic peak and phospholipid characteristic peak; Extracting intensity integral values of adenine characteristic peak, phenylalanine characteristic peak and phospholipid characteristic peak by peak area integration method to form multi-peak joint intensity distribution spectrum. 5.The method for constructing a colorectal cancer early diagnosis model based on Raman spectrum detection according to claim 4, characterized in that: The specific steps of generating patient entropy change index are as follows, Based on the health control group, a data layer is constructed, a modeling layer is constructed based on kernel density estimation, and a calculation layer is constructed based on metabolic disorder quantification mechanism, and a metabolic entropy change model is constructed according to the data layer, the modeling layer and the calculation layer; Inputting multi-peak joint intensity distribution spectrum into metabolic entropy change model to generate patient entropy change index through probability density mapping and negative logarithmic transformation. 6.The method for constructing a colorectal cancer early diagnosis model based on Raman spectrum detection according to claim 5, wherein: The specific steps of generating metabolic plasticity feature matrix are as follows, Performing time axis alignment on metabolic compensation rate curve to generate standardized metabolic compensation rate sequence; The patient entropy index is extended into a time-space sequence synchronized with the metabolic compensation rate sequence to generate a patient entropy index sequence; The metabolic compensation rate sequence and the patient entropy index sequence are fused to generate a joint feature sequence, and the joint feature sequence is converted into a two-dimensional matrix to output a metabolic plasticity feature matrix. 7.The method for constructing a colorectal cancer early diagnosis model based on Raman spectrum detection according to claim 6, wherein: The specific steps of constructing the dual-path neural network are as follows, Based on the dynamic fluctuation characteristics of the metabolic plasticity feature matrix, a time series analysis branch is constructed, based on the spatial coordination characteristics of the multi-peak joint intensity distribution spectrum, a spatial correlation branch is constructed, and an output layer is constructed according to an interface connection mechanism; The dual-path neural network is constructed according to the time series analysis branch, the spatial correlation branch and the output layer. 8.The method for constructing a colorectal cancer early diagnosis model based on Raman spectrum detection according to claim 7, wherein: The specific steps of generating the time series feature encoding and the spatial feature encoding are as follows, The metabolic plasticity matrix is input into the dual-path neural network, the metabolic response features are extracted through the time series analysis branch to generate the time series feature encoding, and the metabolic spatial correlation features are extracted through the spatial correlation branch to generate the spatial feature encoding. The specific steps of generating the three-dimensional metabolic stress map are as follows, 9.The method of constructing a model for early diagnosis of colorectal cancer based on Raman spectroscopy detection according to claim 8, wherein: The time series feature encoding and the spatial feature encoding are intelligently fused through an attention interaction method to generate a fusion feature; The glutamine metabolic pathway activity axis and the fatty acid metabolic pathway activity axis are defined, and a two-dimensional activity gradient plane is constructed; The metabolic stress intensity is calculated according to the fusion feature, the fusion feature is projected to the activity gradient plane through thermal color mapping, and the three-dimensional metabolic stress map is generated by superimposing the metabolic stress intensity. The specific steps of outputting the colorectal cancer early diagnosis model are as follows, 10.The method of constructing a model for early diagnosis of colorectal cancer based on Raman spectroscopy detection according to claim 9, wherein: The spatial grid sampling is performed on the three-dimensional metabolic stress map to extract the feature point coordinates on the activity gradient plane and the corresponding metabolic stress intensity values, and a metabolic pathway activity vector is generated; A residual expansion layer is added in front of the output layer of the dual-path neural network according to the metabolic pathway activity vector to construct a dual-path neural network with expanded structure; The dual-path neural network with expanded structure is executed to perform hierarchical sampling dataset division, cross-entropy loss optimization and diagnosis efficiency verification, and the colorectal cancer early diagnosis model is output.