Intestinal tissue auxiliary analysis system and method based on multi-modal data combination
Through a multimodal data-combined intestinal tissue-assisted analysis system, combined with OCT, Raman, white light and OCE information, the multimodal feature interaction and fusion of intestinal lesions assessment is achieved, solving the low accuracy problems caused by modal independent processing in the prior art, and improving the accuracy and robustness of the assessment.
Patent Information
- Application Number
- CN202510536949.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-04-27
AI Technical Summary
The existing intestinal lesion assessment methods independently process data of each modality, ignoring the physical correlation between different modalities, resulting in low accuracy and inability to accurately guide doctors to diagnose.
The intestinal tissue assisted analysis system based on multimodal data is adopted to collect OCT, Raman, white light and OCE information through a multimodal imaging system, and the feature extraction module and multimodal feature fusion classification module are used to realize the interactive analysis and fusion of OCE, OCT, Raman and white light features.
It improves the accuracy and robustness of intestinal lesions assessment, and through the complementary multimodal information, it enhances the ability to identify and evaluate intestinal lesions, providing more reliable diagnostic support.
Smart Images

Figure CN120078349A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intestinal tissue imaging, and particularly relates to an intestinal tissue assisted analysis system and method based on multi-modal data combination. Background Art
[0002] With the rapid development of social economy, people's diet and some living habits have changed. Irregular work and rest and unhealthy diet have led to an increasing incidence of digestive tract diseases, especially inflammatory bowel disease, and the trend of onset is gradually becoming younger. Intestinal fibrosis caused by chronic inflammation may further develop into intestinal sclerosis, stenosis and obstruction, and these complications cannot be treated by drugs, which means that patients need to undergo surgery for treatment. Therefore, detecting the type of intestinal lesions and evaluating their severity play a crucial role in early management of patients. Clinically, biopsy is an effective means to evaluate intestinal tissue lesions and their severity, but this method is invasive, may cause discomfort or complications to patients, and can only provide information on local tissues, making it difficult to comprehensively reflect the severity of diseases in the entire intestine; at the same time, existing intestinal lesion evaluations only independently extract features from each modal data, without information interaction between different modal features, ignoring the physical relevance between different modalities, resulting in low accuracy and the obtained evaluation results being unable to accurately guide doctors. Summary of the Invention
[0003] The purpose of the present invention is to provide an intestinal tissue assisted analysis system and method based on multi-modal data combination.
[0004] In a first aspect, the present invention provides an intestinal tissue assisted analysis system based on multi-modal data combination, including a multi-modal imaging system and an intestinal lesion evaluation system; the multi-modal imaging system includes an endoscope probe, a signal light acquisition module and an acoustic excitation module; the signal light acquisition module is used to collect OCT, Raman and white light information of intestinal tissue; the acoustic excitation module is used to obtain OCE information of intestinal tissue; the intestinal lesion evaluation system is used to evaluate lesions according to the information collected by the signal light acquisition module.
[0005] The described intestinal lesion assessment system includes a feature extraction module and a multi-modal feature fusion and classification module; the feature extraction module includes a first dual-modal analysis branch, a second dual-modal analysis branch, and a Raman spectroscopy analysis branch; the first dual-modal analysis branch is used to fuse OCE features and OCT features; the first dual-modal analysis branch realizes multi-perspective decoupling of three-dimensional features based on the self-attention mechanism method, obtains two-dimensional orthogonal perspective feature maps of OCE and OCT, and uses these two-dimensional orthogonal perspective feature maps to extract the query matrix, key matrix, and value matrix of OCE features and OCT features respectively, and realizes the interaction between OCE features and OCT features through the cross-attention mechanism, obtaining an OCE fusion spatial feature map focusing on OCE features and an OCT fusion spatial feature map focusing on OCT features from different perspectives;
[0006] The second dual-modal analysis branch is used to fuse white light features and OCE features; the Raman spectroscopy analysis branch is used to extract Raman spectroscopy features; the multi-modal feature fusion and classification module is used to fuse the feature vectors extracted by the first dual-modal analysis branch, the second dual-modal analysis branch, and the Raman spectroscopy analysis branch to obtain the final evaluation result.
[0007] Preferably, the second dual-modal analysis branch includes a cyclic cross-attention module for capturing global context information; the cyclic cross-attention module encodes the feature map input to the cyclic cross-attention module through three convolutional layers to obtain a query matrix, a key matrix, and a value matrix; the query matrix is dot-multiplied with the key matrix to obtain a first attention score; the value matrix is weighted and summed using the first attention score to obtain a first intermediate feature map; the first intermediate feature map is dot-multiplied with the query matrix to obtain a second attention score, and the value matrix is weighted and summed using the second attention score to obtain a second intermediate feature map, which is fused with the feature map input to the cyclic cross-attention module to obtain a fused feature map.
[0008] Preferably, the second dual-modal analysis branch further includes an encoder, an elastic attention guidance module, and a global average pooling layer; the original feature map of the white light image is extracted through the encoder; the elastic attention guidance module processes the OCE fusion spatial feature map through an activation function and multiplies the processing result element-wise with the original feature map to obtain an associated feature map; the associated feature map is processed sequentially through the cyclic cross-attention module and the global average pooling layer to obtain a white light feature vector.
[0009] Preferably, the first dual-modal analysis branch includes two 3D feature encoders, a multi-view cross-attention module, two residual convolution modules, and two global average pooling layers; the two feature encoders respectively extract 3D feature maps of OCT image data and OCE image data, and use the multi-view cross-attention module to interact with the two 3D feature maps to obtain an OCT fusion spatial feature map focusing on OCT features and an OCE fusion spatial feature map focusing on OCE features from different perspectives; the fusion spatial feature maps from different perspectives are further processed through the residual convolution module and the global average pooling layer in sequence to obtain OCT feature vectors and OCE feature vectors from different perspectives; the feature vectors from different perspectives are concatenated to obtain the OCE feature vector and OCT feature vector finally output by the first dual-modal analysis branch.
[0010] Preferably, the multi-view cross-attention module includes two multi-view decoupling modules and three cross-attention modules; the two multi-view decoupling modules respectively decompose the two input 3D feature maps into two-dimensional OCT initial spatial feature maps and two-dimensional OCE initial spatial feature maps in three orthogonal perspectives; the initial spatial feature maps from the same perspective are processed through the three cross-attention modules respectively, and each cross-attention module outputs an OCT fusion spatial feature map focusing on OCT features and an OCE fusion spatial feature map focusing on OCE features from the corresponding perspective.
[0011] Preferably, the Raman spectroscopy analysis branch includes a preprocessing module, a convolutional layer, a multi-scale fusion module, and a residual convolution module connected in sequence.
[0012] Preferably, the acoustic excitation module includes an ultrasonic transducer device; the ultrasonic transducer device includes a plurality of air-coupled ultrasonic transducer elements arranged in an annular array.
[0013] Preferably, the signal light acquisition module includes a common optical path and imaging optical paths for three modalities; the imaging optical paths for the three modalities are an OCT imaging optical path, a Raman imaging optical path, and a white light imaging optical path respectively;
[0014] The common optical path includes a first long-pass dichroic filter, an X-Y scanning galvanometer, a second long-pass dichroic filter, a first focusing lens, and a transparent plane mirror; the light beams input into the common optical path by the OCT imaging optical path and the Raman imaging optical path sequentially pass through the first long-pass dichroic filter, the X-Y scanning galvanometer, the second long-pass dichroic filter, the first focusing lens, and the transparent plane mirror, and the light beams are focused on the intestinal tissue, and the OCT information and Raman information are collected by receiving the reflected light beams.
[0015] The white light imaging optical path includes a white light source and a CMOS camera. The white light imaging is provided with illumination conditions by the white light source, and the signal light is collected by a common optical path. The signal light passes through a transparent plane mirror, a first focusing lens, and a second long-wave pass dichroic filter to enter the CMOS camera, thereby realizing the collection of white light information.
[0016] In a second aspect, the present invention provides an intestinal tissue auxiliary analysis method based on multimodal data combination, which uses the above-mentioned multimodal intestinal tissue detection system; the multimodal intestinal tissue detection method comprises the following steps:
[0017] Step 1: Use a multimodal imaging system to collect multimodal image data of intestinal tissues with different lesions and construct a labeled intestinal lesion dataset; the samples of the lesion dataset include OCE images, OCT images, Raman spectra and white light images of intestinal tissues as well as labels of different lesions and their severity;
[0018] Step 2: Construct an intestinal lesion assessment model; the intestinal lesion assessment model includes a feature extraction module and a multimodal feature fusion classification module; the feature extraction module extracts feature vectors of OCE images, OCT images, Raman spectra and white light images, and the multimodal feature fusion classification module fuses and classifies them to obtain the final assessment result;
[0019] Step 3: Use the intestinal lesion dataset to train the intestinal lesion assessment model;
[0020] Step 4: Input the OCE image, OCT image, Raman spectrum and white light image of the subject into the trained intestinal lesion assessment model to obtain the final assessment result.
[0021] Preferably, the method for collecting OCE images is as follows: use an ultrasonic transducer device to emit ultrasonic waves to the intestinal tissue to excite the intestinal tissue, and collect OCT information through a signal light acquisition module; measure the phase change of the spectrum after the interference of the reflected light inside the tissue and the reference light after ultrasonic excitation to calculate the deformation of the intestinal tissue; model the force on the intestinal tissue and construct a theoretical displacement function of the intestinal tissue; obtain the shear modulus by minimizing the difference between the deformation at the same position of the intestinal tissue and the theoretical displacement of the intestinal tissue, and then reconstruct the OCE image data of the intestinal tissue; the reference light is the light reflected from the collected OCT information.
[0022] The present invention has the following beneficial effects:
[0023] 1. The present invention uses the first dual-modal analysis branch to perform interactive analysis on OCT features and OCE features. By transforming the analysis of three-dimensional feature maps into the analysis of two-dimensional spatial feature maps, it reduces the computational complexity and the model size while improving the feature analysis efficiency. At the same time, covering the three-dimensional space from three orthogonal perspectives also avoids information loss in a single perspective. In addition, due to the different tissue densities of different layers of intestinal tissue, there are certain differences in the tissue elasticity of different layers. By establishing the correlation between the two modalities, the complementarity of structural and elastic information can be achieved.
[0024] 2. The present invention uses the second dual-modal analysis branch to fuse OCE features and white light features. Based on the fact that OCE can better reflect the characteristics of intestinal lesions, the OCE image is used to guide the model to focus on the lesion area in the white light image, thereby strengthening the model's ability to represent the pathological features of the mucosal layer and further improving the overall robustness of the model. At the same time, the present invention extracts features through the cyclic cross-attention module in the second dual-modal analysis branch. Compared with the existing cyclic cross-attention module, the present invention reduces the computational complexity while maintaining the integrity of feature expression.
[0025] 3. The present invention uses an annular multi-element ultrasonic transducer to achieve ultrasonic excitation of the sample. The ultrasonic excitation device emits ultrasonic waves at a certain angle to the sample, realizing precise and uniform excitation of the sample area by ultrasonic waves, and then enhancing the OCE imaging quality. At the same time, the redundant design of the annular array improves the fault tolerance of the system. When a single or multiple elements fail, the remaining elements can still maintain basic functions, which can not only avoid the inconvenience of temporarily replacing instruments but also significantly reduce the overall impact of the signal light acquisition subsystem on the OCE imaging quality.
[0026] 4. The present invention combines optical coherence elastography technology, Raman spectroscopy technology, and endoscopic imaging technology to achieve structural and functional imaging of lesions, and through the tissue elasticity analysis of OCE, the tissue structure analysis of OCT, the molecular property detection of Raman spectroscopy, and the visualization of the intestinal mucosal layer by white light endoscopy, it realizes multi-modal comprehensive detection of intestinal lesions. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 It is a schematic diagram of the multi-modal imaging system in Embodiment 1 of the present invention.
[0028] Figure 2 It is a schematic diagram of the signal light acquisition module in Embodiment 1 of the present invention.
[0029] Figure 3 It is a schematic diagram of the OCT spectrometer structure in Embodiment 1 of the present invention.
[0030] Figure 4 It is a schematic diagram of the Raman spectrometer structure in Embodiment 1 of the present invention.
[0031] Figure 5 Schematic diagram of the intestinal lesion evaluation system in Embodiment 1 of the present invention.
[0032] Figure 6 Schematic diagram of the multi-view cross-attention module structure in Embodiment 1 of the present invention.
[0033] Figure 7 3D feature map input into the multi-view cross-attention module in Embodiment 1 of the present invention.
[0034] Figure 8 Schematic diagram of the residual convolution module structure in Embodiment 1 of the present invention.
[0035] Figure 9 Schematic diagram of the traditional cyclic cross-attention module structure.
[0036] Figure 10 Schematic diagram of the cyclic cross-attention module structure in Embodiment 1 of the present invention.
[0037] Figure 11 Schematic diagram of the multi-scale fusion module structure in Embodiment 1 of the present invention.
[0038] Reference numerals: 1, first long-wave pass dichroic filter; 2, X-Y scanning galvanometer; 3, second long-wave pass dichroic filter; 4, first focusing lens; 5, mirror; 6, OCT light source; 7, sample arm collimating lens; 8, OCT fiber coupler; 9, reference arm collimating lens; 10, reference arm focusing lens; 11, reference arm mirror; 12, OCT spectrometer; 13, Raman light source; 14, first collimating lens; 15, Raman fiber coupler; 16, Raman spectrometer; 17, white light source; 18, CMOS camera; 19, second collimating lens; 20, first transmissive diffraction grating; 21, second focusing lens; 22, first line array CCD; 23, third collimating lens; 24, Rayleigh wave filter; 25, second transmissive diffraction grating; 26, third focusing lens; 27, second line array CCD; 28, camera fixing device; 29, transparent plane mirror; 30, ultrasonic transducer device. Detailed implementation manners
[0039] The present invention will be further described below with reference to the accompanying drawings.
[0040] Embodiment 1
[0041] As Figure 1As shown in the figure, an intestinal tissue assisted analysis system based on multi-modal data combination includes a multi-modal imaging system and an intestinal lesion evaluation system. The multi-modal imaging system includes an endoscope probe, a signal light acquisition module, and an acoustic excitation module installed at the end of the endoscope probe. The signal light acquisition module is used to collect the OCT (Optical Coherence Tomography), Raman, and white light information of the intestinal tissue sample. The signal light acquisition module includes a common optical path and imaging optical paths for three modalities. The common optical path is built into the endoscope probe; the imaging optical paths for the three modalities are the OCT imaging optical path, the Raman imaging optical path, and the white light imaging optical path. In the entire multi-modal imaging system, single-mode optical fibers are used to transmit light.
[0042] The common optical path is used to integrate the imaging optical paths for the three modalities, achieving coaxial unity of the optical paths, ensuring a high degree of spatial consistency of the imaging data in different modalities, and facilitating comprehensive analysis by doctors; at the same time, multi-modal information can be fused without a complex image registration algorithm. As Figure 1 and 2 shown in the figure, the common optical path includes a first long-pass dichroic filter 1, an X-Y scanning galvanometer 2, a second long-pass dichroic filter 3, a first focusing lens 4, and a transparent plane mirror 29. The first long-pass dichroic filter 1 and the second long-pass dichroic filter 3, through their beam splitting characteristics, transmit and reflect different beam wavelength bands, thereby realizing the integration of the imaging optical paths for the three modalities. Among them, the first long-pass dichroic filter 1 combines and splits the OCT light and the Raman light, the OCT beam is transmitted, and the Raman beam is reflected; the second long-pass dichroic filter 3 reflects and transmits the white light signal light and the signal after the combination of OCT and Raman respectively. The beams input into the common optical path of the OCT imaging optical path and the Raman imaging optical path sequentially pass through the first long-pass dichroic filter 1, the X-Y scanning galvanometer 2, the second long-pass dichroic filter 3, the first focusing lens 4, and the transparent plane mirror 29, and the beams are focused on the sample, and the OCT information and the Raman information are collected by receiving the reflected beams. The beam of the white light imaging optical path input into the common optical path passes through the second long-pass dichroic filter 3, the first focusing lens 4, and the transparent plane mirror 29, and the beam is focused on the sample, and the white light information is collected by receiving the reflected beam.
[0043] The OCT imaging optical path includes an OCT light source 6, a sample arm collimating lens 7, an OCT fiber coupler 8, a reference arm collimating lens 9, a reference arm focusing lens 10, a reference arm mirror 11, and an OCT spectrometer 12. The OCT imaging optical path and the common optical path together form an OCT imaging system. In the OCT imaging system, the OCT light source 6 transmits a laser beam through a single-mode fiber to the OCT fiber coupler 8. The OCT fiber coupler 8 divides the beam into two parts. One part of the beam is transmitted through the fiber to the reference arm composed of the reference arm collimating lens 9, the reference arm focusing lens 10, and the reference arm mirror 11. The other part of the beam is transmitted through the fiber to the sample arm composed of the collimating lens 7 and the common optical path. The beam entering the reference arm is collimated by the reference arm collimating lens 9 and then directed towards the reference arm focusing lens 10, which focuses the beam onto the reference arm mirror 11 and then reflects it back along the original path in the reverse direction. The beam entering the sample arm becomes a parallel beam after passing through the sample arm collimating lens 7 and is then directed towards the common optical path. After passing through the common optical path, the OCT beam is reflected back along the original path after irradiating the sample, obtaining an OCT beam carrying sample information. The OCT beam returned in the sample arm interferes with the light returned in the reference arm in the OCT fiber coupler 8 and is then transmitted through the fiber to the OCT spectrometer 12 to obtain OCT information. This OCT information can reflect the depth information of the intestinal tissue, obtain a cross-sectional image (depth image) of the tissue, and facilitate the subsequent identification and analysis of intestinal lesions.
[0044] As Figure 3 shown, the OCT spectrometer 12 includes a second collimating lens 19, a first transmissive diffraction grating 20, a second focusing lens 21, and a first linear array CCD 22. The OCT beam entering the OCT spectrometer 12 is collimated by the second collimating lens 19 and then dispersed by the diffraction grating 20. Finally, the beam is focused onto the first linear array CCD 22 by the second focusing lens 21, thereby realizing OCT imaging.
[0045] The Raman imaging optical path includes a mirror 5, a Raman light source 13, a first collimating lens 14, a Raman fiber coupler 15, and a Raman spectrometer 16. The Raman imaging optical path and the common optical path together form a Raman imaging system. In the Raman imaging system, the Raman light source 13 transmits a laser beam through a single-mode fiber to the Raman fiber coupler 15 and then through the fiber to the first collimating lens 14 to obtain a parallel beam. The parallel beam is reflected by the mirror 5 and enters the common optical path. After passing through the common optical path, the Raman beam is reflected back along the original path after irradiating the sample, obtaining a Raman beam carrying sample information, and the Raman beam is transmitted to the Raman spectrometer 16 through the Raman fiber coupler 15 to obtain spectral information. This spectral information can obtain the spectral peak differences of type I collagen in the intestine, reflect the stress changes in the intestine, and facilitate the subsequent identification and analysis of lesions such as intestinal fibrosis.
[0046] As Figure 4As shown, the Raman spectrometer 16 includes a third collimating lens 23, a Rayleigh wave filter 24, a second transmissive diffraction grating 25, a third focusing lens 26, and a second linear array CCD 27. The Raman beam incident on the Raman spectrometer 16 is collimated by the third collimating lens 23 and then the Rayleigh wave is filtered out by the Rayleigh wave filter 24. After that, it is dispersed by the second transmissive diffraction grating 25, and finally the beam is focused onto the second linear array CCD 27 by the third focusing lens 26, thereby realizing Raman spectral imaging.
[0047] The white light imaging optical path includes a white light source 17, a CMOS camera 18, and a camera fixing device 28. The white light imaging optical path and the common optical path together form a white light imaging system. In the white light imaging system, the white light source 17 irradiates the sample; the CMOS camera 18 fixed on the endoscopic probe through the camera fixing device 28 receives the white light information reflected by the sample irradiated by the white light source 17 based on the spectral splitting property of the second long-pass dichroic filter 3, and then collects the white light information to realize white light imaging in the intestine. This white light information can reflect the condition of the mucosal layer of the intestine, facilitating subsequent identification and analysis of intestinal lesions.
[0048] The acoustic excitation module includes an ultrasonic transducer device 30, which is used to deform the tissue to obtain its elastic information, so as to complete OCE (Optical Coherence Elastography) imaging according to the spectral information collected by the OCT system. The ultrasonic transducer device 30 includes 12 air-coupled ultrasonic transducer elements, and the effective area of each element is about 1 mm × 1 mm, arranged in a circular array. Each element is directionally installed at an inclination angle of 45°. This inclined design ensures that the ultrasonic waves can be accurately focused on the sample surface to achieve non-contact excitation. This structural design has two advantages: First, the symmetric distribution of the 12 air-coupled ultrasonic transducer elements can achieve uniform sound field coverage of the sample surface, ensuring the stability of excitation; Second, the redundant design of the circular array improves the fault tolerance of the system. When a single or multiple elements fail, the remaining elements can still maintain basic functions, which can not only avoid the inconvenience of temporarily replacing the instrument, but also significantly reduce the overall impact of the signal light acquisition subsystem on the OCE imaging quality.
[0049] In this embodiment, the cut-off wavelength of the first long-pass dichroic filter 1 in the common optical path is 990 nm, the transmission band is 1000 nm to 1400 nm, and the reflection band is 780 nm to 980 nm.
[0050] In this embodiment, the cut-off wavelength of the second long-pass dichroic filter 2 in the common optical path is 770 nm, the transmission band is 780 nm to 1400 nm, and the reflection band is 400 nm to 760 nm.
[0051] In this embodiment, the effective focal length of the first focusing lens 4 in the common optical path is 10 mm, and the lens diameter is 5 mm.
[0052] In this embodiment, the OCT light source 6 in the OCT imaging system is a superluminescent light-emitting diode with a central wavelength of 1300 nm, a power of 5 mW, and a full width at half maximum of 75 nm.
[0053] In this embodiment, the OCT fiber coupler 8 in the OCT imaging system is a single-mode 2×2 broadband fiber coupler with a central wavelength of 1300±75 nm and a coupling ratio of 50:50.
[0054] In this embodiment, the working wavelength of the single-mode fiber used in the OCT imaging system is 1200 nm to 1390 nm, and the numerical aperture NA is 0.12.
[0055] In this embodiment, the effective focal lengths of the sample arm collimating lens 7 and the reference arm collimating lens 9 in the OCT imaging system are both 10 mm, and the diameters are 3 mm.
[0056] In this embodiment, the effective focal length of the second collimating lens 19 in the OCT spectrometer 12 is 10 mm, and the diameter is 5 mm.
[0057] In this embodiment, the first transmissive diffraction grating 20 in the OCT spectrometer 12 has 1145 lines / mm and an incident angle of 48.6°.
[0058] In this embodiment, the effective focal length of the second focusing lens 21 in the OCT spectrometer 12 is 25 mm, and the diameter is 30 mm.
[0059] In this embodiment, the first linear array CCD 22 in the OCT spectrometer 12 has 2048 pixels, and the pixel size is 10 μm×10 μm.
[0060] In this embodiment, the spot size of the light beam collimated by the sample arm collimating lens 7 and the reference arm collimating lens 9 of the OCT imaging system is approximately 2.4 mm.
[0061] In this embodiment, the lateral resolution ∆x of the OCT imaging system is approximately 6.90 μm, and the axial resolution ∆z is approximately 4.96 μm. It can be used to obtain high-resolution depth and three-dimensional images of the intestine. The OCT images can visually observe information such as the cross-sectional structure, thickness, lesion depth, and fibrosis tissue distribution of the intestine, which helps to analyze and evaluate the lesion range and invasion depth. In addition, by using the OCT system to collect the phase information of the OCT interference spectrum under the excitation state, the biomechanical parameters of the tissue, such as the elastic properties of the tissue, are obtained, providing important information for the identification and evaluation of intestinal lesions.
[0062] In this embodiment, the Raman light source 13 in the Raman imaging system is a semiconductor laser, with a central wavelength of 785 nm, an effective power of 20 mW, and a spectral range for collecting type I collagen of 350 to 2100 of the spectral range.
[0063] In this embodiment, the Raman fiber coupler 15 in the Raman imaging system is a single-mode 1×2 broadband fiber coupler, with a central wavelength of 865 ± 85 nm and a coupling ratio of 50:50.
[0064] In this embodiment, the single-mode optical fiber used in the Raman imaging system has a working wavelength of 780 - 950 nm and a numerical aperture NA of 0.12.
[0065] In this embodiment, the effective focal length of the first collimating lens 14 in the Raman imaging system is 10 mm and the diameter is 3 mm.
[0066] In this embodiment, the spot size of the light collimated by the first collimating lens 14 in the Raman imaging system is approximately 2.4 mm.
[0067] In this embodiment, the effective focal length of the third collimating lens 23 in the Raman spectrometer 16 is 10 mm and the diameter is 5 mm.
[0068] In this embodiment, the second transmissive diffraction grating 25 in the Raman spectrometer 16 has 1800 lines / mm and an incident angle of 42.8°.
[0069] In this embodiment, the effective focal length of the third focusing lens 26 in the Raman spectrometer 16 is 45 mm and the diameter is 44 mm.
[0070] In this embodiment, the second linear array CCD 27 in the Raman spectrometer 16 has 2048×264 pixels and a pixel size of 15 μm×15 μm.
[0071] In this embodiment, the Raman scattering wavelength of the Raman imaging system is in the range of 807 - 944 nm, that is, the spectral range is 350 - 2100 , and the spectral resolution is about 2 , which helps doctors analyze the degree of intestinal fibrosis and study the molecular mechanism of fibrosis.
[0072] In this embodiment, the swing angle of the Y galvanometer in the X - Y scanning galvanometer 2 is ±2.86°, and the swing angle of the X galvanometer is ±2.04°, achieving a scanning range of 2 mm×2 mm for the sample.
[0073] In this embodiment, the white light source 17 in the white light imaging system is an LED white light source, with a wavelength range of 400 nm - 760 nm.
[0074] In this embodiment, the CMOS camera 18 in the white light imaging system has 2048×2048 pixels, and the pixel size is 1.4μm×1.4μm.
[0075] In this embodiment, the optimal working distance of the endoscope is 3 mm.
[0076] As Figure 5 shown, in order to achieve multi-modal precise analysis, the intestinal lesion evaluation system proposes corresponding feature extraction modules and multi-modal feature fusion classification modules according to the characteristics of different modalities. The feature extraction module includes a first dual-modal analysis branch, a second dual-modal analysis branch, and a Raman spectroscopy analysis branch.
[0077] Since the two modalities of OCE and OCT are similar and have a certain physical correlation, the first dual-modal analysis branch is used to perform interactive analysis on the features of these two modalities. The first dual-modal analysis branch is used to analyze and fuse OCE and OCT features. The first dual-modal analysis branch includes two feature encoders, a multi-view cross-attention module, two residual convolution modules, and two global average pooling layers. The two 3D feature encoders respectively extract the 3D feature maps of OCT image data and OCE image data, and use the multi-view cross-attention module to perform feature fusion on the 3D feature maps extracted by the two 3D feature encoders, obtaining an OCT fusion spatial feature map focusing on OCT features and an OCE fusion spatial feature map focusing on OCE features from different perspectives, realizing the elastic and structural feature interaction of OCE image data and OCT image data at the spatial level; further performing high-level deep feature processing on the fusion spatial feature maps from different perspectives through the residual convolution module, and performing pooling processing through the global average pooling layer to obtain OCT feature vectors and OCE feature vectors from different perspectives; splicing the feature vectors from different perspectives to obtain the OCE feature vectors and OCT feature vectors finally output by the first dual-modal analysis branch to support subsequent classification decisions.
[0078] In this embodiment, the feature encoder adopts an architecture of the ResNet series or an architecture of the Transformer series.
[0079] As Figure 6 shown, the multi-view cross-attention module includes two multi-view decoupling modules and three cross-attention modules. The two multi-view decoupling modules respectively perform axial analysis on the two 3D feature maps based on the self-attention method, decomposing each 3D feature map into two-dimensional OCT initial spatial feature maps and two-dimensional OCE initial spatial feature maps in three orthogonal perspectives of H×W, W×D, and H×D, thereby realizing the analysis of the 3D feature map into the analysis of the two-dimensional spatial feature map, reducing the calculation amount and the model volume while improving the feature analysis efficiency. In addition, covering the three-dimensional space from three orthogonal perspectives also avoids the information loss of a single perspective. AsFigure 7 As shown, it represents two examples of three-dimensional feature maps in the input multi-view decoupling module. Taking the generation process of the H×W two-dimensional spatial feature map as an example, the three-dimensional feature map is flattened into (H×W)×D, and three projection matrices, namely the query matrix, the key matrix, and the value matrix, with a size of (H×W)×D, are obtained through encoding using a 1×1 convolutional layer. Among them, the query matrix is multiplied by the key matrix and the attention weights are calculated through the Softmax function to obtain a D×D attention weight map, which represents the association between the depth directions of the query matrix and the key matrix. The D×D attention weight map is multiplied by the Value and averaged in the depth direction to obtain an H×W two-dimensional initial spatial feature map in the D-axis direction view. Through this method, the dynamic allocation of weights of three-dimensional features in the axial direction can be achieved.
[0080] The cross-attention module uses the cross-attention strategy to establish the association of bimodal features, and then uses the features of one modality to assist in locating the key regions of the other modality. Since the tissue density of different layers of intestinal tissue is different, and the tissue elasticity of different layers itself also has certain differences, by establishing the association of the two modalities, the complementarity of structural and elastic information can be achieved. The cross-attention module encodes the initial spatial feature maps of the same view in the bimodal data into three projection matrices, namely the query, the key, and the value, respectively, using 3 1×1 convolutional layers. The query matrix of one modality is multiplied by the key matrix of the other modality, processed through the Softmax function, and then multiplied by the value matrix of the other modality. Each cross-attention module outputs an OCT fusion spatial feature map that focuses on the OCT features and an OCE fusion spatial feature map that focuses on the OCE features of the corresponding view, thereby establishing the association between one modality and the other modality.
[0081] As Figure 8 shown, the residual convolutional module includes multiple groups of cascaded convolutional layers with residual structures, batch normalization, and ReLU activation functions, which are used to deepen the model hierarchy and further extract high-level features; the convolutional layer in the first bimodal analysis branch uses a two-dimensional convolutional kernel with a size of 3×3.
[0082] Since both the white light image and the OCE image contain mucosal layer information, the second bimodal analysis branch is used to guide the model to focus on the lesion area of the white light image through the mechanical characteristics reflected by the OCE image, thereby strengthening the model's ability to represent the pathological characteristics of the mucosal layer.
[0083] The second dual-modal analysis branch includes a two-dimensional feature encoder, an elastic attention guidance module, a cyclic cross-attention module, and a global average pooling layer connected in sequence. The two-dimensional feature encoder performs multi-scale feature extraction on the white light image to obtain the original feature map, providing a basic representation for subsequent multi-modal interaction. The elastic attention guidance module fuses the original feature map and the OCE fusion spatial feature map in the D-axis direction view (H×W) to obtain an associated feature map, strengthening the model's attention to the lesion area of the white light image. Based on the associated feature map, the cyclic cross-attention module is further used to extract the global features of the white light image to obtain a fused feature map; finally, the fused feature map is processed by the global average pooling layer to obtain the feature vector of the white light image to support subsequent classification decisions.
[0084] In this embodiment, the two-dimensional feature encoder uses ResNet or Transformer as the basic feature extraction network.
[0085] The elastic attention guidance module performs a non-linear mapping on the OCE fusion spatial feature map in the D-axis direction view (H×W) through the Sigmoid activation function to generate an attention weight distribution map along the depth direction (D-axis). Since the regions with higher elastic feature values in the OCE feature map correspond to the parts with greater tissue hardness, the Sigmoid activation function will assign higher weight values to these regions, thereby highlighting the response of the lesion area. Subsequently, multiplying the attention weight distribution map of the OCE modality element-wise with the original feature map can guide the model to pay more attention to the lesion area in the white light image and dynamically enhance the feature representation related to the lesion.
[0086] As Figure 9 and Figure 10 shown, compared with the existing cyclic cross-attention module, the cyclic cross-attention module in the present invention reduces the repeated calculations and simultaneously realizes the global interaction of the OCT, OCE, and white light three-modal features in the spatial features; while reducing the computational complexity, it maintains the integrity of the feature expression.
[0087] The cyclic cross-attention module encodes the associated feature map through three convolutional layers with a kernel size of 1×1 to obtain three parts: a query matrix, a key matrix, and a value matrix. Calculate the first attention score between the query matrix and the key matrix in the horizontal and vertical directions; use the first attention score to perform weighted summation on the value matrix to obtain the first intermediate feature map, realizing the interaction of each pixel on the feature map with the pixel information in its corresponding horizontal and vertical directions. If you want to achieve the interaction with all pixel information, calculate the second attention score between the first intermediate feature map and the query matrix in the horizontal and vertical directions, and perform weighted summation on the value matrix to obtain the second intermediate feature map; fuse the second intermediate feature map with the associated feature map to obtain the fused feature map.
[0088] Since the molecular characteristics reflected by Raman spectroscopy data are quite different from those of other modalities, it is taken as an independent branch in the system. The feature vectors of Raman spectroscopy data are extracted separately by the Raman spectroscopy analysis branch to analyze the spectral data of type I collagen, providing auxiliary support for intestinal lesion identification.
[0089] The Raman spectroscopy analysis branch includes a preprocessing module, a 1×3 one-dimensional convolutional layer, a multi-scale fusion module (Multi-scale Fusion module), and a residual convolutional module connected in sequence. The preprocessing module uses wavelet transform to reduce noise and improve the signal-to-noise ratio of the spectrum; at the same time, methods such as random spectral translation, random spectral cropping, and random Gaussian noise are used to enhance the input data; the convolutional layer in the Raman analysis branch uses a one-dimensional convolutional kernel of size 1×3.
[0090] As Figure 11 shown, to achieve multi-scale feature extraction, the multi-scale fusion module uses one-dimensional convolutional kernels of different sizes (including 1×12, 1×6, and 1×3) to extract features from the shallow local feature maps output by the one-dimensional convolution. Among them, the 1×12 convolutional kernel is suitable for extracting wide peak features, the 1×6 convolutional kernel is suitable for extracting medium peak width features, and the 1×3 convolutional kernel focuses on capturing narrow peak features; for example, features such as peak value, peak position, and peak width in regions such as the amide I band or amide III band are extracted to enhance the richness of spectral features. The extracted multi-scale features are concatenated and then passed through a 1×1 convolutional layer for feature fusion. The residual convolutional module includes multiple groups of convolutional layers, BN, and ReLU activation functions in series with residual structures, which are used to deepen the model hierarchy, further extract high-level features, obtain Raman spectroscopy feature vectors, and provide support for subsequent analysis tasks. The convolutional layer in the residual convolutional module uses a convolutional kernel of size 1×3.
[0091] The multi-modal feature fusion classification module realizes the deep fusion and classification decision of multi-modal information by integrating the feature vectors extracted by the first bimodal analysis branch, the second bimodal analysis branch, and the Raman spectroscopy analysis branch. First, the feature vectors generated by the above modules are concatenated in the channel dimension to form a unified joint feature representation. Subsequently, the concatenated features are input into the fully connected layer, and high-level discriminant features are extracted through non-linear transformation, and finally the classification result is output to obtain the final evaluation result. This module can make full use of the complementary information of different modalities to improve the classification performance.
[0092] Example 2
[0093] A method for auxiliary analysis of intestinal tissue based on the joint of multi-modal data, using the multi-modal intestinal tissue detection system in Example 1; this multi-modal intestinal tissue detection method includes the following steps:
[0094] Step 1. Image acquisition
[0095] Multimodal image data of different diseased intestinal tissues are collected using a multimodal imaging system to construct a labeled intestinal disease dataset; the samples of the disease dataset include OCE images, OCT images, Raman spectra, white light images of intestinal tissues, as well as labels for normal, different diseases and their severities. The diseases mainly include ulcers, erosions, fibrosis, etc., and among them, the fibrosis lesions are divided into mild, moderate, and severe degrees, and their severity is based on the degree and scope of fibrous tissue hyperplasia. OCT image data of intestinal tissues are obtained using an OCT imaging system, and OCE image data of intestinal tissues are obtained based on the OCT imaging system and an acoustic excitation module; Raman spectral image data of type I collagen molecules in intestinal tissues are obtained using a Raman imaging system; white light image data of the intestinal mucosal layer in intestinal tissues are obtained using a white light imaging system. The image data of the four modalities respectively reflect the characteristics of intestinal tissues from different levels: OCE images quantify the degree of tissue sclerosis through elastic modulus, providing a mechanical basis for intestinal disease identification; OCT images reveal the deep structural characteristics of intestinal tissues and can reflect some structural changes of lesions such as fibrosis; Raman spectra provide a molecular-level basis for intestinal disease identification by analyzing the molecular vibration characteristics of type I collagen; white light images directly reflect the information on the mucosal surface layer of intestinal tissues and can be used to identify the morphological and color characteristics of lesions such as tissue surface fibrosis. Based on the complementarity of the above four-modal data, high-accuracy automated assisted analysis can be achieved by fusing multimodal features, providing support for doctors' clinical decisions. The specific process of obtaining OCE image data of intestinal tissues based on the OCT imaging system and the acoustic excitation module is as follows:
[0096] An ultrasonic transducer device 30 is used to emit ultrasonic waves to the intestinal tissue to achieve excitation of the intestinal tissue. At the same time, the signal light acquisition module will collect OCT information. Based on the OCT interference principle, by measuring the phase change of the reflected light inside the tissue after ultrasonic excitation , and according to calculate the deformation amount of the intestinal tissue . Among them, represents the optical refractive index of the intestinal tissue; represents the average wavelength of the light source output by OCE. Based on the motion equation , model the force on the intestinal tissue; where m is the equivalent mass; c is the viscosity coefficient; is the equivalent spring stiffness; is the deformation amount. During the modeling process, the intestinal tissue is regarded as a multi-layer viscoelastic medium, and the following assumptions are made: the structures of each layer of the medium are uniform and incompressible; the ultrasonic wave is an axisymmetric radiation force acting on the upper surface of the medium; the mechanical parameters (Young's elastic modulus and shear viscosity) and density inside the medium are constants. Based on the above assumptions, the theoretical displacement function of the intestinal tissue can be expressed as Yx, y, z =- ∫ 0 ∞ α 2 J (α x 2 + y 2 )[ A 1 e -αz - A 2 e αz +α( B 1 e -βz + B 2 e βz )]dα ; wherein, represents the independent variable from 0 to integral; is the Bessel function of order zero; , , are respectively the lateral distance and height of a certain point of the tissue in the rectangular coordinate system; , , , are obtained from the set boundary conditions; ; ; is the Young's elastic modulus; is the tissue density; is the angular frequency; is the shear viscosity. By minimizing the difference between the deformation amount and the theoretical displacement amount at the same position of the intestinal tissue, the shear viscosity and the Young's elastic modulus can be obtained, and the shear modulus is obtained according to , and then multiple mechanical properties at the same position in the tissue are obtained, and finally the OCE elastic map (OCE image data) of the intestinal tissue is reconstructed.
[0097] Step 2. Construct an intestinal lesion evaluation model; the intestinal lesion evaluation model includes a feature extraction module and a multi-modal feature fusion and classification module; the feature extraction module extracts the feature vectors of the OCE image, OCT image, Raman spectrogram and white light image, and the multi-modal feature fusion and classification module performs fusion and classification to obtain the final evaluation result;
[0098] Step 3. Use the intestinal lesion data set to train the intestinal lesion evaluation model.
[0099] Step 4: Input the OCE image, OCT image, Raman spectrogram, and white light image of the subject into the trained intestinal lesion assessment model, and output the final assessment result for the auxiliary analysis of intestinal tissue.
Claims
1. An intestinal tissue auxiliary analysis system based on multimodal data combination, including a multimodal imaging system and an intestinal lesion assessment system; the multimodal imaging system includes an endoscope probe and a signal light acquisition module; The signal light acquisition module is used to realize the OCT, Raman and white light information acquisition of intestinal tissue; the intestinal lesion evaluation system is used to evaluate the lesion according to the information acquired by the signal light acquisition module; it is characterized in that it also includes an acoustic excitation module installed at the end of the endoscope probe, which is used to obtain the OCE information of the intestinal tissue; The intestinal lesion assessment system includes a feature extraction module and a multimodal feature fusion classification module; the feature extraction module includes a first bimodal analysis branch, a second bimodal analysis branch and a Raman spectroscopy analysis branch; the first bimodal analysis branch is used to fuse OCE features and OCT features; the first bimodal analysis branch extracts the query matrix, key matrix and value matrix of OCT features and OCE features respectively, and realizes the interaction between OCT features and OCE features through a cross-attention mechanism, so as to obtain OCT fusion spatial feature maps focusing on OCT features and OCE fusion spatial feature maps focusing on OCE features under different perspectives; The second bimodal analysis branch is used to fuse white light features and OCE features; the Raman spectrum analysis branch is used to extract Raman spectrum features; the multimodal feature fusion classification module is used to fuse the feature vectors extracted by the first bimodal analysis branch, the second bimodal analysis branch, and the Raman spectrum analysis branch to obtain the final evaluation result.
2. The intestinal tissue auxiliary analysis system based on multimodal data combination according to claim 1, characterized in that: The second bimodal analysis branch includes a recurrent cross-attention module for capturing global contextual information; The cyclic cross attention module encodes the feature map input to the cyclic cross attention module through three convolutional layers to obtain a query matrix, a key matrix, and a value matrix; a dot product is performed between the query matrix and the key matrix to obtain a first attention score; Use the first attention score to perform weighted summation on the value matrix to obtain the first intermediate feature map; The first intermediate feature map is dot-producted with the query matrix to obtain the second attention score, and the value matrix is weighted summed using the second attention score to obtain the second intermediate feature map, which is fused with the feature map of the input cycle cross attention module to obtain a fused feature map.
3. The intestinal tissue auxiliary analysis system based on multimodal data combination according to claim 2 is characterized by: The second bimodal analysis branch also includes an encoder, an elastic attention guidance module and a global average pooling layer; the encoder is used to extract the original feature map of the white light image; the elastic attention guidance module processes the OCE fusion space feature map through an activation function, and multiplies the processing result by the original feature map element by element to obtain a related feature map; the related feature map is processed in turn through a cyclic cross attention module and a global average pooling layer to obtain a white light feature vector.
4. The intestinal tissue auxiliary analysis system based on multimodal data combination according to claim 1, characterized in that: The first bimodal analysis branch includes two feature encoders, a multi-view cross-attention module, two residual convolution modules and two global average pooling layers; the two feature encoders extract three-dimensional feature maps of OCT image data and OCE image data respectively, and use the multi-view cross-attention module to interact with the two three-dimensional feature maps to obtain OCT fusion spatial feature maps focusing on OCT features and OCE fusion spatial feature maps focusing on OCE features under different viewpoints; the fusion spatial feature maps of different viewpoints are further processed through residual convolution modules and global average pooling layers in turn to obtain OCT feature vectors and OCE feature vectors of different viewpoints; the feature vectors of different viewpoints are spliced to obtain the OCE feature vector and OCT feature vector finally output by the first bimodal analysis branch.
5. The intestinal tissue auxiliary analysis system based on multimodal data combination according to claim 4 is characterized in that: The multi-view cross-attention module includes two multi-view decoupling modules and three cross-attention modules; the two multi-view decoupling modules respectively decompose the two input three-dimensional feature maps into two-dimensional OCT initial spatial feature maps and two-dimensional OCE initial spatial feature maps of three orthogonal viewpoints; the initial spatial feature maps of the same viewpoint are processed respectively by the three cross-attention modules, and each cross-attention module outputs an OCT fused spatial feature map focusing on OCT features and an OCE fused spatial feature map focusing on OCE features of the corresponding viewpoint.
6. The intestinal tissue auxiliary analysis system based on multimodal data combination according to claim 1, characterized in that: The acoustic excitation module comprises an ultrasonic transducer device (30); the ultrasonic transducer device (30) comprises a plurality of air-coupled ultrasonic transducer array elements arranged in a ring array.
7. The intestinal tissue auxiliary analysis system based on multimodal data combination according to claim 1, characterized in that: The signal light acquisition module includes a common optical path and three-modal imaging optical paths; the three-modal imaging optical paths are OCT imaging optical path, Raman imaging optical path and white light imaging optical path; the common optical path is built into the endoscope probe; The common optical path comprises a first long-wave-pass dichroic filter (1), an XY scanning galvanometer (2), a second long-wave-pass dichroic filter (3), a first focusing lens (4) and a transparent plane mirror (29); the light beams input into the common optical path by the OCT imaging optical path and the Raman imaging optical path sequentially pass through the first long-wave-pass dichroic filter (1), the XY scanning galvanometer (2), the second long-wave-pass dichroic filter (3), the first focusing lens (4) and the transparent plane mirror (29), so as to focus the light beams on the intestinal tissue and realize the collection of OCT information and Raman information by receiving the light beams reflected therefrom; The white light imaging optical path comprises a white light source (17) and a CMOS camera (18); the white light imaging is provided with illumination conditions by the white light source (17), and the signal light is collected by using a common optical path; the signal light passes through a transparent plane mirror (29), a first focusing lens (4), and a second long-wave pass dichroic filter (3) to enter the CMOS camera (18), thereby realizing the collection of white light information.
8. The intestinal tissue auxiliary analysis system based on multimodal data combination according to claim 1, characterized in that: The Raman spectrum analysis branch includes a preprocessing module, a convolution layer, a multi-scale fusion module and a residual convolution module which are connected in sequence.
9. A method for assisting intestinal tissue analysis based on multimodal data combination, characterized in that: Using an intestinal tissue auxiliary analysis system based on multimodal data combination as described in claim 1; the multimodal intestinal tissue detection method comprises the following steps: collecting multimodal image data of intestinal tissues with different lesions through a multimodal imaging system to construct a labeled intestinal lesion data set; constructing an intestinal lesion assessment model and using the intestinal lesion data set for training; inputting the OCE image, OCT image, Raman spectrum and white light image of the subject into the trained intestinal lesion assessment model, and outputting the final assessment result; The intestinal lesion assessment model includes a feature extraction module and a multimodal feature fusion classification module; the feature extraction module extracts feature vectors of OCE images, OCT images, Raman spectra and white light images, and the multimodal feature fusion classification module fuses and classifies them.
10. The intestinal tissue auxiliary analysis method based on multimodal data combination according to claim 9, characterized in that: The method for collecting OCE images is as follows: using an ultrasonic transducer (30) to transmit ultrasonic waves to intestinal tissue to excite the intestinal tissue, and collecting OCT information through a signal light collection module; measuring the phase change of the interference spectrum between the reflected light inside the tissue and the reference light after ultrasonic excitation to calculate the deformation of the intestinal tissue; modeling the force on the intestinal tissue and constructing a theoretical displacement function of the intestinal tissue; obtaining the shear modulus by minimizing the difference between the deformation at the same position of the intestinal tissue and the theoretical displacement of the intestinal tissue, and then reconstructing the OCE image data of the intestinal tissue.
Citation Information
Patent Citations
Multi-mode hysteroscope system and obtaining method thereof
CN104586344A
Endoscopy Raman spectrum detection system used for cancer early screening
CN113116302A
Multimode imaging device
CN113812929A
Multi-modal imaging system
CN116530935A
Multi-modal imaging and detection system for oral cavity and tongue
CN118697294A
Cited By
Oral and maxillofacial surgery image recognition and diagnosis method and system based on deep learning
CN120280132A