A system and method for auxiliary analysis of intestinal tissue based on multimodal data combination

Through the intestinal tissue-assisted analysis system combined with multimodal data, combined with OCT, OCE, Raman and white light imaging technology, the problems of low accuracy and high invasiveness of intestinal lesions in the prior art are solved, and efficient and accurate intestinal lesions evaluation and detection are achieved.

CN120078349BActive Publication Date: 2025-08-29HANGZHOU DIANZI UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510536949.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-08-29
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

The existing intestinal lesions assessment methods rely on the extraction of independent modal data features, neglecting the information interaction between different modalities, resulting in low evaluation accuracy, unable to fully reflect the severity of intestinal disease, and invasive examinations may bring discomfort and complications.

Method used

The intestinal tissue assisted analysis system combined with multimodal data is adopted, combined with OCT, OCE, Raman and white light imaging technology, and the fusion analysis of different modal features is achieved through the self-attention mechanism and the cross-attention module, including the feature extraction module and the multimodal feature fusion classification module, and the ring multi-array element ultrasonic transducer is used for precise excitation to construct an intestinal lesion evaluation model.

Benefits of technology

It improves the accuracy and robustness of intestinal lesions assessment, reduces the calculation volume and model volume, enhances the quality of OCE imaging, realizes multimodal comprehensive detection of intestinal lesions, and reduces the risk of invasive examination.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120078349B_ABST
    Figure CN120078349B_ABST
Patent Text Reader

Abstract

The present invention discloses an auxiliary analysis system and method for intestinal tissue based on the combination of multimodal data. The multimodal intestinal tissue detection system includes a multimodal imaging system and an intestinal lesion assessment system; the multimodal imaging system organically combines OCE imaging, OCT imaging, white light imaging and Raman imaging systems. Among them, OCE imaging reflects the elastic properties of intestinal lesion tissue, OCT imaging reflects the deep structural characteristics of intestinal lesion tissue, white light imaging reflects the mucosal surface manifestations of intestinal tissue, and Raman imaging reflects the molecular mechanism characteristics of intestinal lesion tissue. Utilizing the different characteristics reflected by the four modalities, the intestinal lesion assessment system extracts OCE feature vectors, OCT feature vectors, white light feature vectors and Raman spectrum feature vectors respectively, and uses a multimodal feature fusion classification module for processing, thereby realizing deep fusion and classification decision-making of multimodal information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intestinal tissue imaging, and specifically relates to an intestinal tissue auxiliary analysis system and method based on multimodal data combination. Background Art

[0002] With rapid socioeconomic development, people's diets and lifestyles have changed. Irregular sleep schedules and unhealthy diets have led to an increasing incidence of gastrointestinal diseases, particularly inflammatory bowel disease, and the incidence is gradually increasing at a younger age. Intestinal fibrosis caused by chronic inflammation can further develop into intestinal sclerosis, stenosis, and obstruction. These complications cannot be treated with medication, requiring surgery. Therefore, detecting the type of intestinal lesions and assessing their severity are crucial for early patient management. Clinically, biopsy is an effective means of assessing intestinal tissue lesions and their severity, but this method is invasive and can cause patient discomfort or complications. It only provides localized information and fails to fully reflect the severity of disease throughout the entire intestine. Furthermore, existing intestinal lesion assessments rely solely on feature extraction from each modality, without leveraging the information interaction between features from different modalities. This approach ignores the physical connections between modalities, resulting in low accuracy and inadequate guidance for physicians. Summary of the Invention

[0003] The purpose of the present invention is to provide an intestinal tissue auxiliary analysis system and method based on multimodal data combination.

[0004] In a first aspect, the present invention provides an intestinal tissue auxiliary analysis system based on multimodal data combination, comprising a multimodal imaging system and an intestinal lesion assessment system; the multimodal imaging system comprises an endoscopic probe, a signal light acquisition module, and an acoustic excitation module; the signal light acquisition module is used to realize the acquisition of OCT, Raman, and white light information of intestinal tissue; the acoustic excitation module is used to obtain OCE information of intestinal tissue; and the intestinal lesion assessment system is used to perform lesion assessment based on the information collected by the signal light acquisition module;

[0005] The intestinal lesion assessment system includes a feature extraction module and a multimodal feature fusion classification module; the feature extraction module includes a first bimodal analysis branch, a second bimodal analysis branch, and a Raman spectroscopy analysis branch; the first bimodal analysis branch is used to fuse OCE features and OCT features; the first bimodal analysis branch implements multi-perspective decoupling of three-dimensional features based on a self-attention mechanism method, obtains two-dimensional orthogonal perspective feature maps of OCE and OCT, and uses these two-dimensional orthogonal perspective feature maps to extract query matrices, key matrices, and value matrices of OCE features and OCT features, respectively, and implements interaction between OCE features and OCT features through a cross-attention mechanism, thereby obtaining OCE fusion spatial feature maps focusing on OCE features and OCT fusion spatial feature maps focusing on OCT features under different perspectives;

[0006] The second bimodal analysis branch is used to fuse white light features and OCE features; the Raman spectrum analysis branch is used to extract Raman spectrum features; the multimodal feature fusion classification module is used to fuse the feature vectors extracted by the first bimodal analysis branch, the second bimodal analysis branch, and the Raman spectrum analysis branch to obtain the final evaluation result.

[0007] Preferably, the second bimodal analysis branch includes a cyclic cross-attention module for capturing global context information; the cyclic cross-attention module encodes the feature map of the input cyclic cross-attention module through three convolutional layers to obtain a query matrix, a key matrix, and a value matrix; the query matrix is ​​dot-producted with the key matrix to obtain a first attention score; the value matrix is ​​weightedly summed using the first attention score to obtain a first intermediate feature map; the first intermediate feature map is dot-producted with the query matrix to obtain a second attention score, the value matrix is ​​weightedly summed using the second attention score to obtain a second intermediate feature map, and the second intermediate feature map is fused with the feature map of the input cyclic cross-attention module to obtain a fused feature map.

[0008] Preferably, the second bimodal analysis branch also includes an encoder, an elastic attention guidance module and a global average pooling layer; the original feature map of the white light image is extracted by the encoder; the elastic attention guidance module processes the OCE fusion spatial feature map through an activation function, and multiplies the processing result with the original feature map element by element to obtain a correlation feature map; the correlation feature map is processed in turn through a cyclic cross attention module and a global average pooling layer to obtain a white light feature vector.

[0009] Preferably, the first bimodal analysis branch includes two three-dimensional feature encoders, a multi-view cross-attention module, two residual convolution modules and two global average pooling layers; the two feature encoders respectively extract three-dimensional feature maps of OCT image data and OCE image data, and use the multi-view cross-attention module to interact with the two three-dimensional feature maps to obtain OCT fusion spatial feature maps focusing on OCT features and OCE fusion spatial feature maps focusing on OCE features under different viewpoints; the fusion spatial feature maps of different viewpoints are further processed by the residual convolution module and the global average pooling layer in turn to obtain OCT feature vectors and OCE feature vectors of different viewpoints; the feature vectors of different viewpoints are spliced ​​to obtain the OCE feature vector and OCT feature vector finally output by the first bimodal analysis branch.

[0010] Preferably, the multi-perspective cross-attention module includes two multi-perspective decoupling modules and three cross-attention modules; the two multi-perspective decoupling modules respectively decompose the two input three-dimensional feature maps into two-dimensional OCT initial spatial feature maps and two-dimensional OCE initial spatial feature maps of three orthogonal perspectives; the initial spatial feature maps of the same perspective are processed respectively by the three cross-attention modules, and each cross-attention module outputs an OCT fusion spatial feature map focusing on OCT features and an OCE fusion spatial feature map focusing on OCE features of the corresponding perspective.

[0011] Preferably, the Raman spectroscopy analysis branch includes a preprocessing module, a convolution layer, a multi-scale fusion module and a residual convolution module connected in sequence.

[0012] Preferably, the acoustic excitation module includes an ultrasonic transducer device; the ultrasonic transducer device includes a plurality of air-coupled ultrasonic transducer array elements arranged in a ring array.

[0013] Preferably, the signal light acquisition module includes a common optical path and three modal imaging optical paths; the three modal imaging optical paths are respectively an OCT imaging optical path, a Raman imaging optical path and a white light imaging optical path;

[0014] The common optical path includes a first long-wavelength pass dichroic filter, an XY scanning galvanometer, a second long-wavelength pass dichroic filter, a first focusing lens, and a transparent plane mirror. The light beams input into the common optical path by the OCT imaging optical path and the Raman imaging optical path pass through the first long-wavelength pass dichroic filter, the XY scanning galvanometer, the second long-wavelength pass dichroic filter, the first focusing lens, and the transparent plane mirror in sequence, focusing the light beams on the intestinal tissue and receiving the reflected light beams to realize the collection of OCT information and Raman information.

[0015] The white light imaging optical path includes a white light source and a CMOS camera. The white light imaging is provided by the white light source to provide illumination conditions, and the signal light is collected by using a common optical path. The signal light passes through a transparent plane mirror, a first focusing lens, and a second long-wavelength dichroic filter to enter the CMOS camera, thereby realizing the collection of white light information.

[0016] In a second aspect, the present invention provides an intestinal tissue auxiliary analysis method based on multimodal data combination, which uses the above-mentioned multimodal intestinal tissue detection system; the multimodal intestinal tissue detection method includes the following steps:

[0017] Step 1: Use a multimodal imaging system to collect multimodal image data of intestinal tissue with different lesions and construct a labeled intestinal lesion dataset. The lesion dataset samples include OCE images, OCT images, Raman spectra, and white light images of intestinal tissue, as well as labels for different lesions and their severity.

[0018] Step 2: Construct an intestinal lesion assessment model; the intestinal lesion assessment model includes a feature extraction module and a multimodal feature fusion classification module. The feature extraction module extracts feature vectors of OCE images, OCT images, Raman spectra, and white light images, and the multimodal feature fusion classification module fuses and classifies them to obtain the final assessment results.

[0019] Step 3: Use the intestinal lesion dataset to train the intestinal lesion assessment model;

[0020] Step 4: Input the subject's OCE image, OCT image, Raman spectrum image and white light image into the trained intestinal lesion assessment model to obtain the final assessment result.

[0021] Preferably, the method for collecting OCE images is as follows: using an ultrasonic transducer device to transmit ultrasonic waves to the intestinal tissue to excite the intestinal tissue, and collecting OCT information through a signal light acquisition module; measuring the phase change of the spectrum after the interference of the reflected light inside the tissue and the reference light after ultrasonic excitation to calculate the deformation of the intestinal tissue; modeling the force on the intestinal tissue and constructing a theoretical displacement function of the intestinal tissue; obtaining the shear modulus by minimizing the difference between the deformation at the same position of the intestinal tissue and the theoretical displacement of the intestinal tissue, and then reconstructing the OCE image data of the intestinal tissue; the reference light is the light reflected back from collecting the OCT information.

[0022] The present invention has the following beneficial effects:

[0023] 1. This invention uses a first bimodal analysis branch to interactively analyze OCT and OCE features. By converting the analysis of three-dimensional feature maps into two-dimensional spatial feature maps, this reduces computational complexity and model size while improving feature analysis efficiency. Simultaneously, covering the three-dimensional space from three orthogonal perspectives avoids information loss from a single perspective. Furthermore, due to the varying structural density and elasticity of different layers of intestinal tissue, establishing a correlation between the two modalities allows for complementary structural and elastic information.

[0024] 2. The present invention adopts the second bimodal analysis branch to fuse OCE features and white light features. Based on the fact that OCE can better reflect the characteristics of intestinal lesions, the OCE image is used to guide the model to focus on the lesion area in the white light image, thereby enhancing the model's ability to characterize the pathological characteristics of the mucosal layer, thereby improving the overall robustness of the model; at the same time, the present invention performs feature extraction through the cyclic cross-attention module in the second bimodal analysis branch. Compared with the existing cyclic cross-attention module, the present invention reduces the computational complexity while maintaining the integrity of the feature expression.

[0025] 3. The present invention utilizes a ring-shaped multi-element ultrasonic transducer to achieve ultrasonic excitation of the sample. The ultrasonic excitation device transmits ultrasonic waves toward the sample at a certain angle, achieving precise and uniform ultrasonic excitation of the sample area, thereby enhancing the OCE imaging quality. At the same time, the redundant design of the ring array improves the fault tolerance of the system. When a single or multiple elements fail, the remaining elements can still maintain basic functions, avoiding the inconvenience of temporary instrument replacement and significantly reducing the overall impact of the signal light acquisition subsystem on the OCE imaging quality.

[0026] 4. The present invention combines optical coherence elastography, Raman spectroscopy, and endoscopic imaging to achieve structural and functional imaging of lesions, and realizes multimodal comprehensive detection of intestinal lesions through tissue elasticity analysis by OCE, tissue structure analysis by OCT, molecular property detection by Raman spectroscopy, and visualization of the intestinal mucosal layer by white light endoscopy. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 Schematic diagram of a multimodal imaging system in Example 1 of the present invention.

[0028] Figure 2 This is a schematic diagram of the signal light acquisition module in Example 1 of the present invention.

[0029] Figure 3 Schematic diagram of the structure of the OCT spectrometer in Example 1 of the present invention.

[0030] Figure 4 Schematic diagram of the structure of the Raman spectrometer in Example 1 of the present invention.

[0031] Figure 5 Schematic diagram of the intestinal lesion assessment system in Example 1 of the present invention.

[0032] Figure 6 This is a schematic diagram of the structure of the multi-view cross-attention module in Example 1 of the present invention.

[0033] Figure 7 This is the three-dimensional feature map input to the multi-view cross-attention module in Example 1 of the present invention.

[0034] Figure 8 Schematic diagram of the residual convolution module structure in Example 1 of the present invention.

[0035] Figure 9 Schematic diagram of the traditional recurrent cross-attention module structure.

[0036] Figure 10 Schematic diagram of the structure of the cyclic cross attention module in Example 1 of the present invention.

[0037] Figure 11 Schematic diagram of the multi-scale fusion module structure in Example 1 of the present invention.

[0038] Reference numerals: 1, first long-wavelength pass dichroic filter; 2, XY scanning galvanometer; 3, second long-wavelength pass dichroic filter; 4, first focusing lens; 5, reflector; 6, OCT light source; 7, sample arm collimating lens; 8, OCT fiber coupler; 9, reference arm collimating lens; 10, reference arm focusing lens; 11, reference arm reflector; 12, OCT spectrometer; 13, Raman light source; 14, first collimating lens; 15, Raman fiber coupler ; 16. Raman spectrometer; 17. White light source; 18. CMOS camera; 19. Second collimating lens; 20. First transmission diffraction grating; 21. Second focusing lens; 22. First linear array CCD; 23. Third collimating lens; 24. Rayleigh wave filter; 25. Second transmission diffraction grating; 26. Third focusing lens; 27. Second linear array CCD; 28. Camera fixing device; 29. ​​Transparent plane mirror; 30. Ultrasonic transducer device. DETAILED DESCRIPTION

[0039] The present invention will be further described below with reference to the accompanying drawings.

[0040] Example 1

[0041] like Figure 1As shown, an intestinal tissue auxiliary analysis system based on multimodal data integration includes a multimodal imaging system and an intestinal lesion assessment system. The multimodal imaging system includes an endoscopic probe, a signal light acquisition module, and an acoustic excitation module installed at the end of the endoscopic probe. The signal light acquisition module is used to acquire OCT (optical coherence tomography), Raman, and white light information from intestinal tissue samples. The signal light acquisition module includes a common optical path and three imaging optical paths. The common optical path is built into the endoscopic probe; the three imaging optical paths are OCT imaging, Raman imaging, and white light imaging. Single-mode optical fiber is used to transmit light throughout the multimodal imaging system.

[0042] The common optical path is used to integrate the imaging optical paths of the three modalities, achieving coaxial unification of the optical paths and ensuring that the imaging data of different modalities are highly consistent in space, which facilitates comprehensive analysis by doctors. At the same time, multimodal information can be integrated without the need for complex image registration algorithms. Figure 1 and 2 As shown, the common optical path includes a first long-wavelength pass dichroic filter 1, an XY scanning galvanometer 2, a second long-wavelength pass dichroic filter 3, a first focusing lens 4, and a transparent plane mirror 29. The first and second long-wavelength pass dichroic filters 1 and 3 transmit and reflect different wavelength bands of the light beam through their spectroscopic properties, thereby integrating the imaging optical paths of the three modalities. The first long-wavelength pass dichroic filter 1 combines and splits the OCT light and Raman light, transmitting the OCT beam and reflecting the Raman beam. The second long-wavelength pass dichroic filter 3 reflects and transmits the white light signal and the combined OCT and Raman light, respectively. The light beams input from the OCT imaging optical path and the Raman imaging optical path into the common optical path pass sequentially through the first long-pass dichroic filter 1, the XY scanning galvanometer 2, the second long-pass dichroic filter 3, the first focusing lens 4, and the transparent plane mirror 29, focusing the light beams on the sample. By receiving the reflected light beams, OCT and Raman information are collected. The light beams input from the white light imaging optical path into the common optical path pass through the second long-pass dichroic filter 3, the first focusing lens 4, and the transparent plane mirror 29, focusing the light beams on the sample. By receiving the reflected light beams, white light information is collected.

[0043] The OCT imaging optical path includes an OCT light source 6, a sample arm collimating lens 7, an OCT fiber coupler 8, a reference arm collimating lens 9, a reference arm focusing lens 10, a reference arm reflector 11, and an OCT spectrometer 12. The OCT imaging optical path and the common optical path together constitute an OCT imaging system. In the OCT imaging system, the OCT light source 6 transmits a laser beam to the OCT fiber coupler 8 via a single-mode optical fiber. The OCT fiber coupler 8 splits the beam into two parts. One part is transmitted via optical fiber to the reference arm, which consists of the reference arm collimating lens 9, the reference arm focusing lens 10, and the reference arm reflector 11. The other part is transmitted via optical fiber to the sample arm, which consists of the collimating lens 7 and the common optical path. The light beam entering the reference arm is collimated by the reference arm collimating lens 9 and then directed toward the reference arm focusing lens 10, which focuses the beam onto the reference arm reflector 11. After reflection, the beam propagates back along its original path. The light beam entering the sample arm is converted into a parallel beam by the sample arm collimating lens 7 and then directed toward the common optical path. After passing through the common optical path, the OCT beam strikes the sample and then propagates back along its original path, resulting in an OCT beam carrying sample information. The OCT beam returning from the sample arm interferes with the light returning from the reference arm at the OCT fiber coupler 8, and is then transmitted via optical fiber to the OCT spectrometer 12 to obtain OCT information. This OCT information reflects the depth of intestinal tissue, producing a cross-sectional image (depth image) of the tissue, facilitating subsequent identification and analysis of intestinal lesions.

[0044] like Figure 3 As shown, the OCT spectrometer 12 includes a second collimating lens 19, a first transmissive diffraction grating 20, a second focusing lens 21, and a first linear CCD array 22. The OCT beam entering the OCT spectrometer 12 is collimated by the second collimating lens 19 and then split by the diffraction grating 20. Finally, the beam is focused by the second focusing lens 21 onto the first linear CCD array 22, thereby achieving OCT imaging.

[0045] The Raman imaging optical path includes a reflector 5, a Raman light source 13, a first collimating lens 14, a Raman fiber coupler 15, and a Raman spectrometer 16. The Raman imaging optical path and the common optical path together constitute a Raman imaging system. In the Raman imaging system, the Raman light source 13 transmits a laser beam via a single-mode optical fiber to the Raman fiber coupler 15, and then transmits it to the first collimating lens 14 via the optical fiber to obtain a parallel beam. The parallel beam is then reflected by the reflector 5 into the common optical path. After irradiating the sample through the common optical path, the Raman beam propagates in the opposite direction along the original optical path to obtain a Raman beam carrying sample information. The Raman beam is then transmitted via the Raman fiber coupler 15 to the Raman spectrometer 16 to obtain spectral information. This spectral information can be used to obtain the spectral peak difference of intestinal type I collagen, reflecting the stress changes in the intestine and facilitating the subsequent identification and analysis of lesions such as intestinal fibrosis.

[0046] like Figure 4As shown, the Raman spectrometer 16 includes a third collimating lens 23, a Rayleigh wave filter 24, a second transmissive diffraction grating 25, a third focusing lens 26, and a second linear CCD array 27. The Raman beam entering the Raman spectrometer 16 is collimated by the third collimating lens 23, filtered out by the Rayleigh wave filter 24, and then split by the second transmissive diffraction grating 25. Finally, the beam is focused by the third focusing lens 26 onto the second linear CCD array 27, thereby achieving Raman spectral imaging.

[0047] The white-light imaging optical path includes a white-light light source 17, a CMOS camera 18, and a camera fixture 28. Together, the white-light imaging optical path and the common optical path constitute a white-light imaging system. In this system, the white-light light source 17 illuminates the sample. The CMOS camera 18, secured to the endoscopic probe via the camera fixture 28, receives white-light information reflected from the sample by the white-light source 17 based on the spectroscopic properties of the second long-wave-pass dichroic filter 3. This information is then collected to achieve white-light imaging of the intestine. This information can reflect the condition of the intestinal mucosa, facilitating subsequent identification and analysis of intestinal lesions.

[0048] The acoustic excitation module includes an ultrasonic transducer 30, which is used to deform tissue to obtain elastic information, thereby performing optical coherence elastography (OCE) imaging based on the spectral information collected by the OCT system. The ultrasonic transducer 30 comprises 12 air-coupled ultrasonic transducer elements, each with an effective area of ​​approximately 1 mm x 1 mm, arranged in a circular array. Each element is oriented at a 45° angle. This tilt ensures precise focusing of ultrasound waves on the sample surface, enabling non-contact excitation. This structural design offers two advantages: First, the symmetrical distribution of the 12 air-coupled ultrasonic transducer elements ensures uniform acoustic field coverage of the sample surface, ensuring excitation stability. Second, the redundant design of the circular array enhances system fault tolerance. If one or more elements fail, the remaining elements can maintain basic functionality, avoiding the inconvenience of temporary instrument replacement and significantly reducing the overall impact of the signal light acquisition subsystem on OCE imaging quality.

[0049] In this embodiment, the first long-wavepass dichroic filter 1 in the common optical path has a cutoff wavelength of 990 nm, a transmission band of 1000 nm to 1400 nm, and a reflection band of 780 nm to 980 nm.

[0050] In this embodiment, the second long-wavepass dichroic filter 2 in the common optical path has a cutoff wavelength of 770 nm, a transmission band of 780 nm to 1400 nm, and a reflection band of 400 nm to 760 nm.

[0051] In this embodiment, the effective focal length of the first focusing lens 4 in the common optical path is 10 mm, and the lens diameter is 5 mm.

[0052] In this embodiment, the OCT light source 6 in the OCT imaging system is a superluminescent diode with a central wavelength of 1300 nm, a power of 5 mW, and a full width at half maximum of 75 nm.

[0053] In this embodiment, the OCT fiber coupler 8 in the OCT imaging system is a single-mode 2×2 broadband fiber coupler with a central wavelength of 1300±75 nm and a coupling ratio of 50:50.

[0054] In this embodiment, the single-mode optical fiber used in the OCT imaging system has an operating wavelength of 1200 nm to 1390 nm and a numerical aperture NA of 0.12.

[0055] In this embodiment, the sample arm collimating lens 7 and the reference arm collimating lens 9 in the OCT imaging system both have an effective focal length of 10 mm and a diameter of 3 mm.

[0056] In this embodiment, the effective focal length of the second collimating lens 19 in the OCT spectrometer 12 is 10 mm, and the diameter is 5 mm.

[0057] In this embodiment, the first transmission diffraction grating 20 in the OCT spectrometer 12 has a wavelength of 1145 lines / mm and an incident angle of 48.6°.

[0058] In this embodiment, the effective focal length of the second focusing lens 21 in the OCT spectrometer 12 is 25 mm, and the diameter is 30 mm.

[0059] In this embodiment, the first linear array CCD 22 in the OCT spectrometer 12 has 2048 pixels, and the pixel size is 10 μm×10 μm.

[0060] In this embodiment, the spot size of the light beam after being collimated by the sample arm collimating lens 7 and the reference arm collimating lens 9 of the OCT imaging system is approximately 2.4 mm.

[0061] In this embodiment, the OCT imaging system has a lateral resolution ∆x of approximately 6.90 μm and an axial resolution ∆z of approximately 4.96 μm. This allows for the acquisition of high-resolution depth and stereoscopic images of the intestine. OCT images provide intuitive visualization of the intestinal cross-sectional structure, thickness, lesion depth, and distribution of fibrotic tissue, facilitating analysis and assessment of lesion extent and invasion depth. Furthermore, the OCT system acquires phase information from the OCT interferometer spectrum under excitation to obtain biomechanical parameters of the tissue, such as its elastic properties, providing crucial information for the identification and assessment of intestinal lesions.

[0062] In this embodiment, the Raman light source 13 in the Raman imaging system is a semiconductor laser with a central wavelength of 785 nm and an effective power of 20 mW. ~2100 spectral range.

[0063] In this embodiment, the Raman fiber coupler 15 in the Raman imaging system is a single-mode 1×2 broadband fiber coupler with a central wavelength of 865±85 nm and a coupling ratio of 50:50.

[0064] In this embodiment, the single-mode optical fiber used in the Raman imaging system has an operating wavelength of 780-950 nm and a numerical aperture NA of 0.12.

[0065] In this embodiment, the first collimating lens 14 in the Raman imaging system has an effective focal length of 10 mm and a diameter of 3 mm.

[0066] In this embodiment, the spot size of the light after the first collimating lens 14 in the Raman imaging system collimates the light is about 2.4 mm.

[0067] In this embodiment, the third collimating lens 23 in the Raman spectrometer 16 has an effective focal length of 10 mm and a diameter of 5 mm.

[0068] In this embodiment, the second transmission diffraction grating 25 in the Raman spectrometer 16 has a wavelength of 1800 lines / mm and an incident angle of 42.8°.

[0069] In this embodiment, the effective focal length of the third focusing lens 26 in the Raman spectrometer 16 is 45 mm, and the diameter is 44 mm.

[0070] In this embodiment, the second linear array CCD 27 in the Raman spectrometer 16 has 2048×264 pixels, and the pixel size is 15 μm×15 μm.

[0071] In this embodiment, the Raman scattering wavelength of the Raman imaging system is 807-944 nm, that is, the spectral range is 350-2100 nm. , the spectral resolution is about 2 , which helps doctors analyze the degree of intestinal fibrosis and study the molecular mechanism of fibrosis.

[0072] In this embodiment, the swing angle of the Y galvanometer mirror in the XY scanning galvanometer mirror 2 is ±2.86°, and the swing angle of the X galvanometer mirror is ±2.04°, thereby achieving a scanning range of 2 mm×2 mm for the sample.

[0073] In this embodiment, the white light source 17 in the white light imaging system is an LED white light source, and its wavelength range is 400 nm to 760 nm.

[0074] In this embodiment, the CMOS camera 18 in the white light imaging system has 2048×2048 pixels, and the pixel size is 1.4 μm×1.4 μm.

[0075] In this embodiment, the optimal working distance of the endoscope is 3 mm.

[0076] like Figure 5 As shown in the figure, to achieve accurate multimodal analysis, the intestinal lesion assessment system proposes corresponding feature extraction modules and multimodal feature fusion classification modules based on the characteristics of different modalities. The feature extraction module includes a first bimodal analysis branch, a second bimodal analysis branch, and a Raman spectroscopy analysis branch.

[0077] Because OCE and OCT modalities are similar and physically related, the first bimodal analysis branch is used to interactively analyze the features of these two modalities. This branch is used to analyze and fuse OCE and OCT features. It consists of two feature encoders, a multi-view cross-attention module, two residual convolution modules, and two global average pooling layers. Two 3D feature encoders extract 3D feature maps of OCT image data and OCE image data respectively, and a multi-view cross-attention module is used to fuse the 3D feature maps extracted by the two 3D feature encoders to obtain OCT fusion spatial feature maps focusing on OCT features and OCE fusion spatial feature maps focusing on OCE features from different viewpoints, thereby realizing the elasticity and structural feature interaction between OCE image data and OCT image data at the spatial level; the fused spatial feature maps of different viewpoints are further processed with advanced deep features through the residual convolution module, and pooled through the global average pooling layer to obtain OCT feature vectors and OCE feature vectors of different viewpoints; the feature vectors of different viewpoints are spliced ​​to obtain the OCE feature vectors and OCT feature vectors finally output by the first bimodal analysis branch to support subsequent classification decisions.

[0078] In this embodiment, the feature encoder adopts the architecture of the ResNet series or the architecture of the Transformer series.

[0079] like Figure 6 As shown in the figure, the multi-view cross-attention module includes two multi-view decoupling modules and three cross-attention modules. The two multi-view decoupling modules perform axial analysis on the two three-dimensional feature maps based on the self-attention method, and decompose each three-dimensional feature map into a two-dimensional OCT initial spatial feature map and a two-dimensional OCE initial spatial feature map of three orthogonal perspectives of H×W, W×D and H×D, thereby realizing the transformation of the analysis of the three-dimensional feature map into the analysis of the two-dimensional spatial feature map, reducing the amount of calculation and the model volume while improving the efficiency of feature analysis. In addition, covering the three-dimensional space from three orthogonal perspectives also avoids the loss of information from a single perspective. As shown in the figure, the multi-view cross-attention module includes two multi-view decoupling modules and three multi-view decoupling modules. Figure 7 The figure shows two examples of three-dimensional feature maps in the input multi-view decoupling module. Taking the generation process of an H×W two-dimensional spatial feature map as an example, the three-dimensional feature map is flattened to (H×W)×D. A 1×1 convolutional layer is used to encode the query matrix, key matrix, and value matrix, each of size (H×W)×D. The query matrix is ​​multiplied by the key matrix and the attention weights are calculated using Softmax, resulting in a D×D attention weight map. This represents the depth-wise correlation between the query matrix and the key matrix. The D×D attention weight map is multiplied by the value matrix and averaged in the depth direction to obtain an H×W two-dimensional initial spatial feature map for the D-axis perspective. This method allows for dynamic weight distribution of three-dimensional features along the axial direction.

[0080] The cross-attention module utilizes a cross-attention strategy to establish associations between bimodal features, thereby using features from one modality to assist in locating key areas in the other modality. Because different layers of intestinal tissue have varying structural density and elasticity, establishing an association between the two modalities allows for complementary structural and elastic information. The cross-attention module encodes the initial spatial feature maps from the same viewpoint in both modalities using three 1×1 convolutional layers, each encoding them into three projection matrices: query, key, and value. The query matrix of one modality is multiplied by the key matrix of the other modality, processed by softmax, and then multiplied by the value matrix of the other modality. Each cross-attention module outputs an OCT fusion spatial feature map focusing on OCT features and an OCE fusion spatial feature map focusing on OCE features for the corresponding viewpoint, thereby establishing the association between one modality and the other.

[0081] like Figure 8 As shown in the figure, the residual convolution module includes multiple groups of convolutional layers with residual structures, batch normalization and ReLU activation functions, which are used to deepen the model level and further extract high-level features; the convolution layer in the first bimodal analysis branch uses a two-dimensional convolution kernel of size 3×3.

[0082] Since both white light images and OCE images contain information about the mucosal layer, the second bimodal analysis branch is used to guide the model to focus on the lesion area in the white light image through the mechanical characteristics reflected by the OCE image, thereby enhancing the model's ability to represent the pathological characteristics of the mucosal layer.

[0083] The second bimodal analysis branch consists of a sequentially connected two-dimensional feature encoder, an elastic attention guidance module, a recurrent cross-attention module, and a global average pooling layer. The two-dimensional feature encoder extracts multi-scale features from the white-light image to obtain a raw feature map, which provides a foundational representation for subsequent multimodal interaction. The elastic attention guidance module fuses the raw feature map with the OCE fusion spatial feature map of the D-axis (H×W) viewing angle to generate a correlation feature map, which strengthens the model's focus on the lesion area in the white-light image. Based on the correlation feature map, a recurrent cross-attention module is further used to extract global features of the white-light image, resulting in a fused feature map. Finally, the fused feature map is processed by a global average pooling layer to obtain a feature vector for the white-light image, supporting subsequent classification decisions.

[0084] In this embodiment, the two-dimensional feature encoder uses ResNet or Transformer as the basic feature extraction network.

[0085] The elastic attention guidance module uses a sigmoid activation function to nonlinearly map the OCE fusion spatial feature map of the D-axis (H×W) viewing angle to generate an attention weight distribution map along the depth direction (D-axis). Since regions with higher elastic eigenvalues ​​in the OCE feature map correspond to areas of greater tissue hardness, the sigmoid activation function assigns higher weights to these regions, thereby highlighting the response of the lesion area. Subsequently, the OCE modality attention weight distribution map is element-wise multiplied with the original feature map, guiding the model to focus more on the lesion area in the white light image and dynamically enhancing the feature representation related to the lesion.

[0086] like Figure 9 and Figure 10 As shown, compared with the existing cyclic cross-attention module, the cyclic cross-attention module in the present invention reduces repeated calculations and realizes the global interaction of the three-modal features of OCT, OCE, and white light in spatial features; while reducing the computational complexity, it maintains the integrity of feature expression.

[0087] The recurrent cross-attention module encodes the associated feature map through three convolutional layers with a kernel size of 1×1, resulting in a query matrix, a key matrix, and a value matrix. First attention scores are calculated horizontally and vertically for the query matrix and the key matrix. The first attention scores are weighted summed on the value matrix to produce a first intermediate feature map, enabling interaction between each pixel on the feature map and its corresponding horizontal and vertical pixel information. To achieve interaction with all pixel information, second attention scores are calculated horizontally and vertically for the first intermediate feature map and the query matrix, and the value matrix is ​​weighted summed to produce a second intermediate feature map. This second intermediate feature map is then fused with the associated feature map to produce a fused feature map.

[0088] Since the molecular characteristics reflected by Raman spectral data are quite different from those of other modalities, it is treated as an independent branch in the system. The Raman spectral analysis branch is used to extract the characteristic vectors of Raman spectral data separately, which is used to analyze the spectral data of type I collagen to provide auxiliary support for the identification of intestinal lesions.

[0089] The Raman spectroscopy analysis branch consists of a preprocessing module, a 1×3 one-dimensional convolutional layer, a multi-scale fusion module, and a residual convolution module. The preprocessing module uses wavelet transform to reduce noise and improve the signal-to-noise ratio of the spectrum. It also enhances the input data using methods such as random spectral shifting, random spectral cropping, and random Gaussian noise. The convolutional layer in the Raman analysis branch uses a 1×3 one-dimensional convolution kernel.

[0090] like Figure 11 As shown, the multi-scale fusion module extracts multi-scale features by using one-dimensional convolution kernels of different sizes (including 1×12, 1×6, and 1×3) to extract features from the shallow local feature maps output by the one-dimensional convolution. The 1×12 convolution kernel is suitable for extracting broad peak features, the 1×6 convolution kernel is suitable for extracting medium-width peak features, and the 1×3 convolution kernel focuses on capturing narrow peak features. For example, it extracts features such as the peak value, peak position, and peak width of regions like the amide I or amide III bands to enhance the richness of spectral features. The extracted multi-scale features are concatenated and then fused through a 1×1 convolution layer. The residual convolution module comprises multiple concatenated convolution layers with a residual structure, batch normalization, and ReLU activation functions. This deepens the model hierarchy and further extracts high-level features, resulting in Raman spectral feature vectors to support subsequent analysis tasks. The convolution layers in the residual convolution module use 1×3 convolution kernels.

[0091] The multimodal feature fusion classification module achieves deep fusion of multimodal information and classification decisions by integrating the feature vectors extracted by the first bimodal analysis branch, the second bimodal analysis branch, and the Raman spectroscopy analysis branch. The feature vectors generated by these modules are first concatenated along the channel dimension to form a unified joint feature representation. These concatenated features are then fed into a fully connected layer, where high-level discriminant features are extracted through a nonlinear transformation. Finally, the classification result is output, yielding the final evaluation result. This module fully leverages the complementary information of different modalities to improve classification performance.

[0092] Example 2

[0093] An intestinal tissue auxiliary analysis method based on multimodal data combination uses the multimodal intestinal tissue detection system in Example 1; the multimodal intestinal tissue detection method includes the following steps:

[0094] Step 1: Image Acquisition

[0095] A multimodal imaging system was used to collect multimodal image data of intestinal tissue with different lesions to construct a labeled intestinal lesion dataset. The lesion dataset samples included OCE images, OCT images, Raman spectra, and white light images of intestinal tissue, as well as labels for normal tissue, different lesions, and their severity. Lesions primarily included ulcers, erosions, and fibrosis, with fibrotic lesions classified as mild, moderate, or severe, depending on the extent and range of fibrous tissue proliferation. OCT image data of intestinal tissue was obtained using an OCT imaging system, and OCE image data of intestinal tissue was obtained based on the OCT imaging system and an acoustic excitation module. Spectral image data of type I collagen molecules in intestinal tissue was obtained using a Raman imaging system, and white light image data of the intestinal mucosal layer in intestinal tissue was obtained using a white light imaging system. The image data of the four modalities reflect the characteristics of intestinal tissue from different levels: OCE images quantify the degree of tissue hardening through elastic modulus, providing a mechanical basis for the identification of intestinal lesions; OCT images reveal the deep structural characteristics of intestinal tissue and can reflect some structural changes of lesions such as fibrosis; Raman spectroscopy provides a molecular basis for the identification of intestinal lesions by analyzing the molecular vibration characteristics of type I collagen; white light images intuitively reflect the information of the mucosal surface of intestinal tissue and can be used to identify the morphological and color characteristics of lesions such as tissue surface fibrosis. Based on the complementarity of the above four modal data, the fusion of multimodal features can achieve highly accurate automated auxiliary analysis to provide support for doctors' clinical decision-making. The specific process of obtaining intestinal tissue OCE image data based on the OCT imaging system and acoustic excitation module is as follows:

[0096] The ultrasonic transducer 30 is used to transmit ultrasonic waves to the intestinal tissue to achieve the stimulation of the intestinal tissue. At the same time, the signal light acquisition module collects OCT information. Based on the OCT interference principle, the phase change of the reflected light inside the tissue after ultrasonic stimulation is measured. , and based on Calculating the deformation of intestinal tissue .in, represents the optical refractive index of intestinal tissue; Indicates the average wavelength of the light source output by the OCE. Based on the equation of motion , the force on the intestinal tissue is modeled; where m is the equivalent mass; c is the viscosity coefficient; is the equivalent spring stiffness; is the deformation variable. In the modeling process, the intestinal tissue is regarded as a multi-layer viscoelastic medium, and the following assumptions are made: the structure of each layer of the medium is uniform and incompressible; ultrasound is an axisymmetric radiation force acting on the surface of the medium; the mechanical parameters (Young's elastic modulus and shear viscosity) and density of the medium are constants. Based on the above assumptions, the theoretical displacement function of the intestinal tissue can be expressed as Y x,y,z =- ∫ 0 ∞ α 2 J (α x 2 + y 2 )[ A 1 e -αz - A 2 e αz +α( B 1 e -βz + B 2 e βz )]dα ;in, Indicates independent variables From 0 to 's points; is the Bessel function of order 0; , , are the lateral distance and height of a certain point of the tissue in the rectangular coordinate system respectively; 、 、 、 Obtained by the set boundary conditions; ; ; is Young's elastic modulus; is tissue density; is the angular frequency; is the shear viscosity. By minimizing the deformation at the same position of the intestinal tissue and theoretical displacement The difference between and Young's modulus of elasticity , and based on Obtain the shear modulus , and then obtain multiple mechanical properties of the same position in the tissue, and finally reconstruct the OCE elastic map (OCE image data) of the intestinal tissue.

[0097] Step 2: Construct an intestinal lesion assessment model; the intestinal lesion assessment model includes a feature extraction module and a multimodal feature fusion classification module. The feature extraction module extracts feature vectors of OCE images, OCT images, Raman spectra, and white light images, and the multimodal feature fusion classification module fuses and classifies them to obtain the final assessment results.

[0098] Step 3: Use the intestinal lesion dataset to train the intestinal lesion assessment model.

[0099] Step 4: Input the subject's OCE image, OCT image, Raman spectrum and white light image into the trained intestinal lesion assessment model, and output the final assessment results for auxiliary analysis of intestinal tissue.

Claims

1. An intestinal tissue auxiliary analysis system based on multimodal data integration, including a multimodal imaging system and an intestinal lesion assessment system; the multimodal imaging system includes an endoscopic probe and a signal light acquisition module; The signal light acquisition module is used to acquire OCT, Raman, and white light information from intestinal tissue. The intestinal lesion assessment system is used to perform lesion assessment based on the information acquired by the signal light acquisition module. The system is characterized by further comprising an acoustic excitation module installed at the end of the endoscope probe for acquiring OCE information from intestinal tissue. The intestinal lesion assessment system includes a feature extraction module and a multimodal feature fusion classification module; the feature extraction module includes a first bimodal analysis branch, a second bimodal analysis branch, and a Raman spectroscopy analysis branch; the first bimodal analysis branch is used to fuse OCE features and OCT features; the first bimodal analysis branch extracts the query matrix, key matrix, and value matrix of the OCT features and the OCE features, respectively, and realizes the interaction between the OCT features and the OCE features through a cross-attention mechanism, thereby obtaining an OCT fusion spatial feature map focusing on the OCT features and an OCE fusion spatial feature map focusing on the OCE features from different perspectives; The second bimodal analysis branch is used to fuse white light features and OCE features; the Raman spectrum analysis branch is used to extract Raman spectrum features; the multimodal feature fusion classification module is used to fuse the feature vectors extracted by the first bimodal analysis branch, the second bimodal analysis branch, and the Raman spectrum analysis branch to obtain the final evaluation result; The second bimodal analysis branch includes a cyclic cross attention module for capturing global context information; the cyclic cross attention module encodes the feature map of the input cyclic cross attention module through three convolutional layers to obtain a query matrix, a key matrix, and a value matrix; the query matrix and the key matrix are dot-producted to obtain a first attention score; the value matrix is ​​weighted summed using the first attention score to obtain a first intermediate feature map; the first intermediate feature map is dot-producted with the query matrix to obtain a second attention score, the value matrix is ​​weighted summed using the second attention score to obtain a second intermediate feature map, and the second intermediate feature map is fused with the feature map of the input cyclic cross attention module to obtain a fused feature map; The second bimodal analysis branch also includes an encoder, an elastic attention guidance module and a global average pooling layer; the encoder extracts the original feature map of the white light image; the elastic attention guidance module processes the OCE fusion spatial feature map through an activation function, and multiplies the processing result with the original feature map element by element to obtain a correlation feature map; the correlation feature map is processed in turn through a cyclic cross attention module and a global average pooling layer to obtain a white light feature vector.

2. The intestinal tissue auxiliary analysis system based on multimodal data combination according to claim 1 is characterized by: The first bimodal analysis branch includes two feature encoders, a multi-view cross-attention module, two residual convolution modules and two global average pooling layers; the two feature encoders respectively extract three-dimensional feature maps of OCT image data and OCE image data, and use the multi-view cross-attention module to interact with the two three-dimensional feature maps to obtain OCT fusion spatial feature maps focusing on OCT features and OCE fusion spatial feature maps focusing on OCE features under different viewpoints; the fusion spatial feature maps of different viewpoints are further processed by the residual convolution module and the global average pooling layer in turn to obtain OCT feature vectors and OCE feature vectors of different viewpoints; the feature vectors of different viewpoints are spliced ​​to obtain the OCE feature vector and OCT feature vector finally output by the first bimodal analysis branch.

3. The intestinal tissue auxiliary analysis system based on multimodal data combination according to claim 2 is characterized by: The multi-view cross-attention module includes two multi-view decoupling modules and three cross-attention modules; the two multi-view decoupling modules respectively decompose the two input three-dimensional feature maps into three orthogonal perspectives of two-dimensional OCT initial spatial feature maps and two-dimensional OCE initial spatial feature maps; the initial spatial feature maps of the same perspective are processed respectively by the three cross-attention modules, and each cross-attention module outputs an OCT fusion spatial feature map focusing on OCT features and an OCE fusion spatial feature map focusing on OCE features of the corresponding perspective.

4. The intestinal tissue auxiliary analysis system based on multimodal data combination according to claim 1, characterized in that: The acoustic excitation module comprises an ultrasonic transducer device (30); the ultrasonic transducer device (30) comprises a plurality of air-coupled ultrasonic transducer array elements arranged in a ring array.

5. The intestinal tissue auxiliary analysis system based on multimodal data combination according to claim 1, characterized in that: The signal light acquisition module includes a common optical path and three modal imaging optical paths; the three modal imaging optical paths are OCT imaging optical path, Raman imaging optical path and white light imaging optical path; the common optical path is built into the endoscope probe; The common optical path includes a first long-wavelength dichroic filter (1), an XY scanning galvanometer (2), a second long-wavelength dichroic filter (3), a first focusing lens (4), and a transparent plane mirror (29); the light beams input into the common optical path by the OCT imaging optical path and the Raman imaging optical path sequentially pass through the first long-wavelength dichroic filter (1), the XY scanning galvanometer (2), the second long-wavelength dichroic filter (3), the first focusing lens (4), and the transparent plane mirror (29), thereby focusing the light beams on the intestinal tissue and realizing the collection of OCT information and Raman information by receiving the light beams reflected therefrom; The white light imaging optical path includes a white light source (17) and a CMOS camera (18); the white light imaging is provided with illumination conditions by the white light source (17), and the signal light is collected by using a common optical path. The signal light passes through a transparent plane mirror (29), a first focusing lens (4), and a second long-wavelength dichroic filter (3) and enters the CMOS camera (18), thereby realizing the collection of white light information.

6. The intestinal tissue auxiliary analysis system based on multimodal data combination according to claim 1, characterized in that: The Raman spectrum analysis branch includes a preprocessing module, a convolution layer, a multi-scale fusion module and a residual convolution module which are connected in sequence.

7. A method for assisting intestinal tissue analysis based on multimodal data integration, characterized by: Using the intestinal tissue auxiliary analysis system based on multimodal data combination as described in claim 1; The intestinal tissue-assisted analysis method includes the following steps: using a multimodal imaging system to collect multimodal image data of intestinal tissue with different lesions to construct a labeled intestinal lesion dataset; constructing an intestinal lesion assessment model and training it using the intestinal lesion dataset; inputting the OCE image, OCT image, Raman spectrum and white light image of the subject into the trained intestinal lesion assessment model, and outputting the final assessment result; The intestinal lesion assessment model includes a feature extraction module and a multimodal feature fusion classification module; The feature vectors of OCE images, OCT images, Raman spectra and white light images are extracted through the feature extraction module, and fused and classified through the multimodal feature fusion classification module.

8. The intestinal tissue auxiliary analysis method based on multimodal data combination according to claim 7, characterized in that: The method for collecting OCE images is as follows: using an ultrasonic transducer device (30) to transmit ultrasonic waves to the intestinal tissue to excite the intestinal tissue, and collecting OCT information through a signal light acquisition module; measuring the phase change of the interference spectrum of the reflected light inside the tissue and the reference light after ultrasonic excitation to calculate the deformation of the intestinal tissue; modeling the force on the intestinal tissue and constructing a theoretical displacement function of the intestinal tissue; obtaining the shear modulus by minimizing the difference between the deformation at the same position of the intestinal tissue and the theoretical displacement of the intestinal tissue, and then reconstructing the OCE image data of the intestinal tissue.

Citation Information

Patent Citations

  • Multi-modal imaging system

    CN116530935A

  • Multi-modal imaging and detection system for oral cavity and tongue

    CN118697294A

  • Esophageal squamous carcinoma detection method and system based on combination of multi-modal information fusion and artificial intelligence

    CN119399126A