Multi-modal multi-view-angle endoscopic image segmentation method, device and equipment and medium
By extracting and fusing the modal features and viewing angle features of narrowband light and visible light images, and performing spatial registration and generating joint features, the problems of modes not registering and viewing angle changes in laminoscopic image segmentation are solved, and segmentation accuracy and robustness are improved.
Patent Information
- Application Number
- CN202510089745.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-21
AI Technical Summary
In laminoscope image segmentation, there is a spatial misalignment problem for images of two modes, narrowband light and visible light, which leads to the inability of traditional segmentation methods to effectively utilize the complementary information of the two modes, affecting the segmentation accuracy. In addition, the feature distribution of multi-view images varies greatly, and existing methods are difficult to cope with the problems caused by changes in perspective.
A multimodal multi-view laminoscope image segmentation method is proposed. By obtaining narrow-band light and visible light images at different viewing angles, modal features and viewing angle features are extracted, spatial registration and feature fusion are performed, and joint features are generated. Deep learning-based segmentation networks use these features for segmentation, improving the accuracy of segmentation.
Through spatial registration and feature fusion, the problem of spatial non-registration between modes of narrowband light and visible light images is solved, and the segmentation accuracy of multimodal images is improved. The adaptive feature learning mechanism enhances the adaptability to multi-view scenes, effectively deals with the impact of viewing angle changes on segmentation results, and improves the detection accuracy of endoscopic image lesion areas.
Smart Images

Figure CN119991621A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a multi-modal and multi-view endoscopic image segmentation method, device, equipment and medium. Background Art
[0002] Endoscopic image segmentation aims to accurately mark the lesion area in an automated way. However, since the endoscope lens often involves different modalities (such as narrowband light NBI and visible light WLI) and images of different perspectives during the acquisition process, the segmentation task faces many challenges. Currently, narrowband light (NBI) and visible light (WLI) are widely used in medical endoscopy. NBI images can enhance tissue details, especially vascular structures, while WLI images provide a more comprehensive field of view and background information. However, due to differences in acquisition equipment and changes in lens perspective, the images of these two modalities usually have spatial misalignment problems, resulting in the inability of traditional segmentation methods to effectively utilize the complementary information of the two modalities, thereby affecting the accuracy of segmentation. In addition, since endoscopic images may produce complex geometric deformations due to changes in lens angle, lighting conditions and tissue surface morphology during the acquisition process, existing segmentation methods are difficult to deal with the above problems. Summary of the invention
[0003] The main purpose of the embodiments of the present application is to propose a multi-modal and multi-view endoscopic image segmentation method, device, equipment and medium to improve the segmentation accuracy of endoscopic images.
[0004] To achieve the above object, an embodiment of the present application provides a multi-modal multi-view endoscopic image segmentation method, which comprises the following steps:
[0005] Acquire narrow-band light images and visible light images of the lesion area captured by the endoscope at different viewing angles;
[0006] extracting modal features of the narrowband light image and the visible light image respectively, and then fusing the modal features to obtain multimodal fusion features;
[0007] Performing spatial registration on the narrowband light image and the visible light image to respectively extract viewing angle features of the narrowband light image and the visible light image, and then fusing the viewing angle features to obtain multi-view fusion features;
[0008] Completing, registering and fusing the multimodal fusion features and the multi-view fusion features to obtain a joint feature;
[0009] The deep learning-based segmentation network obtains a first segmentation area according to the modal features and the joint features; the deep learning-based segmentation network obtains a second segmentation area according to the viewing angle features and the joint features.
[0010] In some embodiments, respectively extracting the modal features of the narrow-band light image and the visible light image comprises the following steps:
[0011] Using a feature extraction network to extract the texture of the lesion area in the narrow-band light image and the visible light image as local detail features;
[0012] Using the feature extraction network to extract the shapes of the lesion areas in the narrow-band light image and the visible light image as global context features;
[0013] The feature extraction network is used to extract the edges of the lesion area in the narrow-band light image and the visible light image as multi-scale features;
[0014] Among them, the local detail features, the global context features and the multi-scale features serve as the modal features.
[0015] In some embodiments, the fusing the modal features to obtain multimodal fusion features comprises the following steps:
[0016] Using a cross attention mechanism, the attention weights of the local detail feature, the global context feature, and the multi-scale feature are calculated respectively;
[0017] The local detail features, the global context features and the multi-scale features are weightedly fused according to the corresponding attention weights to obtain the multimodal fusion features; wherein the multimodal fusion features include the vascular detail information of the narrow-band light image and the overall structural information of the visible light image.
[0018] In some embodiments, spatially registering the narrow-band light image and the visible light image to respectively extract viewing angle features of the narrow-band light image and the visible light image comprises the following steps:
[0019] Calculating a spatial transformation relationship between the narrow-band light image and the visible light image by using a deformation field model or a geometric transformation model based on deep learning;
[0020] The narrow-band light image and the visible light image are mapped to a unified spatial coordinate system according to the spatial transformation relationship, and then the registered image features after the spatial transformation are output as the viewing angle features.
[0021] In some embodiments, the fusing the viewing angle features to obtain multi-view fusion features comprises the following steps:
[0022] The perspective features are jointly modeled through three-dimensional position encoding or perspective self-attention mechanism to eliminate perspective differences and then output the multi-perspective fusion features.
[0023] In some embodiments, the multimodal fusion feature and the multi-view fusion feature are complemented, registered and fused to obtain a joint feature, comprising the following steps:
[0024] For the missing areas in the multimodal fusion features or the multi-view fusion features, dynamically complete the missing areas by using the complete information of the multi-view fusion features or the multimodal fusion features through saliency analysis and non-local attention mechanism;
[0025] Using a deep learning deformation field model to perform spatial registration on the multimodal fusion features and the multi-view fusion features;
[0026] The multimodal fusion features that have undergone dynamic completion and spatial registration and the multi-view fusion are fused through feature weighting and cross-attention mechanism to obtain the joint features.
[0027] In some embodiments, the method further comprises the following steps:
[0028] Calculate a segmentation loss function according to the pixel-level annotation data of the lesion area, the first segmented area, and the second segmented area; wherein the segmentation loss function includes a cross entropy loss, a Dice loss, and a boundary smoothing loss;
[0029] Optimizing the parameters of the multimodal multi-view endoscopic image segmentation model using the AdamW optimizer according to the segmentation loss function and the prediction loss; wherein the multimodal multi-view endoscopic image segmentation model is used to segment the narrow-band light image and the visible light image to obtain the first segmented area and the second segmented area;
[0030] Updating the optimized parameters to the multi-modal multi-view endoscopic image segmentation model through a back-propagation algorithm;
[0031] Setting a dynamically decaying learning rate to accelerate the convergence of the multi-modal multi-view endoscopic image segmentation model;
[0032] Determining whether the multimodal multi-view endoscopic image segmentation model meets the convergence condition through the performance index of the validation set; wherein the convergence condition includes that the rate of change of the segmentation loss function is less than a first preset threshold, or that the segmentation accuracy of the first segmented area and the second segmented area reaches a second preset threshold;
[0033] If the multi-modal multi-view endoscopic image segmentation model satisfies the convergence condition, saving the trained multi-modal multi-view endoscopic image segmentation model;
[0034] If the multimodal multi-view endoscopic image segmentation model does not meet the convergence condition, return to the step of optimizing the parameters of the multimodal multi-view endoscopic image segmentation model using the AdamW optimizer according to the segmentation loss function and the prediction loss until the convergence condition is met.
[0035] To achieve the above object, another aspect of the embodiment of the present application provides a multi-modal multi-view endoscopic image segmentation device, the device comprising:
[0036] An image acquisition unit, used to acquire narrow-band light images and visible light images of the lesion area photographed by the endoscope at different viewing angles;
[0037] A modal feature processing unit, used to extract the modal features of the narrow-band light image and the visible light image respectively, and then fuse the modal features to obtain a multi-modal fusion feature;
[0038] A viewing angle feature processing unit, configured to perform spatial registration on the narrow-band light image and the visible light image, thereby respectively extracting viewing angle features of the narrow-band light image and the visible light image, and then fusing the viewing angle features to obtain a multi-view fusion feature;
[0039] A feature combining unit, used for completing, registering and fusing the multi-modal fusion feature and the multi-view fusion feature to obtain a combined feature;
[0040] An image segmentation unit is used to obtain a first segmentation area based on the modal feature and the joint feature segmentation by a segmentation network based on deep learning; and obtain a second segmentation area based on the viewing angle feature and the joint feature segmentation by the segmentation network based on deep learning.
[0041] To achieve the above-mentioned purpose, another aspect of an embodiment of the present application proposes an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the above-mentioned multi-modal and multi-view endoscopic image segmentation method when executing the computer program.
[0042] To achieve the above-mentioned purpose, another aspect of an embodiment of the present application proposes a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the multi-modal and multi-view endoscopic image segmentation method is implemented.
[0043] The embodiments of the present application include at least the following beneficial effects:
[0044] The present application can obtain narrowband light images and visible light images taken by an endoscope at different viewing angles of the lesion area; extract the modal features of the narrowband light image and the visible light image respectively, and then fuse the modal features to obtain multimodal fusion features; spatially register the narrowband light image and the visible light image and then extract the viewing angle features of the narrowband light image and the visible light image respectively, and then fuse the viewing angle features to obtain multi-view fusion features; complete, register and fuse the multimodal fusion features and the multi-view fusion features to obtain joint features; a segmentation network based on deep learning obtains a first segmentation area according to the modal features and the joint features; a segmentation network based on deep learning obtains a second segmentation area according to the viewing angle features and the joint features. The present application solves the problem of spatial misalignment between the narrowband light image and the visible light image modalities through spatial registration and feature fusion, and improves the segmentation accuracy of the multimodal image; the adaptive feature learning mechanism of the present application enhances the adaptability to multi-view scenes and effectively copes with the impact of viewing angle changes on the segmentation results; the multimodal multi-view segmentation network combined with deep learning effectively improves the detection accuracy of the lesion area of the endoscopic image. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0046] Figure 1 A schematic diagram of a flow chart of a multi-modal and multi-view endoscopic image segmentation method provided in an embodiment of the present application;
[0047] Figure 2 An example flow chart of a multi-modal and multi-view endoscopic image segmentation method provided in an embodiment of the present application;
[0048] Figure 3 An example scene diagram of dual-modal dual-view endoscopic image segmentation provided in an embodiment of the present application;
[0049] Figure 4 A schematic diagram of the structure of a multi-modal and multi-view endoscopic image segmentation device provided in an embodiment of the present application;
[0050] Figure 5 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0051] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below in conjunction with the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are only examples of devices and methods consistent with some aspects of the embodiments of the present application as detailed in the attached claims.
[0052] It is understood that the terms "first", "second", etc. used in this application can be used to describe various concepts in this article, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another concept. For example, without departing from the scope of the embodiment of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein can be interpreted as "at the time of" or "when" or "in response to determination".
[0053] The terms "at least one", "multiple", "each", "any", etc. used in this application, at least one includes one, two or more, multiple includes two or more, each refers to each of the corresponding multiple, and any refers to any one of the multiple.
[0054] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0055] In order to solve the problem of inaccurate endoscopic image segmentation, existing studies have attempted to improve the segmentation performance of multimodal endoscopic images through image registration and feature fusion technology. For example, some methods achieve modality registration through geometric transformation, and use simple fusion strategies (such as feature splicing or weighted summation) to jointly model the features of multimodal images. However, these methods usually ignore the deep complementary information between the modalities and cannot give full play to the synergy of the two modalities in the segmentation task. In addition, the feature distribution of multi-view endoscopic images is quite different, and the existing methods are difficult to effectively deal with the problems caused by the change of perspective, resulting in poor generalization ability of the model in cross-view scenarios. In view of the shortcomings of the prior art, some embodiments of the present application propose a multimodal multi-view endoscopic image segmentation method based on deep learning. The method combines spatial registration, feature fusion and adaptive feature learning technologies to solve the problem of spatial misalignment between NBI and WLI images and inconsistent feature distribution of multi-view images. Specifically, some embodiments of the present application perform weighted fusion of multimodal features through a cross-attention mechanism to fully explore the complementary information of the two modal images, thereby enhancing the detection capability of the lesion area. At the same time, the deformation field model or geometric transformation model of deep learning is used to complete the spatial registration of multi-view images and solve the problem of feature inconsistency caused by changes in perspective. In addition, the adaptive feature learning mechanism effectively handles the differences in feature distribution by dynamically adjusting network parameters, thereby improving the generalization ability of the segmentation model. This application comprehensively solves the core problems in multi-modal and multi-view endoscopic image segmentation through multiple technologies such as multimodal fusion, multi-view spatial registration, and adaptive feature learning, providing efficient and reliable technical support for lesion detection, and has broad application prospects.
[0056] The embodiments of the present application provide a multi-modal and multi-view endoscopic image segmentation method, device, equipment and medium. The technical solution of the present application includes: obtaining narrowband light images and visible light images taken by the endoscope at different viewing angles of the lesion area; extracting the modal features of the narrowband light image and the visible light image respectively, and then fusing the modal features to obtain multi-modal fusion features; spatially registering the narrowband light image and the visible light image and then extracting the viewing angle features of the narrowband light image and the visible light image respectively, and then fusing the viewing angle features to obtain multi-view fusion features; completing, registering and fusing the multi-modal fusion features and the multi-view fusion features to obtain joint features; a segmentation network based on deep learning obtains a first segmentation area according to the modal features and the joint features; a segmentation network based on deep learning obtains a second segmentation area according to the viewing angle features and the joint features. This application solves the problem of spatial misalignment between narrowband light image and visible light image modalities through spatial registration and feature fusion, and improves the segmentation accuracy of multimodal images; the adaptive feature learning mechanism of this application enhances the adaptability to multi-view scenes and effectively copes with the impact of view changes on segmentation results; the multimodal multi-view segmentation network combined with deep learning effectively improves the detection accuracy of lesion areas in endoscopic images.
[0057] The embodiments of the present application provide a multimodal multi-view endoscopic image segmentation method, device, equipment and medium, which relate to the field of artificial intelligence technology. The multimodal multi-view endoscopic image segmentation method, device, equipment and medium provided in the embodiments of the present application can be applied to a terminal, can also be applied to a server, and can also be software running in a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, and a car terminal, etc., but is not limited to this; the server side can be configured as an independent physical server, or it can be configured as a server cluster or distributed system composed of multiple physical servers, and can also be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application that implements a knowledge extraction method, etc., but is not limited to the above forms.
[0058] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0059] Reference Figure 1 The embodiment of the present application provides a multi-modal and multi-view endoscopic image segmentation method, which may include but is not limited to S100 to S140, as follows:
[0060] S100: Acquire narrow-band light images and visible light images of the lesion area captured by the endoscope at different viewing angles.
[0061] S110: extracting modal features of the narrow-band light image and the visible light image respectively, and then fusing the modal features to obtain multi-modal fusion features.
[0062] Furthermore, respectively extracting the modal features of the narrow-band light image and the visible light image in S110 may include the following steps S111 to S113:
[0063] S111: using a feature extraction network to respectively extract textures of the lesion area in the narrow-band light image and the visible light image as local detail features;
[0064] S112: using the feature extraction network to respectively extract the shapes of the lesion areas in the narrow-band light image and the visible light image as global context features;
[0065] S113: using the feature extraction network to respectively extract edges of the lesion area in the narrow-band light image and the visible light image as multi-scale features;
[0066] Among them, the local detail features, the global context features and the multi-scale features serve as the modal features.
[0067] More specifically, fusing the modal features in S110 to obtain multimodal fusion features includes the following steps S114 to S115:
[0068] S114: Calculate the attention weights of the local detail feature, the global context feature, and the multi-scale feature respectively using a cross attention mechanism;
[0069] S115: Performing weighted fusion on the local detail features, the global context features and the multi-scale features according to the corresponding attention weights to obtain the multimodal fusion features; wherein the multimodal fusion features include the vascular detail information of the narrow-band light image and the overall structural information of the visible light image.
[0070] S120: spatially registering the narrow-band light image and the visible light image to extract viewing angle features of the narrow-band light image and the visible light image respectively, and then fusing the viewing angle features to obtain multi-view fusion features.
[0071] Furthermore, the spatial registration of the narrow-band light image and the visible light image in S120 to respectively extract the viewing angle features of the narrow-band light image and the visible light image includes the following steps S121-S122:
[0072] S121: Calculating a spatial transformation relationship between the narrow-band light image and the visible light image by using a deformation field model or a geometric transformation model based on deep learning;
[0073] S122: Mapping the narrow-band light image and the visible light image to a unified spatial coordinate system according to the spatial transformation relationship, and then outputting spatially transformed registration image features as the viewing angle features.
[0074] Furthermore, fusing the viewing angle features in S120 to obtain multi-view fusion features includes the following steps S123:
[0075] S123: Jointly modeling the view features through three-dimensional position encoding or view self-attention mechanism to eliminate view differences, and then output the multi-view fusion features.
[0076] S130: Completing, registering and fusing the multimodal fusion features and the multi-view fusion features to obtain joint features.
[0077] Further, S130 may include the following steps S131 to S133:
[0078] S131: dynamically completing the missing area in the multimodal fusion feature or the multi-view fusion feature by using the complete information of the multi-view fusion feature or the multimodal fusion feature through saliency analysis and a non-local attention mechanism;
[0079] S132: Using a deep learning deformation field model to perform spatial registration on the multimodal fusion feature and the multi-view fusion feature;
[0080] S133: The multimodal fusion features after dynamic completion and spatial registration and the multi-view fusion are fused through feature weighting and cross-attention mechanism to obtain the joint features.
[0081] S140: The segmentation network based on deep learning obtains a first segmentation area according to the modal feature and the joint feature segmentation; the segmentation network based on deep learning obtains a second segmentation area according to the view feature and the joint feature segmentation.
[0082] In some optional embodiments, the present application may use a multimodal multi-view endoscopic image segmentation model to perform S110 to S140, that is, input the narrowband light image and the visible light image into the multimodal multi-view endoscopic image segmentation model, and use the multimodal multi-view endoscopic image segmentation model to output the first segmentation area and the second segmentation area. To improve the segmentation accuracy, the embodiment of the present application may also include the step of training the multimodal multi-view endoscopic image segmentation model, which may specifically include the following steps S151 to S157:
[0083] S151: Calculating a segmentation loss function according to the pixel-level annotation data of the lesion area, the first segmented area, and the second segmented area; wherein the segmentation loss function includes a cross entropy loss, a Dice loss, and a boundary smoothing loss;
[0084] S152: Optimizing the parameters of a multi-modal multi-view endoscopic image segmentation model using an AdamW optimizer according to the segmentation loss function and the prediction loss; wherein the multi-modal multi-view endoscopic image segmentation model is used to segment the narrow-band light image and the visible light image to obtain the first segmented region and the second segmented region;
[0085] S153: updating the optimized parameters to the multi-modal multi-view endoscopic image segmentation model through a back propagation algorithm;
[0086] S154: Setting a dynamically decaying learning rate to accelerate the convergence of the multi-modal multi-view endoscopic image segmentation model;
[0087] S155: judging whether the multi-modal multi-view endoscopic image segmentation model meets the convergence condition through the performance index of the validation set; wherein the convergence condition includes that the rate of change of the segmentation loss function is less than a first preset threshold, or that the segmentation accuracy of the first segmented area and the second segmented area reaches a second preset threshold;
[0088] S156: If the multi-modal multi-view endoscopic image segmentation model satisfies the convergence condition, saving the trained multi-modal multi-view endoscopic image segmentation model;
[0089] S157: If the multimodal multi-view endoscopic image segmentation model does not meet the convergence condition, return to the step of optimizing the parameters of the multimodal multi-view endoscopic image segmentation model using the AdamW optimizer according to the segmentation loss function and the prediction loss until the convergence condition is met.
[0090] Next, the solution of the embodiment of the present application will be introduced and explained in detail with reference to specific application examples.
[0091] This embodiment relates to the field of artificial intelligence and medical image processing, and more specifically, to a multi-modal and multi-view segmentation method for endoscopic images based on deep learning. The method is mainly used for the automatic segmentation of endoscopic images, especially for the spatial registration, feature fusion and automatic detection and positioning of lesion areas of narrow band light (NBI) and visible light (WLI) modal images.
[0092] This embodiment aims to solve the problems of spatial misalignment of narrowband light (NBI) and visible light (WLI) endoscopic images, large differences in multi-view feature distribution, and insufficient robustness of existing segmentation methods in complex scenes, and proposes a multi-modal multi-view endoscopic image segmentation method based on deep learning. This method combines spatial registration, feature fusion, and adaptive feature learning technology to achieve high-precision segmentation of multi-modal multi-view endoscopic images, solving the core problem of existing methods in lesion area detection.
[0093] Reference Figure 2 , Figure 2 This is an example flow chart of a multi-modal and multi-view endoscopic image segmentation method provided in this embodiment. Figure 3 This is an example scene diagram of dual-modal dual-view endoscopic image segmentation in this embodiment.
[0094] It can be understood that the multi-modal and multi-view endoscopic image segmentation model trained in this embodiment simultaneously inputs images of two modalities and perspectives, and outputs lesion segmentation maps corresponding to the two images after inference by the deep network model.
[0095] Exemplarily, the technical solution of this embodiment includes the following steps:
[0096] Step 1: Input narrow-band light (NBI) and visible light (WLI) dual-modality multi-view images, and preprocess the input images, including image size normalization, noise removal, contrast enhancement, and color space conversion to ensure image quality and consistency.
[0097] Specifically, the input image is a dual-modal dual-view image taken by an endoscope (or endoscope) device at the lesion in the human body. Because two modal light sources (visible white light WLI and narrow-band light source NBI) are used during the operation of the endoscope. The two light source environments cannot exist at the same time because shooting the two lights at the same time will cause the light environment to merge because narrow-band light is part of the visible spectrum. The spatial positions of the two images are also mismatched, on the one hand because the position of the handheld endoscope changes, and on the other hand because the movement of the human muscle tissue causes the position of the lesion to change.
[0098] Furthermore, the spatial registration of NBI and WLI images is achieved by combining a convolutional neural network (CNN) with a geometric transformation module to eliminate the spatial inconsistency problem caused by changes in the viewing angle of the imaging device.
[0099] Step 2: Use a feature extraction network (such as a convolutional neural network (CNN) or a Transformer network) to extract features from images of the NBI and WLI modalities, respectively. The extracted features include local detail features, global context features, and multi-scale features, which capture the texture, shape, edge, and other information of the lesion area.
[0100] Specifically, through the implementation of the CNN-based feature extraction module, local detail features of images from different perspectives such as texture, color, edge, shape, etc. are extracted. The channel features and spatial features between modalities are jointly modeled through the cross-attention mechanism to improve the efficiency of the segmentation network in utilizing complementary information between modalities.
[0101] The feature extraction module in step 2 can abstract low-dimensional features into high-dimensional context information encoding through deep neural network reasoning, and obtain multi-scale features through deep multi-level feature extraction.
[0102] Step 3: Use the cross-attention mechanism to perform weighted fusion on the features extracted from NBI and WLI images. In the cross-attention mechanism, the attention weights of different modal features are calculated through adaptive learning to highlight the significant features of the lesion area. The fused features can simultaneously retain the vascular detail information of the NBI image and the overall structural information of the WLI image, thereby improving the segmentation performance.
[0103] Specifically, the attention mechanism in step 3 adaptively fuses the low-dimensional detail features and high-dimensional context features extracted from different modalities through a cross-fusion mechanism, complements the information of the two modal images, and improves the segmentation performance of the model.
[0104] Step 4: Use the spatial registration module to register the NBI and WLI images to solve the spatial inconsistency problem between the two modal images caused by changes in viewing angles. During the spatial registration process, the spatial transformation relationship between the two modalities is calculated through a deformation field model or geometric transformation model based on deep learning, and the NBI and WLI images are mapped to a unified spatial coordinate system, and the registered image features after spatial transformation are output.
[0105] Specifically, in step 4, the image registration module based on CNN and fully connected network realizes the spatial transformation features after dimensionality reduction of the reasoning of the spatial transformation relationship of images with different perspectives, maps the images of different perspective modalities to a unified spatial coordinate system through the mapping relationship between the local targets of the two images, and calculates the image features after spatial transformation through the registration of three-dimensional position coding reasoning.
[0106] In step 4, the adaptive feature learning mechanism dynamically adjusts network parameters to address the differences in feature distribution of endoscopic images under different viewing angles, thereby enhancing the model's ability to handle inconsistent feature distribution.
[0107] Step 5: The multi-view feature fusion module adaptively fuses image features from different viewpoints and jointly models multi-view features by introducing three-dimensional position encoding or viewpoint self-attention mechanism, effectively eliminating the influence of viewpoint differences on feature expression and outputting a fused feature representation that is robust to viewpoint changes.
[0108] Specifically, the multi-view feature fusion module described in step 5 adaptively fuses the image features obtained after transforming images of different viewpoints, introduces three-dimensional position coding into the attention fusion mechanism, jointly models the multi-view features and the image features themselves, minimizes the difference features caused by viewpoint differences, and outputs the viewpoint fusion feature representation.
[0109] Step 6: Input the multi-modal and multi-view fused features into the adaptive learning module, and further optimize the feature representation through feature completion, registration, and deep fusion. For the missing areas that may exist in the features, the complete information of another modality or perspective is used for dynamic completion through saliency analysis and non-local attention mechanism. At the same time, the deformation field model of deep learning is used to spatially register the multi-modal and multi-view features to solve the spatial inconsistency problem. Finally, deep fusion is achieved through feature weighting and cross-attention mechanism to generate a joint feature representation with both spatial consistency and rich semantic expression, providing high-quality input for the segmentation task.
[0110] Specifically, the adaptive learning module in step 6 adaptively learns the fused features based on the prior knowledge of endoscopic image recognition, and dynamically adjusts the prior feature parameters through additional optimization strategies. The feature selection mechanism can automatically select highly relevant features.
[0111] Step 7: Input the feature layer after adaptive learning optimization, predict the lesion area of the two input images respectively, and the segmentation network adopts a deep learning-based segmentation architecture (such as UNet, DeepLab or Transformer-based segmentation network) to accurately segment the lesion area in the image and output the lesion segmentation result of each image.
[0112] Specifically, step 7 can use the feature layer obtained in the previous step to predict the range of the lesion under two viewing angles, and the segmentation network uses a CNN-based deep network to perform pixel-level segmentation on the image.
[0113] In step 7, a comprehensive loss function is designed, including segmentation loss, bimodal image supervision loss, and feature consistency loss, to guide the parameter update of the model and ensure the segmentation accuracy of the lesion area.
[0114] Step 8: According to the pixel-level annotation data and segmentation prediction results of the lesion area, the segmentation loss function is calculated, and the loss function includes cross entropy loss, Dice loss and boundary smoothing loss; the network parameters are optimized using the AdamW optimizer, and the network weights are updated through the back propagation algorithm; the learning rate decay mechanism is set to dynamically adjust the learning rate to accelerate model convergence.
[0115] Specifically, step 8 uses a variety of loss functions for endoscopic image lesion tasks to dynamically optimize the parameters of the model.
[0116] Step 9: Determine whether the model has converged based on the performance indicators of the validation set. The convergence conditions include that the segmentation loss function no longer decreases significantly or the segmentation accuracy reaches the preset threshold. If the model converges, save the trained segmentation model for subsequent use. If not, return to continue optimizing the network parameters until the model meets the convergence conditions.
[0117] In summary, this embodiment provides a multi-modal and multi-view segmentation method for endoscopic images based on deep learning, which aims to solve the problem of spatial misalignment between narrowband light (NBI) and visible light (WLI) modal images, as well as the large difference in the distribution of multi-view image features, and improve the accuracy of endoscopic image segmentation. Endoscopic image segmentation has important application value in medical image analysis, especially in the automatic detection and positioning of lesion areas. However, due to the change in the viewing angle of the endoscopic lens and the spatial difference in multi-modal imaging, the traditional segmentation method cannot effectively combine multi-modal and multi-view information, affecting the segmentation performance. In order to solve these challenges, this embodiment proposes a segmentation framework for multi-modal and multi-view endoscopic images. First, the multi-modal feature fusion of NBI and WLI images is realized through a deep learning algorithm. The feature fusion module uses the cross-attention mechanism to fully mine the complementary information of the two modalities, retaining the microscopic vascular details in the NBI image, and combining the global tissue structure information of the WLI image, significantly improving the overall performance of the segmentation task. Secondly, this embodiment designs a multi-view spatial feature registration and fusion technology. Through the deformation field model or geometric transformation model based on deep learning, the image features under different perspectives are spatially aligned to solve the spatial inconsistency problem caused by the change of lens perspective. Further, the features of different perspectives are adaptively fused by using three-dimensional position encoding and perspective self-attention mechanism to generate a joint feature representation that is robust to perspective changes. In addition, this embodiment proposes a strategy for complementing and fusion of multimodal features and multi-perspective features. In view of the problem of incomplete information of the lesion area caused by perspective occlusion or incomplete modal information, the missing information is complemented by the complete features in another perspective or modality. Through saliency analysis and non-local attention mechanism, dynamic complementation between multimodal and multi-perspective features is achieved to further improve the integrity of the lesion area. Subsequently, the feature weighting mechanism and deep fusion network are used to jointly model the complemented multimodal and multi-perspective features to generate high-quality segmentation feature representations. The innovation of this embodiment is that the core problem in endoscopic image segmentation is solved through multimodal fusion, multi-perspective spatial feature registration and fusion, and dynamic complementation and deep fusion of multimodal and multi-perspective features. This method has high accuracy and robustness in endoscopic image segmentation tasks, provides strong technical support for lesion detection, and has broad application prospects.
[0118] This embodiment can also provide a more specific implementation method, as follows:
[0119] This embodiment proposes a multimodal and multi-view endoscopic image segmentation method, which aims to solve the problems of multimodal image spatial misalignment, inconsistent multi-view feature distribution, and insufficient lesion area detection accuracy in existing endoscopic image segmentation methods. To this end, this embodiment designs a multimodal and multi-view segmentation framework based on deep learning, which makes full use of the complementary information of multimodal images, and adaptively optimizes the distribution characteristics of multi-view images to improve segmentation accuracy and robustness. This embodiment includes the following steps:
[0120] Step 1: Input narrow-band light (NBI) and visible light (WLI) dual-modality multi-view images, and preprocess the input images, including image size normalization, noise removal, contrast enhancement, and color space conversion to ensure image quality and consistency.
[0121] Step 2: Use a feature extraction network (such as a convolutional neural network (CNN) or a Transformer network) to extract features from images of the NBI and WLI modalities, respectively. The extracted features include local detail features, global context features, and multi-scale features, which capture the texture, shape, edge, and other information of the lesion area.
[0122] Step 3: Use the cross-attention mechanism to perform weighted fusion on the features extracted from NBI and WLI images. In the cross-attention mechanism, the attention weights of different modal features are calculated through adaptive learning to highlight the significant features of the lesion area. The fused features can simultaneously retain the vascular detail information of the NBI image and the overall structural information of the WLI image, thereby improving the segmentation performance.
[0123] Step 4: Use the spatial registration module to register the NBI and WLI images to solve the spatial inconsistency problem between the two modal images caused by changes in viewing angles. During the spatial registration process, the spatial transformation relationship between the two modalities is calculated through a deformation field model or geometric transformation model based on deep learning, and the NBI and WLI images are mapped to a unified spatial coordinate system, and the registered image features after spatial transformation are output.
[0124] Step 5: The multi-view feature fusion module adaptively fuses image features from different viewpoints and jointly models multi-view features by introducing three-dimensional position encoding or viewpoint self-attention mechanism, effectively eliminating the influence of viewpoint differences on feature expression and outputting a fused feature representation that is robust to viewpoint changes.
[0125] Step 6: Input the multi-modal and multi-view fused features into the adaptive learning module, and further optimize the feature representation through feature completion, registration, and deep fusion. For the missing areas that may exist in the features, the complete information of another modality or perspective is used for dynamic completion through saliency analysis and non-local attention mechanism. At the same time, the deformation field model of deep learning is used to spatially register the multi-modal and multi-view features to solve the spatial inconsistency problem. Finally, deep fusion is achieved through feature weighting and cross-attention mechanism to generate a joint feature representation with both spatial consistency and rich semantic expression, providing high-quality input for the segmentation task.
[0126] Step 7: Input the feature layer after adaptive learning optimization, predict the lesion area of the two input images respectively, and the segmentation network adopts a deep learning-based segmentation architecture (such as UNet, DeepLab or Transformer-based segmentation network) to accurately segment the lesion area in the image and output the lesion segmentation result of each image.
[0127] Step 8: According to the pixel-level annotation data and segmentation prediction results of the lesion area, the segmentation loss function is calculated, and the loss function includes cross entropy loss, Dice loss and boundary smoothing loss; the network parameters are optimized using the AdamW optimizer, and the network weights are updated through the back propagation algorithm; the learning rate decay mechanism is set to dynamically adjust the learning rate to accelerate model convergence.
[0128] Step 9: Determine whether the model has converged based on the performance indicators of the validation set. The convergence conditions include that the segmentation loss function no longer decreases significantly or the segmentation accuracy reaches the preset threshold. If the model converges, save the trained segmentation model for subsequent use. If not, return to continue optimizing the network parameters until the model meets the convergence conditions.
[0129] This embodiment designs a dual-modality dual-view endoscopic image lesion area segmentation method, which can still be referred to Figure 2 In the training phase, the method of this embodiment first uses paired dual-modal dual-view endoscopic images, which are images of the same lesion area taken under different light environments. In order to improve the generalization performance of the neural network, the method performs image enhancement of the same type and parameters on the two images in the input phase, such as rotation, cropping, flipping, and adding noise.
[0130] In step 2, the multimodal image feature extraction module first performs CNN-based feature extraction on the image. The two groups of modules have the same network structure but do not share parameters. This is to extract image features from imaging structures of different modalities to adapt to the characteristic differences between NBI and WLI images. The extracted features include low-level features of the image such as edges, textures, etc. and high-level features such as regional shape and semantic information. In the process of extracting WLI modal features, spatial features will be fused into the variables of the next layer with the reasoning of the convolution operation. In the process of extracting NBI modal features, the vascular pattern features that are highly correlated with lesion features will be mapped to the variables of the next layer in the extraction of channel features.
[0131] The feature fusion process in step 3 aims to combine the complementary information of NBI and WLI modalities to generate more discriminative joint features. First, the features of the two modalities are unified through 1×1 convolution to ensure the consistency of feature dimensions. Then, the cross-attention mechanism is used to capture the correlation between the two modalities respectively, and the important features of each modality are enhanced through attention weights to highlight the key areas. Subsequently, the updated modality features are concatenated in the channel dimension and further fused through convolutional layers to extract meaningful contextual information in the joint features. The final generated fusion feature fully combines the detailed information of the NBI modality (such as blood vessels, lesion contrast) and the global background information of the WLI modality (such as tissue shape and structure), providing high-quality input for subsequent segmentation tasks.
[0132] In step 4, spatial feature extraction is performed through multi-scale convolution operations, using small convolution kernels to capture local fine texture information (such as edges and small lesions), while using large convolution kernels to extract global structural information (such as the overall shape and distribution of lesions), thereby ensuring that details are retained and the overall layout is understood. In order to further enhance the expression of key areas, a spatial attention mechanism is introduced to extract global information of salient areas through global pooling, generate a spatial weight map, focus on salient locations related to recognition such as lesions or blood vessels, and effectively suppress background interference. In addition, a spatial transformer network is used to spatially adjust the feature map of each viewpoint, and the feature position is corrected by learning affine transformations (such as displacement, rotation, scaling, etc.), reducing the interference of viewpoint changes on feature extraction and making the features more standardized. In order to capture the correlation between distant spatial positions, a non-local module is used to model spatial contextual relationships, and the semantic association between the lesion area and the surrounding tissue is enhanced by calculating the correlation between any two spatial positions. Finally, a dynamic weight allocation mechanism generates an independent weight map for the spatial features of each viewpoint, highlighting the recognition importance of the same area in different viewpoints, and combining the differences in viewpoint features to ensure that the extracted spatial features have a high degree of expression. These processing steps together generate a spatial feature map with detailed accuracy and global understanding capabilities, laying a solid foundation for subsequent view fusion and lesion segmentation.
[0133] In step 6, in the process of completing, registering and finally fusing the bimodal fusion features with the dual-view fusion features, the completion mechanism is first used to solve the problem of incomplete information in the lesion area caused by the difference in view angle. For some lesion areas that may be incomplete in a single view due to occlusion or change in view angle, more complete features in another view are used to compensate. Through saliency analysis and regional difference detection, the missing position of the lesion area is identified, and the lesion features of another view are mapped to the current view using a non-local attention mechanism or feature interpolation method to dynamically fill in the missing features. At the same time, the bimodal fusion features (such as microscopic details in NBI or global semantics in WLI) are used to assist in completion to further improve the representation of the lesion area. After completion, in order to eliminate the spatial inconsistency between the bimodal and dual-view features, the spatial registration technology is used to align the lesion area through key point detection and affine transformation, and the consistency loss function is used to optimize the spatial consistency and semantic expression of the features. On this basis, the features after completion and registration are further integrated through channel splicing and fusion networks, and the key features of the lesion area are highlighted in combination with the attention weighting mechanism, while suppressing redundant information. Finally, the fused feature map has both the detail advantages of dual modalities and the completion capability of dual perspectives, which not only ensures the integrity of the lesion area, but also achieves the unity of spatial consistency and semantic expression, providing high-quality input features for subsequent segmentation tasks.
[0134] In step 8, during the training of the segmentation network, the pixel-level annotation data and prediction results of the lesion area are first used to calculate the comprehensive segmentation loss function, which is composed of cross entropy loss, Dice loss and boundary smoothing loss. Cross entropy loss focuses on the accuracy of global pixel classification, Dice loss focuses on the overlap of lesion areas and solves the segmentation problem of small target areas, while boundary smoothing loss optimizes the edge quality of the segmentation result by constraining the gradient difference between the predicted and annotated boundaries. These three losses are weighted to form a comprehensive loss function, which optimizes the global classification performance, regional overlap and boundary detail performance respectively. Then, the AdamW optimizer is used to optimize the network parameters. The optimizer combines momentum update and weight decay mechanism, which can not only converge quickly, but also effectively prevent overfitting through regularization. Through the back propagation algorithm, the network weights are updated layer by layer according to the gradient calculated by the segmentation loss function, so that the model parameters continue to approach the optimal solution. In addition, in order to further improve the training efficiency and stability, the learning rate decay mechanism is adopted to dynamically adjust the learning rate according to the progress of training. A higher learning rate is set in the early stage to accelerate convergence, and the learning rate is gradually reduced in the later stage through strategies such as exponential decay, cosine annealing, or validation set performance monitoring to avoid model oscillation in the later stage of training. This comprehensive training strategy ensures that the model can achieve a balance between global accuracy, local area overlap, and boundary detail optimization, while improving the accuracy and robustness of the segmentation results.
[0135] In summary, the beneficial effects of this embodiment include:
[0136] This embodiment proposes a multi-modal and multi-view endoscopic image segmentation method, which solves the problem of spatial misalignment between NBI and WLI modalities through spatial registration and feature fusion technology, and improves the segmentation accuracy of multi-modal images. The adaptive feature learning mechanism proposed in this method enhances the adaptability of the model to multi-view scenes and effectively copes with the impact of perspective changes on segmentation results. This embodiment combines the multi-modal and multi-view segmentation framework of deep learning to effectively improve the detection accuracy of endoscopic image lesion areas.
[0137] Reference Figure 4 The embodiment of the present application further provides a multi-modal multi-view endoscopic image segmentation device, which can implement the above-mentioned multi-modal multi-view endoscopic image segmentation method, and the device includes:
[0138] An image acquisition unit, used to acquire narrow-band light images and visible light images of the lesion area photographed by the endoscope at different viewing angles;
[0139] A modal feature processing unit, used to extract the modal features of the narrow-band light image and the visible light image respectively, and then fuse the modal features to obtain a multi-modal fusion feature;
[0140] A viewing angle feature processing unit, configured to perform spatial registration on the narrow-band light image and the visible light image, thereby respectively extracting viewing angle features of the narrow-band light image and the visible light image, and then fusing the viewing angle features to obtain a multi-view fusion feature;
[0141] A feature combining unit, used for completing, registering and fusing the multi-modal fusion feature and the multi-view fusion feature to obtain a combined feature;
[0142] An image segmentation unit is used to obtain a first segmentation area based on the modal feature and the joint feature segmentation by a segmentation network based on deep learning; and obtain a second segmentation area based on the viewing angle feature and the joint feature segmentation by the segmentation network based on deep learning.
[0143] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0144] The embodiment of the present application also provides an electronic device, the electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the above multi-modal multi-view endoscopic image segmentation method when executing the computer program. The electronic device can be any smart terminal including a tablet computer, a car computer, etc.
[0145] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0146] See also Figure 5 , Figure 5 The hardware structure of an electronic device of another embodiment is illustrated, and the electronic device includes:
[0147] The processor 501 may be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;
[0148] The memory 502 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 502 can store an operating system and other application programs. When the technical solution provided in the embodiment of this specification is implemented by software or firmware, the relevant program code is stored in the memory 502, and the processor 501 calls and executes the multi-modal multi-view endoscopic image segmentation method of the embodiment of this application;
[0149] Input / output interface 503, used to implement information input and output;
[0150] Communication interface 504, used to realize communication interaction between the device and other devices, which can be realized by wired mode (such as USB, network cable, etc.) or wireless mode (such as mobile network, WIFI, Bluetooth, etc.);
[0151] A bus 505 that transmits information between the various components of the device (e.g., the processor 501, the memory 502, the input / output interface 503, and the communication interface 504);
[0152] The processor 501 , the memory 502 , the input / output interface 503 and the communication interface 504 are connected to each other in communication within the device via the bus 505 .
[0153] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the multi-modal and multi-view endoscopic image segmentation method is implemented.
[0154] It can be understood that the contents of the above method embodiments are all applicable to the present storage medium embodiments, the functions specifically implemented by the present storage medium embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0155] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0156] The embodiments described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0157] Those skilled in the art will appreciate that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0158] The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0159] Those skilled in the art will appreciate that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices may be implemented as software, firmware, hardware, or a suitable combination thereof.
[0160] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0161] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0162] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the above units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0163] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0164] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0165] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including multiple instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, referred to as ROM), random access memory (Random Access Memory, referred to as RAM), disk or optical disk and other media that can store programs.
[0166] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but the scope of the rights of the present invention is not limited thereto. Any modification, equivalent substitution and improvement made by a person skilled in the art without departing from the scope and essence of the present invention should be within the scope of the rights of the present invention.
Claims
1. A multi-modal and multi-view endoscopic image segmentation method, characterized in that: The method comprises the following steps: Acquire narrow-band light images and visible light images of the lesion area captured by the endoscope at different viewing angles; extracting modal features of the narrowband light image and the visible light image respectively, and then fusing the modal features to obtain multimodal fusion features; Performing spatial registration on the narrowband light image and the visible light image to respectively extract viewing angle features of the narrowband light image and the visible light image, and then fusing the viewing angle features to obtain multi-view fusion features; Completing, registering and fusing the multimodal fusion features and the multi-view fusion features to obtain a joint feature; The deep learning-based segmentation network obtains a first segmentation area according to the modal features and the joint features; the deep learning-based segmentation network obtains a second segmentation area according to the viewing angle features and the joint features.
2. The multi-modal multi-view endoscopic image segmentation method according to claim 1, characterized in that: The extracting of the modal features of the narrow-band light image and the visible light image respectively comprises the following steps: Using a feature extraction network to extract the texture of the lesion area in the narrow-band light image and the visible light image as local detail features; Using the feature extraction network to extract the shapes of the lesion areas in the narrow-band light image and the visible light image as global context features; The feature extraction network is used to extract the edges of the lesion area in the narrow-band light image and the visible light image as multi-scale features; Among them, the local detail features, the global context features and the multi-scale features serve as the modal features.
3. The multi-modal multi-view endoscopic image segmentation method according to claim 2, characterized in that: The fusing of the modal features to obtain multimodal fusion features comprises the following steps: Using a cross attention mechanism, the attention weights of the local detail feature, the global context feature, and the multi-scale feature are calculated respectively; The local detail features, the global context features and the multi-scale features are weightedly fused according to the corresponding attention weights to obtain the multimodal fusion features; wherein the multimodal fusion features include the vascular detail information of the narrow-band light image and the overall structural information of the visible light image.
4. The multi-modal multi-view endoscopic image segmentation method according to claim 1, characterized in that: The spatially registering the narrow-band light image and the visible light image and then respectively extracting the viewing angle features of the narrow-band light image and the visible light image comprises the following steps: Calculating a spatial transformation relationship between the narrow-band light image and the visible light image by using a deformation field model or a geometric transformation model based on deep learning; The narrow-band light image and the visible light image are mapped to a unified spatial coordinate system according to the spatial transformation relationship, and then the registered image features after the spatial transformation are output as the viewing angle features.
5. The multi-modal multi-view endoscopic image segmentation method according to claim 1, characterized in that: The step of fusing the viewing angle features to obtain multi-view fusion features comprises the following steps: The perspective features are jointly modeled through three-dimensional position encoding or perspective self-attention mechanism to eliminate perspective differences and then output the multi-perspective fusion features.
6. The multi-modal multi-view endoscopic image segmentation method according to claim 1, characterized in that: The method of completing, registering and fusing the multimodal fusion features and the multi-view fusion features to obtain a joint feature includes the following steps: For the missing areas in the multimodal fusion features or the multi-view fusion features, dynamically complete the missing areas by using the complete information of the multi-view fusion features or the multimodal fusion features through saliency analysis and non-local attention mechanism; Using a deep learning deformation field model to perform spatial registration on the multimodal fusion features and the multi-view fusion features; The multimodal fusion features that have undergone dynamic completion and spatial registration and the multi-view fusion are fused through feature weighting and cross-attention mechanism to obtain the joint features.
7. The multi-modal multi-view endoscopic image segmentation method according to any one of claims 1 to 6, characterized in that: The method further comprises the following steps: Calculate a segmentation loss function according to the pixel-level annotation data of the lesion area, the first segmented area, and the second segmented area; wherein the segmentation loss function includes a cross entropy loss, a Dice loss, and a boundary smoothing loss; Optimizing the parameters of the multimodal multi-view endoscopic image segmentation model using the AdamW optimizer according to the segmentation loss function and the prediction loss; wherein the multimodal multi-view endoscopic image segmentation model is used to segment the narrow-band light image and the visible light image to obtain the first segmented area and the second segmented area; Updating the optimized parameters to the multi-modal multi-view endoscopic image segmentation model through a back-propagation algorithm; Setting a dynamically decaying learning rate to accelerate the convergence of the multi-modal multi-view endoscopic image segmentation model; Determining whether the multimodal multi-view endoscopic image segmentation model meets the convergence condition through the performance index of the validation set; wherein the convergence condition includes that the rate of change of the segmentation loss function is less than a first preset threshold, or that the segmentation accuracy of the first segmented area and the second segmented area reaches a second preset threshold; If the multi-modal multi-view endoscopic image segmentation model satisfies the convergence condition, saving the trained multi-modal multi-view endoscopic image segmentation model; If the multimodal multi-view endoscopic image segmentation model does not meet the convergence condition, return to the step of optimizing the parameters of the multimodal multi-view endoscopic image segmentation model using the AdamW optimizer according to the segmentation loss function and the prediction loss until the convergence condition is met.
8. A multi-modal and multi-view endoscopic image segmentation device, characterized in that: The device comprises: An image acquisition unit, used to acquire narrow-band light images and visible light images of the lesion area photographed by the endoscope at different viewing angles; A modal feature processing unit, used to extract the modal features of the narrow-band light image and the visible light image respectively, and then fuse the modal features to obtain a multi-modal fusion feature; A viewing angle feature processing unit, configured to perform spatial registration on the narrow-band light image and the visible light image, thereby respectively extracting viewing angle features of the narrow-band light image and the visible light image, and then fusing the viewing angle features to obtain a multi-view fusion feature; A feature combining unit, used for completing, registering and fusing the multi-modal fusion feature and the multi-view fusion feature to obtain a combined feature; An image segmentation unit is used to obtain a first segmentation area based on the modal feature and the joint feature segmentation by a segmentation network based on deep learning; and obtain a second segmentation area based on the viewing angle feature and the joint feature segmentation by the segmentation network based on deep learning.
9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the multi-modal and multi-view endoscopic image segmentation method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the multi-modal and multi-view endoscopic image segmentation method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Agricultural product defect detection method and system based on multi-view image feature fusion
CN116563229A
Multi-modal method for classifying thyroid nodule based on ultrasound and infrared thermal images
US20240282090A1
Cited By
Multi-modal endoscope image fusion analysis method
CN121391627A
A Multimodal Endoscopic Image Fusion Analysis Method
CN121391627B
Disease diagnosis and treatment method, system and equipment based on ear-nose-throat endoscope image and medium
CN121481964A
Ear-nose-throat endoscope image-based lesion segmentation method, system, device and medium
CN121481964B
Target tracking method and system based on cross-view multi-modal feature completion
CN121788860A