Multi-region recognition method, device and equipment for endoscope image and storage medium

By matching endoscopic images with three-dimensional anatomical atlases to generate region and constraint masks, and combining this with feature enhancement techniques, the problem of accuracy in identifying multiple anatomical regions in endoscopic images is solved, improving recognition accuracy and robustness.

CN120976923AActive Publication Date: 2025-11-18ZHEJIANG HEALNOC TECH CO LTD
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202511493910.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-20
Publication Date
2025-11-18
Estimated Expiration
2045-10-20

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately identify multiple anatomical regions in endoscopic images, relying heavily on physician experience and exhibiting insufficient accuracy and poor robustness.

Method used

By matching endoscopic images with three-dimensional anatomical atlases, region masks and constraint masks are generated. Feature enhancement is performed by combining image features and anatomical information, and region recognition is performed using a pre-built network.

Benefits of technology

It enables accurate identification of multiple anatomical regions in endoscopic images, improving recognition accuracy and robustness, and reducing reliance on physician experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976923A_ABST
    Figure CN120976923A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-region recognition method, device and equipment for an endoscope image and a storage medium, and the method comprises the steps: projecting a three-dimensional anatomical atlas matched with the endoscope image to the endoscope image, and generating a region mask of the endoscope image according to a projection result; wherein the region mask is used for indicating an anatomical region corresponding to each pixel position in the endoscope image; generating a constraint mask of each anatomical region based on the region masks; determining a target image feature of the endoscope image based on the first image feature of the endoscope image and the region mask; and enhancing the features of the target image based on the constraint mask of each anatomical region, and performing region identification on the endoscope image based on the enhanced features of the target image. Through the method and the device, the problem that accurate recognition of multiple anatomical regions in the endoscope image is difficult to realize is solved, and accurate recognition of the multiple anatomical regions in the endoscope image is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, in particular to a multi-region identification method, device and equipment of an endoscope image and a storage medium. BACKGROUND

[0002] In modern medical diagnosis, endoscopy is a key means for detecting diseases in the throat and other parts. However, its image analysis is highly dependent on the clinical experience of doctors, and faces challenges such as complex and variable anatomical structures, blurred boundaries of diseased tissues, and easy missed diagnosis of small lesions.

[0003] To solve the above problems, the existing method often uses a pre-trained image recognition model to identify the collected endoscope image, but such a model still has the problems of insufficient recognition accuracy and poor robustness, and it is difficult to achieve accurate identification of multiple anatomical regions in the endoscope image.

[0004] To solve the problem that it is difficult to achieve accurate identification of multiple anatomical regions in the endoscope image in the related art, no effective solution has been proposed so far. SUMMARY

[0005] A multi-region identification method, device, equipment and storage medium of an endoscope image are provided in the present embodiment to solve the problem that it is difficult to achieve accurate identification of multiple anatomical regions in the endoscope image in the related art.

[0006] In a first aspect, a multi-region identification method of an endoscope image is provided in the present embodiment, comprising:

[0007] projecting a three-dimensional anatomical atlas matched with the endoscope image to the endoscope image, and generating a region mask of the endoscope image according to the projection result; wherein the region mask is used to indicate the anatomical region corresponding to each pixel position in the endoscope image;

[0008] generating a constraint mask of each anatomical region based on the region mask;

[0009] determining a target image feature of the endoscope image based on a first image feature of the endoscope image and the region mask;

[0010] enhancing the target image feature based on the constraint mask of each anatomical region, and performing region identification on the endoscope image based on the enhanced target image feature.

[0011] In some embodiments, the projecting a three-dimensional anatomical atlas matched with the endoscope image to the endoscope image, and generating a region mask of the endoscope image according to the projection result comprises:

[0012] establish a correspondence between the endoscope image and the three-dimensional anatomical atlas based on a plurality of feature points in the endoscope image;

[0013] project vertices of each anatomical region patch in the three-dimensional anatomical atlas to the endoscope image using the correspondence to obtain a projected two-dimensional point set;

[0014] determine the anatomical region corresponding to each pixel position in the endoscope image according to the projected two-dimensional point set to generate the region mask.

[0015] In some embodiments, the generating of the constraint mask of each anatomical region based on the region mask comprises:

[0016] determining a core region corresponding to each anatomical region; the core region is a set of pixels in the region mask belonging to the corresponding anatomical region;

[0017] determining a transition region corresponding to each anatomical region; each pixel in the transition region is less than or equal to a preset threshold distance from the boundary of the anatomical region;

[0018] performing maximum fusion of the mask of each core region and the mask of the corresponding transition region to generate the constraint mask of the corresponding anatomical region.

[0019] In some embodiments, after the generating of the constraint mask of each anatomical region based on the region mask, the method further comprises:

[0020] determining a plurality of reflective regions in the endoscope image, and a modulation function matching the reflective intensity of each reflective region;

[0021] performing convolution operation on image blocks corresponding to different modulation functions in the endoscope image according to the modulation functions to obtain second image features of the endoscope image;

[0022] optimizing the second image features based on preset parameters to obtain the first image features of the endoscope image; the preset parameters include reflectivity and compensation intensity coefficient.

[0023] In some embodiments, the determining of the target image features of the endoscope image based on the first image features of the endoscope image and the region mask comprises:

[0024] inputting the first image features and the region mask into a pre-constructed double-branch network for processing to obtain first target features and second target features of the endoscope image; the double-branch network comprises a first branch and a second branch;

[0025] wherein the first branch is configured to splice the first image feature with the region mask to obtain the first target feature; and the second branch is configured to determine the second target feature based on the first image feature and a local morphological feature of a three-dimensional mucosa surface; and the three-dimensional mucosa surface is reconstructed based on the endoscopic image;

[0026] fusing the first target feature and the second target feature of the endoscopic image to obtain a target image feature.

[0027] In some embodiments, the enhancing the target image feature based on the constraint mask comprises:

[0028] determining a region saliency score of each of the anatomical regions based on the target image feature;

[0029] adaptively enhancing the target image feature according to the region saliency score of each of the anatomical regions, the constraint mask, and a preset enhancement factor.

[0030] In some embodiments, the performing region recognition on the endoscopic image based on the enhanced target image feature comprises:

[0031] inputting the enhanced target image feature into a convolutional network for processing to output an anatomical region segmentation result of the endoscopic image.

[0032] In a second aspect, an embodiment of the present disclosure provides an endoscopic image multi-region recognition device, comprising:

[0033] a projection module configured to project a three-dimensional anatomical atlas matched with the endoscopic image to the endoscopic image to generate a region mask of the endoscopic image according to a projection result; wherein the region mask is configured to indicate an anatomical region corresponding to each pixel position in the endoscopic image;

[0034] a generation module configured to generate a constraint mask of each of the anatomical regions based on the region mask;

[0035] a fusion module configured to determine a target image feature of the endoscopic image based on a first image feature of the endoscopic image and the region mask;

[0036] a recognition module configured to enhance the target image feature based on the constraint mask of each of the anatomical regions, and perform region recognition on the endoscopic image based on the enhanced target image feature.

[0037] In a third aspect, a computer device is provided in the embodiments, including a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the multi-region identification method of the endoscopic image according to the first aspect when executing the computer program.

[0038] In a fourth aspect, a storage medium is provided in the embodiments, and the storage medium stores a computer program, and the computer program is executable on a processor to implement the multi-region identification method of the endoscopic image according to the first aspect.

[0039] Compared with the related art, the multi-region identification method of the endoscopic image, the device, the equipment and the storage medium provided in the embodiments project a three-dimensional anatomical atlas matched with the endoscopic image to the endoscopic image, generate a region mask of the endoscopic image according to a projection result, wherein the region mask is used to indicate an anatomical region corresponding to each pixel position in the endoscopic image, generate a constraint mask of each anatomical region based on the region mask, determine a target image feature of the endoscopic image based on a first image feature of the endoscopic image and the region mask, enhance the target image feature based on the constraint mask of each anatomical region, and perform region identification on the endoscopic image based on the enhanced target image feature, thereby solving the problem that it is difficult to accurately identify multiple anatomical regions in the endoscopic image, and achieving accurate identification of multiple anatomical regions in the endoscopic image.

[0040] Details of one or more embodiments of the present application are presented in the following drawings and description to make other features, objects and advantages of the present application more apparent. BRIEF DESCRIPTION OF DRAWINGS

[0041] The drawings described herein are intended to provide further understanding of the present application, and form a part of the present application. The illustrative embodiments of the present application and their description serve to explain the present application, and do not constitute an improper limitation of the present application. In the drawings:

[0042] Figure 1 is a flowchart of the multi-region identification method of the endoscopic image provided in an embodiment of the present application;

[0043] Figure 2 is a flowchart of the region mask generation method provided in an embodiment of the present application;

[0044] Figure 3 is a flowchart of the constraint mask generation method provided in an embodiment of the present application;

[0045] Figure 4 is a flowchart of the cross-scale image feature fusion method provided in an embodiment of the present application;

[0046] Figure 5is a flowchart of an image fusion feature enhancement method provided by an embodiment of the present application;

[0047] Figure 6 is a flowchart of a multi-anatomical region identification method provided by an embodiment of the present application;

[0048] Figure 7 is a structural block diagram of a multi-region identification device for endoscopic images provided by an embodiment of the present application.

[0049] In the figure: 10, projection module; 20, generation module; 30, fusion module; 40, identification module. DETAILED DESCRIPTION

[0050] In order to more clearly understand the purpose, technical scheme and advantages of the present application, the present application is described and explained below in conjunction with the drawings and embodiments.

[0051] Unless otherwise defined, technical terms or scientific terms involved in the present application shall have the general meaning understood by those skilled in the art with general knowledge. In the present application, "one", "a", "an", "the", "these" and similar words do not represent a quantitative limitation, and they can be singular or plural. In the present application, the terms "include", "contain", "have" and any variants thereof have the purpose of covering non-exclusive inclusion; for example, a process, method and system, product or device containing a series of steps or modules (units) are not limited to the listed steps or modules (units), but can include steps or modules (units) not listed, or can include other steps or modules (units) inherent to the process, method, product or device. In the present application, the terms "connected", "connected", "coupled" and similar words are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. In the present application, "multiple" means two or more. The association between the associated objects is described by the term "and / or", which means that there can be three relationships, for example, "A and / or B" can mean that A exists alone, A and B exist together, and B exists alone. In general, the character " / " represents an "or" relationship between the objects before and after. In the present application, the terms "first", "second", "third" and the like are only used to distinguish similar objects, and do not represent a specific order of the objects.

[0052] In the present embodiment, a multi-region identification method for endoscopic images is provided, Figure 1 is a flowchart of the multi-region identification method for endoscopic images of the present embodiment, as Figure 1 shown, the flowchart includes the following steps:

[0053] Step S110, projecting the three-dimensional anatomical atlas matched with the endoscope image to the endoscope image, and generating a region mask of the endoscope image according to a projection result; wherein the region mask is used to indicate an anatomical region corresponding to each pixel position in the endoscope image;

[0054] Specifically, medical scan data of a target site (such as a throat, a digestive tract organ, etc.) of a plurality of healthy individuals is acquired, including Computed Tomography (CT) data / Nuclear Magnetic Resonance Imaging (NMR) data, etc. Subsequently, a three-dimensional surface reconstruction algorithm such as Marching Cubes (MC) or Marching Tetrahedra (MT) is used to process the acquired medical scan data to generate a corresponding three-dimensional surface mesh model, which is composed of a plurality of surface mesh patches. The surface mesh model obtained by reconstruction is labeled to identify a plurality of key regions (such as the epiglottis, vocal cords, and piriform fossa of the throat) of the target site. The labeling process can be automatically completed by a target detection or segmentation model based on deep learning, or manually labeled by a professional. Finally, the labeled three-dimensional surface mesh model is integrated and constructed into a three-dimensional anatomical atlas of the target site.

[0055] Further, an endoscope image to be recognized is acquired, and a corresponding three-dimensional anatomical atlas is matched according to the acquisition site. The spatial coordinates of the matched three-dimensional anatomical atlas are projected to the endoscope image plane, and a region mask of the endoscope image is generated according to the projection result. The mask accurately identifies the anatomical structure (such as the epiglottis, vocal cords, and piriform fossa) to which each pixel point in the image belongs, providing structured prior information for subsequent anatomical positioning and structure recognition, that is, helping to constrain the network to perform segmentation only in the anatomically feasible region, thereby improving the region recognition rate.

[0056] Step S120, generating a constraint mask of each anatomical region based on the region mask;

[0057] Specifically, a core region corresponding to each anatomical region is determined, and the core region is a set of pixels belonging to the corresponding anatomical region in the region mask. At the same time, a transition region corresponding to each anatomical region is determined, and the distance between each pixel in the transition region and the boundary of the anatomical region is less than or equal to a preset threshold.

[0058] Further, the mask of each core region and the mask of the corresponding transition region are fused to generate a constraint mask of the corresponding anatomical region, so as to embed the three-dimensional anatomical atlas as a spatial constraint and establish a topological relationship constraint using the anatomical structure information in the atlas.

[0059] At step S130, a target image feature of the endoscopic image is determined based on the first image feature of the endoscopic image and the region mask;

[0060] The first image feature of the endoscopic image is extracted, and then the target image feature of the endoscopic image is generated in combination with the region mask. The specific implementation includes splicing and fusing the first image feature of the endoscopic image and the region mask to obtain the target image feature, or inputting the first image feature of the endoscopic image and the region mask into a double-branch network, wherein the first branch generates a first target feature by splicing the first image feature and the region mask, and the second branch generates a second target feature in combination with the first image feature and a three-dimensional mucosal surface feature reconstructed based on the endoscopic image, and finally fusing the output features of the two branches to obtain the target image feature. In addition, the endoscopic image can be pre-processed (such as reflection compensation, noise suppression, etc.), and the subsequent processing is based on the optimized endoscopic image feature to improve the quality of the generated feature.

[0061] At step S140, the target image feature is enhanced based on the constraint mask of each anatomical region, and the endoscopic image is regionally recognized based on the enhanced target image feature.

[0062] Specifically, the target image feature is enhanced using the constraint mask of each anatomical region to obtain the enhanced target image feature. In other embodiments, the target image feature can be adaptively enhanced in combination with the region saliency score and the constraint mask of each anatomical region to effectively strengthen the key region feature and suppress the influence of non-key or interference regions, thereby improving the accuracy and robustness of subsequent image region recognition.

[0063] Further, the enhanced target image feature is input into a selected image recognition model (such as a convolutional neural network, a Transformer architecture, or a graph neural network, etc.) for processing to obtain an anatomical region segmentation result of the endoscopic image. For example, the enhanced target image feature is input into a convolutional network to extract local anatomical structure features layer by layer, and the same size of convolution kernel is used at each layer to ensure the consistency of the receptive field, and finally the anatomical region segmentation result of the endoscopic image is output. The anatomical region segmentation result accurately identifies a plurality of specific anatomical structures contained in the current collection site, such as the epiglottis, vocal cords, and piriform fossa of the throat, the fundus, body, and antrum of the stomach, etc.

[0064] It should be noted that the endoscopic image in the present embodiment can be derived from a public medical image database, medical teaching materials, algorithm development and verification data sets, or device test simulation environment, etc., which are not specifically limited here. Correspondingly, the region recognition result of the endoscopic image described above can be used for statistical analysis of medical image data (such as image feature distribution statistics, etc.), training and verification of related image processing algorithms, testing of medical image analysis software, etc.

[0065] In modern medical diagnosis, endoscopy is a key means for detecting diseases in the throat and other parts. However, its image analysis is highly dependent on the clinical experience of doctors, and faces challenges such as complex and variable anatomical structures, blurred boundaries of diseased tissues, and easy missed diagnosis of small lesions. To solve this problem, existing methods often use pre-trained image recognition models to recognize the collected endoscopic images, but such models still have problems of insufficient recognition accuracy and poor robustness, making it difficult to achieve accurate recognition of multiple anatomical regions in endoscopic images.

[0066] However, compared with the prior art, the three-dimensional anatomical atlas matched with the endoscopic image is projected onto the endoscopic image, and a region mask of the endoscopic image is generated according to the projection result; wherein the region mask is used to indicate the anatomical region corresponding to each pixel position in the endoscopic image; based on the region mask, a constraint mask of each anatomical region is generated; based on the first image feature of the endoscopic image and the region mask, a target image feature of the endoscopic image is determined; the target image feature is enhanced based on the constraint mask of each anatomical region, and the endoscopic image is recognized based on the enhanced target image feature. Based on this, by introducing a three-dimensional anatomical atlas as a spatial constraint, the topological relationship constraint is established using the anatomical structure information in the atlas, the feature representation combined with anatomical prior is realized, and the rationality and accuracy of region recognition are significantly improved, thereby solving the problem of difficult accurate recognition of multiple anatomical regions in endoscopic images, and achieving accurate recognition of multiple anatomical regions in endoscopic images.

[0067] In some embodiments, as shown in Figure 2 The projection of the three-dimensional anatomical atlas matched with the endoscopic image onto the endoscopic image in step S110, and the generation of the region mask of the endoscopic image according to the projection result, includes the following steps:

[0068] Step S111, based on a plurality of feature points in the endoscopic image, a correspondence between the endoscopic image and the three-dimensional anatomical atlas is established;

[0069] Step S112, using the correspondence, the vertices of each anatomical region patch in the three-dimensional anatomical atlas are projected onto the endoscopic image to obtain a projected two-dimensional point set;

[0070] Step S113, according to the projected two-dimensional point set, the anatomical region corresponding to each pixel position in the endoscopic image is determined to generate a region mask.

[0071] Specifically, feature operators such as Scale-invariant Feature Transform (SIFT) and Speeded Up Robust Features (SURF) are used to extract features from endoscopic images to obtain multiple feature points in the endoscopic images and establish the correspondence between each feature point in the endoscopic images and the anatomical structures in the three-dimensional anatomical atlas.

[0072] Using correspondences, the vertices of each anatomical region patch in the 3D anatomical atlas are projected onto the endoscopic image plane. Based on the projected 2D point set, the anatomical region corresponding to each pixel position in the endoscopic image is determined to generate a region mask. The region mask... The specific expression is as follows:

[0073] (1)

[0074] In equation (1), This represents the Euclidean distance from the anatomical region patch to the image acquisition camera; This represents the two-dimensional point set after projecting all vertices of the anatomical region k. For example, the pyriform fossa region has 120 vertices, which correspond to 120 two-dimensional points after projection. This involves calculating the convex hull. It's understandable that during mask generation, for each pixel in the image... Find the projected convex hull of all anatomical regions k. Select from the area. The smallest region (i.e., the nearest anatomical structure) is labeled with that region. .

[0075] In this embodiment, a correspondence between the endoscope image and the three-dimensional anatomical atlas is established based on multiple feature points in the endoscope image. Using the correspondence, the vertices of each anatomical region patch in the three-dimensional anatomical atlas are projected onto the endoscope image to obtain a projected two-dimensional point set. Based on the projected two-dimensional point set, the anatomical region corresponding to each pixel position in the endoscope image is determined to generate a region mask, thereby achieving accurate generation of the region mask.

[0076] In some of these embodiments, such as Figure 3 As shown, step S120, which generates a constraint mask for each anatomical region based on the region mask, includes the following steps:

[0077] Step S121: Determine the core region corresponding to each anatomical region; the core region is the set of pixels belonging to the corresponding anatomical region in the region mask;

[0078] Step S122, determine the transition region corresponding to each anatomical region; the distance between each pixel in the transition region and the anatomical region boundary is less than or equal to a preset threshold value;

[0079] Step S123, maximum fusion is performed on the mask of each core region and the mask of the corresponding transition region to generate the constraint mask of the corresponding anatomical region.

[0080] Specifically, based on the generated region mask, the core region corresponding to each anatomical region is determined, and the core region is a pixel set belonging to the corresponding anatomical region in the region mask. The specific expression of the core region mask is as follows:

[0081] (2)

[0082] In formula (2), the pixel mask value belonging to the anatomical region k in the region mask is 1; otherwise, the pixel mask value not belonging to the anatomical region k is 0.

[0083] Based on the generated region mask, the transition region corresponding to each anatomical region is determined, and the distance between each pixel in the transition region and the anatomical region boundary is less than or equal to a preset threshold value. The specific expression of the transition region mask is as follows:

[0084] (3)

[0085] In formula (3), represents the region boundary; represents a preset threshold value, for example is set to 15. If the distance between the pixel and the anatomical region boundary is less than or equal to the preset threshold value , it belongs to the transition region, and the mask value is 0.5; otherwise, the pixel mask value is 0. Here, the distance between the pixel and the anatomical region boundary is defined as the minimum Euclidean distance from the pixel to the anatomical region boundary.

[0086] Further, maximum fusion is performed on the mask of each core region and the mask of the corresponding transition region to generate the constraint mask of the corresponding anatomical region, and the specific expression is as follows:

[0087] (4)

[0088] In formula (4), represents the constraint mask of each anatomical region k; max() represents taking the maximum value.

[0089] ​​By this embodiment, the core region corresponding to each anatomical region is determined, the core region is a pixel set belonging to the corresponding anatomical region in the region mask, and the transition region corresponding to each anatomical region is determined, each pixel in the transition region has a distance less than or equal to a preset threshold from the boundary of the anatomical region, and finally the mask of each core region and the mask of the corresponding transition region are fused by maximum value to generate the constraint mask of the corresponding anatomical region, so that in subsequent image processing, the edge uncertainty can be considered while focusing on the anatomical region, which helps to improve the robustness of region recognition and segmentation.

[0090] In some embodiments, after generating the constraint mask of each anatomical region based on the region mask, the multi-region recognition method of the endoscopic image further includes the following steps:

[0091] Determine a plurality of light reflection regions in the endoscopic image, and determine a modulation function matching the light reflection intensity of each light reflection region;

[0092] According to the modulation functions, perform convolution operation on the image blocks corresponding to different modulation functions in the endoscopic image to obtain the second image features of the endoscopic image;

[0093] Optimize the second image features based on preset parameters to obtain the first image features of the endoscopic image; the preset parameters include reflectivity and compensation intensity coefficient.

[0094] Specifically, the endoscopic image is detected to identify the light reflection regions in the endoscopic image. For example, the pixels with Y channel brightness value greater than 220 and gradient change value less than 5 in the YUV color space of the image are detected as light reflection regions.

[0095] The light reflection intensity of each light reflection region is analyzed, and the calculation method includes but is not limited to the average brightness intensity of the local light reflection region, the brightness distribution intensity, etc. For example, the specific calculation formula of the light reflection intensity is as follows:

[0096] (5)

[0097] In formula (5), G represents the light reflection region; represents the pixel neighborhood selected by a 7x7 window; indicates the number of pixels; is the light reflection intensity, and the calculation logic is to take each pixel position (x, y) as the center, and count the proportion of pixels detected as the light reflection region in the neighborhood window, that is, the local light reflection intensity.

[0098] A filter operator matching the light reflection intensity of each light reflection region is determined, and the dynamic matching mechanism is as follows:

[0099] (6)

[0100] In formula (6), represents a matched filter operator. If , it indicates that the region is a strong light reflection region, and a Laplacian operator is used to retain the structure of the tissue by using its edge enhancement characteristics; if , it indicates that the region is a moderate light reflection region, and and are combined, is used to detect the edge in the vertical direction, is used to detect the edge in the horizontal direction; if , it indicates that the region is a weak light reflection region, and a Gaussian operator is used to retain the mucosa details by using Gaussian smoothing. It should be noted that formula (6) above is an example of a matching mechanism, and the operator selection logic can be flexibly adjusted according to specific needs in actual applications.

[0101] Further, according to the matched filter operator, a corresponding modulation function is determined, and the specific expression is as follows:

[0102] (7)

[0103] In formula (7), represents the modulation function; represents the initial weight coefficient.

[0104] According to the matched modulation functions, the image blocks corresponding to different modulation functions in the endoscope image are subjected to convolution operation, to obtain a second image feature of the endoscope image. The specific expression of the convolution operation is as follows:

[0105] (8)

[0106] In formula (8), represents an image block in the endoscope image, represents the modulation function corresponding to the image block ; represents the image feature after the convolution operation, i.e., the second image feature.

[0107] Then, the second image feature is optimized based on a preset parameter to obtain a first image feature of the endoscope image, and the preset parameter includes reflectivity and a compensation intensity coefficient. The specific expression of the optimization process is as follows:

[0108] (9)

[0109] In formula (9), represents the optimized image feature, i.e., the first image feature; represents the compensation intensity coefficient, such as is set to 0.05; represents reflectivity.

[0110] By this embodiment, multiple light reflection regions in the endoscope image are determined, and a modulation function matched with the light reflection intensity of each light reflection region is determined. According to the modulation functions, the image blocks corresponding to different modulation functions in the endoscope image are subjected to convolution operation, to obtain a second image feature of the endoscope image. Then, the second image feature is optimized based on preset parameters, to obtain a first image feature of the endoscope image. In this way, a dynamic kernel function is designed. By analyzing the light intensity distribution of the local region, the weight parameters of the convolution kernel are adaptively adjusted, and the light reflection region is compensated in the gradient domain. This scheme can fundamentally suppress the feature interference of the light reflection region and compensate the effective lesion feature, realize endoscopic image enhancement, break through the bottleneck of detail loss and structure distortion caused by general image restoration, and help to improve the accuracy of subsequent image region recognition.

[0111] In some embodiments, as shown in FIG. 13B, the step S130 of determining the target image feature of the endoscope image based on the first image feature of the endoscope image and the region mask includes the following steps: Figure 4

[0112] The step S131 inputs the first image feature and the region mask into a pre-constructed double-branch network for processing, to obtain a first target feature and a second target feature of the endoscope image. The double-branch network includes a first branch and a second branch.

[0113] The first branch is used to splice the first image feature and the region mask, to obtain the first target feature. The second branch is used to determine the second target feature based on the first image feature and the local morphological feature of the three-dimensional mucosa surface. The three-dimensional mucosa surface is obtained based on the reconstruction of the endoscope image.

[0114] The step S132 fuses the first target feature and the second target feature of the endoscope image, to obtain the target image feature.

[0115] Specifically, the first image feature and the region mask are input into a pre-constructed double-branch network for processing, to obtain a first target feature and a second target feature of the endoscope image. The double-branch network includes a first branch and a second branch.

[0116] ​The first branch is a macroscopic anatomical branch (processing global anatomical structure) for splicing the first image feature and a region mask to obtain a first target feature; and the second branch is a microscopic lesion branch (focusing on mucosal surface details) for determining a second target feature based on the first image feature and local morphological features of a three-dimensional mucosal surface. The three-dimensional mucosal surface is reconstructed based on the endoscope image, and the local morphological features of the three-dimensional mucosal surface include mucosal surface distance information and surface curvature features.

[0117] Further, the first target feature and the second target feature of the endoscope image are cross-scale feature fused to obtain a target image feature.

[0118] For example, the first branch adopts U-Net network coding, and the second branch adopts HRNet network, and the specific expression is as follows:

[0119] (10)

[0120] (11)

[0121] In formula (10), ; denotes the first image feature; denotes the region mask. In formula (11), denotes the second target feature; denotes the first image feature; denotes the mucosal surface distance map. It should be noted that the mucosal surface distance map is used to indicate the vertical distance (i.e. normal distance) from the point corresponding to the image coordinates (x, y) on the reconstructed three-dimensional mucosal surface to the best reference plane locally fitted near the (x, y) point.

[0122] The first target feature and the second target feature are input into a cross attention (Cross Attention) network for cross-scale feature fusion to obtain a target image feature. The specific expression of the fusion process is as follows:

[0123] (12)

[0124] In formula (12), denotes the target image feature obtained by fusion; ; denotes the second target feature.

[0125] By this embodiment, the first image feature and the region mask are input into a pre-constructed double-branch network for processing to obtain a first target feature and a second target feature of the endoscopic image. The double-branch network includes a first branch and a second branch. The first branch is used for splicing the first image feature with the region mask to obtain the first target feature. The second branch is used for determining the second target feature based on the first image feature and a local morphological feature of a three-dimensional mucosa surface, which is reconstructed based on the endoscopic image. Then, the first target feature and the second target feature of the endoscopic image are fused to obtain a target image feature. In this way, the macroscopic anatomical branch and the microscopic lesion branch are fused to realize cross-scale feature interaction, which can simultaneously identify the macroscopic anatomical region and the microscopic lesion, thereby improving the recognition rate of the difficult sample region by using the constraint of the anatomical atlas and the multi-scale lesion perception mechanism.

[0126] In some embodiments, as shown in FIG. 1, the step S140 of enhancing the target image feature based on the constraint mask includes the following steps: Figure 5

[0127] In step S141, the region saliency score of each anatomical region is determined based on the target image feature.

[0128] In step S142, the target image feature is adaptively enhanced according to the region saliency score of each anatomical region, the constraint mask and a preset enhancement factor.

[0129] Specifically, the region saliency of each anatomical region is evaluated based on the target image feature to obtain the region saliency score of each anatomical region. The specific expression of this evaluation method is as follows:

[0130] (13)

[0131] In formula (13), the saliency is calculated by the differential change rate of the feature. Wherein, represents the region saliency score of the anatomical region k; represents the set of all pixels of the anatomical region k; represents the number of pixels of the anatomical region k; represents the gradient of the target image feature.

[0132] Further, the target image feature is adaptively enhanced according to the region saliency score of each anatomical region, the constraint mask and a preset enhancement factor. The specific expression of this enhancement process is as follows:

[0133] (14)

[0134] In formula (14), represents the enhanced target image feature;​ denotes a target image feature; denotes a preset enhancement factor, such as is set to 0.3; denotes a region saliency score of the anatomical region k; denotes a constraint mask of the anatomical region k.

[0135] Through the embodiment, based on the target image feature, the region saliency score of each anatomical region is determined, and the target image feature is adaptively enhanced according to the region saliency score of each anatomical region, the constraint mask and the preset enhancement factor, so as to effectively strengthen the key region feature and improve the accuracy and robustness of subsequent image region recognition.

[0136] In some embodiments, as shown in Figure 6 the step S140 of performing region recognition on the endoscopic image based on the enhanced target image feature includes the following steps:

[0137] Step S143, inputting the enhanced target image feature into a convolutional network for processing to output an anatomical region segmentation result of the endoscopic image.

[0138] Specifically, the enhanced target image feature is input into the convolutional network for processing. The network extracts local anatomical structure features in the image through layer-by-layer and multi-level convolution operations, while maintaining the consistency of the receptive field to preserve key spatial structure information, and finally outputs the anatomical region segmentation result of the endoscopic image. The specific expression of the recognition process is as follows:

[0139] (15)

[0140] In formula (15), denotes the enhanced target image feature; Softmax() is an activation function; Conv() is a convolution operation; denotes the anatomical region segmentation result.

[0141] The anatomical region segmentation result accurately identifies a plurality of specific anatomical structures contained in the current collection site. For example, in the throat site, key structures such as the epiglottis, vocal cords, and piriform fossa are distinguished and marked; in the stomach, the main anatomical partitions such as the fundus, body, and antrum of the stomach are clearly segmented.

[0142] Through the embodiment, the enhanced target image feature is input into the convolutional network for processing to output the anatomical region segmentation result of the endoscopic image, realizing the multi-region accurate recognition of the endoscopic image, thereby providing an accurate anatomical structure information basis.

[0143] The embodiment will be described and explained below through specific embodiments.

[0144] CT / MRI scan data of 200 healthy volunteers' nasopharyngeal laryngeal sites are collected in advance, the CT / MRI scan data of the nasopharyngeal larynx is reconstructed in three dimensions by the Marching Cubes algorithm, the nasopharyngeal laryngeal surface grid patches are generated, and the key area annotation of the three-dimensional reconstruction result is performed, including the nasal vestibule, nasal septum, inferior turbinate, middle turbinate, eustachian tube pharyngeal opening, nasopharyngeal roof, oropharyngeal posterior wall, epiglottic valley, epiglottis, vocal cord, ventricular band, and piriform fossa, so as to generate a three-dimensional anatomical atlas of the nasopharyngeal laryngeal site.

[0145] The nasopharyngeal laryngeal endoscope image to be recognized is obtained, and the three-dimensional anatomical atlas of the nasopharyngeal laryngeal site is called, a plurality of feature points in the endoscope image are extracted, a corresponding relationship between the endoscope image and the three-dimensional anatomical atlas is established based on each feature point, and the vertices of each anatomical region patch in the three-dimensional anatomical atlas are projected to the endoscope image using the corresponding relationship to obtain a projected two-dimensional point set. Then, according to the projected two-dimensional point set, the anatomical region corresponding to each pixel position in the endoscope image is determined to generate a region mask.

[0146] The core region and the transition region corresponding to each anatomical region are determined, the core region is a pixel set in the region mask belonging to the corresponding anatomical region, and the distance between each pixel in the transition region and the boundary of the anatomical region is less than or equal to a preset threshold. The mask of each core region and the mask of the corresponding transition region are fused by maximum value to generate a constraint mask of the corresponding anatomical region.

[0147] Further, a plurality of reflection regions in the endoscope image are detected, the local reflection intensity of the reflection region is analyzed and calculated, and a modulation function matched with the reflection intensity of each reflection region is selected. According to each modulation function, the image blocks corresponding to different modulation functions in the endoscope image are convolved to obtain a second image feature of the endoscope image, and the second image feature is optimized based on the reflectivity and the compensation intensity coefficient to finally obtain a first image feature of the endoscope image.

[0148] Then, the first image feature and the region mask are input into a pre-constructed double-branch network for processing to obtain a first target feature and a second target feature of the endoscope image. The double-branch network includes a first branch and a second branch, the first branch adopts U-Net network coding to splice the first image feature and the region mask to obtain the first target feature, and the second branch adopts HRNet network to determine the second target feature based on the first image feature and three-dimensional mucosa surface distance information, and the three-dimensional mucosa surface is reconstructed based on the endoscope image. Subsequently, the first target feature and the second target feature of the endoscope image are input into a cross-attention network for cross-scale feature fusion to obtain a target image feature.

[0149] The region saliency of the target image feature is calculated by using the differential change rate of the feature, the region saliency score of each anatomical region is obtained, and the target image feature is adaptively enhanced according to the region saliency score of each anatomical region, the constraint mask and the preset enhancement factor. The high-dimensional feature tensor after multi-scale feature fusion and spatial attention enhancement is input into the convolutional network, the local anatomical structure features are extracted layer by layer and the receptive field consistency is maintained, and finally the segmentation result of specific anatomical structures such as piriform fossa and epiglottis is output, and the lesion position can be identified.

[0150] Through the embodiment, the depth embedding of the nasopharyngeal three-dimensional anatomical topological constraint and the dynamic feature focusing mechanism are realized, and the mirror reflection area interference is adaptively compensated, and the precise synchronous identification of the multi-scale anatomical region and the lesion in the nasopharyngoscope image is realized.

[0151] It should be noted that the steps shown in the above flow or the flowchart of the accompanying drawings can be executed in a computer system such as a group of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that here.

[0152] In the embodiment, an endoscope image multi-region identification device is also provided, which is used to realize the above-mentioned embodiments and preferred embodiments, and will not be described again. The terms "module", "unit", "sub-unit" and the like used below can be a combination of software and / or hardware that realizes a predetermined function. Although the device described in the following embodiments is preferably realized in software, hardware or a combination of software and hardware is also possible and is conceived.

[0153] Figure 7 is a structural block diagram of the endoscope image multi-region identification device of the embodiment, as Figure 7 shown, the device comprises:

[0154] The projection module 10 is configured to project a three-dimensional anatomical atlas matched with the endoscope image to the endoscope image, and generate a region mask of the endoscope image according to a projection result; wherein the region mask is used to indicate an anatomical region corresponding to each pixel position in the endoscope image;

[0155] The generation module 20 is configured to generate a constraint mask of each anatomical region based on the region mask;

[0156] The fusion module 30 is configured to determine a target image feature of the endoscope image based on a first image feature of the endoscope image and the region mask;

[0157] The recognition module 40 is used to enhance the features of the target image based on the constraint mask of each anatomical region, and to perform region recognition of the endoscopic image based on the enhanced target image features.

[0158] The apparatus provided in this embodiment projects a three-dimensional anatomical atlas matching the endoscopic image onto the endoscopic image, and generates a region mask for the endoscopic image based on the projection result. The region mask indicates the anatomical region corresponding to each pixel position in the endoscopic image. Based on the region mask, a constraint mask for each anatomical region is generated. Based on the first image features of the endoscopic image and the region mask, the target image features of the endoscopic image are determined. The target image features are enhanced based on the constraint masks of each anatomical region, and based on the enhanced target image features, region recognition is performed on the endoscopic image. This solves the problem of accurately identifying multiple anatomical regions in pharyngeal endoscopic images, and achieves accurate identification of multiple anatomical regions in pharyngeal endoscopic images.

[0159] In some embodiments, the projection module 10 is further configured to establish a correspondence between the endoscope image and the three-dimensional anatomical atlas based on multiple feature points in the endoscope image; using the correspondence, project the vertices of each anatomical region patch in the three-dimensional anatomical atlas onto the endoscope image to obtain a projected two-dimensional point set; and determine the anatomical region corresponding to each pixel position in the endoscope image based on the projected two-dimensional point set to generate a region mask.

[0160] In some embodiments, the generation module 20 is further configured to determine the core region corresponding to each anatomical region; the core region is the set of pixels belonging to the corresponding anatomical region in the region mask; determine the transition region corresponding to each anatomical region; the distance between each pixel in the transition region and the boundary of the anatomical region is less than or equal to a preset threshold; and fuse the mask of each core region and the mask of the corresponding transition region by maximizing the value to generate the constraint mask of the corresponding anatomical region.

[0161] In some of these embodiments, Figure 7 Based on this, the device also includes an optimization module for determining multiple reflective regions in the endoscopic image and a modulation function matching the reflective intensity of each reflective region; according to each modulation function, convolving the image blocks corresponding to different modulation functions in the endoscopic image to obtain the second image feature of the endoscopic image; optimizing the second image feature based on preset parameters to obtain the first image feature of the endoscopic image; the preset parameters include reflectivity and compensation intensity coefficient; and determining the target image feature of the endoscopic image based on the first image feature and the region mask.

[0162] In some embodiments, the fusion module 30 is further configured to input the first image feature and the region mask into a pre-constructed double-branch network to obtain a first target feature and a second target feature of the endoscopic image; the double-branch network comprises a first branch and a second branch; the first branch is configured to splice the first image feature and the region mask to obtain the first target feature; the second branch is configured to determine the second target feature based on the first image feature and a local morphological feature of a three-dimensional mucosa surface; the three-dimensional mucosa surface is reconstructed based on the endoscopic image; and the first target feature and the second target feature of the endoscopic image are fused to obtain a target image feature.

[0163] In some embodiments, the identification module 40 is further configured to determine a region saliency score of each anatomical region based on the target image feature; and perform adaptive enhancement on the target image feature according to the region saliency score of each anatomical region, the constraint mask and a preset enhancement factor.

[0164] In some embodiments, the identification module 40 is further configured to input the enhanced target image feature into a convolutional network to output an anatomical region segmentation result of the endoscopic image.

[0165] It should be noted that each of the above modules can be a functional module or a program module, and can be implemented by software or hardware. For the modules implemented by hardware, each of the above modules can be located in the same processor; or each of the above modules can be located in different processors in any combination.

[0166] In the present embodiment, a computer device is also provided, which comprises a memory and a processor, the memory stores a computer program, and the processor is configured to execute the computer program to perform the steps in any of the above method embodiments.

[0167] Optionally, the computer device can further comprise a transmission device and an input / output device, wherein the transmission device is connected with the processor, and the input / output device is connected with the processor.

[0168] Optionally, in the present embodiment, the processor can be configured to execute the following steps through the computer program:

[0169] S1, projecting a three-dimensional anatomical atlas matched with the endoscopic image to the endoscopic image, and generating a region mask of the endoscopic image according to a projection result; wherein the region mask is used to indicate an anatomical region corresponding to each pixel position in the endoscopic image;

[0170] S2, generating a constraint mask of each anatomical region based on the region mask;

[0171] S3, determining a target image feature of the endoscopic image based on the first image feature and the region mask of the endoscopic image;

[0172] S4, enhancing the target image feature based on the constraint mask of each anatomical region, and performing region recognition on the endoscopic image based on the enhanced target image feature.

[0173] It should be noted that the specific examples in the embodiment can refer to the examples described in the above embodiments and optional implementation manners, which will not be described herein again.

[0174] In addition, in combination with the multi-region recognition method of the endoscopic image provided in the above embodiments, a storage medium can also be provided in the embodiment to realize the multi-region recognition method of the endoscopic image. The storage medium has a computer program stored thereon; the computer program is executed by a processor to realize any one of the multi-region recognition methods of the endoscopic image in the above embodiments.

[0175] It should be understood that the specific embodiments described herein are only used to explain this application, but not to limit it. According to the embodiments provided in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor are within the scope of protection of the present application.

[0176] Obviously, the drawings are only some examples or embodiments of the present application, and can be applied to other similar situations without creative labor for those of ordinary skill in the art. In addition, it can be understood that although the work done in the development process may be complex and long, some design, manufacture or production changes according to the technical content disclosed in the present application are only routine technical means for those of ordinary skill in the art, and should not be regarded as insufficient disclosure of the present application.

[0177] The word "embodiment" in the present application means that the specific features, structures or characteristics described in combination with the embodiments can be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily mean the same embodiment, nor does it mean independence or alternative to other embodiments. Those of ordinary skill in the art can clearly or implicitly understand that the embodiments described in the present application can be combined with other embodiments without conflict.

[0178] The above described embodiments only express several implementation manners of the present application, which are described in detail and specifically, but should not be understood as limitation to the patent protection scope. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, some modifications and improvements can be made, which are within the scope of protection of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A method for multi-region recognition of endoscopic images, characterized in that, include: A three-dimensional anatomical atlas matching the endoscopic image is projected onto the endoscopic image, and a region mask of the endoscopic image is generated based on the projection result; wherein, the region mask is used to indicate the anatomical region corresponding to each pixel position in the endoscopic image; Based on the region mask, a constraint mask for each of the anatomical regions is generated; Based on the first image features of the endoscope image and the region mask, the target image features of the endoscope image are determined; The target image features are enhanced based on the constraint mask of each anatomical region, and the endoscopic image is region-based based on the enhanced target image features.

2. The multi-region recognition method for endoscopic images according to claim 1, characterized in that, The step of projecting a three-dimensional anatomical atlas matching the endoscopic image onto the endoscopic image, and generating a region mask for the endoscopic image based on the projection result, includes: Based on multiple feature points in the endoscopic image, a correspondence between the endoscopic image and the three-dimensional anatomical atlas is established; Using the aforementioned correspondence, the vertices of each anatomical region patch in the three-dimensional anatomical atlas are projected onto the endoscopic image to obtain a projected two-dimensional point set; Based on the projected two-dimensional point set, the anatomical region corresponding to each pixel position in the endoscopic image is determined to generate the region mask.

3. The multi-region recognition method for endoscopic images according to claim 1, characterized in that, The step of generating a constraint mask for each anatomical region based on the region mask includes: Determine the core region corresponding to each of the anatomical regions; the core region is the set of pixels in the region mask that belong to the corresponding anatomical region. Determine a transition region corresponding to each of the anatomical regions; the distance between each pixel in the transition region and the boundary of the anatomical region is less than or equal to a preset threshold. The mask of each core region and the mask of the corresponding transition region are fused by maximizing the value to generate the constraint mask corresponding to the anatomical region.

4. The multi-region recognition method for endoscopic images according to claim 1, characterized in that, After generating a constraint mask for each anatomical region based on the region mask, the method further includes: Identify multiple reflective regions in the endoscopic image, and a modulation function that matches the reflective intensity of each reflective region; According to each of the modulation functions, the image blocks corresponding to different modulation functions in the endoscope image are convolved to obtain the second image features of the endoscope image; The second image features are optimized based on preset parameters to obtain the first image features of the endoscope image; the preset parameters include reflectivity and compensation intensity coefficient.

5. The multi-region recognition method for endoscopic images according to claim 1 or claim 4, characterized in that, The step of determining the target image features of the endoscope image based on the first image features of the endoscope image and the region mask includes: The first image features and the region mask are input into a pre-constructed dual-branch network for processing to obtain the first target features and the second target features of the endoscope image; the dual-branch network includes a first branch and a second branch; The first branch is used to concatenate the first image features with the region mask to obtain the first target features; the second branch is used to determine the second target features based on the first image features and the local morphological features of the three-dimensional mucosal surface; the three-dimensional mucosal surface is reconstructed based on the endoscopic image; The first target feature and the second target feature of the endoscopic image are fused to obtain the target image feature.

6. The multi-region recognition method for endoscopic images according to claim 1, characterized in that, The enhancement of the target image features based on the constraint mask includes: Based on the target image features, a regional salience score is determined for each of the anatomical regions; The target image features are adaptively enhanced based on the regional saliency score of each anatomical region, the constraint mask, and the preset enhancement factor.

7. The multi-region recognition method for endoscopic images according to claim 1, characterized in that, The process of performing region recognition on the endoscopic image based on the enhanced target image features includes: The enhanced target image features are input into a convolutional network for processing to output the anatomical region segmentation results of the endoscope image.

8. A multi-region recognition device for endoscopic images, characterized in that, include: A projection module is used to project a three-dimensional anatomical atlas matching the endoscopic image onto the endoscopic image, and generate a region mask of the endoscopic image based on the projection result; wherein, the region mask is used to indicate the anatomical region corresponding to each pixel position in the endoscopic image; A generation module is used to generate a constraint mask for each of the anatomical regions based on the region mask; A fusion module is used to determine the target image features of the endoscope image based on the first image features of the endoscope image and the region mask; The recognition module is used to enhance the features of the target image based on the constraint mask of each of the anatomical regions, and to perform region recognition on the endoscopic image based on the enhanced target image features.

9. A computer device, comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to run the computer program to perform the steps of the multi-region recognition method for endoscopic images according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the multi-region recognition method for endoscopic images according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Intraoperative preset area positioning method and device, storage medium and electronic equipment

    CN111415404A

  • Three-dimensional reconstruction method and system based on instance segmentation, storage medium and terminal

    CN112927354A

  • Image enhancement method and device, equipment and storage medium

    CN114140343A

  • Image fusion method and system for endoscope, medium and electronic equipment

    CN117495693A

  • Integrated image restoration method and system based on image mask modeling

    CN119722531A