Oral soft tissue recognition method and device, electronic equipment and storage medium
By constructing a multi-scale feature pyramid and domain-adaptive processing, combined with depth estimation and multi-task collaborative learning, the accuracy and real-time performance issues of oral soft tissue segmentation in existing technologies are solved, achieving high-precision, real-time oral soft tissue segmentation, especially showing excellent performance in small structures such as the gingival margin.
Patent Information
- Application Number
- CN202511762671.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-27
- Publication Date
- 2026-02-27
AI Technical Summary
Existing oral soft tissue segmentation technologies suffer from poor domain adaptability, lack of geometric priors, inaccurate detail segmentation, and insufficient real-time performance, making it difficult to achieve high-precision, real-time oral soft tissue segmentation under complex lighting and occlusion conditions.
We employ multi-level feature extraction to construct a multi-scale feature pyramid, combine domain-adaptive processing and depth estimation, and perform image segmentation through multi-task collaborative learning. We introduce a progressive domain-adaptive mechanism and depth-guided multi-task collaborative learning to enhance semantic accuracy and geometric rationality.
It achieves high-quality, real-time segmentation of oral soft tissues under monocular RGB image conditions, improves the segmentation accuracy of fine structures such as the gingival margin, meets the needs of real-time clinical interaction, and improves the average IoU by more than 12%.
Smart Images

Figure CN121582967A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of oral scanning technology, and more specifically, to a method, device, electronic device, and storage medium for identifying oral soft tissue. Background Technology
[0002] With the rapid development of digital oral medicine, oral endoscopic imaging technology, due to its advantages such as high resolution, non-invasiveness, and real-time visualization, has been widely used in clinical scenarios such as periodontal disease diagnosis, mucosal lesion detection, intraoperative navigation, and postoperative evaluation. Accurate identification and segmentation of oral soft tissues (such as gingiva, buccal mucosa, tongue margin, inflamed areas, or ulcerated areas) is a key prerequisite for achieving automated assisted diagnosis and three-dimensional reconstruction.
[0003] Currently, mainstream oral image analysis methods are mainly based on traditional image processing techniques (such as edge detection, threshold segmentation, and region growing) or shallow machine learning models (such as support vector machines and random forests). These methods rely on manually designed features and have poor robustness under complex lighting, local occlusion, color variation, and individual differences, making it difficult to meet the needs of precision medicine. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a method, apparatus, device and storage medium for oral soft tissue identification.
[0005] To achieve the above objectives, the technical solutions adopted in the embodiments of the present invention are as follows: In a first aspect, the present invention provides a method for identifying oral soft tissue, the method comprising: Acquire raw intraoral endoscopic images; Multi-level feature extraction is performed on the original oral endoscopy images to construct a multi-scale feature pyramid; The multi-scale feature pyramid is subjected to neighborhood adaptive processing, and image segmentation and depth estimation are performed simultaneously based on the processed multi-scale feature pyramid to obtain segmentation features. Based on the segmentation features, the oral soft tissue segmentation results are obtained.
[0006] Optionally, the step of extracting multi-level features from the original oral endoscope image and constructing a multi-scale feature pyramid includes: The original oral endoscope image is subjected to convolution operations of different scales in sequence to obtain shallow edge features, mid-layer texture structure features and deep semantic features; The shallow edge features, the mid-layer texture structure features, and the deep semantic features are sampled to a preset spatial resolution, and the sampled shallow edge features, mid-layer texture features, and deep semantic features are fused to obtain fused multi-scale features. Based on the fused multi-scale features, the multi-scale feature pyramid is constructed.
[0007] Optionally, the step of performing domain adaptation processing on the multi-scale feature pyramid includes: For each layer of features in the multi-scale feature pyramid, discriminative learning of the source and target domains is performed to obtain the differences in feature distribution between domains; Based on the differences in inter-domain feature distribution, the multi-scale feature pyramid is subjected to domain offset correction to obtain the processed multi-scale feature pyramid.
[0008] Optionally, the step of simultaneously performing image segmentation and depth estimation based on the processed multi-scale feature pyramid to obtain segmentation features includes: Each layer of features in the processed multi-scale feature pyramid is sequentially input into a preset decoding network. The first branch of the preset decoding network is used for image segmentation to obtain the first intermediate feature. The second branch of the preset decoding network is used for depth estimation to obtain the second intermediate feature. The first intermediate feature reflects the category of oral soft tissue, and the second intermediate feature reflects the three-dimensional geometric structure of the surface of oral soft tissue. The segmentation features are obtained based on the first intermediate feature and the second intermediate feature.
[0009] Optionally, the step of obtaining the segmentation features based on the first intermediate feature and the second intermediate feature includes: The first intermediate feature and the second intermediate feature are subjected to semantic accuracy coordination processing and geometric rationality coordination processing using a preset loss function to obtain the segmentation feature.
[0010] Optionally, the step of obtaining the oral soft tissue segmentation result based on the segmentation features includes: The segmentation features are subjected to pixel-level semantic decoding to generate initial soft tissue segmentation results; The initial soft tissue segmentation results are corrected for geometric shape and local misjudgment to obtain the oral soft tissue segmentation results.
[0011] Optionally, the step of performing pixel-level semantic decoding on the segmentation features to generate preliminary soft tissue segmentation results includes: The segmentation features are subjected to pixel-by-pixel classification processing to obtain the probability distribution of each pixel in the segmentation features belonging to different soft tissue categories; Based on the probability distribution of each pixel in the segmentation features belonging to different soft tissue categories, the initial segmentation result of the soft tissue is generated.
[0012] In a second aspect, the present invention provides an oral soft tissue recognition device, the device comprising: The acquisition module is used to acquire raw intraoral endoscopic images; The processing module is used to extract multi-level features from the original oral endoscope image and construct a multi-scale feature pyramid; perform neighborhood adaptive processing on the multi-scale feature pyramid, and simultaneously perform image segmentation and depth estimation processing based on the processed multi-scale feature pyramid to obtain segmentation features; and obtain oral soft tissue segmentation results based on the segmentation features.
[0013] Thirdly, the present invention provides an electronic device including a processor and a memory, the memory storing machine-executable instructions executable by the processor, the processor executing the machine-executable instructions to implement the oral soft tissue recognition method described in the first aspect above.
[0014] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the oral soft tissue recognition method as described in the first aspect above.
[0015] The oral soft tissue recognition method, apparatus, device, and storage medium provided in this invention first acquire the original oral endoscope image, then perform multi-level feature extraction on the original oral endoscope image to construct a multi-scale feature pyramid; next, perform domain-adaptive processing on the multi-scale feature pyramid, and simultaneously perform image segmentation and depth estimation based on the processed multi-scale feature pyramid to obtain segmentation features; finally, obtain the oral soft tissue segmentation result based on the segmentation features. Because this invention captures features of different scale structures (such as large areas of gingiva and fine gingival margins) in oral soft tissue by constructing a multi-scale feature pyramid, and combines domain-adaptive processing to bridge the domain differences between natural and medical images, and improves the segmentation accuracy and anatomical rationality of complex anatomical structures through a multi-task collaborative mechanism of simultaneous image segmentation and depth estimation, it ultimately achieves high-quality oral soft tissue segmentation while ensuring real-time performance.
[0016] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This figure shows a schematic block diagram of an electronic device provided by an embodiment of the present invention; Figure 2 A flowchart illustrating a method for identifying oral soft tissue provided by an embodiment of the present invention is shown; Figure 3 The diagram shows a flowchart illustrating an implementation of step S102 according to an embodiment of the present invention. Figure 4 A flowchart illustrating an implementation of step S104 provided in an embodiment of the present invention is shown. Figure 5 A functional block diagram of an oral soft tissue recognition device provided in an embodiment of the present invention is shown.
[0019] Icons: 100 - Electronic device; 110 - Memory; 120 - Processor; 130 - Communication module; 200 - Oral soft tissue recognition device; 201 - Acquisition module; 202 - Processing module. Detailed Implementation
[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0021] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0022] It should be noted that relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0023] With the rapid development of artificial intelligence and medical image analysis technology, deep learning-based image semantic segmentation methods have been widely used in the field of oral medicine. Especially in digital oral diagnosis and treatment, precise segmentation of soft tissues such as the gingiva, buccal mucosa, and lingual frenulum is a crucial foundation for achieving automated diagnosis, treatment planning, and surgical navigation.
[0024] Currently, mainstream oral soft tissue segmentation methods mainly rely on convolutional neural network (CNN) architectures, such as U-Net and its variants. These models are typically trained end-to-end on well-labeled intraoral endoscopic image datasets and can achieve a certain degree of recognition and segmentation of soft tissue regions. However, such methods generally suffer from limited feature representation capabilities, especially exhibiting instability when dealing with complex lighting changes, low-contrast boundaries, and occluded scenes.
[0025] In recent years, some studies have attempted to introduce more powerful visual foundational models, such as Vision Transformer (ViT) or its derivative DINOv3, to improve the model's global modeling capabilities and semantic understanding. While these models demonstrate excellent performance in natural image tasks, they still face significant challenges when directly transferred to medical image segmentation, particularly oral soft tissue segmentation. The main reason is the significant domain gap between natural and medical images, including inconsistencies in imaging equipment, color distribution, texture characteristics, and anatomical structures, which leads to decreased model generalization ability and limited segmentation accuracy.
[0026] Furthermore, some high-precision methods rely on multimodal inputs, such as combining RGB images with real 3D information obtained from depth cameras, and using spatial geometric priors to assist segmentation. While these methods can alleviate segmentation ambiguity in 2D images to some extent, their application is severely limited by dedicated hardware, making them costly and difficult to popularize in ordinary clinical settings.
[0027] Meanwhile, in actual clinical practice, doctors need real-time feedback to guide the scanning path or determine the extent of lesions, thus placing high demands on the inference speed of the algorithm. However, most existing deep learning models have complex structures and large numbers of parameters, making it difficult to meet real-time requirements when running on mobile terminals or embedded devices. They typically cannot achieve processing speeds of more than 25 frames per second, limiting their deployment and application in portable dental scanners.
[0028] More importantly, existing technologies still have significant shortcomings in terms of the precision of segmenting fine anatomical structures. For example, for key areas such as the gingival margin, gingival papilla, and labial / lingual frenulum, due to their blurred edges, minute scale, and variable morphology, traditional methods are prone to breakage, missegmentation, or missed detection, which seriously affects the accuracy of subsequent clinical measurements and analyses.
[0029] In summary, existing oral soft tissue segmentation technologies have many shortcomings, including poor domain adaptability, lack of geometric priors, inaccurate detail segmentation, and insufficient real-time performance. A universal solution that balances high accuracy, strong robustness, and efficient reasoning capabilities has not yet been developed.
[0030] In order to effectively overcome the above-mentioned technical bottlenecks and achieve high-quality, real-time segmentation of oral soft tissue on a mobile platform using only monocular RGB images, embodiments of the present invention provide an oral soft tissue recognition method, device, electronic device, and storage medium, which will be described in detail below.
[0031] Please refer to Figure 1 This is a block diagram of an electronic device 100. The electronic device 100 includes a memory 110, a processor 120, and a communication module 130. The memory 110, processor 120, and communication module 130 are electrically connected to each other directly or indirectly to realize data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses or signal lines.
[0032] The memory 110 is used to store programs or data. The memory 110 may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.
[0033] The processor 120 is used to read / write data or programs stored in the memory 110 and to perform corresponding functions.
[0034] The communication module 130 is used to establish a communication connection between the electronic device 100 and other communication terminals through the network, and to send and receive data through the network.
[0035] It should be understood that, Figure 1 The structure shown is only a schematic diagram of the electronic device 100. The electronic device 100 may also include components that are larger than... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown. Figure 1The components shown can be implemented using hardware, software, or a combination thereof.
[0036] Please refer to Figure 2 The oral soft tissue identification method includes steps S101 to S104.
[0037] S101, acquire raw intraoral endoscopic images.
[0038] In this embodiment of the invention, images of the patient's oral cavity are first acquired using an RGB endoscope lens in an oral scanner. The original oral endoscope image is a 2D color image, typically with a resolution of 1080p or higher, covering various soft tissue areas such as the gingiva, buccal mucosa, tongue, and labial frenulum. This image does not rely on depth sensors or multimodal devices; it can be acquired using only conventional optical imaging, making it suitable for widespread use in ordinary clinical settings.
[0039] Due to interference factors such as uneven lighting, reflections, and saliva obstruction within the oral cavity, the original image often has complex visual noise and low-contrast boundaries.
[0040] S102, multi-level feature extraction is performed on the original oral endoscopy image to construct a multi-scale feature pyramid.
[0041] In order to effectively capture multi-level information from macroscopic anatomical structures to microscopic edge details, this embodiment of the invention adopts an encoder architecture based on DINOv3 as the backbone network to achieve deep semantic analysis of the original oral endoscope images and construct a multi-scale feature pyramid that combines high-resolution details with rich semantic expression capabilities.
[0042] In a possible implementation, step S102 includes the following steps: Figure 3 The sub-steps S102-1 to S102-3 are shown.
[0043] S102-1 performs convolution operations at different scales on the original oral endoscope image to obtain shallow edge features, mid-layer texture structure features, and deep semantic features.
[0044] The raw intraoral endoscopy images are input into the DINOv3 encoder and processed sequentially through multiple Transformer blocks. Three key feature categories are extracted at different depth levels: shallow edge features, mid-level texture structure features, and deep semantic features.
[0045] Shallow edge features correspond to the high-resolution feature maps output from the first few layers of the encoder (subsampled by 4 times), primarily responding to contours, edges, and color abrupt changes in the image.
[0046] The mid-layer texture structure features are derived from the feature representation of the intermediate layer (subsampled 8 or 16 times), including local texture patterns such as gingival surface texture and blood vessel distribution.
[0047] Deep semantic features are low-resolution but highly semantically abstract features generated by high-level networks (sampling 32 times), which can identify the overall organizational category and its spatial context.
[0048] S102-2, shallow edge features, mid-layer texture structure features and deep semantic features are sampled to a preset spatial resolution, and the sampled shallow edge features, mid-layer texture features and deep semantic features are fused to obtain fused multi-scale features.
[0049] To unify the spatial dimensions of features across different levels and promote cross-scale information interaction, this embodiment of the invention introduces a lightweight Feature Pyramid Network (FPN) structure. The shallow, mid-level, and deep features are adjusted to a preset spatial resolution through upsampling operations. For high-level features, bilinear interpolation or transposed convolution is used for upsampling; for low-level features, the original resolution is preserved or appropriate downsampling is performed.
[0050] Subsequently, multi-layer features of the same resolution are fused element-wise through lateral connections. Before each fusion, a 1×1 convolution is used to normalize the number of channels to ensure dimensionality consistency. The final result is a fused multi-scale feature set.
[0051] S102-3, based on the fused multi-scale features, construct a multi-scale feature pyramid.
[0052] Based on the above fusion results, a complete multi-scale feature pyramid is formed. This pyramid not only preserves the spatial details at the bottom level but also integrates semantic information at the higher levels.
[0053] Multi-scale feature pyramids construct multi-scale feature maps with rich semantic information through top-down paths and lateral connections. The high-level feature map is upsampled and fused with the feature map of the previous layer through element-wise addition. This process can be formally represented as:
[0054]
[0055] Here, Conv is a 1x1 convolution used to adjust the number of channels, and Upsample is a 2x upsampling. Finally, This forms a feature pyramid that integrates high-resolution details and deep semantic information, providing a powerful multi-scale feature foundation for subsequent decoding and depth estimation tasks.
[0056] S103 performs domain-adaptive processing on the multi-scale feature pyramid and simultaneously performs image segmentation and depth estimation based on the processed multi-scale feature pyramid to obtain segmentation features.
[0057] To address the challenge of domain transfer in medical image tasks using pre-trained natural image models, this invention proposes a progressive domain adaptation mechanism that efficiently transfers knowledge from the general visual domain to the dental domain while keeping the core parameters essentially frozen.
[0058] In a possible implementation, the process of "performing domain-adaptive processing on the multi-scale feature pyramid" could be as follows: Discriminative learning of the source and target domains is performed on each layer of features in the multi-scale feature pyramid to obtain the inter-domain feature distribution differences. Based on these inter-domain feature distribution differences, domain offset correction is applied to the multi-scale feature pyramid to obtain the processed multi-scale feature pyramid.
[0059] Understandably, a lightweight adapter module is embedded within each Transformer block of DINOv3, with a bottleneck-type feedforward network structure (e.g., dimensionality reduction → GELU activation → dimensionality increase). During training, only the adapter parameters and task head weights are updated, while most parameters of the backbone ViT remain frozen, thus enabling efficient fine-tuning of the parameters.
[0060] An adapter is typically a two-layer feedforward network, and its forward process can be represented as:
[0061] in, and For learnable parameters, It's the bottleneck dimension. The output of the Transformer block becomes... .
[0062] Furthermore, the domain discriminator , with feature extractor Conduct adversarial training. Domain discriminator. It receives feature representations from the natural image source domain and the oral medicine target domain, attempting to distinguish their origins; simultaneously, a feature extractor... It then optimizes its own output to confuse the discriminator's judgment. This is achieved by minimizing the adversarial loss function:
[0063] By minimizing This promotes the alignment of feature distributions between the source and target domains in the latent space, significantly reducing the performance degradation caused by neighborhood offset.
[0064] Based on this, the embodiments of the present invention further adopt a multi-task collaborative learning framework to simultaneously carry out the two tasks of image segmentation and depth estimation.
[0065] In a possible implementation, the process of "simultaneously performing image segmentation and depth estimation based on the processed multi-scale feature pyramid to obtain segmentation features" can be as follows: input each layer of features in the processed multi-scale feature pyramid into a preset decoding network in sequence, perform image segmentation using the first branch of the preset decoding network to obtain the first intermediate feature, perform depth estimation using the second branch of the preset decoding network to obtain the second intermediate feature, and obtain the segmentation features based on the first and second intermediate features.
[0066] The first intermediate feature reflects the category of oral soft tissue, and the second intermediate feature reflects the three-dimensional geometric structure of the oral soft tissue surface.
[0067] In other words, the processed multi-scale feature pyramid is fed layer by layer into a pre-defined decoding network. This decoding network contains two parallel branches: The first branch is the segmented decoder head. It is used to upsample step by step and recover pixel-level classification results, and output the first intermediate feature, which represents the semantic distribution of the soft tissue category to which each pixel belongs (such as gingiva, buccal mucosa, tongue, etc.). The second branch is the depth estimation decoder. The relative depth value corresponding to each pixel is predicted by regression, and a second intermediate feature is generated to reflect the three-dimensional geometric morphology of the soft tissue surface (such as gingival sulcus depression, raised area, etc.).
[0068] Furthermore, to enhance the interaction and complementarity of the two types of information, this embodiment of the invention designs a depth-guided attention module that downsamples the pseudo-depth map to match the current feature map. The same spatial dimensions are then passed through a lightweight convolutional network. Learning Spatial Attention Weight Map :
[0069] in, It is the Sigmoid function.
[0070] This is then applied to segmentation features, resulting in a weighted enhancement of the feature map. This allows for greater attention to areas with significant depth variations (such as the gingival margin and frenulum root), thereby enhancing the characteristic expression of key structures.
[0071] In this embodiment of the invention, a preset loss function can be used to perform semantic accuracy coordination processing and geometric rationality coordination processing on the first intermediate feature and the second intermediate feature to obtain the segmentation feature.
[0072] Understandably, during training, the predicted pseudo-depth map is used to calculate surface normals or curvature information, and a geometric consistency loss is defined:
[0073] in, It is a differentiable geometric property calculation function. For segmentation mask, As a pseudo-depth map, this loss constrains the rationality of the segmentation results in terms of three-dimensional geometry, effectively reducing ambiguous segmentation errors in two-dimensional images.
[0074] By jointly optimizing the above multi-task objectives, the final output is a segmentation feature that integrates semantic accuracy and geometric rationality.
[0075] S104, based on segmentation features, obtain the segmentation results of oral soft tissue.
[0076] After obtaining high-quality segmentation features, they need to be further decoded into accurate segmentation masks that can be used clinically.
[0077] In a possible implementation, step S104 includes the following steps: Figure 4 The sub-steps S104-1 to S104-2 are shown.
[0078] S104-1 performs pixel-level semantic decoding on the segmentation features to generate the initial segmentation result for soft tissue.
[0079] In this embodiment of the invention, the segmentation features can be classified pixel by pixel to obtain the probability distribution of each pixel in the segmentation features belonging to different soft tissue categories; based on the probability distribution of each pixel in the segmentation features belonging to different soft tissue categories, an initial soft tissue segmentation result is generated.
[0080] In other words, the segmentation features are input into the edge-aware decoding head, and pixel-by-pixel classification is performed. The resolution is gradually restored to the original image resolution through a series of upsampling and convolutional layers. Finally, the Softmax function is used in the last layer to output the probability distribution of each pixel belonging to each soft tissue category.
[0081] The category label for each pixel is determined based on the principle of maximum probability, generating an initial segmentation result for soft tissue. This result can roughly delineate the main soft tissue regions, but there may still be breaks or misclassifications at small structures (such as the tip of the gingival papilla) or blurred boundaries.
[0082] S104-2, geometric morphology correction and local misjudgment correction are performed on the initial soft tissue segmentation results to obtain the oral soft tissue segmentation results.
[0083] To further improve edge accuracy and anatomy rationality, this embodiment of the invention introduces a gated feature fusion mechanism and an edge-assisted supervision strategy during the decoding process. During upsampling, the RGB features from the encoder are... Geometric features extracted from pseudo-depth maps Fusion is performed through gating weights. control:
[0084]
[0085] in, This indicates a splicing operation. These are convolution weights.
[0086] The gating feature fusion mechanism can adaptively fuse appearance and geometric information, significantly improving the segmentation accuracy of regions with blurred boundaries.
[0087] Simultaneously, an edge detection branch was added to specifically predict the binary map of soft tissue boundaries. This branch adopts the HED (Holistically-Nested Edge Detection) structure and applies an edge-assisted loss to force the model to focus on key boundaries such as the gingival margin and frenulum attachment point.
[0088] Finally, combining the initial segmentation results with prior edge information, post-processing is performed using a conditional random field or a lightweight RefineNet module to correct isolated noise points, fill in broken edges, smooth unreasonable jagged edges, and output the final oral soft tissue segmentation results.
[0089] The entire model undergoes Neural Architecture Search (NAS) optimization and quantized pruning, achieving an inference speed of 25-30 FPS on mobile chips, meeting the needs of real-time clinical interaction. The total loss function is a weighted sum of the individual losses:
[0090] Compared to existing technologies, this invention achieves high-precision, real-time segmentation of oral soft tissues without depth sensor support by constructing a multi-scale feature pyramid, introducing a progressive domain adaptation mechanism, implementing depth-guided multi-task collaborative learning, and employing an edge-aware decoding strategy. It demonstrates particularly excellent performance in recognizing fine structures such as the gingival margin and frenulum, with an average IoU improvement exceeding 12%, and the segmentation results conform to the laws of three-dimensional anatomical structure, showing promising clinical application prospects.
[0091] To perform the corresponding steps in the above embodiments and various possible methods, an implementation of an oral soft tissue recognition device 200 is given below. Further, please refer to... Figure 5 , Figure 5 This is a functional block diagram of an oral soft tissue recognition device 200 provided in an embodiment of the present invention. It should be noted that the basic principle and technical effects of the oral soft tissue recognition device 200 provided in this embodiment are the same as those in the above embodiments. For the sake of brevity, any parts not mentioned in this embodiment can be referred to the corresponding content in the above embodiments. The oral soft tissue recognition device 200 includes: The acquisition module 201 is used to acquire raw intraoral endoscopic images.
[0092] The processing module 202 is used to extract multi-level features from the original oral endoscope image and construct a multi-scale feature pyramid; perform domain-adaptive processing on the multi-scale feature pyramid, and simultaneously perform image segmentation and depth estimation processing based on the processed multi-scale feature pyramid to obtain segmentation features; and obtain oral soft tissue segmentation results based on the segmentation features.
[0093] Optionally, the above modules can be stored in the form of software or firmware. Figure 1 The memory 110 shown is either stored in or embedded in the operating system (OS) of the electronic device 100, and can be used by... Figure 1 The processor 120 executes the program. Meanwhile, the data and program code required to execute the above modules can be stored in the memory 110.
[0094] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus embodiments described above are merely illustrative; for example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0095] In addition, the functional modules in the various embodiments of the present invention can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0096] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, essentially, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, electronic device 100, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0097] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for identifying oral soft tissue, characterized in that, The method includes: Acquire raw intraoral endoscopic images; Multi-level feature extraction is performed on the original oral endoscopy images to construct a multi-scale feature pyramid; The multi-scale feature pyramid is subjected to neighborhood adaptive processing, and image segmentation and depth estimation are performed simultaneously based on the processed multi-scale feature pyramid to obtain segmentation features. Based on the segmentation features, the oral soft tissue segmentation results are obtained.
2. The oral soft tissue identification method as described in claim 1, characterized in that, The step of extracting multi-level features from the original oral endoscope image and constructing a multi-scale feature pyramid includes: The original oral endoscope image is subjected to convolution operations of different scales in sequence to obtain shallow edge features, mid-layer texture structure features and deep semantic features; The shallow edge features, the mid-layer texture structure features, and the deep semantic features are sampled to a preset spatial resolution, and the sampled shallow edge features, mid-layer texture features, and deep semantic features are fused to obtain fused multi-scale features. Based on the fused multi-scale features, the multi-scale feature pyramid is constructed.
3. The oral soft tissue identification method as described in claim 1, characterized in that, The steps for performing domain-adaptive processing on the multi-scale feature pyramid include: For each layer of features in the multi-scale feature pyramid, discriminative learning of the source and target domains is performed to obtain the differences in feature distribution between domains; Based on the differences in inter-domain feature distribution, the multi-scale feature pyramid is subjected to domain offset correction to obtain the processed multi-scale feature pyramid.
4. The oral soft tissue identification method as described in claim 1, characterized in that, The step of simultaneously performing image segmentation and depth estimation based on the processed multi-scale feature pyramid to obtain segmentation features includes: Each layer of features in the processed multi-scale feature pyramid is sequentially input into a preset decoding network. The first branch of the preset decoding network is used for image segmentation to obtain the first intermediate feature. The second branch of the preset decoding network is used for depth estimation to obtain the second intermediate feature. The first intermediate feature reflects the category of oral soft tissue, and the second intermediate feature reflects the three-dimensional geometric structure of the surface of oral soft tissue. The segmentation features are obtained based on the first intermediate feature and the second intermediate feature.
5. The oral soft tissue identification method as described in claim 4, characterized in that, The step of obtaining the segmentation features based on the first intermediate feature and the second intermediate feature includes: The first intermediate feature and the second intermediate feature are subjected to semantic accuracy coordination processing and geometric rationality coordination processing using a preset loss function to obtain the segmentation feature.
6. The oral soft tissue identification method as described in claim 1, characterized in that, The step of obtaining the oral soft tissue segmentation result based on the segmentation features includes: The segmentation features are subjected to pixel-level semantic decoding to generate initial soft tissue segmentation results; The initial soft tissue segmentation results are corrected for geometric shape and local misjudgment to obtain the oral soft tissue segmentation results.
7. The oral soft tissue identification method as described in claim 6, characterized in that, The step of performing pixel-level semantic decoding on the segmentation features to generate preliminary soft tissue segmentation results includes: The segmentation features are subjected to pixel-by-pixel classification processing to obtain the probability distribution of each pixel in the segmentation features belonging to different soft tissue categories; Based on the probability distribution of each pixel in the segmentation features belonging to different soft tissue categories, the initial segmentation result of the soft tissue is generated.
8. A device for recognizing oral soft tissue, characterized in that, The device includes: The acquisition module is used to acquire raw intraoral endoscopic images; The processing module is used to extract multi-level features from the original oral endoscope image and construct a multi-scale feature pyramid; perform neighborhood adaptive processing on the multi-scale feature pyramid, and simultaneously perform image segmentation and depth estimation processing based on the processed multi-scale feature pyramid to obtain segmentation features; and obtain oral soft tissue segmentation results based on the segmentation features.
9. An electronic device, characterized in that, It includes a processor and a memory, the memory storing machine-executable instructions that can be executed by the processor, the processor executing the machine-executable instructions to implement the oral soft tissue recognition method according to any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the oral soft tissue recognition method as described in any one of claims 1-7.