Endoscopic decayed tooth detection method based on multi-view self-supervised learning
Through the endoscopic caries detection method based on multi-view self-supervised learning, global and local view features are generated, and combined with BFN and MSAM modules, the problems of data scarcity and difficulty in handling multi-scale diseases in the prior art are solved, and higher accuracy of caries detection is achieved.
Patent Information
- Application Number
- CN202510243595.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-20
AI Technical Summary
The existing endoscopic caries detection technology based on deep learning is restricted by the lack of medical endoscopic caries data sets. Self-supervised learning has problems such as insufficient feature representation and difficulty in dealing with multi-scale diseases.
The endoscopic caries detection method based on multi-view self-supervised learning is adopted. By generating global and local views, multi-level comparison loss calculation is performed using MSS strategy, and combined with the BFN backbone network and MSAM module, the model's ability to extract multi-scale caries disease characteristics is enhanced.
It effectively improves the accuracy of endoscopic caries detection, enhances the model's detection performance for multi-scale caries diseases, and reduces the characteristic differences between pre-trained models and downstream applications.
Smart Images

Figure CN120182203A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical image processing, and particularly to an endoscopic caries detection method based on multi-view self-supervised learning. Background Art
[0002] Although progress has been made in endoscopic caries detection based on deep learning, due to the extreme scarcity of medical endoscopic caries datasets, these technologies have been significantly restricted, affecting their further development. The emergence of self-supervised learning (SSL) provides a highly potential solution to address this challenge. However, directly applying traditional SSL technologies to the medical field still faces several challenges.
[0003] Firstly, for the self-supervised pre-trained model, its backbone network is mainly designed based on convolution and focuses on supervising deep features. It cannot fully capture the feature representation from unlabeled endoscopic images and is difficult to adapt to the detection of downstream multi-scale caries targets, making the capabilities learned by the pre-trained model unable to be exerted when applied to downstream detection. Secondly, in dental endoscopic images, the scale differences in the lesion areas are large, and there are particularly some small-scale diseases that are difficult to identify. Therefore, a mechanism capable of effectively processing multi-scale diseases needs to be introduced to improve the accuracy of the downstream model in caries detection.
[0004] In view of this, the present invention proposes an endoscopic caries detection method based on multi-view self-supervised learning. Summary of the Invention
[0005] The purpose of the present invention is to provide an endoscopic caries detection method based on multi-view self-supervised learning for the deficiencies of the prior art.
[0006] To solve the above technical problems, the following technical solutions are adopted:
[0007] An endoscopic caries detection method based on multi-view self-supervised learning, comprising the following steps:
[0008] S1. Generate global images: Input a batch of images. For each original image I in the same batch, generate two different but same-sized global images I O and I M , using two encoders f O and f M to encode the global images I O and I M respectively, and generate four levels of global features and
[0009] S2. Generate global image sample pairs: Consider the global features generated from the global image I O and I M as positive sample pairs, and the other global image features in the global memory bank as negative sample pairs, and calculate the multi-level contrast loss through the MSS strategy;
[0010] S3. Generate local views: Convert each original image I in the same batch in step S1 into local views, and generate two different but same-sized local view sets L O and L M , using two encoders f O and f M to encode the local view sets L O and L M respectively. Obtain the features of 4 local views at each stage, and splice the features of the 4 local views to obtain four-level local view features and
[0011] S4. Generate local view sample pairs: Consider the local view features generated from the global image I O and I M as positive sample pairs, and the other local view features in the local memory bank as negative sample pairs, and calculate the multi-level contrast loss through the MSS strategy;
[0012] S5. Global and local sample pairs: For the same image I in the same batch, consider the global image features generated from the global image and the local image features generated from the global image as positive sample pairs, and the other global image features in the global memory bank as negative sample pairs, and calculate the multi-level contrast loss through the MSS strategy;
[0013] S6. Calculate the loss of the pre-trained model: The total loss is the global image feature loss, the local view feature loss, and the global and local feature loss;
[0014] S7. Fine-tune the pre-trained model: Apply the backbone weights of the pre-trained model trained in the above steps to the backbone of the downstream dental caries detection model, that is, initialize the backbone weights of the detection model as the backbone weights of the pre-trained model, and perform fine-tuning during the training process, and apply the MSAM module to enhance the model's ability to extract multi-scale dental caries disease features.
[0015] Furthermore, the construction methods of the two encoders f O and f M are as follows: Based on the upstream contrast learning pre-trained model, use the BFN bidirectional feature network to construct two encoders f O and f M with the same architecture, and the encoder fO and f M The weights of are random initially. During the training process, the weights of the encoder f M are determined by the encoder f O and are updated by the EMA (Exponential Moving Average) method.
[0016] Furthermore, two encoders f with the same architecture are constructed using the BFN (Bidirectional Feature Network). O and f M Specifically, two encoders f with the same architecture and four stages are constructed using the BFN (Bidirectional Feature Network). O and f M The BFN (Bidirectional Feature Network) combines a CNN branch and a Transformer branch, extracts image features through the CNN branch and the Transformer branch respectively, and performs feature fusion through FOT.
[0017] Furthermore, S1 also includes: The image enhancement methods include random cropping, random horizontal flipping, Gaussian blur, and color jitter related to brightness, contrast, saturation, hue, and grayscale. The encoders f with four stages O and f M are used to encode the global images I O and I M to generate four levels of global features and
[0018] Furthermore, in S2, the process of calculating the multi-level contrast loss through the MSS strategy is as follows: The four levels of global features are respectively used to calculate the contrast loss with to achieve the supervision of multi-level features and help the model learn richer feature representations.
[0019] Furthermore, S3 specifically includes: Input a batch of images. For each original image I in the same batch, 4 local blocks accounting for 60% of the size of the original image are randomly cropped, and then the size of the local blocks is adjusted to 255×255. Then, image enhancement methods including random cropping, random horizontal flipping, Gaussian blur, and color jitter related to brightness, contrast, saturation, hue, and grayscale are applied to generate two different but same-sized local view sets L O and L M , and each local view set contains 4 local blocks; After passing through the encoders f O and f M , each local view set generates four levels of local view set features, and the features of the 4 local views are concatenated to obtain four levels of local view features and
[0020] Further, in the step S4, the calculation process of the multi-level contrast loss through the MSS strategy is as follows: the local view features of four levels respectively with calculate the contrast loss Increasing the model's learning of instance-level images is beneficial to the detection of downstream dental caries.
[0021] Further, the step S6 specifically includes:
[0022] Global loss: w i is the weight of the contrast loss of 4 levels; τ is the temperature coefficient, which is used to balance the learning difficulty of the model and the contrast intensity between positive and negative samples; k is the number of all samples in the global memory bank, and j is the index;
[0023] Local loss: r i is the weight of the contrast loss of 4 levels;
[0024] Global and local loss: t i is the weight of the contrast loss of 4 levels;
[0025] The total loss function is:
[0026] The present invention proposes another technical solution: an endoscopic dental caries detection device based on multi-view self-supervised learning, including a memory and a processor. The memory stores a computer program, and the computer program can be executed by the processor to implement the above-mentioned endoscopic dental caries detection method based on multi-view self-supervised learning.
[0027] The present invention proposes another technical solution: a computer-readable storage medium, characterized in that it stores a computer program, and the computer program can be executed by the processor of the device where the computer-readable storage medium is located to implement the above-mentioned endoscopic dental caries detection method based on multi-view self-supervised learning.
[0028] Due to the adoption of the above technical solution, the following beneficial effects are achieved:
[0029] The present invention provides an endoscopic caries detection method based on multi-view self-supervised learning. Specifically, the present invention is divided into an upstream pre-training model based on contrastive learning and a downstream caries detection model. In the pre-training model, the present invention proposes a new backbone network BFN. BFN combines a CNN branch and a Transformer branch, and fuses the local features extracted by CNN and the global representation information proposed by Transformer through FOT, reducing the feature differences between previous pre-training models and downstream applications, thereby improving the accuracy of caries detection. The present invention reconstructs the views in contrastive learning into global images and local views. Through the MSS strategy, it not only realizes the supervision between the features at different stages of the backbone network, but also conducts the contrast between the global image and the local view, which can effectively enhance the features of each layer of the backbone network and is also beneficial to the performance of the downstream multi-scale caries detector. In addition, the present invention proposes an MSAM module to enhance the downstream model's ability to extract multi-scale caries disease features and enhance the model's detection of caries.
[0030] The present invention reconstructs the contrast views of the contrastive learning pre-training model into global image pairs and local view pairs, and adds the contrast between the global and local views. This method not only trains the image-level representation, but also strengthens the instance-level representation suitable for downstream caries detection, realizing the mutual supervision and enhancement of the two. The features extracted by BFN and FMM, which combine a CNN branch and a Transformer branch, contain both rich local features and global representations, enhancing the downstream model's detection performance for caries. The MSS strategy simultaneously supervises the multi-level feature representations, increases the strong discriminability of each layer of features, and enhances the downstream model's detection of multi-scale caries diseases.
[0031] The BFN backbone network designed by the present invention is suitable for both downstream detection models with a CNN network as the backbone and downstream detection models with a Transformer network as the backbone, improving the applicability of the pre-training model. The global and local views are mutually supervised at multiple levels through the MSS strategy, significantly improving the model's feature representation and being more conducive to downstream multi-scale caries detection. Brief Description of the Drawings
[0032] The present invention will be further described below with reference to the accompanying drawings:
[0033] Figure 1 It is a schematic diagram of the overall process of the present invention.
[0034] Figure 2 It is a schematic diagram of the MSS strategy of the present invention. Detailed Embodiments
[0035] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. However, it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the scope of the present invention. In addition, in the following description, the descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present invention.
[0036] Refer to Figure 1-2 , the first embodiment of the present invention provides an endoscope caries detection method based on multi-view self-supervised learning, which can be executed by an endoscope caries detection device (hereinafter referred to as the detection device), and particularly, executed by one or more processors in the detection device to at least achieve the following steps:
[0037] S1. Generate global images: Input a batch of images. For each original image I in the same batch, generate two different but same-sized global images I O and I M by two sets of random augmentation methods, and use two encoders f O and f M to encode the global images I O and I M respectively, to generate four levels of global features and
[0038] In this embodiment, the detection device can be a terminal with data processing and analysis capabilities such as a desktop computer, a laptop computer, a server, or a workstation. Among them, the corresponding operating system and application software can be installed in the detection device, and the functions required in this embodiment are realized through the combination of the operating system and the application software.
[0039] It should be noted that, refer to Figure 1 , the original image I can be taken by the patient's terminal, such as a terminal with an image shooting function like a mobile phone, a camera, or a tablet computer. After the patient sends the original image I to the doctor, the doctor inputs it into the system.
[0040] As a further description of this embodiment, the S1 further includes: the image augmentation methods include random cropping, random horizontal flipping, Gaussian blur, and color jitter related to brightness, contrast, saturation, hue, and grayscale. Use encoders f O and f M with four stages to encode the global images I O and I M respectively, to generate four levels of global features and
[0041] S2. Generate global image sample pairs: Consider the global images I O and I M generated as a positive sample pair, and the other global image features in the global memory bank as negative sample pairs, and calculate the multi-level contrast loss through the MSS strategy.
[0042] As a further illustration of this embodiment, in the S2, the process of calculating the multi-level contrast loss through the MSS strategy is as follows: The global features of four levels are respectively compared with to calculate the contrast loss to achieve the supervision of multi-level features and help the model learn richer feature representations.
[0043] Specifically, for the global features generated from two different global images I O and I M of each original image in the same batch, consider them as a positive sample pair, and the other global image features in the global memory bank as negative sample pairs. The MSS strategy is as Figure 2 shown. Calculate the multi-level contrast loss through the MSS strategy, that is, the after passing through a projection layer and normalization are respectively compared with to calculate the contrast loss to achieve the supervision of multi-level features and help the model learn richer feature representations. The global loss function is defined by the following formula:
[0044]
[0045] S3. Generate local views: Input a batch of images. For each original image I in the same batch, randomly crop out 4 local blocks that account for 60% of the size of the original image, then resize the local blocks to 255×255, and then apply image enhancement methods including random cropping, random horizontal flipping, Gaussian blur, and color jitter related to brightness, contrast, saturation, hue, and grayscale to generate two different but same-sized local view sets L O and L M , and each local view set contains 4 local blocks; after passing through the encoders f O and f M , each local view set generates local view set features of four levels, and after splicing the features of the 4 local views, local view features of four levels and
[0046] S4. Generate local view sample pairs: Consider the global images I O and I MThe generated local view features are regarded as positive sample pairs, and the other local view features in the local memory bank are negative sample pairs. The multi-level contrast loss is calculated through the MSS strategy.
[0047] As a further illustration of this embodiment, in step S4, the process of calculating the multi-level contrast loss through the MSS strategy is as follows: the local view features of four levels are respectively used to calculate the contrast loss Increasing the model's learning of instance-level images is beneficial to the downstream detection of dental caries.
[0048] Specifically, for two different global images I O and I M of each original image in the same batch, the generated local view features are regarded as positive sample pairs, and the other local view features in the local memory bank are negative sample pairs. Similar to step 2, the multi-level contrast loss is calculated through the MSS strategy, that is, after passing through a projection layer and normalization processing and used to calculate the contrast loss Increasing the model's learning of instance-level images is more beneficial to the downstream detection of dental caries. The local loss function is defined by the following formula:
[0049]
[0050] S5. Global and local sample pairs: For the same image I in the same batch, the global image features generated by the global image and the local image features generated by the global image are regarded as positive sample pairs, and the other global image features in the global memory bank are negative sample pairs. The multi-level contrast loss is calculated through the MSS strategy;
[0051] Specifically, for the same image I in the same batch, the global image features generated by I M and the local image features generated by L O are regarded as positive sample pairs, and the other global image features in the global memory bank are negative sample pairs. The multi-level contrast loss is calculated through the MSS strategy The global and local contrast loss function is defined by the following formula:
[0052]
[0053] τ is the temperature coefficient, which is used to balance the learning difficulty of the model and the contrast intensity between positive and negative samples; k is the number of all samples in the global memory bank, and j is the index.
[0054] S6. Calculation of the pre-trained model loss: The total loss is the global image feature loss, the local view feature loss, and the global and local feature loss.
[0055] As a further illustration of this embodiment, S6 specifically includes:
[0056] Global loss: w i is the weight of the 4-level contrast loss;
[0057] Local loss: r i is the weight of the 4-level contrast loss;
[0058] Global and local loss: t i is the weight of the 4-level contrast loss;
[0059] The total loss function is:
[0060] Specifically, the total loss function of the upstream contrast pre-training model is The pre-training model is trained for 100 epochs, and the trained encoder f o is used as the backbone network for downstream dental caries detection.
[0061] S7. Fine-tuning of the pre-training model: Apply the backbone network weights of the pre-training model trained in the above steps to the backbone network of the downstream dental caries detection model, that is, initialize the backbone network weights of the detection model to the backbone network weights of the pre-training model, and perform fine-tuning during the training process, and apply the MSAM module to enhance the model's ability to extract multi-scale dental caries disease features.
[0062] Specifically, initialize the backbone network weights of the detection model to the backbone network weights of the pre-training model, and use the labeled endoscopic dental caries images to train and fine-tune the detection model. In order to enhance the model's ability to extract multi-scale dental caries disease features, the present invention proposes an MSAM module to capture effective feature representations. The present invention divides the attention weights of self-attention into four parts, and the mathematical formula is as follows:
[0063]
[0064] Among them, ⊙ represents dot product, is the linear transformation of the features captured by the backbone network. By downsampling twice, the resolution size is reduced to half of the original to reduce memory consumption. is the sum of the four attention weights, and m is the number of attention heads. ξ1 is used to measure the similarity, and the formula is where U m and V m c is a learnable embedding matrix. ξ2 depends on the content and relative position of the query, and is expressed as where R k-q Projects the relative position k-q into a high-dimensional representation by calculating sine and cosine functions of different wavelengths. Vm R Is a learnable embedding matrix. ξ2 helps the model to adaptively assign high attention weights to the disease prospect features according to the query content. ξ3 and ξ4 are independent of the query content and are represented as and ξ4 = vm T Vm R R k-q , where um T and vm T Are learnable vectors. They capture the key disease-related content and context dependence that the task should focus on.
[0065] The present invention mainly aims at the problem of endoscopic multi-scale dental caries detection, and proposes an endoscopic dental caries detection method and device based on multi-view self-supervised learning. Specifically, this method is divided into an upstream pre-training model based on contrast learning and a downstream dental caries detection model. In the upstream pre-training model, the present invention reconstructs the view into a global image and a local view, and uses BFN and MSS to align the gap between the pre-training model and the downstream detection model, so that the trained backbone network is more suitable for the downstream dental caries detection model. In the downstream dental caries detection model, this method uses the MSAM module to help the model capture the feature representation of multi-scale dental caries and enhance the detection performance of the model for dental caries.
[0066] As a further description of this embodiment, before performing step S1, it is necessary to first construct an encoder, and the two encoders f O and f M are constructed as follows: Based on the upstream contrast learning pre-training model, two encoders f with the same architecture are constructed using the BFN bidirectional feature network O and f M , and the weights of the encoder f O and f M are random initially. During the training process, the weights of the encoder f M are updated by the weights of the encoder f O through the EMA exponential moving average method.
[0067] As a further description of this embodiment, the specific construction of two encoders f with the same architecture using the BFN bidirectional feature network o and f M includes: Using the BFN bidirectional feature network to construct two encoders f with the same architecture and four stages o and f M, the BFN bidirectional feature network combines a CNN branch and a Transformer branch, extracts image features through the CNN branch and the Transformer branch respectively, and performs feature fusion through FOT. Among them, each CNN branch includes a 1×1 down-projection convolution, a 3×3 deformable convolution, and a 1×1 up-projection convolution. The deformable convolution can better handle the geometric deformation of the disease lesion area in the image. Each Transformer branch includes a multi-head self-attention module and a feed-forward layer (FFN). The feature maps generated by the CNN branch and the Transformer branch are of dimension C×H×W. The feature map generated by the CNN branch passes through an SE attention mechanism to extract more important caries foreground features, and is added element-wise to the features generated by the Transformer branch processed by the MLP.
[0068] The second embodiment of the present invention provides an endoscope caries detection device based on multi-view self-supervised learning, including a memory and a processor. A computer program is stored in the memory and can be executed by the processor to implement a method for endoscope caries detection based on multi-view self-supervised learning as described in any one of the above.
[0069] The third embodiment of the present invention provides a computer-readable storage medium storing a computer program, which can be executed by the processor of the device where the computer-readable storage medium is located to implement a method for endoscope caries detection based on multi-view self-supervised learning as described in any one of the above.
[0070] The present invention reconstructs the contrast views of the contrast learning pre-training model into global image pairs and local view pairs, and increases the contrast between the global and local views. This method trains both the image-level representation and strengthens the instance-level representation suitable for downstream caries detection, realizing mutual supervision and enhancement between the two. The features extracted by the BFN and FMM that combine the CNN branch and the Transformer branch contain both rich local features and global representations, enhancing the caries detection performance of the downstream model. The MSS strategy simultaneously supervises the multi-level feature representations, increases the strong discriminability of each layer of features, and enhances the detection of multi-scale caries diseases by the downstream model.
[0071] Exemplarily, the computer program described in the third and fourth embodiments of the present invention can be divided into one or more modules. The one or more modules are stored in the memory and executed by the processor to complete the present invention. The one or more modules can be a series of computer program instruction segments capable of completing specific functions, and these instruction segments are used to describe the execution process of the computer program in implementing a prediction device for occlusal coverage. For example, the device described in the second embodiment of the present invention.
[0072] The so-called processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the method for predicting a composite coverage, and uses various interfaces and circuits to connect the whole to implement various parts of the method for predicting a composite coverage.
[0073] The memory can be used to store the computer program and / or module. The processor realizes various functions of the method for predicting a composite coverage by running or executing the computer program and / or module stored in the memory, and by calling the data stored in the memory. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, a text conversion function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, text message data, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0074] Among them, if the implemented module is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-described embodiment methods of the present invention, it can also be completed by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0075] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines. Those of ordinary skill in the art can understand and implement it without creative effort.
[0076] As described above, the above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. An endoscopic caries detection method based on multi-view self-supervised learning, characterized in that The following steps are involved: S1. Generate global image: Input a batch of images, and for each original image I of the same batch, generate two different but same-size global images I through two sets of random enhancement methods. O and I M , using two encoders f O and f M For the global image I O and I M Encode and generate four levels of global features and S2. Generate global image sample pairs: O and I M The generated global features are regarded as positive sample pairs, and other global image features in the global memory library are regarded as negative sample pairs. The multi-level contrast loss is calculated through the MSS strategy; S3. Generate local views: Convert each original image I of the same batch in step S1 into a local view, and generate two local view sets L of different but same size through two sets of local random enhancement. O and L M , using two encoders f O and f M For each local view set L O and L M Encoding is performed to obtain features of four local views at each stage, and the features of the four local views are concatenated to obtain local view features of four levels. and S4. Generate local view sample pairs: transform the global image I O and I M The generated local view features are regarded as positive sample pairs, and other local view features in the local memory library are negative sample pairs. The multi-level contrast loss is calculated through the MSS strategy; S5. Global and local sample pairs: For the same image I in the same batch, the global image features generated by the global image and the local image features generated by the global image are regarded as positive sample pairs, and the other global image features in the global memory library are negative sample pairs. The multi-level contrast loss is calculated using the MSS strategy; S6. Pre-training model loss calculation: The total loss is the global image feature loss, local view feature loss, and global and local feature loss; S7. Fine-tuning of pre-trained model: Apply the backbone network weights of the pre-trained model trained in the above steps to the backbone network of the downstream caries detection model, that is, initialize the backbone network weights of the detection model to the backbone network weights of the pre-trained model, and fine-tune it during the training process, and apply the MSAM module to enhance the model's ability to extract multi-scale caries disease features.
2. The endoscopic caries detection method based on multi-view self-supervised learning according to claim 1 is characterized in that: The two encoders f O and f M The construction method is as follows: Based on the upstream contrastive learning pre-training model, the BFN bidirectional feature network is used to build two encoders with the same architecture f O and f M , encoder f O The weights of fM and fM are random at the beginning. During the training process, the encoder f M The weight of is updated by the weight of encoder fO through the EMA exponential moving average method.
3. The endoscopic caries detection method based on multi-view self-supervised learning according to claim 2 is characterized in that: Use the BFN bidirectional feature network to build two encoders with the same architecture f O and f M Specifically, the BFN bidirectional feature network is used to construct two encoders with the same architecture and four stages. O and f M The BFN bidirectional feature network combines the CNN branch and the Transformer branch, extracts image features through the CNN branch and the Transformer branch respectively, and performs feature fusion through FOT.
4. The endoscopic caries detection method based on multi-view self-supervised learning according to claim 3 is characterized in that: The S1 also includes: the image enhancement method includes random cropping, random horizontal flipping, Gaussian blur and color jitter related to brightness, contrast, saturation, hue and grayscale, using an encoder with four stages f O and f M For the global image I O and I M Encode and generate four levels of global features and 5. The endoscopic caries detection method based on multi-view self-supervised learning according to claim 1 is characterized in that: In S2, the process of calculating the multi-level contrast loss by the MSS strategy is as follows: The global features of the four levels Respectively Computing contrast loss Realize supervision of multi-level features to help the model learn richer feature representations.
6. The endoscopic caries detection method based on multi-view self-supervised learning according to claim 1 is characterized in that: S3 specifically includes: inputting a batch of images, for each original image I of the same batch, randomly cropping out 4 local blocks that occupy 60% of the size of the original image, and then adjusting the size of the local blocks to 255×255, and then applying image enhancement methods including random cropping, random horizontal flipping, Gaussian blurring, and color jittering related to brightness, contrast, saturation, hue, and grayscale to generate two different but same-size local view sets L O and L M , each local view set contains 4 local blocks; after the encoder f O and f M Each local view set generates four levels of local view set features, and the features of the four local views are spliced to obtain four levels of local view features. and 7. The endoscopic caries detection method based on multi-view self-supervised learning according to claim 1, characterized in that: In S4, the calculation process of multi-level contrast loss by MSS strategy is as follows: the local view features of four levels Respectively Computing contrast loss Increasing the model’s learning of instance-level images is beneficial for downstream caries detection.
8. The endoscopic caries detection method based on multi-view self-supervised learning according to claim 1 is characterized in that: The S6 specifically includes: Global loss: w i is the weight of the contrast loss of the four levels; τ is the temperature coefficient, which is used to balance the learning difficulty of the model and the contrast strength between positive and negative samples; k is the number of all samples in the global memory library, and j is the index; Local losses: r i is the weight of the contrast loss of the four levels; Global and local losses: t i is the weight of the contrast loss of the four levels; The total loss function is:
9. An endoscopic caries detection device based on multi-view self-supervised learning, characterized in that: It comprises a memory and a processor, wherein the memory stores a computer program, and the computer program can be executed by the processor to implement an endoscopic caries detection method based on multi-view self-supervised learning as described in any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that: A computer program is stored, and the computer program can be executed by a processor of the device where the computer-readable storage medium is located to implement an endoscopic caries detection method based on multi-view self-supervised learning as described in any one of claims 1 to 8.