Dermatophyte fluorescence recognition method based on deep learning

By combining multimodal image acquisition and generative adversarial network virtual reconstruction technology with discriminative fusion network, the problems of data dependence and insufficient generalization ability of existing dermatophyte fluorescence identification methods are solved, realizing accurate, stable and interpretable automated identification of dermatophytes and improving the identification rate of rare species.

CN121640461APending Publication Date: 2026-03-10FIRST HOSPITAL AFFILIATED TO GENERAL HOSPITAL OF PLA
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-25
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing methods for identifying dermatophytes rely excessively on single-modal fluorescence image information, leading to inherent limitations of deep learning models in terms of data dependence, generalization ability, and decision interpretability. These models are vulnerable to complex real-world samples, especially exhibiting low recognition rates for rare species.

Method used

By employing multimodal image acquisition and generative adversarial network virtual reconstruction technology, combined with discriminative fusion network, we can achieve accurate, stable, and interpretable automated identification of dermatophytes. Multimodal images are acquired through dual-modal imaging channels, and virtual modal samples are generated through adversarial training of the generator and discriminator. Multimodal features are then fused for identification.

Benefits of technology

It improves the generalization ability to staining differences and rare morphologies, enhances the recognition rate of rare fungal species, and achieves efficient identification of dermal fungi under single fluorescent image input conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121640461A_ABST
    Figure CN121640461A_ABST
Patent Text Reader

Abstract

The invention provides a dermatophyte fluorescence recognition method based on deep learning, and relates to the technical field of image recognition, and the method comprises the steps: carrying out the multi-modal image collection of a target recognition region of dermatophyte, and obtaining a multi-modal image; analyzing a recognition rate of a second mode in the virtual first mode based on the multi-mode image to obtain an image recognition rate; and constructing a fluorescence recognition system based on the image recognition rate, and carrying out fluorescence recognition. According to the application, the technical problem of low recognition rate of skin fungus fluorescence recognition on rare strains in the prior art can be solved, and multi-modal information is virtually reconstructed through the generative adversarial network under the condition of only using single fluorescence image input and is cooperated with the discriminant fusion network; the technical target of accurate, stable and explainable automatic identification of the dermatophyte is completed, and the technical effects of improving the generalization ability of dyeing differences and rare forms and improving the identification rate of rare strains are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image recognition technology, and in particular to a deep learning-based method for fluorescent recognition of skin fungi. Background Technology

[0002] In the detection of dermatophyte infections, traditional microscopic examination and culture methods, while serving as the gold standard and possessing reference value, suffer from inherent limitations such as low sensitivity and lengthy processing times. To overcome these shortcomings, identification methods combining fluorescent staining and digital image analysis have emerged, significantly improving identification efficiency and representing an important direction for development in this field. However, in the process of this technological evolution, its inherent defects have gradually become apparent.

[0003] Currently, existing methods for identifying dermatophytes using fluorescence primarily rely on end-to-end model training and recognition of single-modality fluorescently stained images. While fluorescent staining significantly enhances the contrast between the target and background through specific labeling, this technique inherently contains a creative contradiction: the double-edged sword effect of "information simplification." Specifically, while highlighting specific components such as fungal chitin, fluorescent labeling strips away the original morphological information of hyphae—natural pigments, fine textures, and interactions with surrounding tissues. This information is precisely the key multidimensional feature that experienced laboratory technicians rely on for species identification. Consequently, the model learning essentially involves simplified "brightness and shape patterns" rather than a complete biological morphological representation, resulting in lower generalization ability and robustness compared to human experts when faced with uneven staining, non-specific fluorescence, or rare morphological variations.

[0004] In summary, existing technologies suffer from technical problems such as vulnerability to complex real-world samples and low recognition rates for rare bacterial species due to over-reliance on simplified single-modal fluorescence image information and the inherent limitations of deep learning models in terms of data dependence, generalization ability, and decision interpretability. Summary of the Invention

[0005] The purpose of this application is to provide a deep learning-based fluorescent identification method for dermatophytes, in order to solve the technical problems in the prior art, which are that the identification system is vulnerable to complex real-world samples and has a low recognition rate for rare fungi due to the over-reliance on simplified single-modal fluorescence image information and the inherent limitations of deep learning models in terms of data dependence, generalization ability and decision interpretability.

[0006] In view of the above problems, this application provides a deep learning-based fluorescent recognition method for dermatophytes, comprising: performing multimodal image acquisition on the target recognition region of dermatophytes to obtain a multimodal image; analyzing the recognition rate of a second modality under a virtual first modality based on the multimodal image to obtain an image recognition rate; and constructing a fluorescent recognition system based on the image recognition rate to perform fluorescent recognition.

[0007] Preferably, the deep learning-based method for fluorescent identification of skin fungi further includes: configuring a dual-modal imaging channel, wherein the dual-modal imaging channel is switched between a first imaging channel and a second imaging channel via a modality switching device; acquiring a first modal image of the target recognition region based on the first imaging channel to obtain a first modal image; switching the channel via the modality switching device to acquire a second modal image of the target recognition region based on the second imaging channel to obtain a second modal image; and pairing the second modal image and the second modal image to obtain the multimodal image.

[0008] Preferably, the deep learning-based method for identifying dermal fungi with fluorescence further includes: constructing a paired sample set, an unpaired sample set, a single-fluorescence sample set, and a single-bright-field sample set for the sample identification region; constructing an image translation framework, including a first generator, a second generator, a first discriminator, and a second discriminator; jointly training the first generator and the second generator using an adversarial training method based on the paired sample set and the unpaired sample set to obtain a trained first generator and a trained second generator; inputting the single-fluorescence sample set and the single-bright-field sample set into the trained first generator and the trained second generator to generate virtual images, outputting virtual first modality samples and virtual second modality samples; comparing the virtual second modality samples with the single-fluorescence sample set in the first discriminator and outputting a fluorescence discrimination value; comparing the virtual first modality samples with the single-bright-field sample set in the second discriminator and outputting a bright-field discrimination value; and obtaining an image translation module based on the fluorescence discrimination value and the bright-field discrimination value.

[0009] Preferably, the deep learning-based dermatophyte fluorescence identification method further includes: the sample identification region includes positive features of dermatophytes.

[0010] Preferably, the deep learning-based method for identifying dermal fungi fluorescence further includes: the first generator being used to convert a first modality image into a second modality image, and the second generator being used to convert a second modality image into a first modality image.

[0011] Preferably, the deep learning-based method for identifying dermatophytes using fluorescence further includes: performing a first-stage training on the multimodal fusion network framework using the paired sample set to obtain a first-stage training framework; and performing a second-stage training on the first-stage training framework using the single fluorescence sample set, the virtual first modality sample, and the single bright-field sample set to obtain the multimodal fusion network.

[0012] Preferably, the deep learning-based method for identifying dermatophyte fluorescence further includes: the multimodal fusion network framework is constructed by concatenating the first feature extracted from the first branch and the second feature extracted from the second branch by a tail-joining feature fusion module.

[0013] Preferably, the deep learning-based method for identifying dermal fungi fluorescence further includes: the second stage training includes replacing the single bright field sample set with the virtual first modality sample with a preset probability.

[0014] Preferably, the deep learning-based method for identifying dermal fungi fluorescence further includes: inputting the multimodal image into the image translation module to generate a virtual image, obtaining a virtual first modality image and a virtual second modality image; inputting the multimodal image into the multimodal fusion network for feature extraction to obtain real features; inputting the virtual first modality image and the virtual second modality image into the multimodal fusion network for feature extraction to obtain virtual features; comparing the virtual features and the real features, and obtaining the image recognition rate based on the feature comparison value.

[0015] Preferably, the deep learning-based method for fluorescent identification of skin fungi further includes: establishing the fluorescence identification system based on the image recognition rate, combined with the image translation module and the multimodal fusion network, wherein the fluorescence identification system is used to access the second imaging channel in the cloud and identify a single fluorescence image set.

[0016] The technical solution provided in this application has at least the following technical effects or advantages: by achieving the technical goal of accurate, stable and interpretable automated identification of dermatophytes by virtually reconstructing multimodal information through generative adversarial networks and cooperating with discriminative fusion networks under the condition of using only a single fluorescent image input, the technical effect of improving the generalization ability of staining differences and rare morphologies and improving the identification rate of rare fungi is achieved.

[0017] The above description is merely an overview of the technical solution of this application. To enable a clearer understanding of the technical means of this application and to facilitate its implementation according to the description, and to make the above and other objects, features, and advantages of this application more apparent, specific embodiments of this application are described below. It should be understood that the content described in this section is not intended to identify key or important features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent through the following description. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely exemplary. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0019] Figure 1 This is a flowchart illustrating a deep learning-based fluorescent recognition method for skin fungi according to this application.

[0020] Figure 2 This is a schematic diagram illustrating the process of obtaining multimodal images in a deep learning-based method for fluorescence recognition of skin fungi according to this application. Detailed Implementation

[0021] This application provides a deep learning-based fluorescence identification method for dermatophytes, addressing the technical problems in existing technologies. These problems stem from over-reliance on simplified single-modal fluorescence image information and the inherent limitations of deep learning models in data dependence, generalization ability, and decision interpretability. These limitations lead to the vulnerability of identification systems to complex real-world samples and low recognition rates for rare species. The method achieves the technical goal of accurate, stable, and interpretable automated identification of dermatophytes using only a single fluorescence image input. This is accomplished through the virtual reconstruction of multimodal information via a generative adversarial network and its collaboration with a discriminative fusion network. This improves the generalization ability to staining differences and rare morphologies, and enhances the recognition rate of rare species.

[0022] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. It should be understood that this application is not limited to the exemplary embodiments described herein. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application. It should also be noted that, for ease of description, only the parts related to this application are shown in the accompanying drawings, not all of them.

[0023] Please see Figure 1 and Figure 2 This application provides a deep learning-based fluorescent identification method for dermatophytes, which specifically includes the following steps:

[0024] S1: Perform multimodal image acquisition on the target recognition region of skin fungi to obtain multimodal images.

[0025] Furthermore, this application also includes: configuring a dual-modal imaging channel, wherein the dual-modal imaging channel is switched between a first imaging channel and a second imaging channel via a modal switching device; acquiring a first modal image of the target recognition region based on the first imaging channel to obtain a first modal image; switching the channel via the modal switching device to acquire a second modal image of the target recognition region based on the second imaging channel to obtain a second modal image; and pairing the second modal image and the second modal image to obtain the multimodal image.

[0026] Specifically, configuring a dual-modal imaging channel refers to setting up two different physical or logical imaging paths for an image acquisition device. A dual-modal imaging channel includes a first imaging channel capable of generating a first-modal image and a second imaging channel capable of generating a second-modal image. These two channels correspond to different optical imaging principles or signal acquisition mechanisms. To achieve the orderly operation of the two imaging modes, a modality switching device is integrated. This device can be a mechanical filter wheel, an electric turntable, or an electronically controlled optical path selector. It can accurately and quickly switch between the first and second imaging channels according to instructions, enabling the same device to acquire image data from different modalities sequentially.

[0027] Furthermore, acquiring the first modal image of the target recognition region based on the first imaging channel refers to activating the specific imaging conditions corresponding to that channel. For example, when the first modality is defined as bright-field imaging, a transmitted illumination method is used, and a visible light source is used to illuminate the sample. The image sensor captures the light passing through the sample, thereby obtaining a grayscale or color image that reflects the natural absorption and refraction characteristics of the sample, i.e., the first modal image, which then presents the basic morphological structure and tissue background of the target recognition region.

[0028] A channel switching operation is performed via a modality switching device, transitioning from the first imaging channel to the second imaging channel. Subsequently, a second modality image of the target recognition region is acquired based on the second imaging channel, applying a different set of imaging parameters. Taking fluorescence imaging as an example, an excitation light source is activated to illuminate the sample labeled with fluorescent dye. A matching emission filter is used, and the image sensor selectively captures the fluorescence signal emitted by the sample after excitation, resulting in a high-contrast second modality image. Structures with specific markings within the target recognition region will exhibit bright signals, significantly distinguishing them from the background.

[0029] Finally, the second modality image, acquired sequentially from the same target recognition region, was correlated and integrated with the first modality image to ensure that the two images correspond completely in spatial viewpoint, representing two different information representations of the same microscopic region. The resulting set of images constitutes a multimodal image, providing a complementary data source that contains both rich morphological details and highly specific markers for subsequent analysis.

[0030] S2: Based on the multimodal image analysis, the recognition rate of the second mode under the virtual first mode is obtained to obtain the image recognition rate.

[0031] Furthermore, this application also includes: constructing a paired sample set, an unpaired sample set, a single fluorescence sample set, and a single bright-field sample set for the sample recognition region; constructing an image translation framework, including a first generator, a second generator, a first discriminator, and a second discriminator; jointly training the first generator and the second generator using an adversarial training method based on the paired sample set and the unpaired sample set to obtain a trained first generator and a trained second generator; inputting the single fluorescence sample set and the single bright-field sample set into the trained first generator and the trained second generator to generate virtual images, and outputting virtual first modality samples and virtual second modality samples; comparing the virtual second modality samples with the single fluorescence sample set in the first discriminator and outputting a fluorescence discrimination value; comparing the virtual first modality samples with the single bright-field sample set in the second discriminator and outputting a bright-field discrimination value; and obtaining an image translation module based on the fluorescence discrimination value and the bright-field discrimination value.

[0032] Furthermore, this application also includes: the sample identification region includes positive features of dermatophytes.

[0033] Furthermore, this application also includes: the first generator is used to convert the first modal image into a second modal image, and the second generator is used to convert the second modal image into the first modal image.

[0034] Furthermore, this application also includes: performing a first-stage training on the multimodal fusion network framework using the paired sample set to obtain a first-stage training framework; and performing a second-stage training on the first-stage training framework using the single fluorescence sample set, the virtual first modality sample, and the single bright-field sample set to obtain a multimodal fusion network.

[0035] Furthermore, this application also includes: the multimodal fusion network framework is formed by splicing the first feature extracted from the first branch and the second feature extracted from the second branch by the tail feature fusion module.

[0036] Furthermore, this application also includes: the second stage of training includes replacing the single bright field sample set with the virtual first modality sample with a preset probability.

[0037] Furthermore, this application also includes: inputting the multimodal image into the image translation module to generate a virtual image, obtaining a virtual first modal image and a virtual second modal image; inputting the multimodal image into the multimodal fusion network to extract features, obtaining real features; inputting the virtual first modal image and the virtual second modal image into the multimodal fusion network to extract features, obtaining virtual features; comparing the virtual features and the real features, and obtaining the image recognition rate based on the feature comparison value.

[0038] Specifically, the sample identification region refers to the area observable under a microscope and verified by mycology to contain dermatophyte infection, i.e., the area with positive characteristics of dermatophytes. A dataset of sample identification regions for model training and evaluation is constructed, including paired sample sets, unpaired sample sets, single-fluorescence sample sets, and single-bright-field sample sets. Paired sample sets refer to paired image data, where each pair of images is acquired for the same confirmed sample identification region. Each pair of images contains one first-modal image obtained through a first imaging channel and one second-modal image obtained through a second imaging channel, spatially aligned. Unpaired sample sets refer to independent first-modal and second-modal images that are not precisely spatially paired, originating from sample identification regions with positive characteristics, but without a one-to-one correspondence between them. Single-fluorescence sample sets refer to collections containing only second-modal (fluorescent) images. Single-bright-field sample sets refer to collections containing only first-modal (bright-field) images.

[0039] Furthermore, an image translation framework is constructed based on the principles of generative adversarial networks (GANs). The first generator is a deep neural network model that receives a first modality image as input and learns to convert it into an output that visually resembles a second modality image. Correspondingly, the second generator is another deep neural network model that receives a second modality image as input and learns to convert it into an output that visually resembles a first modality image. To guide and optimize the training process of the first and second generators, the image translation framework also includes a first discriminator and a second discriminator. Both the first and second discriminators are also neural network models. The first discriminator's task is to distinguish between real second modality images and virtual second modality images generated by the first generator, while the second discriminator's task is to distinguish between real first modality images and virtual first modality images generated by the second generator. The first and second generators are jointly trained using an adversarial training method based on paired and unpaired sample sets. The adversarial training method is a machine learning training paradigm used to enable the generator and discriminator to jointly optimize through competition. The generator strives to produce realistic images sufficient to fool the discriminator, while the discriminator continuously improves its ability to distinguish between real and fake images. Paired sample sets provide an explicit input-output mapping for supervised learning, while unpaired sample sets help improve the generalization ability of the image translation framework. Through joint training with both paired and unpaired sample sets, the first and second generators after parameter optimization are finally obtained, each possessing relatively reliable modality transfer capabilities.

[0040] Subsequently, to evaluate the generator's performance and solidify the image translation module, independent single-fluorescence sample sets and single-bright-field sample sets were input into the trained generator. Specifically, the single-fluorescence sample set was input into the trained second generator to process each second modality image and output an image simulating the features of the first modality, serving as a virtual first modality sample. Inputting the single-bright-field sample set into the trained first generator would output the corresponding virtual second modality sample.

[0041] Next, the virtual second modality samples are compared with the real single-fluorescent sample set in the first discriminator. The first discriminator evaluates the authenticity of each input image and outputs a probability value that the image is judged as a real second modality image, i.e., the fluorescence discrimination value, which reflects how close the virtual second modality sample is to the real fluorescence image in terms of visual features.

[0042] Simultaneously, the virtual first modality sample is compared with the real single brightfield sample set in the second discriminator. The evaluation process performed by the second discriminator outputs a probability value indicating that the image is judged as a real first modality image, i.e., the brightfield discrimination value, which reflects the similarity between the virtual first modality sample and the real brightfield image.

[0043] Finally, based on the fluorescence discrimination value and the bright field discrimination value, the overall performance of the image translation framework can be quantitatively evaluated, reflecting the quality and realism of the images generated by the two generators. When the fluorescence discrimination value and the bright field discrimination value reach the preset threshold standard, it indicates that the system composed of the first generator, the second generator, the first discriminator and the second discriminator has a stable and high-quality image mode conversion capability, and at this time it can be confirmed as the finally usable image translation module.

[0044] Furthermore, the paired sample set refers to a dataset containing pairs of first-modality images and second-modality images. The first-modality images and second-modality images originate from the same sample recognition region and are spatially precisely aligned. The multimodal fusion network framework refers to a pre-defined initial architecture of a deep neural network with a dual-branch parallel processing structure. The first-stage training framework is obtained by training the multimodal fusion network framework using the paired sample set. The first-stage training refers to the initial parameter optimization of the network framework using supervised learning with the paired sample set. Effective feature representations are extracted separately, and the correlation between the two modalities is initially established. The optimized network state is called the first-stage training framework.

[0045] After completing the first stage of training, the first-stage training framework is trained in the second stage using a single-fluorescence sample set, virtual first-modality samples, and a single bright-field sample set to obtain the final multimodal fusion network. The single-fluorescence sample set contains only second-modality images, the single bright-field sample set contains only first-modality images, and the virtual first-modality samples specifically refer to image data simulating first-modality features generated by the second generator in the trained image translation module after processing the single-fluorescence sample set. The second-stage training is a further optimization process based on the first-stage training framework's existing basic feature extraction capabilities. By introducing a combination of virtual and single-modality data, the aim is to adapt the network to more complex data input scenarios that closely resemble real-world applications. The network model that completes this stage of training is the multimodal fusion network.

[0046] The multimodal fusion network framework consists of a concatenated feature fusion module that combines the first features extracted from the first branch and the second features extracted from the second branch. Specifically, the first branch is a sub-network in the multimodal fusion network framework used to process the first modality image. It extracts a high-level, abstract feature representation, i.e., the first feature, from the input image through operations such as convolutional layers. The second branch is a parallel sub-network in the network used to process the second modality image and extract the second feature. The feature fusion module is a network layer located at the concatenation point between the first and second branches. It receives the first and second features as input and then performs a concatenation operation to form a composite feature representation that integrates bimodal information for subsequent classification or recognition tasks.

[0047] The second phase of training replaces the single bright-field sample set with virtual first-modality samples with a preset probability. This preset probability is a value between 0 and 1, controlling the frequency of the replacement event. During each training iteration or batch of data preparation in the second phase, decisions are made randomly based on this preset probability. When a decision is triggered, the real first-modality image that should have been input into the multimodal fusion network framework—the image from the single bright-field sample set—is replaced by the corresponding virtual first-modality sample. This introduces simulated data created by the generative model into the training process, forcing the multimodal fusion network framework to learn not only feature mappings based on real paired data during training, but also how to utilize and trust the information provided by the virtual first-modality image generated from the second-modality image. This significantly enhances the network's robustness and generalization ability in real-world application scenarios with only a single second-modality image input.

[0048] The multimodal images are input into the image translation module for virtual image generation, resulting in virtual first-modal images and virtual second-modal images. Specifically, the real first-modal image from the multimodal images is input into the first generator, which outputs an image that corresponds to the original image in content but imitates the second modality in imaging style and features—that is, a virtual second-modal image. The real second-modal image from the multimodal images is input into the second generator, which obtains the corresponding virtual first-modal image. Through the learned mapping relationship, virtual image pairs that are paired with the original image but with interchanged modalities are created.

[0049] Next, the same set of multimodal images, i.e., real image pairs, are input into the multimodal fusion network for feature extraction to obtain real features. The multimodal fusion network is a classification and recognition network that has been trained in two stages and has the ability to integrate dual-branch inputs and feature fusion. Feature extraction refers to the dual-branch processing flow of the multimodal fusion network. The first branch receives the real first modality image and extracts its deep feature representation, called the first real feature; the second branch receives the real second modality image and extracts its deep feature representation, called the second real feature. Subsequently, the feature fusion module at the end of the multimodal fusion network integrates the virtual first modality image and the virtual second modality image. The final output fused feature vector is the real feature, representing a comprehensive information representation for recognition and judgment distilled from the original real multimodal data.

[0050] Then, the virtual first modality image and the virtual second modality image are jointly input into the multimodal fusion network for feature extraction to obtain virtual features. The virtual image pairs are then used as new inputs to the multimodal fusion network. The first branch of the multimodal fusion network receives the virtual first modality image and extracts the corresponding first virtual feature; the second branch receives the virtual second modality image and extracts the second virtual feature. The feature fusion module processes and merges the two sets of virtual features, ultimately generating a fused feature vector, which serves as the virtual feature, representing the information representation learned from the simulated multimodal data generated by the image translation module.

[0051] Finally, by comparing virtual and real features, the image recognition rate is obtained based on the feature comparison values. The comparison operation refers to calculating the similarity or distance metric between virtual and real features within the feature space. The feature comparison value is one or more quantitative indicators, such as cosine similarity, the reciprocal of Euclidean distance, or a confidence score calculated through a dedicated evaluation subnetwork. By statistically analyzing the feature comparison values ​​of a large number of test samples, the extent to which virtual features can replace or approximate the classification information carried by real features can be evaluated. The image recognition rate is the core performance indicator derived from this. When using virtual features for classification decisions, it represents the accuracy or recall achieved in classifying the data. The recognition rate quantitatively reflects the effectiveness of the virtual modal data generated by the image translation module in maintaining the original recognition performance of the multimodal fusion network, and is a key basis for measuring whether it can still work reliably when real multimodal data is missing.

[0052] S3: Construct a fluorescence recognition system based on the image recognition rate to perform fluorescence recognition.

[0053] Furthermore, this application also includes: based on the image recognition rate, combined with the image translation module and the multimodal fusion network, establishing the fluorescence recognition system, wherein the fluorescence recognition system is used to access the second imaging channel in the cloud and recognize a single fluorescence image set.

[0054] Specifically, image recognition rate is a quantitative performance evaluation metric that reflects the accuracy of completing recognition tasks with the assistance of virtual modal data. Based on the image recognition rate, an image translation module and a multimodal fusion network are combined—that is, the image translation module responsible for cross-modal image generation and the multimodal fusion network responsible for final classification and recognition—to establish a fluorescence recognition system according to a preset logic and data flow. This builds a complete and executable software application or service platform. The process includes determining the calling interfaces between modules, designing the data processing flow, and encapsulating the overall functionality.

[0055] The fluorescence recognition system is designed to access a cloud-based second imaging channel to identify single fluorescence image sets. Access to the cloud-based second imaging channel means that the fluorescence recognition system has the ability to communicate with image acquisition devices or image storage services deployed on a cloud server, and can receive image data streams or files uploaded from remote terminals via the network. The second imaging channel specifically refers to the logical path or data interface configured in the cloud for acquiring or receiving second-modality, i.e., fluorescence modality images. Identifying single fluorescence image sets is the core function of this system, meaning that the fluorescence recognition system can automatically analyze and diagnose received sets containing only fluorescence modality images. Specifically, when one or more fluorescence images are input through the cloud channel, the fluorescence recognition system first processes them using a second generator in the integrated image translation module to generate a corresponding virtual first-modality image. Subsequently, the original fluorescence image and its corresponding virtual bright-field image are jointly input into a multimodal fusion network. By extracting and fusing dual-path features, the final output is the classification result of the dermatophytes in the image.

[0056] In summary, the deep learning-based fluorescence identification method for dermatophytes provided in this application has the following technical effects: by achieving the technical goal of accurate, stable, and interpretable automated identification of dermatophytes through the virtual reconstruction of multimodal information using a generative adversarial network and in collaboration with a discriminative fusion network under the condition of using only a single fluorescence image input, it achieves the technical effect of improving the generalization ability of staining differences and rare morphologies and improving the identification rate of rare fungal species.

[0057] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0058] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of this application and its equivalents, this application also intends to include such modifications and variations.

Claims

1. A deep learning-based skin fungus fluorescence recognition method, characterized in that, The method comprises the following steps: performing multi-modal image acquisition on a target recognition area of skin fungus to obtain multi-modal images; analyzing the recognition rate of the second modality under the virtual first modality based on the multi-modal images to obtain an image recognition rate; constructing a fluorescence recognition system based on the image recognition rate to perform fluorescence recognition.

2. The skin fungus fluorescence recognition method based on deep learning according to claim 1, wherein, Performing multi-modal image acquisition on a target recognition area of skin fungus to obtain multi-modal images, comprising: configuring a dual-modality imaging channel, which switches between a first imaging channel and a second imaging channel through a modality switching device; performing first modality image acquisition of the target recognition area based on the first imaging channel to obtain a first modality image; switching channels through the modality switching device, and performing second modality image acquisition of the target recognition area based on the second imaging channel to obtain a second modality image; pairing the second modality image and the second modality image to obtain the multi-modal image.

3. The deep learning-based fluorescent recognition method for skin fungus according to claim 2, wherein, Before analyzing the recognition rate of the second modality under the virtual first modality based on the multi-modal images, comprising: constructing a paired sample set, an unpaired sample set, a single fluorescence sample set and a single bright field sample set of a sample recognition area; constructing an image translation framework, including a first generator, a second generator, a first discriminator and a second discriminator; based on the paired sample set and the unpaired sample set, the first generator and the second generator are jointly trained by an adversarial training method to obtain a trained first generator and a trained second generator; inputting the single fluorescence sample set and the single bright field sample set into the trained first generator and the trained second generator to generate virtual images, and outputting virtual first modality samples and virtual second modality samples; comparing the virtual second modality samples with the single fluorescence sample set at the first discriminator to output fluorescence discrimination values; comparing the virtual first modality samples with the single bright field sample set at the second discriminator to output bright field discrimination values; based on the fluorescence discrimination values and the bright field discrimination values, an image translation module is obtained.

4. The skin fungus fluorescence recognition method based on deep learning according to claim 3, characterized in that, The sample recognition area includes positive features with skin fungus.

5. The skin fungus fluorescence recognition method based on deep learning according to claim 3, characterized in that, The first generator is used to convert the first modality image into the second modality image, and the second generator is used to convert the second modality image into the first modality image.

6. The skin fungus fluorescence recognition method based on deep learning according to claim 3, characterized in that, Before analyzing the recognition rate of the second modality under the virtual first modality based on the multi-modal images, further comprising: training a multi-modal fusion network framework with the paired sample set in a first stage to obtain a first stage training framework; training the first stage training framework with the single fluorescence sample set, the virtual first modality sample and the single bright field sample set in a second stage to obtain a multi-modal fusion network.

7. The skin fungus fluorescence recognition method based on deep learning according to claim 6, characterized in that, The multi-modal fusion network framework splices the first feature extracted by the first branch and the second feature extracted by the second branch through a tailing feature fusion module.

8. The skin fungus fluorescence recognition method based on deep learning according to claim 6, wherein, The second stage training includes replacing the single bright field sample set with the virtual first modality sample with a preset probability.

9. The skin fungus fluorescence recognition method based on deep learning according to claim 6, wherein, Analyzing the recognition rate of the second modality under the virtual first modality based on the multi-modal images to obtain an image recognition rate, comprising: Input the multi-modal image into the image translation module for virtual image generation to obtain a virtual first modality image and a virtual second modality image; Input the multi-modal image into the multi-modal fusion network for feature extraction to obtain real features; Input the virtual first modality image and the virtual second modality image into the multi-modal fusion network for feature extraction to obtain virtual features; Compare the virtual features and the real features, and obtain the image recognition rate based on a feature comparison value.

10. The skin fungus fluorescence recognition method based on deep learning according to claim 9, wherein, Based on the image recognition rate, a fluorescence recognition system is constructed for fluorescence recognition, including: Based on the image recognition rate, the image translation module and the multi-modal fusion network are combined to establish the fluorescence recognition system, and the fluorescence recognition system is used for accessing a cloud second imaging channel to recognize a single fluorescence image set.

Citation Information

Patent Citations

  • Target re-identification method based on generative multi-modal image fusion

    CN116824625A

  • CycleGan-based colposcope image modal conversion method

    CN117437514A

  • Mark-free single-cell two-dimensional light scattering imaging modal amplification method and system

    CN120182965A

  • Fluorescence image recognition system for detecting fungi based on deep learning algorithm

    CN120823597A

  • Method and system for digital staining of microscopy images using deep learning

    US20230030424A1