Training Method Based on Complementary Attention

The complementary attention mechanism in a three-layered neural network architecture addresses the issue of CNNs overlooking lesions in medical images, improving detection accuracy by optimizing the network with a total loss function.

CN114972278BActive Publication Date: 2025-07-15SHENZHEN SIBRIGHT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210631057.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-11-28
Filing Date
2020-11-27
Publication Date
2025-07-15
Estimated Expiration
2040-11-27

AI Technical Summary

Technical Problem

When identifying lesion sites such as the fundus, existing convolutional neural networks are prone to ignore lesion areas with low attention, resulting in low accuracy in tissue lesion recognition.

Method used

Using a training method based on complementary attention mechanism, by preparing the training data set, multiple artificial neural networks are used for feature extraction and recognition, combining attention heat maps and complementary attention heat maps, the artificial neural network module is optimized to improve recognition accuracy.

Benefits of technology

The accuracy of identification of tissue lesions is improved, especially in the recognition of fundus lesions, and the recognition effect of lesion areas is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114972278B_ABST
    Figure CN114972278B_ABST
Patent Text Reader

Abstract

The present disclosure describes a training method based on complementary attention, including: preparing a training data set, which includes multiple examination images with lesions and annotation images associated with the examination images and having lesion annotation results; extracting features from the examination images to obtain feature maps, and processing the examination images based on an attention mechanism to obtain attention heat maps; classifying the examination images using a first artificial neural network, and combining with the annotation images to obtain a first loss function, using a second artificial neural network module to classify the examination images based on the feature maps and attention heat maps, and combining with the annotation results to obtain a second loss function, using a third artificial neural network to perform disease-free discrimination on the examination images to obtain a third loss function. By combining these three loss functions, the accuracy of tissue lesion recognition for tissue lesions can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the application with the application date of November 27, 2020, application number 2020113599702, and invention name of "Training Method and Training System for Tissue Lesion Recognition Based on Artificial Neural Network". Technical Field

[0002] This disclosure generally relates to a training method based on complementary attention. Background Art

[0003] With the development and maturity of artificial intelligence technology, artificial intelligence technology has been gradually popularized in various aspects of the medical field. In particular, medical imaging in medicine is a relatively popular field for the application of artificial intelligence technology currently. Medical imaging is a useful tool for diagnosing many diseases. A large amount of medical image data will be generated during the medical imaging process. It takes a lot of time for doctors to process and identify these image data, and it is difficult to ensure the accuracy of the identification. In medical images, artificial intelligence technology is mainly used to identify tissue lesions in the images to improve the accuracy of tissue lesion identification.

[0004] Currently, a convolutional neural network (CNN) is usually adopted in the application of artificial intelligence technology to identify medical images. The convolutional structure of the convolutional neural network can reduce the memory occupied by the deep network. It has three key operations. One is the local receptive field, the second is weight sharing, and the third is the pooling layer. Thus, it can effectively reduce the number of network parameters and alleviate the overfitting problem of the convolutional neural network. The structure of the convolutional neural network can better adapt to the structure of medical images and extract and identify features.

[0005] However, for some lesion sites such as fundus lesion sites, the lesion areas are relatively small and irregularly distributed. Generally, a convolutional neural network applying an attention mechanism tends to ignore the lesion areas with low attention in the attention heat map, resulting in misjudgment, so that the accuracy of tissue lesion identification in these lesion areas is relatively low. Summary of the Invention

[0006] In view of the above-mentioned status of the prior art, this disclosure is completed, and its purpose is to provide a training method and training system for tissue lesion recognition based on artificial neural network that can effectively improve the accuracy of tissue lesion recognition.

[0007] To this end, a first aspect of the present disclosure provides a training method for tissue lesion recognition based on an artificial neural network, which is characterized by including: preparing a training data set, the training data set including a plurality of inspection images and annotation images associated with the inspection images, the annotation images including annotation results of lesions or annotation results of no lesions; inputting the training data set into an artificial neural network module to extract features from the inspection images to obtain feature maps, and processing the feature maps based on an attention mechanism to obtain attention heat maps, and processing the attention heat maps based on a complementary attention mechanism to obtain complementary attention heat maps; the artificial neural network module includes a first artificial neural network, a second artificial neural network, and a third artificial neural network; using the first artificial neural network to extract features from the inspection images to obtain the feature maps, using the second artificial neural network to obtain the attention heat maps indicating the lesion regions and the complementary attention heat maps indicating the non-lesion regions, the inspection images being composed of the lesion regions and the non-lesion regions, using the third artificial neural network to recognize the inspection images based on the feature maps to obtain a first recognition result, using the third artificial neural network to recognize the inspection images based on the feature maps and the attention heat maps to obtain a second recognition result, using the third artificial neural network to recognize the inspection images based on the feature maps and the complementary attention heat maps to obtain a third recognition result; combining the first recognition result with the annotation images to obtain a first loss function when the attention mechanism is not used, combining the second recognition result with the annotation images to obtain a second loss function when the attention mechanism is used, combining the third recognition result with the annotation images having the annotation results of no lesions to obtain a third loss function when the complementary attention mechanism is used, using the first loss function, the second loss function, and the third loss function to obtain a total loss function including a first loss term based on the first loss function, a second loss term based on the difference between the second loss function and the first loss function, and a third loss term based on the third loss function, and using the total loss function to optimize the artificial neural network module. In this case, a first recognition result, a second recognition result, and a third recognition result can be obtained, and a total loss function can be obtained based on the first recognition result, the second recognition result, and the third recognition result, so that the artificial neural network module can be optimized using the total loss function, thereby improving the accuracy of tissue lesion recognition of the artificial neural network module.

[0008] In addition, in the training method for tissue lesion recognition based on an artificial neural network according to the first aspect of the present disclosure, optionally, the total loss function further includes a total area term of the attention heat map, and the total area term is used to evaluate the area of the lesion region. In this case, the fifth loss term can be used to evaluate the area of the lesion region in the attention heat map and control the number of pixels in the attention heat map that have a greater impact on the recognition result, so that the attention of the network is restricted to the pixels that have a greater impact on the recognition result.

[0009] In addition, in the training method for tissue lesion recognition based on an artificial neural network according to the first aspect of the present disclosure, optionally, the total loss function further includes a regularization term for the attention heat map. In this case, overfitting of the artificial neural network module can be suppressed.

[0010] In addition, in the training method for tissue lesion recognition based on an artificial neural network according to the first aspect of the present disclosure, optionally, the first artificial neural network, the second artificial neural network, and the third artificial neural network are trained simultaneously. In this case, the training speed can be accelerated.

[0011] In addition, in the training method for tissue lesion recognition based on an artificial neural network according to the first aspect of the present disclosure, optionally, the third artificial neural network includes an input layer, an intermediate layer, and an output layer connected in sequence, and the output layer is configured to output a recognition result reflecting the inspection image. In this case, the third artificial neural network can be used to output a recognition result reflecting the tissue image.

[0012] In addition, in the training method for tissue lesion recognition based on an artificial neural network according to the first aspect of the present disclosure, optionally, the training method of the artificial neural network module is weakly supervised. In this case, a recognition result with more information can be obtained through the artificial neural network module using a labeling result with less information.

[0013] In addition, in the training method for tissue lesion recognition based on an artificial neural network according to the first aspect of the present disclosure, optionally, the first loss function is used to evaluate the degree of inconsistency between the recognition result of the inspection image without using the attention mechanism and the labeling result. In this case, the accuracy of tissue lesion recognition of the artificial neural network module without using the attention mechanism can be improved.

[0014] Additionally, in the training method for tissue lesion recognition based on an artificial neural network according to the first aspect of the present disclosure, optionally, the second loss function is used to evaluate the degree of inconsistency between the recognition result of the examination image when using the attention mechanism and the annotation result. In this case, the accuracy of tissue lesion recognition by the artificial neural network module when using the attention mechanism can be improved.

[0015] Additionally, in the training method for tissue lesion recognition based on an artificial neural network according to the first aspect of the present disclosure, optionally, the third loss function is used to evaluate the degree of inconsistency between the recognition result of the examination image when using the complementary attention mechanism and the annotation result of no lesion. In this case, the accuracy of tissue lesion recognition by the artificial neural network module when using the complementary attention mechanism can be improved.

[0016] Additionally, in the training method for tissue lesion recognition based on an artificial neural network according to the first aspect of the present disclosure, optionally, the total loss function is used to optimize the artificial neural network module to minimize the total loss function. In this case, the total loss function can be minimized to improve the accuracy of tissue lesion recognition by the artificial neural network module.

[0017] Additionally, in the training method for tissue lesion recognition based on an artificial neural network according to the first aspect of the present disclosure, optionally, the tissue lesion is an eye fundus lesion. In this case, the artificial neural network module can be used to obtain the recognition result of the eye fundus image regarding the eye fundus lesion.

[0018] The second aspect of the present disclosure provides a training system for tissue lesion recognition based on an artificial neural network, characterized in that the training method provided by the first aspect of the present disclosure is used for training. In this case, the artificial neural network module can be trained using the training system.

[0019] According to the present disclosure, a training method and a training system for tissue lesion recognition based on an artificial neural network that can effectively improve the accuracy of tissue lesion recognition can be provided. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Embodiments of the present disclosure will now be further explained in detail only by way of examples with reference to the accompanying drawings, where:

[0021] Figure 1 is a schematic diagram of an electronic device involved in an example of the present disclosure.

[0022] Figure 2 is a tissue image involved in an example of the present disclosure.

[0023] Figure 3It is a block diagram showing the structure of an identification system for tissue lesion identification based on an artificial neural network according to an example of the present disclosure.

[0024] Figure 4 It is a block diagram showing an example of an artificial neural network module according to an example of the present disclosure.

[0025] Figure 5 It is a block diagram showing a modified example of an artificial neural network module according to an example of the present disclosure.

[0026] Figure 6 It is a schematic diagram showing the structure of a first artificial neural network according to an example of the present disclosure.

[0027] Figure 7 It is a block diagram showing the structure of a training system for tissue lesion identification based on an artificial neural network according to an example of the present disclosure.

[0028] Figure 8 It is a flowchart showing a training method for tissue lesion identification based on an artificial neural network according to an example of the present disclosure.

[0029] Figure 9(a) is a schematic diagram showing an example of a fundus image obtained by training without using an attention mechanism according to an example of the present disclosure.

[0030] Figure 9(b) is a schematic diagram showing an example of a lesion area of a fundus image obtained by training using a complementary attention mechanism according to an example of the present disclosure.

[0031] Main reference numerals: 1... electronic device, 10... processor, 20... memory, 30... computer program, 40... identification system, 410... acquisition module, 4200... backbone neural network, 420... artificial neural network module, 421... first artificial neural network, 422... second artificial neural network, 423... third artificial neural network, 424... feature combination module, 430... training system, 431... storage module, 432... processing module, 433... optimization module, C1... first convolutional layer, C2... second convolutional layer, C3... third convolutional layer, S1... first pooling layer, S2... second pooling layer, S3... third pooling layer Detailed implementation manners

[0032] Hereinafter, preferred implementation manners of the present disclosure will be described in detail with reference to the drawings. In the following description, the same reference numerals are given to the same components, and repeated descriptions are omitted. In addition, the drawings are only schematic diagrams, and the dimensional ratios between components or the shapes of components may be different from the actual ones.

[0033] Figure 1 It is a schematic diagram of an electronic device according to an embodiment of the present disclosure.

[0034] As Figure 1 shown, the recognition system 40 for tissue lesion recognition based on an artificial neural network according to the present disclosure may be carried by an electronic device 1 (such as a computer). In some examples, the electronic device 1 may include one or more processors 10, a memory 20, and a computer program 30 arranged in the memory 20. Among them, the one or more processors 10 may include a central processing unit, a graphics processing unit, and any other electronic components capable of processing data. For example, the processor 10 may execute instructions stored on the memory 20.

[0035] In some examples, the memory 20 may be a computer-readable medium capable of carrying or storing data. In some examples, the memory 20 may include, but is not limited to, non-volatile memory or flash memory (Flash Memory), etc. In some examples, the memory 20 may also be, for example, ferroelectric random access memory (FeRAM), magnetic random access memory (MRAM), phase change random access memory (PRAM), or resistive random access memory (RRAM). Thereby, the possibility of data loss caused by sudden power failure can be reduced.

[0036] In other examples, the memory 20 may also be other types of readable storage media, such as read-only memory (Read-Only Memory, ROM), random access memory (Random Access Memory, RAM), programmable read-only memory (Programmable Read-only Memory, PROM), erasable programmable read-only memory (Erasable Programmable Read Only Memory, EPROM), one-time programmable read-only memory (One-time Programmable Read-Only Memory, OTPROM), electrically-erasable programmable read-only memory (Electrically-Erasable Programmable Read-Only Memory, EEPROM), compact disc read-only memory (Compact Disc Read-Only Memory, CD-ROM).

[0037] In some examples, the memory 20 may be an optical disc memory, a magnetic disk memory, or a tape memory. Thereby, the appropriate memory 20 can be selected according to different situations.

[0038] In some examples, the computer program 30 may include instructions executed by one or more processors 10, and by executing the instructions, the recognition system 40 can perform tissue lesion recognition on tissue images. In some examples, the computer program 30 may be deployed within a local computer or on a server in the cloud.

[0039] In some examples, the computer program 30 may be stored in a computer-readable medium. The computer-readable storage medium may include one or more of a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), or a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, and a magnetic storage device.

[0040] Figure 2 is a tissue image related to an example of the present disclosure. Figure 3 is a structural block diagram of the recognition system 40 for tissue lesion recognition based on an artificial neural network related to an example of the present disclosure.

[0041] In some examples, the tissue lesion recognition of the tissue image can be implemented using the recognition system 40 for tissue lesion recognition based on an artificial neural network, and a recognition result can be obtained. In some examples, the recognition system 40 for tissue lesion recognition may also be referred to as the recognition system 40.

[0042] In some examples, as Figure 3 shown, the recognition system 40 may include an acquisition module 410, an artificial neural network module 420, and a training system 430 for tissue lesion recognition based on an artificial neural network. In some examples, the acquisition module 410 may be configured to acquire tissue images. In some examples, the artificial neural network module 420 may be configured to perform processing such as feature extraction and tissue lesion recognition on the tissue images, and obtain a recognition result of the tissue lesion recognition. In some examples, the training system 430 for tissue lesion recognition based on an artificial neural network may be configured to train the artificial neural network module 420. In some examples, the training system 430 may utilize the first recognition result, the second recognition result, and the third recognition result obtained by the artificial neural network module 420, and obtain a total loss function based on the first recognition result, the second recognition result, and the third recognition result to optimize the artificial neural network module 420. In this case, the first recognition result, the second recognition result, and the third recognition result can be obtained, and a total loss function can be obtained based on the first recognition result, the second recognition result, and the third recognition result, so that the artificial neural network module 420 can be optimized using the total loss function, thereby improving the accuracy of tissue lesion recognition of the artificial neural network module 420.

[0043] In some examples, the training system 430 for tissue lesion recognition based on an artificial neural network may also be referred to as the training system 430.

[0044] In some examples, the recognition system 40 may further include a preprocessing module and a judgment module (not shown).

[0045] In some examples, the tissue image may be from a CT scan, a PET-CT scan, a SPECT scan, an MRI, an ultrasound, an X-ray, a mammogram, an angiogram, a fluorogram, an image of a tissue cavity taken by a capsule endoscope, or a combination thereof. In some examples, the tissue image may be acquired by the acquisition module 410.

[0046] In some examples, the acquisition module 410 may be configured to acquire a tissue image, which may be a tissue image acquired by a collection device such as a camera, an ultrasonic imager, or an X-ray scanner.

[0047] In some examples, the tissue image may be, for example, a fundus image, an esophageal image, a gastric image, a large intestine image, a colon image, or a small intestine image. As Figure 2 shown, the tissue image may be a fundus image. In this case, the fundus lesion recognition can be performed on the fundus image by the recognition system 40.

[0048] In some examples, the tissue lesion recognition may be to recognize the tissue lesion of the tissue image to obtain a recognition result.

[0049] In some examples, when the tissue image is a fundus image, the tissue lesion may be a fundus lesion. In this case, the recognition result of the fundus image regarding the fundus lesion can be obtained by using the artificial neural network module 420.

[0050] In some examples, the tissue image may be composed of a lesion area and a non-lesion area.

[0051] In some examples, a tissue image (color image) with a tissue lesion generally contains obvious features such as erythema and swelling. Therefore, the trained artificial neural network can be used to automatically extract and recognize these features to help the patient identify possible lesions. Thus, the accuracy and speed of recognition can be improved, and at the same time, problems such as large errors and long time consumption caused by human physicians reviewing images one by one based on their own experience can be reduced.

[0052] In some examples, when the tissue image is a fundus image, the tissue image can be classified according to its function. For example, in the training step, the tissue image can be an inspection image, a labeled image (described later).

[0053] In some examples, the image input to the artificial neural network module 420 can be a tissue image. In this case, the artificial neural network module 420 can be used to identify tissue lesions in the tissue image.

[0054] In some examples, the recognition system 40 can be used for identifying tissue lesions in tissue images. In some examples, when a tissue image enters the recognition system 40, operations such as preprocessing, feature extraction, and tissue lesion identification can be performed on the tissue image.

[0055] In some examples, the recognition system 40 may further include a preprocessing module and a judgment module. The preprocessing module can be used to preprocess the tissue image and input the preprocessed tissue image into the artificial neural network module 420.

[0056] In some examples, the preprocessing module can preprocess the tissue image. In some examples, the preprocessing may include at least one of region of interest detection, image cropping, size adjustment, and normalization. In this case, it is convenient for the subsequent artificial neural network module 420 to identify and judge tissue lesions in the tissue image. In some examples, the tissue image can be, for example, a fundus image, an esophagus image, a stomach image, a large intestine image, a colon image, or a small intestine image.

[0057] In some examples, the preprocessing module may include a region detection unit, an adjustment unit, and a normalization unit.

[0058] In some examples, the region detection unit can detect a region of interest from the tissue image. For example, if the tissue image is a fundus image, a fundus region centered on the optic disc, or a fundus region that includes the optic disc and is centered on the macula, etc., can be detected from the fundus image. In some examples, the region detection unit can detect the region of interest in the tissue image by, for example, a sampling threshold method, Hough transform.

[0059] In some examples, the adjustment unit can be used to crop and adjust the size of the tissue image. Since the devices used to acquire tissue images or the shooting conditions are different, the acquired tissue images may vary in terms of resolution, size, etc. In this case, these tissue images can be cropped and size-adjusted to reduce the differences. In some examples, the tissue image can be cropped according to a specific shape. In some examples, the specific shape can include but is not limited to square, rectangle, circle, or ellipse, etc.

[0060] In some other examples, the size of the tissue image can be adjusted to a specified size by an adjustment unit. For example, the specified size can be 256×256, 512×512, 1024×1024, etc. However, the examples of the present disclosure are not limited thereto. In some other examples, the size of the tissue image can also be of any other specification. For example, the size of the tissue image can be 128×128, 768×768, 2048×2048, etc.

[0061] In some examples, the preprocessing module may include a normalization unit. The normalization unit can be used to perform normalization processing on multiple tissue images.

[0062] In some examples, there is no particular limitation on the normalization method of the normalization unit. For example, zero mean, unit standard deviation, etc. can be adopted. Additionally, in some examples, it can also be normalized within the range of [0,1]. In this case, through normalization, the differences between different tissue images can be overcome.

[0063] In some examples, normalization includes normalization of image format, image slice interval, image intensity, image contract, and image orientation. In some examples, the tissue image can be normalized to DICOM format, NIfTI format, or raw binary format.

[0064] Figure 4 It is a block diagram showing an example of the artificial neural network module related to the examples of the present disclosure.

[0065] As described above, the recognition system 40 may include an artificial neural network module 420. In some examples, the artificial neural network module 420 can be used to identify tissue lesions in tissue images. In some examples, the artificial neural network module 420 may include multiple artificial neural networks. In some examples, one or more processors 10 can be used to train the artificial neural network. Generally speaking, an artificial neural network can include artificial neurons or nodes, which can be used to receive tissue images and perform operations on the tissue images based on weights, and then selectively transfer the operation results to other neurons or nodes. Among them, the weights can be associated with the artificial neurons or nodes and simultaneously constrain the outputs of other artificial neurons. The weights (i.e., network parameters) can be determined by iteratively training the artificial neural network with a training data set (described later).

[0066] In some examples, as Figure 4 shown, the artificial neural network module 420 may include a backbone neural network 4200 and a second artificial neural network 422.

[0067] In some examples, the backbone neural network 4200 may include a first artificial neural network 421, a third artificial neural network 423, and a feature combination module 424.

[0068] In some examples, the first artificial neural network 421 may receive a tissue image and perform feature extraction on the tissue image to obtain a feature map.

[0069] In some examples, the second artificial neural network 422 may receive the feature map and the recognition result from the third artificial neural network 423 and obtain an attention heat map indicating the lesion area and a complementary attention heat map indicating the non-lesion area. It should be noted that in some other examples, the above-mentioned attention heat map or complementary attention heat map may also be regarded as a kind of feature map.

[0070] In some examples, the feature combination module 424 may receive the feature map, the attention heat map, and the complementary attention heat map and output a feature combination set. In some examples, the feature combination module 424 may also directly output the feature map.

[0071] In some examples, the third artificial neural network 423 may receive the feature map or the feature combination set and output the recognition result of the tissue lesion recognition of the tissue image.

[0072] In some examples, the tissue image (such as the preprocessed tissue image) input to the artificial neural network module 420 may enter the first artificial neural network 421 and finally the recognition result may be output by the third artificial neural network 423.

[0073] Figure 5 It is a block diagram showing a modification example of the artificial neural network module related to the examples of the present disclosure.

[0074] In addition, in some examples, as Figure 5 shown, the artificial neural network module 420 may include a backbone neural network 4200 and a second artificial neural network 422.

[0075] In some examples, as Figure 5 shown, the backbone neural network 4200 may include a first artificial neural network 421 and a third artificial neural network 423.

[0076] In some examples, the third artificial neural network 423 may have a feature combination function. For specific content, refer to the relevant description in the feature combination module 424.

[0077] In some examples, the first artificial neural network 421 may receive a tissue image and perform feature extraction on the tissue image to obtain a feature map.

[0078] In some examples, the second artificial neural network 422 may also obtain a complementary attention heat map indicating non-lesion regions based on the attention heat map.

[0079] In some examples, the third artificial neural network 423 may receive the feature map, the attention heat map, and the complementary attention heat map and output an identification result of tissue lesion identification of the tissue image. In some examples, the attention heat map may be a heat map indicating lesion regions obtained based on an attention mechanism. In some examples, the attention heat map may show the importance degree of each pixel point in the tissue image when forming the feature map.

[0080] In some examples, the complementary attention heat map may be a heat map indicating non-lesion regions obtained based on a complementary attention mechanism.

[0081] In some examples, the complementary attention heat map may be a complementary image of the attention heat map. In some examples, the size and format of the complementary attention heat map may be the same as those of the attention heat map.

[0082] As described above, the artificial neural network module 420 may include the first artificial neural network 421 (see Figure 5 ).

[0083] In some examples, the first artificial neural network 421 may use one or more deep neural networks to automatically identify features in the tissue image.

[0084] In some examples, the first artificial neural network 421 may be used to receive the tissue image preprocessed by the preprocessing module and generate one or more feature maps. In some examples, the first artificial neural network 421 may, for example, combine multiple layers of low-level features (pixel-level features). In this case, an abstract description of the tissue image can be achieved.

[0085] In some examples, the first artificial neural network 421 may include an input layer, an intermediate layer, and an output layer connected in sequence. The input layer may be configured to receive the tissue image preprocessed by the preprocessing module. The intermediate layer is configured to be able to extract the feature map based on the tissue image, and the output layer is configured to be able to output the feature map.

[0086] In some examples, the tissue image input to the artificial neural network module 420 may be converted into a pixel matrix, for example, a three-dimensional pixel matrix. The length and width of the three-dimensional matrix may represent the size of the image, and the depth of the three-dimensional matrix represents the color channels of the image. In some examples, the depth may be 1 (i.e., the tissue image is a grayscale image), and in some examples, the depth may be 3 (i.e., the tissue image is a color image in the RGB color mode).

[0087] In some examples, the first artificial neural network 421 may adopt a convolutional neural network. Since convolutional neural networks have advantages such as local receptive fields and weight sharing, they can greatly reduce the training of parameters, thus improving the processing speed and saving hardware overhead. In addition, convolutional neural networks can more effectively identify tissue images.

[0088] Figure 6 FIG. 4 is a schematic structural diagram of the first artificial neural network 421 involved in the examples of the present disclosure.

[0089] In some examples, the first artificial neural network 421 may include multiple intermediate layers. The intermediate layers may include multiple neurons or nodes. An activation function (such as a ReLU (rectified linear unit) function, a sigmoid function, or a tanh function, etc.) may be applied to the output of each neuron or node in each neuron or node in the intermediate layer. The activation functions applied by different neurons affect the activation functions applied by other neurons.

[0090] In some examples, as Figure 6 shown, the intermediate layer of the first artificial neural network 421 may include multiple convolutional layers and multiple pooling layers. In some examples, the convolutional layers and the pooling layers may be alternately combined. In some examples, the tissue image may sequentially pass through the first convolutional layer C1, the first pooling layer S1, the second convolutional layer C2, the second pooling layer S2, the third convolutional layer C3, and the third pooling layer S3. In this case, convolutional processing and pooling processing can be alternately performed on the tissue image.

[0091] In other examples, the first artificial neural network 421 may not include a pooling layer, thereby avoiding data loss during the pooling process and simplifying the network structure.

[0092] In some examples, the convolutional layer may use a convolution kernel to perform convolution on the tissue image in the convolutional neural network. In this case, features with a higher degree of abstraction can be obtained, making the matrix depth deeper.

[0093] In some examples, the size of the convolution kernel may be 3*3. In other examples, the size of the convolution kernel may be 5*5. In some examples, a 5×5 convolution kernel may be used in the first convolutional layer C1, and 3×3 convolution kernels may be used in other convolutional layers. In this case, the training efficiency can be improved. In some examples, the size of the convolution kernel can be set to any size. In this case, the size of the convolution kernel can be selected according to the size of the image and the computational cost.

[0094] In some examples, the pooling layer can also be referred to as the downsampling layer. In some examples, pooling methods such as max-pooling, mean-pooling, or stochastic-pooling can be used to process the input tissue images. In this case, through the pooling operation, on the one hand, the feature dimension can be reduced and the operation efficiency can be improved. On the other hand, the convolutional neural network can also extract more abstract high-level features to improve the accuracy of tissue lesion recognition.

[0095] In addition, in some examples, in the above-mentioned convolutional neural network, the number of convolutional layers and pooling layers can also be correspondingly increased according to the situation. In this case, the convolutional neural network can also extract more abstract high-level features to further improve the accuracy of tissue lesion recognition.

[0096] In some examples, after the preprocessed tissue image passes through the first artificial neural network 421, a feature map corresponding to the tissue image can be output. In some examples, the feature map can have multiple depths. In some examples, after the preprocessed tissue image passes through the first artificial neural network 421, multiple feature maps can be output. In some examples, multiple feature maps can each correspond to a feature. In some examples, tissue lesion recognition can be performed on the tissue image based on the feature corresponding to the feature map.

[0097] In some examples, before the first artificial neural network 421 outputs the feature map, deconvolution and upsampling processes can be sequentially performed on the feature map. In some examples, the feature map can undergo multiple deconvolution and upsampling processes. For example, the feature map can sequentially pass through the first deconvolution layer, the first upsampling layer, the second deconvolution layer, the second upsampling layer, the third deconvolution layer, and the third upsampling layer. In this case, the size of the feature map can be changed and part of the data information of the tissue image can be retained.

[0098] In some examples, the number of deconvolution layers can be the same as the number of convolutional layers, and the number of pooling layers (downsampling layers) can be the same as the number of upsampling layers. Thus, the size of the feature map can be made the same as that of the tissue image.

[0099] In some examples, before the feature map passes through the deconvolution layer (upsampling), the tissue image processed by the convolutional layer (pooling layer) can be selected for convolution. For example, before the feature map enters the second deconvolution layer (second upsampling layer), the feature map can be convolved with the output image of the second convolutional layer C2 (second pooling layer S2). Before the feature map enters the third deconvolution layer (third upsampling layer), the feature map can be convolved with the output image of the first convolutional layer C1 (first pooling layer S1). In this case, the data information lost during passing through the pooling layer or convolutional layer can be supplemented.

[0100] In some examples, after generating a feature map through the first artificial neural network 421, an attention heat map matching the feature map can be generated through the second artificial neural network 422.

[0101] In this embodiment, the second artificial neural network 422 is an artificial neural network with an attention mechanism. In some examples, the output image of the second artificial neural network 422 may include an attention heat map and a complementary attention heat map.

[0102] In some examples, the second artificial neural network 422 may include an input layer, an intermediate layer, and an output layer connected in sequence. The input layer is configured to receive the feature map and partial weights of the third artificial neural network 423 or the recognition result of tissue lesion recognition. The intermediate layer can be configured to obtain feature weights based on the partial weights of the third artificial neural network 423 or the tissue lesion recognition result. The intermediate layer can be configured to generate an attention heat map and / or a complementary attention heat map based on the feature map and the feature weights. The output layer is configured to output the attention heat map and / or the complementary attention heat map. In some examples, the feature map can be generated by the first artificial neural network 421.

[0103] In some examples, the attention mechanism can selectively filter out a small amount of important information from a large amount of information in the input feature map and focus on this important information.

[0104] In some examples, the attention heat map can be an image representing attention in the form of a heat map. Generally, the pixels at the corresponding positions in the attention heat map that are red or white have a greater impact on the recognition of tissue image lesions. The pixels at the corresponding positions in the attention heat map that are blue or black have a smaller impact on the recognition of tissue image lesions.

[0105] In some examples, the feature maps can be weighted using the feature weights to obtain an attention heat map. In some examples, the feature weights can be obtained through the attention mechanism. In some examples, the attention mechanism can include, but is not limited to, a channel attention mechanism (channel attention module, CAM), a gradient-based channel attention mechanism (Grad-CAM), a gradient-based enhanced channel attention mechanism (Grad-CAM++), or a spatial attention mechanism (spatial attention module, SAM), etc.

[0106] In some examples, when the third artificial neural network 423 has a global pooling layer and a fully connected layer. In some examples, the feature weights can be the weights from the fully connected layer to the output layer of the third artificial neural network 423 in the third artificial neural network 423. For example, in the case where the tissue image is a fundus image, the third artificial neural network 423 can receive the feature map of the fundus image and obtain a first recognition result (described later). If the first recognition result is "macula", then the weights from each neuron or node of the global pooling layer to the recognition result of "macula" in the fully connected layer are extracted as the feature weights.

[0107] In some examples, the feature weights can be calculated based on the tissue lesion recognition result in the third artificial neural network 423. In some examples, the partial derivatives of the first recognition result (e.g., the probability of tissue lesion) of the third artificial neural network 423 with respect to all pixels in a feature map can be calculated, and global pooling processing is performed on the partial derivatives of all pixels in this feature map to obtain the feature weights corresponding to this feature map.

[0108] In some examples, the second artificial neural network 422 can generate an attention heat map that matches the feature map. In some examples, the second artificial neural network 422 can generate a complementary attention heat map. In some examples, the two pixel values corresponding to the pixels at the same position in the attention heat map and the complementary attention heat map are anti-correlated. In some examples, the attention heat map and / or the complementary attention heat map can be normalized. In some examples, the sum or product of the two pixel values corresponding to the pixels at the same position in the attention heat map and the complementary attention heat map is a constant value.

[0109] In some examples, total variation can be used to regularize the attention heat map and / or the complementary attention heat map.

[0110] In some examples, a feature combination module 424 can be connected to the output layers of the first artificial neural network 421 and the second artificial neural network 422.

[0111] In some examples, the feature combination module 424 can have an input layer and an output layer. In some examples, the output layer of the feature combination module 424 can be a feature map or a feature combination set. In some examples, the input layer of the feature combination module 424 can receive a feature map, an attention heat map, or a complementary attention heat map.

[0112] In some examples, the feature combination module 424 can perform feature combination on the feature map output by the first artificial neural network 421 and the attention heat map or the complementary attention heat map output by the second artificial neural network 422 to form a feature combination set.

[0113] In some examples, the set of feature combinations may include at least one of a first set of feature combinations and a second set of feature combinations.

[0114] In some examples, the feature combination module 424 may perform feature combination on the feature map output by the first artificial neural network 421 and the attention heat map output by the second artificial neural network 422 to form a first set of feature combinations.

[0115] In some examples, the feature combination module 424 may perform feature combination on the feature map output by the first artificial neural network 421 and the complementary attention heat map output by the second artificial neural network 422 to form a second set of feature combinations.

[0116] In some examples, the feature combination module 424 may directly output the feature map.

[0117] In some examples, the feature combination module 424 may also calculate the difference between the feature map and the attention heat map to obtain a first set of feature combinations.

[0118] In some examples, the feature combination module 424 may also calculate the difference between the feature map and the complementary attention heat map to obtain a second set of feature combinations

[0119] In some examples, the feature combination module 424 may also calculate the convolution of the feature map and the attention heat map to obtain a first set of feature combinations.

[0120] In some examples, the feature combination module 424 may also calculate the convolution of the feature map and the complementary attention heat map to obtain a second set of feature combinations.

[0121] In some examples, the feature combination module 424 may also calculate the mean of the feature map and the attention heat map to obtain a first set of feature combinations.

[0122] In some examples, the feature combination module 424 may also calculate the mean of the feature map and the complementary attention heat map to obtain a second set of feature combinations.

[0123] In addition, in some other examples, the feature combination module 424 may perform a linear or non - linear transformation on the feature map and the attention heat map to obtain a first set of feature combinations.

[0124] In addition, in some other examples, the feature combination module 424 may perform a linear or non - linear transformation on the feature map and the complementary attention heat map to obtain a second set of feature combinations.

[0125] In some examples, the output layer of the feature combination module 424 can output a feature map, a first feature combination set, and a second feature combination set. In some examples, the feature map, the first feature combination set, and the second feature combination set output by the feature combination module 424 can be input into the third artificial neural network 423, and the third artificial neural network 423 can perform tissue lesion recognition on them.

[0126] In some examples, the feature combination module 424 can be incorporated into the third artificial neural network 423 and serve as a part of the third artificial neural network 423. In this case, the artificial neural network module 420 can include the first artificial neural network 421, the second artificial neural network 422, and the third artificial neural network 423.

[0127] In some examples, when the feature combination module 424 is incorporated into the third artificial neural network 423, the input layer of the third artificial neural network 423 can receive a feature map, an attention heat map, or a complementary attention heat map.

[0128] In some examples, the third artificial neural network 423 can include an input layer, an intermediate layer, and an output layer connected in sequence. In some examples, the output layer is configured to be able to output the recognition result reflecting the tissue image. In this case, the third artificial neural network 423 can be used to output the recognition result reflecting the tissue image. In some examples, the output layer of the third artificial neural network 423 can include a Softmax layer. In some examples, the intermediate layer of the third artificial neural network 423 can be a fully connected layer.

[0129] In some examples, the final classification can be performed by a fully connected layer, and finally the probability that the tissue image belongs to each tissue lesion category can be obtained through the Softmax layer. In this case, the recognition result of the tissue lesion recognition of the tissue image can be obtained based on the probability.

[0130] In some examples, the third artificial neural network 423 can include various linear classifiers, such as a single-layer fully connected layer.

[0131] In some examples, the third artificial neural network 423 can include various non-linear classifiers. For example, Logistic Regression, Random Forest, or Support Vector Machines, etc.

[0132] In some examples, the third artificial neural network 423 may include multiple classifiers. In some examples, the classifier may give an identification result of the tissue lesion identification of the tissue image. For example, in the case where the tissue image is a fundus image, an identification result of the fundus lesion identification of the fundus image may be given. In this case, it is possible to perform fundus lesion identification on the fundus image.

[0133] In some examples, the output of the third neural network 423 may be a value between 0 and 1, and these values can be used to represent the probability that the tissue image belongs to each tissue lesion category.

[0134] In some examples, when the probability that the tissue image belongs to a certain tissue lesion category is the highest, then this category is taken as the identification result of the tissue lesion identification of the tissue image. For example, among the probabilities that the tissue image belongs to each tissue lesion category, if the probability of the category being no lesion is the highest, then the identification result of the tissue lesion identification of this tissue image may be no lesion. Another example is that in the process of performing fundus lesion identification on a fundus image, if the predicted probabilities of the macula and no lesion output by the third artificial neural network 423 are 0.8 and 0.2 respectively, then it can be considered that there is a macula lesion in this fundus image.

[0135] In some examples, the third artificial neural network 423 may output an identification result that matches the tissue image. In some examples, the identification result may include a first identification result when the attention mechanism is not used, a second identification result when the attention mechanism is used, and a third identification result when the attention mechanism and complementary attention are used.

[0136] In some examples, the third artificial neural network 423 may perform tissue lesion identification on the feature map output by the feature combination module 424 and obtain a first identification result.

[0137] In some examples, the third artificial neural network 423 may perform tissue lesion identification on the first feature combination set output by the feature combination module 424 and obtain a second identification result.

[0138] In some examples, the third artificial neural network 423 may perform tissue lesion identification on the second feature combination set output by the feature combination module 424 and obtain a third identification result.

[0139] In some examples, the recognition results may include two results: with lesions and without lesions. In some examples, the recognition results may also include without lesions or specific types of lesions. For example, in the case where the tissue image is a fundus image, the recognition results may include, but are not limited to, one of without lesions, hypertensive retinopathy, or diabetic retinopathy. In this case, the recognition results of the fundus lesion recognition of the fundus image can be obtained. In some examples, the recognition results of a tissue image can be multiple. For example, the recognition results can be two results: hypertensive retinopathy and diabetic retinopathy.

[0140] In some examples, the recognition system 40 may further include a judgment module.

[0141] In some examples, the judgment module can receive the output of the artificial neural network module 420. In this case, the judgment module can synthesize the output results of the artificial neural network module 420 and output the final recognition results, so as to be able to generate a summary report.

[0142] In some examples, the first recognition result can be used as the final recognition result of the tissue image. In this case, when using the artificial neural network module 420 to perform tissue lesion recognition on the tissue image, the tissue lesion recognition of the tissue image can be performed through the backbone neural network 4200 including the first artificial neural network 421 and the third artificial neural network 423, thereby accelerating the recognition speed.

[0143] In some examples, the second recognition result can be used as the final recognition result of the tissue image.

[0144] As described above, the third recognition result can be obtained based on the complementary attention mechanism. In some examples, the final recognition result of the tissue image can be obtained based on the first recognition result, the second recognition result, and the third recognition result. For example, in some examples, the final recognition result may include the second recognition result and the third recognition result. In some examples, the final recognition result may include the first recognition result and the third recognition result.

[0145] In some examples, the summary report generated by the judgment module may include at least one of the first recognition result, the second recognition result, the third recognition result, and the final recognition result. In some examples, the judgment module can perform color coding on the tissue image based on the attention heat map to generate a lesion indication map to indicate the lesion area. The summary report generated by the judgment module may include the lesion indication map.

[0146] In some examples, the summary report generated by the judgment module may include the location of the corresponding lesion, and use a bounding box to mark the location.

[0147] In some examples, the summary report generated by the judgment module can display the lesion area of the tissue image in the form of a heat map. Specifically, in the heat map, areas with a high likelihood of lesions can be colored red or white, while areas with a low likelihood of lesions can be colored blue or black. In this case, the lesion area can be indicated in an intuitive manner.

[0148] In some examples, the judgment module can also be used to frame the lesion area. In some examples, the lesion area can be framed by a fixed shape (such as regular shapes like triangles, circles, quadrilaterals, etc.). In some examples, the lesion area can also be outlined. In this case, the lesion area can be visually displayed.

[0149] In some examples, the judgment module can also be used to outline the lesion area. For example, the values corresponding to each pixel point in the attention heat map can be analyzed, and the pixel points with values greater than the first preset value can be classified as the lesion area, while the pixel points with values less than the first preset value can be classified as the non-lesion area.

[0150] The recognition method for tissue lesion recognition based on artificial neural network involved in the present disclosure is implemented through the recognition system 40.

[0151] In some examples, the recognition method includes: obtaining a tissue image and using the artificial neural network module 420 to obtain the recognition result of tissue lesion recognition. In some examples, the tissue image can be a tissue image collected by a collection device. In some examples, the artificial neural network module 420 is trained by the training system 430. In this case, the recognition result of tissue lesion recognition can be obtained using the artificial neural network module 420, and the artificial neural network module 420 can be optimized using the total loss function, thereby improving the accuracy of tissue lesion recognition.

[0152] Hereinafter, the training method (which can sometimes be simply referred to as the training method) and the training system for tissue lesion recognition based on artificial neural network involved in the present embodiment will be specifically described with reference to the accompanying drawings.

[0153] In some examples, the training method can be implemented using the training system 430 for tissue lesion recognition based on artificial neural network. In this case, the artificial neural network module 420 can be trained using the training system 430.

[0154] Figure 7 FIG. is a block diagram showing the structure of the training system 430 for tissue lesion recognition based on artificial neural network according to the example of the present disclosure.

[0155] In some examples, as Figure 7As shown, the training system 430 may include a storage module 431, a processing module 432, and an optimization module 433. In some examples, the storage module 431 may be configured to store a training data set. In some examples, the processing module 432 may perform operations such as feature extraction, generating an attention heat map and a complementary attention heat map, and tissue lesion recognition by using the artificial neural network module 420. In some examples, the optimization module 433 may obtain a total loss function based on the recognition results of tissue lesion recognition (including the first recognition result, the second recognition result, and the third recognition result) to optimize the artificial neural network module 420. In this case, the recognition results of tissue lesion recognition can be obtained by using the attention mechanism and the complementary attention mechanism, and the total loss function can be obtained based on the recognition results of tissue lesion recognition, so that the artificial neural network module 420 can be optimized by using the total loss function, thereby improving the accuracy of tissue lesion recognition of the artificial neural network module 420.

[0156] In some examples, the training method of the artificial neural network module 420 may be weakly supervised. In this case, more informative recognition results can be obtained by using the artificial neural network module 420 with less informative annotation results. In some examples, when the annotation result is a text annotation, the recognition result may include the location and size of the lesion area. In some examples, the training method of the artificial neural network module 420 may also be unsupervised, semi-supervised, reinforcement learning, etc.

[0157] In some examples, the artificial neural network module 420 may be trained by using a first loss function, a second loss function, and a third loss function. It should be noted that since the training models and loss functions involved are generally complex, there is generally no analytical solution for this model. In some examples, the value of the loss function can be minimized as much as possible by iteratively updating the model parameters a finite number of times through an optimization algorithm (such as batch gradient descent (BGD), stochastic gradient descent (SGD), etc.), that is, finding the analytical solution of this model. In some examples, the artificial neural network module 420 can be trained by using the backpropagation algorithm. In this case, the network parameters with the minimum error can be achieved, thereby improving the recognition accuracy.

[0158] Figure 8 is a flowchart showing a training method for tissue lesion recognition based on an artificial neural network according to an example of the present disclosure.

[0159] In some examples, such as Figure 8As shown, the training method may include preparing a training data set (step S100); inputting the training data set into the artificial neural network module 420, and obtaining a first recognition result, a second recognition result, and a third recognition result that match each inspection image (step S200); calculating a total loss function based on the first recognition result, the second recognition result, and the third recognition result (step S300); and optimizing the artificial neural network module 420 using the total loss function (step S400). In this case, the first recognition result, the second recognition result, and the third recognition result can be obtained, and the total loss function can be obtained based on the first recognition result, the second recognition result, and the third recognition result, so that the artificial neural network module 420 can be optimized using the total loss function, thereby improving the accuracy of tissue lesion recognition of the artificial neural network module 420.

[0160] In step S100, a training data set can be prepared. In some examples, the training data set may include multiple inspection images and labeled results of having lesions or not having lesions associated with the inspection images.

[0161] In some examples, the training data set may include multiple inspection images and labeled images associated with the inspection images.

[0162] In some examples, the inspection images can be 50,000 - 200,000 tissue images from partner hospitals with patient information removed. In some examples, the inspection images can be from CT scans, PET-CT scans, SPECT scans, MRIs, ultrasounds, X-rays, mammograms, angiograms, fluorograms, tissue images taken by capsule endoscopes, or combinations thereof. In some examples, the inspection images can be fundus images. In some examples, the inspection images can be composed of a lesion area and a non-lesion area. In some examples, the inspection images can be used for the training of the artificial neural network module 420.

[0163] In some examples, the inspection images can be obtained by the acquisition module 410.

[0164] In some examples, the labeled images may include labeled results of having lesions or not having lesions. In some examples, the labeled results can be used as the ground truth to measure the magnitude of the loss function.

[0165] In some examples, the labeled results can be image labels or text labels. In some examples, the image labels can be labeled boxes manually labeled for bounding the lesion area.

[0166] In some examples, the labeled boxes can be of a fixed shape, such as regular shapes like triangles, circles, or quadrilaterals. In some examples, the labeled boxes can also be irregular shapes outlined based on the lesion area.

[0167] In some examples, the text annotation can be the determination result of checking whether there is a lesion in the image. For example, "there is a lesion" or "there is no lesion". In some examples, the text annotation can also be the type of the lesion. For example, when the image being examined is a fundus image, the text annotation can be "macular lesion", "hypertensive retinopathy", or "diabetic retinopathy", etc.

[0168] In some examples, the training data set can be stored in the storage module 431. In some examples, the storage module 431 can be configured to store the training data set.

[0169] In some examples, the training data set can include 30%-60% of the examination images with annotation results of no-lesion results. In some examples, the training data set can include 10%, 20%, 30%, 40%, 50%, or 60% of the examination images with annotation results of no-lesion results.

[0170] In some examples, the storage module 431 can be used to store the training data set. In some examples, the storage module 431 can include the memory 20.

[0171] In some examples, the storage module 431 can be configured to store the examination images and the annotation images associated with the examination images.

[0172] In some examples, the artificial neural network module 420 can receive the training data set stored in the storage module 431.

[0173] In some examples, the training data set can be preprocessed.

[0174] In step S200, the training data set can be input into the artificial neural network module 420, and the first recognition result, the second recognition result, and the third recognition result matching each examination image can be obtained. In some examples, the training data set can be input into the artificial neural network module 420 to obtain the feature map, the attention heat map, and the complementary attention heat map. In some examples, feature extraction can be performed on the examination image to obtain the feature map. In some examples, the feature map can be processed based on the attention mechanism to obtain the attention heat map. In some examples, the attention heat map can be processed based on the complementary attention mechanism to obtain the complementary attention heat map.

[0175] In some examples, step S200 can be implemented by using the processing module 432. In some examples, the processing module 432 can include at least one processor 10.

[0176] In some examples, as described above, the artificial neural network module 420 can include the first artificial neural network 421, the second artificial neural network 422, and the third artificial neural network 423.

[0177] In some examples, the processing module 432 may be configured to extract features from the inspection image using the first artificial neural network 421 to obtain a feature map. In some examples, the processing module 432 may be configured to obtain an attention heat map indicating a lesion area and a complementary attention heat map indicating a non-lesion area using the second artificial neural network 422.

[0178] In some examples, the processing module 432 may be configured to obtain an identification result including tissue lesion identification using the third artificial neural network 423. As described above, the third artificial neural network 423 may include an output layer. In some examples, the output layer may be configured to output an identification result reflecting the inspection image. In this case, the third artificial neural network 423 can output an identification result reflecting the inspection image.

[0179] In some examples, the processing module 432 may use the third artificial neural network 423 to identify the inspection image based on the feature map to obtain a first identification result.

[0180] In some examples, the processing module 432 may use the third artificial neural network 423 to identify the inspection image based on the feature map and the attention heat map to obtain a second identification result.

[0181] In some examples, the processing module 432 may use the third artificial neural network 423 to identify the inspection image based on the feature map and the complementary attention heat map to obtain a third identification result.

[0182] In some examples, the tissue lesion may be an eye fundus lesion. In this case, the artificial neural network module 420 can be used for eye fundus lesion identification of eye fundus images.

[0183] In step S300, the total loss function may be calculated based on the first identification result, the second identification result, and the third identification result.

[0184] In some examples, step S300 may be implemented using the optimization module 433.

[0185] In some examples, the optimization module 433 may obtain the total loss function of the artificial neural network module 420 based on the first loss function, the second loss function, and the third loss function. In this case, the artificial neural network module 420 can be optimized using the total loss function.

[0186] In some examples, the optimization module 433 can combine the first recognition result with the annotated image to obtain a first loss function when the attention mechanism is not used. In some examples, the first loss function can be used to evaluate the degree of inconsistency between the recognition result and the annotation result of the inspection image when the attention mechanism is not used. In this case, the accuracy of tissue lesion recognition of the artificial neural network module 420 when the attention mechanism is not used can be improved.

[0187] In some examples, the optimization module 433 can combine the second recognition result with the annotated image to obtain a second loss function when the attention mechanism is used. In some examples, the second loss function can be used to evaluate the degree of inconsistency between the recognition result and the annotation result of the inspection image when the attention mechanism is used. In this case, the accuracy of tissue lesion recognition of the artificial neural network module 420 when the attention mechanism is used can be improved.

[0188] In some examples, the optimization module 433 can combine the third recognition result with the annotated image with a lesion-free annotation result to obtain a third loss function when the complementary attention mechanism is used. In some examples, the third loss function can be used to evaluate the degree of inconsistency between the recognition result of the inspection image when the complementary attention mechanism is used and the lesion-free recognition. In this case, the accuracy of tissue lesion recognition of the artificial neural network module 420 when the complementary attention mechanism is used can be improved.

[0189] In some examples, the first loss function, the second loss function, and the third loss function can be obtained through an error loss function. In some examples, the error loss function can be a correlation function, an L1 loss function, an L2 loss function, or a Huber loss function, etc., which are functions used to evaluate the correlation between the true value (i.e., the annotation result) and the predicted value (i.e., the recognition result).

[0190] In some examples, the total loss function can include a first loss term, a second loss term, and a third loss term.

[0191] In some examples, the first loss term can be positively correlated with the first loss function. In this case, the first loss term can be used to evaluate the degree of inconsistency between the recognition result and the annotation result of the inspection image when the attention mechanism is not used, thereby improving the accuracy of tissue lesion recognition.

[0192] In some examples, the second loss term can be positively correlated with the difference between the second loss function and the first loss function. In some examples, when the second loss function is less than the first loss function, the second loss term can be a constant value. In this case, the second loss term can be used to evaluate the degree of inconsistency between the recognition result of the inspection image when the attention mechanism is used and the recognition result when the attention mechanism is not used.

[0193] In some examples, the second loss term can be positively correlated with the difference between the second loss function and the first loss function. Specifically, when the difference between the second loss function and the first loss function is greater than zero, the difference between the second loss function and the first loss function can be used as the second loss term. When the difference between the second loss function and the first loss function is less than zero, the second loss term can be set to zero. In this case, the degree of inconsistency between the first recognition result and the second recognition result can be evaluated using the second loss term, so that the second recognition result can be closer to the annotation result relative to the first recognition result.

[0194] In some examples, the third loss term can be positively correlated with the third loss function. In this case, the degree of inconsistency between the third recognition result of the inspection image when using the complementary attention mechanism and the annotation result of no lesion can be evaluated using the third loss term, so that the occurrence of misjudgment or missed judgment can be reduced.

[0195] In some examples, the total loss function can further include a fourth loss term. In some examples, the fourth loss term can be a regularization term. In some examples, the fourth loss term can be a regularization term for the attention heat map. In some examples, the regularization term can be obtained based on the total variation. In this case, overfitting of the artificial neural network module 420 can be suppressed.

[0196] In some examples, the total loss function can include loss term weight coefficients that match each loss term. In some examples, the total loss function can further include a first loss term weight coefficient that matches the first loss term, a second loss term weight coefficient that matches the second loss term, a third loss term weight coefficient that matches the third loss term, and a fourth loss term weight coefficient that matches the fourth loss term, etc.

[0197] In some examples, the first loss term can be multiplied by the first loss term weight coefficient, the second loss term can be multiplied by the second loss term weight coefficient, the third loss term can be multiplied by the third loss term weight coefficient, the fourth loss term can be multiplied by the fourth loss term weight coefficient, and the fifth loss term can be multiplied by the fifth loss term weight coefficient. Thus, the influence degree of each loss term on the total loss function can be adjusted through the loss term weight coefficients.

[0198] In some examples, the loss term weight coefficient can be set to 0. In some examples, the loss term weight coefficient can be set to a positive number. In this case, since each loss term is non - negative, the value of the total loss function can be made not less than zero.

[0199] In some examples, the functional formula of the total loss function can be:

[0200]

[0201] Wherein, L is the total loss function, λ1 is the weight coefficient of the first loss term, λ2 is the weight coefficient of the second loss term, λ3 is the weight coefficient of the third loss term, λ4 is the weight coefficient of the fourth loss term, f is the error loss function, X is the inspection image, F(X) is the feature map generated after the inspection image X passes through the first artificial neural network 421, l(X) is the annotation result of the inspection image X, max is the maximum value function, C is the classifier function that outputs the recognition result based on the input feature map or feature combination set, margin is the preset parameter, l0 is the annotation result of no lesion, M(X) is the attention heat map matching the inspection image X, is the complementary attention heat map matching the inspection image X. The "·" in the functional formula of the total loss function is the dot product operation of matrices, and Regularize(M) is the regularization term for the attention heat map M. In some examples, the classifier function can be implemented by the third artificial neural network 423.

[0202] In some examples, the optimization module 433 can obtain the total loss function including the first loss term based on the first loss function, the second loss term based on the difference between the second loss function and the first loss function, and the third loss term based on the third loss function by using the first loss function, the second loss function, and the third loss function, and optimize the artificial neural network module 420 by using the total loss function.

[0203] In some examples, the total loss function can further include a fifth loss term. In some examples, the fifth loss term can be the total area term of the attention heat map. Specifically, the total area term of the attention heat map can be the area of the region determined to be a lesion within the attention heat map. In some examples, the total area term of the attention heat map M(X) can be represented by the formula SUM(M(X)). In some examples, the artificial neural network module 420 can be trained by using the fourth loss term to make the lesion region within the attention heat map smaller. In this case, the fifth loss term can be used to evaluate the area of the lesion region within the attention heat map and control the number of pixels that have a greater impact on the recognition result in the attention heat map, so that the attention of the network is limited to the pixels that have a greater impact on the recognition result. Thereby, the accuracy of lesion region recognition can be increased.

[0204] In some examples, the total loss function can further include a sixth loss term. In some examples, the sixth loss term can be used to evaluate the degree of inconsistency between the boxed region of the lesion region in the recognition result and the annotation box of the lesion region manually annotated in the annotation image.

[0205] In step S400, the artificial neural network module 420 can be optimized by using the total loss function.

[0206] In some examples, step S400 may be implemented by using the optimization module 433.

[0207] In some examples, the optimization module 433 may optimize the artificial neural network module 420 by using a total loss function to minimize the total loss function. In this case, the total loss function can be minimized to improve the accuracy of tissue lesion recognition of the artificial neural network module 420.

[0208] In some examples, the optimization module 433 may obtain a total loss function based on the first loss term, the second loss term, the third loss term, and the total area term of the attention heat map, and optimize the artificial neural network module 420 by using the total loss function to obtain an artificial neural network module 420 that can be used for tissue lesion recognition. Thereby, the accuracy of tissue lesion recognition of the artificial neural network module 420 can be further improved.

[0209] In some examples, the optimization module 433 may adjust the total loss function by changing the weights of the first loss term, the second loss term, the third loss term, and the fourth loss term.

[0210] In some examples, the optimization module 433 may use the first loss term and the sixth loss term as the total loss function (i.e., setting the loss term weight coefficients of other loss terms to zero) to optimize the artificial neural network module 420. Thereby, the accuracy of the attention heat map and the complementary attention heat map generated by the second artificial neural network 422 can be improved.

[0211] In some examples, during the optimization process, the loss term weight coefficients in the total loss function may be modified.

[0212] In some examples, the optimization module 433 may use an optimization algorithm to perform multiple iterations on the parameters in the total loss function to reduce the value of the total loss function. For example, in this embodiment, the mini-batch stochastic gradient descent algorithm may be used. By randomly selecting a set of input function parameters, and then performing multiple iterations on the parameters to reduce the value of the loss function.

[0213] In some examples, when the total loss function is less than a second preset value or the number of iterations exceeds a third preset value, the training is paused.

[0214] In some examples, the optimization module 433 may first pre-train the artificial neural network module 420 without using the attention mechanism, and then train the artificial neural network module 420 by using the attention mechanism. In this case, the training speed can be accelerated.

[0215] In some examples, the optimization module 433 can train the first artificial neural network 421, the second artificial neural network 422, and the third artificial neural network 423 simultaneously. In this case, the training speed can be accelerated.

[0216] In some examples, after the training is completed, the optimization module 433 can adopt, for example, 0 - 20,000 tissue images (such as fundus images) as test tissue images to form a test set.

[0217] In some examples, the test tissue images can be used for the post - training test of the artificial neural network module 420.

[0218] FIG. 9(a) is a schematic diagram showing an example of a lesion area of a fundus image obtained by training without using an attention mechanism according to an example of the present disclosure. FIG. 9(b) is a schematic diagram showing an example of a lesion area of a fundus image obtained by training using a complementary attention mechanism according to an example of the present disclosure.

[0219] In some examples, the accuracy of tissue lesion recognition of the fundus image obtained by training using a complementary attention mechanism is higher. As an example of not using an attention mechanism, FIG. 9(a) shows the lesion area A of the fundus image obtained by training without using an attention mechanism. As an example of a complementary attention mechanism, FIG. 9(b) shows the lesion area B of the fundus image obtained by training using a complementary attention mechanism.

[0220] Although the present disclosure has been specifically described above in conjunction with the drawings and embodiments, it can be understood that the above description does not limit the present disclosure in any form. Those skilled in the art can make deformations and changes to the present disclosure according to needs without departing from the essence and scope of the present disclosure, and these deformations and changes all fall within the protection scope of the present disclosure.

Claims

1. A training method based on complementary attention, characterized in that Including: Preparing a training data set, where the training data set includes multiple inspection images composed of lesion regions and non-lesion regions, and annotation images associated with the inspection images; Inputting the training data set into an artificial neural network module to extract features from the inspection images to obtain feature maps; Processing the feature maps based on an attention mechanism to obtain an attention heat map indicating the lesion regions, where the attention heat map is a heat map indicating the lesion regions obtained based on the attention mechanism and is configured to display the importance of each pixel point in the inspection images when forming the feature maps; Processing the attention heat map based on a complementary attention mechanism to obtain a complementary attention heat map indicating the non-lesion regions, where the complementary attention heat map is a heat map indicating the non-lesion regions obtained based on the complementary attention mechanism; Obtaining a first recognition result based on the feature maps, obtaining a second recognition result based on the feature maps and the attention heat map, and obtaining a third recognition result based on the feature maps and the complementary attention heat map; Obtaining a total loss function based on the first recognition result, the second recognition result, the third recognition result, and the annotation images; Combining the first recognition result with the annotation images to obtain a first loss function when the attention mechanism is not used, combining the second recognition result with the annotation images to obtain a second loss function when the attention mechanism is used, and combining the third recognition result with the annotation images with a non-lesion annotation result to obtain a third loss function when the complementary attention mechanism is used; Using the first loss function, the second loss function, and the third loss function to obtain a total loss function including a first loss term based on the first loss function, a second loss term based on the difference between the second loss function and the first loss function, and a third loss term based on the third loss function; Optimizing the artificial neural network module using the total loss function.

2. The training method according to claim 1, wherein: The annotation images include an annotation result of having a lesion or an annotation result of having no lesion.

3. The training method according to claim 2, wherein: The annotation result is an image annotation or a text annotation, and the image annotation is an annotation box manually annotated and used to frame the lesion regions.

4. The training method according to claim 1, wherein: The artificial neural network module includes a first artificial neural network, a second artificial neural network, and a third artificial neural network; Using the first artificial neural network to extract features from the inspection images to obtain the feature maps; Using the second artificial neural network to obtain the attention heat map indicating the lesion regions and the complementary attention heat map indicating the non-lesion regions; Using the third artificial neural network to obtain a first recognition result based on the feature maps, obtain a second recognition result based on the feature maps and the attention heat map, and obtain a third recognition result based on the feature maps and the complementary attention heat map.

5. The training method according to claim 4, wherein in the attention heat map and the complementary attention heat map, the sum of the pixel values of the pixels at the same position is a constant value.

6. The training method according to claim 4, wherein after obtaining the attention heat map and the complementary attention heat map, perform normalization processing on the attention heat map and / or the complementary attention heat map.

7. The training method according to claim 1, wherein the first loss function, the second loss function, and the third loss function are obtained through an error loss function, and the error loss function is at least one of a correlation function, an L1 loss function, an L2 loss function, or a Huber loss function.

8. The training method according to claim 1, wherein the first loss term is positively correlated with the first loss function.

9. The training method according to claim 1, wherein when the second loss function is less than the first loss function, the second loss term is zero.

Citation Information

Patent Citations

  • Expression recognition method, device and system

    CN109815924A

  • Fundus image recognition method and device, computer equipment and storage medium

    CN110348543A