Image recognition method and system, diagnosis method and system and storage medium

By adopting image recognition methods in gastrointestinal endoscopy, using multi-scale convolutional kernels and global information combined with attention mechanisms, the misdiagnosis and misdiagnosis problems caused by traditional endoscopy relying on artificial judgments is solved, and a higher recognition accuracy and accurate detection of targets of different sizes are achieved.

CN120107558APending Publication Date: 2025-06-06SHENZHEN EACHON BIO-TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510252382.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Traditional gastrointestinal endoscopy relies on the doctor's experience and observation ability, poses a risk of misdiagnosis and missed diagnosis, and is time-consuming and labor-intensive, which may lead to a decrease in diagnostic accuracy due to doctor's fatigue or insufficient experience.

Method used

By using the image recognition method, by obtaining the input features of the target image, using multiple convolution kernels of different sizes for convolution operations, obtaining feature representations and splicing of multiple scales, combining global information and attention mechanisms, a new tensor is formed to be used for inputs of neural network models to improve the recognition accuracy.

Benefits of technology

Through SK convolution, the image input by the model is modified and the attention mechanism is introduced, so that the neural network model can adaptively adjust the receptive field of view, improve the recognition accuracy, ensure that different proportions of targets can be accurately detected, and improve the recognition accuracy of targets of different sizes and the recognition accuracy of disease sites.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107558A_ABST
    Figure CN120107558A_ABST
Patent Text Reader

Abstract

The invention relates to an image recognition method and system, a diagnosis method and system and a storage medium. The method comprises the steps of obtaining and sending real-time image data of a focus part; performing noise reduction on the obtained real-time image data, and taking the real-time image data as a target image of the SK convolution image recognition method; the input features of the image are modified through an SK convolution image recognition method, an attention mechanism is introduced into the neural network model, then up-sampling and feature splicing are carried out on input values to lead out recognition results of different scales, and abnormal positions are marked; and generating a diagnosis report according to the abnormal condition and / or the abnormal position. An image input by the model is modified through SK convolution, and an attention mechanism is introduced into the neural network model, so that the neural network model can adaptively adjust the size of a feeling visual field according to a scene, the recognition accuracy is improved, targets of different proportions in image recognition can be accurately detected, the recognition precision of the targets of different sizes is improved, and the recognition efficiency is improved. And the recognition accuracy of the model on the disease of the focus part is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image recognition methods, and in particular to an image recognition method, a diagnosis method, a diagnosis system and a computer-readable storage medium. Background Art

[0002] With the continuous development of global medical technology, endoscopic technology plays an increasingly important role in the diagnosis and treatment of gastrointestinal diseases. Traditional gastrointestinal endoscopy mainly relies on the doctor's experience and observation ability. By inserting an endoscope to visualize the digestive tract, various lesions including gastritis, gastric ulcers, intestinal polyps, tumors, etc. are detected and diagnosed. However, this method relies on the doctor's subjective judgment to a certain extent, and there is a risk of misdiagnosis and missed diagnosis. This mode that relies on human operation and judgment is not only time-consuming and laborious, but may also lead to a decrease in diagnostic accuracy due to the doctor's fatigue or lack of experience. At the same time, since the image quality of endoscopic examination may be affected by factors such as operating techniques or the patient's physical condition, it further increases the difficulty of accurate diagnosis. Summary of the invention

[0003] The technical problem to be solved by the present invention is that traditional gastrointestinal endoscopy mainly relies on the doctor's experience and observation ability. By inserting an endoscope to perform a visual inspection of the digestive tract, various lesions including gastritis, gastric ulcer, intestinal polyps, tumors, etc. are discovered and diagnosed. This method relies on the doctor's subjective judgment to a certain extent, and there is a risk of misdiagnosis and missed diagnosis. This mode that relies on human operation and judgment is not only time-consuming and laborious, but may also reduce the diagnostic accuracy due to the doctor's fatigue or lack of experience. In view of the above-mentioned defects of the prior art, an image recognition method, a diagnostic method, a system and a storage medium are provided.

[0004] In order to solve the above technical problems, the technical solution adopted by the present invention is:

[0005] Construct an image recognition method, including:

[0006] Obtain at least one input feature of the target image,

[0007] Use multiple convolution kernels of different sizes to perform convolution operations on the input features to obtain feature representations of multiple scales, and then concatenate the feature representations;

[0008] Add corresponding elements in tensors of the same shape to get global information;

[0009] After reducing the dimension of the global information and then increasing the dimension, the channel descriptors corresponding to K scales are obtained to form a new tensor;

[0010] The weight of each scale is summed with the corresponding structure weight after the previous convolution to obtain different branch weight combinations to form the modified input features;

[0011] The modified input features are used as the input values ​​of the neural network model and up-sampled and feature concatenated to induce recognition results of different scales.

[0012] Preferably, the step of adding corresponding elements in tensors of the same shape to obtain global information includes:

[0013] First, multiple branch elements are summed up, and then global average pooling is performed to compress the summed elements into feature vectors with the same number of channels and capture global information.

[0014] Preferably, after said capturing the global information to obtain the global information;

[0015] The global information is reduced and then increased in dimension to obtain channel descriptors corresponding to K scales, and the feature vector after the increase in dimension is reshaped to the same size as the input, and the reshaped feature vector is stacked according to the K dimensions to form a new tensor.

[0016] Preferably, the step of summing the weight of each scale with the corresponding structure after the previous convolution to obtain different branch weight combinations to form the modified input features further includes:

[0017] The weights corresponding to each scale are normalized through the softmax function, so that the size of the feature map is changed after the image data is downsampled.

[0018] Preferably, the step of using the modified input feature as an input value of the neural network model includes:

[0019] Introducing the attention mechanism into the neural network model enables the neural network model to adaptively adjust the field of view size.

[0020] Construct a diagnostic method, including:

[0021] Obtaining real-time image data of the lesion site and sending it;

[0022] The obtained real-time image data is de-noised and used as the target image of the SK convolution image recognition method;

[0023] Recognize the target image by one of the above-mentioned image recognition methods, determine the abnormality of the image, and mark the abnormal position;

[0024] Generate a diagnostic report based on the anomaly condition and / or anomaly location.

[0025] Preferably, in the method of using the obtained real-time image data after denoising as the target image of the image recognition method of SK convolution, the method includes:

[0026] The real-time image data is cropped after denoising, and the cropped image data is used as the target image of the SK convolution image recognition method.

[0027] Preferably, the step of cropping the real-time image data after reducing noise further includes:

[0028] The real-time image is cropped and normalized so that the pixel value is normalized from [0,255] to [0,1].

[0029] Construct a diagnostic system, including:

[0030] An image acquisition module, for acquiring image data of the lesion site;

[0031] The image preprocessing module reduces noise and crops the collected image data to make it conform to the neural network input format;

[0032] The image recognition module uses the modified neural network model or the trained neural network model to recognize the image, determine the abnormal situation and mark the abnormal location;

[0033] The image reporting module generates a diagnosis report based on the analysis results of the image recognition module.

[0034] A storage medium is constructed, on which program instructions are stored, characterized in that: the program instructions are used to execute an image recognition method as described above when running.

[0035] The beneficial effects of the present invention are as follows: by modifying the image input to the model through SK convolution and introducing the attention mechanism into the neural network model, the neural network model can adaptively adjust the size of the perception field of view according to the scene, thereby improving the recognition accuracy, so that targets of different proportions in image recognition can be accurately detected, thereby improving the recognition accuracy of targets of different sizes, and further improving the model's recognition accuracy for diseases at the lesion site. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the present invention will be further described below in conjunction with the accompanying drawings and embodiments. The drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative work:

[0037] Figure 1 A schematic diagram of a method flow of a diagnostic method according to a preferred embodiment of the present invention;

[0038] Figure 2 A flowchart of the image input and the convolution layer of the SK convolution modification model in a preferred embodiment of the present invention;

[0039] Figure 3 A schematic diagram of the principle of a diagnostic system according to a preferred embodiment of the present invention;

[0040] Figure 4 It is a structural diagram of the improved yolo10n model of a preferred embodiment of the present invention;

[0041] Figure 5 SKNet structure schematic diagram of a preferred embodiment of the present invention. DETAILED DESCRIPTION

[0042] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the following will be described clearly and completely in combination with the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are partial embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work belong to the protection scope of the present invention.

[0043] A diagnostic method according to a preferred embodiment of the present invention; Figure 1 As shown, it is a flowchart of the endoscope operation method provided by an embodiment of the present invention. The method can be executed by a device, and the device can be implemented by software and / or hardware.

[0044] Specifically, in this embodiment, the diagnostic method includes:

[0045] S20: Acquire and send real-time image data of the lesion area;

[0046] In a preferred embodiment of the present invention, real-time image data of the lesion site is first obtained. In the present invention, obtaining real-time image data of the gastrointestinal lesion site is used as an illustration. Real-time image data of the patient's gastrointestinal site is collected by a gastrointestinal endoscope handle and a gastroenteroscope and the real-time image data is sent to an image processing module. The gastrointestinal endoscope handle operates the movement of the gastroenteroscope and obtains real-time image data of the gastrointestinal site. The gastroenteroscope then performs photoelectric conversion on the real-time image data, and the converted data enters the endoscope handle for further decoding to form original image format data. The size of the original image format data collected is 1920X1080 resolution, and the original image format data is sent to the image processing module through the gastrointestinal endoscope handle.

[0047] S21: receiving real-time image data and performing noise reduction and cropping to obtain image data that conforms to the neural network input format;

[0048] In a preferred embodiment of the present invention, after receiving the real-time original image format data sent by the endoscope handle, the image format data is first subjected to denoising to eliminate image noise. Since the original model input image is a square, and the real-time collected image format data is a rectangle, it is not suitable to be directly placed in the subsequent neural network, so the image needs to be properly cropped to conform to the neural network input format. And since the cropped image size is 1600X1024, some corners will be blocked during actual display, so the cropped image also needs to be normalized, and the pixels of the cropped image data are normalized from [0,255] to [0,1]. The specific normalization formula is:

[0049]

[0050] The real-time image data of gastrointestinal images are processed in the above manner so that the pixels of the processed image data meet the requirements of the subsequent neural network input format.

[0051] S22: Modify the input image through SK convolution on the cropped image data and introduce the attention mechanism into the neural network model;

[0052] In a preferred embodiment of the present invention, a trained neural network model is used to identify the processed graphic data to determine whether there are abnormalities and mark the abnormal locations. The yolo10n neural network model can be used to identify lesions, wherein the original paper of yolo10n was published on arxiv, and the paper address is: https: / / arxiv.org / pdf / 2405.14458, which will not be described in detail here. However, the original model training of the yolo10n neural network model uses image data of size 640X640 as the input training image. The image data of this size is small and often cannot meet the needs of actual production. Therefore, the present invention uses SK convolution to modify the input image and the convolution layer of the neural network model, and replaces PSA with the SKAttention mechanism, that is, using Figure 5 The SKNet structure shown replaces the original PSA structure, so that the attention mechanism is introduced into the neural network model, so that objects of different scales in the endoscopic image can be accurately detected, so as to identify different lesion sizes and regions of different sizes, and enhance the recognition rate of objects of different scales. From the neural network model structure of yolo10n, it can be seen that it performs a maximum of 6 downsamplings, and each time the downsampled feature image becomes half of the original, that is, the width and height of the changed image data need to be The width and height of the image data must be integers, and the initial size data of the real-time image obtained is 1920X1080. At the same time, when the image is actually displayed, there will be a small part of the two sides blocked, so the final cropped image size is determined to be 1600X1024.

[0053] S23: The lesion location and size of the image are identified through the modified neural network model, and the lesion is marked and judged whether there is an abnormality;

[0054] In a preferred embodiment of the present invention, the input image is modified by SK convolution, and a part of the original yolo10n model is replaced to introduce an attention mechanism into the neural network model to enhance the recognition rate of the neural network model for targets of different scales. Since the PSA in the original yolo10n model is a fixed adjustment of the perception field size, such as it can only achieve a good effect on fixed perception fields such as 4×4, 20×20, 40×40, etc., but it cannot provide a good effect for other perception fields, so the SK convolution replaces the PSA, so that the model can adaptively adjust the perception field size, so that it is not limited to one or several fixed perception fields. Then, recognition results of different scales are derived through upsampling and feature splicing, thereby realizing the result recognition of the image of the lesion site. Since the SK convolution has an attention mechanism, the scene is improved through the modified neural network model to improve the recognition accuracy.

[0055] S24: The identified lesion location and size information are converted into diagnostic data, and a specific disease report is generated;

[0056] In a preferred embodiment of the present invention, when the lesion site information is identified and a diagnosis report is formed, it is divided into two parts, one part is a manual information writing module, and the other part is an intelligent information generation module. The intelligent information generation module is mainly used to display and summarize the data after diagnosis to form a report on a specific disease, including the diagnosis results, the credibility of the diagnosis results, the diagnosis suggestions, etc. The manual information writing module is used by medical staff to write the corresponding patient information, including name, age, gender, examination number, examination date and department, etc. At the same time, the manual information writing module can also further modify the information in the disease report generated by the intelligent information module, and store pictures or cases with different diagnosis results between doctors and patients for further analysis.

[0057] Specifically, in the above embodiments, Figure 2 As shown, the above step S22: modifying the cropped image data through SK convolution and introducing the attention mechanism into the neural network model includes:

[0058] S220: performing a convolution operation on the input features using multiple convolution kernels of different sizes to obtain feature representations of multiple scales, and concatenating the obtained feature representations;

[0059] In a preferred embodiment of the present invention, the feature map I transmitted in the previous process H×W×C , first perform two transformations and Make their convolution kernels 3x3 and 5x5 respectively, and divide them into and

[0060] S221: Use fully connected layer, global average pooling and Relu activation function to get a new tensor;

[0061] In a preferred embodiment of the present invention, the convolution kernel in step S220 and Summing each element, we get U:

[0062]

[0063] The corresponding elements in the tensor of the same shape are added by formula ②. Then the global average pooling F gp Compress (B, C, H, W) to (B, C) to become a single feature vector s c .as follows:

[0064]

[0065] The global average pooling is performed through formula ③, which is compressed into feature vectors with the same number of channels to capture global information. Then, the fully connected layer F is pooled through formula ④. fc Perform dimensionality reduction or dimensionality increase to obtain channel descriptors corresponding to K scales, and reshape the feature vector after dimensionality increase to the same size as the input. The reshaped feature vector is stacked according to the 0th dimension (K dimension) to form a new tensor.

[0066] z=F fc (s)+δ(B(Ws)) ④

[0067] S222: normalize the weight corresponding to each scale through the softmax function so that the sum is 1;

[0068] In a preferred embodiment of the present invention, the weights of each feature scale are obtained by Softmax. Softmax is used in the multi-classification process, and it maps the outputs of multiple neurons to the interval (0,1).

[0069] Suppose we have an array V, Vi represents the i-th element in V, then the Softmax value of this element is:

[0070]

[0071] From formula ⑤, we can see that the weights corresponding to each scale are normalized by the softmax function so that their sum is 1, thereby changing the size of the feature map after downsampling the image data. For example, the original image data is 1600×1024, and it becomes 800×512 after downsampling by the softmax function.

[0072] S223: weighting and summing the weight of each scale with the result after the corresponding previous convolution to obtain different branch weight combinations to obtain a modified model input image and corresponding convolution layer;

[0073] In a preferred embodiment of the present invention, the weight after Softmax processing is weighted and summed with the convolution structure obtained in operation one to obtain different branch weight combinations, which affect the effective receptive field size of the fused level V.

[0074] The input image and convolutional layer of the yolo10n model are improved by the above method, and the attention mechanism is introduced into the neural network model, so that the model can adaptively fix the size of the perception field of view, improve the recognition accuracy of targets of different sizes, and further improve the recognition accuracy of the model for gastric diseases.

[0075] Corresponding to the above-mentioned diagnostic method, the present invention also provides a diagnostic system, specifically, as Figure 3 As shown, the diagnostic system includes: an image acquisition module 100 , an image preprocessing module 110 , an image recognition module 120 and an image reporting module 130 .

[0076] An image acquisition module 100 is used to acquire image data of the lesion site;

[0077] Specifically, real-time image data of the lesion site is first acquired, and the real-time image data is photoelectrically converted, and the converted data is decoded to form original image format data. The size of the original image format data collected is 1920X1080 resolution, and the original image format data is sent to the image preprocessing module.

[0078] An image preprocessing module 110 performs noise reduction and cropping on the collected image data to make it conform to the neural network input format;

[0079] Specifically, after receiving the real-time original image format data sent by the endoscope handle, the image format data is first subjected to denoising to eliminate image noise. Since the original model input image is square, and the real-time collected image format data is rectangular, it is not suitable to be directly put into the subsequent neural network, so the image needs to be properly cropped to make it conform to the neural network input format.

[0080] The real-time image data is processed in the above manner so that the pixels of the processed image data meet the requirements of the subsequent neural network input format.

[0081] The image recognition module 120 uses the modified neural network model or the trained neural network model to recognize the image, determine the abnormal situation and mark the abnormal position;

[0082] Specifically, the input image of the model is modified through SK convolution, and the attention mechanism is introduced into the neural network model. The processed graphic data is identified through the improved neural network model to determine whether it has an abnormality and mark the abnormal location. The input image of the neural network model is modified by SK convolution, and PSA is replaced by SKAttention mechanism, that is, using Figure 5 The SKNet structure shown replaces the original PSA structure to introduce an attention mechanism into the neural network model, so that targets of different proportions in the endoscope can be accurately detected to identify different lesion sizes and areas of different sizes. From the neural network model structure of yolo10n, it can be seen that it performs a maximum of 6 downsamplings, and each time the downsampled feature image becomes half of the original, that is, the width and height of the changed image data need to be The width and height of the image data must be integers, and the initial size data of the real-time image is 1920X1080. At the same time, when the image is actually displayed, there will be a small part of the two sides blocked, so the final cropped image size is determined to be 1600X1024. In order to identify different lesion sizes and different sized areas, so that targets of different proportions in endoscopic images can be accurately detected, SK convolution is used to modify the image input to the model and introduce the attention mechanism to obtain a modified neural network model, identify the image, determine whether there is an abnormality, and mark the abnormal position. At the same time, the neural network model is trained during the judgment to improve the accuracy of the neural network model. The specific modification method is as described above and will not be repeated here.

[0083] An image reporting module 130 generates a diagnosis report based on the analysis results of the image recognition module;

[0084] Specifically, when the lesion site information is identified and a diagnosis report is formed, it is divided into two parts, one is a manual information writing module, and the other is an intelligent information generation module. The intelligent information generation module is mainly used to display and summarize the data after diagnosis to form a report on a specific disease, including the diagnosis results, the credibility of the diagnosis results, diagnosis suggestions, etc. The manual information writing module is used by medical staff to write the corresponding patient information, including name, age, gender, examination number, examination date and department, etc. At the same time, the manual information writing module can also further modify the information in the disease report generated by the intelligent information module, and store pictures or cases with different diagnosis results between doctors and patients for further analysis.

[0085] Those skilled in the art will appreciate that all or part of the functions of the various methods in the above-mentioned embodiments can be implemented by hardware or by computer programs. When all or part of the functions in the above-mentioned embodiments are implemented by computer programs, the program can be stored in a computer-readable storage medium, and the storage medium can include: read-only memory, random access memory, disk, optical disk, hard disk, etc., and the program is executed by a computer to implement the above-mentioned functions. For example, the program is stored in the memory of the device, and when the program in the memory is executed by the processor, all or part of the above-mentioned functions can be implemented. In addition, when all or part of the functions in the above-mentioned embodiments are implemented by computer programs, the program can also be stored in a storage medium such as a server, another computer, disk, optical disk, flash disk or mobile hard disk, and can be downloaded or copied and saved in the memory of the local device, or the system of the local device is updated, and when the program in the memory is executed by the processor, all or part of the functions in the above-mentioned embodiments can be implemented.

[0086] Furthermore, the program instructions are used to perform the following steps when running:

[0087] Acquire real-time image data of the lesion area;

[0088] sending the acquired real-time image data and performing noise reduction on the received real-time image data;

[0089] The denoised image data is cropped to obtain image data that conforms to the neural network input format;

[0090] Use multiple convolution kernels of different sizes to perform convolution operations on the input features to obtain feature representations at multiple scales;

[0091] Concatenate the obtained feature representations;

[0092] Compressed into feature vectors with the same number of channels through global average pooling to capture global information;

[0093] The feature vector is first reduced in dimension and then increased in dimension to obtain channel descriptors corresponding to K scales;

[0094] Reshape the dimensionally upgraded feature vector to the same size as the input;

[0095] Stack the reshaped feature vectors to form a new tensor;

[0096] The weights of each feature scale are obtained through the softmax function;

[0097] The weights after softmax processing are weighted and summed with the obtained convolution structure to obtain the image input to the improved model and introduce the attention mechanism into the neural network model;

[0098] After SK convolution, the image is upsampled and features are concatenated to produce recognition results of different scales;

[0099] Generate and store a diagnostic report based on the identification results.

[0100] It should be understood that the present invention is described by some embodiments, and those skilled in the art will appreciate that various changes or equivalent substitutions may be made to these features and embodiments without departing from the spirit and scope of the present invention. In addition, under the teachings of the present invention, these features and embodiments may be modified to adapt to specific circumstances and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application are within the scope of protection of the present invention.

Claims

1. An image recognition method, characterized in that: include: Obtain at least one input feature of the target image, Use multiple convolution kernels of different sizes to perform convolution operations on the input features to obtain feature representations of multiple scales, and then concatenate the feature representations; Add corresponding elements in tensors of the same shape to get global information; After reducing the dimension of the global information and then increasing the dimension, the channel descriptors corresponding to K scales are obtained to form a new tensor; The weight of each scale is summed with the corresponding structure weight after the previous convolution to obtain different branch weight combinations to form the modified input features; The modified input features are used as the input values ​​of the neural network model and up-sampled and feature concatenated to induce recognition results of different scales.

2. The image recognition method according to claim 1, characterized in that: In the step of adding corresponding elements in tensors of the same shape to obtain global information, the method includes: First, multiple branch elements are summed up, and then global average pooling is performed to compress the summed elements into feature vectors with the same number of channels and capture global information.

3. The image recognition method according to claim 2, characterized in that: After said capturing the global information to obtain the global information; The global information is reduced and then increased in dimension to obtain channel descriptors corresponding to K scales, and the feature vector after the increase in dimension is reshaped to the same size as the input, and the reshaped feature vector is stacked according to the K dimensions to form a new tensor.

4. The image recognition method according to claim 1, characterized in that: The weight of each scale is summed with the corresponding structure after the previous convolution to obtain different branch weight combinations to form a modified input feature, and further includes: The weights corresponding to each scale are normalized through the softmax function, so that the size of the feature map is changed after the image data is downsampled.

5. The image recognition method according to claim 1, characterized in that: The step of using the modified input features as input values ​​of the neural network model includes: Introducing the attention mechanism into the neural network model enables the neural network model to adaptively adjust the field of view size.

6. A diagnostic method, characterized in that: include: Obtaining real-time image data of the lesion site and sending it; The obtained real-time image data is de-noised and used as the target image of the SK convolution image recognition method; Recognize the target image by using an image recognition method according to any one of claims 1 to 5, determine the abnormality of the image, and mark the abnormal position; Generate a diagnostic report based on the anomaly condition and / or anomaly location.

7. The diagnostic method according to claim 6, characterized in that: In the method of reducing the noise of the obtained real-time image data as a target image of the image recognition method of SK convolution, the method includes: The real-time image data is cropped after denoising, and the cropped image data is used as the target image of the SK convolution image recognition method.

8. The diagnostic method according to claim 7, characterized in that: The step of cutting the real-time image data after reducing noise also includes: The real-time image is cropped and normalized so that the pixel value is normalized from [0,255] to [0,1].

9. A diagnostic system, characterized in that: include: An image acquisition module, for acquiring image data of the lesion site; The image preprocessing module reduces noise and crops the collected image data to make it conform to the neural network input format; The image recognition module uses the modified neural network model or the trained neural network model to recognize the image, determine the abnormal situation and mark the abnormal location; The image reporting module generates a diagnosis report based on the analysis results of the image recognition module.

10. A storage medium having program instructions stored thereon, characterized in that: The program instructions are used to execute an image recognition method as claimed in any one of claims 1 to 5 when running.