Image recognition method, diagnosis method, diagnosis system and storage medium

By introducing image recognition methods and attention mechanisms in traditional gastrointestinal endoscopy, the neural network model is improved, and misdiagnosis and missed diagnosis caused by relying on doctors' subjective judgment in traditional methods is solved, and accurate detection of targets at different scales and high accuracy diagnosis are achieved.

CN120107627APending Publication Date: 2025-06-06SHENZHEN EACHON BIO-TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510252442.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Traditional gastrointestinal endoscopy relies on the doctor's experience and subjective judgment, poses a risk of misdiagnosis and missed diagnosis, and is time-consuming and labor-intensive. It may cause a decrease in diagnosis accuracy due to doctor's fatigue or insufficient experience.

Method used

An image recognition method is adopted to obtain the input features and convolution kernel of the target image at the sampling position, combine the learnable offset field to generate offsets dynamically, perform sampling and feature stitching, introduce attention mechanisms, and improve neural network models to achieve accurate detection of targets at different scales.

Benefits of technology

It improves the flexibility and adaptability of the neural network model, enhances the accuracy of detection of small targets, greatly improves the computing speed and accuracy of the model, and provides high-level diagnostic support for the intelligent diagnosis of endoscopic images of ear, nose and throat.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107627A_ABST
    Figure CN120107627A_ABST
Patent Text Reader

Abstract

The invention relates to an image recognition method, a diagnosis method, a diagnosis system and a storage medium. The method comprises the following steps: acquiring real-time image data of a focus part; performing noise reduction and cutting on the obtained image data; dynamically generating offset by combining a learnable offset field with a convolutional layer; sampling values of all offset sampling positions are obtained according to the offset sampling positions and the offset; calculating output values of the sampling positions according to the sampling values of all the offset positions; and taking the output value as an input value of the neural network model to perform up-sampling and feature splicing so as to lead out recognition results of different scales, judging an image abnormal condition, and marking an abnormal position. AK convolution is introduced to modify the neural network model, so that the detection of target images with different shapes is realized, and the flexibility of the neural network model is enhanced. And meanwhile, an attention mechanism is introduced, so that the small target detection is more accurate, the adaptive capacity, the operation speed and the accuracy of the model are greatly improved, and high-level diagnosis is provided for intelligent diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image recognition methods, and in particular to an image recognition method, a diagnosis method, a diagnosis system and a computer-readable storage medium. Background Art

[0002] With the continuous development of global medical technology, endoscopic technology plays an increasingly important role in the diagnosis and treatment of gastrointestinal diseases. Traditional gastrointestinal endoscopy mainly relies on the doctor's experience and observation ability. By inserting an endoscope to visualize the digestive tract, various lesions including gastritis, gastric ulcers, intestinal polyps, tumors, etc. are detected and diagnosed. However, this method relies on the doctor's subjective judgment to a certain extent, and there is a risk of misdiagnosis and missed diagnosis. This mode that relies on human operation and judgment is not only time-consuming and laborious, but may also lead to a decrease in diagnostic accuracy due to the doctor's fatigue or lack of experience. At the same time, since the image quality of endoscopic examination may be affected by factors such as operating techniques or the patient's physical condition, it further increases the difficulty of accurate diagnosis. Summary of the invention

[0003] The technical problem to be solved by the present invention is that traditional gastrointestinal endoscopy mainly relies on the doctor's experience and observation ability, and visual inspection of the digestive tract is performed by inserting an endoscope to detect and diagnose various lesions including gastritis, gastric ulcer, intestinal polyps, tumors, etc. In addition, this method relies on the doctor's subjective judgment to a certain extent, and there is a risk of misdiagnosis and missed diagnosis. This mode that relies on human operation and judgment is not only time-consuming and laborious, but may also lead to a decrease in diagnostic accuracy due to fatigue or lack of experience of the doctor. In view of the above-mentioned defects of the prior art, an image recognition method, a diagnostic method, a diagnostic system and a storage medium are provided.

[0004] In order to solve the above technical problems, the technical solution adopted by the present invention is:

[0005] Construct an image recognition method, including:

[0006] Obtaining at least one input feature and at least one convolution kernel of a target image at a sampling position;

[0007] And dynamically generate offsets through learnable offset fields combined with convolutional layers;

[0008] Obtaining sampling values ​​of all offset sampling positions according to the offset sampling positions and offsets;

[0009] And calculate the output value of the sampling position according to the sampling values ​​of all offset positions;

[0010] The output values ​​are used as the input values ​​of the neural network model for upsampling and feature concatenation to induce recognition results of different scales.

[0011] Preferably, the process of obtaining the sampling values ​​of all offset sampling positions according to the offset sampling positions and offsets includes:

[0012] The actual sampling values ​​of all offset positions are calculated by linear interpolation, and the output value of the sampling position is calculated, and each position of the feature map of the output value is obtained by convolving the sampling position on the input feature map dynamically.

[0013] Preferably, after calculating the output value of the sampling position according to the sampling values ​​of all offset positions, the method further includes:

[0014] The attention mechanism is introduced into the neural network model to modify the neural network model, and then upsampling and feature concatenation are performed to induce recognition results of different scales.

[0015] Preferably, introducing an attention mechanism into the neural network model to modify the neural network model comprises:

[0016] Insert the query vector and the key vector into the neural network model, perform similarity calculation using the query vector and the key vector to obtain a weight, and perform normalization based on the similarity calculation score to obtain an attention weight.

[0017] Preferably, after obtaining the attention weight, the method further includes:

[0018] Insert the value vector into the neural network model, perform weighted summation on the attention weight and the value vector to obtain the final output, and upsample and feature concatenate the final output result to derive recognition results of different scales.

[0019] Construct a diagnostic method, including:

[0020] Obtaining real-time image data of the lesion site and sending it;

[0021] The obtained real-time image data is de-noised and used as the target image of the AK convolution image recognition method;

[0022] Recognize the target image by using an image recognition method as described above, determine the abnormality of the image, and mark the abnormal position;

[0023] Generate a diagnostic report based on the anomaly condition and / or anomaly location.

[0024] Preferably, in the target image of the image recognition method using the obtained real-time image data after denoising as AK convolution, it includes:

[0025] The real-time image data is cropped after denoising, and the cropped image data is used as the target image of the AK convolution image recognition method.

[0026] Preferably, the step of cropping the real-time image data after reducing noise further includes:

[0027] The real-time image is cropped and normalized so that the pixel value is normalized from [0,255] to [0,1].

[0028] Construct a diagnostic system, including:

[0029] An image acquisition module, for acquiring image data of the lesion site;

[0030] The image preprocessing module reduces noise and crops the collected image data to make it conform to the neural network input format;

[0031] Image recognition module, which uses a neural network model or a trained neural network model to recognize images, determine abnormal situations and mark abnormal locations;

[0032] Image reporting module, which generates diagnostic reports based on the analysis results of the image recognition module

[0033] A storage medium is constructed, on which program instructions are stored, characterized in that: the program instructions are used to execute an image recognition method as described above when running.

[0034] The beneficial effects of the present invention are: introducing AK convolution into the yolo10n neural network model to modify the neural network model, so that it can detect target images of different shapes, and enhance the flexibility of the neural network model. At the same time, introducing the attention mechanism into the neural network model makes it more accurate to detect small targets, greatly improving the adaptability, operation speed and accuracy of the model, and providing a high level of diagnosis for the intelligent diagnosis of ear, nose and throat endoscopic images. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the present invention will be further described below in conjunction with the accompanying drawings and embodiments. The drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative work:

[0036] Figure 1 A schematic diagram of a method flow of a diagnostic method according to a preferred embodiment of the present invention;

[0037] Figure 2 A flow chart of a method for introducing AK convolution and attention mechanism in a preferred embodiment of the present invention;

[0038] Figure 3 A schematic diagram of the principle of a diagnostic system according to a preferred embodiment of the present invention;

[0039] Figure 4 A structural diagram of an improved neural network model according to a preferred embodiment of the present invention;

[0040] Figure 5 Schematic diagram of the structure of a 5x5 convolution kernel in a preferred embodiment of the present invention;

[0041] Figure 6 A schematic diagram of a neural network model without introducing an attention mechanism in a preferred embodiment of the present invention;

[0042] Figure 7 A schematic diagram of a neural network model that introduces an attention mechanism in a preferred embodiment of the present invention;

[0043] Figure 8 Schematic diagram of the structure of the attention mechanism of a preferred embodiment of the present invention. DETAILED DESCRIPTION

[0044] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the following will be described clearly and completely in combination with the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are partial embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work belong to the protection scope of the present invention.

[0045] A diagnostic method according to a preferred embodiment of the present invention; Figure 1 As shown, it is a flowchart of the endoscope operation method provided by an embodiment of the present invention. The method can be executed by a device, and the device can be implemented by software and / or hardware.

[0046] Specifically, in this embodiment, the diagnostic method includes:

[0047] S20: Acquire and send real-time image data of the lesion area;

[0048] In a preferred embodiment of the present invention, real-time image data of the lesion site is first obtained. In the present invention, obtaining real-time image data of the lesion site such as the ear, nose and throat is used as an explanation. The real-time image data of the patient's ear, nose and throat are collected by the ENT endoscope handle and the ear, nose and laryngoscope and the real-time image data is sent to the image processing module. The ENT endoscope handle operates the movement of the ear, nose and laryngoscope and obtains the real-time image data of the ear, nose and throat. Then the ear, nose and laryngoscope performs photoelectric conversion on the real-time image data, and the converted data enters the endoscope handle for further decoding to form the original image format data. The size of the collected original image format data is 1280X720 resolution, and the original image format data is sent to the image processing module through the ENT endoscope handle.

[0049] S21: receiving real-time image data and performing noise reduction and cropping to obtain image data that conforms to the neural network input format;

[0050] In a preferred embodiment of the present invention, after receiving the real-time original image format data sent by the endoscope handle, the image format data is first subjected to denoising to eliminate image noise. Since the original model input image is a square, and the real-time collected image format data is a rectangle, it is not suitable to be directly placed in the subsequent neural network, so the image needs to be properly cropped to conform to the neural network input format. And since the size of the cropped image is 960X768, some corners will be blocked during actual display, so the cropped image needs to be normalized, and the pixels of the cropped image data are normalized from [0,255] to [0,1]. The specific normalization formula is:

[0051]

[0052] The real-time image data of gastrointestinal images are processed in the above manner so that the pixels of the processed image data meet the requirements of the subsequent neural network input format.

[0053] S22: Modify the input image of the cropped image data through AK convolution and introduce the attention mechanism into the neural network model;

[0054] In a preferred embodiment of the present invention, a trained neural network model is used to identify the processed graphic data to determine whether there are abnormalities and mark the abnormal locations. The yolo10n neural network model can be used to identify lesions, wherein the original paper of yolo10n was published on arxiv, and the paper address is: https: / / arxiv.org / pdf / 2405.14458, which will not be described in detail here. However, the original model training of the yolo10n neural network model uses image data of size 640X640 as the input training image. The image data of this size is small and often cannot meet the needs of actual production. Therefore, the present invention introduces AK convolution to modify the input image and the convolution layer of the neural network model, and introduces the AKConv mechanism in the yolo10n neural network model, that is, using Figure 4 The AKConv mechanism structure shown in the figure introduces the attention mechanism into the neural network model, so that objects of different scales in the endoscopic image can be accurately detected to identify different lesion sizes and areas of different sizes, and enhance the recognition rate of objects of different scales. From the neural network model structure of yolo10n, it can be seen that it performs a maximum of 6 downsamplings, and each time the downsampled feature image becomes half of the original, that is, the width and height of the changed image data need to be The width and height of the image data must be integers, and the initial size data of the real-time image obtained is 1280X720. At the same time, when the image is actually displayed, the horizontal area will be partially blocked, and the vertical area will be filled to meet the requirements, so the final determined image size is 960×768.

[0055] S23: The lesion location and size of the image are identified through the modified neural network model, and the lesion is marked and judged whether there is an abnormality;

[0056] In a preferred embodiment of the present invention, the input image is modified by introducing a changeable kernel convolution, and the original yolo10n model is modified to introduce an attention mechanism into the neural network model to enhance the flexibility of the neural network model for the convolution kernel. Since the convolution operation in the original yolo10n model has two inherent defects, on the one hand, the convolution operation is limited to the local window, and the information of other positions cannot be captured, and its sampling shape is fixed. On the other hand, the size of the convolution kernel is k×k, which is a fixed square, and the number of parameters tends to grow squared with the size. The shapes and sizes of targets in different data sets and different positions are different. Therefore, the convolution kernel with a fixed sample shape and square cannot adapt well to the changing targets. Therefore, by introducing AK convolution to modify the neural network model, the model can better adapt to the changing targets and enhance the flexibility of the convolution kernel. Then, recognition results of different scales are derived by upsampling and feature splicing, thereby realizing the result recognition of the image of the lesion site. Since AK convolution has an attention mechanism, the modified neural network model greatly improves the adaptability, operation speed and accuracy of the model, and provides a high level of diagnosis for the intelligent diagnosis of ear, nose and throat endoscopic images.

[0057] S24: The identified lesion location and size information are converted into diagnostic data, and a specific disease report is generated;

[0058] In a preferred embodiment of the present invention, when the lesion site information is identified and a diagnosis report is formed, it is divided into two parts, one part is a manual information writing module, and the other part is an intelligent information generation module. The intelligent information generation module is mainly used to display and summarize the data after diagnosis to form a report on a specific disease, including the diagnosis results, the credibility of the diagnosis results, the diagnosis suggestions, etc. The manual information writing module is used by medical staff to write the corresponding patient information, including name, age, gender, examination number, examination date and department, etc. At the same time, the manual information writing module can also further modify the information in the disease report generated by the intelligent information module, and store pictures or cases with different diagnosis results between doctors and patients for further analysis.

[0059] Specifically, in the above embodiments, Figure 2As shown, the above step S22: modifying the cropped image data through AK convolution and introducing the attention mechanism into the neural network model includes:

[0060] S220: Dynamically generate offsets through learnable offset fields combined with convolutional layers;

[0061] In the preferred embodiment of the present invention, since the convolution operation in the yolo10n neural network model has two defects, AK convolution is introduced into the neural network model to enhance the flexibility of the convolution kernel, so that it can collect the shapes and sizes of targets in different data sets and different positions. A more flexible convolution kernel can better detect targets of different shapes. In the operation of standard convolution, given an input feature map X and a convolution kernel W, each position y(p) of the output feature map Y is obtained by performing a dot product calculation at a fixed sampling position:

[0062]

[0063] Where p represents the output position, Δp k is the offset of the convolution kernel relative to position p, and K is the size of the convolution kernel.

[0064] The kernel convolution method can be changed to enhance the flexibility of the convolution kernel, the offset Δp k is not fixed, but rather through a learnable offset field This offset field is learned through an additional convolutional layer:

[0065]

[0066] Among them, W offet is the convolution kernel used to learn the offset, and f(*) is a convolution operation.

[0067] S221: Obtain sampling values ​​of all offset sampling positions through the offset sampling positions and offsets, and output a deformable convolution kernel to modify the neural network model;

[0068] In a preferred embodiment of the present invention, due to the offset sampling position It may not be on the discrete grid of the input feature map, so linear interpolation is needed to calculate the actual sample value. Falling on four pixels p 0 ,p 1 ,p 2 ,p 3 The sampling value is between Calculated by the following formula:

[0069]

[0070] Among them, w i is the interpolation weight, given by is determined by the relative position of .

[0071] After obtaining the values ​​of all offset sampling positions, the output of the deformable convolution kernel can be calculated by the following formula:

[0072]

[0073] It can be seen from formula ⑤ that each position y(p) of the output feature map is obtained by convolution of the dynamically adjusted position on the input feature map.

[0074] During training, the offset of the deformable convolution kernel and the convolution kernel W are all learnable parameters. The gradient of the offset is also calculated using the chain rule:

[0075]

[0076] By introducing the AK convolution into the neural function network model through the above method, the improved neural network model can improve the detection of targets of different shapes and improve the adaptability of the model.

[0077] S222: Calculate the similarity between the query vector and the key vector to obtain a weight;

[0078] In the preferred embodiment of the present invention, since many diseases are in the early stage of onset during endoscopic examination, the abnormal area is small, that is, the target in the image is very small. In order to more accurately identify the abnormal state in the early stage of onset, an attention mechanism is added to the SPPF of the neural network model. Figure 6 The SPPF in the original model is given in Figure 7 The improved SPPF attention is given in.

[0079] like Figure 8 The principle diagram of the attention mechanism shown in Figure 1 shows that for the attention mechanism, the similarity between the query vector and the key vector is first calculated to obtain the weight. The formula is as follows:

[0080] score(q,k i )=q·k i ⑦

[0081] Among them: the query vector (Query), abbreviated as q, represents the content we want to focus on; the key vector (Key), abbreviated as k, is compared with the query vector to determine its relevance.

[0082] S223: Normalize the similarity scores through the softmax function to obtain the attention weight;

[0083] In a preferred embodiment of the present invention, the similarity score in step S222 is normalized by a softmax function to obtain the attention weight:

[0084]

[0085] Here α i is the attention weight corresponding to the i-th key, and n is the number of key-value pairs.

[0086] S224: Use the attention weight to sum all value vectors to obtain the final output, so that the attention mechanism is introduced into the neural network model;

[0087] In a preferred embodiment of the present invention, weighted summation is performed using the attention weight α i For all value vectors v i Perform weighted summation to get the final output: The formula is as follows:

[0088]

[0089] Among them: the value vector (Value), abbreviated as v, is aggregated according to the relevance weight to obtain the final output, thus obtaining the final result output. Thus, the attention mechanism is introduced into the modified neural network model, the detection ability of the neural network model for small targets is enhanced, and the adaptability of the model is improved.

[0090] By introducing AK convolution into the yolo10n neural network model through the above method, it can detect target images of different shapes and enhance the flexibility of the neural network model. At the same time, the attention mechanism is introduced into the neural network model to make it more accurate in detecting small targets, which greatly improves the adaptability, computing speed and accuracy of the model, and provides a high level of diagnosis for the intelligent diagnosis of ear, nose and throat endoscope images.

[0091] Corresponding to the above-mentioned diagnostic method, the present invention also provides a diagnostic system, specifically, as Figure 3 As shown, the diagnostic system includes: an image acquisition module 100 , an image preprocessing module 110 , an image recognition module 120 and an image reporting module 130 .

[0092] An image acquisition module 100 is used to acquire image data of the lesion site;

[0093] Specifically, real-time image data of the lesion site is first acquired, and the real-time image data is photoelectrically converted, and the converted data is decoded to form original image format data. The size of the original image format data collected is 1280X720 resolution, and the original image format data is sent to the image preprocessing module.

[0094] An image preprocessing module 110 performs noise reduction and cropping on the collected image data to make it conform to the neural network input format;

[0095] Specifically, after receiving the real-time original image format data sent by the endoscope handle, the image format data is first subjected to denoising to eliminate image noise. Since the original model input image is square, and the real-time collected image format data is rectangular, it is not suitable to be directly put into the subsequent neural network, so the image needs to be properly cropped to make it conform to the neural network input format.

[0096] The real-time image data is processed in the above manner so that the pixels of the processed image data meet the requirements of the subsequent neural network input format.

[0097] The image recognition module 120 uses the modified neural network model or the trained neural network model to recognize the image, determine the abnormal situation and mark the abnormal position;

[0098] Specifically, the input image of the model is modified by introducing AK convolution into the neural network model, and the attention mechanism is introduced into the neural network model. The processed graphic data is identified by the improved neural network model to determine whether it is abnormal and mark the abnormal position. The input image of the neural network model is modified by using AK convolution, and the AKConv mechanism is introduced into the yolo10n neural network model, that is, using Figure 4 The AKConv mechanism structure shown in the figure introduces the attention mechanism into the neural network model, so that objects of different scales in the endoscopic image can be accurately detected to identify different lesion sizes and areas of different sizes, and enhance the recognition rate of objects of different scales. From the neural network model structure of yolo10n, it can be seen that it performs a maximum of 6 downsamplings, and each time the downsampled feature image becomes half of the original, that is, the width and height of the changed image data need to be The width and height of the image data must be integers, and the initial size data of the real-time image obtained is 1280X720. At the same time, when the image is actually displayed, there will be a part of the occlusion in the horizontal area, and the vertical area is filled to meet the requirements, so the final determined image size is 960×768. By introducing AK convolution into the neural network model, the shapes and sizes of targets in different data sets and different positions can be collected, so that targets of different shapes in endoscopic images can be detected, and the flexibility of the neural network model is improved. The attention mechanism is introduced into the modified neural network model to identify small target images, so that it can more accurately identify the abnormal state in the early stage of the disease, determine whether there is an abnormality, and mark the abnormal position. At the same time, the neural network model is trained during judgment to improve the accuracy of the neural network model, and the adaptability, computing speed and accuracy of the neural network model are greatly improved. The specific modification method is as described above, and will not be repeated here.

[0099] An image reporting module 130 generates a diagnosis report based on the analysis results of the image recognition module;

[0100] Specifically, when the lesion site information is identified and a diagnosis report is formed, it is divided into two parts, one is a manual information writing module, and the other is an intelligent information generation module. The intelligent information generation module is mainly used to display and summarize the data after diagnosis to form a report on a specific disease, including the diagnosis results, the credibility of the diagnosis results, diagnosis suggestions, etc. The manual information writing module is used by medical staff to write the corresponding patient information, including name, age, gender, examination number, examination date and department, etc. At the same time, the manual information writing module can also further modify the information in the disease report generated by the intelligent information module, and store pictures or cases with different diagnosis results between doctors and patients for further analysis.

[0101] Those skilled in the art will appreciate that all or part of the functions of the various methods in the above-mentioned embodiments can be implemented by hardware or by computer programs. When all or part of the functions in the above-mentioned embodiments are implemented by computer programs, the program can be stored in a computer-readable storage medium, and the storage medium can include: read-only memory, random access memory, disk, optical disk, hard disk, etc., and the program is executed by a computer to implement the above-mentioned functions. For example, the program is stored in the memory of the device, and when the program in the memory is executed by the processor, all or part of the above-mentioned functions can be implemented. In addition, when all or part of the functions in the above-mentioned embodiments are implemented by computer programs, the program can also be stored in a storage medium such as a server, another computer, disk, optical disk, flash disk or mobile hard disk, and can be downloaded or copied and saved in the memory of the local device, or the system of the local device is updated, and when the program in the memory is executed by the processor, all or part of the functions in the above-mentioned embodiments can be implemented.

[0102] Furthermore, the program instructions are used to perform the following steps when running:

[0103] Acquire real-time image data of the lesion area;

[0104] sending the acquired real-time image data and performing noise reduction on the received real-time image data;

[0105] The denoised image data is cropped to obtain image data that conforms to the neural network input format;

[0106] Dynamically generate offsets through learnable offset fields combined with convolutional layers;

[0107] Obtain the sampling values ​​of all offset sampling positions through the offset sampling position and offset;

[0108] And output deformable convolution kernel to modify the neural network model;

[0109] Calculate the similarity between the query vector and the key vector to obtain the weight;

[0110] Normalize the similarity scores through the softmax function to obtain the attention weight;

[0111] Use the attention weights to sum all value vectors to get the final output;

[0112] The AK convolution and attention mechanism are introduced into the output neural network model;

[0113] Upsampling and feature concatenation of images lead to recognition results of different scales, and the improved neural network model is trained;

[0114] Generate and store a diagnostic report based on the identification results.

[0115] It should be understood that the present invention is described by some embodiments, and those skilled in the art will appreciate that various changes or equivalent substitutions may be made to these features and embodiments without departing from the spirit and scope of the present invention. In addition, under the teachings of the present invention, these features and embodiments may be modified to adapt to specific circumstances and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application are within the scope of protection of the present invention.

Claims

1. An image recognition method, It is characterized in that include: Obtaining at least one input feature and at least one convolution kernel of a target image at a sampling position; And dynamically generate offsets through learnable offset fields combined with convolutional layers; Obtaining sampling values ​​of all offset sampling positions according to the offset sampling positions and offsets; And calculate the output value of the sampling position according to the sampling values ​​of all offset positions; The output values ​​are used as the input values ​​of the neural network model for upsampling and feature concatenation to induce recognition results of different scales.

2. The image recognition method according to claim 1, Features: The process of obtaining the sampling values ​​of all offset sampling positions according to the offset sampling positions and the offsets includes: The actual sampling values ​​of all offset positions are calculated by linear interpolation, and the output value of the sampling position is calculated, and each position of the feature map of the output value is obtained by convolving the sampling position on the input feature map dynamically.

3. The image recognition method according to claim 1, Features: After calculating the output value of the sampling position according to the sampling values ​​of all offset positions, the method further includes: The attention mechanism is introduced into the neural network model to modify the neural network model, and then upsampling and feature concatenation are performed to induce recognition results of different scales.

4. The image recognition method according to claim 3, Features: Introducing an attention mechanism into the neural network model to modify the neural network model includes: Insert the query vector and the key vector into the neural network model, perform similarity calculation using the query vector and the key vector to obtain a weight, and perform normalization based on the similarity calculation score to obtain an attention weight.

5. The image recognition method according to claim 4, Features: After obtaining the attention weight, the method further includes: Insert the value vector into the neural network model, perform weighted summation on the attention weight and the value vector to obtain the final output, and upsample and feature concatenate the final output result to derive recognition results of different scales.

6. A diagnostic method, It is characterized in that include: Obtaining real-time image data of the lesion site and sending it; The obtained real-time image data is de-noised and used as the target image of the AK convolution image recognition method; Recognize the target image by using an image recognition method according to any one of claims 1 to 5, determine the abnormality of the image, and mark the abnormal position; Generate a diagnostic report based on the anomaly condition and / or anomaly location.

7. The diagnostic method according to claim 6, Features: In the method of using the obtained real-time image data after noise reduction as the target image of the image recognition method of AK convolution, the method includes: The real-time image data is cropped after denoising, and the cropped image data is used as the target image of the AK convolution image recognition method.

8. The diagnostic method according to claim 7, Features: The step of cutting the real-time image data after reducing noise also includes: The real-time image is cropped and normalized so that the pixel value is normalized from [0,255] to [0,1].

9. A diagnostic system, It is characterized in that include: An image acquisition module, for acquiring image data of the lesion site; The image preprocessing module reduces noise and crops the collected image data to make it conform to the neural network input format; Image recognition module, which uses a neural network model or a trained neural network model to recognize images, determine abnormal situations and mark abnormal locations; The image reporting module generates a diagnosis report based on the analysis results of the image recognition module.

10. A storage medium having program instructions stored thereon, Features: The program instructions are used to execute an image recognition method as claimed in any one of claims 1 to 5 when running.