An artificial intelligence-based gastrointestinal lesion localization method and device
Through the artificial intelligence-based gastrointestinal lesions positioning method, combined with local and global features, multi-scale feature extraction and image fusion are solved, and the problem of small differentiation of feature information and excessive redundancy in fuzzy boundaries is achieved, and high-accurate gastrointestinal lesions positioning is achieved.
Patent Information
- Application Number
- CN202411498097.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-25
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-10-25
AI Technical Summary
When facing fuzzy boundaries, the characteristic information is small and the redundant is too differentiated, resulting in inaccurate positioning of gastrointestinal lesions and reducing the accuracy of gastrointestinal lesions.
A method of gastrointestinal lesions based on artificial intelligence is proposed. By obtaining the captured images of the gastrointestinal tract of the target patient, image preprocessing and feature extraction are performed on the image, multi-scale feature extraction is performed based on local and global features, and the target feature map is obtained through image fusion to finally determine the lesion type and regional location.
By extracting local and global features and performing multi-scale feature extraction, features at different scales can be effectively captured, the ability to describe lesions can be improved, and the lesion area positioning is accurately performed, which improves the accuracy and accuracy of gastrointestinal lesions.
Smart Images

Figure CN119444855B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to a gastrointestinal lesion localization method and device based on artificial intelligence. Background Art
[0002] With the rising incidence of digestive tract diseases, gastroscopy, as an important diagnostic method, has been gradually popularized. Many patients choose to undergo gastroscopy due to symptoms such as indigestion, abdominal pain, and bleeding, in order to detect potential lesions at an early stage. However, traditional gastroscopy relies on doctors' experience and professional skills. During long-term endoscopic examinations, doctors need to concentrate and process a large amount of video data, which not only easily causes visual fatigue but also may affect the accurate judgment of lesion localization, increasing the risk of missed diagnosis and misdiagnosis.
[0003] To solve the above problems, technologies based on artificial intelligence and machine learning have begun to be applied to the recognition and localization of lesions. These emerging technologies can train on a large amount of endoscopic image data, extract effective features, and identify lesion areas, thereby reducing doctors' workload and improving the accuracy and efficiency of examinations. Although existing machine learning methods have theoretical advantages, in practical applications, the accuracy of lesion localization is still challenged. Especially when facing fuzzy boundaries, the differentiation of feature information is small and there is too much redundancy, resulting in inaccurate lesion localization and reducing the accuracy of gastrointestinal lesion localization. Summary of the Invention
[0004] The object of the present invention is to solve the problem that when facing fuzzy boundaries, the differentiation of feature information is small and there is too much redundancy, resulting in inaccurate lesion localization and reducing the accuracy of gastrointestinal lesion localization, and to propose a gastrointestinal lesion localization method and device based on artificial intelligence.
[0005] In the first aspect of the implementation of the present invention, a gastrointestinal lesion localization method based on artificial intelligence is first proposed. The method includes:
[0006] Obtain a captured image of the gastrointestinal tract of a target patient, perform image preprocessing on the captured image to obtain a first image, extract local features of the first image to obtain a local feature map, extract global features of the first image, and obtain a global feature map according to the local feature map and the global features;
[0007] Perform multi-scale feature extraction on the local feature map to obtain a first feature set, and perform multi-scale feature extraction on the global feature map to obtain a second feature set; the number of images in the first feature set is the same as that in the second feature set and the scales correspond one by one;
[0008] The first target feature map and the second target feature map are obtained according to the first feature map set and the second feature map set, and the first target feature map and the second target feature map are subjected to image fusion to obtain a target feature map;
[0009] The target feature map is substituted into a preset database to determine the lesion type, and the lesion area is located according to the lesion type and the target feature map.
[0010] Optionally, extracting the local features of the first image to obtain a local feature map includes:
[0011] Substituting the first image into a preset convolutional neural network model to obtain a first local feature map and a second local feature map;
[0012] Extracting the high-frequency feature map of the first local feature map to obtain a first high-frequency feature map, and performing feature superposition on the first high-frequency feature map and the first local feature map to obtain a first target local feature map;
[0013] Performing multi-scale feature extraction on the second local feature map to obtain a second-scale local feature map set, performing non-local operations on each second-scale local feature map in the second-scale local feature map set to obtain non-local feature maps, and superimposing all non-local feature maps to obtain a second target local feature map;
[0014] Superimposing the first target local feature map and the second target local feature map to obtain a local feature map.
[0015] Optionally, extracting the global features of the first image to obtain a global feature map includes:
[0016] Cutting the first image to obtain a first sub-image set, and substituting a target sub-image into an attention mechanism to obtain a target query vector, a target key vector, and a target value vector corresponding to the target sub-image; the target sub-image is any one in the first sub-image set;
[0017] Calculating the dot product of the target query vector and the key vector of each sub-image in the second sub-image set to obtain a weight value set corresponding to the target sub-image; the second sub-image set is the sub-image set obtained by removing the target sub-image from the first sub-image set; obtaining an output vector corresponding to the target sub-image according to the weight value set and the target value vector of each sub-image in the second sub-image set;
[0018] Performing weighted summation on the output vectors corresponding to all sub-images in the first sub-image set to obtain global features.
[0019] Optionally, obtaining the first target feature map and the second target feature map according to the first feature map set and the second feature map set includes:
[0020] The number of times of performing multi-scale feature extraction on the obtained local feature map is denoted as the extraction times K. The loop times C is obtained by subtracting one from the extraction times K. Target loop operations are performed according to the loop times C.
[0021] The target loop operations include
[0022] Step 1: Process the image and the image on the image to obtain the image and enter Step 2;
[0023] Step 2: Process the image and the image on the image to obtain the image where i is greater than 2 and less than K, and i is an integer;
[0024] Step 3: If i is not equal to K + 1, then increment i by 1, and repeat Step 2 until the image and the image are used to process the image to obtain the image and enter Step 4;
[0025] Step 4: Process the image and the image on the image to obtain the image and enter Step 5; Step 5: Process the image and the image on the image to obtain the image where i is greater than 2 and less than K, and i is an integer;
[0026] Step 6: If i is not equal to K + 1, then increment i by 1, and repeat Step 5 until the image and the image are used to process the image to obtain the image
[0027] Step 7: If j is not equal to C, then increment j by 1, and repeat Step 4 until j is equal to C, and enter Step 8
[0028] Step 8: If the sum of the superscript value and the subscript value of image A is K + 1, then denote this image as the target superimposed map, and superimpose all the target superimposed maps to obtain the final target feature map;
[0029] If image A is an image in the first feature map set, then image B is an image in the second feature map set. At this time, the final target feature map is denoted as the first target feature map; the is the image A obtained by the image of the i-th channel during the (j - 1)-th fusion; the is the image B obtained by the image of the i-th channel during the (j - 1)-th fusion;
[0030] Conversely, if image A is an image in the second feature map set, then image B is an image in the first feature map set. At this time, the final target feature map is denoted as the first target feature map.
[0031] Optionally, processing the image and the image to obtain the image includes: superimposing the image and the image and the image to obtain a target correction image, and performing first multi-core convolution on the target correction image to obtain a first corrected convolution image, a second corrected convolution image, and a third corrected convolution image;
[0032] Performing second multi-core convolution on the image to obtain a first convolution image, a second convolution image, and a third convolution image; the convolution kernels in the second multi-core convolution and the first multi-core convolution are transposed convolution kernels of each other;
[0033] Subtracting the first convolution image from the first corrected convolution image to obtain a first target image; subtracting the second convolution image from the second corrected convolution image to obtain a second target image; subtracting the third convolution image from the third corrected convolution image to obtain a third target image;
[0034] Superimposing the first target image, the second target image, and the third target image to obtain the image
[0035] In the second aspect of the implementation of the present invention, an artificial intelligence-based gastrointestinal lesion localization device is proposed, including: an image preprocessing module, configured to obtain a captured image of the gastrointestinal tract of a target patient, perform image preprocessing on the captured image to obtain a first image, extract local features of the first image to obtain a local feature map, extract global features of the first image, and obtain a global feature map according to the local feature map and the global features;
[0036] A scale feature extraction module, configured to perform multi-scale feature extraction on the local feature map to obtain a first feature map set, and perform multi-scale feature extraction on the global feature map to obtain a second feature map set; the number of images in the first feature map set is the same as that in the second feature map set and the scales correspond one by one;
[0037] A feature map fusion module, configured to obtain a first target feature map and a second target feature map according to the first feature map set and the second feature map set, and perform image fusion on the first target feature map and the second target feature map to obtain a target feature map; a lesion localization module, configured to substitute the target feature map into a preset database to determine the lesion type, and perform lesion area localization according to the lesion type and the target feature map.
[0038] Optionally, the image preprocessing module includes:
[0039] A local feature map generation module, configured to substitute the first image into a preset convolutional neural network model to obtain a first local feature map and a second local feature map;
[0040] A feature superposition module, configured to extract a high-frequency feature map of the first local feature map to obtain a first high-frequency feature map, and perform feature superposition on the first high-frequency feature map and the first local feature map to obtain a first target local feature map;
[0041] A non-local feature map superposition module, configured to perform multi-scale feature extraction on the second local feature map to obtain a second-scale local feature map set, perform non-local operations on each second-scale local feature map in the second-scale local feature map set to obtain non-local feature maps, and perform superposition on all non-local feature maps to obtain a second target local feature map;
[0042] A local feature superposition module, configured to perform superposition on the first target local feature map and the second target local feature map to obtain a local feature map.
[0043] Optionally, the image preprocessing module further includes:
[0044] An image cutting module, configured to cut the first image to obtain a first sub-image set, and substitute a target sub-image into an attention mechanism to obtain a target query vector, a target key vector, and a target value vector corresponding to the target sub-image; the target sub-image is any one in the first sub-image set;
[0045] A weight value calculation module, configured to calculate the dot product of the target query vector and the key vector of each sub-image in a second sub-image set to obtain a weight value set corresponding to the target sub-image; the second sub-image set is the sub-image set obtained by removing the target sub-image from the first sub-image set;
[0046] An output vector determination module, configured to obtain an output vector corresponding to the target sub-image according to the weight value set and the target value vector of each sub-image in the second sub-image set;
[0047] A weighted summation module, configured to perform weighted summation on the output vectors corresponding to all sub-images in the first sub-image set to obtain a global feature.
[0048] Optionally, the feature map fusion module includes:
[0049] A multi-scale feature extraction times determination module, configured to obtain the number of times of performing multi-scale feature extraction on the local feature map, denoted as the extraction times K, subtract one from the extraction times K to obtain the loop times C, and perform a target loop operation according to the loop times C; A loop module, where the target loop operation includes,
[0050] Step 1: Process image and image on image to obtain image and enter Step 2;
[0051] Step 2: Process image and image on image to obtain image where i is greater than 2 and less than K, and i is an integer;
[0052] Step 3: If i is not equal to K + 1, then increment i by 1, and repeat Step 2 until image and image are used to process image to obtain image and enter Step 4;
[0053] Step 4: Process image and image on image to obtain image and enter Step 5; Step 5: Process image and image on image to obtain image where i is greater than 2 and less than K, and i is an integer;
[0054] Step 6: If i is not equal to K + 1, then increment i by 1, and repeat Step 5 until image and image are used to process image to obtain image
[0055] Step 7: If j is not equal to C, then increment j by 1, and repeat Step 4 until j is equal to C, and enter Step 8
[0056] Step 8: If the sum of the superscript value and the subscript value of Image A is K + 1, then mark this image as the target superimposed image, and superimpose all the target superimposed images to obtain the final target feature map;
[0057] If Image A is an image in the first feature map set, then Image B is an image in the second feature map set. At this time, the final target feature map is marked as the first target feature map; the is the Image A obtained by the image of the i-th channel at the (j - 1)-th fusion; the is the Image B obtained by the image of the i-th channel at the (j - 1)-th fusion;
[0058] Conversely, if Image A is an image in the second feature map set, then Image B is an image in the first feature map set. At this time, the final target feature map is marked as the first target feature map.
[0059] Optionally, the loop module includes:
[0060] The first multi-core convolution module is used to superimpose Image and Image to obtain the target corrected image, and perform the first multi-core convolution on the target corrected image to obtain the first corrected convolution image, the second corrected convolution image, and the third corrected convolution image; the second multi-core convolution module is used to perform the second multi-core convolution on Image to obtain the first convolution image, the second convolution image, and the third convolution image; the convolution kernels in the second multi-core convolution and the first multi-core convolution are transposed convolution kernels; the convolution image fusion module is used to subtract the first convolution image from the first corrected convolution image to obtain the first target image; subtract the second convolution image from the second corrected convolution image to obtain the second target image; subtract the third convolution image from the third corrected convolution image to obtain the third target image;
[0061] The target image superimposing module is used to superimpose the first target image, the second target image, and the third target image to obtain Image
[0062] The beneficial effects of the present invention:
[0063] The present invention proposes a gastrointestinal lesion localization method based on artificial intelligence. By obtaining the captured images of the gastrointestinal tract of the target patient, performing image preprocessing on the captured images to obtain the first image, extracting the local features of the first image to obtain the local feature map, extracting the global features of the first image, and obtaining the global feature map according to the local feature map and the global features; performing multi-scale feature extraction on the local feature map to obtain the first feature map set, and performing multi-scale feature extraction on the global feature map to obtain the second feature map set; the number of images in the first feature map set is the same as that in the second feature map set and the scales correspond one by one; obtaining the first target feature map and the second target feature map according to the first feature map set and the second feature map set, and performing image fusion on the first target feature map and the second target feature map to obtain the target feature map; substituting the target feature map into the preset database to determine the lesion type, and performing lesion area localization according to the lesion type and the target feature map. By separately extracting local features and global features, it is possible to provide context information while capturing the subtle changes of the lesion, helping to understand the position of the lesion in the entire image. Performing multi-scale feature extraction on the local and global feature maps can effectively capture features at different scales, especially for lesions with blurred boundaries. Obtaining the target feature map through the method of image fusion for the target feature map can integrate local and global information, improve the description ability of the lesion, and then determine the lesion type in combination with historical data, so as to accurately perform lesion area localization, improving the accuracy and precision of gastrointestinal lesion localization. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] The present invention will be further described below with reference to the accompanying drawings.
[0065] Figure 1 FIG. is a flowchart of a gastrointestinal lesion localization method based on artificial intelligence provided by an embodiment of the present invention;
[0066] Figure 2 FIG. is a schematic structural diagram of a gastrointestinal lesion localization device based on artificial intelligence provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0067] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the descriptions such as "first" and "second" in the present invention are only for descriptive purposes, and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In addition, the technical solutions between the various embodiments can be combined with each other, but it must be based on what can be achieved by those of ordinary skill in the art. When the combination of technical solutions conflicts or cannot be realized, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.
[0068] Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0069] An embodiment of the present invention provides a method for locating gastrointestinal lesions based on artificial intelligence. See Figure 1 , Figure 1 is a flowchart of a method for locating gastrointestinal lesions based on artificial intelligence provided by an embodiment of the present invention. The method includes the following steps:
[0070] S101, obtain a captured image of the gastrointestinal tract of a target patient, perform image preprocessing on the captured image to obtain a first image, extract local features of the first image to obtain a local feature map, extract global features of the first image, and obtain a global feature map according to the local feature map and the global features;
[0071] S102, perform multi-scale feature extraction on the local feature map to obtain a first feature map set, and perform multi-scale feature extraction on the global feature map to obtain a second feature map set;
[0072] S103, obtain a first target feature map and a second target feature map according to the first feature map set and the second feature map set, and perform image fusion on the first target feature map and the second target feature map to obtain a target feature map;
[0073] S104, substitute the target feature map into a preset database to determine the lesion type, and perform lesion area location according to the lesion type and the target feature map.
[0074] Among them, the number of images in the first feature map set is the same as that in the second feature map set, and the scales correspond one by one.
[0075] Based on an artificial-intelligence-based gastrointestinal lesion localization method provided by an embodiment of the present invention, by separately extracting local features and global features, it is possible to capture subtle changes in lesions and also provide context information to help understand the position of the lesions in the entire image. By performing multi-scale feature extraction on the local and global feature maps, features at different scales can be effectively captured. Especially for lesions with blurred boundaries, the target feature map is obtained through an image fusion method, which can integrate local and global information, improve the ability to describe the lesions, and then determine the lesion type in combination with historical data, so as to accurately locate the lesion area and improve the accuracy and precision of gastrointestinal lesion localization.
[0076] In one implementation, through multi-scale feature extraction, lesions of different sizes and shapes can be captured more comprehensively, improving the accuracy of lesion localization, especially in the case of blurred boundaries; the combination of local features and global features enables the model to comprehensively consider details and the overall structure during analysis, enhancing the ability to identify lesions and reducing the risk of missed diagnosis.
[0077] In one implementation, the automated image analysis process reduces the visual fatigue of doctors during long-term examinations, helping to improve work efficiency and accuracy; the captured images of the gastrointestinal tract of the target patient can be obtained through techniques such as endoscopy, CT scan, or MRI, and the target patient is the patient seeking medical treatment.
[0078] In one implementation, the global feature map is obtained by combining the enhanced global features with the local features of the CNN branch. By combining global features and local features, the overall structure and detailed information of the image can be captured simultaneously, making the global feature map more comprehensive for subsequent lesion recognition.
[0079] In one implementation, the preset database stores images corresponding to different lesions, with the lesion types marked on these images, as well as the target feature maps corresponding to different lesions. By calculating the cosine similarity between the target feature map and the target feature maps corresponding to different lesions in the preset database, the lesion type closest to the target feature is found, and through the target feature map and the identified lesion type, the lesion area is obtained by applying bounding box regression.
[0080] In one embodiment, extracting the local features of the first image to obtain the local feature map includes:
[0081] Substituting the first image into a preset convolutional neural network model to obtain a first local feature map and a second local feature map;
[0082] Extracting the high-frequency feature map of the first local feature map to obtain a first high-frequency feature map, and performing feature superposition on the first high-frequency feature map and the first local feature map to obtain a first target local feature map;
[0083] Perform multi-scale feature extraction on the second local feature map to obtain a second-scale local feature map set. Perform non-local operations on each second-scale local feature map in the second-scale local feature map set to obtain non-local feature maps, and superimpose all non-local feature maps to obtain a second target local feature map;
[0084] Superimpose the first target local feature map and the second target local feature map to obtain a local feature map.
[0085] In one implementation, the preset convolutional neural network model uses ResNet-34 as the encoder; the first local feature map is a feature map obtained through a 3×3 convolutional kernel, and the second local feature map is a feature map obtained through a 7×7 convolutional kernel.
[0086] In one implementation, extracting the high-frequency feature map of the first local feature map to obtain the first high-frequency feature map specifically means performing Fourier transform on the first local feature map to extract high-frequency features to obtain the first high-frequency features, and performing inverse Fourier transform on the first high-frequency features to obtain the first high-frequency feature map.
[0087] In one implementation, the non-local operation is an image processing and feature extraction technique used to capture the global relationship between pixels in an image.
[0088] In one implementation, by extracting high-frequency feature maps and performing feature superposition, the detail expression ability of the local feature map can be improved, making it more sensitive when identifying lesions; the non-local operation can consider global context information, which helps to enhance the correlation between features, thereby enhancing the model's recognition ability for similar structures. Especially in the case of complex backgrounds, superimposing feature maps from different convolutional kernels can effectively integrate various information and reduce the possibility of misjudgment; by combining high-frequency and non-local features, the boundaries and shapes of lesions can be captured more precisely, thereby improving the positioning accuracy, especially in the case of blurred boundaries.
[0089] In one embodiment, extracting the global feature of the first image to obtain the global feature map includes:
[0090] Cut the first image to obtain a first sub-image set, substitute the target sub-image into the attention mechanism to obtain the target query vector, target key vector, and target value vector corresponding to the target sub-image; the target sub-image is any one in the first sub-image set; calculate the dot product of the target query vector and the key vectors of each sub-image in the second sub-image set to obtain the weight value set corresponding to the target sub-image; the second sub-image set is the sub-image set obtained by removing the target sub-image from the first sub-image set;
[0091] Obtain the output vector corresponding to the target sub-image according to the weight value set and the target value vector of each sub-image in the second sub-image set; perform weighted summation on the output vectors corresponding to all sub-images in the first sub-image set to obtain the global feature.
[0092] In one implementation, each sub-image in the first sub-image set is with a position encoding; the attention mechanism can more effectively extract the features related to the target sub-image and improve the sensitivity of the model to important information.
[0093] In one implementation, after calculating the dot product of the target query vector and the key vector, use the softmax function for all dot product values to ensure that the sum of the similarity values of all keys corresponding to each query is 1, forming a set of probability distributions to prepare for the calculation of the subsequent output vector.
[0094] In one implementation, for each sub-image in the second sub-image set, calculate the product of the weight value of the sub-image and the target value vector, and then sum up the products calculated for all sub-images to obtain the output vector.
[0095] In one embodiment, obtaining the first target feature map and the second target feature map from the first feature map set and the second feature map set includes:
[0096] Record the number of times of multi-scale feature extraction for obtaining the local feature map as the extraction times K, obtain the loop times C by subtracting one from the extraction times K, and perform the target loop operation according to the loop times C;
[0097] The target loop operation includes,
[0098] Step 1: Process the image and the image on the image to obtain the image and enter Step 2;
[0099] Step 2: Process the image and the image on the image to obtain the image where i is greater than 2 and less than K, and i is an integer;
[0100] Step 3: If i is not equal to K + 1, then increment i by 1, and repeat Step 2 until the image and the image are used to process the image to obtain the image and enter Step 4;
[0101] Step 4: Process the image and the image on the image to obtain the image Proceed to step 5; Step 5: Through the image and the image process the image to obtain the image where i is greater than 2 and less than K, and i is an integer;
[0102] Step 6: If i is not equal to K + 1, then increment i by 1 and repeat Step 5 until through the image and the image process the image to obtain the image
[0103] Step 7: If j is not equal to C, then increment j by 1 and repeat Step 4 until j is equal to C, then proceed to Step 8
[0104] Step 8: If the sum of the superscript value and the subscript value of image A is K + 1, then mark this image as the target superimposed image, and superimpose all the target superimposed images to obtain the final target feature map;
[0105] If image A is an image in the first feature map set, then image B is an image in the second feature map set. At this time, the final target feature map is marked as the first target feature map; is the image A obtained for the image of the i-th channel during the (j - 1)-th fusion; is the image B obtained for the image of the i-th channel during the (j - 1)-th fusion;
[0106] Conversely, if image A is an image in the second feature map set, then image B is an image in the first feature map set. At this time, the final target feature map is marked as the first target feature map.
[0107] In one implementation, through multiple loop processes, features of different scales can be fully extracted and fused, so that the final feature map contains more information and can better represent the image content; fusing features of different channels allows the model to utilize the characteristics of each channel to generate a more comprehensive and effective feature representation, improving the accuracy of lesion localization.
[0108] In one implementation, the images in the first feature map set are saved in ascending order of scale, and the images in the second feature map set are saved in ascending order of scale. When image A is an image in the first feature map set, the set of image A is Similarly, the set of image B is For example, when the multi-scale feature extraction of the local feature map is performed 3 times, and they are 16*16, 32*32, and 64*64 respectively, the set of image A is The set of image B is represents the 16*16 image in the first feature map set, Represents an image of 32*32 in the first feature map set, Represents an image of 64*64 in the first feature map set. The same applies to the set of image B. At this time, through image and image corrects image to obtain image Then image is the image obtained during the first fusion of the 16*16 image in the first feature map set.
[0109] In one embodiment, through image and image processes image to obtain image including:
[0110] Superimposes image and image to obtain a target corrected image, and performs a first multi-core convolution on the target corrected image to obtain a first corrected convolution image, a second corrected convolution image, and a third corrected convolution image;
[0111] Performs a second multi-core convolution on image to obtain a first convolution image, a second convolution image, and a third convolution image; the convolution kernels in the second multi-core convolution and the first multi-core convolution are transposed convolution kernels of each other;
[0112] Subtracts the first convolution image from the first corrected convolution image to obtain a first target image; subtracts the second convolution image from the second corrected convolution image to obtain a second target image; subtracts the third convolution image from the third corrected convolution image to obtain a third target image;
[0113] Superimposes the first target image, the second target image, and the third target image to obtain image
[0114] In one implementation, by superimposing and processing images A and B to generate a target corrected image, the features of these two images can be effectively fused, thereby improving the quality and expression ability of the features; using multi-core convolution can extract features from different angles, generate multiple convolution images, and provide a richer and more diverse feature representation.
[0115] In one implementation, setting the convolution kernels in the first multi-core convolution and the second multi-core convolution to be transposed to each other can effectively capture the spatial information of the image, enhance the learning ability of the model. By subtracting the convolution images, redundant or noise information can be effectively eliminated, making the target image more focused on important features and improving the accuracy of the model; superimposing multiple target images can effectively integrate features at different levels and scales, enhancing the expression ability of the final feature map.
[0116] In one implementation, for the image and the image overlaying them can integrate local features and global features together, and then subtracting the image to obtain the image enhances the image contour; through the differential processing of correcting the convolutional image and the original convolutional image, redundant information and noise are effectively removed, making the target features more accurate and improving the clarity of image expression.
[0117] In one implementation, for example, if the first multi-core convolution is 1*1, 1*3, 1*5, then the first corrected convolutional image is the image obtained from the target corrected image through a 1*1 convolutional kernel, the second corrected convolutional image is the image obtained from the target corrected image through a 1*3 convolutional kernel, and the third corrected convolutional image is the image obtained from the target corrected image through a 1*5 convolutional kernel. And if the convolutional kernels in the second multi-core convolution and the first multi-core convolution are transposed convolutional kernels, then the second multi-core convolution is 1*1, 3*1, 5*1. Then the first convolutional image is the image obtained through a 1*1 convolutional kernel, the second convolutional image is the image obtained through a 3*1 convolutional kernel, and the third convolutional image is the image obtained through a 5*1 convolutional kernel.
[0118] Based on the same inventive concept, the embodiment of the present invention also provides an artificial intelligence-based gastrointestinal lesion localization device. Refer to Figure 2 , Figure 2 which is a schematic structural diagram of an artificial intelligence-based gastrointestinal lesion localization device provided by the embodiment of the present invention, including:
[0119] An image preprocessing module, configured to obtain a captured image of the gastrointestinal tract of a target patient, perform image preprocessing on the captured image to obtain a first image, extract local features of the first image to obtain a local feature map, extract global features of the first image, and obtain a global feature map according to the local feature map and the global features;
[0120] A scale feature extraction module, configured to perform multi-scale feature extraction on the local feature map to obtain a first feature map set, and perform multi-scale feature extraction on the global feature map to obtain a second feature map set; the number of images in the first feature map set is the same as that in the second feature map set and the scales correspond one by one;
[0121] A feature map fusion module, configured to obtain a first target feature map and a second target feature map according to the first feature map set and the second feature map set, and perform image fusion on the first target feature map and the second target feature map to obtain a target feature map;
[0122] A lesion localization module, which is used to substitute the target feature map into a preset database to determine the lesion type, and perform lesion area localization according to the lesion type and the target feature map.
[0123] Based on an artificial intelligence-based gastrointestinal lesion localization device provided by an embodiment of the present invention, by separately extracting local features and global features, it can provide context information while capturing subtle changes in the lesion, helping to understand the position of the lesion in the entire image. Performing multi-scale feature extraction on local and global feature maps can effectively capture features at different scales. Especially for lesions with blurred boundaries, obtaining the target feature map through the method of image fusion for the target feature map can integrate local and global information, improve the description ability of the lesion, and then combine historical data to determine the lesion type, so as to accurately perform lesion area localization, improving the accuracy and precision of gastrointestinal lesion localization.
[0124] In one embodiment, the image preprocessing module includes:
[0125] A local feature map generation module, which is used to substitute the first image into a preset convolutional neural network model to obtain a first local feature map and a second local feature map;
[0126] A feature superposition module, which is used to extract the high-frequency feature map of the first local feature map to obtain a first high-frequency feature map, and perform feature superposition on the first high-frequency feature map and the first local feature map to obtain a first target local feature map;
[0127] A non-local feature map superposition module, which is used to perform multi-scale feature extraction on the second local feature map to obtain a second-scale local feature map set, perform non-local operations on each second-scale local feature map in the second-scale local feature map set to obtain non-local feature maps, and perform superposition on all non-local feature maps to obtain a second target local feature map;
[0128] A local feature superposition module, which is used to perform superposition on the first target local feature map and the second target local feature map to obtain a local feature map.
[0129] In one embodiment, the image preprocessing module further includes:
[0130] An image cutting module, which is used to cut the first image to obtain a first sub-image set, and substitute the target sub-image into the attention mechanism to obtain a target query vector, a target key vector, and a target value vector corresponding to the target sub-image; the target sub-image is any one in the first sub-image set;
[0131] A weight value calculation module, configured to calculate the dot product of the target query vector and the key vector of each sub-image in the second sub-image set to obtain a weight value set corresponding to the target sub-image; the second sub-image set is the sub-image set obtained by removing the target sub-image from the first sub-image set; an output vector determination module, configured to obtain an output vector corresponding to the target sub-image according to the weight value set and the target value vector of each sub-image in the second sub-image set;
[0132] A weighted summation module, configured to perform weighted summation on the output vectors corresponding to all sub-images in the first sub-image set to obtain a global feature.
[0133] In one embodiment, the feature map fusion module includes:
[0134] A multi-scale feature extraction times determination module, configured to obtain the number of times of multi-scale feature extraction of the local feature map, denoted as the extraction times K, subtract one from the extraction times K to obtain the loop times C, and perform a target loop operation according to the loop times C;
[0135] A loop module, where the target loop operation includes,
[0136] Step 1: Process the image and the image on the image to obtain the image and enter Step 2;
[0137] Step 2: Process the image and the image on the image to obtain the image where i is greater than 2 and less than K, and i is an integer;
[0138] Step 3: If i is not equal to K + 1, then increment i by 1, and repeat Step 2 until the image and the image are used to process the image to obtain the image and enter Step 4;
[0139] Step 4: Process the image and the image on the image to obtain the image and enter Step 5; Step 5: Process the image and the image on the image to obtain the image where i is greater than 2 and less than K, and i is an integer;
[0140] Step 6: If i is not equal to K + 1, then increment i by 1, and repeat Step 5 until the image and the image For the image Obtain the image
[0141] Step 7: If j is not equal to C, then increment j by 1, and repeat Step 4 until j is equal to C, then proceed to Step 8
[0142] Step 8: If the sum of the superscript value and the subscript value of image A is K + 1, then mark this image as the target superimposed image, and superimpose all the target superimposed images to obtain the final target feature map;
[0143] If image A is an image in the first feature map set, then image B is an image in the second feature map set. At this time, the final target feature map is marked as the first target feature map; Is the image A obtained for the i-th channel during the (j - 1)-th fusion; Is the image B obtained for the i-th channel during the (j - 1)-th fusion;
[0144] Conversely, if image A is an image in the second feature map set, then image B is an image in the first feature map set. At this time, the final target feature map is marked as the first target feature map.
[0145] In one embodiment, the loop module includes:
[0146] The first multi-core convolution module is used to superimpose the image and the image to obtain the target corrected image, and perform the first multi-core convolution on the target corrected image to obtain the first corrected convolution image, the second corrected convolution image, and the third corrected convolution image;
[0147] The second multi-core convolution module is used to perform the second multi-core convolution on the image to obtain the first convolution image, the second convolution image, and the third convolution image; The convolution kernels in the second multi-core convolution and the first multi-core convolution are transposed convolution kernels of each other;
[0148] The convolution image fusion module is used to subtract the first convolution image from the first corrected convolution image to obtain the first target image; subtract the second convolution image from the second corrected convolution image to obtain the second target image; subtract the third convolution image from the third corrected convolution image to obtain the third target image;
[0149] The target image superimposing module is used to superimpose the first target image, the second target image, and the third target image to obtain the image
[0150]
[0151] The above has described in detail an embodiment of the present invention, but the content is only a preferred embodiment of the present invention and cannot be considered as defining the scope of implementation of the present invention. All equivalent changes and improvements made according to the scope of the application of the present invention shall still fall within the scope covered by the patent of the present invention.
Claims
1. A gastrointestinal lesion localization method based on artificial intelligence, characterized in that: The method comprises: Acquire a captured image of the gastrointestinal tract of a target patient, perform image preprocessing on the captured image to obtain a first image, extract local features of the first image to obtain a local feature map, extract global features of the first image, and obtain a global feature map based on the local feature map and the global features; Performing multi-scale feature extraction on the local feature map to obtain a first feature map set, and performing multi-scale feature extraction on the global feature map to obtain a second feature map set; the number of images in the first feature map set and the number of images in the second feature map set are the same and the scales correspond one to one; Obtain a first target feature map and a second target feature map according to the first feature map set and the second feature map set, and perform image fusion on the first target feature map and the second target feature map to obtain a target feature map; Substituting the target feature map into a preset database to determine the type of lesion, and locating the lesion area according to the lesion type and the target feature map; Extracting the local features of the first image to obtain a local feature map includes: Substituting the first image into a preset convolutional neural network model to obtain a first local feature map and a second local feature map; Extracting a high-frequency feature map of the first local feature map to obtain a first high-frequency feature map, and superimposing features of the first high-frequency feature map and the first local feature map to obtain a first target local feature map; Performing multi-scale feature extraction on the second local feature map to obtain a second-scale local feature map set, performing a non-local operation on each second-scale local feature map in the second-scale local feature map set to obtain a non-local feature map, and superimposing all non-local feature maps to obtain a second target local feature map; The first target local feature map and the second target local feature map are superimposed to obtain a local feature map.
2. The method for locating gastrointestinal lesions based on artificial intelligence according to claim 1, characterized in that: Extracting the global features of the first image to obtain a global feature map includes: The first image is cut to obtain a first sub-image set, and a target sub-image is substituted into an attention mechanism to obtain a target query vector, a target key vector, and a target value vector corresponding to the target sub-image; the target sub-image is any one of the first sub-image set; Calculating the dot product of the target query vector and the key vector of each sub-image in a second sub-image set to obtain a weight value set corresponding to the target sub-image; the second sub-image set is a sub-image set obtained by removing the target sub-image from the first sub-image set; Obtaining an output vector corresponding to the target sub-image according to the weight value set and the target value vector of each sub-image in the second sub-image set; The global feature is obtained by performing weighted summation on the output vectors corresponding to all sub-images in the first sub-image set.
3. The method for locating gastrointestinal lesions based on artificial intelligence according to claim 1, characterized in that: Obtaining a first target feature map and a second target feature map according to the first feature map set and the second feature map set includes: The number of times the local feature map is obtained for multi-scale feature extraction is recorded as the extraction number K, the number of loops C is obtained by subtracting one from the extraction number K, and the target loop operation is performed according to the loop number C; The target cycle operation includes: Step 1: Through the image and images For images Processing to obtain image Go to step 2; Step 2: Through the image and images For images Processing to obtain image Where i is greater than 2 and i is less than K, i is an integer; Step 3: If i is not equal to K plus 1, then add 1 to i and repeat step 2 until the image is passed. and images For images Get the image Go to step 4; Step 4: Through the Image and images For images Processing to obtain image Go to step 5; Step 5: Through the Image and images For images Processing to obtain image Where i is greater than 2 and i is less than K, i is an integer; Step 6: If i is not equal to K plus 1, then i is increased by 1 and step 5 is repeated until the image is passed. and images For images Get the image Step 7: If j is not equal to C, add 1 to j and repeat step 4 until j is equal to C, then go to step 8 Step 8: If the sum of the superscript value and the subscript value of image A is K plus 1, then the image is recorded as the target overlay map, and all target overlay maps are superimposed to obtain the final target feature map; If image A is an image in the first feature atlas, then image B is an image in the second feature atlas, and the final target feature map is recorded as the first target feature map; is the image A obtained when the image of the i-th channel is fused for the j-1th time; is the image B obtained when the image of the i-th channel is fused for the j-1th time; On the contrary, if image A is an image in the second feature atlas, then image B is an image in the first feature atlas, and the final target feature map is recorded as the first target feature map.
4. The method for locating gastrointestinal lesions based on artificial intelligence according to claim 3, characterized in that: via image and images For images Processing to obtain image include: For images and images Superimposing to obtain a target corrected image, performing a first multi-core convolution on the target corrected image to obtain a first corrected convolution image, a second corrected convolution image and a third corrected convolution image; For images Performing a second multi-core convolution to obtain a first convolution image, a second convolution image, and a third convolution image; the convolution kernels in the second multi-core convolution and the first multi-core convolution are transposed convolution kernels of each other; The first rectified convolution image is subtracted from the first convolution image to obtain a first target image; the second rectified convolution image is subtracted from the second convolution image to obtain a second target image; the third rectified convolution image is subtracted from the third convolution image to obtain a third target image; The first target image, the second target image and the third target image are superimposed to obtain an image 5. A gastrointestinal lesion localization device based on artificial intelligence, characterized in that: The device comprises: An image preprocessing module, used to obtain a captured image of the gastrointestinal tract of a target patient, perform image preprocessing on the captured image to obtain a first image, extract local features of the first image to obtain a local feature map, extract global features of the first image, and obtain a global feature map based on the local feature map and the global features; A scale feature extraction module, configured to perform multi-scale feature extraction on the local feature map to obtain a first feature map set, and perform multi-scale feature extraction on the global feature map to obtain a second feature map set; the number of images in the first feature map set is the same as that in the second feature map set, and the scales thereof correspond one to one; A feature map fusion module, used to obtain a first target feature map and a second target feature map according to the first feature map set and the second feature map set, and perform image fusion on the first target feature map and the second target feature map to obtain a target feature map; A lesion localization module, used to substitute the target feature map into a preset database to determine the lesion type, and locate the lesion area according to the lesion type and the target feature map; The image preprocessing module comprises: A local feature map generation module, used for substituting the first image into a preset convolutional neural network model to obtain a first local feature map and a second local feature map; A feature superposition module, used for extracting a high-frequency feature map of the first local feature map to obtain a first high-frequency feature map, and performing feature superposition on the first high-frequency feature map and the first local feature map to obtain a first target local feature map; a non-local feature map superposition module, configured to perform multi-scale feature extraction on the second local feature map to obtain a second-scale local feature map set, perform a non-local operation on each second-scale local feature map in the second-scale local feature map set to obtain a non-local feature map, and superimpose all non-local feature maps to obtain a second target local feature map; A local feature superposition module is used to superimpose the first target local feature map and the second target local feature map to obtain a local feature map.
6. The gastrointestinal lesion localization device based on artificial intelligence according to claim 5, characterized in that: The image preprocessing module also includes: An image cutting module, configured to cut the first image to obtain a first sub-image set, and substitute a target sub-image into an attention mechanism to obtain a target query vector, a target key vector, and a target value vector corresponding to the target sub-image; the target sub-image is any one of the first sub-image set; A weight value calculation module, used for calculating the dot product of the target query vector and the key vector of each sub-image in a second sub-image set to obtain a weight value set corresponding to the target sub-image; the second sub-image set is a sub-image set after removing the target sub-image from the first sub-image set; an output vector determination module, configured to obtain an output vector corresponding to the target sub-image according to the weight value set and a target value vector of each sub-image in the second sub-image set; The weighted summation module is used to perform weighted summation on the output vectors corresponding to all sub-images in the first sub-image set to obtain global features.
7. The gastrointestinal lesion localization device based on artificial intelligence according to claim 5, characterized in that: The feature map fusion module includes: A multi-scale feature extraction times determination module is used to obtain the times of performing multi-scale feature extraction on the local feature map as the extraction times K, subtract one from the extraction times K to obtain the cycle times C, and perform the target cycle operation according to the cycle times C; A loop module, for said target loop operation, comprises: Step 1: Through the image and images For images Processing to obtain image Go to step 2; Step 2: Through the image and images For images Processing to obtain image Where i is greater than 2 and i is less than K, i is an integer; Step 3: If i is not equal to K plus 1, then add 1 to i and repeat step 2 until the image is passed. and images For images Get the image Go to step 4; Step 4: Through the Image and images For images Processing to obtain image Go to step 5; Step 5: Through the Image and images For images Processing to obtain image Where i is greater than 2 and i is less than K, i is an integer; Step 6: If i is not equal to K plus 1, then i is increased by 1 and step 5 is repeated until the image is passed. and images For images Get the image Step 7: If j is not equal to C, add 1 to j and repeat step 4 until j is equal to C, then go to step 8 Step 8: If the sum of the superscript value and the subscript value of image A is K plus 1, then the image is recorded as the target overlay map, and all target overlay maps are superimposed to obtain the final target feature map; If image A is an image in the first feature atlas, then image B is an image in the second feature atlas, and the final target feature map is recorded as the first target feature map; is the image A obtained when the image of the i-th channel is fused for the j-1th time; is the image B obtained when the image of the i-th channel is fused for the j-1th time; On the contrary, if image A is an image in the second feature atlas, then image B is an image in the first feature atlas, and the final target feature map is recorded as the first target feature map.
8. The gastrointestinal lesion localization device based on artificial intelligence according to claim 7, characterized in that: The circulation module comprises: The first multi-core convolution module is used to and images Superimposing to obtain a target corrected image, performing a first multi-core convolution on the target corrected image to obtain a first corrected convolution image, a second corrected convolution image and a third corrected convolution image; The second multi-core convolution module is used to Performing a second multi-core convolution to obtain a first convolution image, a second convolution image, and a third convolution image; the convolution kernels in the second multi-core convolution and the first multi-core convolution are transposed convolution kernels of each other; A convolution image fusion module, configured to obtain a first target image by subtracting the first convolution image from the first rectified convolution image; obtain a second target image by subtracting the second convolution image from the second rectified convolution image; and obtain a third target image by subtracting the third convolution image from the third rectified convolution image; A target image superposition module is used to superimpose the first target image, the second target image and the third target image to obtain an image
Citation Information
Patent Citations
Children pneumonia classification system and method based on hierarchical multi-scale feature fusion
CN116433973A