Lung cancer image classification method, system and medium based on multi-source heterogeneous information fusion
Through the exchange of multi-source heterogeneous images and the amplification and segmentation of images, rich training images are formed, and multi-model training technology is used to solve the problems of insufficient training data and low classification accuracy in the existing technology, which significantly improves the classification accuracy of lung cancer images.
Patent Information
- Application Number
- CN202411002971.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-25
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-07-25
AI Technical Summary
The prior art is difficult to expand the training data when the number of training images is small, and it is not effectively segmented and processed when the input images are large, resulting in low classification accuracy, especially for pathological types with high similarity, it is difficult to improve the recognition accuracy.
By collecting multi-source heterogeneous images, partial exchanges are performed to obtain more kinds of images, and the original image is enlarged and segmented into multiple sub-images to form a rich training image. Then, the model is trained using the neural network convolution algorithm and the recognition accuracy of pathological types is improved by creating multiple models (such as the iA-th model and the iB-th model).
Through rich training images and multi-model training, the classification accuracy of lung cancer images is significantly improved, especially among similar pathological types, which can be more accurately distinguished and identified.
Smart Images

Figure CN118864975B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and more specifically to a lung cancer image classification method, system and medium based on multi-source heterogeneous information fusion. Background Art
[0002] With the continuous advancement of science and technology, the medical field is also developing continuously. Among them, the application of medical imaging technology is becoming more and more extensive. By analyzing the medical images of patients, the workload of medical staff can be reduced, and the recognition accuracy of medical images can be improved. For example: Chinese patent CN114267432B, the invention discloses a method, device, and medium for generating alternating training and medical image classification of a generative adversarial network. By measuring the convergence degree of the generator and the discriminator, the training convergence degree of the generator and the discriminator can be quantified; by formulating an adaptive alternating training strategy, it can be adaptively determined whether to train the generator or the discriminator next time according to the convergence degree of the current generator and the discriminator; by classifying medical images based on an adaptive generative adversarial network, the classification accuracy and recognition effect can be improved, and medical staff can be assisted in diagnosing diseases. For example, in US Patent US20230146953A1, a trained machine learning image generator is used to generate a set of training images based on three-dimensional patient imaging data, wherein each training image is marked with a corresponding two-dimensional projection projection angle. Using the training image set, a machine learning image classifier model is trained to identify the patient rotation angle in an X-ray image. The X-ray images are processed using a machine learning image classifier model to identify the patient's rotation angle. The machine learning medical condition classifier model is trained to identify medical conditions using labeled X-ray images. The machine learning medical condition classifier model determines the indication of medical conditions in the patient's X-ray images. Both of the above patents implement the classification of medical images, but neither considers expanding the training images and increasing the amount of training data when the number of training images is small, nor considers dividing the image into multiple parts when the input image is large, nor does it create models and train for pathological types with high similarity to improve classification accuracy. Summary of the invention
[0003] In order to better solve the above problems, the present invention provides a lung cancer image classification method based on multi-source heterogeneous information fusion, the method comprising the following steps:
[0004] Step S1: acquiring a plurality of multi-source heterogeneous images of known pathological types from a storage unit through a collection unit, and acquiring an original image based on the multi-source heterogeneous images;
[0005] Step S2: Enlarging the original image to a preset size, and using the enlarged original image as a first image; when the area of the first image is greater than or equal to a set area, dividing the first image into a plurality of second images through a segmentation unit, using the first image whose area is smaller than the set area as a third image, and removing the backgrounds of the second image and the third image and enlarging them to a plurality of preset multiples to obtain training images;
[0006] Step S3: Generate a first model through a model creation unit, and use the training image to train the first model, magnify the patient's target image to a plurality of the preset multiples and input it into the first model to obtain a plurality of target sub-images;
[0007] Step S4: acquiring the training image corresponding to each pathological type according to the pathological type corresponding to the original image, and comparing the training images of any two pathological types to acquire two pathological types with the highest similarity to the i-th pathological type, and creating an iA model and an iB model corresponding to the i-th pathological type by the model creation unit, and training the iA model and the iB model based on the i-th pathological type and the training images corresponding to the two pathological types;
[0008] Step S5: Input the multiple target sub-images into the iA model and the iB model corresponding to the i-th pathological type respectively, and judge the pathological type to which the target image belongs according to the sum of the matching degrees between the target image and the i-th pathological type respectively output by the iA model and the iB model.
[0009] As a preferred technical solution of the present invention, in step S1, the pathological information contained in the multi-source heterogeneous image is multi-source heterogeneous information. The multi-source heterogeneous image obtains the number of images from each source based on images including multiple different sources. When the number of images is less than a set number, the effective positions of any two multi-source heterogeneous images with the same pathological type in the images of the source are partially exchanged to obtain a new original image.
[0010] As a preferred technical solution of the present invention, step S3 includes: creating the first model through the model creation unit, and using the training image to train the first model through a neural network convolution algorithm, and magnifying the target image of the patient to multiple preset multiples, and using the enlarged image as an input image, inputting multiple input images into the first model and obtaining multiple groups of output images, wherein the output image is an image after all the input images are segmented, and obtaining the target sub-image based on the clustering result of the output image, wherein the target sub-image includes the entire content of the target image.
[0011] As a preferred technical solution of the present invention, step S4 includes the following steps:
[0012] Step S41: acquiring the training image corresponding to each pathological type according to the pathological type corresponding to the original image, and comparing any two training images of the pathological type to obtain a comparison result;
[0013] Step S42: according to the comparison result, obtaining the jth pathological type and the kth pathological type having the highest similarity to the ith pathological type;
[0014] Step S43: Create the iA model and the iB model through the model creation unit, and train the iA model through the first learning image corresponding to the i-th pathological type and the j-th pathological type in the training image, and train the iB model through the second learning image corresponding to the i-th pathological type and the k-th pathological type in the training image, wherein the value range of i is a positive integer greater than or equal to 1 and less than or equal to N, and N is the number of all pathological types corresponding to the original image.
[0015] As a preferred technical solution of the present invention, step S5 comprises the following steps:
[0016] Step S51: inputting a plurality of the target sub-images into the iA model and the iB model respectively, and obtaining a first value and a second value output by the iA model, and a third value and a fourth value output by the iB model, wherein the first value is the matching degree between the target image output by the iA model and the i-th pathological type, and the second value is the matching degree between the target image output by the iA model and the j-th pathological type; the third value is the matching degree between the target image output by the iB model and the i-th pathological type, and the fourth value is the matching degree between the target image output by the iB model and the k-th pathological type;
[0017] Step S52: calculating the sum of the first value and the third value as the i-th matching degree cumulative value, and when the m-th matching degree cumulative value is the largest among all the i-th matching degree cumulative values, and the first value and the second value corresponding to the m-th matching degree cumulative value are both greater than a first threshold, taking the m-th pathological type corresponding to the m-th matching degree cumulative value as the pathological type corresponding to the target image;
[0018] The value range of j, k, and m is a positive integer greater than or equal to 1 and less than or equal to N, where N is the number of all pathological types corresponding to the original image.
[0019] As a preferred technical solution of the present invention, the step S5 also includes: when the cumulative value of the matching degree corresponding to the nth pathological type is the largest, and the difference between the first value and the third value output by the nA model or the nB model corresponding to the nth pathological type is greater than the preset difference, that is, one of the first value or the third value is greater than or equal to the first threshold and the other is less than or equal to the second threshold, repeating the method of step S4 to retrain the nA model and the nB model corresponding to the nth pathological type, wherein the first threshold is greater than the second threshold.
[0020] As a preferred technical solution of the present invention, when the cumulative value of the matching degree corresponding to the nth pathological type is the largest, and the first value and the third value corresponding to the nth pathological type are both less than the second threshold, the method of step S2-step S3 is repeated to reacquire multiple target sub-images.
[0021] As a preferred technical solution of the present invention, the training images include images of all lung cancer pathological types and normal images.
[0022] The present invention also provides a lung cancer image classification system based on multi-source heterogeneous information fusion, the system is used to implement the above method, the system comprises:
[0023] A collecting unit, used for acquiring a plurality of multi-source heterogeneous images of known pathological types from a storage unit, and acquiring an original image based on the multi-source heterogeneous images;
[0024] The segmentation unit is configured to: enlarge the original image to a preset size, and use the enlarged original image as a first image; when the area of the first image is greater than or equal to a set area, the first image is segmented into a plurality of second images by the segmentation unit, and the first image whose area is smaller than the set area is used as a third image, and the backgrounds of the second image and the third image are removed and enlarged to a plurality of preset multiples to obtain a training image;
[0025] A model creation unit, configured to generate a first model, train the first model using the training image, magnify the patient's target image to a plurality of the preset multiples and input the magnified image into the first model to obtain a plurality of target sub-images;
[0026] The calculation unit is configured to: obtain the training image corresponding to each pathological type according to the pathological type corresponding to the original image, and compare the training images of any two pathological types to obtain the two pathological types with the highest similarity to the i-th pathological type;
[0027] The model creation unit is further used to create an iA model and an iB model corresponding to the i-th pathological type, and train the iA model and the iB model based on the i-th pathological type and the training images corresponding to the two pathological types;
[0028] The classification unit is configured to: input the multiple target sub-images into the iA model and the iB model corresponding to the i-th pathological type respectively, and judge the pathological type to which the target image belongs based on the sum of the matching degrees between the target image and the i-th pathological type respectively output by the iA model and the iB model.
[0029] The present invention also provides a computer storage medium, wherein the storage medium stores program instructions, wherein when the program instructions are executed, the device where the storage medium is located is controlled to execute the above method.
[0030] Compared with the prior art, the beneficial effects of the present invention are at least as follows:
[0031] The present invention obtains a plurality of multi-source heterogeneous images of known pathological types from a storage unit, and obtains more types of images through partial exchange by using the types of the multi-source heterogeneous images whose number is less than a set number, thereby enriching the sample images, amplifying the original image, segmenting the image whose area is larger than a set area, and using the segmented image and the unsegmented image as the third image, and amplifying the third image and the second image to a plurality of preset multiples to obtain a comprehensive and rich training image, and also obtains two pathological types with the highest similarity to the i-th pathological type by comparing the training images of any two of the pathological types, and creates the iA model and the iB model corresponding to the i-th pathological type through the model creation unit, and trains the iA model and the iB model to improve the recognition accuracy of the i-th pathological type, and respectively inputs the plurality of the target sub-images into the iA model and the iB model corresponding to each of the pathological types, and judges the pathological type to which the target image belongs according to the sum of the matching degrees of the target image and the i-th pathological type respectively output by the iA model and the iB model, thereby improving the classification accuracy of the target image. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 The present invention is a flow chart of a lung cancer image classification method based on multi-source heterogeneous information fusion;
[0033] Figure 2 This is a structural diagram of the lung cancer image classification system based on multi-source heterogeneous information fusion of the present invention. DETAILED DESCRIPTION
[0034] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0035] The present invention provides a lung cancer image classification method based on multi-source heterogeneous information fusion, such as Figure 1 As shown, the method comprises the following steps:
[0036] Step S1: acquiring a plurality of multi-source heterogeneous images of known pathological types from a storage unit through a collection unit, and acquiring an original image based on the multi-source heterogeneous images;
[0037] Specifically, a plurality of multi-source heterogeneous images of known pathological types are read from the storage unit through the collection unit, the multi-source heterogeneous images including CT images, X-ray images and examination result images, etc., wherein the multi-source heterogeneous images correspond to multi-source heterogeneous information describing the images, and according to the sources of the multi-source heterogeneous images, i.e., whether the multi-source heterogeneous images are from CT images or X-ray images, etc., in order to obtain richer original images, when the number of images from one source is less than a set number, the effective parts of any two multi-source heterogeneous images of the same pathological type in the sources are partially exchanged to obtain new original images, wherein the positions of the identified lesions in the multi-source heterogeneous images are effective positions. Through the technical solution, a foundation is laid for further obtaining richer training images and improving the accuracy of the first model, the iA model and the iB model.
[0038] Step S2: Enlarging the original image to a preset size, and using the enlarged original image as a first image; when the area of the first image is greater than or equal to a set area, dividing the first image into a plurality of second images by a segmentation unit, using the first image with an area smaller than the set area as a third image, and removing the backgrounds of the second image and the third image and enlarging them to a plurality of preset multiples to obtain training images;
[0039] Specifically, by enlarging the above-mentioned original image, a clearer above-mentioned first image is obtained. When the above-mentioned first image is more complex, the corresponding area is also relatively large. In order to more accurately identify the pathological features in the above-mentioned first image, when the area of the above-mentioned first image is larger than the above-mentioned set area, the above-mentioned first image is divided into multiple above-mentioned second images, and the background of the above-mentioned second image and the unsegmented above-mentioned first image is removed to obtain the background-removed image in which only the corresponding lung has the pathological features, and the above-mentioned background-removed image is enlarged to multiple above-mentioned preset multiples to obtain the above-mentioned training image, for example: 2 times, 4 times and 6 times, etc. Through the mutual cooperation of the above-mentioned step S1 and the above-mentioned step S2, the above-mentioned rich and comprehensive training images are obtained, which lays a foundation for further improving the accuracy of the first model, the iA model and the iB model.
[0040] Step S3: Generate a first model through a model creation unit, and use the training image to train the first model, magnify the patient's target image to a plurality of the preset multiples and input it into the first model to obtain a plurality of target sub-images;
[0041] Specifically, the above-mentioned first model is created by the above-mentioned model creation unit, and the above-mentioned first model is trained based on the above-mentioned training image through the neural network convolution algorithm. The segmentation method is learned through the above-mentioned training image. In order to obtain a better segmentation effect, the patient's target image is enlarged to multiple of the above-mentioned preset multiples to obtain multiple of the above-mentioned enlarged target images, and the multiple enlarged target images are input into the above-mentioned first model to obtain multiple of the above-mentioned target sub-images. Through the above-mentioned technical solution, the above-mentioned target image is accurately segmented into multiple of the above-mentioned target sub-images, so as to further improve the pathological classification accuracy of the above-mentioned target image.
[0042] Step S4: acquiring the training image corresponding to each pathological type according to the pathological type corresponding to the original image, and comparing the training images of any two pathological types to acquire two pathological types with the highest similarity to the i-th pathological type, and creating an iA model and an iB model corresponding to the i-th pathological type by the model creation unit, and training the iA model and the iB model based on the i-th pathological type and the training images corresponding to the two pathological types;
[0043] Specifically, by comparing the above-mentioned training images corresponding to any two of the above-mentioned pathological types, the similarity of the above-mentioned training images corresponding to the two pathological types is obtained, and the two pathological types with the highest similarity to the i-th pathological type training image, namely the j-th pathological type and the k-th pathological type, are obtained, and the above-mentioned iA model and the above-mentioned iB model are respectively created by the above-mentioned model creation unit. Since the training image of the i-th pathological type is most similar to the training images of the j-th pathological type and the k-th pathological type, the above-mentioned iA model is trained by the above-mentioned i-th pathological type training image and the above-mentioned j-th pathological type training image, and the above-mentioned i-th pathological type training image is trained by the above-mentioned i-th pathological type training image. The iB model is trained by using the iA model and the iB model, and the iA model and the iB model are respectively made to learn the difference between the i-th pathology type training image and the j-th pathology type image, and the difference between the i-th pathology type training image and the k-th pathology type image. Through the above technical solution, when the target sub-image is input into the iA model and the iB model, the matching degree between the target image corresponding to the target sub-image and the i-th pathology type can be obtained more accurately, thereby laying a foundation for obtaining the accurate classification of the above-mentioned target image.
[0044] Step S5: Input the multiple target sub-images into the iA model and the iB model corresponding to the i-th pathological type respectively, and judge the pathological type to which the target image belongs according to the sum of the matching degrees between the target image and the i-th pathological type respectively output by the iA model and the iB model.
[0045] Specifically, by calculating the sum of the matching degrees of the multiple target sub-images output by the iA model and the iB model and the i-th pathological type, wherein the larger the sum of the matching degrees is, the greater the probability that the multiple target sub-images are the i-th pathological type, therefore, by calculating the sum of the matching degrees of the multiple target sub-images and any pathological type, and taking the i-th pathological type with the largest sum of the matching degrees and the matching degrees output by the iA model and the iB model both greater than the first threshold as the pathological type of the target image, that is, the matching degrees of the multiple target images output by the iA model and the iB model and the i-th pathological type are both high, and the i-th pathological type is also the one with the highest matching degree with the multiple target sub-images among all the pathological types, therefore, by screening the above two conditions, it can be determined that the pathological type of the target image is the i-th pathological type, and through the above technical solution, the pathological type to which the target image belongs can be obtained more accurately.
[0046] Furthermore, in step S1, the pathological information contained in the multi-source heterogeneous image is multi-source heterogeneous information. The multi-source heterogeneous image obtains the number of images from each source based on images including multiple different sources. When the number of images is less than a set number, the effective positions of any two images of the source are partially exchanged, and the newly acquired multi-source heterogeneous image and the multi-source heterogeneous image whose number of images is greater than or equal to the set number are used as the original images.
[0047] Specifically, in order to obtain richer original images, when the number of images from one source is less than a set number, the effective parts of any two of the multi-source heterogeneous images belonging to the same pathological type in the above sources are partially exchanged to obtain the new original images, wherein the positions of the identified lesions in the multi-source heterogeneous images are effective positions. Through the above technical solution, a foundation is laid for further obtaining richer training images and improving the accuracy of the first model, the iA model and the iB model.
[0048] Furthermore, the step S3 includes: creating the first model through the model creation unit, and using the training image to train the first model through a neural network convolution algorithm, and magnifying the target image of the patient to multiple preset multiples, and using the enlarged image as an input image, inputting multiple input images into the first model and obtaining multiple groups of output images, wherein the output image is an image after all the input images are segmented, and obtaining the target sub-image based on the clustering result of the output image, wherein the target sub-image includes the entire content of the target image.
[0049] Specifically, the target image of the patient is enlarged to multiple of the preset multiples, and the image after the method is input as the input image into the first model, so that a more suitable segmentation method can be found for any part of the target image. The multiple groups of output images of the first model are also used. Since the input images are repeated, multiple segmented images may be output for the same part of the target image. Therefore, through the clustering results, the image with the highest clarity and the most complete edge is selected from the same image after segmentation as the target sub-image of each part, wherein the target sub-image includes all images in the target image that are useful for determining the type of lung cancer. Through the above technical solution, the best target sub-image after segmentation of the target image can be obtained, thereby laying the foundation for obtaining accurate classification of the target image.
[0050] Further, step S4 comprises the following steps:
[0051] Step S41: acquiring the training image corresponding to each pathological type according to the pathological type corresponding to the original image, and comparing any two training images of the pathological type to obtain a comparison result;
[0052] Step S42: according to the comparison result, obtaining the jth pathological type and the kth pathological type having the highest similarity to the ith pathological type;
[0053] Specifically, the above-mentioned original images include all types of lung cancer and lung images in normal conditions. In order to obtain accurate classification of the above-mentioned target images, it is necessary to be able to accurately identify even when the target images corresponding to the two pathological types have a high degree of similarity. Therefore, through the above-mentioned technical solution, the two pathological types with the highest similarity to the training images corresponding to each of the above-mentioned pathological types are obtained, laying the foundation for further learning the differences between the images of the above-mentioned pathological types and the two pathological types with the highest similarity through model training.
[0054] Step S43: Create the iA model and the iB model through the model creation unit, and train the iA model through the first learning image corresponding to the i-th pathological type and the j-th pathological type in the training image, and train the iB model through the second learning image corresponding to the i-th pathological type and the k-th pathological type in the training image, wherein the value range of i is a positive integer greater than or equal to 1 and less than or equal to N, and N is the number of all pathological types corresponding to the original image.
[0055] Specifically, the training images include images corresponding to all pathological types of lung cancer, and the pathological type corresponding to each of the training images is known. The iA model is trained by the training images corresponding to the i-th pathological type and the training images corresponding to the j-th pathological type, and the iB model is trained by the training images corresponding to the i-th pathological type and the training images corresponding to the k-th pathological type, so that the iA model and the iB model learn the difference between the i-th pathological type training image and the j-th pathological type image and the difference between the i-th pathological type training image and the k-th pathological type image, respectively. Through the technical solution, when the target sub-image is input into the iA model and the iB model, a more accurate matching degree between the target image corresponding to the target sub-image and the i-th pathological type can be obtained, thereby laying a foundation for obtaining accurate classification of the target image.
[0056] Furthermore, step S5 comprises the following steps:
[0057] Step S51: inputting a plurality of the target sub-images into the iA model and the iB model respectively, and obtaining a first value and a second value output by the iA model, and a third value and a fourth value output by the iB model, wherein the first value is the matching degree between the target image output by the iA model and the i-th pathological type, and the second value is the matching degree between the target image output by the iA model and the j-th pathological type; the third value is the matching degree between the target image output by the iB model and the i-th pathological type, and the fourth value is the matching degree between the target image output by the iB model and the k-th pathological type;
[0058] Specifically, the multiple target sub-images are images after segmentation of the target image and the multiple target sub-images include all useful information related to the pathological type of the target image. By inputting the multiple target sub-images into the iA model and the iB model corresponding to each i-th pathological type, the matching degrees of the multiple target sub-images output by the iA model with the i-th pathological type and the j-th pathological type and the matching degrees of the multiple target sub-images output by the iB model with the i-th pathological type and the k-th pathological type are obtained respectively, thereby providing a data basis for further judging the pathological type of the target image by the sum of the matching degrees of the two i-th pathological types output by the two models.
[0059] Step S52: calculating the sum of the first value and the third value as the i-th matching degree cumulative value, and when the m-th matching degree cumulative value is the largest among all the i-th matching degree cumulative values, and the first value and the second value corresponding to the m-th matching degree cumulative value are both greater than a first threshold, taking the m-th pathological type corresponding to the m-th matching degree cumulative value as the pathological type corresponding to the target image;
[0060] The value range of j, k, and m is a positive integer greater than or equal to 1 and less than or equal to N, where N is the number of all pathological types corresponding to the original image.
[0061] Specifically, the i-th matching degree cumulative value is obtained by calculating the sum of the matching degrees of the multiple target sub-images output by the i-th model and the i-th model and the i-th pathological type, that is, the sum of the first value and the second and third values. The larger the i-th matching degree cumulative value is, the greater the probability that the multiple target sub-images are the i-th pathological type. Therefore, by calculating the matching degree cumulative values of the multiple target sub-images and any pathological type, and taking the pathological type with the largest matching degree cumulative value and the corresponding first value and third value both greater than the first threshold as the pathological type of the target image, that is, the matching degrees of the multiple target images output by the i-th model and the i-th model and the i-th pathological type are relatively high, and the i-th pathological type is also the one with the highest matching degree with the multiple target sub-images among all the pathological types. Therefore, by screening the above two conditions, it can be determined that the pathological type of the target image is the i-th pathological type. Through the above technical solution, the pathological type to which the target image belongs can be obtained more accurately.
[0062] Furthermore, the step S5 also includes: when the cumulative value of the matching degree corresponding to the nth pathological type is the largest, and the difference between the first value and the third value output by the nA model or the nB model corresponding to the nth pathological type is greater than a preset difference, that is, one of the first value or the third value is greater than or equal to the first threshold and the other is less than or equal to the second threshold, repeating the method of step S4 to retrain the nA model and the nB model corresponding to the nth pathological type, wherein the first threshold is greater than the second threshold.
[0063] Specifically, when the cumulative value of the matching degree corresponding to the nth pathological type is the largest, when there is a difference in the matching degrees between the multiple target sub-images respectively output by the nth A model and the nth B model and the nth pathological type, since they are output by two different models, it is possible that the difference in the matching degrees between the outputs of the two models and the nth pathological type is normal. However, since both models output the matching degrees between the multiple target sub-images and the nth pathological type, when there is a large difference in the matching degrees between the multiple target sub-images output by the two models and the nth pathological type, for example, the first numerical value is 90%, the third numerical value is 20%, the first threshold is 80%, and the second threshold is 50%. The value is 40%. At this time, there must be one that is inaccurate. Since the above training images include all lung cancer pathology images and normal images, the multiple target sub-images can always match the corresponding pathology type, that is, the m-th pathology type. When the matching degree of the multiple target sub-images output by the above two models and the n-th pathology type is greatly different, at least one of the nA model and the nB model is not accurate enough. Therefore, the nA model and the nB model are retrained by the method of step S4, so as to improve the progress of the nA model and the nB model, and the accurate classification of the target image is obtained again by step S5.
[0064] Furthermore, when the cumulative value of the matching degree corresponding to the nth pathological type is the largest, and the first value and the third value corresponding to the nth pathological type are both less than the second threshold, the method of step S2 to step S3 is repeated to reacquire multiple target sub-images.
[0065] Specifically, when the cumulative value of the matching degree of the above-mentioned nth pathological type is the largest, it means that the above-mentioned target image is closest to the training image corresponding to the above-mentioned nth pathological type. However, since the matching degrees with the above-mentioned nth pathological type output by the above-mentioned nA model and the above-mentioned nB model are both less than the above-mentioned second threshold, that is, the matching degrees with the nth pathological type output by the above-mentioned two models are both small, therefore, it may be caused by the inaccuracy in segmenting the above-mentioned target image into the above-mentioned multiple target sub-images. Therefore, repeat steps S2-step S3 to re-acquire the above-mentioned multiple target sub-images, and re-acquire the accurate classification of the above-mentioned target image through step S5.
[0066] Furthermore, the training images include all lung cancer pathology type images and normal images.
[0067] The present invention also provides a lung cancer image classification system based on multi-source heterogeneous information fusion, the system is used to implement the above method, such as Figure 2 As shown, the system comprises:
[0068] A collecting unit, used for acquiring a plurality of multi-source heterogeneous images of known pathological types from a storage unit, and acquiring an original image based on the multi-source heterogeneous images;
[0069] The segmentation unit is configured to: enlarge the original image to a preset size, and use the enlarged original image as a first image; when the area of the first image is greater than or equal to a set area, the first image is segmented into a plurality of second images by the segmentation unit, and the first image whose area is smaller than the set area is used as a third image, and the backgrounds of the second image and the third image are removed and enlarged to a plurality of preset multiples to obtain a training image;
[0070] A model creation unit, configured to generate a first model, train the first model using the training image, magnify the patient's target image to a plurality of the preset multiples and input the magnified image into the first model to obtain a plurality of target sub-images;
[0071] The calculation unit is configured to: obtain the training image corresponding to each pathological type according to the pathological type corresponding to the original image, and compare the training images of any two pathological types to obtain the two pathological types with the highest similarity to the i-th pathological type;
[0072] The model creation unit is further used to create an iA model and an iB model corresponding to the i-th pathological type, and train the iA model and the iB model based on the i-th pathological type and the training images corresponding to the two pathological types;
[0073] The classification unit is configured to: input the multiple target sub-images into the iA model and the iB model corresponding to the i-th pathological type respectively, and judge the pathological type to which the target image belongs based on the sum of the matching degrees between the target image and the i-th pathological type respectively output by the iA model and the iB model.
[0074] The present invention also provides a computer storage medium, wherein the storage medium stores program instructions, wherein when the program instructions are executed, the device where the storage medium is located is controlled to execute the above method.
[0075] In summary, the present invention obtains a plurality of multi-source heterogeneous images of known pathological types from a storage unit, and obtains more types of images through partial exchange by using the types of the multi-source heterogeneous images whose number is less than the set number, and enriches the sample images, and also amplifies the original image, and segments the image whose area is larger than the set area, and uses the segmented image and the unsegmented image as the third image, and amplifies the third image and the second image to multiple preset multiples to obtain comprehensive and rich training images, and also obtains two pathological types with the highest similarity to the i-th pathological type by comparing the training images of any two of the pathological types, and creates the iA model and the iB model corresponding to the i-th pathological type through the model creation unit, and trains the iA model and the iB model to improve the recognition accuracy of the i-th pathological type, and inputs the plurality of the target sub-images into the iA model and the iB model corresponding to each of the pathological types, and judges the pathological type to which the target image belongs according to the sum of the matching degrees of the target image and the i-th pathological type respectively output by the iA model and the iB model, thereby improving the classification accuracy of the target image.
[0076] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0077] The above-mentioned embodiments only express several implementation methods of the present invention, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.
[0078] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A lung cancer image classification method based on multi-source heterogeneous information fusion, characterized in that: The method comprises the following steps: Step S1: acquiring a plurality of multi-source heterogeneous images of known pathological types from a storage unit through a collection unit, and acquiring an original image based on the multi-source heterogeneous images; Step S2: Enlarging the original image to a preset size, and using the enlarged original image as a first image; when the area of the first image is greater than or equal to a set area, dividing the first image into a plurality of second images through a segmentation unit, using the first image whose area is smaller than the set area as a third image, and removing the backgrounds of the second image and the third image and enlarging them to a plurality of preset multiples to obtain training images; Step S3: Generate a first model by a model creation unit, and use the training image to train the first model, magnify the patient's target image to a plurality of the preset multiples and input the magnified image into the first model to obtain a plurality of target sub-images; Step S4: acquiring the training image corresponding to each pathological type according to the pathological type corresponding to the original image, and comparing the training images of any two pathological types to acquire two pathological types with the highest similarity to the i-th pathological type, and creating an iA model and an iB model corresponding to the i-th pathological type by the model creation unit, and training the iA model and the iB model based on the i-th pathological type and the training images corresponding to the two pathological types; Step S5: inputting the plurality of target sub-images into the ith model and the ith model corresponding to each pathological type respectively, and judging the pathological type to which the target image belongs according to the matching degree output by the ith model and the ith model; The step S4 comprises the following steps: Step S41: acquiring the training image corresponding to each pathological type according to the pathological type corresponding to the original image, and comparing any two training images of the pathological type to obtain a comparison result; Step S42: according to the comparison result, obtaining the jth pathological type and the kth pathological type having the highest similarity to the ith pathological type; Step S43: creating an iA model and an iB model through the model creation unit, and training the iA model through a first learning image corresponding to the i-th pathological type and the j-th pathological type in the training image, and training the iB model through a second learning image corresponding to the i-th pathological type and the k-th pathological type in the training image, wherein the value range of i is a positive integer greater than or equal to 1 and less than or equal to N, and N is the number of all pathological types corresponding to the original image; The step S5 comprises the following steps: Step S51: inputting a plurality of the target sub-images into the iA model and the iB model respectively, and obtaining a first value and a second value output by the iA model, and a third value and a fourth value output by the iB model, wherein the first value is the matching degree between the target image output by the iA model and the i-th pathological type, and the second value is the matching degree between the target image output by the iA model and the j-th pathological type; the third value is the matching degree between the target image output by the iB model and the i-th pathological type, and the fourth value is the matching degree between the target image output by the iB model and the k-th pathological type; Step S52: calculating the sum of the first value and the third value as the i-th matching degree cumulative value, and when the m-th matching degree cumulative value is the largest among all the i-th matching degree cumulative values, and the first value and the second value corresponding to the m-th matching degree cumulative value are both greater than a first threshold, taking the m-th pathological type corresponding to the m-th matching degree cumulative value as the pathological type corresponding to the target image; The value range of j, k, and m is a positive integer greater than or equal to 1 and less than or equal to N, where N is the number of all pathological types corresponding to the original image.
2. The method according to claim 1, characterized in that In step S1, the pathological information contained in the multi-source heterogeneous image is multi-source heterogeneous information. The multi-source heterogeneous image obtains the number of images from each source based on images including multiple different sources. When the number of images is less than a set number, the effective positions of any two multi-source heterogeneous images with the same pathological type in the images of the sources are partially exchanged to obtain a new original image.
3. The method according to claim 1, characterized in that The step S3 includes: creating the first model through the model creation unit, and using the training image to train the first model through a neural network convolution algorithm, and enlarging the target image of the patient to multiple preset multiples, and using the enlarged image as an input image, inputting multiple input images into the first model and obtaining multiple groups of output images, wherein the output image is an image after all the input images are segmented, and obtaining the target sub-image according to the clustering result of the output image, wherein the target sub-image includes the entire content of the target image.
4. The method according to claim 1, characterized in that: The step S5 further includes: when the accumulated value of the matching degree corresponding to the nth pathological type is the largest, and the difference between the first value and the third value output by the nth A model or the nth B model corresponding to the nth pathological type is greater than a preset difference, and and When one of the first value or the third value is greater than or equal to the first threshold and the other is less than or equal to the second threshold, repeat the method of step S4 to retrain the nA model and the nB model corresponding to the n pathological type, wherein the first threshold is greater than the second threshold.
5. The method according to claim 4, characterized in that When the cumulative value of the matching degree corresponding to the nth pathological type is the largest, and the first value and the third value corresponding to the nth pathological type are both smaller than the second threshold, the method of step S2 to step S3 is repeated to reacquire a plurality of the target sub-images.
6. The method according to claim 1, characterized in that The training images include images of all lung cancer pathological types and normal images.
7. A lung cancer image classification system based on multi-source heterogeneous information fusion, the system is used to implement the method according to any one of claims 1-6, characterized in that: The system comprises: A collecting unit, used for acquiring a plurality of multi-source heterogeneous images of known pathological types from a storage unit, and acquiring an original image based on the multi-source heterogeneous images; The segmentation unit is configured to: enlarge the original image to a preset size, and use the enlarged original image as a first image; when the area of the first image is greater than or equal to a set area, the first image is segmented into a plurality of second images by the segmentation unit, and the first image whose area is smaller than the set area is used as a third image, and the backgrounds of the second image and the third image are removed and enlarged to a plurality of preset multiples to obtain a training image; A model creation unit, configured to generate a first model, train the first model using the training image, magnify the patient's target image to a plurality of the preset multiples and input the magnified image into the first model to obtain a plurality of target sub-images; The calculation unit is configured to: obtain the training image corresponding to each pathological type according to the pathological type corresponding to the original image, and compare the training images of any two pathological types to obtain the two pathological types with the highest similarity to the i-th pathological type; The model creation unit is further used to create an iA model and an iB model corresponding to the i-th pathological type, and train the iA model and the iB model based on the i-th pathological type and the training images corresponding to the two pathological types; The classification unit is configured to: input the plurality of target sub-images into the iAth model and the iBth model, and determine the pathological type to which the target image belongs according to the matching degree output by the iAth model and the iBth model.
8. A computer storage medium, characterized in that: The storage medium stores program instructions, wherein when the program instructions are executed, the device where the storage medium is located is controlled to execute the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Generative Adversarial Network Alternating Training and Medical Image Classification Methods, Devices, and Media
CN114267432B
Classification of medical images using machine learning to account for body orientation
US20230146953A1
Method and device for predicting stage of diabetes and computer device
CN110197724A
Auxiliary diagnosis method, device, equipment and system and storage medium
CN116030958A
Camera
JP2021100186A