Pathological image classification methods, devices, equipment, storage media and program products

By constructing a classification model based on multiple pathological image databases, the problem of poor correlation between pathological image sets was solved, achieving centralized and standardized classification of pathological images and simplifying the management and use of pathological images.

CN114332526BActive Publication Date: 2026-04-03TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-11
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

In existing technologies, the correlation between pathological image sets is poor, which makes it impossible for a pathological image to be classified to obtain the correct classification label in multiple pathological image sets at the same time.

Method used

By acquiring sample pathological images from multiple pathological image databases, matching sample labels with the label sets of the pathological image databases, constructing a classification model based on multiple pathological image databases, and training the model using sample prediction probabilities and loss values, a pathological image classification model is obtained, which is used to simultaneously output prediction results from multiple pathological image databases.

Benefits of technology

It enhances the correlation between pathological image databases, enabling better centralized, professional, and standardized classification of pathological images, and simplifies the management and use of pathological images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114332526B_ABST
    Figure CN114332526B_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, device, storage medium, and program product for classifying pathological images, relating to the field of machine learning. The method includes: acquiring sample pathological images from a pathological image database, the sample pathological images being labeled with sample tags; matching the sample tags with corresponding pathological image tag sets in the pathological image database to determine the sample pathological image database; inputting the sample pathological images into a classification model to obtain sample prediction probabilities; determining a loss value based on the sample prediction probabilities and sample tags corresponding to the sample pathological image database, training the classification model to obtain a pathological image classification model, which is used to classify target pathological images to obtain prediction results. Through this approach, the correlation between pathological image databases can be strengthened, achieving better centralized pathological image classification and facilitating better management and use of pathological images. This application can be applied to various scenarios such as cloud technology, artificial intelligence, and intelligent transportation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of machine learning, and in particular to a method, apparatus, device, storage medium, and program product for classifying pathological images. Background Technology

[0002] Pathological image classification is the process of categorizing pathological images into different levels based on their attributes or characteristics. Identifying the value of pathological images at different levels, and thus correctly analyzing them, is fundamental to obtaining accurate medical analysis results.

[0003] In related technologies, pathological images are usually classified based on specific classification labels corresponding to a specific set of pathological images. Each set of pathological images is often trained with a corresponding model. By inputting the pathological image to be classified into the model corresponding to the specific set of pathological images, the classification label of the pathological image to be classified in the specific set of pathological images is determined.

[0004] However, the model trained using the above method is independent for each pathological image set, and the correlation between different pathological image sets is poor. When a pathological image to be classified simultaneously matches the classification labels of multiple pathological image sets, it is impossible to obtain the classification labels of the pathological image to be classified in multiple pathological image sets at the same time. Summary of the Invention

[0005] This application provides a method, apparatus, device, storage medium, and program product for classifying pathological images, which can enhance the correlation between pathological image databases and better achieve centralized pathological image classification. The technical solution is as follows.

[0006] On the one hand, a pathological image classification method is provided, the method comprising:

[0007] Sample pathological images are obtained from at least two pathological image libraries, wherein the sample pathological images are labeled with sample tags, and the at least two pathological image libraries each correspond to a pathological image tag set;

[0008] The sample label is matched with the pathological image label sets corresponding to the at least two pathological image libraries respectively, and the sample pathological image library that matches the sample label is determined from the at least two pathological image libraries;

[0009] The sample pathological images are input into the classification model to obtain the sample prediction probability corresponding to the sample pathological image database. The classification model is a training model constructed based on the at least two pathological image databases.

[0010] Based on the sample prediction probability corresponding to the sample pathological image library and the sample label, the loss value is determined, and the classification model is trained to obtain a pathological image classification model. The pathological image classification model is used to classify the target pathological image and obtain the prediction results corresponding to the at least two pathological image libraries respectively.

[0011] On the other hand, a pathological image classification device is provided, the device comprising:

[0012] The acquisition module is used to acquire sample pathological images from at least two pathological image libraries, wherein the sample pathological images are labeled with sample tags, and the at least two pathological image libraries each correspond to a pathological image tag set;

[0013] The matching module is used to match the sample label with the pathological image label sets corresponding to the at least two pathological image libraries respectively, and to determine the sample pathological image library that matches the sample label from the at least two pathological image libraries;

[0014] A model is determined to input the sample pathological images into the classification model to obtain the sample prediction probability corresponding to the sample pathological image library. The classification model is a training model constructed based on the at least two pathological image libraries.

[0015] The training module is used to determine the loss value based on the sample prediction probability corresponding to the sample pathological image library and the sample label, and to train the classification model to obtain a pathological image classification model. The pathological image classification model is used to classify the target pathological image and obtain the prediction results corresponding to the at least two pathological image libraries respectively.

[0016] On the other hand, a computer device is provided, the computer device including a processor and a memory, the memory storing at least one instruction, at least one program, code set or instruction set, the at least one instruction, the at least one program, the code set or instruction set being loaded and executed by the processor to implement any of the pathological image classification methods described in the embodiments of this application above.

[0017] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction, at least one program, code set, or instruction set is stored therein, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the pathological image classification method as described in any of the embodiments of this application above.

[0018] On the other hand, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform any of the pathological image classification methods described in the above embodiments.

[0019] The beneficial effects of the technical solutions provided in this application include at least the following:

[0020] Sample pathological images are obtained from a pathological image database. The sample labels corresponding to the sample pathological images are matched with the corresponding pathological image label sets in the database to determine the database to which the sample pathological image belongs. Then, the sample pathological image is input into a classification model built based on the aforementioned database. The output is the predicted probability of the sample corresponding to the database. Based on the sample labels and the predicted probabilities, a loss value is determined, thus training the classification model and obtaining a pathological image classification model. In application, the target pathological image to be classified is input into the classification model, and multiple pathological image labels related to the database are output as the prediction result corresponding to the target pathological image, achieving classification of the target pathological image. This method avoids the problem of needing to input the target pathological image into multiple classification models corresponding to multiple pathological image databases to obtain prediction results. The pathological image classification model in this application strengthens the correlation between pathological image databases, better achieving centralized, professional, and standardized pathological image classification, and facilitating better management and use of pathological images. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is a schematic diagram of an implementation environment provided by an exemplary embodiment of this application;

[0023] Figure 2 This is a flowchart of a pathological image classification method provided in an exemplary embodiment of this application;

[0024] Figure 3 This is a flowchart of a pathological image classification method provided in another exemplary embodiment of this application;

[0025] Figure 4This is a flowchart of a pathological image classification method provided in another exemplary embodiment of this application;

[0026] Figure 5 This is a schematic diagram of a hematoxylin-eosin staining method provided in an exemplary embodiment of this application;

[0027] Figure 6 This is a flowchart of a pathological image classification method provided in another exemplary embodiment of this application;

[0028] Figure 7 This is a schematic diagram illustrating the processing of nuclear pathology images provided in an exemplary embodiment of this application;

[0029] Figure 8 This is a schematic diagram illustrating the application of a pathological image classification method to obtain cell nucleus classification labels, provided in an exemplary embodiment of this application.

[0030] Figure 9 This is a structural block diagram of a pathological image classification device provided in an exemplary embodiment of this application;

[0031] Figure 10 This is a structural block diagram of a pathological image classification device provided in another exemplary embodiment of this application;

[0032] Figure 11 This is a structural block diagram of a server provided in an exemplary embodiment of this application. Detailed Implementation

[0033] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0034] First, a brief introduction to the terms used in the embodiments of this application will be given.

[0035] Artificial Intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.

[0036] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0037] Machine Learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instruction-based learning.

[0038] In related technologies, pathological images are typically classified based on specific classification labels corresponding to specific pathological image sets. Each pathological image set usually corresponds to a trained model. By inputting the pathological image to be classified into the model corresponding to the specific pathological image set, the classification label of the image to be classified in that specific pathological image set is determined. However, with the above method, the models trained for each pathological image set are independent, and the correlation between different pathological image sets is poor. When the pathological image to be classified simultaneously matches the classification labels of multiple pathological image sets, it is impossible to utilize the correlation between multiple pathological image sets to simultaneously obtain the classification label of the image to be classified in multiple pathological image sets.

[0039] This application provides a pathological image classification method that considers the correlation between multiple pathological image databases. It analyzes a pathological image to be classified using a pathological image classification model, simultaneously obtaining the classification labels of the image in multiple pathological image databases. Optionally, the application of the pathological image classification method trained in this application to the medical field will be used as an example for illustration.

[0040] In the medical field, pathological image databases are generally difficult to annotate, resulting in a relatively small volume of pathological images stored in these databases. Each database corresponds to a training classification model. When multi-label classification of a pathological image is required, the image is typically input into multiple classification models corresponding to different databases to obtain the corresponding medical pathological image labels. Illustratively, the pathological image classification method provided in this application constructs a classification model using multiple medical pathological image databases. Medical pathological images from these databases are then used as sample images to train the model, resulting in a trained classification model. In application, inputting the pathological image to be classified into this model simultaneously yields multiple medical labels, making the process of obtaining medical pathological image labels more convenient.

[0041] Secondly, the implementation environment involved in the embodiments of this application will be described, for illustrative purposes only. Please refer to [the relevant documentation]. Figure 1 The implementation environment involves a terminal 110 and a server 120, which are connected via a communication network 130.

[0042] In some embodiments, the terminal 110 is equipped with an application that has pathological image acquisition capabilities. In some embodiments, the terminal 110 is used to send a target pathological image to the server 120. The server 120 can predict the probability through a pathological image classification model, classify the target pathological image according to the predicted probability, output the prediction result, and feed the prediction result back to the terminal 110 for display.

[0043] The pathological image classification model is trained using this method by acquiring sample pathological images from a pathological image database. Illustratively, after acquiring sample pathological images from n pathological image databases, the sample labels corresponding to the sample pathological images are matched with the n pathological image label sets corresponding to the n pathological image databases. This determines the pathological image database to which the sample pathological image belongs. The sample pathological image is then input into the classification model to be trained, obtaining the predicted probability of the sample corresponding to the pathological image database. Based on the predicted probability and the sample labels, a loss value is determined, and the classification model is trained using this loss value to obtain the pathological image classification model. The above process is an example of a non-unique case of the pathological image classification model training process.

[0044] It is worth noting that the aforementioned terminals include, but are not limited to, mobile terminals such as mobile phones, tablets, portable laptops, smart voice interaction devices, smart home appliances, and in-vehicle terminals, and can also be implemented as desktop computers, etc.

[0045] The aforementioned servers can be independent physical servers, server clusters or distributed systems composed of multiple physical servers, or cloud servers that provide basic cloud computing services such as cloud services, cloud pathology image libraries, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and large pathology image and artificial intelligence platforms.

[0046] Cloud technology refers to a hosting technology that unifies hardware, applications, and network resources within a wide area network (WAN) or local area network (LAN) to achieve the computation, storage, processing, and sharing of pathological images. Cloud technology is a collective term for network technology, information technology, integration technology, management platform technology, and application technology applied in the cloud computing business model. It can form resource pools, providing flexible and convenient on-demand access. Cloud computing technology will become a crucial support. Backend services of technical network systems require substantial computing and storage resources, such as video websites, image websites, and many portal websites. With the rapid development and application of the internet industry, every item may have its own identification mark in the future, requiring transmission to backend systems for logical processing. Pathological images of different levels will be processed separately, and various industry-specific pathological images will all require robust system support, which can only be achieved through cloud computing.

[0047] In some embodiments, the aforementioned server can also be implemented as a node in a blockchain system. Blockchain is a novel application model of computer technologies such as distributed pathological image storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized pathological image library, a chain of pathological image blocks linked using cryptographic methods. Each pathological image block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include a blockchain underlying platform, a platform product service layer, and an application service layer.

[0048] Based on the above-mentioned terminology and application scenarios, the pathological image classification method provided in this application will be explained, taking the application of this method to a server as an example. Figure 2 As shown, the method includes the following steps.

[0049] Step 210: Obtain sample pathological images from at least two pathological image libraries.

[0050] A pathological image library is a collection of pathological images used to store pathological images. It can store one type of pathological image (e.g., epithelial cell pathological images or cancer cell pathological images) or multiple types of pathological images (e.g., epithelial cell pathological images and cancer cell pathological images). Optionally, the pathological image library can be divided according to the types of pathological images stored within it; for example, a pathological image library storing cellular pathological images is called a cellular pathological image library. At least two pathological image libraries can be those storing the same type of pathological image, those storing different types of pathological image, or those storing multiple types of pathological image.

[0051] Optionally, obtaining sample pathological images from at least two pathological image libraries includes at least one of the following methods.

[0052] 1. Select at least one pathological image from at least two pathological image libraries as the sample pathological image.

[0053] In illustrative terms, multiple pathological images are stored in at least two pathological image libraries. At least one pathological image is randomly selected from the multiple pathological images as a sample pathological image, that is, at least one pathological image is selected from the multiple pathological images as a sample pathological image in an equal probability manner.

[0054] 2. Select pathological images from at least two pathological image libraries in a round-robin fashion as sample pathological images.

[0055] Polling involves sequentially selecting pathological images from a pathological image library, and then repeating this process. Illustratively, three pathological image libraries are pre-determined as the source of sample pathological images: Pathological Image Library 1, Pathological Image Library 2, and Pathological Image Library 3. When selecting a sample pathological image, one image is selected from Pathological Image Library 1, then from Pathological Image Library 2, and then from Pathological Image Library 3. This process is considered one selection cycle, and then the selection process is repeated: selecting another image from Pathological Image Library 1, then another from Pathological Image Library 2, and so on. The selection process stops when the number of polling iterations reaches a threshold or when no suitable pathological image is available in the selected image library. It should be noted that the number of sample pathological images selected from different image libraries is not unique, and the conditions for ending the polling are also not unique. The above is merely an illustrative example, and this embodiment does not limit the scope of the application.

[0056] Optionally, each pathological image library corresponds to a pathological image tag set. The pathological image tag set includes at least one pathological image tag, used to indicate the common features of some or all pathological images in the pathological image library. For example, pathological image library A includes pathological images A1, A2, A3, A4, and A5. The pathological image tag set corresponding to pathological image library A includes pathological image tags a1 and a2. Pathological images A1, A2, and A3 are labeled with pathological image tag a1, meaning that pathological images A1, A2, and A3 have the same features, which can be reflected by pathological image tag a1. Similarly, pathological images A4 and A5 are labeled with pathological image tag a2, meaning that pathological images A4 and A5 have the same features, which can be reflected by pathological image tag a2.

[0057] The pathological images of the samples were obtained from at least two pathological image databases and were labeled with sample tags.

[0058] Step 220: Match the sample label with the pathological image label sets corresponding to at least two pathological image libraries respectively, and determine the sample pathological image library that matches the sample label from at least two pathological image libraries.

[0059] In an optional embodiment, the sample label is matched with pathological image labels in the pathological image label sets corresponding to at least two pathological image libraries; in response to the pathological image label set corresponding to the pathological image library including the sample label, the pathological image library is determined as the sample pathological image library.

[0060] The sample pathological images are obtained from at least two pathological image databases. Therefore, there is a certain correspondence between the sample labels corresponding to the sample pathological images and the at least two pathological image label sets corresponding to the at least two pathological image databases. Optionally, the sample labels are matched one by one with each pathological image label in the pathological image label sets corresponding to the at least two pathological image databases to obtain label matching results. The label matching results include at least one of the following cases.

[0061] 1. The tag matching result is "tag matching successful".

[0062] In a schematic manner, after matching a sample label with a pathological image label in the pathological image label set, a successful label match is considered achieved when the sample label and the pathological image label are identical. For example, if the sample label is "lymphocyte" and the pathological image label is also "lymphocyte," the condition for a successful label match is met. Alternatively, a successful label match is considered achieved when the similarity between the sample label and the pathological image label exceeds a preset similarity threshold. For example, if the preset similarity threshold is 0.8, and the sample label and the pathological image label are text information, the process of determining the similarity between the sample label and the pathological image label is essentially a text similarity comparison. A successful label match is achieved when the similarity exceeds 0.8. Optionally, after matching the sample label with the pathological image label, the sample label can be further matched with other pathological image labels in the pathological image label set.

[0063] 2. The tag matching result is "tag matching failed".

[0064] To illustrate, after matching a sample label with a pathological image label from the pathological image label set, a label match is considered to have failed if the sample label does not match the pathological image label. For example, if the sample label is "epithelial cells" and the pathological image label is "lymphocytes," then the sample label does not match the pathological image label, which meets the condition for label match failure.

[0065] In an optional embodiment, a successful tag match is defined as identical tags. When a sample tag is matched with a pathological image tag set corresponding to a pathological image library, a successful tag match is considered to have occurred if the sample tag matches at least one pathological image tag in the pathological image tag set. For example, let the sample tag be m. The sample tag m is matched with pathological image tag sets corresponding to three pathological image libraries. These three pathological image tag sets are: pathological image tag set M (including pathological image tag m, pathological image tag l, and pathological image tag n); pathological image tag set L (including pathological image tag q, pathological image tag l, and pathological image tag n); and pathological image tag set N (including pathological image tag m, pathological image tag o, and pathological image tag n). If the sample tag m matches a pathological image tag m in both pathological image tag sets M and N, the sample tag is considered to have successfully matched both pathological image tag sets M and N. Therefore, the pathological image libraries corresponding to pathological image tag sets M and N are determined as the sample pathological image library. The above is merely an illustrative example, and the embodiments of this application do not limit the scope of the application.

[0066] Optionally, when the sample pathological image is a plurality of pathological images selected from at least two pathological image libraries, the sample label corresponding to the sample pathological image is matched sequentially with the pathological image label set corresponding to the at least two pathological image libraries respectively. Based on the sample label corresponding to the sample pathological image, the sample label is matched with the pathological image label in the pathological image label set. For example, when the sample label and the pathological image label are the same, it is considered that the label is successfully matched. The pathological image label set to which the pathological image label belongs is determined, and the pathological image library corresponding to the pathological image label set is determined accordingly. The pathological image library is used as the sample pathological image library.

[0067] Step 230: Input the sample pathological images into the classification model to obtain the sample prediction probability corresponding to the sample pathological image library.

[0068] The classification model is a training model built from at least two pathological image databases.

[0069] Classification is the process of categorizing existing pathological images based on their attributes or features. A classification model, also known as a classifier, is the model used to implement this classification process. Classification models can divide pathological images according to their attributes and features. Common classification models include logistic regression models, decision tree models, and multilayer perceptron models.

[0070] Optionally, there are many different types of pathological images. Pathological images of the same type are often stored in the same pathological image library. Different classification models are usually trained for different types of pathological image libraries. For example, an epithelial cell pathological image classification model is trained based on epithelial cell pathological images in an epithelial cell pathological image library; a cancer cell pathological image classification model is trained based on cancer cell pathological images in a cancer cell pathological image library, etc.

[0071] In illustrative terms, the classification model is a general classification model with certain pathological image classification capabilities. That is, the general classification domain can be used to perform basic classification of pathological images from multiple different domains or a specific domain. Different general classification models may correspond to different pathological image databases. Optionally, based on the correlation between different pathological image databases, model fusion can be performed on the general classification models corresponding to different pathological image databases to obtain a classification model applicable to multiple pathological image databases; alternatively, a framework can be provided for the general classification models corresponding to multiple different pathological image databases. This framework can utilize the relationships between multiple pathological image sets to perform multi-label classification of pathological images, and this framework can be used as a classification model applicable to multiple pathological image databases.

[0072] After determining the classification model, in order to improve the classification accuracy of the classification model, it is usually trained on the basis of the pathological image database. For example, the classification model is used as the model to be trained, and the classification model is trained on one or more pathological image databases, so that the classification model can better identify the classification category of the pathological image and improve the classification accuracy of the classification model.

[0073] In an optional embodiment, at least two pathological image libraries are pathological image libraries of the same type storing the same type of pathological images. For example, at least two pathological image libraries are pathological image libraries storing cellular pathological images, including epithelial cell pathological images, cancer cell pathological images, lymphocyte pathological images, etc. The classification model is a general classification model of the same type as the pathological image libraries, and the classification model is constructed based on at least two pathological image libraries. The sample pathological image is a pathological image randomly obtained from at least two pathological image libraries. The sample pathological image is input into the classification model, which can match the sample pathological image with the pathological images in at least two pathological image libraries, and output a pathological image result including the sample prediction probability. Here, the pathological image result is the prediction result output by the classification model, and the sample prediction probability is the prediction result corresponding to the sample pathological image and multiple sample pathological image labels output by the classification model. The sample prediction probability corresponds to the sample pathological image library and is subordinate to the pathological image result.

[0074] Step 240: Determine the loss value based on the sample prediction probability and sample label corresponding to the sample pathological image library, train the classification model, and obtain the pathological image classification model.

[0075] Optionally, when determining the loss value based on the sample prediction probability and the sample label, at least one of the following methods may be included:

[0076] The first method involves obtaining the predicted probability of samples corresponding to the sample pathological image library from the pathological image results of the classification model. The predicted probability of samples includes the predicted probability corresponding to the sample pathological image labels in the sample pathological image library. The label prediction results of the predicted probability of samples corresponding to the sample pathological image library and the sample labels are input into the loss function to obtain the loss value corresponding to the sample pathological image.

[0077] For illustration, the sample label corresponding to pathological image A is a1. The sample pathological image library has a sample pathological image label set, which includes three sample pathological image labels, a2, b, and c. When the sample pathological image A is input into the classification model, the classification model outputs the sample prediction probabilities corresponding to the three pathological image labels a2, b, and c, which are 0.7, 0.1, and 0.2, respectively. Therefore, the label prediction result of the sample pathological image A corresponding to the sample pathological image library is determined to be a2. Then, the sample label a1 and the label prediction result a2 corresponding to the sample pathological image A are input into the loss function to obtain the loss value corresponding to the sample pathological image A.

[0078] The second approach involves matching a sample label with a pathological image label in the sample pathological image database (a successful match is indicated by a match result of 1). Matching a sample label with a different label from the database results in a failed match (a match result of 0). The loss value for each pathological image is calculated based on the matching results of the pathological images in the database and the predicted probabilities of the corresponding images in the classification model's pathological image results. The predicted probabilities include the predicted probabilities corresponding to the pathological image labels in the database.

[0079] To illustrate, the sample label for pathological image A is 'a'. The pathological image database contains a set of sample pathological image labels, which includes three labels: a, b, and c. Matching the sample label with the labels in the sample pathological image label set yields matching results of "1, 0, 0". Inputting pathological image A into a classification model, the model outputs predicted probabilities for the three pathological image labels a, b, and c, respectively: "0.7, 0.1, 0.2". Distance calculations are performed between "1, 0, 0" and "0.7, 0.1, 0.2" to obtain the loss value corresponding to the pathological image.

[0080] It is worth noting that, in this embodiment of the application, the calculation of the loss value using sample labels and label prediction results is used as an example for illustration.

[0081] Furthermore, the above sample prediction probability is illustrated using the soft label output by the classification model as an example. In some embodiments, the sample prediction probability output by the classification model can also be implemented as a hard label. That is, when the sample pathological image A is input into the classification model, the classification model outputs the sample prediction probabilities corresponding to the three pathological image labels a, b, and c, respectively, as "1, 0, 0". This embodiment does not limit this.

[0082] Optionally, the loss value is calculated using a pre-defined loss function. The sample labels and the predicted labels obtained from the sample pathological image database are substituted into this loss function to obtain the loss value. The classification model is then trained based on this loss value to obtain a pathological image classification model. The pathological image classification model is obtained after training the classification model.

[0083] In an optional embodiment, a pathological image classification model is used to classify the target pathological image to obtain prediction results corresponding to at least two pathological image databases.

[0084] Optionally, after obtaining the pathological image classification model, the target pathological image can be classified using the model. The target pathological image can be a pathological image from at least two pathological image databases, or it can be a pathological image from outside of at least two such databases. That is, the application of the pathological image classification model is not limited to the at least two pathological image databases used when constructing the model. Illustratively, the pathological image classification model can perform relatively accurate classification predictions for pathological images from pathological image databases of the same type as the at least two pathological image databases. By inputting the target pathological image into the model, the prediction results corresponding to the target pathological image and the at least two pathological image databases are obtained. For example, the pathological image label corresponding to the target pathological image in the at least two pathological image databases can be determined.

[0085] Optionally, before acquiring the target pathological image, a reference pathological image is acquired, which is the target pathological image before preprocessing. Based on the pathological image type of the reference pathological image, preprocessing is performed on the reference pathological image to obtain a target pathological image more suitable for model recognition. Optionally, the preprocessing of the reference pathological image includes cropping, sharpening, and denoising. For example, taking cropping as an illustration, the reference pathological image is a nuclear pathological image obtained after acquiring the target acquisition point. Then, the region corresponding to the target acquisition point is used to acquire the nuclear pathological image, resulting in at least two target nuclear pathological images, thus obtaining the target pathological image. The above is merely an illustrative example, and the embodiments of this application do not limit this scope.

[0086] In summary, this method involves obtaining sample pathological images from a pathological image database, matching the sample labels corresponding to the sample pathological images with the corresponding pathological image label sets in the database, thereby determining the pathological image database to which the sample pathological image belongs. Then, the sample pathological image is input into a classification model built based on the aforementioned pathological image database, and the output is the predicted probability of the sample corresponding to the pathological image database. Based on the sample labels and the predicted probability, a loss value is determined, thus completing the training process of the classification model and obtaining a pathological image classification model. In application, the target pathological image to be classified is input into the aforementioned pathological image classification model, and multiple pathological image labels related to the pathological image database are output as the prediction result corresponding to the target pathological image, achieving the classification of the target pathological image. This method avoids the problem of needing to input the target pathological image into multiple classification models corresponding to multiple pathological image databases to obtain prediction results. The pathological image classification model in this application strengthens the correlation between pathological image databases, better achieving centralized, professional, and standardized pathological image classification, and facilitating better management and use of pathological images.

[0087] In an optional embodiment, the process of training the classification model to obtain a pathological image classification model is achieved through a loss value. (Illustrative example, such as...) Figure 3 As shown above, Figure 2 Step 240 in the illustrated embodiment can also be implemented as steps 310 to 340.

[0088] Step 310: Obtain the predicted probability of the sample corresponding to the sample pathological image library in the pathological image results of the classification model.

[0089] Optionally, after inputting the sample pathological images into the classification model, the classification model outputs pathological image results, for example, the pathological image results are output in probability form. The pathological image results include both the predicted probability of the sample corresponding to the sample pathological image library and the predicted probability of the pathological images corresponding to other pathological image libraries besides the sample pathological image library.

[0090] The sample prediction probability includes the prediction probabilities corresponding to the sample pathological image labels in the sample pathological image library. That is, the sample prediction probability is the prediction probability output by the classification model, which is the prediction probability corresponding to the sample pathological image and multiple sample pathological image labels respectively.

[0091] Step 320: The label prediction result of the predicted probability of the corresponding sample in the sample pathological image library is input into the loss function along with the sample label to obtain the loss value corresponding to the sample pathological image.

[0092] Optionally, the loss value can be calculated using a loss function, which is a predefined function. For example, the expression for the loss function is shown below.

[0093]

[0094] Where L0 represents the loss function; i represents the i-th pathological image, B represents the pathological image sampling group; j represents the j-th pathological image database, and D represents the entire pathological image database; d i,j It is a 0 / 1 indicator function. A value of 1 indicates that pathological image i has a pathological image label from the corresponding pathological image label set in pathological image library j; a value of 0 indicates that pathological image i does not have a pathological image label from the corresponding pathological image label set in pathological image library j. m represents the m-th label in the j-th pathological image library. j The pathological image label set has a total of S. j Pathological image labels; y i,j (m) represents the matching probability of the m-th label in the j-th pathological image database; y′ i,j (m) represents the predicted probability of the m-th label in the j-th label set output by the model. The pathological image sampling group B can be either the above-mentioned at least two pathological image libraries, or a set of pathological images randomly obtained from the above-mentioned at least two pathological image libraries.

[0095] Optionally, the above loss function expression is a cross-entropy loss function with a 0 / 1 indicator function. The indicator function is used during the training of the classification model to determine the sample pathological image library to which the sample pathological image belongs; that is, it indicates whether the sample label is a pathological image label corresponding to a certain pathological image library. Illustratively, i represents the sample pathological image. When the indicator function is 1, it means that the sample pathological image has a pathological image label from the pathological image label set corresponding to pathological image library j, and this pathological image label is considered the true label. Optionally, when applying the above loss function, when the indicator function is 1, the label prediction result of the sample prediction probability corresponding to the sample pathological image library is included in the loss function calculation; when the indicator function is 0, the sample pathological image does not match the pathological image label in the pathological image label set corresponding to that pathological image library, and the pathological image prediction probabilities corresponding to other pathological image libraries are not included in the loss function calculation.

[0096] Optionally, represents a sample pathological image; j represents a sample pathological image library; y i,j (m) represents the sample label. The label prediction result y′ is obtained based on the classification model and the sample pathological image database. i,j (m), the sample label y i,j (m) and the label prediction result y′ of the sample prediction probability i,j(m) Substitute into the above loss function to calculate the loss value. The loss value is used to indicate the degree of difference between the sample label and the label prediction result. Optionally, the degree of difference between the sample label and the label prediction result is represented by probability.

[0097] Based on the above loss function, the loss value corresponding to the sample pathological image is determined. Optionally, when obtaining pathological images from the pathological image sampling group for prediction, if the selected sample pathological images are multiple pathological images, the loss value is calculated for each sample pathological image using the above method. Alternatively, after selecting a sample pathological image once, the classification model needs to be trained multiple times. It can either continue to use the previously selected sample pathological images to train the classification model, or reselect other sample pathological images from at least two pathological image databases to train the classification model. The above are only illustrative examples, and the embodiments of this application do not limit this.

[0098] Step 330: Based on the loss value corresponding to the sample label, adjust the model parameters of the classification model to obtain the candidate classification model.

[0099] This illustrates how to adjust the model parameters of a classification model with the goal of reducing the loss value, such as by using gradient descent or backpropagation.

[0100] Optionally, based on the loss value corresponding to a sample label, the model parameters of the classification model can be adjusted at least once. When there are multiple sample labels corresponding to multiple pathological images, the model parameters of the classification model need to be adjusted multiple times. The purpose of adjusting the model parameters of the classification model is to obtain a well-trained pathological image classification model. Schematic, in the process of adjusting the model parameters of the classification model to obtain a pathological image classification model, the model whose parameters have been adjusted but have not yet reached the conditions for a pathological image classification model can be called a candidate classification model. That is, the candidate classification model is the model obtained after adjusting the model parameters of the classification model. Because it has not been fully trained, the candidate classification model is an intermediate state model.

[0101] Step 340: In response to the training objective being achieved by training the candidate classification model based on the loss value, a pathological image classification model is obtained.

[0102] As an illustration, multiple sample pathological images are randomly selected from at least two pathological image databases. The classification model is trained using each sample pathological image as a prerequisite. For example, after performing label matching, probability prediction, and loss calculation for each sample pathological image, a loss value corresponding to each sample pathological image is obtained. Then, the classification model is initially adjusted using the first loss value corresponding to the first sample pathological image to obtain a candidate classification model. Subsequently, the candidate classification model is trained using the second loss value corresponding to the second sample pathological image, and so on. Optionally, the same sample pathological image can be used to train the classification model once or multiple times. The above is merely an illustrative example, and the embodiments of this application do not limit this practice.

[0103] Optionally, during the training of the candidate classification model with the loss value, a pathological image classification model will be obtained as the training of the candidate classification model reaches the training objective. Indicatively, the training objective includes at least one of the following situations.

[0104] 1. In response to the convergence of the loss value, the candidate classification model obtained from the most recent iteration of training is used as the pathological image classification model.

[0105] Indicatively, the convergence of the loss value indicates that the value of the loss obtained through the loss function no longer changes or the change is less than a preset threshold. For example, if the loss value corresponding to the nth pathological image is 0.1 and the loss value corresponding to the (n+1)th pathological image is also 0.1, this can be considered as the loss value reaching a convergence state. The candidate classification model with the adjusted loss value corresponding to the nth or (n+1)th pathological image is then used as the pathological image classification model to achieve the training process of the classification model.

[0106] 2. In response to the number of times the loss value is obtained reaching the threshold, the candidate classification model obtained from the most recent iteration of training is used as the pathological image classification model.

[0107] In a schematic representation, each acquisition yields one loss value. The number of times the loss value is acquired for training the classification model is pre-defined. When one pathological image corresponds to one loss value, the number of acquisitions equals the number of pathological images; conversely, when one pathological image corresponds to multiple loss values, the number of acquisitions equals the number of loss values. For example, if one acquisition yields one loss value, and the threshold for the number of acquisitions is 10, then when the threshold is reached, the candidate classification model with the most recent loss value adjustment is used as the pathological image classification model, or the candidate classification model with the smallest loss value adjustment within the 10 loss value adjustments is used as the pathological image classification model, thus completing the training process for the classification model.

[0108] The above are merely illustrative examples, and the embodiments of this application are not intended to limit the scope of the application.

[0109] In summary, by matching sample labels with the corresponding pathological image label sets in a pathological image database, the corresponding database for a sample pathological image is determined. Then, the sample pathological image is input into a classification model, which outputs the predicted probability of the corresponding sample in the database. The loss value determined by the sample labels and the predicted probability is used to train the classification model, resulting in a pathological image classification model. In application, the target pathological image is input into the classification model, and multiple pathological image labels are output as the predicted results for the target pathological image, thus achieving the classification of the target pathological image. The pathological image classification model in this application strengthens the correlation between pathological image databases, better achieves centralized, professional, and standardized pathological image classification, and facilitates better management and use of pathological images.

[0110] In this embodiment, the process of training a classification model based on a loss value to obtain a pathological image classification model is described. Based on the label prediction results of the sample pathological image obtained by inputting the sample pathological image into the classification model and the sample label corresponding to the sample pathological image, the degree of difference between the true label of the sample pathological image and the predicted label of the sample pathological image obtained by the model to be trained can be known, i.e., the loss value. The loss value can be calculated using the aforementioned loss function. Based on the loss value, the model parameters of the classification model are adjusted until the training objective is achieved to obtain the pathological image classification model. Through the above method, the robustness of the classification model can be improved by utilizing the loss value, resulting in a pathological image classification model that can better predict the classification of pathological images.

[0111] In an optional embodiment, after obtaining the pathological image classification model, the target pathological image can be classified using the pathological image classification model to obtain the pathological image classification result. (Illustrative example, such as...) Figure 4 As shown above, Figure 2 The illustrated embodiment includes steps 410 to 430 after step 240.

[0112] Step 410: Input the target pathological image into the pathological image classification model to determine the predicted probability of the target pathological image and the pathological image labels in the total set of pathological image labels.

[0113] In an optional embodiment, the target pathological image is input into a pathological image classification model; the pathological image classification model predicts the prediction results corresponding to the target pathological image and the pathological image labels in the total set of pathological image labels, respectively, to obtain the label prediction probability.

[0114] In illustration, a pathological image classification model is obtained by training a classification model with sample pathological images. Target pathological images can include various forms such as text, images, and audio. Target pathological images can be obtained through extraction, cropping, and synthesis. After obtaining the target pathological image, it can be input into the pathological image classification model; alternatively, the target pathological image can be preprocessed (e.g., text segmentation; image cropping; audio decoding and encoding) before being input into the pathological image classification model.

[0115] The pathological image label set is a collection of pathological image label sets corresponding to at least two pathological image databases. Illustratively, each pathological image database corresponds to one pathological image label set, and each pathological image label set includes at least one pathological image label. The total pathological image label set includes all pathological image labels corresponding to at least two pathological image databases. Optionally, a target pathological image is labeled with a target label. After the target pathological image is input into a pathological image classification model, the classification model outputs multiple label prediction probabilities equal to the number of pathological image labels. Illustratively, the total pathological image label set contains x pathological image labels. After a target pathological image is input into a pathological image classification model, the model outputs x label prediction probabilities.

[0116] Step 420: Based on the division criteria between at least two pathological image databases, the total set of pathological image labels is divided to obtain the predicted label probabilities of the target pathological image in at least two pathological image databases.

[0117] This is illustrative, suggesting that at least two pathological image databases exist under certain criteria for classification. For example, the databases may originate from different sources; they may store different types of pathological images; or the relationships between the pathological images in the databases may differ. Optionally, when acquiring at least two pathological image databases, the classification criteria between them are determined accordingly. For instance, if pathological image database Y originates from research institution Y and pathological image database Z originates from research institution Z, the difference in their origins constitutes the classification criterion between database Y and database Z.

[0118] Optionally, different pathological image databases correspond to different pathological image label sets. The total set of pathological image labels is a collection of pathological image label sets. The process of dividing the total set of pathological image labels according to the classification criteria between at least two pathological image databases is the process of dividing the total set of pathological image labels to determine the pathological image label sets. A pathological image label set includes multiple pathological image labels. After the target pathological image is input into the pathological image classification model, the pathological image classification model outputs multiple label prediction probabilities. Based on the classification criteria of the pathological image databases, the multiple label prediction probabilities corresponding to a pathological image database are grouped together. Illustratively, the classification model is trained based on P pathological image databases, and the multiple label prediction probabilities corresponding to each of the P pathological image databases are subsequently divided into P groups.

[0119] Step 430: Based on the predicted probability of the label corresponding to the target pathological image in at least two pathological image databases, determine the pathological image classification result corresponding to the target pathological image.

[0120] The label prediction probability can be used to determine the degree of difference between the target label (the true label of the target pathological image) and multiple pathological image labels. Optionally, after obtaining the label prediction probability of the target pathological image in at least two pathological image databases, the pathological image classification result corresponding to the target pathological image can be determined based on the label prediction probability. The pathological image classification result is the pathological image label in at least two pathological image databases corresponding to the target pathological image.

[0121] Schematic, in response to the prediction probability of the label corresponding to the target pathological image in at least two pathological image databases reaching the prediction probability criterion, the pathological image classification results corresponding to the target pathological image in at least two pathological image databases are determined. Wherein, depending on the prediction probability criterion, the pathological image classification results include at least one of the following cases.

[0122] 1. The prediction probability standard is to take the label prediction probability that is the highest.

[0123] To illustrate, after inputting the target pathological image into the pathological image classification model, d label prediction probabilities are obtained. The d label prediction probabilities are compared numerically, and the pathological image label corresponding to the highest label prediction probability is taken as the pathological image classification result, thus determining the pathological image classification of the target pathological image.

[0124] 2. The prediction probability standard is to reach a pre-set probability threshold.

[0125] In a schematic manner, after inputting the target pathological image into the pathological image classification model, e label prediction probabilities are obtained. A probability threshold of 0.8 is preset. The pathological image labels corresponding to the prediction probabilities exceeding 0.8 among the e label prediction probabilities are taken as the pathological image classification results, thus determining the pathological image classification of the target pathological image. Optionally, when determining the pathological image classification result based on the probability threshold, there may be cases where the target pathological image corresponds to 0, 1, or multiple pathological image labels in a pathological image library. In such cases, either all corresponding pathological image labels can be retained to obtain multiple pathological image classification results, or a unique pathological image label can be determined as the unique pathological image classification result. This embodiment does not limit this approach.

[0126] In summary, by matching sample labels with the corresponding pathological image label sets in a pathological image database, the corresponding database for a sample pathological image is determined. Then, the sample pathological image is input into a classification model, which outputs the predicted probability of the corresponding sample in the database. The loss value determined by the sample labels and the predicted probability is used to train the classification model, resulting in a pathological image classification model. In application, the target pathological image is input into the classification model, and multiple pathological image labels are output as the predicted results for the target pathological image, thus achieving the classification of the target pathological image. The pathological image classification model in this application strengthens the correlation between pathological image databases, better achieves centralized, professional, and standardized pathological image classification, and facilitates better management and use of pathological images.

[0127] In this embodiment, the pathological image classification model obtained by the above-mentioned pathological image classification method is used to classify the target pathological image to be classified, determine the target label corresponding to the target pathological image and the label prediction probability corresponding to each pathological image label, divide the label prediction probabilities based on the division criteria between pathological image databases, and then determine the pathological image classification result corresponding to the target pathological image based on the label prediction probabilities corresponding to different pathological image databases. This achieves the purpose of inputting the target pathological image into a pathological image classification model to obtain multiple labels at the same time, realizing the process of multi-label classification of the target pathological image.

[0128] In an optional embodiment, the aforementioned pathological image classification model can perform multi-label classification on pathological images, meaning that each target pathological image can be predicted with one or more classification labels, thereby achieving multi-label classification based on a set of multiple pathological images. Illustratively, the pathological image classification method is applied to cell nucleus classification, and the implementation of the pathological image classification method includes both training and application processes.

[0129] I. Training Process

[0130] Optionally, the untrained model used for classification is called a classification model; the trained classification model that performs multi-label classification of cell nuclei is called a cell nucleus classification model, i.e., the cell nucleus classification model is obtained after training the classification model using a pathological image database. When training the cell nucleus classification model, the cell nucleus pathological image database and the corresponding pathological image label set are first determined. Optionally, the classification model is constructed based on five databases, each storing pathological images; therefore, the databases can be considered as pathological image databases. The five databases are: NuCLS (Nucleus Classification), BreCaHAD (Breast Cancer Histopathological Annotation and Diagnosis), CoNSeP (Colorectal Nucleus Segmentation and Phenotype), MoNuSAC (Multi-organ Nuclei Segmentation and Classification), and panNuke (pan-cancer histology dataset for Nucleiinstance segmentation and classification). Illustratively, each database (which can be considered a pathological image library) corresponds to a data label set (which can be considered a pathological image label set). Each data label set includes multiple data labels (which can be considered pathological image labels). Data labels in different data label sets may be the same or different. Illustratively, each of the above five databases corresponds to a data label set, as shown below. Arabic numerals are used to indicate different data labels.

[0131] NuCLS data tag set: 0. Tumor cells; 1. Fibroblasts; 2. Lymphocytes; 3. Plasma cells; 4. Macrophages; 5. Mitotic figures; 6. Vascular endothelium; 7. Myoepithelial cells; 8. Apoptotic bodies; 9. Neutrophils; 10. Ductal epithelial cells; 11. Eosinophils

[0132] BreCaHAD data tag set: 0. Mitotic cells; 1. Non-mitotic cells; 2. Apoptotic cells; 3. Tumor cells; 4. Non-tumor cells; 5. Lumen cells; 6. Non-lumen cells

[0133] CoNSeP data label set: 0. Other cells; 1. Inflammatory cells; 2. Healthy epithelial cells; 3. Malignant epithelial cells; 4. Fibroblasts; 5. Connective cells; 6. Endothelial cells

[0134] MoNuSAC data label set: 0, epithelial cells; 1, lymphocytes; 2, neutrophils; 3, macrophages

[0135] PanNuke data label set: 0, tumor cells; 1, inflammatory cells; 2, connective cells; 3, dead cells; 4, epithelial cells

[0136] The above 5 data tag sets correspond to a total of 35 data tags. Among them, the NuCLS data tag set corresponds to 12 data tags; the BreCaHAD data tag set corresponds to 7 data tags; the CoNSeP data tag set corresponds to 7 data tags; the MoNuSAC data tag set corresponds to 4 data tags; and the PanNuke data tag set corresponds to 5 data tags.

[0137] When training the classification model, a sample data point is randomly selected from the five databases mentioned above and input into the model. The model matches the sample data with 35 data labels, outputting a total of 35 probabilities in the final layer to indicate the differences between the sample data and the 35 data labels. However, during training, the focus is on the true label of the sample data itself (i.e., the sample label). The cross-entropy of the predicted probability of the sample data label is used as the loss function, and other probabilities besides the predicted probability are ignored. Illustratively, the probabilities output by the classification model are sequentially output in the order of the dataset and data labels. If the sample data is obtained from the NuCLS dataset, the cross-entropy loss function is only calculated for the first 12 predicted probabilities to obtain the loss value.

[0138] Optionally, the data from the above five datasets can be processed through a classification model to determine the loss values ​​corresponding to different data. These loss values ​​can then be used to train the classification model, thus achieving the process of training a classification model using the data in the datasets. The loss function is the function used to train the classification model. By determining the loss value through the loss function, the encoding part before the classification model can be trained more effectively, improving the accuracy of the classification model's predictions and ultimately resulting in a cell nucleus classification model with higher accuracy and the ability to perform multi-label classification.

[0139] II. Application Process

[0140] HE staining (Hematoxylin-Eosin staining) is one of the most commonly used staining methods in microscopic medical analysis. It consists of two staining solutions: hematoxylin, which is alkaline, primarily stains the chromatin in the cell nucleus and nucleic acids in the cytoplasm a purple-blue color; and eosin, which is acidic, primarily stains components in the cytoplasm and extracellular matrix a red color. (Illustrative example, such as...) Figure 5 The image shown is the imaging result obtained by staining different cells using the HE staining method.

[0141] Optionally, after obtaining the cell nucleus classification model based on the training process, multi-label classification of the target cell nucleus is performed, that is, using different category labels corresponding to multiple databases to classify the same cell nucleus in multiple ways. (Illustrative example follows.) Figure 6 As shown, the application process of multi-label classification of cell nuclei based on the cell nucleus classification model includes the following steps.

[0142] Step 610, cell detection and segmentation.

[0143] Indicative, such as Figure 7 As shown, when analyzing cell nuclei, HE staining is applied to facilitate observation. Subsequently, a pathological image 710 with stained cell nuclei is acquired. This image 710 is then input into a general image detection and segmentation model, such as the Hovernet model (used for simultaneous segmentation and classification of cell nuclei in multi-tissue histological images); or, a general image detection and segmentation algorithm, such as Mask R-CNN (Mask Region-Convolution Neural Network) and Fast R-CNN (Fast Region-Convolution Neural Network). Schematively, the pathological image 710 is input into the Hovernet model to obtain a detection image 720, thereby obtaining the location information of the cells to be detected. Optionally, a cell image 730 of size 64×64 is cropped from the cell to be detected in the detection image 720. The cell image 730 is then enlarged to a size of 224×224 and input into a trained cell nucleus classification model for classification processing.

[0144] The detection image 720 is cut into a 64×64 cell image 730 to present the cell state of the cell to be detected more comprehensively and accurately. The cell image 730 is enlarged to 224×224 because the 224×224 image is more convenient for the cell nucleus classification model to perform classification prediction.

[0145] Step 620, Cell multi-label classification.

[0146] In a schematic representation, inputting a target cell image into a trained cell nucleus classification model enables multi-label cell classification. The target cell image is the cell image to be classified and predicted. The cell nucleus classification model yields 35 predicted probabilities for the target cell image. These 35 probabilities are then divided according to database partitioning criteria: 12 probabilities for the NuCLS database; 7 probabilities for the BreCaHAD database; 7 probabilities for the CoNSeP database; 4 probabilities for the MoNuSAC database; and 5 probabilities for the PanNuke database. Optionally, the label with the highest predicted probability from each database is used as the predicted label for the target cell image; therefore, one target cell image corresponds to five predicted labels.

[0147] Optionally, the predicted label can be represented by a corresponding data label set and a corresponding data label. Illustratively, the data label set corresponding to the predicted label is represented by the name of the data label set; the data label corresponding to the predicted label is represented by its Arabic numeral number as a subscript (where the Arabic numeral number corresponds to the Arabic numeral number corresponding to different data labels in the aforementioned data label set). For example, if the predicted label for the target cell image is NuCLS0, it means that the data label set corresponding to this predicted label is the NuCLS data label set, and the data label corresponding to this predicted label is the label result corresponding to data label 0 in the NuCLS data label set—tumor cells. Similarly, if the predicted label for the target cell image is BreCaHAD3, it means: tumor cells under the BreCaHAD data label set; if the predicted label for the target cell image is CoNSeP3, it means: malignant epithelial cells under the CoNSeP data label set; if the predicted label for the target cell image is MoNuSAC0, it means: epithelial cells under the MoNuSAC data label set; if the predicted label for the target cell image is panNuke0, it means: tumor cells under the panNuke data label set.

[0148] Optionally, the above data classification method can be applied to a terminal or server to achieve real-time automatic multi-label classification of cell nuclei in the extracted pathological images. When using the five databases mentioned above, five labels will be output for each cell nucleus. The input and output results are illustrated below. Figure 8 As shown.

[0149] Pathological image 810 corresponds to pathological label 811. Inputting pathological image 810 into the data classification model will yield the label predictions for pathological image 810 and the five datasets mentioned above. For example: pathological image 810 corresponds to the NuCLS data label set with a NuCLS label prediction of 821; pathological image 810 corresponds to the BreCaHAD data label set with a BreCaHAD label prediction of 831; pathological image 810 corresponds to the CoNSeP data label set with a CoNSeP label prediction of 841; pathological image 810 corresponds to the MoNuSAC data label set with a MoNuSAC label prediction of 851; and pathological image 810 corresponds to the panNuke data label set with a panNuke label prediction of 861.

[0150] In summary, by matching sample labels with the corresponding pathological image label sets in a pathological image database, the corresponding database for a sample pathological image is determined. Then, the sample pathological image is input into a classification model, which outputs the predicted probability of the corresponding sample in the database. The loss value determined by the sample labels and the predicted probability is used to train the classification model, resulting in a pathological image classification model. In application, the target pathological image is input into the classification model, and multiple pathological image labels are output as the predicted results for the target pathological image, thus achieving the classification of the target pathological image. The pathological image classification model in this application strengthens the correlation between pathological image databases, better achieves centralized, professional, and standardized pathological image classification, and facilitates better management and use of pathological images.

[0151] In this embodiment, based on the above-mentioned pathological image classification method, the target label corresponding to the target cell image is matched with the data label in the data label set corresponding to each training dataset. Based on the predicted probability obtained by matching, the data label with the highest predicted probability is determined from the data label set corresponding to each dataset as the classification prediction result corresponding to the target cell image. Thus, multiple data labels corresponding to a target cell image can be obtained, achieving the purpose of multi-label classification of a target cell image.

[0152] Figure 9 This is a structural block diagram of a pathological image classification device provided in an exemplary embodiment of this application, such as... Figure 9 As shown, the device includes the following parts:

[0153] The acquisition module 910 is used to acquire sample pathological images from at least two pathological image libraries, wherein the sample pathological images are labeled with sample tags, and the at least two pathological image libraries each correspond to a pathological image tag set;

[0154] The matching module 920 is used to match the sample label with the pathological image label sets corresponding to the at least two pathological image libraries respectively, and to determine the sample pathological image library that matches the sample label from the at least two pathological image libraries;

[0155] The determination module 930 is used to input the sample pathological image into the classification model to obtain the sample prediction probability corresponding to the sample pathological image library. The classification model is a training model constructed based on the at least two pathological image libraries.

[0156] The training module 940 is used to train the classification model based on the sample prediction probability corresponding to the sample pathological image library and the sample label to determine the loss value, thereby obtaining a pathological image classification model. The pathological image classification model is used to classify the target pathological image and obtain prediction results corresponding to the at least two pathological image libraries respectively.

[0157] like Figure 10 As shown, in an optional embodiment, the sample pathology image tag set corresponding to the sample pathology image library includes at least one sample pathology image tag;

[0158] The training module 940 includes:

[0159] The acquisition unit 941 is used to acquire the sample prediction probability corresponding to the sample pathological image library in the pathological image results of the classification model, wherein the sample prediction probability includes the prediction probability corresponding to the sample pathological image label of the sample pathological image library respectively.

[0160] Input unit 942 is used to input the label prediction result of the sample prediction probability corresponding to the sample pathological image library and the sample label into the loss function to obtain the loss value corresponding to the sample pathological image.

[0161] In an optional embodiment, the pathological image tag set includes at least one pathological image tag;

[0162] The matching model 920 is used to match the sample label with the pathological image labels in the pathological image label sets corresponding to the at least two pathological image libraries respectively; in response to the pathological image label set corresponding to the pathological image library including the sample label, the pathological image library is determined as the sample pathological image library.

[0163] In an optional embodiment, the apparatus further includes:

[0164] The probability determination module 950 is used to input the target pathological image into the pathological image classification model and determine the predicted probability of the target pathological image and the pathological image label in the pathological image label set, respectively. The pathological image label set is the set of pathological image label sets corresponding to the at least two pathological image libraries.

[0165] The segmentation module 960 is used to segment the total set of pathological image labels based on the segmentation criteria between the at least two pathological image libraries, and to obtain the predicted probability of the label corresponding to the target pathological image in the at least two pathological image libraries respectively.

[0166] The result determination module 970 is used to determine the pathological image classification result corresponding to the target pathological image based on the predicted probability of the label corresponding to the target pathological image in the at least two pathological image libraries.

[0167] In an optional embodiment, the probability determination module 950 is further configured to input the target pathological image into the pathological image classification model; and predict the prediction results corresponding to the target pathological image and the pathological image labels in the total set of pathological image labels through the pathological image classification model to obtain the label prediction probability.

[0168] In an optional embodiment, the result determination module 970 is further configured to determine the pathological image classification results of the target pathological image corresponding to the at least two pathological image libraries in response to the prediction probability of the label corresponding to the target pathological image in the at least two pathological image libraries reaching the prediction probability standard.

[0169] In an optional embodiment, the training module 940 is further configured to determine a loss value based on the sample prediction probability corresponding to the sample pathological image library and the sample label; adjust the model parameters of the classification model based on the loss value corresponding to the sample label to obtain a candidate classification model; and obtain the pathological image classification model in response to the training of the candidate classification model based on the loss value reaching the training objective.

[0170] In an optional embodiment, the training module 940 is further configured to, in response to the loss value reaching a convergence state, use the candidate classification model obtained in the most recent iteration of training as the pathological image classification model; or, in response to the number of times the loss value is acquired reaching a threshold, use the candidate classification model obtained in the most recent iteration of training as the pathological image classification model.

[0171] In an optional embodiment, the acquisition module 910 is further configured to randomly select at least one pathological image from the at least two pathological image libraries as the sample pathological image; or, to poll and select pathological images from the at least two pathological image libraries as the sample pathological image.

[0172] It should be noted that the pathological image classification device provided in the above embodiments is only an example of the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the pathological image classification device and the pathological image classification method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0173] Figure 11 This illustration shows a schematic diagram of a server provided in an exemplary embodiment of this application. The server 1100 includes a Central Processing Unit (CPU) 1101, a system memory 1104 including Random Access Memory (RAM) 1102 and Read Only Memory (ROM) 1103, and a system bus 1105 connecting the system memory 1104 and the CPU 1101. The server 1100 also includes a mass storage device 1106 for storing an operating system 1113, application programs 1114, and other program modules 1115.

[0174] Mass storage device 1106 is connected to central processing unit 1101 via a mass storage controller (not shown) connected to system bus 1105. Mass storage device 1106 and its associated computer-readable media provide non-volatile storage for server 1100. That is, mass storage device 1106 may include computer-readable media (not shown) such as hard disk or compact disc read-only memory (CD-ROM) drives.

[0175] Without loss of generality, computer-readable media can include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include RAM, ROM, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other solid-state storage technologies, CD-ROM, digital versatile disc (DVD) or other optical storage, magnetic tape cassettes, magnetic tape, disk storage, or other magnetic storage devices. Of course, those skilled in the art will recognize that computer storage media are not limited to the above-mentioned types. The system memory 1104 and mass storage device 1106 described above can be collectively referred to as memory.

[0176] According to various embodiments of this application, server 1100 can also be connected to a remote computer on a network, such as the Internet. That is, server 1100 can be connected to network 1112 via network interface unit 1111 connected to system bus 1105, or it can also use network interface unit 1111 to connect to other types of networks or remote computer systems (not shown).

[0177] The aforementioned memory also includes one or more programs, which are stored in the memory and configured to be executed by the CPU.

[0178] Embodiments of this application also provide a computer device, which includes a processor and a memory. The memory stores at least one instruction, at least one program, code set, or instruction set. The at least one instruction, at least one program, code set, or instruction set is loaded and executed by the processor to implement the pathological image classification method provided in the above-described method embodiments.

[0179] Embodiments of this application also provide a computer-readable storage medium storing at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, at least one program, code set, or instruction set is loaded and executed by a processor to implement the pathological image classification method provided in the above-described method embodiments.

[0180] Embodiments of this application also provide a computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform any of the pathological image classification methods described in the above embodiments.

[0181] Optionally, the computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), solid-state drives (SSDs), or optical discs, etc. The random access memory may include resistive random access memory (ReRAM) and dynamic random access memory (DRAM). The sequence numbers of the embodiments described above are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0182] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0183] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for classifying pathological images, characterized in that, The method includes: Sample pathological images are obtained from at least two pathological image libraries, wherein the sample pathological images are labeled with sample tags, and the at least two pathological image libraries each correspond to a pathological image tag set; The sample label is matched with the pathological image label sets corresponding to the at least two pathological image libraries respectively, and the sample pathological image library that matches the sample label is determined from the at least two pathological image libraries. The sample pathological image label set corresponding to the sample pathological image library includes at least one sample pathological image label. The sample pathological images are input into the classification model to obtain the sample prediction probability corresponding to the sample pathological image database. The classification model is a training model constructed based on the at least two pathological image databases. Obtain the sample prediction probability corresponding to the sample pathological image library in the pathological image results of the classification model, wherein the sample prediction probability includes the prediction probability corresponding to the sample pathological image label of the sample pathological image library respectively; The label prediction result of the sample prediction probability corresponding to the sample pathological image library is input into the loss function along with the sample label to obtain the loss value corresponding to the sample pathological image. The classification model is then trained to obtain a pathological image classification model. The pathological image classification model is used to classify the target pathological image and obtain prediction results corresponding to the at least two pathological image libraries respectively.

2. The method according to claim 1, characterized in that, The pathological image tag set includes at least one pathological image tag; The step of matching the sample label with the pathological image label sets corresponding to the at least two pathological image libraries, and determining the sample pathological image library that matches the sample label from the at least two pathological image libraries, includes: The sample label is matched with the pathological image labels in the pathological image label sets corresponding to the at least two pathological image libraries respectively; In response to the fact that the pathological image tag set corresponding to the pathological image library includes the sample tag, the pathological image library is determined as the sample pathological image library.

3. The method according to claim 1 or 2, characterized in that, After training the classification model to obtain the pathological image classification model, the process further includes: The target pathological image is input into the pathological image classification model to determine the label prediction probability corresponding to the target pathological image and the pathological image label in the pathological image label set, wherein the pathological image label set is the set of pathological image label sets corresponding to the at least two pathological image libraries respectively. Based on the division criteria between the at least two pathological image databases, the total set of pathological image labels is divided to obtain the predicted label probabilities of the target pathological image in the at least two pathological image databases respectively. Based on the predicted probability of the label corresponding to the target pathological image in the at least two pathological image databases, the pathological image classification result corresponding to the target pathological image is determined.

4. The method according to claim 3, characterized in that, The step of inputting the target pathological image into the pathological image classification model and determining the label prediction probabilities corresponding to the target pathological image and the pathological image labels in the total set of pathological image labels includes: The target pathological image is input into the pathological image classification model; The pathological image classification model is used to predict the prediction results of the target pathological image and the pathological image labels in the total set of pathological image labels, respectively, to obtain the label prediction probability.

5. The method according to claim 3, characterized in that, The step of determining the pathological image classification result corresponding to the target pathological image based on the predicted probability of the label corresponding to the target pathological image in the at least two pathological image databases includes: In response to the prediction probability of the label corresponding to the target pathological image in the at least two pathological image databases reaching the prediction probability standard, the pathological image classification results corresponding to the target pathological image in the at least two pathological image databases are determined.

6. The method according to claim 1 or 2, characterized in that, The process of training the classification model to obtain a pathological image classification model includes: Based on the loss value corresponding to the pathological image of the sample, the model parameters of the classification model are adjusted to obtain a candidate classification model; In response to the training objective being achieved by training the candidate classification model based on the loss value, the pathological image classification model is obtained.

7. The method according to claim 6, characterized in that, The step of obtaining the pathological image classification model in response to achieving the training objective based on the loss value of the candidate classification model includes: In response to the loss value reaching a convergent state, the candidate classification model obtained from the most recent iteration of training is used as the pathological image classification model; or, When the number of times the loss value is obtained reaches a threshold, the candidate classification model obtained from the most recent iteration of training is used as the pathological image classification model.

8. The method according to claim 1 or 2, characterized in that, The process of obtaining sample pathological images from at least two pathological image libraries includes: At least one pathological image is randomly selected from the at least two pathological image libraries as the sample pathological image; or, Pathological images are selected from the at least two pathological image libraries in a round-robin fashion as the sample pathological images.

9. A pathological image classification device, characterized in that, The device includes: The acquisition module is used to acquire sample pathological images from at least two pathological image libraries, wherein the sample pathological images are labeled with sample tags, and the at least two pathological image libraries each correspond to a pathological image tag set; The matching module is used to match the sample label with the pathological image label sets corresponding to the at least two pathological image libraries respectively, and to determine the sample pathological image library that matches the sample label from the at least two pathological image libraries. The sample pathological image label set corresponding to the sample pathological image library includes at least one sample pathological image label. A model is determined to input the sample pathological images into the classification model to obtain the sample prediction probability corresponding to the sample pathological image library. The classification model is a training model constructed based on the at least two pathological image libraries. The training module is used to obtain the sample prediction probability corresponding to the sample pathological image library in the pathological image results of the classification model. The sample prediction probability includes the prediction probability corresponding to the sample pathological image label of the sample pathological image library. The label prediction result of the sample pathological image library corresponding to the sample prediction probability and the sample label are input into the loss function to obtain the loss value corresponding to the sample pathological image. The classification model is then trained to obtain a pathological image classification model. The pathological image classification model is used to classify the target pathological image and obtain the prediction result corresponding to the at least two pathological image libraries.

10. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing at least one program, which is loaded and executed by the processor to implement the pathological image classification method as described in any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that, The storage medium stores at least one program, which is loaded and executed by a processor to implement the pathological image classification method as described in any one of claims 1 to 8.

12. A computer program product, characterized in that, It includes computer instructions that, when executed by a processor, implement the pathological image classification method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Neural network model training method and device for pathological image sample

    CN112232407A

  • Disease type discrimination method based on pathology and living tissue clinical diagnosis big data

    CN112863665A