Cell image analysis method and program product based on morphological thought chain
Patent Information
- Application Number
- CN202610979087.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-02
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]本发明提供了一种基于形态学思维链的细胞图像分析方法及程序产品,以解决相关技术中基于人工对细胞图像进行识别存在一致性低,可靠性差的问题
[0010]本发明实施例的技术方案,首先,通过接收图像分析指令,以及,获取细胞图像;其中,所述细胞图像包括区域标注信息,所述区域标注信息用于指示所述细胞图像中待检测的图像区域,所述图像分析指令用于指示图像分析模型针对所述细胞图像中区域标注信息进行思维链推理的输出内容;获取图像分析指令约束模型的输出方向,确保图像处理的定向性,获取细胞图像以及其中的区域标注信息精准确定细胞图像内需要分析的目标区域,避免无关背景带的干扰,降低无效数据计算量,提高图像处理效率;接着,通过将所述细胞图像和所述图像分析指令输入至所述图像分析模型中,以获取与所述细胞图像对应的图像描述信息;其中,所述图像描述信息至少用于指示所述区域标注信息对应的细胞形态学描述以及与所述细胞形态学描述对应的病变类别;通过图像分析模型实现对细胞图像特征的精准提取,输出细胞形态学描述以及对应的病变类别,避免相关技术中分类模型仅输出单一病变类别标签的局限性,不仅输出病变判定结果,还输出模型判定所根据的细胞形态细节依据,便于工作人员追溯分类判定逻辑,实现对细胞图像的可靠分类。
Smart Images

Figure CN122820601A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and in particular to a cell image analysis method and program product based on morphological thought chain. Background Technology
[0002] Cell image recognition, as a key branch of the medical field, can perform quantitative analysis and objective judgment of cell morphology and tissue structure, providing objective, repeatable and quantifiable data support for clinical disease screening, disease grading, personalized treatment plan formulation, targeted drug development, and efficacy evaluation.
[0003] Current technologies largely rely on manual identification of cell images. However, the varying experience levels of different physicians make it difficult to standardize lesion classification criteria. This can easily lead to different classifications of the same cell image by different personnel, significantly increasing the risk of missed diagnoses and misdiagnoses, and compromising the consistency and reliability of cell image analysis. Therefore, there is an urgent need for a cell image analysis method based on morphological thought processes to improve the accuracy and reliability of cell image classification. Summary of the Invention
[0004] This invention provides a cell image analysis method and program product based on morphological thinking chain to solve the problems of low consistency and poor reliability in related technologies that rely on manual cell image recognition.
[0005] According to one aspect of the present invention, a cell image analysis method based on morphological thought chain is provided, the method comprising: The system receives an image analysis instruction and acquires a cell image; wherein the cell image includes region annotation information, the region annotation information is used to indicate the image region to be detected in the cell image, and the image analysis instruction is used to instruct the image analysis model to perform thought chain reasoning output based on the region annotation information in the cell image; The cell image and the image analysis instructions are input into the image analysis model to obtain image description information corresponding to the cell image; wherein, the image description information is at least used to indicate the cell morphology description corresponding to the region annotation information and the lesion category corresponding to the cell morphology description.
[0006] According to another aspect of the present invention, a cell image analysis device based on morphological thought chain is provided, the device comprising: An image acquisition module is used to receive image analysis instructions and acquire cell images; wherein, the cell images include region annotation information, the region annotation information is used to indicate the image regions to be detected in the cell images, and the image analysis instructions are used to instruct the image analysis model to perform thought chain reasoning output based on the region annotation information in the cell images; An image processing module is used to input the cell image and the image analysis instructions into the image analysis model to obtain image description information corresponding to the cell image; wherein, the image description information is used at least to indicate the cell morphology description corresponding to the region annotation information and the lesion category corresponding to the cell morphology description.
[0007] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the cell image analysis method based on morphological thought chain as described in any embodiment of the present invention.
[0008] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the cell image analysis method based on morphological thought chain as described in any embodiment of the present invention.
[0009] According to another aspect of the present invention, embodiments of this disclosure also provide a computer program product, including a computer program that, when executed by a processor, implements the cell image analysis method based on morphological thought chain as described in any of the embodiments of this disclosure.
[0010] The technical solution of this invention embodiment firstly involves receiving an image analysis instruction and acquiring a cell image; wherein the cell image includes region annotation information, the region annotation information is used to indicate the image region to be detected in the cell image, and the image analysis instruction is used to instruct the image analysis model to perform thought chain reasoning output based on the region annotation information in the cell image; the image analysis instruction constrains the output direction of the model, ensuring the orientation of image processing, and accurately determining the target region to be analyzed within the cell image by acquiring the cell image and its region annotation information, avoiding interference from irrelevant background bands, reducing the amount of invalid data computation, and improving image processing efficiency; then, by... The cell image and the image analysis instructions are input into the image analysis model to obtain image description information corresponding to the cell image; wherein, the image description information is at least used to indicate the cell morphology description corresponding to the region annotation information and the lesion category corresponding to the cell morphology description; the image analysis model achieves accurate extraction of cell image features, outputs cell morphology description and corresponding lesion category, avoids the limitation of classification models in related technologies that only output a single lesion category label, and outputs not only the lesion judgment result, but also the cell morphology details on which the model judgment is based, which makes it easier for staff to trace the classification judgment logic and achieve reliable classification of cell images.
[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a flowchart of a cell image analysis method based on morphological thought chain provided in Embodiment 1 of the present invention; Figure 2 This is a flowchart of a cell image analysis method based on morphological thought chain according to Embodiment 2 of the present invention; Figure 3 This is a schematic diagram of the image analysis model of a cell image analysis method based on morphological thought chain provided in Embodiment 2 of the present invention, which processes cell images. Figure 4This is a schematic diagram of the structure of an image analysis model for a cell image analysis method based on morphological thought chain provided in Embodiment 2 of the present invention; Figure 5 This is a schematic diagram of the labeled cell morphological description corresponding to the region labeling information in the sample cell image of a cell image analysis method based on morphological thought chain provided in Embodiment 2 of the present invention; Figure 6 This is a schematic diagram of the first instruction set of a cell image analysis method based on morphological thought chain provided in Embodiment 2 of the present invention; Figure 7 This is a schematic diagram of the second instruction set of a cell image analysis method based on morphological thought chain provided in Embodiment 2 of the present invention; Figure 8 This is a schematic diagram of cell image processing according to a cell image analysis method based on morphological thought chain provided in Embodiment 2 of the present invention; Figure 9 This is a schematic diagram of the structure of a cell image analysis device based on morphological thought chain according to Embodiment 3 of the present invention; Figure 10 This is a schematic diagram of the structure of an electronic device that implements the cell image analysis method based on morphological thought chain according to an embodiment of the present invention. Detailed Implementation
[0014] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0015] It should be noted that the terms "first," "second," "target," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0016] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0017] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0018] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0019] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0020] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0021] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0022] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0023] Example 1 Figure 1 This document provides a flowchart of a cell image analysis method based on morphological thought chain analysis, as described in Embodiment 1 of the present invention. This embodiment is applicable to situations where morphological thought chain analysis is performed on cell images to generate image description information. This method can be executed by a cell image analysis device based on morphological thought chain analysis. This device can be implemented in hardware and / or software, optionally through an electronic device, such as a mobile terminal, PC, or server. Figure 1 As shown, the method may specifically include: S110. Receive image analysis instructions and acquire cell images; wherein, the cell images include region annotation information, the region annotation information is used to indicate the image regions to be detected in the cell images, and the image analysis instructions are used to instruct the image analysis model to perform thought chain reasoning output based on the region annotation information in the cell images.
[0024] In this embodiment of the invention, the image analysis instruction can be an instruction used to instruct an image analysis model to analyze a cell image. The image analysis instruction can be used to instruct the image analysis model to directionally analyze a specified region in the cell image and to limit the model's output dimension and format; that is, to instruct the image analysis model to perform thought chain reasoning based on the region annotation information in the cell image. The cell image can be understood as a tissue cell image obtained after scanning and imaging a pathological slide; for example, a cell image can be a cervical liquid-based cytology image. The region annotation information can be location annotation data on the cell image, which can be in the form of rectangles, outlines, etc., used to select local regions in the cell image. The image analysis model can be a model used to detect cell images to obtain corresponding image description information.
[0025] Optionally, the system can receive image analysis instructions and acquire cell images, which may include region annotation information. Specifically, it can directly receive input cell images, or obtain the storage location of cell images according to image analysis instructions and retrieve cell images from the storage location.
[0026] S120. Input the cell image and the image analysis command into the image analysis model to obtain image description information corresponding to the cell image; wherein, the image description information is at least used to indicate the cell morphology description corresponding to the region annotation information and the lesion category corresponding to the cell morphology description.
[0027] The image description information can be descriptive information generated by an image analysis model after analyzing the cell image. The image description information is used at least to indicate the cell morphology description corresponding to the region annotation information and the lesion category corresponding to the cell morphology description.
[0028] Optionally, cell morphology description can be descriptive information about cells. For example, cell morphology description may include, but is not limited to, evaluation of nuclear atypia (degree of nuclear enlargement, irregularity of nuclear membrane, chromatin characteristics), analysis of cytoplasmic characteristics and nuclear-cytoplasmic ratio, identification of cell arrangement structure and background components (inflammatory cells, flora, mucus, etc.), and identification analysis of key points (such as "whether the perinuclear halo is typical" and "whether atrophic changes are excluded"), in order to obtain multi-dimensional microscopic pathological features of cells. The lesion category can be a classification derived from the above-mentioned cellular morphology description. For example, for cervical liquid-based cytology images, the lesion category can be classified using pre-set standards. For example, the lesion category may include, but is not limited to, various categories such as Negative for Intraepithelial Lesion or Malignancy (NILM), Atypical Squamous Cells of Undetermined Significance (ASC-US), Low-grade Squamous Intraepithelial Lesion (LSIL), and High-grade Squamous Intraepithelial Lesion (HSIL).
[0029] Optionally, cell images and image analysis instructions can be input into an image analysis model to obtain image description information corresponding to the cell images. The image description information is used at least to indicate the cell morphology description corresponding to the region annotation information and the lesion category corresponding to the cell morphology description.
[0030] Optionally, the image analysis model can perform category determination step by step based on the inference chain. First, it extracts the microscopic morphological features of cells from the region annotation information to generate cell morphological descriptions. Then, it predicts the final classification result based on the cell morphological descriptions to obtain the lesion category corresponding to the cell morphological descriptions.
[0031] For example, such as Figure 8 As shown, users can input cell images and image analysis commands into the interactive interface, and then perform analysis based on the image analysis model to generate image description information corresponding to the cell images. The corresponding response information is then generated in the interactive interface to achieve visual detection.
[0032] The technical solution of this invention embodiment firstly involves receiving an image analysis instruction and acquiring a cell image; wherein the cell image includes region annotation information, the region annotation information is used to indicate the image region to be detected in the cell image, and the image analysis instruction is used to instruct the image analysis model to perform thought chain reasoning output based on the region annotation information in the cell image; the image analysis instruction constrains the output direction of the model, ensuring the orientation of image processing; acquiring the cell image and its region annotation information accurately determines the target region to be analyzed within the cell image, avoiding interference from irrelevant background bands, reducing the amount of invalid data computation, and improving image processing efficiency; then, by... The cell image and the image analysis instructions are input into the image analysis model to obtain image description information corresponding to the cell image; wherein, the image description information is at least used to indicate the cell morphology description corresponding to the region annotation information and the lesion category corresponding to the cell morphology description; the image analysis model achieves accurate extraction of cell image features, outputs cell morphology description and corresponding lesion category, avoiding the limitation of classification models in related technologies that only output a single lesion category label, outputting not only the lesion judgment result, but also the cell morphology details on which the model judgment is based, making it easier for relevant personnel to trace the classification judgment logic and achieve reliable classification of cell images.
[0033] Example 2 Figure 2 This flowchart illustrates a cell image analysis method based on morphological thought chain, as provided in Embodiment 2 of the present invention, further describing the training process of the image analysis model. Specific implementation details can be found in the description of this embodiment. Technical features that are the same as or similar to those in the foregoing embodiments will not be repeated here. Figure 2 As shown, the method may specifically include: S210. Acquire multiple sample cell images and corresponding sample training sets; wherein, the sample training set includes labeled cell morphological descriptions, labeled lesion categories, a first instruction set, and a second instruction set corresponding to the region labeling information in the sample cell images.
[0034] The sample cell images can be cell images used for model training. The sample training set can be a training set corresponding to the sample cell images.
[0035] Optionally, the sample training set may include, but is not limited to, labeled cell morphology descriptions, labeled lesion categories, a first set of instructions, and a second set of instructions corresponding to the region annotation information in the sample cell images. The labeled cell morphology descriptions may be standard cell morphology description text labeled for the sample cell images, which can be used as standard cell morphology description answers for model training. The labeled lesion categories may be lesion labels labeled for the region annotation information in the sample cell images, which can be used as standard lesion classification labels for model training. The first set of instructions may be a set of instructions used to train the model to generate cell morphology descriptions based on the region annotation information in the sample cell images. The second set of instructions may be a set of instructions used to train the model to perform hierarchical reasoning of cell morphology descriptions and lesion categories based on the region annotation information in the sample cell images.
[0036] In related technologies, the training data is mostly composed of natural images, while medical pathology images are scarce, especially those related to liquid-based cytology. This leads to the model's visual-text feature space being biased towards natural images, resulting in insufficient alignment accuracy of cytology pathology features. In order to guide the alignment of visual-text feature spaces in the field of cytology, embed professional domain knowledge, and avoid the pathological cognitive bias of general models, multiple cell morphology feature description text pairs can be constructed.
[0037] Specifically, multiple sample cell images can be acquired and annotated to obtain the annotated cell morphological descriptions and lesion categories corresponding to the annotated regions, such as... Figure 3 As shown, the labeled cell morphology description can be "medium-sized, mononuclear, with an irregular nuclear membrane, unevenly stained, deeply stained nuclear chromatin in a coarse aggregated state, increased nucleoplasm-to-cytoplasm ratio, and no obvious nucleoli; low cytoplasmic weight and light staining; presence of nuclear atypia and koilocytes, but no abnormal keratinization." The lesion category can then be labeled as "Considering the above morphological characteristics, it should be LSIL (low-grade squamous intraepithelial lesion)." Simultaneously, the first and second instruction sets can be obtained to provide learning directions for model training.
[0038] For example, suppose the sample cell image is a cervical liquid-based cytology image, such as Figure 5 The image shown may be an example of a labeled cell morphological description corresponding to the region annotation information in a sample cell image, such as... Figure 6 As shown, the first instruction in the first instruction set is a morphological description instruction, mainly used for extracting cell microscopic features. Instruction examples may include "Please describe the morphological features of the cells in the box.", "What are the morphological features of the cells in the box?", "Please describe the morphological features of the cells in the box.", "Please summarize the morphological features of the cells in the box.", "What morphological features do the cells in the box exhibit?", "Please comment on the morphological features of the cells in the red box.", etc. Figure 7 As shown, the second instruction in the second instruction set is a pathological analysis and reasoning instruction, mainly used for causal determination from characteristics to lesion categories. Instruction examples may include: "Please analyze the cells in the red box of this cervical cytology image and provide your analysis opinion.", "In this cervical liquid-based cytology image, what is the lesion category of the cells in the red box?", "To which category should the cells in the red box in the image be classified?", "Please determine the classification result of the cells in the red box based on their morphological characteristics.", "This is a cervical cytology image; determine the lesion category of the cells in the red box.", "Please determine the category of the cells in the red box.", etc.
[0039] S220. The deep learning model is trained based on multiple sample cell images and their corresponding labeled cell morphological descriptions and a first instruction set to obtain an initial image analysis model; wherein, the first instruction set is used to train the deep learning model to generate cell morphological descriptions corresponding to the region labeling information in the sample cell images.
[0040] The initial image analysis model can be a model trained based on a deep learning model that has the ability to map cell morphology features to text feature space, and can initially realize the generation of cell morphology descriptions.
[0041] Optionally, based on multiple sample cell images and labeled cell morphological descriptions, one instruction can be randomly sampled from the first instruction set to form instruction-answer single-turn dialogue data. The deep learning model is fine-tuned using an autoregressive loss cross-entropy function and combined with low-rank adapter (LoRA), i.e. supervised labeled cell morphological descriptions, to obtain cell morphology perception ability while maintaining the original general capabilities of the model, so as to obtain an initial image analysis model.
[0042] Based on the above scheme, optionally, the deep learning model includes a visual encoder, a fully connected mapping layer, a word segmenter, and a text decoder; the first instruction set includes multiple first instructions; training the deep learning model based on multiple sample cell images and their corresponding labeled cell morphology descriptions and the first instruction set to obtain an initial image analysis model includes: inputting sample cell images and their corresponding labeled cell morphology descriptions and first instructions into the deep learning model to obtain a sequence of labeled description lexical units and predicted probability values of lexical units corresponding to multiple time steps in the predicted cell morphology description; wherein, the sequence of labeled description lexical units is determined based on the labeled cell morphology description; calculating a first cross-entropy loss based on the predicted probability values of lexical units corresponding to multiple time steps in the predicted cell morphology description and the lexical units corresponding to multiple time steps in the labeled description lexical unit sequence; adding a low-rank adapter to multiple layers of the visual encoder, the fully connected mapping layer, and the text decoder, fixing the parameters of the word segmenter, the visual encoder, the text decoder, and the fully connected mapping layer other than the low-rank adapter, and performing backpropagation based on the first cross-entropy loss to update the parameters of the low-rank adapter to obtain an initial image analysis model. .
[0043] The visual encoder can be a module used to decompose cell images, extract deep visual features such as cell texture and contours, and perform visual feature encoding. The fully connected mapping layer can be a mapping layer that aligns visual feature dimensions based on image visual features and matches the input dimension of the text decoder. The word segmenter can be a module used to decompose the input instruction text into the smallest unit of word that the model can compute, and perform text encoding. The low-rank adapter can be a fine-tuning parameter component embedded in multiple layers of the visual encoder, fully connected mapping layer, and text decoder.
[0044] The first instruction set includes multiple first instructions. First instructions can be extracted from this set. The sample cell image and the first instructions are input into the deep learning model. Based on the first instructions, the sample cell image is analyzed, and a predicted cell morphology description is output step-by-step. The predicted probability values of corresponding terms at multiple time steps in the predicted cell morphology description are obtained. The predicted cell morphology description can be the cell morphology description output by the deep learning model. The predicted probability value indicates the confidence probability of a term output by the model at a single time step; a higher value indicates higher accuracy in the model's judgment of that term.
[0045] Simultaneously, the labeled cell morphology descriptions corresponding to the sample cell images can be input into the deep learning model to obtain a labeled description lexical sequence. This labeled description lexical sequence can be a standardized lexical sequence constructed by the word segmenter in the deep learning model after splitting the labeled cell morphology descriptions, and it can be used to provide a standard answer lexical sequence for model training.
[0046] Furthermore, the first cross-entropy loss can be calculated based on the predicted probability values of the corresponding words at multiple time steps in the predicted cell morphology description and the corresponding words at multiple time steps in the labeled description word sequence. The low-rank adapter is then added to multiple layers of the visual encoder, the fully connected mapping layer, and the text decoder. The parameters of the word segmenter, as well as the visual encoder, the text decoder, and the fully connected mapping layer, are fixed except for the low-rank adapter. Only the parameters of the low-rank adapter are retained for iterative updates. Backpropagation is performed based on the first cross-entropy loss to update the parameters of the low-rank adapter, thereby reducing the model prediction error and obtaining the initial image analysis model.
[0047] Based on the above scheme, optionally, the step of inputting the sample cell image and its corresponding labeled cell morphology description and the first instruction into the deep learning model to obtain the labeled description lexical sequence and the predicted probability values of lexical units corresponding to multiple time steps in the predicted cell morphology description includes: inputting the sample cell image into a visual encoder for image segmentation to obtain an initial visual label sequence; inputting the initial visual label sequence into a fully connected mapping layer for feature mapping to obtain a visual label sequence; inputting the first instruction into a word segmenter for word segmentation to obtain a first text label sequence, and inputting the labeled cell morphology description into the word segmenter for word segmentation to obtain a labeled description lexical sequence; inputting the visual label sequence and the first text label sequence into a text decoder to obtain the predicted probability values of lexical units corresponding to multiple time steps in the predicted cell morphology description.
[0048] The initial visual marker sequence can be the original image feature encoding sequence extracted by the visual encoder after image segmentation, without dimensionality adaptation transformation. Alternatively, the visual marker sequence can be the image feature encoding sequence after transformation by a fully connected mapping layer. The first text marker sequence can be the word encoding sequence generated after word segmentation of the first instruction.
[0049] Specifically, the sample cell image can be input into the visual encoder to first perform native dynamic resolution processing on the input image, dividing the image into grid blocks of patch_size=14, extracting local features of the cell nucleus and cytoplasm, and obtaining a preliminary initial visual label sequence. Then, the initial visual label sequence is input into a fully connected mapping layer, such as a multilayer perceptron (MLP) compression module, which maps it to the text feature space through the MLP layer to obtain the visual label sequence. At the same time, the first instruction can be input into the word segmenter for word segmentation processing to obtain the first text label sequence.
[0050] Furthermore, the visual marker sequence and the first text marker sequence can be directly concatenated at the input space level to jointly construct a unified multimodal input sequence, which can then be embedded as input data and sent to the text decoder.
[0051] For positional encoding, the model employs Multimodal Rotary Position Embedding (M-RoPE), decomposing the rotational embedding into three independent components: time, height, and width. The first text tag sequence is represented using one-dimensional positional encoding, while the visual tag sequence is assigned two-dimensional height and width IDs based on its actual spatial position in the image. For video frames, a time ID is additionally added, thus unifying the encoding of spatial and temporal relationships across different modalities in text, images, and videos. It is important to note that cytology does not have a video modality. Since the input is a static pathological cell image with no dynamic image requirements, a time ID is not added to reduce computational overhead. Finally, a text decoder, based on a causal language modeling framework, uses the concatenated multimodal input as a prefix and generates the target text output token-by-token in an autoregressive manner, obtaining the predicted probability values of the corresponding terms at multiple time steps in the predicted cell morphology description.
[0052] Optionally, the visual encoder extracts image features, which are then concatenated with the text embedding generated by the first instruction after passing through a fully connected mapping layer. These features are then input into the text decoder to generate an output sequence in an autoregressive manner. Alternatively, the labeled cell morphology description can be input into the word segmenter for word segmentation to obtain a labeled description word sequence. This labeled description word sequence is then used as a supervision signal to calculate the autoregressive cross-entropy loss of the answer part.
[0053] Based on the above scheme, optionally, the step of calculating the first cross-entropy loss based on the predicted probability values of the corresponding words at multiple time steps in the predicted cell morphology description and the corresponding words at multiple time steps in the labeled description word sequence includes: ; in, denoted as the first cross-entropy loss, used to measure the difference between the word sequence corresponding to the predicted cell morphology description generated by the model and the labeled description word sequence; This indicates the number of terms included in the labeled descriptive term sequence, i.e., the total number of terms obtained after the labeled cell morphology description is processed and split by the model; Indicates the first The predicted probability value of the word at each time step. It can be the index of the lexical position currently being calculated; Indicates the first term in the labeled descriptive lexical sequence. The terminology corresponding to the time step is used in conjunction with the first term in the predicted cell morphology description. The corresponding lexical units at each time step are compared; Represents a sequence of visual markers; Indicates the first text tag sequence; Indicates the first denominator in the sequence of labeled descriptive terms. All lexical units prior to a given time step are used to provide context for the model to generate lexical units.
[0054] In this process, the token loss of the input instruction part is masked, and the model is only trained to learn to generate morphological descriptions. Each fully connected layer in the visual encoder, fully connected mapping layer and text decoder is injected with a LoRA adapter, and then the cross-entropy loss is backpropagated to update the LoRA parameters of all fully connected layers in the model, i.e. LoRA fine-tuning. The backbone parameters of the model's visual encoder, text decoder and word segmenter are frozen throughout the process to ensure training stability.
[0055] S230. The initial image analysis model is trained based on multiple sample cell images and their corresponding labeled cell morphology descriptions, labeled lesion categories, and a second set of instructions to obtain an image analysis model; wherein, the second set of instructions is used to train the initial image analysis model to generate lesion categories corresponding to the region labeling information in the sample cell images based on the cell morphology descriptions.
[0056] Optionally, the initial image analysis model can be trained based on multiple sample cell images and their corresponding labeled cell morphological descriptions, labeled lesion categories, and a second set of instructions to obtain an image analysis model.
[0057] Based on the above scheme, optionally, the step of training the initial image analysis model based on multiple sample cell images and their corresponding labeled cell morphology descriptions, labeled lesion categories, and a second set of instructions to obtain an image analysis model includes: constructing thought chain reasoning text based on the labeled cell morphology descriptions and labeled lesion categories corresponding to the sample cell images; wherein, the thought chain reasoning text is used to describe the processing logic of the image analysis model on the sample cell images; and training the initial image analysis model based on multiple sample cell images and their corresponding thought chain reasoning text and a second set of instructions to obtain an image analysis model.
[0058] The reasoning text in the thought chain can be text used to describe the processing logic of the image analysis model on the sample cell image. For example, it may include reasoning link text from cell morphology description to matching the corresponding lesion type.
[0059] Specifically, a thought chain reasoning text can be constructed based on the labeled cell morphology description and labeled lesion category corresponding to the sample cell image. An instruction is randomly sampled from the second instruction set, and the image features of the sample cell image are spliced together as the context. The structured thought chain reasoning text is used as the answer to construct a single-round instruction-answer dialogue data. Among them, the thought chain reasoning text is a hierarchical thought chain format that conforms to the pre-set cell image analysis standard to train the image analysis model to reason step by step to obtain image description information. The model is forced to output the following reasoning steps in sequence: (1) evaluation of cell nuclear atypia (nuclear enlargement, irregular nuclear membrane, chromatin characteristics); (2) analysis of cytoplasmic characteristics and nuclear-cytoplasmic ratio; (3) identification of cell arrangement structure and background components (inflammatory cells, flora, mucus, etc.); (4) identification analysis of key points (such as "whether the perinuclear halo is typical" and "whether atrophic changes are excluded"); (5) final classification (such as NILM, ASC-US, LSIL, HSIL, etc.) to obtain the cell morphology description corresponding to the region labeling information and the lesion category corresponding to the cell morphology description.
[0060] Compared to the initial image analysis model, the above steps add a reasoning summary from cell morphological description of cell morphological features to lesion category. Simultaneously, based on the initial image analysis model, it can be trained using multiple sample cell images and their corresponding thought chain reasoning text and second instruction set. Fine-tuning is performed using autoregressive loss (cross-entropy loss) combined with a low-rank adapter, freezing the main parameters, updating the LoRA parameters injected into all fully connected layers, and directly supervising the output complete reasoning text. This allows the model to explicitly learn the causal reasoning process from morphological microscopic features to macroscopic image analysis conclusions, rather than simply outputting lesion category labels. By using morphological descriptions as an additional supervisory signal to constrain the model to first perceive morphological features and using these features as a prerequisite for inputting classification labels, the model's discriminative ability is improved.
[0061] Based on the above scheme, optionally, the initial image analysis model includes a visual encoder, a fully connected mapping layer, a word segmenter, and a text decoder. Multiple layers of the visual encoder, fully connected mapping layer, and text decoder include low-rank adapters. The second instruction set includes multiple second instructions. Training the initial image analysis model based on multiple sample cell images and their corresponding thought chain reasoning texts and the second instruction set to obtain an image analysis model includes: inputting sample cell images and their corresponding thought chain reasoning texts and second instructions into the initial image analysis model to obtain the predicted probability values of thought chain reasoning word sequences and corresponding words at multiple time steps in the predicted image description information; wherein the thought chain reasoning word sequence is determined based on the thought chain reasoning text; calculating a second cross-entropy loss based on the predicted probability values of corresponding words at multiple time steps in the predicted image description information and the corresponding words at multiple time steps in the thought chain reasoning word sequence; fixing the parameters of the word segmenter, and the visual encoder, the text decoder, and the fully connected mapping layer (excluding the low-rank adapter), and performing backpropagation based on the second cross-entropy loss to update the parameters of the low-rank adapter to obtain the image analysis model.
[0062] The second instruction can be a training instruction for the model to generate the corresponding lesion category based on the cell morphology description. The thought chain reasoning word sequence can be the standard answer word sequence obtained after the thought chain reasoning text is split by a word segmenter. The predicted image description information can be the descriptive text output by the model during the training phase, containing cell morphology description and lesion category. The second cross-entropy loss can be a loss value used to measure the degree of difference between the word sequence obtained by the model reasoning and the labeled thought chain reasoning word sequence.
[0063] Specifically, the sample cell image and the second instruction can be input into the initial image analysis model. A visual encoder extracts image features, and a word segmenter extracts the text features of the second instruction. The image features are then embedded and concatenated with the text features after passing through a fully connected mapping layer. This concatenation is then input into a text decoder to generate an output sequence in an autoregressive manner, obtaining the predicted probability values of corresponding words at multiple time steps in the predicted image description information. Simultaneously, the word segmenter can also process the thought chain reasoning text to obtain a thought chain reasoning word sequence.
[0064] Furthermore, the lexical sequence corresponding to the predicted image description information can be supervised. Based on the predicted probability values of the lexical corresponding to multiple time steps in the predicted image description information, and the lexical corresponding to multiple time steps in the thought chain reasoning lexical sequence, the second cross-entropy loss is calculated. The parameters of the fixed word segmenter, as well as the visual encoder, text decoder and fully connected mapping layer, except for the low-rank adapter, are backpropagated based on the second cross-entropy loss to update the parameters of the low-rank adapter, so as to obtain the image analysis model.
[0065] Based on the above scheme, optionally, the training process of the image analysis model further includes: performing augmentation processing on the sample cell image to obtain at least two augmented cell images, and obtaining a second instruction from a second instruction set; for each augmented cell image, inputting the augmented cell image, the labeled cell morphology description corresponding to the sample cell image, and the second instruction into the image analysis model to obtain the labeled description lexical sequence and the predicted probability values of lexical units corresponding to multiple time steps in the predicted cell morphology description; performing loss calculation based on the predicted probability values of lexical units corresponding to multiple time steps corresponding to the at least two augmented cell images and the lexical units corresponding to multiple time steps in the labeled description lexical sequence to obtain a loss value, and adjusting the parameters of the image analysis model based on the loss value to obtain an adjusted image analysis model.
[0066] The augmentation process can be a pixel-enhancing operation on the cell image, including but not limited to fine-tuning image brightness, fine-tuning staining hue, local minor cropping, Gaussian blurring, and minor angular rotation, without altering the true morphology and lesion properties of the cells. The augmented cell image can be a cell image obtained by applying at least one augmentation process to a sample cell image.
[0067] Optionally, during training, an inference consistency regularization term, i.e., relative entropy (Kullback-Leibler Divergence, KL divergence), can be added to the autoregressive loss to construct a constraint loss. This allows for the application of two different augmentation processes to the same sample cell image to obtain two augmented cell images, constraining the model's output to have similar prediction distributions at each step. Image augmentation methods can include adaptive saturation adjustment, adaptive contrast adjustment, brightness fine-tuning, pixel histogram normalization, and simulation of imaging sensitivity noise, among others. This regularization term constrains the model to maintain consistency in its inference path and decision logic across different augmented images, suppressing prediction fluctuations caused by local interference or irrelevant changes.
[0068] Specifically, assuming there exists a sample cell image The corresponding thought chain reasoning word sequence is ,in , To determine the number of lexical units included in the thought chain reasoning lexical sequence, the sample cell image is used. Two independent random data augmentation processes were applied to obtain two augmented cell images. and At time step Two augmented cell images are each input to a visual encoder to extract image features. After passing through a fully connected mapping layer, these features are randomly sampled with the second instruction set to obtain the text features and prefixes generated by the second instruction processing. The text embeddings are concatenated and input into the text decoder, which outputs a vector of logistic units (logits). After softmax normalization, the probability distribution for predicting tokens in the next step is obtained. .
[0069] ; in, Indicates that in the known prefix and visual input Under the condition that the model predicts the first Each time step generates word elements The probability value; Indicates the first The first time step The logits value corresponding to each term; Used to convert logits values to non-negative numbers; This represents the set of all possible lexical terms that the model can predict.
[0070] Finally, for the two augmented cell images, the predicted probability distribution value for each step is obtained. and Next, loss is calculated based on the predicted probability values of the corresponding words at multiple time steps in at least two augmented cell images and the corresponding words at multiple time steps in the labeled descriptive word sequence to obtain loss values. Then, the parameters of the image analysis model are adjusted based on the loss values to obtain the adjusted image analysis model.
[0071] Based on the above scheme, optionally, the loss calculation based on the predicted probability values of words corresponding to multiple time steps in at least two augmented cell images and the words corresponding to multiple time steps in the labeled descriptive word sequence to obtain the loss value includes: ; ; ; in, Indicates the loss value; This represents the first loss value, used to measure the morphological term prediction error in augmented cell images; This represents the second loss value, used to constrain the consistency of the output probability distribution of multiple augmented images from the same source; Indicates hyperparameters, ; This indicates the sequence length of the labeled descriptive lexical sequence; Indicates based on augmented cell images In the The prediction at time step is the first time step in the labeled descriptive lexical sequence. The probability value of the word corresponding to each time step; Indicates based on augmented cell images In the The prediction at time step is the first time step in the labeled descriptive lexical sequence. The probability value of the word corresponding to each time step; Represents relative entropy, or KL divergence, used to calculate the difference in word prediction distribution between two augmented images from the same source, thus avoiding model augmentation prediction bias. Indicates based on augmented cell images In the The predicted distribution at each time step; Indicates based on augmented cell images In the The predicted distribution at each time step.
[0072] S240, Receive image analysis instructions and acquire cell images; wherein, the cell images include region annotation information, the region annotation information is used to indicate the image regions to be detected in the cell images, and the image analysis instructions are used to instruct the image analysis model to perform thought chain reasoning output content based on the region annotation information in the cell images.
[0073] S250. Input the cell image and the image analysis command into the image analysis model to obtain image description information corresponding to the cell image; wherein, the image description information is at least used to indicate the cell morphology description corresponding to the region annotation information and the lesion category corresponding to the cell morphology description.
[0074] In this embodiment of the invention, firstly, multiple sample cell images and corresponding sample training sets are acquired. The sample training set includes labeled cell morphological descriptions, labeled lesion categories, a first instruction set, and a second instruction set corresponding to the region annotation information in the sample cell images. Constructing the sample training set ensures the integrity and standardization of the training data, achieving the association and binding of image spatial location annotations, cell morphological text descriptions, lesion classification labels, and training instructions. Next, a deep learning model is trained based on the multiple sample cell images and their corresponding labeled cell morphological descriptions and the first instruction set to obtain an initial image analysis model. The first instruction set is used to train the deep learning model to generate cell morphological descriptions corresponding to the region annotation information in the sample cell images. This enables the model to learn to extract fine-grained visual features such as texture, structure, and size from the region annotation information, establishing a mapping relationship between cell visual features and natural language morphological descriptions. Compared to multi-task joint training, training the morphological description branch separately first avoids the gradient interference of the classification task in the visual feature extraction process, making the model more efficient. The model fully learns subtle pathological features of cells, effectively improving the richness and accuracy of subsequent cell morphology descriptions, providing a reliable feature foundation for the next stage of lesion category prediction. Finally, the initial image analysis model is trained based on multiple sample cell images and their corresponding labeled cell morphology descriptions, labeled lesion categories, and a second set of instructions to obtain an image analysis model. The second set of instructions is used to train the initial image analysis model to generate lesion categories corresponding to the region annotation information in the sample cell images based on the cell morphology descriptions. Based on the morphologically pre-trained initial model, the model is trained by constraining the learning of cell morphology description features to lesion categories through the second set of instructions. This fully reuses the learned fine-grained visual features and semantic representation capabilities of cells, eliminating the need to retrain the feature extraction network from scratch, significantly reducing training computational consumption and convergence time. At the same time, by using morphological descriptions as an intermediate reasoning link, the model simulates the analytical logic of a pathologist who first observes cell morphology and then determines the lesion type, making the model's classification reasoning process interpretable and improving the accuracy of lesion category recognition.
[0075] Example 3 Figure 9 This is a schematic diagram of a cell image analysis device based on morphological thought chain, provided in Embodiment 3 of the present invention. This device is used to execute the cell image analysis method based on morphological thought chain provided in any of the above embodiments. This device and the cell image analysis method based on morphological thought chain in the above embodiments belong to the same inventive concept. Details not described in detail in the embodiments of the cell image analysis device based on morphological thought chain can be referred to the embodiments of the cell image analysis method based on morphological thought chain described above. Figure 9As shown, the device includes an image acquisition module 310 and an image processing module 320.
[0076] The image acquisition module 310 is used to receive image analysis instructions and acquire cell images. The cell images include region annotation information, which indicates the image regions to be detected in the cell images. The image analysis instructions instruct the image analysis model to perform thought chain reasoning based on the region annotation information in the cell images. The image processing module 320 is used to input the cell images and the image analysis instructions into the image analysis model to obtain image description information corresponding to the cell images. The image description information at least indicates the cell morphological description corresponding to the region annotation information and the lesion category corresponding to the cell morphological description.
[0077] The technical solution of this invention embodiment firstly involves receiving an image analysis instruction through an image acquisition module 310, and then acquiring a cell image. The cell image includes region annotation information, which indicates the image region to be detected within the cell image. The image analysis instruction instructs the image analysis model to perform thought chain reasoning based on the region annotation information in the cell image. The image analysis instruction constrains the output direction of the model, ensuring the directionality of image processing. Acquiring the cell image and its region annotation information accurately determines the target region to be analyzed within the cell image, avoiding interference from irrelevant background bands, reducing the amount of invalid data computation, and improving image processing efficiency. Next, the image processing module... The processing module 320 inputs the cell image and the image analysis command into the image analysis model to obtain image description information corresponding to the cell image; wherein, the image description information is at least used to indicate the cell morphology description corresponding to the region annotation information and the lesion category corresponding to the cell morphology description; the image analysis model realizes the accurate extraction of cell image features, outputs cell morphology description and corresponding lesion category, avoids the limitation of classification models in related technologies that only output a single lesion category label, not only outputs the lesion judgment result, but also outputs the cell morphology details on which the model judgment is based, which makes it easier for relevant personnel to trace the classification judgment logic and realize reliable classification of cell images.
[0078] Optionally, based on the above scheme, the device further includes a first model training module, wherein the first model training module includes a training set acquisition submodule, a first model training submodule, and a second model training submodule. The training set acquisition submodule is used to acquire multiple sample cell images and corresponding sample training sets; wherein the sample training set includes labeled cell morphology descriptions, labeled lesion categories, a first instruction set, and a second instruction set corresponding to region labeling information in the sample cell images; the first model training submodule is used to train a deep learning model based on the multiple sample cell images and their corresponding labeled cell morphology descriptions and the first instruction set to obtain an initial image analysis model; wherein the first instruction set is used to train the deep learning model to generate cell morphology descriptions corresponding to region labeling information in the sample cell images; the second model training submodule is used to train the initial image analysis model based on the multiple sample cell images and their corresponding labeled cell morphology descriptions, labeled lesion categories, and the second instruction set to obtain an image analysis model; wherein the second instruction set is used to train the initial image analysis model to generate lesion categories corresponding to region labeling information in the sample cell images based on the cell morphology descriptions.
[0079] Based on the above scheme, optionally, the deep learning model includes a visual encoder, a fully connected mapping layer, a word segmenter, and a text decoder; the first instruction set includes multiple first instructions; the first model training submodule includes a prediction probability value determination unit, a loss calculation unit, and a parameter adjustment unit. The system includes a prediction probability determination unit, which inputs a sample cell image, its corresponding labeled cell morphology description, and a first instruction into the deep learning model to obtain prediction probability values for the labeled description lexical sequence and the corresponding lexical units at multiple time steps in the predicted cell morphology description; wherein the labeled description lexical sequence is determined based on the labeled cell morphology description; a loss calculation unit, which calculates a first cross-entropy loss based on the prediction probability values of the corresponding lexical units at multiple time steps in the predicted cell morphology description and the corresponding lexical units at multiple time steps in the labeled description lexical sequence; and a parameter adjustment unit, which adds a low-rank adapter to multiple layers of the visual encoder, the fully connected mapping layer, and the text decoder, fixes the parameters of the word segmenter, the visual encoder, the text decoder, and the fully connected mapping layer except for the low-rank adapter, and performs backpropagation based on the first cross-entropy loss to update the parameters of the low-rank adapter to obtain an initial image analysis model.
[0080] Based on the above scheme, optionally, the prediction probability value determination unit includes an image segmentation subunit, a feature mapping subunit, a word segmentation subunit, and a prediction probability value determination subunit. Specifically, the image segmentation subunit is used to input the sample cell image into a visual encoder for image segmentation to obtain an initial visual label sequence; the feature mapping subunit is used to input the initial visual label sequence into a fully connected mapping layer for feature mapping to obtain a visual label sequence; the word segmentation subunit is used to input a first instruction into a word segmenter for word segmentation processing to obtain a first text label sequence, and input the labeled cell morphology description into the word segmenter for word segmentation processing to obtain a labeled description lexical sequence; the prediction probability value determination subunit is used to input the visual label sequence and the first text label sequence into a text decoder to obtain prediction probability values for lexical units corresponding to multiple time steps in the predicted cell morphology description.
[0081] Based on the above scheme, optionally, the loss calculation unit is used to calculate the first cross-entropy loss according to the following formula, based on the predicted probability values of the corresponding words at multiple time steps in the predicted cell morphology description and the corresponding words at multiple time steps in the labeled description word sequence: ; in, This represents the first cross-entropy loss; This indicates the number of terms included in the labeled descriptive term sequence; Indicates the first The predicted probability value corresponding to the word at each time step; This indicates the first term in the labeled descriptive lexical sequence. Each time step corresponds to a word element; Represents a sequence of visual markers; Indicates the first text tag sequence; Indicates the first denominator in the sequence of labeled descriptive terms. All lexical units prior to the given time step.
[0082] Based on the above scheme, optionally, the second model training submodule includes an inference text construction unit and a model training unit. The inference text construction unit is used to construct a thought chain inference text based on the labeled cell morphology descriptions and labeled lesion categories corresponding to the sample cell images; wherein the thought chain inference text describes the processing logic of the image analysis model on the sample cell images; the model training unit is used to train the initial image analysis model based on multiple sample cell images and their corresponding thought chain inference texts and a second instruction set to obtain the image analysis model.
[0083] Based on the above scheme, optionally, the initial image analysis model includes a visual encoder, a fully connected mapping layer, a word segmenter, and a text decoder, wherein multiple layers of the visual encoder, the fully connected mapping layer, and the text decoder include low-rank adapters; the second instruction set includes multiple second instructions; the model training unit includes a prediction probability value determination subunit, a loss calculation subunit, and a parameter adjustment subunit. The model includes a prediction probability determination subunit, which inputs the sample cell image and its corresponding thought chain reasoning text and second instruction into the initial image analysis model to obtain the prediction probability values of the thought chain reasoning lexical sequence and the corresponding lexical units at multiple time steps in the predicted image description information; wherein the thought chain reasoning lexical sequence is determined based on the thought chain reasoning text; a loss calculation subunit, which calculates the second cross-entropy loss based on the prediction probability values of the corresponding lexical units at multiple time steps in the predicted image description information and the corresponding lexical units at multiple time steps in the thought chain reasoning lexical sequence; and a parameter adjustment subunit, which fixes the parameters of the word segmenter, the visual encoder, the text decoder, and the fully connected mapping layer, except for the low-rank adapter, and performs backpropagation based on the second cross-entropy loss to update the parameters of the low-rank adapter to obtain the image analysis model.
[0084] Optionally, based on the above scheme, the device further includes a second model training module, wherein the second model training module includes an image augmentation submodule, a prediction probability value determination submodule, and a parameter adjustment submodule. The image augmentation submodule is used to perform augmentation processing based on the sample cell image to obtain at least two augmented cell images and obtain a second instruction from a second instruction set. The prediction probability value determination submodule is used to input the augmented cell image, the corresponding labeled cell morphology description, and the second instruction into the image analysis model for each augmented cell image to obtain the predicted probability values of the labeled description term sequence and the corresponding terms at multiple time steps in the predicted cell morphology description. The parameter adjustment submodule is used to perform loss calculation based on the predicted probability values of the corresponding terms at multiple time steps corresponding to the at least two augmented cell images and the corresponding terms at multiple time steps in the labeled description term sequence to obtain a loss value, and to adjust the parameters of the image analysis model based on the loss value to obtain an adjusted image analysis model.
[0085] Based on the above scheme, optionally, the parameter adjustment submodule is used to perform loss calculation according to the following formula, based on the predicted probability values of words corresponding to multiple time steps of at least two augmented cell images and words corresponding to multiple time steps in the labeled descriptive word sequence, to obtain the loss value: ; ; ; in, Indicates the loss value; This represents the first loss value; Indicates the second loss value; Indicates hyperparameters, ; This indicates the sequence length of the labeled descriptive lexical sequence; Indicates based on augmented cell images In the The prediction at time step is the first time step in the labeled descriptive lexical sequence. The probability value of the word corresponding to each time step; Indicates based on augmented cell images In the The prediction at time step is the first time step in the labeled descriptive lexical sequence. The probability value of the word corresponding to each time step; Represents relative entropy; Indicates based on augmented cell images In the The predicted distribution at each time step; Indicates based on augmented cell images In the The predicted distribution at each time step.
[0086] The cell image analysis device based on morphological thought chain provided in the embodiments of the present invention can execute the cell image analysis method based on morphological thought chain provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the method execution.
[0087] Example 4 Figure 10 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0088] like Figure 10As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0089] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0090] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as cell image analysis methods based on morphological thought chains.
[0091] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication unit 19, or installed from storage unit 18, or installed from ROM 12. When the computer program is executed by processor 11, it performs the functions defined in the methods of the embodiments of the present invention.
[0092] In some embodiments, the morphological thought chain-based cell image analysis method can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the morphological thought chain-based cell image analysis method described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the morphological thought chain-based cell image analysis method by any other suitable means (e.g., by means of firmware).
[0093] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0094] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0095] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0096] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0097] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0098] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0099] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0100] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A cell image analysis method based on morphological thought chain, characterized in that, include: The system receives an image analysis instruction and acquires a cell image; wherein the cell image includes region annotation information, the region annotation information is used to indicate the image region to be detected in the cell image, and the image analysis instruction is used to instruct the image analysis model to perform thought chain reasoning output based on the region annotation information in the cell image; The cell image and the image analysis instructions are input into the image analysis model to obtain image description information corresponding to the cell image; wherein, the image description information is at least used to indicate the cell morphology description corresponding to the region annotation information and the lesion category corresponding to the cell morphology description.
2. The cell image analysis method based on morphological thought chain according to claim 1, characterized in that, The training process of the image analysis model includes: Acquire multiple sample cell images and corresponding sample training sets; wherein, the sample training set includes labeled cell morphological descriptions, labeled lesion categories, a first instruction set, and a second instruction set corresponding to the region labeling information in the sample cell images; The deep learning model is trained based on multiple sample cell images and their corresponding labeled cell morphological descriptions and a first set of instructions to obtain an initial image analysis model; wherein, the first set of instructions is used to train the deep learning model to generate cell morphological descriptions corresponding to the region labeling information in the sample cell images; The initial image analysis model is trained based on multiple sample cell images and their corresponding labeled cell morphology descriptions, labeled lesion categories, and a second set of instructions to obtain an image analysis model; wherein, the second set of instructions is used to train the initial image analysis model to generate lesion categories corresponding to the region labeling information in the sample cell images based on the cell morphology descriptions.
3. The cell image analysis method based on morphological thought chain according to claim 2, characterized in that, The deep learning model includes a visual encoder, a fully connected mapping layer, a word segmenter, and a text decoder; the first instruction set includes multiple first instructions; training the deep learning model based on multiple sample cell images and their corresponding labeled cell morphological descriptions and the first instruction set to obtain an initial image analysis model includes: The sample cell image, its corresponding labeled cell morphology description, and the first instruction are input into the deep learning model to obtain the labeled description lexical sequence and the predicted probability values of the lexical corresponding to multiple time steps in the predicted cell morphology description; wherein, the labeled description lexical sequence is determined based on the labeled cell morphology description; The first cross-entropy loss is calculated based on the predicted probability values of the words corresponding to multiple time steps in the predicted cell morphology description and the words corresponding to multiple time steps in the labeled description word sequence. A low-rank adapter is added to multiple layers of the visual encoder, the fully connected mapping layer, and the text decoder. The parameters of the word segmenter, the visual encoder, the text decoder, and the fully connected mapping layer, except for the low-rank adapter, are fixed. Backpropagation is performed based on the first cross-entropy loss to update the parameters of the low-rank adapter to obtain an initial image analysis model.
4. The cell image analysis method based on morphological thought chain according to claim 3, characterized in that, The step of inputting the sample cell image and its corresponding labeled cell morphology description and the first instruction into the deep learning model to obtain the labeled description lexical sequence and the predicted probability values of the corresponding lexical units at multiple time steps in the predicted cell morphology description includes: The sample cell image is input into a visual encoder for image segmentation to obtain an initial visual label sequence; The initial visual label sequence is input into a fully connected mapping layer for feature mapping to obtain a visual label sequence. The first instruction is input into the word segmenter for word segmentation to obtain the first text tag sequence, and the labeled cell morphology description is input into the word segmenter for word segmentation to obtain the labeled description word sequence. The visual marker sequence and the first text marker sequence are input into a text decoder to obtain the predicted probability values of the corresponding words at multiple time steps in the predicted cell morphology description.
5. The cell image analysis method based on morphological thought chain according to claim 4, characterized in that, The calculation of the first cross-entropy loss based on the predicted probability values of the corresponding words at multiple time steps in the predicted cell morphology description and the corresponding words at multiple time steps in the labeled description word sequence includes: ; in, This represents the first cross-entropy loss; This indicates the number of terms included in the labeled descriptive term sequence; Indicates the first The predicted probability value corresponding to the word at each time step; This indicates the first term in the labeled descriptive lexical sequence. Each time step corresponds to a word element; Represents a sequence of visual markers; Indicates the first text tag sequence; Indicates the first denominator in the sequence of labeled descriptive terms. All lexical units prior to the given time step.
6. The cell image analysis method based on morphological thought chain according to claim 2, characterized in that, The initial image analysis model is trained based on multiple sample cell images and their corresponding labeled cell morphological descriptions, labeled lesion categories, and a second set of instructions to obtain an image analysis model, including: A thought chain reasoning text is constructed based on the labeled cell morphology description and labeled lesion category corresponding to the sample cell image; wherein, the thought chain reasoning text is used to describe the processing logic of the image analysis model on the sample cell image; The initial image analysis model is trained based on multiple sample cell images and their corresponding thought chain reasoning text and second instruction set to obtain the image analysis model.
7. The cell image analysis method based on morphological thought chain according to claim 6, characterized in that, The initial image analysis model includes a visual encoder, a fully connected mapping layer, a word segmenter, and a text decoder. Multiple layers in the visual encoder, fully connected mapping layer, and text decoder include low-rank adapters. The second instruction set includes multiple second instructions. Training the initial image analysis model based on multiple sample cell images and their corresponding thought chain reasoning texts and the second instruction set to obtain an image analysis model includes: The sample cell image and its corresponding thought chain reasoning text and second instruction are input into the initial image analysis model to obtain the predicted probability values of words corresponding to multiple time steps in the thought chain reasoning word sequence and the predicted image description information; wherein, the thought chain reasoning word sequence is determined based on the thought chain reasoning text; The second cross-entropy loss is calculated based on the predicted probability values of the words corresponding to multiple time steps in the predicted image description information and the words corresponding to multiple time steps in the thought chain reasoning word sequence. By fixing the parameters of the word segmenter, the visual encoder, the text decoder, and the fully connected mapping layer, excluding the low-rank adapter, backpropagation is performed based on the second cross-entropy loss to update the parameters of the low-rank adapter, thereby obtaining an image analysis model.
8. The cell image analysis method based on morphological thought chain according to claim 2, characterized in that, The training process of the image analysis model also includes: Augmentation processing is performed based on the sample cell images to obtain at least two augmented cell images, and a second instruction is obtained from the second instruction set; For each augmented cell image, the augmented cell image and the labeled cell morphology description corresponding to the sample cell image and the second instruction are input into the image analysis model to obtain the labeled description word sequence and the predicted probability values of words corresponding to multiple time steps in the predicted cell morphology description; Loss is calculated based on the predicted probability values of words corresponding to multiple time steps in at least two augmented cell images and words corresponding to multiple time steps in the labeled descriptive word sequence to obtain a loss value. The parameters of the image analysis model are then adjusted based on the loss value to obtain an adjusted image analysis model.
9. The cell image analysis method based on morphological thought chain according to claim 8, characterized in that, The loss value is obtained by calculating the predicted probability values of words corresponding to multiple time steps in at least two augmented cell images and the words corresponding to multiple time steps in the labeled descriptive word sequence. include: ; ; ; in, Indicates the loss value; This represents the first loss value; Indicates the second loss value; Indicates hyperparameters, ; This indicates the sequence length of the labeled descriptive lexical sequence; Indicates based on augmented cell images In the The prediction at time step is the first time step in the labeled descriptive lexical sequence. The probability value of the word corresponding to each time step; Indicates based on augmented cell images In the The prediction at time step is the first time step in the labeled descriptive lexical sequence. The probability value of the word corresponding to each time step; Represents relative entropy; Indicates based on augmented cell images In the The predicted distribution at each time step; Indicates based on augmented cell images In the The predicted distribution at each time step.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the cell image analysis method based on morphological thought chain as described in any one of claims 1-9.