Multi-mode-based training method and system for cervical pathology image classification model
By constructing a three-distribution framework and environmental field simulation for a multimodal cervical pathology image classification model, the problem of insufficient structural dynamic modeling in cervical pathology image analysis in existing technologies is solved, and efficient and stable classification and cross-center deployment of cervical pathology images are achieved.
Patent Information
- Application Number
- CN202510893404.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-06-30
AI Technical Summary
Existing technologies struggle to accurately capture the evolution of tissue structures in cervical pathology image analysis. They lack the ability to dynamically model the structural distribution relationships between different tissue levels, resulting in weak task transferability, limited generalization ability, and an inability to provide a temporally consistent processing path, which can easily lead to unstable cross-center performance.
By training a multimodal cervical pathology image classification model, a three-distribution framework of nucleus-stromal-epithelium is constructed. Lifelines are identified using historical data, a simulated environmental field is generated, and hierarchical evolution prediction is performed. Combined with a visual basic model, image-text fusion modeling is carried out to achieve cross-center deployment.
It improves the structural decoupling capability and temporal consistency of cervical pathology image analysis, enhances the stability and controllability of the model under different data domains and task label structures, and ensures the consistency of image detail texture information and the accuracy of structural deformation.
Smart Images

Figure CN120808067A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image classification, in particular to a training method and system of a cervical pathological image classification model based on multi-modal. BACKGROUND
[0002] Traditional cervical pathological image analysis methods mostly rely on shallow feature extraction of static images and single-label classification, which is difficult to accurately capture the evolution process of tissue structure in actual multi-stage pathological tasks. The existing technology generally uses a single-modal visual network to analyze the whole image of the slice image, which cannot distinguish the structural distribution relationship between different tissue levels, and lacks dynamic modeling capability for fine-grained regions such as nuclear dense areas and epithelial boundaries. In the evolution process of cervical lesions, there is a lack of modeling logic for evolution trends, which can easily cause stage classification ambiguity and structural recognition error accumulation. Especially when facing multi-center image sources, multi-task type fusion and structural morphology diversification of real data, the existing technology cannot provide a time sequence consistency processing path, nor can it use text and image fusion to guide the task for non-structured description. Most systems lack modeling and response mechanisms for development processes, evolution trends and regional level differences, and it is difficult to generate reasonable hypothesis models for rare pathological states or unknown evolution paths, resulting in weak task migration ability and limited generalization ability. After system deployment, it is prone to unstable performance across centers. SUMMARY
[0003] Therefore, it is necessary to provide a training method and system of a cervical pathological image classification model based on multi-modal to solve at least one of the above technical problems.
[0004] To achieve the above purpose, the training method of the cervical pathological image classification model based on multi-modal includes the following steps:
[0005] Step S1: obtaining a cervical tissue image; performing tissue structure segmentation on the cervical tissue image to obtain a segmented cervical structure block; and dividing the cervical tissue into a three-distribution framework of nucleus-stromal-epithelium based on the segmented cervical structure block;
[0006] Step S2: collecting cervical tissue development history data; identifying the development lifelines of the three-distribution framework based on the cervical tissue development history data; and confirming the predicted development trends of each layer in the three-distribution framework through the three-distribution framework development lifelines;
[0007] Step S3: performing cervical tissue environment field simulation through the cervical tissue image to generate a simulated cervical environment field; and performing layered evolution prediction on the three-distribution framework of nucleus-stromal-epithelium based on the simulated cervical environment field and the predicted development trends of each layer to obtain layered evolution prediction data of each layer framework;
[0008] Step S4: Evolution projection is performed on the cervical tissue image according to the layer framework evolution prediction data, so as to generate an evolution mapping cervical tissue image; and the cervical tissue image is classified according to the evolution process category based on the evolution mapping cervical tissue image, so as to generate a cervical classification image.
[0009] Step S5: Hierarchical joint modeling training is performed according to the cervical classification image, so as to obtain a visual basic model; graph-text fusion modeling is performed based on the visual basic model, task migration processing is performed, and the system is deployed to an integrated system, so as to generate a cross-center deployment model system.
[0010] The application can construct highly structured spatial information representation in the original pixel space of the image, enhance the structure decoupling and multi-level channel modeling capability in the subsequent processing, ensure that different types of tissue structures have traceable and separable high-resolution expression basis in the data processing flow, use the development history data to perform time sequence fusion on the structure distribution and identify the three-distribution evolution life line, reconstruct the dynamic evolution path of the spatial stable structure in the non-continuous image frame, effectively complete the time sequence consistency alignment processing of the three-distribution structure, make the tissue evolution trajectory have data continuity and structure resolution, further extract the hierarchical development trend of each layer structure, provide directional change signal for subsequent tasks, improve the direction fitting accuracy and deformation rationality in the distribution prediction stage, construct an environment field simulation model based on the cervical tissue image, complete the composite generation of the tension disturbance field, the structure pressure projection field and the trend expansion array through the non-dynamic mechanism, ensure that the structure simulation data consider the realism, heterogeneous disturbance coverage and spatial structure controllability, drive the evolution prediction processing of the three-distribution framework through the environment field data and the trend data, can introduce fitting evolution disturbance flow while keeping the continuity of the original structure organization boundary, make the simulation evolution results of each layer structure have direction guiding property and feature fidelity, map the hierarchical prediction results to the original image domain and generate evolution mapping images, can complete the full-process back-projection conversion of the structure deformation and modal disturbance in the high-resolution image, keep the consistency of the image detail texture information and the channel distribution structure, introduce the evolution process category on this basis for image classification, can convert the potential trend change into image-level structure class label output, provide high-confidence input source for subsequent classification supervision, perform hierarchical joint modeling training based on the classified images, which is beneficial to construct a visual basic model with structure evolution perception ability, and through the fusion training of the image-text input, the encoding ability of the multi-modal representation structure on the context semantic clue and the structure dynamic trajectory is enhanced, the task migration processing and multi-center deployment are performed based on the model structure, the structure consistency response and semantic-level generalization distribution feature regulation can be realized under different data domains and different task label structures, so that a unified cross-modal, cross-time sequence and cross-institutional visual-language joint modeling system is constructed, and the stability and controllability of the overall model structure in image space deformation modeling, semantic cross-modal alignment construction and task-level response flow scheduling are improved.
[0011] The application also provides a training system of a cervical pathological image classification model based on multiple modes, which is used for executing the training method of the cervical pathological image classification model based on multiple modes.
[0012] The three-decomposition module is used for acquiring a cervical tissue image; the cervical tissue image is subjected to tissue structure segmentation, so that a segmented cervical structure block is obtained; and the cervical tissue is divided into a three-distribution framework of nucleus-stroma-epithelium based on the segmented cervical structure block;
[0013] The trend identification module is used for collecting cervical tissue development history data; identifying a three-distribution framework development life line based on the cervical tissue development history data; and confirming the predicted development trend of each layer in the three-distribution framework through the three-distribution framework development life line;
[0014] The environment prediction module is used for simulating a cervical tissue environment field through the cervical tissue image, so that a simulated cervical environment field is generated; and performing layered evolution prediction on the three-distribution framework of nucleus-stroma-epithelium based on the simulated cervical environment field and the predicted development trend of each layer, so as to obtain layered framework evolution prediction data;
[0015] The evolution classification module is used for performing evolution projection on the cervical tissue image according to the layered framework evolution prediction data, so that an evolution mapping cervical tissue image is generated; and performing image classification on the cervical tissue image according to an evolution process category based on the evolution mapping cervical tissue image, so that a cervical classification image is generated;
[0016] The fusion modeling module is used for performing hierarchical joint modeling training according to the cervical classification image, so that a visual basic model is obtained; performing graphic-text fusion modeling based on the visual basic model, and performing task migration processing, and being deployed to an integrated system, so that a cross-center deployment model system is generated.
[0017] The three-decomposition module constructs a three-channel structure framework of nucleus-stroma-epithelium, which is beneficial to improve the spatial decoupling capability of tissue partitioning and the structural clarity of subsequent modeling; the trend identification module extracts a structure evolution trajectory through development history data, establishes a structure correspondence relationship between multiple frames of images, enhances the temporal consistency and trend traceability; the environment prediction module generates a simulated environment field containing tension disturbance, pressure mapping and trend array, realizes non-biological mechanism driven modeling of the structure evolution process, improves the directionality and structure fidelity of evolution simulation; the evolution classification module generates a mapping image and divides a process category based on prediction, so that the image classification has trend perception capability and change phase distinguishability; the fusion modeling module constructs a visual model through a classification image and completes graphic-text alignment training, realizes unified modeling of tissue structure, trend semantics and graphic-text modalities, and the system has clear structure, separated modules, supports cross-modal input, task-level scheduling and multi-center deployment, and has expandability, combinability and structural-level universality. BRIEF DESCRIPTION OF DRAWINGS
[0018] Fig. 1 It is a step flowchart of a training method of a cervical pathological image classification model based on multiple modalities;
[0019] Fig. 2 For the detailed implementation step flow chart of step S2;
[0020] The implementation, functional features and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION
[0021] The technical method of the present application will be described clearly and completely below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0022] In addition, the accompanying drawings are only schematic illustrations of the present application and are not necessarily drawn to scale. Identical reference numerals in the drawings represent identical or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities, which do not necessarily have to correspond to physically or logically independent entities. The functional entities can be implemented in the form of software, or in one or more hardware modules or integrated circuits, or in different network and / or processor methods and / or microcontroller methods.
[0023] It should be understood that although the terms "first", "second" and the like can be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the exemplary embodiments, a first element can be referred to as a second element, and similarly a second element can be referred to as a first element. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0024] To achieve the above-mentioned purpose, please refer to Figs. 1-2 A training method of a multi-modal based cervical pathological image classification model, comprising the following steps:
[0025] Step S1: Obtain a cervical tissue image; segment the cervical tissue image according to the tissue structure, thereby obtaining a segmented cervical structure block; and divide the cervical tissue into a three-distribution framework of nucleus-stromal-epithelium based on the segmented cervical structure block;
[0026] Step S2: Collect cervical tissue development history data; identify the three-distribution framework development lifelines based on the cervical tissue development history data; and confirm the predicted development trend of each layer in the three-distribution framework through the three-distribution framework development lifelines;
[0027] Step S3: simulate the cervical tissue environment field through the cervical tissue image, thereby generating a simulated cervical environment field; based on the simulated cervical environment field and the predicted development trend of each layer, perform hierarchical evolution prediction on the three distribution framework of nucleus-stroma-epithelium to obtain each layer framework evolution prediction data;
[0028] Step S4: according to the hierarchical framework evolution prediction data, perform evolution projection on the cervical tissue image, thereby generating an evolution mapping cervical tissue image; based on the evolution mapping cervical tissue image, classify the cervical tissue image according to the evolution process category, thereby generating a cervical classification image;
[0029] Step S5: according to the cervical classification image, perform hierarchical joint modeling training, thereby obtaining a visual basic model; based on the visual basic model, perform image-text fusion modeling and task migration processing, and deploy to an integrated system, thereby generating a cross-center deployment model system.
[0030] The application can construct highly structured spatial information representation in the original pixel space of the image, enhance the structure decoupling and multi-level channel modeling capability in the subsequent processing, ensure that different types of tissue structures have traceable and separable high-resolution expression basis in the data processing flow, use the development history data to perform time sequence fusion on the structure distribution and identify the three-distribution evolution life line, reconstruct the dynamic evolution path of the spatial stable structure in the non-continuous image frame, effectively complete the time sequence consistency alignment processing of the three-distribution structure, make the tissue evolution trajectory have data continuity and structure resolution, further extract the hierarchical development trend of each layer structure, provide directional change signal for subsequent tasks, improve the direction fitting accuracy and deformation rationality in the distribution prediction stage, construct an environment field simulation model based on the cervical tissue image, complete the composite generation of the tension disturbance field, the structure pressure projection field and the trend expansion array through the non-dynamic mechanism, ensure that the structure simulation data consider the realism, heterogeneous disturbance coverage and spatial structure controllability, drive the evolution prediction processing of the three-distribution framework through the environment field data and the trend data, while keeping the continuity of the original structure organization boundary, introduce the fitting evolution disturbance flow, make the simulation evolution results of each layer structure have direction guiding and feature fidelity, map the hierarchical prediction results to the original image domain and generate evolution mapping images, complete the full-process back-projection conversion of the structure deformation and modal disturbance in the high-resolution image, keep the consistency of the image detail texture information and the channel distribution structure, introduce the evolution process category on this basis for image classification, can convert the potential trend change into image-level structure class label output, provide high-confidence input source for subsequent classification supervision, perform hierarchical joint modeling training based on the classified images, which is conducive to constructing a visual basic model with structure evolution perception ability, and through the fusion training of the image-text input, the encoding ability of the multi-modal representation structure on the context semantic clues and the structure dynamic trajectory is enhanced, the task migration processing and multi-center deployment are performed on the basis of the model structure, the structure consistency response and semantic-level generalization distribution feature regulation can be realized under different data domains and different task label structures, so as to construct a unified cross-modal, cross-time sequence and cross-institutional visual-language joint modeling system, and improve the stability and controllability of the overall model structure in image space deformation modeling, semantic cross-modal alignment construction and task-level response flow scheduling.
[0031] In the embodiment of the application, the training method of the cervical pathological image classification model based on multiple modalities comprises the following steps:
[0032] Step S1: obtaining a cervical tissue image; performing tissue structure segmentation on the cervical tissue image to obtain a segmented cervical structure block; and dividing the cervical tissue into a three-distribution framework of nucleus-stromal-epithelial based on the segmented cervical structure block;
[0033] In this embodiment, the original pathological whole slide image (WSI) collected from the multi-center data collaboration platform is taken as the input sample. All the slice images need to be processed by uniform magnification standardization. The scanning magnification is uniformly set to 20 times. The image quantization coding is performed on each WSI image with a scanning resolution of 0.24 μm / pixel. The highest resolution layer of the WSI image is read by using the OpenSlide tool and block-level division is performed. The division window size is set to 2048 pixels x 2048 pixels and the step size is 512 pixels. A pre-trained semantic segmentation model based on the U-Net structure is used to extract the foreground organization region of each image block. The input channel number of the segmentation model is 3 and the output channel number is 4, corresponding to the nuclear region, epithelial region, interstitial region and background region. The output results of each image block are post-processed and optimized by using a fully connected conditional random field (CRF) to obtain a continuous boundary structure mask. Then, the structure mask of each block is divided into a three-channel distribution framework. The nuclear region is mapped to the red channel, the epithelial region is mapped to the green channel, and the interstitial region is mapped to the blue channel. The corresponding three-channel organization distribution tensor image of the block is saved, and an index file is generated to record the patient ID, spatial coordinates and pixel density information in the channel of the block.
[0034] Step S2: Collecting cervical tissue development history data; identifying three distribution framework development life lines based on cervical tissue development history data; confirming the prediction development trend of each layer in the three distribution framework through the three distribution framework development life lines;
[0035] In this embodiment, a case data set containing a follow-up time sequence is constructed. Each case contains not less than three cervical tissue images collected at different time nodes. The image time interval is set to not less than 6 months. Each image must be scanned at 20x magnification and retain the structure integrity region. An image structure hashing algorithm is used to filter the similarity of the images to ensure that each time point image has an organization structure change feature. Image structure matching is performed on each case time sequence. The block-level feature representation of the three distribution framework in each frame image is extracted through a multi-scale feature pyramid structure. A joint matching standard is formed by using the block center coordinates and texture vectors. It is considered that the structure block has temporal consistency when the Euclidean distance is less than 20 pixels and the cosine similarity is greater than 0.85. The structure development life line is constructed according to the spatial trajectory of the structure block. Statistical analysis is performed on the area behavior of the same structure block at three or more time points. The evolution trend is judged by the structure area change slope. The growth slope threshold is set to +0.15, the decline slope threshold is set to -0.15, and the stable range is set to [-0.15, +0.15]. The structure development type is marked as expansion type, contraction type and stable type, respectively. The trend label of each structure block at each time node is recorded and input into the subsequent simulation modeling stage.
[0036] Step S3: simulate the cervical tissue environment field through the cervical tissue image to generate a simulated cervical environment field; and based on the simulated cervical environment field and the predicted development trend of each layer, perform hierarchical evolution prediction on the three-distribution framework of the nucleus-stroma-epithelium to obtain evolution prediction data of each layer framework;
[0037] In this embodiment, a cervical tissue image with clear structure division is selected as input data, and a tiling process is performed thereon, with a tile size of 128 pixels x 128 pixels. The tissue density distribution map and gradient direction map inside each tile are calculated, the Sobel operator is used to extract the edge direction distribution information of the tile, and a tension field simulation map is constructed according to the structure density of the nuclear region. The area with a pixel density higher than 0.4 in the nuclear region is assigned an initial tension value of 0.6, and the tension value of other areas is set to 0.2. The tension field structure is taken as the basic input of the environmental tension feature map. A horizontal structure tensor and a radial structure tensor are injected into each tile respectively. The horizontal tensor takes the x-axis direction of the tile as the value direction, and the intensity is 60% of the tension value. The radial tensor diffuses from the center of the tile to the surrounding area, and the intensity decreases with the distance according to a square inverse function. The maximum gradient direction extraction process is performed on the fusion map of all tensors. The dominant direction vector is calculated based on the tension gradient field, and the residual guided filling operation is performed on the dominant direction. The non-continuous structure direction is interpolated and completed using the convolution kernel. Finally, the tension sub-tiles are output and integrated into the whole simulated cervical environment field map.
[0038] Step S4: perform evolution projection on the cervical tissue image according to the evolution prediction data of each layer framework to generate an evolution mapping cervical tissue image; and based on the evolution mapping cervical tissue image, classify the cervical tissue image according to the evolution process category to generate a cervical classification image.
[0039] In this embodiment, the structure prediction result and the simulated environment field are used as joint input source data to perform direction projection operation on the nuclear region, stroma region and epithelium region of the three-distribution framework respectively. The target translation position of the tile is calculated according to the structure trend vector and the dominant tension direction vector. The maximum translation distance is set to not more than 16 pixels. The affine transformation matrix is used to perform coordinate mapping on the structure tile and rewrite it into the original image coordinate system. The pixel value of the reconstructed image is resampled and smoothed to maintain the boundary continuity and texture consistency. After the mapping is completed, each tile is assigned an evolution process category according to its evolution trend label, including three categories of expansion, contraction and stability. The center area texture density change value is used as the classification basis. The change rate greater than 20% is marked as expansion type, the change rate less than -20% is marked as contraction type, and the absolute change rate less than 5% is marked as stable type. All image blocks are constructed into a training image set according to the labels. The image size is unified to 224 pixels x 224 pixels.
[0040] Step S5: hierarchical joint modeling training is performed according to the cervical classification image, so as to obtain a visual base model; image-text fusion modeling is performed based on the visual base model, task migration processing is performed, and the cross-center deployment model system is generated by deploying to the integrated system.
[0041] In this embodiment, ViT (Vision Transformer) is used as the image encoder backbone structure, each classification image is divided into patches, each patch has a size of 16 pixels x 16 pixels, each image can be divided into 196 patches, each patch is mapped to a vector embedding with a dimension of 768 using linear transformation, and position encoding is added, and input to a 12-layer Transformer module, each layer is configured with 12 self-attention heads, and a DINOv2 self-supervised pre-training mechanism is used to train the base visual encoder CerS-V under the image alignment and feature consistency optimization target, the optimizer uses AdamW, the initial learning rate is set to 5e-4, the training batch is 300 rounds, and each round of training sample is not less than 2048 images, after the training is completed, the visual encoder weight is frozen, the output feature is used as the input of the image-text modeling, the LoRA (Low-Rank Adaptation) method is used to connect the visual output to the multi-modal embedding layer of the large language model Qwen2.5-VL through a 128-dimensional low-rank adapter, and the image-text fusion base model CerS-M is constructed, the instruction fine-tuning is performed on the fusion model using the instruction data set containing structure description, typing label, dialogue question and answer and other task types, the language model parameters are frozen during the training process, only the parameters of the visual connection module and the instruction decoding head are updated, after the model is trained, the multi-modal deployment module is packaged and the API interface is bound to the integrated platform, the platform supports heterogeneous image input, task instruction loading and structure prediction result output, and provides a structured report interface and a system task management console for deployment of hospitals to use, and finally forms a cross-modal structure classification system model which can be deployed and used in a multi-center environment.
[0042] Preferably, step S1 comprises the following steps:
[0043] Step S11: obtaining a cervical tissue image; performing color normalization on the cervical tissue image to obtain a normalized cervical image; and performing deep contrast enhancement on the normalized cervical image to generate an enhanced cervical image.
[0044] Step S12: identifying cervical tissue structure features based on the enhanced cervical image, and performing spatial collaborative segmentation on the enhanced cervical image according to the cervical tissue structure features to obtain a segmented cervical structure block.
[0045] Step S13: performing kernel density mapping processing on the segmented cervical structure block to obtain kernel density distribution data; performing interstitial region separation based on the kernel density distribution data to generate interstitial distribution data;
[0046] Step S14: performing epithelial structure decoupling through the interstitial distribution data to obtain epithelial distribution data; fusing the epithelial distribution data, the kernel density distribution data, and the interstitial distribution data, and constructing a kernel-interstitial-epithelial three-distribution framework based on the fused data.
[0047] In this embodiment, 20 times magnification scanning of the whole cervical section images is derived from multiple medical image centers as the input image source, the resolution of each image is 0.24 microns per pixel, the number of image channels is 3, and the image format is unified as TIFF. First, Macenko normalization algorithm is used for color standardization processing of the input image. Channel matrix decomposition is needed when constructing the color normalization function. The absorption spectrum characteristic value of each channel is extracted. Then the normalization matrix is applied to the original image color space transformation. After the target color standard template mapping is completed, the normalized cervical image is generated. The normalized cervical image is used as input to perform deep contrast enhancement operation. The enhancement process is based on local contrast enhancement algorithm to perform contrast stretching processing within the image block. The image block is set to 64 pixels x 64 pixels, and the stretching factor is set to 1.4. Adaptive local histogram equalization operation is performed on the nuclear region in the image to enhance the gray scale contrast between the nuclear region and the non-nuclear region. Finally, the enhanced cervical image is output as the input data before structure segmentation. The multi-scale tissue structure features are extracted by the image feature extraction model based on the EfficientNet-B3 backbone network. The input image size of the model is 512 pixels x 512 pixels. After feature extraction, the spatial attention mechanism module is constructed. This module constructs a structure weight matrix according to the spatial density of the structure region in the image. The weight matrix has the same dimension as the input image, and the value range is 0 to 1, representing the response degree of the corresponding pixel to the structure region. The structure feature map is input into the multi-channel spatial collaborative segmentation network. The segmentation network adopts the Encoder-Decoder structure based on depth separable convolution. The output channel number is set to 3, corresponding to the prediction map of the nuclear region, epithelial region and interstitial region respectively. After the maximum channel attribution processing of the output result, three types of segmentation maps are generated. The empty convolution smoothing operation is performed on each channel output and combined with the structure feature map for cross supervision learning. The image of the segmented cervical structure block is output. The image format is a three-channel floating point image. Each channel records the response intensity and spatial position index of the structure region. The center point position of all nuclear blocks is extracted based on the nuclear positioning network. The nuclear density map is generated using the heat map construction method based on density estimation. Each nuclear position is used as the center point of the two-dimensional Gaussian function in the calculation formula. The standard deviation is set to 4 pixels. The Gaussian response of all nuclear points is superimposed to generate the nuclear density map. The density map is a single-channel grayscale image with a value range of 0 to 1, representing the nuclear structure density intensity per unit area. Then, the interstitial region is extracted based on the nuclear density map using the density threshold separation method. The density is less than 0.The region of 15 is defined as a mesenchyme candidate region, small-area pseudo-mesenchyme regions are removed in combination with boundary connectivity constraints, a morphological closing operation is used to complete the region discontinuity, and a generated mesenchyme distribution map is a binary mask map, with black regions corresponding to the nucleus region and the epithelial region, and white regions corresponding to the extracted mesenchyme distribution data. The center coordinates, boundary length and connection component index of each mesenchyme block are recorded in the map. After obtaining the mesenchyme distribution data, the epithelial structure is extracted by a spatial structure decoupling method. In the decoupling process, the epithelial structure is defined as a region with a nucleus density higher than a threshold and a non-central core. A connected region labeling algorithm is used to label all connected nucleus regions in the nucleus density map, and a central core with a diameter greater than 200 pixels is excluded. The peripheral dense core structure is reserved as an epithelial candidate region. The epithelial candidate region and the mesenchyme distribution data are subjected to logical operation. The nucleus density map and the mesenchyme mask map are compared pixel by pixel. Only the nucleus region adjacent to the mesenchyme region is reserved as the final epithelial distribution. The output image format is a single-channel grayscale mask map representing the epithelial region distribution. In constructing the three-distribution framework, the obtained epithelial distribution data, nucleus density distribution data and mesenchyme distribution data are respectively mapped into the RGB three channels. The nucleus region is written into the red channel, the mesenchyme region is written into the green channel, and the epithelial region is written into the blue channel. The combined three-distribution framework image maintains the original image size and records the structure label information corresponding to each pixel.
[0048] Preferably, step S2 comprises the following steps:
[0049] Step S21: collecting cervical tissue development history data; and performing time sequence label fusion on the three-distribution framework based on the cervical tissue development history data, to generate time sequence three-distribution framework data;
[0050] Step S22: performing evolution trajectory tracking according to the time sequence three-distribution framework data, to obtain three-distribution evolution trajectories, wherein each evolution trajectory contains no less than 4 minimum nodes, and the structure block drift distance between nodes is controlled within 64 pixels, and the trajectory tracking chain is automatically interrupted when the threshold is exceeded;
[0051] Step S23: constructing a three-distribution framework development life line through the three-distribution evolution trajectories; and extracting a stratified development trend in the three-distribution framework development life line, wherein a sliding trend window width is set to 3 frames for trend identification, and a trend turning point change slope needs to be greater than 0.4 to be defined as an effective mutation;
[0052] Step S24: performing trend distribution mapping based on the stratified development trend, to confirm the predicted development trend of each layer in the three-distribution framework.
[0053] In this embodiment, a total of 1250 cervical cases from six pathology centers are selected, each case contains at least three images of different time points, the time span is controlled between 6 months and 36 months, all images are digitized by a unified scanning instrument at 20 times, the image format is SVS or TIFF, the resolution is set to 0.24 microns per pixel, after performing standardized resampling and spatial normalization operation on all slices, numbering according to the slice timestamp, and calling the three distribution framework structure graph constructed in the previous step as the structure input for each frame image, performing structure label alignment operation on all time series images, using a structure consistency based probability fusion model (Structure Consistency Fusion Module) to perform channel level label fusion on the three distribution structure mask of each frame image during the alignment process, and setting the region matching confidence threshold to 0.75, below the threshold area will be rejected, the three-channel structure diagram of the fusion output is sorted according to the time stamp to form the time sequence three-distribution framework data, the input data is the time sequence three-distribution framework data labeled by time sequence, in the trajectory tracking stage, the structural block is taken as a unit for tracking, the center coordinates, the main axis direction and the channel label of each structural block are taken as node definition parameters, a joint matching algorithm based on spatial proximity and structural feature cosine similarity is used for bidirectional pairing and matching of structural blocks in adjacent frames, the matching distance control threshold is set to 64 pixels, if the Euclidean distance of two structural blocks exceeds the threshold, the trajectory connection relationship is immediately interrupted, each trajectory requires at least four continuous frame nodes, the tracking algorithm executes the direction consistency constraint, if the main axis direction of the structural block changes by more than 45 degrees, it is classified as a non-continuous evolution path and the trajectory chain breaking operation is performed, all valid trajectories are numbered and the image frame number, structure type and center coordinates of each node are recorded, finally forming a three-distribution evolution trajectory library, each trajectory corresponds to a unique structure number in its life cycle, the input data is the time sequence three-distribution framework data labeled by time sequence, in the trajectory tracking stage, the structural block is taken as a unit for tracking, the center coordinates, the main axis direction and the channel label of each structural block are taken as node definition parameters, a joint matching algorithm based on spatial proximity and structural feature cosine similarity is used for bidirectional pairing and matching of structural blocks in adjacent frames, the matching distance control threshold is set to 64 pixels, if the Euclidean distance of two structural blocks exceeds the threshold, the trajectory connection relationship is immediately interrupted, each trajectory requires at least four continuous frame nodes, the tracking algorithm executes the direction consistency constraint, if the main axis direction of the structural block changes by more than 45 degrees, it is classified as a non-continuous evolution path and the trajectory chain breaking operation is performed, all valid trajectories are numbered and the image frame number, structure type and center coordinates of each node are recorded, finally forming a three-distribution evolution trajectory library, each trajectory corresponds to a unique structure number in its life cycle and is continuously called in the subsequent stage, all trajectory sequences in the three-distribution evolution trajectory library are called as basic data, the structure attribute change trend with time of each trajectory sequence is extracted according to time sorting, a single-variable structure index sequence of each trajectory is constructed, the indexes include structure area, gray mean and channel density, a sliding trend window is used to extract the trend of the index sequence, the sliding window size is set to 3 frames, that is, every three consecutive time points are taken as a group for local slope calculation, when the change slope between any two frames exceeds 0.4 is defined as a structural mutation point, and the mutation point position is recorded as a turning frame index. Continuous trends of change in the same direction are merged into a single trend segment. The trend segment is marked as expansion, stability or contraction according to type, and the start and end frame indexes and trend direction values of each trend segment are recorded. The trend direction value is the slope value of the structural indicator change. The lifeline data structure contains the trajectory number, structure label, number of trend segments, trend segment direction value array and corresponding frame range identifier. All results are stored in a structured trend library. The labeled trend segment data in the structured trend library are called, and each type of structure (nuclear area, interstitial area, epithelial area) is analyzed separately. The frequency and direction of trend types appearing in all trajectories were calculated, and a trend probability matrix was constructed. Each element in the matrix represents the probability value of the trend direction of the corresponding structure within a specific frame range. The probability value is calculated by weighted normalization of the trend segment direction value. The structure segment with the highest probability is marked as the dominant trend segment. The trend value of each frame in the dominant trend segment is encoded as the final predicted development trend. The encoding structure format is a three-channel structural trend graph, with each channel corresponding to the trend direction value of the nuclear region, interstitial region, and epithelial region. The image size is consistent with the original slice, and the trend value range is normalized to [-1, 1].
[0054] Preferably, performing cervical tissue environment field simulation using the cervical tissue image in step S3 includes:
[0055] Extracting cervical environment data from cervical tissue images;
[0056] Fitting the asymmetric tissue tension field to the cervical environmental data to obtain the environmental tension characteristic field;
[0057] Injecting the horizontal-radial structural flow tensor into each tension sub-block of the environmental tension feature field, and performing residual conduction on the maximum gradient direction of the environmental tension feature field, thereby generating environmental direction sensing guidance data;
[0058] The cervical tissue environment field is simulated based on the environmental direction sensing guidance data and the cervical environment data, thereby generating a simulated cervical environment field.
[0059] In this embodiment, the original full-slice cervical image with a resolution of 0.24 microns per pixel and a magnification of 20 times is selected as the input data source. The image is preprocessed by tiling, with a tile size of 512 pixels x 512 pixels. Batch sliding window processing is performed with a sliding step size of 128 pixels. The background region is filtered using an image hue region extraction function based on the HSV color space. The local texture indicators, including energy value, inertia moment, entropy value, and contrast, are calculated for the retained tiles using the gray level co-occurrence matrix (GLCM) method. The region with intense texture changes and tissue differentiation characteristics in the foreground region is selected as the environmental data extraction block. Independent gray level histogram analysis is performed on the stromal channel region in the tile. The histogram kurtosis and skewness are used to determine the tissue arrangement direction and density trend. Tiles with a kernel density below 0.25 and a mean texture contrast greater than 0.6 are included in the cervical environment data set. The tiles extracted in the previous step are used as input for tension feature fitting. Local gradient direction extraction and gray structure response fusion processing are performed on each tile. First, the Sobel gradient operator is used to calculate the local gradient map in the x and y directions. Then, the local structure response function is used to integrate the gray level distribution and gradient direction. The structure response function is set as R(i,j) = αGx(i,j) + βGy(i,j) + γI(i,j), where α, β, and γ are the harmonic coefficients of gradient direction and gray response, with values of 0.4, 0.4, and 0.2, respectively. Direction inversion symmetry detection is performed on all tiles. Regions that meet the criteria of a symmetry center shift greater than 16 pixels or a gradient angle difference exceeding 45 degrees are marked as non-symmetrical tension areas. Finally, the environmental tension feature field map is generated based on the direction field results of each tile. The image format is a two-channel tensor map. The first channel records the tension direction angle of each pixel, and the second channel records the tension amplitude value. The tension direction range is limited to 0 to 180 degrees, and the tension amplitude is standardized to [0, 1]. The tension feature field map is divided into 128 pixel x 128 pixel tension sub-block regions. Horizontal structure flow tensor and radial structure flow tensor injection operations are performed on each sub-block. The horizontal structure flow tensor direction is fixed as the positive x-axis within the sub-block, and the intensity is set to 0.5. The radial structure flow tensor is symmetrically expanded in the four quadrants with the sub-block center as the origin. The tension value decreases gradually from the origin to the edge, with a maximum intensity of 0.8 and a decreasing factor of 0.12, After pixel-level superposition of the two types of tensors, a double-tensor merged image is formed, and the maximum gradient direction calculation process is performed on each tensor sub-block. A 5x5 convolution window is used to search for the local maximum value direction on the direction field image, and the gradient dominant direction vector is extracted. The residual conduction operation is performed on the positions where the direction jump is greater than 30 degrees. The specific method is to construct a 3x3 structure guide convolution kernel to perform gradient field guided diffusion on the tensor direction image, repair the dominant direction interruption area and form a continuous sensing link. Finally, a structural tensor flow guide image is output, each pixel contains complete tensor field and direction vector value. The above structural tensor flow guide image and the original cervical environment data are used as input to construct a tensor disturbance simulation model for each block area. The disturbance simulation takes the direction guided tensor as the core injection source, and adopts the disturbance superposition model to realize the diffusion process of the disturbance field. The disturbance injection range is controlled within 64 pixels x 64 pixels, and the disturbance strength is set according to the structural density change value. The disturbance strength in the core area is set to 0.7, the disturbance strength in the epithelial area is set to 0.5, and the disturbance strength in the interstitial area is set to 0.3. When performing disturbance injection, three layers of structural disturbance flow field channels are constructed, and the direction consistency of each channel is regulated. Local tensor similarity filtering is used to guide the disturbance direction to achieve harmonic regulation. A Gaussian smoothing function is introduced to the structural boundary area to realize disturbance buffer transition. After the superposition of all disturbance channels is completed, a simulated cervical environment field image is generated, which has the same size as the original tissue block. The output image is a multi-channel tensor image, including the original texture image, the direction guide image, the structural disturbance image, and the channel distribution mask image.
[0060] Especially important is that a horizontal-radial structural flow tensor is injected into each tensor sub-block of the environmental tension characteristic field, and residual conduction is performed on the maximum gradient direction of the environmental tension characteristic field, including:
[0061] The environmental tension characteristic field is divided into a plurality of non-overlapping tensor sub-blocks, and a tensor sub-block index is constructed based on the plurality of non-overlapping tensor sub-blocks;
[0062] A horizontal-radial structural flow tensor is injected into each non-overlapping tensor sub-block based on the tensor sub-block index, and structural tensor superposition data is obtained;
[0063] Inter-block tensor continuity balancing processing is performed on the structural tensor superposition data to obtain tensor fusion data;
[0064] The maximum gradient direction in the tensor fusion data is extracted to generate a dominant structural flow direction;
[0065] The high gradient mutation area in the dominant structural flow direction is extracted;
[0066] The continuity of the high gradient mutation area is completed to generate environmental direction sensing guide data.
[0067] In this embodiment, the input data is the environmental tension feature field tensor graph generated by the previous stage, the size of the tensor graph is 2048 pixels x 2048 pixels, in order to ensure the structural alignment of subsequent tensor injection, an equal interval block division operation is performed on the tension feature field, the division scale is set to 128 pixels x 128 pixels, the block traversal is performed in a sliding window manner, the window sliding step is 128 pixels to ensure no overlapping area, a unique block number is assigned to each sub-block and a two-dimensional index mapping matrix is established, which records the starting coordinate position of each sub-block in the original tensor graph, the tension average value, the direction principal axis angle and the structural region channel label, the tension sub-block index structure is used for accurate block-level positioning and tensor propagation path management in the subsequent structural tensor injection process, the index information is stored in a structured JSON format data, the key-value pair includes sub-block ID, center coordinates, tension direction statistical histogram and density classification label, forming a block-level injection control interface of the environmental tension feature field, a horizontal structural tensor template and a radial structural tensor template are constructed respectively, the horizontal tensor direction is the positive direction of the x-axis, the tensor intensity is symmetrically distributed on the center line of the sub-block, the maximum value is 0.75, the edge line decreases to 0.3, an exponential decay function is used for intensity fitting, the radial tensor takes the geometric center of the sub-block as the origin, the Euclidean distance of each pixel point to the center is calculated, and a radial intensity distribution graph is constructed according to a linear normalization function, the maximum tensor value is set to 0.8 and the minimum value is 0.2, the two structural flow tensors are superimposed at the pixel layer, and are combined into a unified tensor field matrix, the tensor field retains the vector components of each pixel in two directions, and writes the tensor superposition result into the sub-block region of the tension feature field original graph, all sub-blocks perform structural flow injection operation in the same way, finally generate a structural tensor superposition data tensor graph, each pixel point contains a two-dimensional vector structure value and a regional structure label index, in the block interconnection continuity balancing processing stage according to the structural tensor superposition data, a block interconnection boundary coordination module is constructed to perform structural tensor edge docking operation on each pair of adjacent tension sub-blocks, first, a 3-pixel wide intersection area is established at the adjacent boundary position, and the tensor direction difference value of the area is calculated, if the direction angle is greater than 30 degrees, it is considered that there is a tensor discontinuity, a transition tensor band is introduced to the discontinuous area, the tensor value is set to the weighted average direction of the tension direction between the two blocks, the intensity is 90% of the lower value of the two tensor intensities, and at the same time, a Gaussian blur kernel is used for smoothing in the area, the blur kernel radius is set to σ = 1.5To ensure the smooth transition of tension direction, after completing the balancing operation for all tension sub-block boundary intersection areas, the full image tensor graph is merged and the tensor fusion data image is generated, which establishes block continuity response on the basis of the original tensor structure, forms a unified structure expression of the full image tensor direction field in the image space, performs gradient response scanning on the tensor fusion image, extracts the tensor direction distribution gradient at each pixel point using a 5x5 Sobel direction response kernel, calculates the direction value change rate of each point in the x direction and the y direction, calculates the tensor gradient amplitude and direction angle of each pixel point, and the maximum gradient direction corresponds to the position direction angle of the local maximum amplitude value. The dominant structure flow direction image extracted is direction filtered, and the area with a gradient angle change of less than 10 degrees is suppressed to retain the significant direction mutation area and establish a direction index map. Finally, each pixel in the output image has a dominant structure flow direction value. The direction map is stored as a single-channel direction angle map, with an angle range limited to 0 to 180 degrees, as the basis for subsequent high-gradient mutation area extraction and conduction repair. A local difference window is used to calculate the direction jump image, a 3x3 neighborhood window is established around each pixel, and the maximum and minimum difference values of the direction values in the area are calculated. The pixel points with a difference value greater than 45 degrees are marked as mutation boundary points. Boundary connectivity analysis is performed on the mutation boundary points, and all mutation blocks are marked using a four-neighbor connection component. Discrete areas with an area less than 25 pixels are excluded, and significant mutation connected regions are retained as high-gradient mutation areas and output. A mutation mask image is established, with each pixel value being 0 or 1, indicating whether it belongs to a mutation area. The mutation area boundary is vectorized and encoded, and the main direction vector of the mutation area is extracted and used for direction constraint in subsequent conduction operations. The mutation mask image is subjected to boundary interpolation expansion processing, and the mutation boundary is fitted with a dominant direction vector path using a Bezier curve. A linear guide path based on the direction consistency principle is established on each mutation path, and a tensor direction completion value is inserted at each break point along the fitted curve direction. The interpolation area uses a trilinear interpolation method to calculate the transition tension vector, and the interpolation tensor intensity is set to the average of the tension values of the two consecutive points. A one-dimensional direction diffusion operation is performed on each completed path, with a diffusion width of ±2 pixels. Fuzzy direction fusion is performed on the diffusion area to alleviate the discontinuity problem of the interpolation edge. All completed results are merged into the dominant structure flow direction image to form a complete and coherent direction guide image. Finally, the environmental direction sensing guide data image is output, with each pixel point containing a tension direction angle and a structure response intensity value. The format is a two-channel floating-point tensor image, with the first channel being the direction angle image and the second channel being the direction sensing intensity image. The image size is the same as the original tension feature. Fig. 1 Therefore, the number of channels is 2, and the data type is float32 format. Each channel retains four decimal precision to improve the accuracy and stability of the conduction path fitting.
[0068] Preferably, the hierarchical evolution prediction of the nuclear-stromal-epithelial tri-distribution framework in step S3 based on the simulated cervical environment field and the predicted development trend of each layer includes:
[0069] projecting the predicted development trend of each layer to the nuclear-stromal-epithelial tri-distribution framework into the simulated cervical environment field to obtain projected tri-distribution data;
[0070] performing hierarchical dynamic migration simulation on the projected tri-distribution data to generate hierarchical dynamic migration data;
[0071] performing inter-layer coupling response speculation based on the hierarchical dynamic migration data to generate coupling response prediction data;
[0072] performing evolution driving processing on the projected tri-distribution data according to the coupling response prediction data to obtain each layer framework evolution prediction data.
[0073] In this embodiment, the input data is a three-channel image in the structural distribution map, each channel corresponding to a nuclear region channel, a stromal region channel and an epithelial region channel, the corresponding predicted development trend being provided by the previous stage trend distribution mapping, and each channel being attached with a structural trend tensor, the tensor dimension being the same as the distribution Fig. 1The trend value is in the range of [-1, 1], where a positive value indicates an expansion trend, a negative value indicates a contraction trend, and a zero value indicates a stable state. The cervical environment field is a multi-channel image that includes a direction guide tensor map, a tension disturbance map, and a structure positioning channel map. After performing pixel-level alignment processing on the three distribution maps and the environment field, the target migration direction of each pixel is calculated based on the trend tensor. The migration direction is based on the dominant direction angle in the environmental direction sensing map, and the trend value is multiplied by a standard offset coefficient, which is set to 8 pixel units. Affine projection transformation is performed on each pixel of the three distribution maps, and the trend intensity in the core area is multiplied by a coefficient of 1.2, the interstitial area is multiplied by a coefficient of 1.0, and the epithelial area is multiplied by a coefficient of 0.9. The trend-guided direction offset map is constructed, and the offset position is redrawn as a projected structure map in the image. Bilinear interpolation remapping is performed on all pixels to ensure the continuity of the structure edge. The final output image is a projected three-distribution data map, which records the original structure label, trend direction vector, and projection position coordinates. The core area, interstitial area, and epithelial area are separated from the projected three-distribution map as independent structure channel maps, and a hierarchical structure migration unit is established. The dynamic evolution process is constructed using continuous frame simulation for each channel, with a simulation time step of Δt = 1. Local tensor response adjustment is performed on each structure block at each time step, and the migration path direction is provided by the tension direction sensing channel in the environment field. The center point of each structure block is shifted by 1 step in its local direction gradient field, and the step size is weighted and adjusted based on the structure trend value. When the trend value is 0.5, the step size is 4 pixels, and when the trend value is -0.5, the step size is -4 pixels. Linear scaling is used for incremental and decremental adjustment. The time window is set to 6 frames, and the structure block coordinates, contour boundary, and pixel grayscale map are recorded for each frame. Each pixel is labeled with its original structure type and current migration direction. The hierarchical dynamic migration data is formed by combining all frame sequences, and the corresponding independent core area migration map, interstitial area migration map, and epithelial area migration map are formed for each channel. The structure index is consistent on the unified time axis. The dynamic migration maps of the three channels are subjected to structure index reconstruction. The spatial position change vector and area change ratio of each structure block are labeled on the time axis. The structure boundary relationship between the core area and the interstitial area, and between the interstitial area and the epithelial area, is extracted at the same time step. The block interconnection edge is established through the spatial adjacency graph, and the coupling index function is constructed for the structure pairs on the interconnection edge. The function is composed of three parts: contact area ratio, trend direction difference, and local tension field direction difference. When the function threshold is greater than 0.6, it is determined as a coupled pair. All coupled pairs perform structure interaction calculation in the time sequence, and the coupling response strength is simulated through the cross-migration convolution network. The convolution kernel size in the network is set to 5x5, the channel number is 3, the input is the boundary direction map and the trend map of adjacent structure blocks, and the output is the response strength map. The strength in the response strength map is greater than 0.The region of 3 is recorded as the effective coupling response region, ultimately forming a coupling response prediction data map. This map records the coupling type, response strength, and direction of action for each pixel. Using the coupling response prediction map as the driving source, a response mapping matrix is constructed for each structural block in the projected three-distribution map. The matrix dimension is the tension-sensing region range of the structural block area, with a range diameter set to 64 pixels. The driving direction is based on the direction vector field provided in the coupling response map, and the driving strength is the response strength value multiplied by the structural trend value. A driving influence factor function is established. Affine transformation superposition is performed on each structural block, directional extension is performed on high-response regions, and area compression is performed on low-response regions. Structural point set mapping is performed on all structural regions and pixel region boundaries are reconstructed. After structural correction, all structural blocks are recombined into three-channel images corresponding to the nuclear, interstitial, and epithelial regional channels. Finally, the framework evolution prediction data map for each layer is obtained. Each channel image retains the pixel label, response direction value, and area change coefficient. The structural block coordinate index and simulation time step sequence number are also recorded. This is used as the structural inference input layer for the subsequent evolution mapping and image classification generation steps.
[0074] Of particular importance is the inference of inter-layer coupling responses based on layered dynamic migration data, including:
[0075] Rebuild cross-layer indexes in layered dynamic migration data and confirm inter-layer structure data based on the cross-layer indexes;
[0076] Performing interlayer tensor interference processing on interlayer structural data according to the cross-layer index to obtain coupled interference data;
[0077] Confirm the interlayer interference direction based on the coupled interferometry data;
[0078] Project the inter-layer interference direction to the previous time sequence of each time node to obtain the reverse response path;
[0079] Directional response values are aggregated based on the reverse response pathway to generate coupled response prediction data.
[0080] In this embodiment, the input data is a dynamic migration sequence image data generated by three structural channels, which are the nuclear region dynamic image, the interstitial region dynamic image and the epithelial region dynamic image, the image size is unified to 512 pixels x 512 pixels, each channel contains structural block identification, pixel mask, structural centroid coordinates and trend direction vector at each time step, the cross-channel spatial correspondence is constructed through spatial overlap relationship, the inter-channel overlap degree analysis is performed on the three distribution maps at any time node, the IOU (Intersection over Union) index is used to calculate the structural block pairing relationship between the nuclear region and the interstitial region, the interstitial region and the epithelial region, the connection relationship is established for the structure pair with IOU value greater than 0.3, and the cross-layer structure pair index table is constructed, each structure pair records the structure block ID, channel number, structure outline, center point distance and structure area difference, the index table takes time stamp as primary key and structure pair ID as sub key, the cross-layer structure data mapping records the structure interaction relationship and its time sequence topological state, after constructing the cross-layer structure index, the tensor interference processing operation is performed on the inter-layer structure data using the index, each structure pair is taken as an interference unit, the tension vector action path is constructed inside the structure pair, the path direction is defined by the centroid vector from the upper structure to the lower structure, the interference channel region with a width of 16 pixels is set around the path, the tensor fusion calculation is performed on the upper and lower structure tension fields in the region, the tensor value is derived from the structure trend direction graph and the tension guide graph in the structure dynamic migration graph, the pixel-by-pixel tension vector superposition rule is used for fusion, the vector merging superposition processing is performed on the vectors with direction difference angle less than 20 degrees, the direction decomposition and weighted attenuation are performed on the vectors with direction difference angle greater than 20 degrees, the generated fusion tensor is called interference tensor, the interference tensor field is stored according to the pixel position, the tensor intensity is normalized in the range of [0.0, 1.0], and the interference intensity is less than 0.2The region is marked as a weak response area, and finally all the interference tensor outputs of the structure pairs are coupled interference data images. The data structure is a four-dimensional tensor, including time steps, x coordinates, y coordinates, and tensor vector values. After generating the coupled interference data, the dominant interference direction identification operation is performed on the interference tensor image at each time step. The maximum response direction of the tensor field in each pixel of the image is calculated, which is the dominant interference direction. The extraction method of the dominant direction is to find the direction of the tensor vector resultant force. After statistical analysis of all directions, the overall dominant direction angle of each structure pair is calculated. The direction range is [0, 180] degrees. Each dominant direction angle is written into the direction mapping image according to the structure block region. The interlayer interference direction image is generated. The image channel number is 1, which records the current structure coupling relationship and the dominant interference direction value of each pixel. The interference direction image is used as the direction control image for subsequent time reversal projection. After the interference direction image is constructed, the time reversal mapping operation is performed on the direction image to obtain the reverse response path of the structure motion trend. For the direction image in each time node t, the interference direction is projected into the corresponding coordinate position in the previous time frame t-1. The mapping method is to move the current pixel position as the starting point in the t frame along the reverse direction vector by d pixels, and d is set to 4 pixels. The reverse path marker is recorded at the corresponding projection point position in the t-1 frame. The reverse response path image is constructed in the t-1 frame. The path image is based on the structure block number and indexed by the direction vector. The position after the direction projection is marked as the potential response position. After the reverse projection of all structure blocks is completed, the full-image direction response path image is constructed. Each pixel position of the image records the path direction and the transfer tensor identifier. Based on the direction response path image, the direction response value aggregation operation is performed. Each path in the path image is input as a response trajectory. The weighted average of the interference intensity of all pixels on the trajectory is calculated. The weight function is constructed according to the path step and the direction consistency factor. The path direction consistency factor is the cosine value of the direction angle difference between the front and rear pixels. The aggregation function is the weighted average tensor intensity. The response values of all paths are accumulated to the path endpoint position. Each path endpoint records the aggregated tensor response value and the direction vector. The coupled response prediction data image is formed by summarizing all path endpoints. The image structure is a two-channel tensor image. The first channel is the response intensity value, and the second channel is the response direction angle value. All response values are normalized to [0, 1]. The response direction angle accuracy is kept to one decimal place. The image size remains consistent with the original three-frame framework.
[0081] Preferably, the step S4 of performing evolution projection on the cervical tissue image according to the evolution prediction data of each layer framework comprises:
[0082] The evolution prediction data of each layer framework is spatially position encoded to obtain position encoded prediction data;
[0083] The position encoded prediction data is structured and projected into the cervical tissue image to generate a structure projection image.
[0084] Evolutionary states in the fusion structure projection image are obtained, thereby obtaining a state fusion image;
[0085] Fine-grained evolution rendering is performed on the state fusion image, thereby obtaining a fine-grained evolution mapping image;
[0086] Pseudo-color in the fine-grained evolution mapping image is enhanced, thereby obtaining an evolution mapping cervical tissue image.
[0087] In this embodiment, the input data is the evolution prediction image corresponding to the three channels of nuclear region, interstitial region and epithelial region. Each structural region is represented by a single-channel mask image, with a size of 512 pixels x 512 pixels. Each pixel in the prediction data contains a structure class, a structure contour index, and a response direction vector. When performing spatial position encoding processing on each structure block, a coordinate indexing method based on structure centroid positioning is used. First, a connected region analysis is performed on the structure mask to extract all structure block boundary coordinates and center positions. Then, a two-dimensional position embedding matrix is constructed for each structure block in the original image coordinate system. The matrix has a dimension of 512 x 512 x 2, recording the relative distance coordinates of each pixel on the x-axis and y-axis. At the same time, the structure block trend direction vector and response intensity are extracted as dynamic features. The dynamic feature values are encoded into a single-channel layer, which is fused with the coordinate embedding image to form a three-channel position encoding prediction image. In the position encoding, each pixel contains the current position coordinates, the belonging structure block ID, and the dominant trend direction value and tension response index of the corresponding structure. The original cervical tissue image is selected as the background projection map, which needs to maintain the same resolution and image dimension as the prediction image. The affine position transformation is performed on each structure block center point coordinate in the structure position encoding image. The transformation matrix determines the translation amount and direction based on the dominant trend direction vector and response intensity. The structure center point translation distance d is the response intensity multiplied by the standard offset coefficient, which is set to 16 pixels. After performing pixel-level coordinate mapping transformation on each structure block, the structure contour is reconstructed at the target position, and the original texture grayscale value is preserved. All structure blocks are written on the background map in sequence and edge fusion processing is performed. The fusion method is to establish a 4-pixel transition band at the structure edge and perform linear blending to reduce boundary fracture phenomena. Finally, the structure projection image is generated, in which each structure block is repositioned and retains the dominant direction field and original structure identification. The structure projection image is called with the corresponding response direction image and trend tensor image to construct the structure state channel in the image channel. This channel is used to record the evolution state identification of each structure block. The evolution state is determined according to the trend tensor value. A value greater than 0.3 is defined as an expansion state, a value less than -0.3 is defined as a contraction state, and a value between them is defined as a stable state. For each structure block, a state mapping channel is constructed in its position region. The expansion region is marked with a red channel, the contraction region is marked with a blue channel, and the stable region is marked with a green channel. The state identification intensity is normalized to [0.2, 1] according to the absolute value of the trend tensor.0], the fusion image is constructed as an RGB three-channel image, which records structural contour texture, structural state color and state intensity respectively, a maximum state intensity priority fusion strategy is used for the overlapping area of multiple structural blocks, and state edge smoothing processing is performed on the whole image, with a smoothing kernel size of 3*3. The output state fusion image contains structural information, trend information and spatial evolution layer. A structural region sub-block is established, and the size of each block is set to 64 pixels*64 pixels. A multi-layer texture reconstruction network is used to render and model the structural region inside each block. The network input is the structural state label, response direction vector and trend intensity value. The network contains three convolutional layers and two skip connection modules, and the output is a refined texture layer, which contains structural edge blur degree, internal density change and direction gradient texture. The rendered results are subjected to block splicing and reorganization to establish a complete fine-grained structure image. The structure image increases three layers of refined texture on the original basis, including morphological expansion map, density stripe map and direction gradient map. Finally, a fine-grained evolution mapping image is synthesized. In the image, each structure not only retains the spatial position and structural type, but also retains the texture information of the simulated tissue evolution. The three-channel layer records the texture grayscale map, trend direction map and structural state map respectively. Standard color mapping table is used for color conversion processing, the input is the state map and texture map channel in the fine-grained evolution image, and Jet pseudo-color mapping function is used to map the grayscale texture value to a color image, in which 0 value is mapped to dark blue, 1 value is mapped to red, and intermediate values are transitioned to green and yellow. The mapping table retains 256-level color mapping precision. Edge sharpening processing is performed on the mapping image, and a 5*5 high-pass filter kernel is used to enhance the structural boundary definition. Color gradient superposition operation is performed on the trend direction layer, and a transparency layer is constructed using the structural direction angle value and the state intensity value. After superposition, a structural state response map is formed. The three-layer fusion constructs the final evolution mapping cervical tissue image, with an image size of 512 pixels*512 pixels and an image format of RGB three-channel floating point image. Each channel stores color value precision to 0.001 in float32 format. The output image is used as the final input data structure for classification modeling and visual alignment. The meaning of the marked image is the evolution of the structure mapping image.
[0088] Preferably, the image classification of the cervical tissue image according to the evolution process category based on the evolution mapping cervical tissue image in step S4 comprises:
[0089] Extracting process features from the evolution mapping cervical tissue image;
[0090] Clustering based on process features and constructing image category feature labels;
[0091] Classifying the evolution mapping cervical tissue image based on the image category feature labels, thereby obtaining multiple evolution process image categories;
[0092] determining the image category of the cervical tissue image corresponding to the cervical tissue image through the evolutionary process image category determination; and
[0093] performing image classification on the cervical tissue image according to the image category of the cervical tissue image, thereby generating a cervical classification image.
[0094] In this embodiment, the input image is an evolved mapping cervical tissue image after completion of evolutionary projection, state fusion and fine-grained evolution rendering, the image size is 512 pixels x 512 pixels, the channel structure includes structure state channel, structure direction channel and evolution texture channel, for the state channel, first perform region slicing operation to obtain the average state intensity value of all structure blocks in the image, and calculate the standard deviation and extreme interval, then perform gradient direction histogram extraction operation on the direction channel, the angle range is divided into 18 equal intervals, the number of pixel points in each interval is accumulated to obtain the direction distribution vector, in the texture channel, extract the gray level co-occurrence matrix (GLCM) parameters of the image block, extract the contrast, entropy, energy and correlation of four directions, i.e. 0 degrees, 45 degrees, 90 degrees and 135 degrees, a total of 16 texture feature values, finally, the state statistics, direction distribution vector and texture parameters are spliced into a 64-dimensional process feature vector, each image corresponds to a vector data, the 64-dimensional process feature vector matrix is input into the K-means clustering model for image category division, the clustering number is set to 5 categories, each category represents a different evolution trend process, the model uses k-means++ initialization strategy to improve the stability of the initial center, the distance measurement function uses Euclidean distance function, the maximum iteration number is set to 300 rounds, the convergence criterion is to stop iteration when all center points move less than 1e-4, the silhouette coefficient is used for category number verification during clustering, it is found that the average silhouette coefficient reaches 0.59 when k = 5 in the verification stage, which meets the clustering stability standard, after clustering, each sample is assigned a clustering category label, the label naming method is "EVT-0" to "EVT-4", and the category statistical parameters are constructed for each label, the output label structure includes image ID, label category and feature vector index, read the image category feature label structure, analyze the category identification corresponding to each image, group all images according to the category field, write the images into different path directories respectively, the directory structure is set to "EVT-0" to "EVT-4" five folders, each folder corresponds to a process category, at the same time, an independent JSON label file is generated for each image, the label content includes image path, state mean value, dominant direction distribution and feature vector abstract, at the same time, a representative image index list file is constructed for each category for subsequent model training sample extraction, the total number of samples in each category is not less than 500, if the image samples of a category are insufficient, use affine rotation ± 15 degrees, brightness disturbance ± 10%, Gaussian blur σ = 0.8Data augmentation is performed to balance the classes, and the final evolved images are labeled, grouped, and written to the image classification structure. An image number mapping table is used for index matching. This table is generated during the evolution mapping phase and contains fields for the original image ID, evolved image ID, and class label fields. The evolved image ID and label are read from the image class label, and the corresponding original image ID field is queried through a hash index structure. Then, an original image and its corresponding evolved class label mapping table is constructed. For each original image, its final evolved classification label is recorded, along with the average state intensity, direction principal axis angle, and structure complexity level of the evolved image. The structure complexity level is calculated by combining the number of structure blocks and the standard deviation of the structure spacing, with a total of five levels, from Level-0 to Level-4. Finally, all original images are assigned a one-to-one classification label. The input is the original cervical tissue image and its corresponding class label mapping table. For each image, a label renaming operation is performed, adding a class field prefix to the original image filename. The naming rule is "EVT-x_image original name". At the same time, a structure information layer is generated for the image. This layer extracts the structure boundary mask from the corresponding evolved image and overlays it onto the original image. The boundary color is encoded in RGB according to the classification label, with EVT-0 being blue, EVT-1 being green, EVT-2 being yellow, EVT-3 being orange, and EVT-4 being red. The structure layer is overlaid with an alpha transparency of 0.4. The output is a three-channel floating-point PNG image. All classification images are written to different directories according to the label classification and bound to JSON label files. The label file records the class ID, image ID, main structure area contour coordinates, and feature parameter abstract. Finally, a cervical classification image set containing more than 5000 classification images is formed. This image set constitutes the data source for visual basic model training and the graph-text fusion input reference graph structure set.
[0095] Preferably, the hierarchical joint modeling training according to the cervical classification images in step S5 includes:
[0096] The cervical classification images are divided into cervical classification pre-cropped tiles at a standard magnification.
[0097] The cervical classification pre-cropped tiles are automatically screened for tissue structure regions to obtain tissue region tiles.
[0098] The tissue region tiles are processed with coordinate index encoding to obtain tile basic data.
[0099] The coverage of the tile basic data is calculated, and high-quality tiles are selected based on the coverage.
[0100] The high-quality tiles are processed for multi-center structure unification, and their labels are integrated to construct a large-scale training data set.
[0101] generating initial visual feature data based on self-supervised visual pre-training based on the DINOv2 architecture according to a large-scale training data set;
[0102] constructing a visual base model with organizational structure representation capability based on the initial visual feature data.
[0103] In this embodiment, the input data is the cervical classification image obtained in the evolution stage. The image is processed by pseudo-color processing, evolution mapping reconstruction, and process label superposition. The image resolution is unified to 0.24 microns per pixel, and the image size is standardized to 4096 pixels×4096 pixels. Before performing tile segmentation, the image needs to be converted to a three-channel matrix representation in floating-point format. Each channel value is normalized to the range [0, 1]. Then, a fixed window cropping method is used to perform equidistant sliding window segmentation on the image. The tile size is set to 512 pixels×512 pixels, and the step size is set to 384 pixels. This ensures that there is an overlapping area of 128 pixels between tiles for boundary compensation. When cropping the image edge to an uneven block area, mirror padding is performed on the insufficient size of the tile. The padding mode is “reflect” edge processing. After tile segmentation, about 63 tiles are generated. Each tile is named by the original image name plus the tile position index and saved in a structured directory. The original image position coordinates and the corresponding image class label are marked for each tile. The generated data structure is named cervical classification pre-cropped tile data. The input data is the pre-cropped tile image set obtained after standard magnification segmentation. The tissue structure region segmentation model is used to perform prediction operations on each tile. The segmentation model is a pre-trained tissue structure region detection model based on the UNet++ architecture. The model input tile size is 512 pixels×512 pixels, and the output is a single-channel region mask image. In the mask image, a pixel value of 0 represents the background, and a pixel value of 1 represents the valid tissue region. When the model performs inference on each tile, a batch size of 8 is used for batch processing. Morphological closing operation and maximum connected region extraction are performed on each mask image to eliminate edge false response regions. Tiles with a tissue proportion less than 60% in the predicted mask are directly excluded. The structure region index map is extracted from the retained tiles. The continuous regions in the mask are labeled as structure blocks, and the centroid coordinates, contour boundaries, and area values of each block are recorded. The structure block data format is a three-tuple of coordinates, structure mask, and label. This data is saved in the tile index structure and forms the tissue region tile dataset. The position coordinates of each tissue region tile in the original image are read. The absolute position coordinates of the structure block on the full image are obtained by adding the structure block centroid coordinates to the pre-cropped tile naming rule. The coordinates are jointly encoded with the tissue channel class and image class label to construct the position index structure. This structure uses a five-tuple format definition, which includes tile ID, structure ID, X coordinate, Y coordinate, and class ID. A two-dimensional coordinate mapping matrix is also constructed for reverse mapping of tiles to the original image. Each tile encoding structure is stored in a separate coordinate index file in HDF5 format. Each HDF5 file corresponds to a tile, and the internal dataset stores structure block coordinate data and tile texture vector information. Each tile file is approximately 1.2MB, the coordinate index encoding output is the tile base data, which is used for coverage calculation, tile screening and structure alignment modeling input layer calling. The area information of the corresponding structure block in the coordinate index is read. The structure area is counted by the pixel in the mask image. The tissue coverage index is calculated by the ratio of the structure area to the total area of the tile. The coverage threshold is set to 0.65. The tile with a coverage lower than the threshold is defined as a low-quality tile and is removed. The structure complexity index of the remaining tile is further extracted. The structure complexity is constructed by the number of different structure blocks and the standard deviation of structure density in the tile. The complexity score function outputs in the range of [0, 1]. The complexity score lower than 0.The tiles of 35 will also be excluded, and the final remaining tiles are defined as high-quality tiles and stored in named directories according to the structure channel category, with nuclear zone tiles, interstitial tiles and epithelial tiles saved in different folders. A total of 128,000 tiles from the six data centers are imported into the data processing platform, which uses the center ID and tile ID to establish an index table and performs cross-center tile style normalization processing. The color standardization module based on the Macenko algorithm is used to map the colors of all tiles, and the target template is derived from the center sample tile with the clearest structure. Label files are generated for all tiles according to image categories and structure channel categories, and label alignment operations are performed. The label content includes image category labels, structure channel labels, center attribution labels and structure outline mask images. The label structure is stored in YAML format, and each tile image is stored in a unified named directory corresponding to the label file. Finally, a large-scale training dataset is constructed, and the DINOv2 model structure based on the Vision Transformer is used. The input image size of the model is 224 pixels x 224 pixels, and each tile image is subjected to data enhancement processing before training. The combination of random flipping, random cropping, color disturbance and Gaussian blur is used to construct pre-training image pairs. Each image pair is fed into the main encoder and auxiliary encoder for self-supervised feature matching. The training process is optimized using a joint loss function composed of feature reconstruction loss and contrast loss. The training batch is set to 256 images per batch, and the training rounds are 400 rounds, each containing about 1000 batches. The optimizer uses AdamW, the initial learning rate is set to 3e-4, and the learning rate is decayed using Cosine Annealing. The model output is a feature vector corresponding to each image, with a vector dimension of 768. The feature vector is extracted and saved as initial visual feature data after training. The DINOv2 main structure is used as the encoder to build the model main architecture, and a structure region prediction head and a multi-task feature branch module are added to the output end. The structure region prediction head uses a combination of bilinear upsampling and spatial attention mechanism. The prediction head output is aligned with the structure channel mask image for feature supervision. The loss function is a combination of cross-entropy loss and Dice loss. The model input is an enhanced tile image, and the output is a structure region prediction map and a structure category feature vector. The training sample is the high-quality tile data selected, and the training method is supervised fine-tuning. The training rounds are 120 rounds, with 1000 batches per round and 64 tile images per batch. The training optimizer is LAMB, and the initial learning rate is 2e-4. The final model is saved as the CerS-V visual base model.
[0104] Preferably, in step S5, the image-text fusion modeling is performed based on the visual base model, and task transfer processing is performed, and the model is deployed to an integrated system, including:
[0105] The visual feature output in the visual base model is matched with the image-text annotation sample one by one, so as to obtain visual language pairing data;
[0106] Based on the visual language pairing data, the image-text modal mapping is bound, so as to generate multi-modal hybrid input data;
[0107] Based on the multi-modal hybrid input data, the image-text joint training processing is performed, so as to obtain a multi-modal base model with alignment capability;
[0108] The image-text base model data and the visual base model in the multi-modal base model are respectively subjected to 25-class core task scene configuration and module encapsulation processing, so as to generate multi-task cooperative processing module data;
[0109] The multi-task cooperative processing module data is injected into the decision intelligent agent planner constructed by the language model, so as to obtain task planning and scheduling control data;
[0110] Based on the task planning and scheduling control data, the modal conversion mapping processing is performed, so as to obtain unified inference path data;
[0111] Based on the unified inference path data, the structured interface is encapsulated, and the user interaction module is docked, so as to generate deployable integrated system data;
[0112] The deployable integrated system data is subjected to forward-looking verification and evaluation in six independent centers, and finally a cross-center deployment model system is obtained.
[0113] In this embodiment, the input image is large-scale high-quality tile data, the label file contains real organization description, pathological description language, structure annotation, cervical type label, text and question answer pair, etc. Text annotation, 768-dimensional visual feature vector is extracted for each tile input CerS-V visual base model, and language vector representation is extracted for the text part using the embedding layer of Qwen2.5-VL model. The language embedding dimension is unified to 4096 dimensions. In order to ensure one-to-one correspondence matching accuracy, a unique pairing index is established between image and text, a bidirectional mapping structure of image ID-text ID is adopted, if a certain image corresponds to multiple text samples, multiple visual language pairing data items are generated respectively, and a unified hybrid key-value pair identifier is generated for each pair of samples. Finally, a training data set containing 2.5 million pairs of visual language pairing data is constructed, each pair of data contains image path, visual feature tensor, text content, text embedding vector and multi-modal index structure. Taking visual language pairing data as input, a modal fusion preprocessing module is constructed. The visual input is the standardized image tensor with a size of 3×224×224, and the language input is the token sequence after word segmentation processing, with a maximum length of no more than 128 tokens. The visual input is extracted by CerS-V model to obtain 768-dimensional patch feature embedding sequence, and the language input is generated by Qwen2.5-VL word embedding layer to obtain 4096-dimensional token feature vector sequence. The two modal vectors are linearly mapped by LoRA low-rank adapter to unify to 1024-dimensional hybrid space, and a learnable cross-modal attention bridge layer is used to connect the two embedding channels. In each batch, 128 pairs of image-text binding are performed. After fusion, each pair of data forms a unified multi-modal hybrid input structure, and the output tensor is (batch, token_length, 1024), wherein token_length is the total number of image-text tokens. The final hybrid input data is used for subsequent multi-modal model training processing. Taking the multi-modal hybrid input data as the model input, the cross-modal language modeling network with Qwen2.5-VL as the language backbone is trained. In the training structure, the CerS-V visual coding layer is connected as the image prefix embedding channel, and the visual patch feature is injected into the first three layers of attention module of the language Transformer backbone through three layers of LoRA low-rank connection module. The training task includes image-text matching task, image-text generation task and cross-modal question answering task. The loss function consists of three parts, namely matching loss (InfoNCE), generation loss (CE) and question answering accuracy loss (F1-score). The training batch is 256 samples, the total number of training rounds is 20 rounds, each round contains 3000 steps of iteration, and the AdamW optimizer and 1e-4 initial learning rate are used. The multi-modal base model obtained by the above training process is named as CerS-M model, which has the ability to perform text generation, structure description, pathological interpretation and question answering after inputting image.According to the actual scene requirements of cervical pathological examination tasks, 25 core task lists are constructed, each task including input modal type, target task label type, judgment output field and structured report interface. For each task, a configuration file is defined, and the configuration fields include input channel mapping structure, task target function, label index mapping and module IO interface protocol. Based on the structure recognition ability of CerS-V model, image task modules such as structure recognition module, squamous carcinoma grading module and rare carcinoma identification module are constructed. Based on the multi-modal question answering ability of CerS-M model, text-image question answering module, structure description generation module and dialogue report generation module are constructed. All modules are saved as pytorchcheckpoint and JSON configuration pairs in the form of independent encapsulation. Finally, 25 decision module data are generated, each module corresponding to a unique module ID and scheduling label. A decision-making agent framework based on Qwen2.5 large language model is used. After loading the task configuration file and module IO interface protocol, the framework performs task analysis, module matching and path planning. The input is user dialogue instruction, task call request or image-text joint input. The planner first performs input content embedding encoding, and then performs top-k similarity matching with the task label word vector of the registered module. After successful matching, a task scheduling tree is generated, which defines the reasoning path node, module call sequence, input-output data flow mapping relationship, and generates a standardized task execution plan file. The file format is JSON structure, and the content includes: task ID, called module ID sequence, input modal path, output target format and task level parameter template. Task planning and scheduling control data are provided to the integrated system reasoning module for automatic calling of decision modules and data flow conversion. From the task scheduling control data, the required modal input, called module ID and data flow path of each task node are extracted. The input modal path performs format conversion and distributed preloading processing. The image modal conversion module completes image size rearrangement, normalization and color channel matching based on the joint preprocessing interface of OpenCV and PIL. The language modal performs word segmentation pre-encoding and Token ID mapping. All conversion results are packaged into a unified tensor structure and bound to the module input slot. The system constructs a modal mapping graph, with each node as a module input interface and each edge as a data conversion channel. The scheduling sequence field and execution timestamp are added to the graph structure. The execution engine triggers the reasoning operation according to the graph sequence, and finally generates unified reasoning path data with a nested embedding structure, including input data tensor, module flow path and output structure distribution plan. The unified reasoning path data is used to define the system service interface specification, and the interface is encapsulated in the RESTful API structure. The input is an image URL or a JSON instruction. The system builds task scheduling API and data input preprocessing API through Flask, and at the same time builds WebSocket channel to support long-session multi-round task question answering call.The structured output interface maps the model output to a JSON structure, fields including module output text, structure region annotation coordinates, judgment level and key structure description, interface logic nesting configuration is performed on the user interaction module, image display, structure drawing and decision dialogue are supported, and the interface state is synchronized to the backend state machine, six tertiary centers with cervical tissue full slice scanning and digital pathology systems are selected as verification places, the same version of the integrated system container is arranged in each center, the deployment mode adopts a local area network remote API access and local cache cooperative mechanism, 500 independent case image samples are selected in each center, including cervical cancer screening images, SILVA classification images and structured question and answer image task input samples, system evaluation indexes include four indexes of task call success rate, structure recognition accuracy, structure description generation consistency and question and answer accuracy, the evaluation period is three weeks, the system call log is recorded by the path engineer every week and is uploaded to a unified server for performance aggregation analysis, all log data is used in the training system for fine tuning compensation optimization, finally, a six-center deployment verification report is formed, and a cross-center deployment model system version after deployment is output as a model integration result, and model precision, stability and portability index data are recorded.
[0114] The application further provides a multi-modal based cervical pathological image classification model training system for executing the multi-modal based cervical pathological image classification model training method described above, the multi-modal based cervical pathological image classification model training system comprising:
[0115] The three-decomposition structure module is configured to acquire a cervical tissue image, perform tissue structure segmentation on the cervical tissue image to obtain a segmented cervical structure block, and divide the cervical tissue into a three-distribution framework of nucleus-stromal-epithelium based on the segmented cervical structure block.
[0116] The trend identification module is configured to collect cervical tissue development history data, identify a three-distribution framework development life line based on the cervical tissue development history data, and confirm a predicted development trend of each layer in the three-distribution framework through the three-distribution framework development life line.
[0117] The environment prediction module is configured to perform cervical tissue environment field simulation on the cervical tissue image to generate a simulated cervical environment field, and perform layered evolution prediction on the three-distribution framework of nucleus-stromal-epithelium based on the simulated cervical environment field and the predicted development trend of each layer to obtain layered framework evolution prediction data.
[0118] The evolution classification module is configured to perform evolution projection on the cervical tissue image according to the layered framework evolution prediction data to generate an evolution mapping cervical tissue image, and perform image classification on the cervical tissue image according to an evolution process category based on the evolution mapping cervical tissue image to generate a cervical classification image.
[0119] The fusion modeling module is used for hierarchical joint modeling training according to the cervical classification image, so as to obtain a visual basic model; the fusion modeling is carried out based on the visual basic model, task migration processing is carried out, and the cross-center deployment model system is generated by deploying to the integrated system.
[0120] The three-decomposition structure module constructs the nuclear-mesenchyme-epithelial three-channel structure framework, which is beneficial to improve the spatial decoupling ability of tissue partition and the structural clarity of subsequent modeling. The trend identification module extracts the structure evolution track through the development historical data, establishes the structure correspondence relationship between multiple images, enhances the temporal consistency and trend traceability. The environment prediction module generates a simulation environment field containing tension disturbance, pressure mapping and trend array, realizes the non-biological mechanism driven modeling of the structure evolution process, improves the directionality and structure fidelity of the evolution simulation. The evolution classification module generates a mapping image and divides the process category based on the prediction, so that the image classification has trend perception ability and change stage distinguishability. The fusion modeling module constructs a visual model through the classification image and completes the image-text alignment training, realizes the unified modeling of the organizational structure, trend semantics and image-text modal, and has expandability, combinability and structural level universality.
Claims
1. A training method for a multimodal cervical pathology image classification model, characterized in that: The following steps are involved: Step S1: Acquire a cervical tissue image; perform tissue structure segmentation on the cervical tissue image to obtain segmented cervical structure blocks; and divide the cervical tissue into a three-distribution framework of nuclear-stromal-epithelial based on the segmented cervical structure blocks. Step S2: collecting historical data on the development of cervical tissue; identifying the development lifeline of the three-distribution framework based on the historical data on the development of cervical tissue; and confirming the predicted development trend of each layer in the three-distribution framework through the development lifeline of the three-distribution framework; Step S3: simulating the cervical tissue environment field using the cervical tissue image to generate a simulated cervical environment field; performing layered evolution prediction on the three-distribution framework of nucleus-stroma-epithelium based on the simulated cervical environment field and the predicted development trend of each layer to obtain the predicted evolution data of each layer framework; Step S4: performing evolutionary projection on the cervical tissue image according to the evolutionary prediction data of each layer of the framework, thereby generating an evolutionary mapped cervical tissue image; classifying the cervical tissue image according to the evolutionary process category based on the evolutionary mapped cervical tissue image, thereby generating a cervical classification image; Step S5: Perform hierarchical joint modeling training based on the cervical classification images to obtain a visual basic model; perform image-text fusion modeling based on the visual basic model, perform task migration processing, and deploy it to the integrated system at the same time, thereby generating a cross-center deployment model system.
2. The training method of the multimodal cervical pathology image classification model according to claim 1, characterized in that: Step S1 includes the following steps: Step S11: acquiring a cervical tissue image; performing color normalization on the cervical tissue image to obtain a normalized cervical image; performing depth contrast enhancement processing on the normalized cervical image to generate an enhanced cervical image; Step S12: identifying cervical tissue structural features based on the enhanced cervical image, and performing spatial collaborative segmentation on the enhanced cervical image according to the cervical tissue structural features, thereby obtaining segmented cervical structure blocks; Step S13: performing kernel density mapping processing on the segmented cervical structure block to obtain kernel density distribution data; performing interstitial region separation based on the kernel density distribution data to generate interstitial distribution data; Step S14: decoupling the epithelial structure through the interstitial distribution data to obtain epithelial distribution data; fusing the epithelial distribution data, nuclear density distribution data and interstitial distribution data, and constructing a nuclear-interstitial-epithelial three-distribution framework based on the fused data.
3. The training method of the multimodal cervical pathology image classification model according to claim 1, characterized in that: Step S2 includes the following steps: Step S21: collecting historical data on the development of cervical tissue; performing time series label fusion on the three distribution frameworks based on the historical data on the development of cervical tissue, thereby generating time series three distribution framework data; Step S22: Track the evolution trajectory based on the time series three-distribution framework data to obtain the three-distribution evolution trajectory, where the minimum number of nodes contained in each evolution trajectory is not less than 4, and the drift distance of the structural blocks between nodes is controlled within 64 pixels. If the threshold is exceeded, the trajectory tracking chain is automatically interrupted; Step S23: constructing a three-distribution framework development lifeline through the three-distribution evolution trajectory; extracting the hierarchical development trend in the three-distribution framework development lifeline, wherein trend identification uses a sliding trend window with a width of 3 frames, and the trend turning point change slope must be greater than 0.4 to be defined as a valid mutation; Step S24: Perform trend distribution mapping based on the hierarchical development trends to confirm the predicted development trends of each layer in the three-distribution framework.
4. The training method of the multimodal cervical pathology image classification model according to claim 1, characterized in that: The cervical tissue environment field simulation using the cervical tissue image in step S3 includes: Extracting cervical environment data from cervical tissue images; Fitting the asymmetric tissue tension field to the cervical environmental data to obtain the environmental tension characteristic field; Injecting the horizontal-radial structural flow tensor into each tension sub-block of the environmental tension feature field, and performing residual conduction on the maximum gradient direction of the environmental tension feature field, thereby generating environmental direction sensing guidance data; The cervical tissue environment field is simulated based on the environmental direction sensing guidance data and the cervical environment data, thereby generating a simulated cervical environment field.
5. The training method of the multimodal cervical pathology image classification model according to claim 1, characterized in that: In step S3, the hierarchical evolution prediction of the three-distribution framework of nucleus, stroma, and epithelium is performed based on the simulated cervical environment field and the predicted development trend of each layer, including: Project the predicted development trend of each layer onto the three-distribution framework of nucleus, stroma and epithelium into the simulated cervical environment field to obtain the projected three-distribution data; Performing hierarchical dynamic migration simulation on the projected three-distribution data to generate hierarchical dynamic migration data; The inter-layer coupling response is inferred based on the layered dynamic migration data, thereby generating coupling response prediction data; The projected three-distribution data are evolution-driven according to the coupled response prediction data to obtain the evolution prediction data of each layer of the framework.
6. The training method of the multimodal cervical pathology image classification model according to claim 1, characterized in that: In step S4, the cervical tissue image is subjected to evolutionary projection according to the evolutionary prediction data of each layer of the framework, including: Perform spatial position encoding on the framework evolution prediction data of each layer to obtain position-encoded prediction data; Structurally projecting the position-encoded prediction data into the cervical tissue image to generate a structural projection image; The evolution state in the fusion structure projection image is obtained to obtain a state fusion image; Perform fine-grained evolution rendering on the state fusion image to obtain a fine-grained evolution mapping image; The pseudo color in the fine-grained evolution map image is enhanced to obtain the evolution map cervical tissue image.
7. The training method of the multimodal cervical pathology image classification model according to claim 1, characterized in that: In step S4, classifying the cervical tissue image according to the evolution process category based on the evolution mapping includes: Extracting process features from evolutionary mapped cervical tissue images; Clustering is performed based on process features, and image category feature labels are constructed; The evolutionary mapping cervical tissue images are classified based on the image category feature labels, thereby obtaining multiple evolutionary process image categories; Determining the category of the cervical tissue image corresponding to the evolutionary mapped cervical tissue image through the evolutionary process image category; The cervical tissue image is classified according to the category to which the image belongs, thereby generating a cervical classification image.
8. The training method of the multimodal cervical pathology image classification model according to claim 1, characterized in that: In step S5, hierarchical joint modeling training is performed based on the cervical classification image, including: Perform standard magnification block segmentation on the cervical classification image to obtain pre-cut cervical classification blocks; Automatically screen the tissue structure area of the cervical classification pre-cut image blocks to obtain tissue area blocks; Carry out coordinate index encoding processing on the tissue area blocks to obtain the basic data of the blocks; Perform coverage statistics based on basic tile data and select high-quality tiles based on coverage; Based on high-quality tiles, multi-center structures are uniformly processed and their labels are integrated to construct a large-scale training dataset; Perform self-supervised visual pre-training based on the DINOv2 architecture based on a large-scale training dataset to generate initial visual feature data; A basic visual model with the ability to represent organizational structure is constructed based on the initial visual feature data.
9. The training method of the multimodal cervical pathology image classification model according to claim 1, characterized in that: In step S5, image-text fusion modeling is performed based on the basic visual model, task migration is performed, and deployment to the integrated system includes: Match the visual feature outputs of the visual basic model with the image and text annotation samples one by one to obtain visual language pairing data; Perform image-text modality mapping and binding based on visual language pairing data to generate multimodal chimeric input data; Based on multimodal mosaic input data, image and text joint training is performed to obtain a multimodal basic model with alignment capabilities; The graphic and text basic model data and the visual basic model in the multimodal basic model are configured and module-encapsulated for 25 core task scenarios, thereby generating multi-task collaborative processing module data. Injecting the multi-task collaborative processing module data into the decision-making agent planner built by the language model to obtain task planning and scheduling control data; Perform modal conversion mapping processing based on task planning and scheduling control data to obtain unified reasoning path data; Based on the unified reasoning path data, structured interface encapsulation is performed and user interaction modules are connected to generate deployable integrated system data. The deployable integrated system data will be prospectively validated and evaluated in six independent centers, ultimately resulting in a cross-center deployment model system.
10. A training system for a multimodal cervical pathology image classification model, characterized in that: The method for training a multimodal cervical pathology image classification model according to claim 1 is configured to include: A three-part decomposition module is used to obtain cervical tissue images; perform tissue structure segmentation on the cervical tissue images to obtain segmented cervical structural blocks; and divide the cervical tissue into a three-part distribution framework of nuclear-stromal-epithelial based on the segmented cervical structural blocks; A trend identification module is used to collect historical data on the development of cervical tissue; identify the development lifeline of the three-distribution framework based on the historical data on the development of cervical tissue; and confirm the predicted development trend of each layer in the three-distribution framework through the development lifeline of the three-distribution framework; The environmental prediction module is used to simulate the cervical tissue environmental field using cervical tissue images to generate a simulated cervical environmental field; based on the simulated cervical environmental field and the predicted development trends of each layer, the layered evolution prediction of the nuclear-stroma-epithelial three-distribution framework is performed to obtain the evolution prediction data of each layer framework; An evolution classification module is used to perform evolution projection on the cervical tissue image according to the evolution prediction data of each layer of the framework, thereby generating an evolution mapping cervical tissue image; based on the evolution mapping cervical tissue image, the cervical tissue image is classified according to the evolution process category, thereby generating a cervical classification image; The fusion modeling module is used to perform hierarchical joint modeling training based on cervical classification images to obtain a visual basic model; based on the visual basic model, image and text fusion modeling is performed, and task migration processing is performed, and the module is deployed to the integrated system at the same time, thereby generating a cross-center deployment model system.
Citation Information
Patent Citations
Multi-type cell nucleus labeling and multi-task processing method for cervical TCT section
CN117496512A
Cervical tissue pathology all-slide image multi-classification method, system and device
CN118196516A
Cervical panoramic image few-sample classification method based on visual guidance and language prompt
CN118230052A
Cervical cancer diagnosis enhancing system based on artificial intelligence
CN118888127A
Cervical image processing method based on double attention and multi-scale fusion
CN118967718A
Cited By
Elevator monitoring image splicing method for behavior monitoring
CN121095057A
Elevator monitoring image stitching method for behavior monitoring
CN121095057B
Scoliosis identification method, device and equipment for back image
CN121353814A
Low-rank adaptation-based few-sample coronal mass ejection segmentation method and system
CN121616830A
Specimen data labeling method and system applied to water ecology laboratory
CN121637148A