Training method and system of multi-modal based cervical pathological image classification model

By using a three-distribution frame segmentation and environmental field simulation in a multimodal cervical pathology image classification model, the dynamic modeling problem of structural evolution in cervical pathology image analysis was solved, achieving cross-center stability and generalization ability, and improving the accuracy and consistency of image classification.

CN120808067BActive Publication Date: 2026-02-24GUANGZHOU JINRUI TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510893404.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2026-02-24
Estimated Expiration
2045-06-30

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately capture the evolution of tissue structures in cervical pathology image analysis, lack the ability to dynamically model fine-grained regions, and cannot provide time-consistent processing paths. This results in weak task transfer capabilities, limited generalization capabilities, and unstable cross-center performance after system deployment.

Method used

By using a multimodal cervical pathology image classification model, a three-distribution framework (nucleus-stromal-epithelium) is adopted for tissue structure segmentation and hierarchical modeling. Combined with environmental field simulation and image-text fusion, a cross-modal, cross-temporal, and cross-institutional visual-language joint modeling system is constructed to realize structural evolution perception and task transfer.

Benefits of technology

It improves the structural decoupling and multi-level channel modeling capabilities of cervical pathology image analysis, enhances the directional fitting accuracy and deformation rationality of image classification, ensures the stability and controllability of the model under different data domains and task label structures, and realizes cross-center structural consistency response and semantic-level generalization distribution feature regulation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808067B_ABST
    Figure CN120808067B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of image classification, and more particularly to a training method and system for a cervical pathological image classification model based on multi-modal. The method comprises the following steps: obtaining cervical tissue images and performing tissue structure segmentation to form a three-distribution framework of nucleus-stroma-epithelium, collecting development history data, confirming the prediction trend of each layer, simulating an environment field through images, generating a simulated cervical environment field, and performing layered evolution prediction on the framework, generating an evolution mapping image according to the evolution data and performing classification, and finally obtaining a visual basic model through joint modeling training, realizing image-text fusion and generating a cross-center deployment model system. The present application realizes visual-language joint modeling, improves the stability and controllability of the overall model structure in image space deformation modeling, semantic cross-modal alignment construction and task-level response flow scheduling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image classification technology, and in particular to a training method and system for a multimodal cervical pathology image classification model. Background Technology

[0002] Traditional cervical pathology image analysis methods largely rely on shallow feature extraction and single-label classification of static images. In actual multi-stage pathological tasks, they are difficult to accurately capture the evolution of tissue structures. Existing technologies generally use single-modal visual networks to perform full-image analysis of slice images, which cannot distinguish the structural distribution relationships between different tissue levels and lacks the ability to dynamically model fine-grained regions such as nuclear dense areas and epithelial boundaries. They lack the modeling logic for evolutionary trends in the evolution of cervical lesions, which can easily lead to ambiguous stage classification and accumulated structural recognition errors. Especially when faced with real data from multiple image sources, multiple task types, and diverse structural morphologies, existing technologies cannot provide a temporally consistent processing path, nor can they use image-text fusion to guide tasks for unstructured descriptions. Most systems lack modeling and response mechanisms for development processes, evolutionary trends, and regional hierarchical differences. At the same time, they are unable to generate reasonable hypothesis models for rare pathological states or unknown evolutionary paths, resulting in weak task transferability and limited generalization ability. After system deployment, they are prone to cross-center performance instability. Summary of the Invention

[0003] Therefore, it is necessary to provide a training method and system for a multimodal cervical pathology image classification model to solve at least one of the above-mentioned technical problems.

[0004] To achieve the above objectives, a training method for a multimodal cervical pathology image classification model includes the following steps:

[0005] Step S1: Obtain cervical tissue images; perform tissue structure segmentation on the cervical tissue images to obtain segmented cervical structural blocks; based on the segmented cervical structural blocks, divide the cervical tissue into a three-distribution framework of nucleus-stromal-epithelial.

[0006] Step S2: Collect historical data on cervical tissue development; identify the three-distribution framework development lifeline based on the historical data on cervical tissue development; confirm the predicted development trend of each layer in the three-distribution framework through the three-distribution framework development lifeline;

[0007] Step S3: Simulate the cervical tissue environment field using cervical tissue images to generate a simulated cervical environment field; based on the simulated cervical environment field and the predicted development trends of each layer, perform hierarchical evolution prediction of the three-layer distribution framework of nucleus-stromal-epithelium to obtain the evolution prediction data of each layer framework.

[0008] Step S4: Based on the evolution prediction data of each framework layer, perform evolutionary projection on the cervical tissue image to generate an evolutionary mapping cervical tissue image; based on the evolutionary mapping cervical tissue image, classify the cervical tissue image according to the evolutionary process category to generate a cervical classification image;

[0009] Step S5: Perform hierarchical joint modeling training based on cervical classification images to obtain a visual base model; perform image-text fusion modeling based on the visual base model, and perform task transfer processing, while deploying to the integration system to generate a cross-center deployment model system.

[0010] This invention, through structural segmentation and tri-distribution framework division of cervical tissue images, can construct a highly structured spatial information representation within the original pixel space of the image. This enhances the structural decoupling and multi-level channel modeling capabilities in subsequent processing, ensuring that different types of tissue structures have a traceable and separable high-resolution representation basis in the data processing flow. Utilizing historical data, it performs temporal fusion of structural distributions and identifies the evolutionary lifelines of the tri-distribution framework. This enables the reconstruction of the dynamic evolution path of spatially stable structures in discontinuous image frames, effectively completing the temporal consistency alignment processing of the tri-distribution tissues. This ensures that the tissue evolution trajectory possesses data continuity and structural discriminability. Furthermore, it extracts the hierarchical development trends of each layer of structure, providing directional change signals for subsequent tasks and improving the direction fitting accuracy and deformation rationality in the distribution prediction stage. Based on cervical tissue images, an environmental field simulation model is constructed. Through non-dynamic mechanisms, it completes the composite generation of tension perturbation field, structural pressure projection field, and trend extension array, ensuring that the structural simulation data balances realism, heterogeneous perturbation coverage, and spatial structure controllability. By driving the evolution prediction processing of the tri-distribution framework through environmental field data and trend data, it can introduce... By fitting the evolutionary perturbation flow, the simulated evolution results of each layer structure possess directional guidance and feature fidelity. The layered prediction results are then back-mapped to the original image domain to generate an evolutionary mapping image. This enables a full-process back-projection transformation of structural deformation and modal perturbation in high-resolution images, maintaining consistency between image detail texture information and channel distribution structure. Furthermore, introducing evolutionary process categories for image classification transforms potential trend changes into image-level structural class labels, providing a high-confidence input source for subsequent classification supervision. Hierarchical joint modeling training based on the classified images facilitates the construction of a structured... This paper constructs a visual foundation model for evolutionary perception capabilities and enhances the encoding ability of multimodal representation structures for contextual semantic cues and structural dynamic trajectories through fusion training on image and text inputs. Based on this model structure, task transfer processing and multi-center deployment are carried out, which can achieve structural consistency response and semantic-level generalization distribution feature regulation under different data domains and different task label structures. This enables the construction of a unified cross-modal, cross-temporal, and cross-institutional vision-language joint modeling system, improving the stability and controllability of the overall model structure in image space deformation modeling, semantic cross-modal alignment construction, and task-level response flow scheduling.

[0011] The present invention also provides a training system for a multimodal cervical pathology image classification model, used to execute the training method for the multimodal cervical pathology image classification model described above. The training system for the multimodal cervical pathology image classification model includes:

[0012] The three-part decomposition module is used to acquire cervical tissue images; the cervical tissue images are segmented into tissue structures to obtain segmented cervical structural blocks; based on the segmented cervical structural blocks, the cervical tissue is divided into a three-part distribution framework of nucleus-stromal-epithelial.

[0013] The trend recognition module is used to collect historical data on cervical tissue development; based on the historical data on cervical tissue development, it identifies the developmental lifeline of the three-distribution framework; and through the developmental lifeline of the three-distribution framework, it confirms the predicted development trend of each layer in the three-distribution framework.

[0014] The environmental prediction module is used to simulate the cervical tissue environment field through cervical tissue images, thereby generating a simulated cervical environment field; based on the simulated cervical environment field and the predicted development trend of each layer, the three-layer distribution framework of nucleus-mesenchyma-epithelium is predicted to evolve in layers to obtain the evolution prediction data of each layer framework.

[0015] The evolution classification module is used to perform evolutionary projection on cervical tissue images based on evolutionary prediction data of each framework layer, thereby generating evolutionary mapping cervical tissue images; based on the evolutionary mapping cervical tissue images, the cervical tissue images are classified according to the evolutionary process category, thereby generating cervical classification images;

[0016] The fusion modeling module is used to perform hierarchical joint modeling training based on cervical classification images to obtain a visual base model; based on the visual base model, image-text fusion modeling is performed, and task transfer processing is carried out, while deploying to the integration system to generate a cross-center deployment model system.

[0017] This invention constructs a three-channel structural framework of nucleus-mesenchyma-epithelium through a three-part decomposition module, which is beneficial to improve the spatial decoupling capability of tissue partitioning and the structural clarity of subsequent modeling. The trend recognition module extracts the structural evolution trajectory through historical data and establishes structural correspondences between multiple frames of images, enhancing temporal consistency and trend traceability. The environment prediction module generates a simulated environment field containing tension perturbations, pressure mapping, and trend arrays, realizing non-biological mechanism-driven modeling of the structural evolution process and improving the directionality and structural fidelity of evolution simulation. The evolution classification module generates mapped images based on predictions and classifies process categories, enabling image classification to have trend perception capabilities and distinguishable change stages. The fusion modeling module constructs a visual model through classified images and completes image-text alignment training, realizing unified modeling of tissue structure, trend semantics, and image-text modality. The system has a clear structure, separate modules, supports cross-modal input, task-level scheduling, and multi-center deployment, and has scalability, composability, and structural universality. Attached Figure Description

[0018] Figure 1 This is a flowchart illustrating the steps of training a multimodal cervical pathology image classification model.

[0019] Figure 2 This is a detailed flowchart illustrating the implementation steps of step S2;

[0020] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0021] The technical method of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0022] Furthermore, the accompanying drawings are merely illustrative of the invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor methods and / or microcontroller methods.

[0023] It should be understood that although the terms "first," "second," etc., may be used herein to describe various units, these units should not be limited by these terms. These terms are used merely to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, a first unit may be referred to as a second unit, and similarly, a second unit may be referred to as a first unit. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0024] To achieve the above objectives, please refer to Figures 1 to 2 A training method for a cervical pathology image classification model based on multimodality includes the following steps:

[0025] Step S1: Obtain cervical tissue images; perform tissue structure segmentation on the cervical tissue images to obtain segmented cervical structural blocks; based on the segmented cervical structural blocks, divide the cervical tissue into a three-distribution framework of nucleus-stromal-epithelial.

[0026] Step S2: Collect historical data on cervical tissue development; identify the three-distribution framework development lifeline based on the historical data on cervical tissue development; confirm the predicted development trend of each layer in the three-distribution framework through the three-distribution framework development lifeline;

[0027] Step S3: Simulate the cervical tissue environment field using cervical tissue images to generate a simulated cervical environment field; based on the simulated cervical environment field and the predicted development trends of each layer, perform hierarchical evolution prediction of the three-layer distribution framework of nucleus-stromal-epithelium to obtain the evolution prediction data of each layer framework.

[0028] Step S4: Based on the evolution prediction data of each framework layer, perform evolutionary projection on the cervical tissue image to generate an evolutionary mapping cervical tissue image; based on the evolutionary mapping cervical tissue image, classify the cervical tissue image according to the evolutionary process category to generate a cervical classification image;

[0029] Step S5: Perform hierarchical joint modeling training based on cervical classification images to obtain a visual base model; perform image-text fusion modeling based on the visual base model, and perform task transfer processing, while deploying to the integration system to generate a cross-center deployment model system.

[0030] This invention, through structural segmentation and tri-distribution framework division of cervical tissue images, can construct a highly structured spatial information representation within the original pixel space of the image. This enhances the structural decoupling and multi-level channel modeling capabilities in subsequent processing, ensuring that different types of tissue structures have a traceable and separable high-resolution representation basis in the data processing flow. Utilizing historical data, it performs temporal fusion of structural distributions and identifies the evolutionary lifelines of the tri-distribution framework. This enables the reconstruction of the dynamic evolution path of spatially stable structures in discontinuous image frames, effectively completing the temporal consistency alignment processing of the tri-distribution tissues. This ensures that the tissue evolution trajectory possesses data continuity and structural discriminability. Furthermore, it extracts the hierarchical development trends of each layer of structure, providing directional change signals for subsequent tasks and improving the direction fitting accuracy and deformation rationality in the distribution prediction stage. Based on cervical tissue images, an environmental field simulation model is constructed. Through non-dynamic mechanisms, it completes the composite generation of tension perturbation field, structural pressure projection field, and trend extension array, ensuring that the structural simulation data balances realism, heterogeneous perturbation coverage, and spatial structure controllability. By driving the evolution prediction processing of the tri-distribution framework through environmental field data and trend data, it can introduce... By fitting the evolutionary perturbation flow, the simulated evolution results of each layer structure possess directional guidance and feature fidelity. The layered prediction results are then back-mapped to the original image domain to generate an evolutionary mapping image. This enables a full-process back-projection transformation of structural deformation and modal perturbation in high-resolution images, maintaining consistency between image detail texture information and channel distribution structure. Furthermore, introducing evolutionary process categories for image classification transforms potential trend changes into image-level structural class labels, providing a high-confidence input source for subsequent classification supervision. Hierarchical joint modeling training based on the classified images facilitates the construction of a structured... This paper constructs a visual foundation model for evolutionary perception capabilities and enhances the encoding ability of multimodal representation structures for contextual semantic cues and structural dynamic trajectories through fusion training on image and text inputs. Based on this model structure, task transfer processing and multi-center deployment are carried out, which can achieve structural consistency response and semantic-level generalization distribution feature regulation under different data domains and different task label structures. This enables the construction of a unified cross-modal, cross-temporal, and cross-institutional vision-language joint modeling system, improving the stability and controllability of the overall model structure in image space deformation modeling, semantic cross-modal alignment construction, and task-level response flow scheduling.

[0031] In this embodiment of the invention, the training method for the multimodal cervical pathology image classification model includes the following steps:

[0032] Step S1: Obtain cervical tissue images; perform tissue structure segmentation on the cervical tissue images to obtain segmented cervical structural blocks; based on the segmented cervical structural blocks, divide the cervical tissue into a three-distribution framework of nucleus-stromal-epithelial.

[0033] In this embodiment, original whole-slide images (WSI) acquired from a multi-center data collaboration platform are retrieved as input samples. All slide images undergo uniform magnification normalization, with the scanning magnification uniformly set to 20x. Each WSI image is quantized and encoded using a scanning resolution with a spatial scale of 0.24 μm / pixel. The highest resolution layer of the WSI image is read using the OpenSlide tool, and block-level partitioning is performed. The partitioning window size is set to 2048 pixels × 2048 pixels, with a stride of 512 pixels. A pre-trained semantic segmentation model based on a U-Net structure is used to extract the foreground tissue region for each image block. The segmentation model has 3 input channels and 4 output channels, corresponding to the nuclear region, epithelial region, mesenchymal region, and background region. The output results of each image block are post-processed and optimized using a fully connected conditional random field (CRF) to obtain a continuous boundary structure mask. Then, the structure mask of each image block is divided into a three-channel distribution frame, mapping the nuclear region to the red channel, the epithelial region to the green channel, and the mesenchymal region to the blue channel. The three-channel tissue distribution tensor map corresponding to the image block is saved, and an index file is generated to record the patient ID to which the image block belongs, the spatial coordinates of the image block, and the pixel density information within the channel.

[0034] Step S2: Collect historical data on cervical tissue development; identify the three-distribution framework development lifeline based on the historical data on cervical tissue development; confirm the predicted development trend of each layer in the three-distribution framework through the three-distribution framework development lifeline;

[0035] In this embodiment, a case dataset containing follow-up time series is constructed. Each case includes at least three cervical tissue images acquired at different time points, with an image time interval of at least 6 months. Each image must be scanned at 20x magnification and retain structural integrity regions. An image structure hashing algorithm is used to filter the images for similarity, ensuring that each time point image has tissue structure change characteristics. Image structure matching is performed on the time series of each case. Block-level feature representations of three-distribution frames in each frame image are extracted using a multi-scale feature pyramid structure. The block center coordinates and texture vectors are used to construct a joint matching criterion, which is suitable for cases with small Euclidean distances. If the structural blocks have a temporal consistency of 20 pixels and a cosine similarity greater than 0.85, the structural development lifeline is constructed based on the spatial trajectory of the structural blocks. The regional behavior of the same structural block at three or more time points is statistically analyzed. The evolution trend is judged by the slope of the structural area change. The growth slope threshold is set to +0.15, the decline slope threshold is set to -0.15, and the stability range is set to [-0.15, +0.15]. The structural development type is marked as expansion type, contraction type and stable type respectively. The trend label of each structural block at each time point is recorded and used as input to the subsequent simulation modeling stage.

[0036] Step S3: Simulate the cervical tissue environment field using cervical tissue images to generate a simulated cervical environment field; based on the simulated cervical environment field and the predicted development trends of each layer, perform hierarchical evolution prediction of the three-layer distribution framework of nucleus-stromal-epithelium to obtain the evolution prediction data of each layer framework.

[0037] In this embodiment, a clearly defined cervical tissue image is selected as input data, and it undergoes tile-based processing. The tile size is set to 128 pixels × 128 pixels. For each tile, the internal tissue density distribution map and gradient direction map are calculated. The Sobel operator is used to extract the edge direction distribution information of the tile, and a tension field simulation map is constructed based on the density of the core region structure. Regions with a pixel density higher than 0.4 in the core region are assigned an initial tension value of 0.6, while other regions have a tension value of 0.2. The tension field structure is used as the basic input for the environmental tension feature map. Each tile is injected with a horizontal and a radial structural tensor. The horizontal tensor is oriented along the x-axis of the tile and has an intensity of 60% of the tension value. The radial tensor diffuses outward from the center of the tile. The intensity decrease function is set as the original tension value multiplied by the inverse square distance function. The maximum gradient direction is extracted for all tensor fusion maps. The dominant direction vector is calculated based on the tension gradient field, and residual-guided filling is performed on the dominant direction. The discontinuous structural directions are interpolated and completed using convolution kernels. Finally, the outputs of each tension sub-block are synthesized into a complete simulated cervical environment field map.

[0038] Step S4: Based on the evolution prediction data of each framework layer, perform evolutionary projection on the cervical tissue image to generate an evolutionary mapping cervical tissue image; based on the evolutionary mapping cervical tissue image, classify the cervical tissue image according to the evolutionary process category to generate a cervical classification image;

[0039] In this embodiment, the structural prediction results and the simulated environment field are used as joint input source data. Directional projection operations are performed on the core region, mesenchymal region, and epithelial region of the three-distribution framework, respectively. The translation position of the target image patch is calculated based on the structural trend vector and the dominant tension direction vector. The maximum translation distance is set to not exceed 16 pixels. An affine transformation matrix is ​​used to perform coordinate mapping on the structural image patch and rewrite it into the original image coordinate system. The reconstructed image is subjected to pixel value smoothing and resampling processing to maintain boundary continuity and texture consistency. After the mapping is completed, each image patch is assigned an evolutionary process category according to its evolutionary trend label, including three categories: expansion, contraction, and stability. The change value of the texture density in the central region is used as the classification basis. Those with a change rate greater than 20% are marked as expansion type, those less than -20% are marked as contraction type, and those with an absolute change rate less than 5% are marked as stability type. All image patches are used to construct a training image set according to the labels, and the image size is uniformly 224 pixels × 224 pixels.

[0040] Step S5: Perform hierarchical joint modeling training based on cervical classification images to obtain a visual base model; perform image-text fusion modeling based on the visual base model, and perform task transfer processing, while deploying to the integration system to generate a cross-center deployment model system.

[0041] In this embodiment, ViT (Vision Transformer) is used as the backbone structure of the image encoder. Each classification image is divided into patches, with each patch being 16 pixels × 16 pixels. Each image can be divided into 196 patches. A linear transformation is used to map each patch into a 768-dimensional vector embedding, and positional encoding is added. The embeddings are then fed into a 12-layer Transformer module, with 12 self-attention heads configured in each layer. The DINOv2 self-supervised pre-training mechanism is used to train the basic visual encoder CerS-V under the optimization objectives of image contrast alignment and feature consistency. The optimizer is AdamW, with an initial learning rate of 5e-4. The training batch size is 300 epochs, with no less than 2048 images per epoch. After training, the visual encoder weights are frozen, and its output features are used as input for image-text modeling. LoRA (Low-Rank) is then applied. The Adaptation method connects the visual output to the multimodal embedding layer of the large language model Qwen2.5-VL through a 128-dimensional low-rank adapter, constructing the image-text fusion basic model CerS-M. On the fusion model, image-text alignment instruction datasets containing task types such as structural description, classification labels, and dialogue question answering are used to perform instruction fine-tuning. During training, the language model parameters are frozen, and only the parameters of the visual connection module and instruction decoding head are updated. After the model is trained, it is encapsulated as a multimodal deployment module and an API interface is built and bound to the integration platform. The platform supports heterogeneous image input, task instruction loading, and structural prediction result output, and provides a structured reporting interface and a system task management console for deployment hospitals to use. Finally, a cross-modal structural classification system model that can be deployed and used in a multi-center environment is formed.

[0042] Preferably, step S1 includes the following steps:

[0043] Step S11: Obtain cervical tissue image; normalize the color of cervical tissue image to obtain normalized cervical image; perform depth contrast enhancement processing on normalized cervical image to generate enhanced cervical image.

[0044] Step S12: Identify cervical tissue structure features based on enhanced cervical images, and perform spatial collaborative segmentation on the enhanced cervical images according to the cervical tissue structure features to obtain segmented cervical structure blocks;

[0045] Step S13: Perform kernel density mapping on the segmented cervical structural blocks to obtain kernel density distribution data; perform stromal region separation based on the kernel density distribution data to generate stromal distribution data;

[0046] Step S14: Decouple the epithelial structure using mesenchymal distribution data to obtain epithelial distribution data; fuse the epithelial distribution data, nuclear density distribution data, and mesenchymal distribution data, and construct a three-distribution framework of nuclear-mesenchymal-epithelial based on the fused data.

[0047] In this embodiment, cervical whole-slice images scanned at 20x magnification from multiple medical image centers are used as input image sources. Each image has a resolution of 0.24 micrometers per pixel, 3 image channels, and is in TIFF format. First, the Macenko normalization algorithm is used to standardize the colors of the input images. When constructing the color normalization function, the image needs to be decomposed into a channel matrix to extract the absorption spectral feature values ​​of each channel. Then, the normalization matrix is ​​applied to the original image's color space transformation. After completing the target color standard template mapping, a normalized cervical image is generated. Finally, the normalized cervical image is used as input to perform deep learning. The contrast enhancement operation performs contrast stretching within image blocks based on a local contrast enhancement algorithm. The image block size is set to 64 pixels × 64 pixels, and the stretching factor is set to 1.4. Adaptive local histogram equalization is performed on the kernel region of the image to enhance the gray-level contrast between the kernel and non-kernel regions. The final output is an enhanced cervical image, which serves as input data before structural segmentation. Multi-scale tissue structure features are extracted using an image feature extraction model based on the EfficientNet-B3 backbone network. The model's input image size is 512 pixels × 512 pixels. After feature extraction, a spatial attention mechanism module is constructed. The module constructs a structure weight matrix based on the spatial density of structural regions in the image. The weight matrix has the same dimension as the input image and a value range of 0 to 1, representing the response degree of the corresponding pixel to the structural region. This structural feature map is then input into a multi-channel spatial collaborative segmentation network. The segmentation network adopts an Encoder-Decoder structure based on depthwise separable convolution, with 3 output channels, corresponding to the predicted maps of the nuclear region, epithelial region, and mesenchymal region, respectively. After processing the maximum value channel of the output results, three-class segmentation maps are generated. Dilated convolution smoothing is performed on each channel output, and cross-supervised learning is combined with the structural feature map to output the final result. Images of segmented cervical structural blocks were generated in a three-channel floating-point image format. Each channel recorded the response intensity and spatial location index of the structural region. The center point positions of all kernel blocks were extracted using a kernel localization network. A kernel density map was generated using a density estimation-based heatmap construction method. In the calculation formula, each kernel position was used as the center point of a two-dimensional Gaussian function with a standard deviation of 4 pixels. The Gaussian responses of all kernel points were superimposed to generate the kernel density map. The density map was a single-channel grayscale image with a value range of 0 to 1, representing the kernel structural density intensity per unit area. Subsequently, a density threshold separation method was used to extract the stroma region based on the kernel density map, separating regions with densities below 0.Region 15 is defined as a candidate interstitial region. Small pseudo-interstitial regions are removed using boundary connectivity constraints. Morphological closing operations are used to complete region discontinuities. The generated interstitial distribution map is a binary mask image, with black areas corresponding to the nuclear and epithelial regions, and white areas corresponding to the extracted interstitial distribution data. The center coordinates, boundary length, and connectivity component index of each interstitial block are recorded in this image. After obtaining the interstitial distribution data, the epithelial structure is extracted using a spatial structure decoupling method. During decoupling, the epithelial structure is defined as a region with a nuclear density higher than a threshold and not a central nucleus. A connected component labeling algorithm is used to label all connected nuclear regions in the nuclear density map, excluding central nuclei with a diameter greater than 200 pixels. The densely packed nuclear bands around the epithelial region are retained as candidate epithelial regions. Logical operations are then performed on these candidate regions and the mesenchymal distribution data. A pixel-by-pixel comparison is made between the nuclear density map and the mesenchymal mask map, retaining only the nuclear regions adjacent to the mesenchymal region as the final epithelial distribution. The output image is a single-channel grayscale mask image representing the epithelial region distribution. When constructing the three-distribution framework, the obtained epithelial distribution data, nuclear density distribution data, and mesenchymal distribution data are mapped to the RGB channels respectively: the nuclear region is written to the red channel, the mesenchymal region to the green channel, and the epithelial region to the blue channel. The merged three-distribution framework image retains the original image size and records the structural label information corresponding to each pixel.

[0048] Preferably, step S2 includes the following steps:

[0049] Step S21: Collect historical data on cervical tissue development; perform time-series label fusion on the three-distribution framework based on the historical data on cervical tissue development to generate time-series three-distribution framework data;

[0050] Step S22: Perform evolution trajectory tracking based on the time-series three-distribution framework data to obtain the three-distribution evolution trajectory. Each evolution trajectory contains a minimum of 4 nodes, and the drift distance between structural blocks between nodes is controlled within 64 pixels. If this threshold is exceeded, the trajectory tracking chain will be automatically interrupted.

[0051] Step S23: Construct a three-distribution framework development lifeline through the three-distribution evolution trajectory; extract the hierarchical development trend in the three-distribution framework development lifeline, wherein the trend identification adopts a sliding trend window width set to 3 frames, and the slope of the trend inflection point change must be greater than 0.4 to be defined as a valid mutation;

[0052] Step S24: Perform trend distribution mapping based on hierarchical development trends to confirm the predicted development trends of each layer in the three-distribution framework.

[0053] In this embodiment, continuous time-series slide images from 1250 cervical cases across six pathology centers were selected. Each case included at least three tissue images acquired at different time points, with a time span between 6 and 36 months. All images were digitized using a unified scanning instrument at 20x magnification, in SVS or TIFF format, with a resolution of 0.24 micrometers per pixel. After performing standardized resampling and spatial normalization on all slides, they were numbered according to their timestamps. The three-distribution framework structure diagram constructed in the preceding steps was used as the structural input for each image frame. Structural label alignment was performed on all time-series images. During the alignment process, a probabilistic fusion model based on structural consistency (Structure Consistency Fusion Module) was used to perform channel-level annotation fusion of the three-distribution structural mask for each image frame. The region matching confidence threshold was set to 0 during the structural alignment process.75. Regions below this threshold will be discarded. The fused output three-channel structure map is sorted by timestamp to form a temporal three-distribution framework data. The input data is the temporal three-distribution framework data after time series labeling. During the trajectory tracking stage, tracking is performed on a block-by-block basis. The center coordinates, principal axis direction, and channel label of each block are used as node definition parameters. A joint matching algorithm based on spatial proximity and cosine similarity of structural features is used to perform bidirectional pairing matching of blocks in adjacent frames. The matching distance control threshold is set to 64 pixels. If the Euclidean distance between two blocks exceeds this threshold, the trajectory connection is immediately interrupted. Each trajectory is required to contain structural block nodes from at least four consecutive frames. The tracking algorithm enforces directional consistency constraints. If the principal axis direction of a structural block changes by more than 45 degrees, it is classified as a discontinuous evolution path and a trajectory chain break operation is performed. All valid trajectories are numbered, and the image frame number, structural type, and center coordinates of each node are recorded, ultimately forming a three-distribution evolution trajectory library. Each trajectory corresponds to a unique structural number within its lifecycle. The input data is time-series labeled three-distribution framework data. During the trajectory tracking phase, tracking is performed on a structural block basis. The center coordinates, principal axis direction, and channel coordinates of each structural block are recorded. Using tags as node definition parameters, a joint matching algorithm based on spatial proximity and structural feature cosine similarity is employed to perform bidirectional pairing matching of structural blocks in adjacent frames. The matching distance control threshold is set to 64 pixels. If the Euclidean distance between two structural blocks exceeds this threshold, the trajectory connection is immediately interrupted. Each trajectory is required to contain structural block nodes from at least four consecutive frames. The tracking algorithm enforces directional consistency constraints. If the principal axis direction of a structural block changes by more than 45 degrees, it is classified as a discontinuous evolution path and a trajectory chain break operation is performed. All valid trajectories are numbered, and the image frame number, structural type, and center coordinate of each node are recorded. The goal is to ultimately form a three-distribution evolution trajectory library. Each trajectory has a unique structure number within its lifecycle and is continuously called in subsequent stages. All trajectory sequences in the three-distribution evolution trajectory library are used as basic data. The structural attributes of each trajectory sequence are extracted by sorting them by time and the trend of change over time is extracted. A univariate structural index sequence is constructed for each trajectory. The index includes structural area, gray-scale mean, and density within the channel. A sliding trend window is used to extract the trend of the index sequence. The sliding window size is set to 3 frames, that is, every three consecutive time points are used as a group for local slope calculation. When the slope of change between any two frames exceeds 0.4 o'clock is defined as a structural abrupt change point. The location of the abrupt change point is recorded as a turning frame index. Continuous trends in the same direction are merged into a single trend segment. Trend segments are labeled as expansion, stability, or contraction, and the start and end frame indices and trend direction values ​​of each trend segment are recorded. The trend direction value is the slope value of the structural index change. The lifeline data structure includes trajectory number, structural label, number of trend segments, trend segment direction value array, and corresponding frame range identifier. All results are stored in a structured trend library. The labeled trend segment data in the structured trend library is called, and each type of structure (nuclear region, mesenchymal region, epithelial region) is analyzed separately. The frequency and direction of trend types appearing in all trajectories were calculated to construct a trend probability matrix. Each element in the matrix represents the probability value of the trend direction of the corresponding structure within a specific frame range. The probability value was calculated by weighted normalization of the trend segment direction values. The segment with the highest probability was marked as the dominant trend segment. The trend value of each frame in the dominant trend segment was encoded as the final predicted development trend. The encoding structure format was a three-channel structural trend map, with each channel corresponding to the trend direction value of the nuclear region, mesenchymal region, and epithelial region. The image size was consistent with the original slice, and the trend value range was normalized to [-1, 1].

[0054] Preferably, step S3, which involves simulating the cervical tissue environment using cervical tissue images, includes:

[0055] Extract cervical environment data from cervical tissue images;

[0056] Asymmetric tissue tension field fitting was performed on cervical environmental data to obtain the environmental tension characteristic field;

[0057] A horizontal-radial structural flow tensor is injected into each tension sub-block of the environmental tension characteristic field, and residual conduction is performed on the maximum gradient direction of the environmental tension characteristic field to generate environmental orientation sensing guidance data.

[0058] Based on environmental orientation sensing guidance data and cervical environment data, a cervical tissue environment field simulation is performed to generate a simulated cervical environment field.

[0059] In this embodiment, a raw whole-slice cervical image with a resolution of 0.24 micrometers per pixel and a magnification of 20x is selected as the input data source. The image undergoes tile-based preprocessing, with the tile size set to 512 pixels × 512 pixels. Batch sliding window processing is then performed with a sliding step size of 128 pixels. An image tone region extraction function based on the HSV color space is used to filter the background region. For the retained tiles, the Gray-Level Co-occurrence Matrix (GLCM) method is used to calculate local texture indices, including energy value, moment of inertia, entropy value, and contrast. Foreground regions with drastic texture changes and tissue differentiation are selected. The regions with distinctive features are used as environmental data extraction blocks. Simultaneously, independent gray-level histogram analysis is performed on the intermediate channel regions of the blocks. The kurtosis and skewness values ​​of the histograms are used to determine the tissue arrangement direction and density trend within the region. Blocks meeting the criteria of kernel density below 0.25 and mean texture contrast greater than 0.6 are included in the cervical environmental data set. The blocks extracted in the previous step are used as input for tension feature fitting. Local gradient direction extraction and gray-level structural response fusion processing are performed on each block. First, the Sobel gradient operator is used to calculate the local gradient maps in the x and y directions respectively. Then, the local structural response function is used to fuse the gray-level features. The degree distribution and gradient direction are weighted and integrated, and the structural response function is set as R(i,j)=αGx(i,j)+βGy(i,j)+γI(i,j), where α, β, and γ represent the harmonic coefficients of gradient direction and gray-level response, with values ​​of 0.4, 0.4, and 0.2, respectively. Orientation reversal symmetry detection is performed on all patches. Regions with a symmetry center drift greater than 16 pixels or a gradient angle difference exceeding 45 degrees are marked as asymmetric tension regions. Finally, an environmental tension feature field map is generated based on the orientation field results of each patch. The image format is a two-channel tensor map, with the first channel recording... For each pixel, the tension direction angle is recorded. The second channel records the tension amplitude value. The tension direction range is limited to 0 to 180 degrees, and the tension amplitude is normalized to [0,1]. The tension feature field map is divided into 128-pixel × 128-pixel tension sub-block regions. For each sub-block, horizontal structure flow tensor and radial structure flow tensor injection operations are performed. The direction of the horizontal structure flow tensor is fixed to the positive x-axis within the sub-block, and the intensity is set to 0.5. The radial structure flow tensor expands symmetrically in four quadrants with the center of the sub-block as the origin. Its tension value gradually decreases from the origin to the edge, with the maximum intensity set to 0.8 and the decrease factor being 0.12. After pixel-level superposition of the two types of tensors to form a merged dual-tensor image, the maximum gradient direction is calculated for each tension sub-block. A 5×5 convolutional window is used on the orientation field map to search for local maxima directions, extracting the dominant gradient direction vector. For positions with orientation jumps greater than 30 degrees, residual conduction is performed. Specifically, a 3×3 structure-guided convolution kernel is constructed to perform gradient field-guided diffusion on the tensor orientation map, repairing interrupted regions of the dominant direction and forming continuous induction links. The final output is a structure tensor flow guided image, where each pixel contains a complete tensor field and orientation vector value. Using the above structure tensor flow guided image and the original cervical environment data as input, a tension perturbation simulation model is constructed for each image region. The perturbation simulation uses the orientation-guided tensor as the core injection source and employs a perturbation superposition model. The diffusion process of the tension perturbation field is realized. The perturbation injection range is controlled within 64 pixels × 64 pixels. The perturbation intensity is set according to the change in structural density, with the perturbation intensity set to 0.7 in the core region, 0.5 in the epithelial region, and 0.3 in the mesenchymal region. During perturbation injection, a three-layer perturbation flow field channel is constructed, and directional consistency control is applied to each channel. Local tensor similarity filtering is used to guide and harmonize the perturbation direction. A Gaussian smoothing function is introduced into the structural boundary region to achieve perturbation buffering transition. After all perturbation channels are superimposed, a simulated cervical environment field map is generated. The image size is consistent with the original tissue patch. The output image is a multi-channel tensor map, including the original environmental texture map, the directional guidance map, the structural perturbation map, and the channel distribution mask map.

[0060] Of particular importance is the injection of a horizontal-radial structured flow tensor into each tension sub-block of the environmental tension characteristic field, and the residual conduction of the maximum gradient direction of the environmental tension characteristic field, including:

[0061] The environmental tension feature field is divided into multiple non-overlapping tension sub-blocks, and a tension sub-block index is constructed based on these multiple non-overlapping tension sub-blocks.

[0062] Based on the tension sub-block index, a horizontal-radial structure flow tensor is injected into each non-overlapping tension sub-block to obtain the structure tensor superposition data.

[0063] Tensor continuity balancing is performed between blocks based on the superimposed structural tensor data to obtain tensor fusion data;

[0064] Extract the direction of maximum gradient from the tensor fusion data to generate the dominant structural flow direction;

[0065] Extract high-gradient abrupt change regions in the dominant structural flow direction;

[0066] By completing the coherence of high-gradient abrupt change regions, environmental orientation sensing guidance data is generated.

[0067] In this embodiment, the input data is an environmental tension feature field tensor map generated in the previous stage. The tensor map has a size of 2048 pixels × 2048 pixels. To ensure the structural alignment of subsequent tensor injection, the tension feature field is divided into equally spaced blocks with a division scale of 128 pixels × 128 pixels. A sliding window is used to perform block traversal, with a window sliding step of 128 pixels to ensure no overlapping areas. A unique block number is assigned to each sub-block, and a two-dimensional index mapping matrix is ​​established. This matrix records the starting coordinate position, average tension, principal axis angle, and structure of each sub-block in the original tensor map. A regional channel label is constructed. This tension sub-block index structure is used for precise block-level positioning and tensor propagation path management in the subsequent structural tensor injection process. The index information is stored in a structured JSON format, with key-value pairs including sub-block ID, center coordinates, tension direction statistical histogram, and density grading label. This forms a block-level injection control interface for the environmental tension feature field. Horizontal and radial structural tensor templates are constructed separately. The horizontal tensor direction is the positive x-axis, and the tensor intensity is symmetrically distributed along the sub-block centerline, with a maximum value of 0.75, decreasing to 0.3 at the edge lines. An exponential decay function is used for intensity fitting. The radial tensor... The origin is used as the geometric center of the sub-block. The Euclidean distance from each pixel to the center is calculated, and a radial intensity distribution map is constructed based on a linear normalization function. The maximum tensor value is set to 0.8, and the minimum value to 0.2. The two structure flow tensors are superimposed at the pixel layer and merged into a unified tensor field matrix. This tensor field retains the vector components of the projected tension in both directions for each pixel. The result of this tensor superposition is written into the sub-block region of the original tension feature field map. All sub-blocks undergo the same structure flow injection operation, ultimately generating a structure tensor superposition data tensor map. Each pixel contains a two-dimensional vector structure value and a region structure tag. In the stage of balancing the inter-block tensor continuity based on the superimposed structural tensor data, an inter-block boundary coordination module is constructed to perform structural tensor edge docking operations on each pair of adjacent tension sub-blocks. First, a 3-pixel-wide intersection region is established at the adjacent boundary position, and the tensor direction difference in this region is calculated. If the direction angle is greater than 30 degrees, it is considered that there is a tensor discontinuity. A transition tensor band is introduced into the discontinuous region, and the tensor value is set as the weighted average direction of the tension direction between the two blocks, with an intensity of 90% of the lower value of the tensor intensity on both sides. At the same time, a Gaussian blur kernel is used for smoothing in the region, and the blur kernel radius is set to σ = 1.5. To ensure a smooth transition in the tension direction, after balancing the boundary intersections of all tension sub-blocks, the entire tensor map is merged to generate a tensor fusion data image. This image establishes inter-block continuity responses based on the original tensor structure, forming a unified structural expression of the tensor direction field across the entire image space. Gradient response scanning is performed on the tensor fusion image, using a 5×5 Sobel direction response kernel to extract the tensor direction distribution gradient at each pixel. The rate of change of direction values ​​in the x and y directions at each point is calculated. For each pixel, its tensor gradient amplitude and direction angle are calculated. The direction of maximum gradient corresponds to the position direction angle where the amplitude value is locally maximum. The dominant structural flow direction image is subjected to directional filtering to suppress regions with gradient angle changes less than 10 degrees, retaining regions with significant directional abrupt changes and establishing a directional index map. The final output image contains a dominant structural flow direction value for each pixel. This directional map is stored as a single-channel directional angle map, with the angle range limited to 0 to 180 degrees, serving as the directional basis for subsequent high-gradient abrupt change region extraction and connectivity repair. A local difference window is used to calculate the directional jump image. A 3×3 neighborhood window is established around each pixel, and the maximum and minimum differences in direction values ​​within this region are calculated. Pixels with differences greater than 45 degrees are marked as abrupt change boundary points, and boundary connectivity analysis is performed on these boundary points. All mutation blocks are labeled using a four-neighbor connectivity component. Discrete regions with an area less than 25 pixels are excluded, and significant mutation connected regions are retained as high-gradient mutation regions for output. A mutation mask map is constructed, where each pixel in the image has a value of 0 or 1, indicating whether it belongs to a mutation region. The boundaries of the mutation regions are vectorized and encoded. The principal direction vector of the mutation region is extracted and used for directional constraints in subsequent conduction operations. The mutation region mask map is expanded by boundary interpolation. The principal direction vector path is fitted to the mutation boundaries using Bézier curves, and a linear guiding path based on the principle of directional consistency is established on each mutation path. Tensor direction completion is inserted along the fitted curve direction at each break point in the path. The transition tension vector in the interpolation region is calculated using trilinear interpolation. The interpolation tensor strength is set to the average tension value of two consecutive points. A one-dimensional directional diffusion operation is performed on each completion path, with a diffusion width of ±2 pixels. Blur directional fusion is performed on the diffusion region to alleviate the discontinuity problem at the interpolation edges. All completion results are merged into the dominant structural flow direction map to form a complete and coherent directional guidance image. The final output is an environmental directional sensing guidance data map. Each pixel in the image contains the tension direction angle and the structural response intensity value. The format is a dual-channel floating-point tensor map, with the first channel being the direction angle map and the second channel being the directional sensing intensity map. The image size is the same as the original tension features. Figure 1 The system has two channels, uses float32 data type, and retains four decimal places for each channel to improve the accuracy of the conduction path and the stability of the fitting.

[0068] Preferably, step S3, which involves predicting the hierarchical evolution of the nucleus-mesenchyma-epithelial three-layer distribution framework based on the simulated cervical environment field and the predicted development trends of each layer, includes:

[0069] The predicted development trends of each layer are projected onto the three-distribution framework of the nucleus-mesenchyma-epithelium into the simulated cervical environment field to obtain projected three-distribution data;

[0070] Hierarchical dynamic migration simulation is performed on the projected three-distribution data to generate hierarchical dynamic migration data;

[0071] Inter-layer coupling response inference is performed based on hierarchical dynamic migration data, thereby generating coupling response prediction data;

[0072] Evolution-driven processing of the projected tri-distribution data is performed based on the coupled response prediction data to obtain the evolution prediction data of each frame layer.

[0073] In this embodiment, the input data is a three-channel image from the structural distribution map. Each channel corresponds to the nuclear region channel, the mesenchymal region channel, and the epithelial region channel, respectively. The corresponding predicted development trend is provided by the trend distribution map of the previous stage. Each channel is accompanied by a structural trend tensor, and the tensor dimension and distribution are specified. Figure 1The value range is [-1, 1], where positive values ​​represent expansion trends, negative values ​​represent contraction trends, and zero values ​​represent a stable state. The simulated cervical environment field is a multi-channel image, which includes a direction-guiding tensor map, a tension perturbation map, and a structural localization channel map. After pixel-level alignment processing of the three distribution maps and the environment field, the target migration direction of each pixel is calculated based on the trend tensor. The migration direction is based on the dominant direction angle in the environmental direction sensing map. The trend value is multiplied by a standard offset coefficient, which is set to 8 pixels. An affine projection transformation is performed on each pixel of the three distribution maps. The trend intensity of the nuclear region is multiplied by a coefficient of 1.2, the stroma region by a coefficient of 1.0, and the epithelial region by a coefficient of 0.9, which constitutes the trend-guiding direction offset. The image is shifted, and the offset positions are redrawn in the image as a projected structural map. Bilinear interpolation remapping is performed on all pixels to ensure structural edge continuity. The final output image is a projected tri-distribution data map. Each pixel records the original structural label, trend direction vector, and projection position coordinates. The core, mesenchymal, and epithelial regions are separated from the projected tri-distribution map into independent structural channel maps, and hierarchical structural migration units are established. A dynamic evolution process is constructed for each channel using continuous frame simulation, with a simulation time step set to Δt = 1. Local tensor response adjustment is performed on each structural block at each time step. The migration path direction is provided by the tension direction sensing channel in the environmental field, and the center point of each structural block is... A 1-step offset is performed in its local directional gradient field. The step size is weighted and adjusted according to the structural trend value, where the step size is 4 pixels when the trend value is 0.5 and -4 pixels when the trend value is -0.5. The increment and decrement are linearly proportional. The time window is set to 6 frames. For each frame, the coordinates, contour boundaries, and pixel grayscale images of the structural blocks are recorded. Each pixel is labeled with its original structural type and current migration direction. All frame sequences are combined to form hierarchical dynamic migration data. For each channel, independent nuclear region migration map, mesenchymal region migration map, and epithelial region migration map are formed. The time axis is unified to maintain the consistency of the structural block index. Structural index reconstruction is performed on the dynamic migration maps of the three channels. Each structural block is labeled on the time axis. The ratio of spatial location change vector to area change is used to extract structural boundary relationships between the core and mesenchymal regions, and between the mesenchymal and epithelial regions, at the same time step. Inter-block connectivity edges are established using a spatial adjacency graph. A coupling index function is constructed for structural pairs on the connectivity edges. This function consists of three parts: the contact area ratio, the trend direction difference, and the local tension field direction difference. A function threshold greater than 0.6 is considered a coupling pair. All coupling pairs undergo structural interaction calculations in the time series. The coupling response intensity is simulated using a cross-transfer convolutional network with a kernel size of 5×5 and 3 channels. The inputs are the boundary direction maps and trend maps of adjacent structural blocks, and the output is a response intensity map, where the intensity is greater than 0.Region 3 is recorded as the effective coupling response region, ultimately forming a coupling response prediction data map. This map records the coupling type, response intensity, and direction of action for each pixel. Using the coupling response prediction map as the driving source, a response mapping matrix is ​​constructed for each structural block in the projected three-distribution map. The matrix dimension is the tension sensing area range of the structural block region, with a range diameter set to 64 pixels. The driving direction is based on the direction vector field provided in the coupling response map, and the driving intensity is the response intensity value multiplied by the structural trend value. A driving influence factor function is established, and affine transformation superposition is performed on each structural block. High-response regions are extended in direction, and low-response regions are compressed in area. Structural point set mapping is performed on all structural regions, and pixel region boundaries are reconstructed. After structural correction, all structural blocks are merged back into a three-channel image, corresponding to the nuclear, mesenchymal, and epithelial regions, respectively. Finally, the evolution prediction data map of each layer of the framework is obtained. Each channel image retains the pixel label, response direction value, and area change coefficient, and simultaneously records the structural block coordinate index and simulation time step sequence number, which are used as the structural inference input layer for subsequent evolution mapping and image classification generation steps.

[0074] Of particular importance is the ability to infer inter-layer coupling response based on hierarchical dynamic migration data, including:

[0075] Reconstruct cross-layer indexes in hierarchical dynamic migration data, and confirm inter-layer structural data based on cross-layer indexes;

[0076] Interlayer tensor interferometry is performed on the interlayer structural data based on the cross-layer index to obtain coupled interferometric data;

[0077] The direction of interlayer interference was confirmed based on coupled interferometric data;

[0078] The interlayer interference direction is projected onto the time sequence preceding each time node to obtain the reverse response path;

[0079] Directional response values ​​are aggregated based on the reverse response path to generate coupled response prediction data.

[0080] In this embodiment, the input data consists of dynamic migration sequence image data generated by three structural channels: dynamic images of the nuclear region, the mesenchymal region, and the epithelial region. The image size is uniformly 512 pixels × 512 pixels. Each channel at each time step includes a structural block identifier, pixel mask, structural centroid coordinates, and trend direction vector. A cross-channel spatial correspondence is constructed through spatial overlap. Inter-channel overlap analysis is performed on the three distribution maps at any given time point, using IOU (Intersection over Union). The IOU index calculates the pairing relationships between structural blocks in the nuclear region and the mesenchymal region, and between the mesenchymal region and the epithelial region. Connections are established for structural pairs with an IOU value greater than 0.3, and a cross-layer structural pair index table is constructed. Each structural pair records the structural block ID, channel number, structural outline, center point distance, and structural area difference. The index table uses the timestamp as the primary key and the structural pair ID as the subkey, constructing a cross-layer structural data mapping to record structural interaction relationships and their temporal topological state. After constructing the cross-layer structural index, tensor interferometry is used to perform tensor interference processing on the interlayer structural data. Each structural pair is treated as an interference unit, and a tension vector action path is constructed within the structural pair. The radial direction is defined by the centroid vector from the upper structure to the lower structure. An interference channel region with a width of 16 pixels is set around this path. Tensor fusion calculation is performed on the tension fields of the upper and lower structures within the region. The tensor values ​​are derived from the structural trend direction map and tension guidance map in the structural dynamic migration map. The fusion is performed using a pixel-by-pixel tension vector superposition rule. Vectors with a directional difference angle of less than 20 degrees are processed by vector merging and superposition. Vectors with a directional difference angle of more than 20 degrees are decomposed and weighted attenuated. The generated fused tensor is called the interference tensor. The interference tensor field is stored according to pixel position. The tensor intensity normalization range is [0.0, 1.0], and the interference intensity is less than 0.Region 2 was marked as a weak response region. The final output of the interference tensor for all structure pairs was a coupled interference data image. The data structure was a four-dimensional tensor, with dimensions including time step, x-coordinate, y-coordinate, and tensor vector value. After generating the coupled interference data, a dominant interference direction identification operation was performed on the interference tensor image at each time step. The maximum response direction within the tensor field of each pixel in the image was calculated; this direction is the dominant interference direction. The dominant direction was extracted by finding the resultant force direction of the tensor vectors. After statistical analysis of all directions, the overall dominant direction angle for each structure pair was calculated, with a range of [0, 180] degrees. Each dominant direction angle is written into the orientation mapping map according to the structural block region to generate an interlayer interference orientation map. This map has 1 image channel and records the current structural coupling relationship of each pixel and its dominant interference direction value. The interference orientation map serves as the orientation control image for subsequent temporal back-projection. After the interference orientation map is constructed, a temporal back-projection operation is performed on the orientation map to obtain the reverse response path of the structural motion trend. For the orientation map at each time node t, the interference direction is projected to the corresponding coordinate position in the previous time frame t-1. The mapping method is to start from the current pixel position in frame t and reverse along the dominant direction vector. The image is shifted forward by 'd' pixels, with 'd' set to 4 pixels. In frame 't-1', the reverse path marker is recorded at the corresponding projection point position, and a reverse response path map is constructed. This path map is based on the structure block number and indexed by the direction vector, marking the projected positions as potential response positions. After back-projection of all structure blocks, a full-image directional response path map is constructed. The path direction and transfer tensor identifier are recorded at each pixel position in the image. A directional response value aggregation operation is performed based on the directional response path map. Each path in the path map serves as the response trajectory input, and a weighted average of the interference intensity of all pixels on the trajectory is calculated. The weighting function is constructed based on the path step size and the orientation consistency factor, where the orientation consistency factor is the cosine of the angle difference between preceding and following pixels. The aggregation function is a weighted average tensor intensity. Response values ​​from all paths are accumulated to the path endpoint. Each path endpoint records the aggregated tensor response value and orientation vector. All path endpoints are aggregated to form a coupled response prediction data map. The image structure is a two-channel tensor map: the first channel represents the response intensity value, and the second channel represents the response orientation angle value. All response values ​​are normalized to [0, 1], and the response orientation angle precision is maintained to one decimal place. The image size remains consistent with the original three-distribution framework.

[0081] Preferably, step S4, which involves evolutionary projection of the cervical tissue image based on the evolutionary prediction data of each framework layer, includes:

[0082] Spatial location encoding is performed on the evolution prediction data of each framework layer to obtain location-encoded prediction data;

[0083] The location-encoded prediction data is structured and projected onto cervical tissue images to generate structured projection images.

[0084] The evolutionary states in the fused structure projection image are obtained to obtain the state fusion image;

[0085] Fine-grained evolution rendering is performed on the state fusion image to obtain a fine-grained evolution mapping image;

[0086] Enhance the false color in the fine-grained evolutionary mapping image to obtain an evolutionary mapping image of cervical tissue.

[0087] In this embodiment, the input data consists of evolution prediction images corresponding to the three channels of the nuclear region, mesenchymal region, and epithelial region, respectively. Each structural region is represented by a single-channel mask image with a size of 512 pixels × 512 pixels. Each pixel in the prediction data contains the structural category, structural contour index, and response direction vector. When performing spatial location encoding processing on each structural block, a coordinate indexing method based on structural centroid localization is adopted. First, connected component analysis is performed on the structural mask to extract the boundary coordinates and center positions of all structural blocks. Then, a two-dimensional position embedding matrix is ​​constructed for each structural block in the original image coordinate system. This matrix has a dimension of 512 × 5. A 12×2 matrix is ​​used to record the relative distance coordinates of each pixel on the x and y axes. Simultaneously, the trend direction vector and response intensity of the structural blocks are extracted as dynamic structural features. These dynamic feature values ​​are encoded into a single-channel layer and fused with the coordinate embedding map to form a three-channel position-coded prediction image. In the position encoding, each pixel contains its current position coordinates, the ID of its assigned structural block, and the dominant trend direction value and tension response index of the corresponding structure. The original cervical tissue image is selected as the background projection map, which must maintain the same resolution and image dimensions as the prediction image. An affine position transformation is performed on the coordinates of the center point of each structural block in the structural position-coded image. The matrix determines the translation amount and direction based on the dominant trend direction vector and the response intensity. The translation distance *d* of the structure center point is the response intensity multiplied by a standard offset coefficient, which is set to 16 pixels. After performing pixel-level coordinate mapping transformation on each structural block, the structural outline is reconstructed at the target location while retaining the original texture grayscale values. All structural blocks are sequentially written onto the base image and edge blending is performed. The blending method involves creating a 4-pixel transition band at the structural edges and performing linear blending to reduce boundary breakage. Finally, a structural projection image is generated, in which each structural block is repositioned, and the dominant direction field and the original... The structure is identified by calling the structure projection image and the corresponding response orientation map and trend tensor map. A structure state channel is constructed in the image channel to record the evolution state identifier of each structure block. The evolution state is determined according to the trend tensor value: a value greater than 0.3 is defined as an expansion state, a value less than -0.3 is defined as a contraction state, and a value between -0.3 and -0.3 is defined as a stable state. A state mapping channel is constructed for each structure block in its location region. The expansion region is marked with a red channel, the contraction region with a blue channel, and the stable region with a green channel. The state identifier intensity is normalized and mapped to [0.2, 1] according to the absolute value of the trend tensor.[0] The fused image is constructed as an RGB three-channel image, which records the structural contour texture, structural state color, and state intensity, respectively. For overlapping areas of multiple structural blocks, a maximum state intensity priority fusion strategy is adopted. State edge smoothing is performed on the entire image with a smoothing kernel size of 3×3. The output state fused image contains structural information, trend information, and spatial evolution layers. Structural region sub-tiles are established, with each tile size set to 64 pixels × 64 pixels. A multi-layer texture reconstruction network is used to render and model the structural region within each tile. The network input is the structural state label, response direction vector, and trend intensity value. The network contains three convolutional layers and two skip connection modules. The output is a refined texture layer, which contains the degree of blurring of structural edges, density changes within the tissue, and directional gradient texture. The rendering result is processed by tile stitching and recombination to establish a complete fine-grained structural image. Three refined texture layers are added to the original structural image, including a morphological expansion map, a density stripe map, and a directional gradient map. Finally, a fine-grained evolutionary mapping image is synthesized. In the image, each structure not only retains its spatial location and structural type but also its basic structure. Based on the tissue evolution texture information generated by evolutionary trend simulation, three-channel layers record the texture grayscale image, trend direction map, and structural state map, respectively. A standard color mapping table is used for color conversion. The input consists of the state map and texture map channels from the fine-grained evolution image. The Jet pseudo-color mapping function is used to map grayscale texture values ​​to color images, where 0 values ​​are mapped to dark blue, 1 values ​​to red, and intermediate values ​​transition to green and yellow. The mapping table retains 256 levels of color mapping precision. Edge sharpening is performed on the mapped image, and a 5×5 high-pass filter kernel is used to enhance the clarity of structural boundaries. A color gradient overlay operation is then performed on the trend direction layer, and a transparency layer is constructed using structural direction angle values ​​and state intensity values. After overlay, a structural state response map is formed. The three layers are then fused to construct the final evolutionary mapped cervical tissue image. The image size is 512 pixels × 512 pixels, and the image format is an RGB three-channel floating-point image. Each channel stores color values ​​in float32 format with a precision of 0.001. The output image serves as the final input data structure for classification modeling and visual alignment. The image is labeled as a tissue evolution structure mapping map.

[0088] Preferably, step S4, which classifies cervical tissue images according to evolutionary process categories based on evolutionary mapping, includes:

[0089] Extracting process features from evolutionary mapping cervical tissue images;

[0090] Clustering is performed based on process features, and image category feature labels are constructed.

[0091] Based on image category feature labels, cervical tissue images of evolutionary mapping are classified into categories, thereby obtaining multiple evolutionary process image categories;

[0092] The image category of the cervical tissue image corresponding to the evolutionary mapping cervical tissue image is determined by the evolutionary process image category;

[0093] Cervical tissue images are classified according to their category to generate cervical classification images.

[0094] In this embodiment, the input image is an evolution-mapped cervical tissue image after evolutionary projection, state fusion, and fine-grained evolutionary rendering. The image size is 512 pixels × 512 pixels. The channel structure includes a structural state channel, a structural orientation channel, and an evolutionary texture channel. For the state channel, a region slicing operation is first performed to obtain the average state intensity value of all structural blocks in the image, and the standard deviation and extreme value interval are calculated. Then, a gradient orientation histogram extraction operation is performed on the orientation channel, dividing the angle range into 18 equally divided intervals. The number of pixels in each interval is accumulated to obtain the orientation distribution vector. In the texture channel, the image block gray-level co-occurrence matrix (G) is extracted. The LCM (Latent Computational Model) parameters were extracted in four directions: 0°, 45°, 90°, and 135°. For each direction, 16 texture feature values ​​(contrast, entropy, energy, and correlation) were extracted. Finally, the state statistics, orientation distribution vector, and texture parameters were concatenated into a 64-dimensional process feature vector, with one vector corresponding to each image. This 64-dimensional process feature vector matrix was then input into a K-means clustering model for image classification. Five clusters were set, each representing a different evolutionary trend. The model employed a k-means++ initialization strategy to improve the stability of the initial centers, and the Euclidean distance metric was used. The distance function is obtained, and the maximum number of clustering iterations is set to 300 rounds. The convergence criterion is to stop iteration when all center points move less than 1e-4. Silhouette coefficients are used to verify the number of categories during clustering. In the verification phase, it was found that the average silhouette coefficient reached 0.59 when k=5, meeting the clustering stability standard. After clustering, a clustering category label is assigned to each sample, named from "EVT-0" to "EVT-4". Category statistical parameters are constructed for each label, and the output label structure includes image ID, label category, and feature vector index. The image category feature label structure is read, and the category identifier corresponding to each image is parsed. Images are grouped by category field and written to different path directories. The directory structure is set as five folders from "EVT-0" to "EVT-4", with each folder corresponding to a process category. At the same time, an independent JSON tag file is generated for each image. The tag content includes image path, state mean, dominant orientation distribution, and feature vector summary. A representative image index list file is also built for each category for subsequent model training sample extraction. The total number of samples for each category is no less than 500. If the number of image samples for a certain category is insufficient, the existing images are subjected to affine rotation ±15 degrees, brightness perturbation ±10%, and Gaussian blur σ=0.Eight methods were used for data augmentation to increase the number of images until class balance was achieved. Finally, all evolution-mapped images were labeled, grouped, and written into an image classification structure. An image ID mapping table, generated during the evolution mapping stage, was used for index matching. This table contains fields for the original image ID, the evolved image ID, and the class label. The evolved image ID and label were read from the image class label, and the corresponding original image ID field was retrieved using a hash index structure. Subsequently, a mapping table between the original images and their corresponding evolution category labels was constructed. For each original image, its final evolution classification label was recorded, along with the mean state intensity, principal axis angle, and structural complexity level. The structural complexity level was calculated by combining the number of structural blocks and the standard deviation of the structural spacing, with five levels from Level-0 to Level-4. Ultimately, all original images obtained a one-to-one corresponding classification label. The input was the original cervical tissue image and its corresponding class label mapping. The system performs a label renaming operation on each image, adding a category field prefix to the original image filename, following the naming rule "EVT-x_original image name". Simultaneously, a structural information layer is generated for each image. This layer extracts the structural boundary mask from the corresponding evolutionary image and overlays it onto the original image. The boundary colors are RGB-encoded according to the classification label: EVT-0 is blue, EVT-1 is green, EVT-2 is yellow, EVT-3 is orange, and EVT-4 is red. The structural layer is overlaid with an alpha transparency of 0.4. The output is a three-channel floating-point PNG image. All classified images are written to different directories according to their labels and bound to a JSON tag file. The tag file records the category ID, image ID, main structural region contour coordinates, and feature parameter summary. This ultimately forms a cervical classification image set containing over 5000 images. This image set constitutes the data source for training the visual basic model and the baseline image structure set for image-text fusion input.

[0095] Preferably, step S5, which involves hierarchical joint modeling training based on cervical classification images, includes:

[0096] The cervical classification image is segmented into standard magnification blocks to obtain cervical classification pre-cut blocks;

[0097] The cervical classification pre-cut map is automatically screened for tissue structure regions to obtain tissue region maps;

[0098] The organization area map tiles are processed by coordinate index encoding to obtain the basic data of the map tiles;

[0099] Coverage statistics are performed using basic tile data, and high-quality tiles are selected based on the coverage.

[0100] A large-scale training dataset is constructed by performing unified processing of multi-center structures based on high-quality map tiles and integrating their labels.

[0101] Self-supervised visual pre-training based on the DINOv2 architecture is performed on a large-scale training dataset to generate initial visual feature data.

[0102] A basic visual model capable of representing organizational structure is constructed based on initial visual feature data.

[0103] In this embodiment, the input data is a cervical classification image obtained during the evolutionary stage. This image undergoes pseudo-color processing, evolutionary mapping reconstruction, and process label overlay processing. The image resolution is uniformly set to 0.24 micrometers per pixel, and the image size is standardized to 4096 pixels × 4096 pixels. Before performing patch segmentation, the image needs to be converted to a floating-point format three-channel matrix representation, with each channel value normalized to [0,1]. Subsequently, a fixed-window cropping method is used to perform equidistant sliding window segmentation on the image. The patch size is set to 512 pixels × 512 pixels, and the step size is set to 384 pixels, ensuring a 128-pixel overlap area between patches for boundary compensation. When cropping incomplete areas at the image edges, incomplete patches are filled using a mirror filling mode of "reflection". The "CT" reflection edge processing generates approximately 63 images after image segmentation. Each image is named with the original image name plus its location index and stored in a structured directory. Each image is also labeled with its original image location coordinates and corresponding image category label. The generated data structure is named "Cervical Classification Pre-cropped Image Data." The input data is a set of pre-cropped image images obtained after standard magnification segmentation. For each image, a tissue structure region segmentation model is performed for prediction. The segmentation model is a pre-trained tissue structure region detection model based on the UNet++ architecture. The model's input image size is 512 pixels × 512 pixels, and the output is a single-channel region mask image. In the mask image, a pixel value of 0 represents the background, and a value of 1 represents the effective tissue region. The model uses batch processing for inference on each image. A batch processing method with a size of 8 is used. Morphological closing operations and maximum connected region extraction are performed on each mask image to eliminate edge pseudo-response regions. Patches with a tissue proportion of less than 60% in the predicted mask are directly removed. For the retained patches, their structural region index maps are extracted. Continuous regions in the mask are labeled as structural blocks, and the centroid coordinates, contour boundaries, and area values ​​of each block are recorded. The structural block data format is a triplet of coordinates + structural mask + label. This data is stored in the patch index structure to form a tissue region patch dataset. The position coordinates of each tissue region patch in the original image are read from the patch, and the results are analyzed according to the pre-cropped patch naming rules and the centroid of the structural block. The coordinates are summed to obtain the absolute position coordinates of the structural block on the entire image. These coordinates are then jointly encoded with the corresponding tissue channel category and image category label to construct a position index structure. This structure is defined using a quintuple format, consisting of the tile ID, structure ID, X coordinate, Y coordinate, and category ID. A two-dimensional coordinate mapping matrix is ​​also constructed for the reverse mapping from the tile to the original image. Each tile encoding structure is stored in a separate coordinate index file in HDF5 format. Each HDF5 file corresponds to one tile, internally storing the structural block coordinate data and tile texture vector information in dataset form. Each tile file is approximately 1 unit in size.The 2MB coordinate index encoding output is the basic tile data, used for coverage calculation, tile selection, and structural alignment modeling input layer calls. It reads the area information of the structural block corresponding to each tile in the coordinate index. The structural area is obtained by counting pixels in the mask image, and then the ratio with the total area of ​​the tile is calculated to obtain the organization coverage index. The coverage threshold is set to 0.65. Tiles with a value lower than this value are defined as low-quality tiles and are removed. The structural complexity index is further extracted from the retained tiles. The structural complexity is constructed using the number of different structural blocks in the tile and the standard deviation of the structural density as the core indicators. The output range of the scoring function is [0,1], and the complexity score is lower than 0.Tiles with a value of 35 or less will also be excluded. The remaining tiles are defined as high-quality tiles and stored in named directories according to their structural channel categories. Nuclear, mesenchymal, and epithelial tiles are stored in different folders. A total of 128,000 tiles from six data centers are imported into the data processing platform. The platform uses a center attribution ID and tile ID to establish an index lookup table and performs cross-center tile style normalization. A color normalization module based on the Macenko algorithm is used to color map all tiles. The target template is derived from the center sample tiles with the clearest structure in the distribution. Label files are generated for all tiles according to image category and structural channel category, and label alignment is performed. The label content includes image category label, structural channel label, center attribution label, and structural contour mask. This label structure is stored in YAML format. Each tile image is stored in a unified named directory with a one-to-one correspondence with a label file. Finally, a large-scale training dataset is constructed and used with Vision... The DINOv2 model architecture is built upon the Transformer. The input image size is 224 pixels × 224 pixels. Each image patch undergoes data augmentation before training, employing a combination of augmentation strategies such as random flipping, random cropping, color perturbation, and Gaussian blur to construct pre-trained image pairs. Each image pair is then fed into the backbone encoder and auxiliary encoder for self-supervised feature matching. The training process uses a joint loss function composed of feature reconstruction loss and contrastive loss for optimization. Training batches are set to 256 images per batch, with 400 training epochs, each containing approximately 1000 batches. The optimizer used is AdamW, with an initial learning rate of 3e-4 and Cosine optimization. Annealing is used to decay the learning rate. The model output is a feature vector corresponding to each image, with a vector dimension of 768. These feature vectors are extracted and saved as initial visual feature data after training. The DINOv2 backbone structure is used as the encoder to construct the main model architecture. A structure region prediction head and a multi-task feature branch module are added to the output. The structure region prediction head uses a combination of bilinear upsampling and spatial attention mechanisms. The prediction head output is aligned with the structure channel mask image for supervision. The loss function is a combination of cross-entropy loss and Dice loss. The model input is an enhanced patch image, and the output is a structure region prediction map and a structure category feature vector. The training samples are selected high-quality patch data. The training method is supervised fine-tuning, with 120 training epochs, 1000 batches per epoch, and 64 patches per batch. The training optimizer is LAMB, with an initial learning rate of 2e-4. The final model is saved as the CerS-V visual base model.

[0104] Preferably, step S5, which involves performing image-text fusion modeling based on a visual foundation model, conducting task migration processing, and deploying it to the integrated system, includes:

[0105] The visual feature outputs in the visual base model are matched one by one with the image and text annotation samples to obtain visual language pairing data.

[0106] Based on visual language pairing data, image-text modality mapping and binding are performed to generate multimodal chimeric input data;

[0107] Based on multimodal chimeric input data, image and text joint training processing is performed to obtain a multimodal basic model with alignment capabilities;

[0108] The image and text base model data and the visual base model in the multimodal base model are configured and module-encapsulated for 25 core task scenarios to generate multi-task collaborative processing module data.

[0109] Data from the multi-task collaborative processing module is injected into the decision-making intelligent agent planner constructed from the language model, thereby obtaining task planning and scheduling control data;

[0110] Modality conversion mapping is performed based on task planning and scheduling control data to obtain unified inference path data;

[0111] Structured interfaces are encapsulated based on unified inference path data, and user interaction modules are connected to generate deployable and integrated system data.

[0112] The deployable integrated system data was prospectively validated and evaluated in six independent centers, ultimately resulting in a cross-center deployment model system.

[0113] In this embodiment, the input image is a large-scale, high-quality patch data. The label file contains text annotations such as real tissue descriptions, pathological descriptions, structural annotations, cervical typing labels, and image-text question-and-answer pairs. A 768-dimensional visual feature vector is extracted from each patch using the CerS-V visual base model. Simultaneously, the text portion is extracted using the embedding layer of the Qwen2.5-VL model, with a uniform 4096-dimensional language embedding dimension. To ensure accurate one-to-one matching, a unique pairing index is established between the image and text, employing a bidirectional mapping structure of image ID and text ID. If an image corresponds to multiple text samples, multiple visual-language pairing data items are generated, and a unified chimeric key-value pair identifier is generated for each pair of samples. Finally, a... A training dataset containing 2.5 million pairs of visual-language pairings is provided. Each pair includes image path, visual feature tensor, text content, text embedding vectors, and a multimodal index structure. Using the visual-language pairing data as input, a modality fusion preprocessing module is constructed. The visual input is a standardized image tensor with a size of 3×224×224, and the language input is a segmented text tag sequence with a maximum length of 128 tokens. The visual input is processed by the CerS-V model to extract a 768-dimensional patch feature embedding sequence, and the language input is processed by the Qwen2.5-VL word embedding layer to generate a 4096-dimensional token feature vector sequence. The two modality vectors are then linearly mapped using a LoRA low-rank adapter. A 1- to 1024-dimensional chimera space is used, and a learnable cross-modal attention bridge layer connects two embedding channels. 128 image-text bindings are performed in each batch. After fusion, each pair of data forms a unified multimodal chimera input structure, with an output tensor of (batch, token_length, 1024), where token_length is the sum of image and text tokens. The resulting chimera input data is used for subsequent multimodal model training. Using this multimodal chimera input data as model input, a cross-modal language modeling network with Qwen2.5-VL as the language backbone is trained. A CerS-V visual encoding layer is integrated into the training structure as an image prefix embedding channel, and three LoRA low-rank connections are used. The module injects visual patch features into the first three attention layers of the language Transformer backbone. Training tasks include image-text matching, image-text generation, and cross-modal question answering. The loss function consists of three parts: matching loss (InfoNCE), generation loss (CE), and question-answering accuracy loss (F1-score). Training batches consist of 256 pairs of samples, with a total of 20 training rounds, each containing 3000 iterations. The AdamW optimizer and an initial learning rate of 1e-4 are used. The resulting multimodal base model, named CerS-M, is capable of generating text, describing structures, interpreting pathological information, and performing question-answering reasoning upon input images.Based on the actual needs of cervical pathology examination tasks, a list of 25 core tasks was constructed. Each task category includes input modality type, target task label type, judgment output fields, and structured report interface. A configuration file was defined for each task category, with configuration fields including input channel mapping structure, task objective function, label index mapping, and module IO interface protocol. Image task modules such as structure recognition, squamous cell carcinoma grading, and rare cancer recognition were constructed based on the structure recognition capability of the CerS-V model. Image-text question answering, structure description generation, and conversational report generation modules were constructed based on the multimodal question answering capability of the CerS-M model. All modules were independently encapsulated as PyTorch checkpoints and JSON configuration pairs, ultimately generating 25 decision module data. Each module corresponds to a unique module ID and scheduling label. A decision-making intelligent agent framework based on the Qwen2.5 large language model was used. This framework loads the task configuration file and module IO interface protocol, then performs task parsing, module matching, and path planning. The input is the user... For dialogue commands, task invocation requests, or combined text and image inputs, the planner first performs embedding encoding of the input content, then performs top-k similarity matching with the task tag word vectors of registered modules. Upon successful matching, a task scheduling tree is generated. This scheduling tree defines the inference path nodes, module invocation order, input-output data flow mapping relationship, and generates a standardized task execution plan file in JSON format. The file includes: task ID, sequence of invoked module IDs, input modal path, output target format, and task-level parameter templates. After execution, the task planning and scheduling control data is provided to the integrated system's inference module for automatic invocation of decision-making modules and data flow. The module extracts the required modal inputs, invoked module IDs, and data flow paths for each task node from the task scheduling control data. Format conversion and distributed preloading processing are performed on the input modal paths. The image modality conversion module, based on a preprocessing interface jointly built with OpenCV and PIL, completes image resizing, normalization, and color channel matching. The language modality performs word segmentation precoding and tokenization. ID mapping is used, and all transformation results are packaged into a unified tensor structure and bound to module input slots. The system constructs a modal mapping graph, where each node is a module input interface and each edge is a data transformation channel. Scheduling order fields and execution timestamps are added to the graph structure. The execution engine triggers inference operations according to the graph order, ultimately generating unified inference path data. The data structure is a nested embedded structure, including input data tensors, module flow paths, and output structure distribution plans. The unified inference path data defines the system service interface specifications, which are encapsulated in a RESTful API structure. Inputs are image URLs or JSON commands. The system uses Flask to build task scheduling APIs and data input preprocessing APIs, and also establishes a WebSocket channel to support long-session, multi-turn task question-and-answer calls.The structured output interface maps model output to a JSON structure, with fields including module output text, structural region coordinates, judgment level, and key structural descriptions. The user interaction module features nested interface logic configuration, supporting image display, structural drawing, and decision-making dialogue. The interface state is synchronized to the backend state machine. Six tertiary centers equipped with cervical tissue whole-slice scanning and digital pathology systems were selected as validation sites. The same version of the integrated system container was deployed at each center, using a LAN-based remote API access and local caching mechanism. Each center selected 500 independent case image samples, including cervical cancer screening images, SILVA typing images, and structured question-and-answer image task input samples. System evaluation metrics included task call success rate, structural recognition accuracy, structural description generation consistency, and question-and-answer accuracy. The evaluation period was three weeks, with path engineers recording system call logs weekly and uploading them to a unified server for performance aggregation and analysis. All log data was used for fine-tuning and compensation optimization in the training system. Finally, a six-center deployment validation report was generated, and the cross-center deployed model system version was output as the model integration result, recording model accuracy, stability, and portability metrics.

[0114] The present invention also provides a training system for a multimodal cervical pathology image classification model, used to execute the training method for the multimodal cervical pathology image classification model described above. The training system for the multimodal cervical pathology image classification model includes:

[0115] The three-part decomposition module is used to acquire cervical tissue images; the cervical tissue images are segmented into tissue structures to obtain segmented cervical structural blocks; based on the segmented cervical structural blocks, the cervical tissue is divided into a three-part distribution framework of nucleus-stromal-epithelial.

[0116] The trend recognition module is used to collect historical data on cervical tissue development; based on the historical data on cervical tissue development, it identifies the developmental lifeline of the three-distribution framework; and through the developmental lifeline of the three-distribution framework, it confirms the predicted development trend of each layer in the three-distribution framework.

[0117] The environmental prediction module is used to simulate the cervical tissue environment field through cervical tissue images, thereby generating a simulated cervical environment field; based on the simulated cervical environment field and the predicted development trend of each layer, the three-layer distribution framework of nucleus-mesenchyma-epithelium is predicted to evolve in layers to obtain the evolution prediction data of each layer framework.

[0118] The evolution classification module is used to perform evolutionary projection on cervical tissue images based on evolutionary prediction data of each framework layer, thereby generating evolutionary mapping cervical tissue images; based on the evolutionary mapping cervical tissue images, the cervical tissue images are classified according to the evolutionary process category, thereby generating cervical classification images;

[0119] The fusion modeling module is used to perform hierarchical joint modeling training based on cervical classification images to obtain a visual base model; based on the visual base model, image-text fusion modeling is performed, and task transfer processing is carried out, while deploying to the integration system to generate a cross-center deployment model system.

[0120] This invention constructs a three-channel structural framework of nucleus-mesenchyma-epithelium through a three-part decomposition module, which is beneficial to improve the spatial decoupling capability of tissue partitioning and the structural clarity of subsequent modeling. The trend recognition module extracts the structural evolution trajectory through historical data and establishes structural correspondences between multiple frames of images, enhancing temporal consistency and trend traceability. The environment prediction module generates a simulated environment field containing tension perturbations, pressure mapping, and trend arrays, realizing non-biological mechanism-driven modeling of the structural evolution process and improving the directionality and structural fidelity of evolution simulation. The evolution classification module generates mapped images based on predictions and classifies process categories, enabling image classification to have trend perception capabilities and distinguishable change stages. The fusion modeling module constructs a visual model through classified images and completes image-text alignment training, realizing unified modeling of tissue structure, trend semantics, and image-text modality. The system has a clear structure, separate modules, supports cross-modal input, task-level scheduling, and multi-center deployment, and has scalability, composability, and structural universality.

Claims

1. A training method for a multimodal cervical pathology image classification model, characterized in that, Includes the following steps: Step S1: Obtain cervical tissue images; perform tissue structure segmentation on the cervical tissue images to obtain segmented cervical structural blocks; based on the segmented cervical structural blocks, divide the cervical tissue into a three-distribution framework of nucleus-stromal-epithelial. Step S2: Collect historical data on cervical tissue development; identify the three-distribution framework development lifeline based on the historical data on cervical tissue development; confirm the predicted development trend of each layer in the three-distribution framework through the three-distribution framework development lifeline; Step S3: Simulate the cervical tissue environment field using cervical tissue images to generate a simulated cervical environment field; based on the simulated cervical environment field and the predicted development trends of each layer, perform hierarchical evolution prediction of the three-layer distribution framework of nucleus-stromal-epithelium to obtain the evolution prediction data of each layer framework. Step S4: Based on the evolution prediction data of each framework layer, perform evolutionary projection on the cervical tissue image to generate an evolutionary mapping cervical tissue image; based on the evolutionary mapping cervical tissue image, classify the cervical tissue image according to the evolutionary process category to generate a cervical classification image; Step S5: Perform hierarchical joint modeling training based on cervical classification images to obtain a visual base model; perform image-text fusion modeling based on the visual base model, and perform task transfer processing, while deploying to the integration system to generate a cross-center deployment model system; The hierarchical joint modeling training based on cervical classification images includes: The cervical classification image is segmented into standard magnification blocks to obtain cervical classification pre-cut blocks; The cervical classification pre-cut map is automatically screened for tissue structure regions to obtain tissue region maps; The organization area map tiles are processed by coordinate index encoding to obtain the basic data of the map tiles; Coverage statistics are performed using basic tile data, and high-quality tiles are selected based on the coverage. A large-scale training dataset is constructed by performing unified processing of multi-center structures based on high-quality map tiles and integrating their labels. Self-supervised visual pre-training based on the DINOv2 architecture is performed on a large-scale training dataset to generate initial visual feature data. A basic visual model capable of representing organizational structure is constructed based on initial visual feature data.

2. The training method for the cervical pathology image classification model based on multimodal imaging according to claim 1, characterized in that, Step S1 includes the following steps: Step S11: Obtain cervical tissue image; normalize the color of cervical tissue image to obtain normalized cervical image; perform depth contrast enhancement processing on normalized cervical image to generate enhanced cervical image. Step S12: Identify cervical tissue structure features based on enhanced cervical images, and perform spatial collaborative segmentation on the enhanced cervical images according to the cervical tissue structure features to obtain segmented cervical structure blocks; Step S13: Perform kernel density mapping on the segmented cervical structural blocks to obtain kernel density distribution data; perform stromal region separation based on the kernel density distribution data to generate stromal distribution data; Step S14: Decouple the epithelial structure using mesenchymal distribution data to obtain epithelial distribution data; fuse the epithelial distribution data, nuclear density distribution data, and mesenchymal distribution data, and construct a three-distribution framework of nuclear-mesenchymal-epithelial based on the fused data.

3. The training method for the cervical pathology image classification model based on multimodal imaging according to claim 1, characterized in that, Step S2 includes the following steps: Step S21: Collect historical data on cervical tissue development; perform time-series label fusion on the three-distribution framework based on the historical data on cervical tissue development to generate time-series three-distribution framework data; Step S22: Perform evolution trajectory tracking based on the time-series three-distribution framework data to obtain the three-distribution evolution trajectory. Each evolution trajectory contains a minimum of 4 nodes, and the drift distance between structural blocks between nodes is controlled within 64 pixels. If this threshold is exceeded, the trajectory tracking chain will be automatically interrupted. Step S23: Construct a three-distribution framework development lifeline through the three-distribution evolution trajectory; extract the hierarchical development trend in the three-distribution framework development lifeline, wherein the trend identification adopts a sliding trend window width set to 3 frames, and the slope of the trend inflection point change must be greater than 0.4 to be defined as a valid mutation; Step S24: Perform trend distribution mapping based on hierarchical development trends to confirm the predicted development trends of each layer in the three-distribution framework.

4. The training method for the cervical pathology image classification model based on multimodal imaging according to claim 1, characterized in that, Step S3, which involves simulating the cervical tissue environment using cervical tissue images, includes: Extract cervical environment data from cervical tissue images; Asymmetric tissue tension field fitting was performed on cervical environmental data to obtain the environmental tension characteristic field; A horizontal-radial structural flow tensor is injected into each tension sub-block of the environmental tension characteristic field, and residual conduction is performed on the maximum gradient direction of the environmental tension characteristic field to generate environmental orientation sensing guidance data. Based on environmental orientation sensing guidance data and cervical environment data, cervical tissue environment field simulation is performed to generate a simulated cervical environment field. Specifically, injecting a horizontal-radial structural flow tensor into each tension sub-block of the environmental tension characteristic field and performing residual conduction on the maximum gradient direction of the environmental tension characteristic field includes: The environmental tension feature field is divided into multiple non-overlapping tension sub-blocks, and a tension sub-block index is constructed based on these multiple non-overlapping tension sub-blocks. Based on the tension sub-block index, a horizontal-radial structure flow tensor is injected into each non-overlapping tension sub-block to obtain the structure tensor superposition data. Tensor continuity balancing is performed between blocks based on the superimposed structural tensor data to obtain tensor fusion data; Extract the direction of maximum gradient from the tensor fusion data to generate the dominant structural flow direction; Extract high-gradient abrupt change regions in the dominant structural flow direction; By completing the coherence of high-gradient abrupt change regions, environmental orientation sensing guidance data is generated.

5. The training method for the cervical pathology image classification model based on multimodal imaging according to claim 1, characterized in that, Step S3 involves predicting the hierarchical evolution of the nucleus-stromal-epithelial three-layer distribution framework based on the simulated cervical environment field and the predicted development trends of each layer. The predicted development trends of each layer are projected onto the three-distribution framework of the nucleus-mesenchyma-epithelium into the simulated cervical environment field to obtain projected three-distribution data; Hierarchical dynamic migration simulation is performed on the projected three-distribution data to generate hierarchical dynamic migration data; Inter-layer coupling response inference is performed based on hierarchical dynamic migration data, thereby generating coupling response prediction data; Evolution-driven processing of the projected three-distribution data is performed based on the coupling response prediction data to obtain the evolution prediction data of each frame layer.

6. The training method for the cervical pathology image classification model based on multimodality according to claim 1, characterized in that, Step S4 involves performing evolutionary projection on the cervical tissue image based on the evolutionary prediction data of each framework layer, including: Spatial location encoding is performed on the evolution prediction data of each framework layer to obtain location-encoded prediction data; The location-encoded prediction data is structured and projected onto cervical tissue images to generate structured projection images. The evolutionary states in the fused structure projection image are obtained to obtain the state fusion image; Fine-grained evolution rendering is performed on the state fusion image to obtain a fine-grained evolution mapping image; Enhance the false color in the fine-grained evolutionary mapping image to obtain an evolutionary mapping image of cervical tissue.

7. The training method for the cervical pathology image classification model based on multimodal imaging according to claim 1, characterized in that, Step S4, which classifies cervical tissue images according to their evolutionary process based on evolutionary mapping, includes: Extracting process features from evolutionary mapping cervical tissue images; Clustering is performed based on process features, and image category feature labels are constructed. Based on image category feature labels, cervical tissue images of evolutionary mapping are classified into categories, thereby obtaining multiple evolutionary process image categories; The image category of the cervical tissue image corresponding to the evolutionary mapping cervical tissue image is determined by the evolutionary process image category; Cervical tissue images are classified according to their category to generate cervical classification images.

8. The training method for the cervical pathology image classification model based on multimodal imaging according to claim 1, characterized in that, Step S5 involves performing image-text fusion modeling based on the visual foundation model, carrying out task migration processing, and deploying it to the integration system, including: The visual feature outputs in the visual base model are matched one by one with the image and text annotation samples to obtain visual language pairing data. Based on visual language pairing data, image-text modality mapping and binding are performed to generate multimodal chimeric input data; Based on multimodal chimeric input data, image and text joint training processing is performed to obtain a multimodal basic model with alignment capabilities; The image and text base model data and the visual base model in the multimodal base model are configured and module-encapsulated for 25 core task scenarios to generate multi-task collaborative processing module data. Data from the multi-task collaborative processing module is injected into the decision-making intelligent agent planner constructed from the language model, thereby obtaining task planning and scheduling control data; Modality conversion mapping is performed based on task planning and scheduling control data to obtain unified inference path data; Structured interfaces are encapsulated based on unified inference path data, and user interaction modules are connected to generate deployable and integrated system data. The deployable integrated system data was prospectively validated and evaluated in six independent centers, ultimately resulting in a cross-center deployment model system.

9. A training system for a multimodal cervical pathology image classification model, characterized in that, The training system for the multimodal cervical pathology image classification model as described in claim 1, used to execute the training method, comprises: The three-part decomposition module is used to acquire cervical tissue images; the cervical tissue images are segmented into tissue structures to obtain segmented cervical structural blocks; based on the segmented cervical structural blocks, the cervical tissue is divided into a three-part distribution framework of nucleus-stromal-epithelial. The trend recognition module is used to collect historical data on cervical tissue development; based on the historical data on cervical tissue development, it identifies the developmental lifeline of the three-distribution framework; and through the developmental lifeline of the three-distribution framework, it confirms the predicted development trend of each layer in the three-distribution framework. The environmental prediction module is used to simulate the cervical tissue environment field through cervical tissue images, thereby generating a simulated cervical environment field; based on the simulated cervical environment field and the predicted development trend of each layer, the three-layer distribution framework of nucleus-mesenchyma-epithelium is predicted to evolve in layers to obtain the evolution prediction data of each layer framework. The evolution classification module is used to perform evolutionary projection on cervical tissue images based on evolutionary prediction data of each framework layer, thereby generating evolutionary mapping cervical tissue images; based on the evolutionary mapping cervical tissue images, the cervical tissue images are classified according to the evolutionary process category, thereby generating cervical classification images; The fusion modeling module is used to perform hierarchical joint modeling training based on cervical classification images to obtain a visual base model; based on the visual base model, image-text fusion modeling is performed, and task transfer processing is carried out, while deploying to the integration system to generate a cross-center deployment model system.

Citation Information

Patent Citations

  • Multi-type cell nucleus labeling and multi-task processing method for cervical TCT section

    CN117496512A

  • Cervical panoramic image few-sample classification method based on visual guidance and language prompt

    CN118230052A