Manifold alignment based model training method and device
By using a manifold alignment-based model training method, a visual language model is iteratively trained on gastric pathology images and text features, which solves the problem that existing models cannot effectively integrate cross-modal information and improves the accuracy of gastric pathology detection.
Patent Information
- Application Number
- CN202511254799.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-04
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-09-04
AI Technical Summary
Existing gastric pathology detection models are unable to effectively capture the complex nonlinear relationships between images and text in their respective latent manifold spaces during training, resulting in limited cross-modal information integration and poor detection accuracy.
A manifold alignment-based model training method is adopted, which iteratively trains gastric pathological image and text features through a visual language model, including image feature extraction, feature enhancement, cross-modal feature optimization, manifold space mapping and feature fusion, and integrates image and text features using an active spatial focusing mechanism.
The detection accuracy of the gastric pathology detection model was improved, and the training effect of the model was improved by effectively integrating cross-modal information.
Smart Images

Figure CN120805067B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of machine learning technology, and in particular to a model training method and apparatus based on manifold alignment. Background Technology
[0002] With the development of machine learning technology, it is gradually being applied in the medical field for assisted diagnosis. One common application is training a gastric pathology detection model to detect gastric pathology images and obtain diagnostic results about pathological types, lesion locations, and other information.
[0003] In the context of gastric pathology examination, multimodal data processing is often involved, such as image data and text data. In the training process of existing gastric pathology examination models, multimodal data is typically directly concatenated into features, and model training is completed based on these concatenated features.
[0004] The method of directly stitching features from multimodal data usually cannot effectively capture the complex nonlinear relationship between images and text in their respective potential manifold spaces. Therefore, the gastric pathology detection model trained based on the existing method has limited integration effect on cross-modal information, resulting in poor detection accuracy. Summary of the Invention
[0005] In view of this, embodiments of the present invention provide a model training method based on manifold alignment to solve the problem that existing model training methods have limited integration effects on cross-modal information, resulting in poor model detection accuracy.
[0006] This invention also provides a model training device based on manifold alignment to ensure the practical implementation and application of the above method.
[0007] To achieve the above objectives, the embodiments of the present invention provide the following technical solutions:
[0008] The first aspect of this invention provides a model training method based on manifold alignment, comprising:
[0009] Determine the training dataset; the training dataset includes multiple gastric pathology images and corresponding descriptive text and annotation labels for each gastric pathology image;
[0010] Based on the training dataset, the pre-constructed visual language model is iteratively trained according to a preset number of training iterations. When entering the current iterative training cycle, the visual language model is applied to extract image features from each of the gastric pathological images to obtain the image features corresponding to each of the gastric pathological images. Text features are also extracted from the descriptive text corresponding to each of the gastric pathological images to obtain the text features corresponding to each of the gastric pathological images.
[0011] For each of the gastric pathological images, the visual language model is applied to enhance the image features corresponding to the gastric pathological image to obtain the enhanced image features corresponding to the gastric pathological image, and the text features corresponding to the gastric pathological image are enhanced to obtain the enhanced text features corresponding to the gastric pathological image.
[0012] For each of the gastric pathology images, based on its corresponding enhanced text features, cross-modal feature optimization is performed on the enhanced image features corresponding to the gastric pathology image to obtain the cross-modal optimized image features corresponding to the gastric pathology image;
[0013] For each of the gastric pathological images, manifold space mapping is performed on the cross-modal optimized image features and enhanced text features corresponding to the gastric pathological image to obtain the mapped image features and mapped text features corresponding to the gastric pathological image;
[0014] For each of the gastric pathological images, based on a preset active spatial focusing mechanism, the mapping image features and mapping text features corresponding to the gastric pathological image are fused to obtain the image-text fusion features corresponding to the gastric pathological image.
[0015] Based on the image-text fusion features corresponding to each of the gastric pathology images, the visual language model is trained to obtain a gastric pathology detection model.
[0016] A second aspect of this invention provides a model training apparatus based on manifold alignment, comprising:
[0017] The first determining unit is used to determine the training dataset; the training dataset includes multiple gastric pathology images and corresponding descriptive text and annotation labels for the gastric pathology images;
[0018] The feature extraction unit is used to iteratively train the pre-built visual language model based on the training dataset according to a preset number of training iterations; when entering the current iterative training cycle, the visual language model is used to extract image features from each of the gastric pathological images to obtain the image features corresponding to each of the gastric pathological images, and to extract text features from the descriptive text corresponding to each of the gastric pathological images to obtain the text features corresponding to each of the gastric pathological images.
[0019] The feature enhancement unit is used to apply the visual language model to each of the gastric pathological images to enhance the image features corresponding to the gastric pathological image, thereby obtaining the enhanced image features corresponding to the gastric pathological image, and to enhance the text features corresponding to the gastric pathological image, thereby obtaining the enhanced text features corresponding to the gastric pathological image.
[0020] The feature optimization unit is used to perform cross-modal feature optimization on the enhanced image features corresponding to each gastric pathological image based on its corresponding enhanced text features, so as to obtain the cross-modal optimized image features corresponding to the gastric pathological image.
[0021] The feature mapping unit is used to perform manifold space mapping on the cross-modal optimized image features and enhanced text features corresponding to each gastric pathological image to obtain the mapped image features and mapped text features corresponding to the gastric pathological image.
[0022] The feature fusion unit is used to perform feature fusion on the mapped image features and mapped text features corresponding to each gastric pathological image based on a preset active spatial focusing mechanism, so as to obtain the image-text fusion features corresponding to the gastric pathological image.
[0023] The feature training unit is used to train the visual language model based on the image-text fusion features corresponding to each of the gastric pathology images to obtain a gastric pathology detection model.
[0024] A manifold alignment-based model training method provided by the above embodiments of the present invention includes: determining a training dataset; the training dataset includes multiple gastric pathological images and corresponding descriptive text and annotation labels for each gastric pathological image; iteratively training a pre-constructed visual language model based on the training dataset according to a preset number of training iterations; upon entering the current iterative training cycle, applying the visual language model to extract image features from each gastric pathological image to obtain image features corresponding to each gastric pathological image, and extracting text features from the descriptive text corresponding to each gastric pathological image to obtain text features corresponding to each gastric pathological image; for each gastric pathological image, applying the visual language model to enhance the image features corresponding to the gastric pathological image to obtain enhanced image features corresponding to the gastric pathological image, and enhancing the text features corresponding to the gastric pathological image. Feature enhancement is performed on the features to obtain the enhanced text features corresponding to the gastric pathology image. For each gastric pathology image, based on its corresponding enhanced text features, cross-modal feature optimization is performed on the enhanced image features corresponding to the gastric pathology image to obtain the cross-modal optimized image features corresponding to the gastric pathology image. For each gastric pathology image, manifold space mapping is performed on the cross-modal optimized image features and enhanced text features corresponding to the gastric pathology image to obtain the mapped image features and mapped text features corresponding to the gastric pathology image. For each gastric pathology image, based on a preset active spatial focusing mechanism, feature fusion is performed on the mapped image features and mapped text features corresponding to the gastric pathology image to obtain the image-text fusion features corresponding to the gastric pathology image. Based on the image-text fusion features corresponding to each gastric pathology image, a visual language model is trained to obtain a gastric pathology detection model. By applying the method provided in this embodiment of the invention, during the training process of the gastric pathology detection model, cross-modal feature optimization of image features can be performed first through text features. Then, for the text features and image features mapped to the manifold space, feature fusion can be further performed based on the active spatial focusing mechanism. This is beneficial for integrating the cross-modal information of text features and image features, improving the training effect of the model, and thereby improving the detection accuracy of the gastric pathology detection model. Attached Figure Description
[0025] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0026] Figure 1 A flowchart illustrating a model training method based on manifold alignment provided in an embodiment of the present invention;
[0027] Figure 2 This is another flowchart of a model training method based on manifold alignment provided in an embodiment of the present invention;
[0028] Figure 3 Another flowchart of a model training method based on manifold alignment provided for an embodiment of the present invention;
[0029] Figure 4 This is a schematic diagram of a model training device based on manifold alignment provided in an embodiment of the present invention. Detailed Implementation
[0030] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0031] In this application, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0032] This invention provides a model training method based on manifold alignment. The method can be applied to a model training platform, and its execution entity can be the processor of the model training platform. The flowchart of the method is as follows: Figure 1 As shown, it includes:
[0033] S101: Determine the training dataset; the training dataset includes multiple gastric pathology images and their corresponding descriptive text and annotation labels.
[0034] In the method provided by this invention, when training a gastric pathology detection model, a gastric pathology image dataset can be obtained first, which includes gastric pathology images containing different gastric pathologies. Descriptive text can be configured for each gastric pathology image using either a large model configuration or manual configuration to describe the gastric pathological features. Simultaneously, each gastric pathology image can be annotated using annotation tools or manual configuration, that is, the pathological type and lesion location of the gastric pathology image can be annotated. This annotated information serves as the annotation label for the gastric pathology image, which can also be understood as the sample label in model training. The pre-acquired gastric pathology images, along with the pre-configured descriptive text and annotation labels corresponding to each gastric pathology image, are used as a training dataset to train the gastric pathology model.
[0035] S102: Based on the training dataset, the pre-built visual language model is iteratively trained according to the preset number of training iterations; when entering the current iterative training cycle, the visual language model is applied to extract image features from each gastric pathological image to obtain the image features corresponding to each gastric pathological image, and text features are extracted from the descriptive text corresponding to each gastric pathological image to obtain the text features corresponding to each gastric pathological image.
[0036] In the method provided by this invention, a visual language model can be pre-constructed as the initial model for the gastric pathology model according to actual needs. This visual language model can be configured based on existing visual language algorithms. Furthermore, the number of training iterations can be pre-configured according to actual needs.
[0037] During model training, the pre-built visual language model is iteratively trained based on the training dataset and according to a preset number of training iterations. At the beginning of each training iteration, the visual language model can be used to detect various gastric pathological images and their corresponding descriptive texts to train the visual language model.
[0038] When using a visual language model to detect gastric pathological images and their corresponding descriptive text, the image encoder in the visual language model can be used to extract image features from the gastric pathological image. The extracted image features are then used as the image features corresponding to the gastric pathological image. Similarly, the text encoder in the visual language model can be used to extract text features from the descriptive text corresponding to the gastric pathological image. The extracted text features are then used as the text features corresponding to the gastric pathological image.
[0039] S103: For each gastric pathological image, apply a visual language model to enhance the image features corresponding to the gastric pathological image to obtain the enhanced image features corresponding to the gastric pathological image, and enhance the text features corresponding to the gastric pathological image to obtain the enhanced text features corresponding to the gastric pathological image.
[0040] In the method provided by this invention, when using a visual language model to detect gastric pathological images and their corresponding descriptive text, the image feature extraction branch structure in the visual language model is used to perform feature enhancement processing on the gastric pathological image object, and the image features output by the image feature extraction branch structure are used as the enhanced image features corresponding to the gastric pathological image. Similarly, the text feature extraction branch structure in the visual language model is used to perform feature enhancement processing on the text features corresponding to the gastric pathological image, and the text features output by the text feature extraction branch structure are used as the enhanced text features corresponding to the gastric pathological image.
[0041] S104: For each gastric pathology image, based on its corresponding enhanced text features, perform cross-modal feature optimization on the enhanced image features corresponding to the gastric pathology image to obtain the cross-modal optimized image features corresponding to the gastric pathology image.
[0042] In the method provided by this invention, when using a visual language model to detect gastric pathological images and their corresponding descriptive text, cross-modal feature optimization can be performed on the enhanced image features corresponding to the gastric pathological image based on the enhanced text features. The optimized enhanced image features are then used as the cross-modal optimized image features corresponding to the gastric pathological image, thereby enhancing the image features based on the text features and fusing the text features with the image features.
[0043] S105: For each gastric pathology image, perform manifold space mapping on the cross-modal optimized image features and enhanced text features corresponding to the gastric pathology image to obtain the mapped image features and mapped text features corresponding to the gastric pathology image.
[0044] In the method provided by this invention, when using a visual language model to detect gastric pathological images and their corresponding descriptive text, manifold space mapping can be performed on the cross-modal optimized image features corresponding to the gastric pathological image. This mapping transforms the cross-modal optimized image features to a Riemannian manifold space, and the resulting cross-modal optimized image features mapped to the Riemannian manifold space are used as the mapped image features corresponding to the gastric pathological image. Simultaneously, manifold space mapping can be performed on the enhanced text features corresponding to the gastric pathological image, transforming these enhanced text features to a Riemannian manifold space, and the resulting enhanced text features are used as the mapped text features corresponding to the gastric pathological image.
[0045] S106: For each gastric pathological image, based on a preset active spatial focusing mechanism, the mapped image features and mapped text features corresponding to the gastric pathological image are fused to obtain the image-text fusion features corresponding to the gastric pathological image.
[0046] In the method provided by the embodiments of the present invention, when the visual language model is used to detect the gastric pathological image and its corresponding descriptive text, a preset active spatial focusing mechanism can be used to perform depth alignment and fusion of the mapping image features corresponding to the gastric pathological image and the mapping text features corresponding to the gastric pathological image, and the fused features are used as the image fusion features corresponding to the gastric pathological image.
[0047] S107: Based on the image-text fusion features corresponding to each gastric pathology image, a visual language model is trained to obtain a gastric pathology detection model.
[0048] In the method provided by this invention, a visual language model is used to perform pathological diagnosis based on the image-text fusion features corresponding to each gastric pathological image. Based on the diagnosis results of the visual language model and the annotation labels corresponding to each gastric pathological image, the visual language model is trained and then enters the next training iteration cycle until the training effect of the model reaches the expected level or the number of training iterations reaches the preset number of training iterations, thus completing the training process of the visual language model. The trained visual language model is then used as a gastric pathological detection model.
[0049] By applying the method provided in this embodiment of the invention, during the training process of the gastric pathology detection model, cross-modal feature optimization of image features can be performed first through text features. Then, for the text features and image features mapped to the manifold space, feature fusion can be further performed based on the active spatial focusing mechanism. This is beneficial for integrating the cross-modal information of text features and image features, improving the training effect of the model, and thereby improving the detection accuracy of the gastric pathology detection model.
[0050] exist Figure 1 Based on the method shown, such as Figure 2 As shown, in the method provided by this embodiment of the invention, the process of enhancing the image features corresponding to the gastric pathological image mentioned in step S103 to obtain the enhanced image features corresponding to the gastric pathological image includes:
[0051] S201: Determine whether the current iterative training cycle meets the preset dynamic adjustment conditions.
[0052] In the method provided by this invention, the image regions that the model needs to focus on can be dynamically adjusted according to the model iteration process, so that the model focuses on image regions that are difficult to detect. Dynamic adjustment conditions can be preset according to actual needs, and these conditions are used to determine whether to adjust the image feature enhancement mechanism in the current iteration training cycle.
[0053] S202: If the current iterative training cycle meets the dynamic adjustment conditions, then determine the number of pixels to be filtered for the current iterative training cycle.
[0054] In the method provided by this invention embodiment, if the current iterative training cycle meets the dynamic adjustment conditions, the corresponding number of pixels to be selected can be calculated based on the current training progress and the number of training cycles corresponding to the current iterative training cycle. The number of pixels to be selected can be calculated based on a proportional function of the total pixel value of the image. The value of this proportional function is dynamically changing. As training progresses, the proportion calculated in different iterative training cycles may be different. Usually, the number of pixels to be selected is small in the early stage of training so that the model focuses on a small number of the most representative pixels. In the later stage of training, the number of pixels to be selected is large so that the model focuses on more difficult samples.
[0055] Specifically, for example, let k represent the number of pixels to be filtered, N represent the total number of pixels in the image, and γ(t) represent the filtering ratio function. The calculation method of γ(t) is as follows:
[0056] (Equation 1).
[0057] Where t represents the number of cycles in the current training iteration, i.e., the current training epoch. max This represents the preset number of training iterations, i.e., the total number of training rounds. γ start This represents the proportion of pixels retained during the initial training phase. It can be set according to actual needs, but is usually set to a low value, such as 0.2, to avoid premature sensitivity to noise. γ end This represents the proportion of pixels retained during the later stages of training. It can be set according to actual needs, and is usually set to a higher value, such as 0.5, to accelerate the model's learning of difficult regions.
[0058] The pixel selection quantity k can be calculated as follows:
[0059] k=[γ(t)·N] (Equation 2).
[0060] In the method provided by this embodiment of the invention, if the current iterative training cycle does not meet the dynamic adjustment conditions, global enhancement can be performed on the image features of the gastric pathological image, and the globally enhanced image features can be used as the enhanced image features corresponding to the gastric pathological image.
[0061] S203: Determine the comprehensive score corresponding to each pixel in the gastric pathological image. Based on the preset loss function, the number of pixels selected, and the comprehensive score corresponding to each pixel in the gastric pathological image, determine the difficult region features corresponding to the gastric pathological image from the image features corresponding to the gastric pathological image.
[0062] In the method provided by this invention, during the process of image encoding of a gastric pathological image using a visual language model to obtain image features, a corresponding score is also obtained for each pixel in the gastric pathological image. In this invention, the score corresponding to each pixel is used as the comprehensive score corresponding to that pixel. The calculation method for the comprehensive score corresponding to each pixel is as follows:
[0063] (Equation 3).
[0064] Where j represents a pixel in the image, s j This represents the overall score corresponding to pixel j in the image. This represents the normalized pixel-level loss value for pixel j, used to measure how difficult it is for the model to predict that pixel. This represents the boundary perception intensity of pixel j, indicating whether the pixel is close to the edge of the lesion. This represents the prediction confidence of pixel j, used to measure the model's confidence in that pixel. , as well as , is the model's confidence score or other evaluation metric for a pixel, calculated based on the model's predictions for that pixel, typically based on the difference between the model's current prediction and the true label. α, β, and δ are three weighted hyperparameters representing the combined importance of these three factors, satisfying the following relationship: α + β + δ = 1. This comprehensive score considers difficult samples, boundary pixels, and uncertain regions; pixels with high scores are considered more critical for model training.
[0065] In this embodiment of the invention, a loss function for feature extraction of difficult samples is pre-set according to actual needs. This loss function can be called the LSL loss function.
[0066] The method provided in this embodiment of the invention can locate corresponding difficult region features in the image features based on the comprehensive score of each pixel in the gastric pathological image, the LSL loss function, and the number of pixels screened, i.e., regions that are relatively difficult for the model to detect. Let T represent the set of pixels of the difficult region features, and the calculation method of this pixel set is as follows:
[0067] (Equation 4).
[0068] Here, `Select` represents the mechanism used to construct the pixel set `T`, which is defined as selecting the k highest-scoring pixels from all pixels in the image. The meanings of the other parameters are explained above. The pixel set `T` is derived from the image based on the overall pixel scores `s`. j The top k pixels are selected in descending order of resolution. These pixels are typically the most representative or challenging areas, potentially including small lesions, marginal regions, or other difficult-to-distinguish areas.
[0069] The method provided in this embodiment of the invention can calculate the loss value for feature extraction in difficult regions based on the following method:
[0070] (Equation 5).
[0071] Where Loss represents the loss function, p j y represents the model's predicted value for pixel j during the feature extraction stage. j This represents the true label of pixel j.
[0072] S204: Perform feature enhancement processing on the difficult region features corresponding to the gastric pathological image, and use the processed features as the enhanced image features corresponding to the gastric pathological image, so that the visual language model focuses on training the difficult region.
[0073] In the method provided by the embodiments of the present invention, the features of difficult regions corresponding to gastric pathological images can be enhanced by existing image feature enhancement methods, and the image features obtained after processing can be used as the enhanced image features corresponding to the gastric pathological image.
[0074] Based on the method provided in this invention, a dynamic adjustment mechanism allows the model to focus on the global structure in the early stages of training to prevent noise interference, and gradually shift its attention to more difficult regions, especially those with blurred boundaries and small lesions, in the later stages of training. This strategy helps the model better adapt to the needs of different training stages, thereby improving its learning performance on difficult samples.
[0075] exist Figure 2 Based on the method shown, such as Figure 3 As shown, in the method provided by this embodiment of the invention, the process of determining whether the current iterative training cycle meets the preset dynamic adjustment conditions mentioned in step S201 includes:
[0076] S301: Determine the training loss value and the validation set mean intersection-union ratio (MIU) for the current iteration training cycle.
[0077] In the method provided by this invention, a future trend perception module can be defined. This module predicts the trend of the indicators in the next round of model training, and then evaluates whether the current iteration training cycle meets the dynamic adjustment mechanism. First, the future trend perception module can calculate the training loss value and the validation set mean intersection-over-union ratio (MIU) index of the current iteration training cycle by using the model's predicted values and labeled tags (i.e., true labels) for each pixel in the gastric pathological image in the current iteration training cycle.
[0078] S302: Based on the training loss value and the validation set mean intersection-union ratio (MIU) index, predict the index trend to obtain the predicted training loss value and the predicted MIU value for the current iteration training cycle.
[0079] In the method provided by the embodiments of the present invention, a sliding window model or a lightweight sequence prediction model can be applied to predict the trend of the training loss value and the validation set mean intersection and union ratio (MIRR) index for the next iteration training period, thereby obtaining the predicted value of the training loss value and the predicted value of the validation set MIRR index for the next iteration training period. The predicted value of the training loss value for the next period is used as the predicted value of the training loss for the current iteration training period, and the predicted value of the validation set MIRR index for the next period is used as the predicted value of the validation set MIRR for the current iteration training period.
[0080] S303: Based on the predicted training loss and the predicted intersection-union ratio of the validation set, determine whether the current iterative training cycle meets the preset first trigger condition, and obtain the first trigger condition judgment result.
[0081] In the method provided by this invention, triggering conditions for the training loss and the validation set average intersection-union ratio (AUC) can be set according to actual needs. That is, when the development trends of the training loss value and the AUC meet certain conditions, a dynamic adjustment mechanism is considered to be triggered. For example, a first triggering condition can be set based on the variation range of the training loss value and the AUC. Specifically, when the variation range of the training loss value and the AUC meets certain conditions, the first triggering condition is considered satisfied. t This represents the training loss value for the current training iteration, expressed in mIoU. t This represents the average intersection-union ratio (AUC) of the validation set in the current iteration of the training cycle. This represents the predicted value of the training loss. This represents the predicted value of the validation set mean intersection-union ratio (MIRR). The difference between the predicted training loss and the actual training loss is used to obtain the training loss difference. Similarly, the difference between the predicted MIRR and the MIRR metric is used to obtain the MIRR difference. The specific calculation method is as follows:
[0082] (Formula 6).
[0083] (Equation 7).
[0084] in, This represents the difference in training loss. This represents the difference between the mean intersection and union ratio of the validation sets.
[0085] In this embodiment of the invention, a first triggering condition is defined as the difference in training loss being greater than or equal to zero and the difference in the average intersection-union (AUC) of the validation set being less than zero. That is, if the difference in training loss and the difference in the AUC of the validation set satisfy the following formula, then the current iterative training cycle is considered to satisfy the preset first triggering condition.
[0086] (Equation 8).
[0087] In the method provided by this embodiment of the invention, corresponding parameters can be set to characterize the judgment result of the first trigger condition. The judgment result of the first trigger condition indicates whether the current iteration training cycle meets or does not meet the first trigger condition. For example, parameter S FATE The result of the first trigger condition judgment is represented. If the first trigger condition is met in the current iteration training cycle, then S is set. FATE =1, if the first trigger condition is not met, then set S. FATE =0.
[0088] Furthermore, the pixel selection ratio in the LSL loss function can also be selected in the following way:
[0089] (Equation 9).
[0090] Where Loγ(t) represents the pixel selection ratio, γ low and γ high In this embodiment of the invention, γ can be set according to actual needs. low =0.2, γ high =0.5.
[0091] S304: Determine the curvature anomaly ratio corresponding to the current iterative training cycle.
[0092] In the method provided by this invention, a training anomaly detection module can be defined. This module detects whether the model training is abnormal, and then evaluates whether the current iteration training cycle meets the dynamic adjustment mechanism. The training anomaly detection module can evaluate the curvature anomaly index of the current iteration training cycle through the gradient and derivative of the current training loss, and obtain the curvature anomaly ratio. Specifically, the curvature anomaly ratio A... t The calculation method is as follows:
[0093] (Equation 10).
[0094] in, This represents the first-order gradient of the total loss in the current training iteration. It represents the second derivative of the total loss in the current iteration training cycle.
[0095] S305: Based on the curvature anomaly ratio, determine whether the current iterative training cycle meets the preset second trigger condition, and obtain the second trigger condition judgment result.
[0096] In the method provided by this invention, triggering conditions for the curvature anomaly ratio can be set according to actual needs. That is, when the curvature anomaly ratio meets certain conditions, a dynamic adjustment mechanism is considered to need to be triggered. For example, a second triggering condition can be set based on the relationship between the current curvature anomaly ratio and the historical curvature anomaly ratio. Whether the current iterative training cycle meets the second triggering condition can be determined using the following formula:
[0097] (Equation 11).
[0098] Where δ represents the preset curvature anomaly detection threshold, and EMA(A) represents the exponential moving average of historical curvature, which refers to the curvature anomaly ratio corresponding to the iterative training cycles before the current cycle.
[0099] In the method provided by this embodiment of the invention, corresponding parameters can be set to characterize the judgment result of the second trigger condition. The judgment result of the second trigger condition indicates whether the current iteration training cycle meets or does not meet the second trigger condition. For example, parameter S ASL This indicates the result of the second trigger condition judgment. If the second trigger condition is met in the current iteration training cycle, then S is set. ASL =1, if the second trigger condition is not met, then set S. ASL =0.
[0100] S306: Determine the average loss conflict degree corresponding to the current iterative training cycle.
[0101] In the method provided by this invention, a multi-task conflict detection module (MOCL) can be defined. This module monitors whether there is an optimization direction conflict among multiple loss tasks during model training, and then evaluates whether the current iteration training cycle meets the dynamic adjustment mechanism. Specifically, the average loss conflict degree can be calculated as follows:
[0102] (Equation 12).
[0103] Among them, Conf tThe average loss conflict degree is represented by M, which is the number of loss functions currently involved in the optimization, i.e., the various loss functions involved in the current task. This represents the gradient of the i-th loss function. This represents the cosine similarity along the gradient direction.
[0104] S307: Based on the average loss conflict degree, determine whether the current iterative training cycle meets the preset third trigger condition, and obtain the third trigger condition judgment result.
[0105] In the method provided by this invention, triggering conditions for the average loss conflict degree can be set according to actual needs. That is, when the average loss conflict degree meets certain conditions, a dynamic adjustment mechanism is considered to need to be triggered. For example, based on the relationship between the average loss conflict degree and a preset threshold, a third triggering condition can be set, such as the average loss conflict degree being greater than the preset threshold. Whether the current iteration training cycle meets the third triggering condition can be determined using the following formula:
[0106] Conf t >0.8 (Equation 13).
[0107] In the method provided by this invention, corresponding parameters can be set to characterize the judgment result of the third triggering condition. The judgment result of the third triggering condition indicates whether the current iteration training cycle meets or does not meet the third triggering condition. For example, a conflict triggering flag S... MOCL This indicates the result of the third trigger condition judgment. If the third trigger condition is met in the current iteration training cycle, then S is set. MOCL =1, if the third trigger condition is not met, then set S. MOCL =0.
[0108] S308: Based on the results of the first trigger condition judgment, the second trigger condition judgment, and the third trigger condition judgment, determine the dynamic adjustment activation probability corresponding to the current iterative training cycle.
[0109] In the method provided by this embodiment of the invention, an activation probability fusion module can be defined. This module evaluates the probability of activating the dynamic adjustment mechanism based on the judgment results of the first trigger condition, the second trigger condition, and the third trigger condition, to obtain the dynamically adjusted activation probability corresponding to the current iteration training cycle. Specifically, the calculation method of the dynamically adjusted activation probability is as follows:
[0110] P t =σ(γ1·S FATE +γ2·S ASL +γ3·S MOCL (Equation 14).
[0111] Among them, P tThe activation probability is dynamically adjusted for the current iteration training cycle. σ(·) is the Sigmoid function, used to smoothly map the data to the (0,1) interval. γ1, γ2, and γ3 represent the weights corresponding to each judgment result, satisfying γ1+γ2+γ3=1. In this embodiment of the invention, the following values are used: γ1=0.5, γ2=0.3, γ3=0.2.
[0112] S309: Determine whether the dynamically adjusted activation probability is greater than the preset probability threshold.
[0113] In the method provided by the embodiments of the present invention, a probability threshold can be set in advance according to actual needs, and the dynamically adjusted activation probability can be compared with the preset probability threshold. For example, the probability threshold can be set to 0.5.
[0114] S310: If the dynamically adjusted activation probability is greater than the probability threshold, then the current iterative training cycle is determined to meet the dynamic adjustment condition.
[0115] In the method provided by this embodiment of the invention, if the dynamically adjusted activation probability is greater than a preset probability threshold, the current iterative training cycle is determined to meet the dynamic adjustment condition. If the dynamically adjusted activation probability is less than the preset probability threshold, the current iterative training cycle is determined to not meet the dynamic adjustment condition. When the dynamically adjusted activation probability is equal to the preset probability threshold, the current iterative training cycle can be configured to meet the dynamic adjustment condition. Specifically, the selection of the loss function during the training process can be determined based on the following formula:
[0116] (Equation 15).
[0117] Among them, Loss t P(t) represents the loss function applied in the current training iteration, which is the dynamically adjusted activation probability P. t CE represents the CE (cross-entropy) loss function, and Dice represents the Dice loss function. LSL γ(t) This represents the LSL loss function.
[0118] exist Figure 1 Based on the method shown, the method provided in this embodiment of the invention, in step S103, involves feature enhancement of the text features corresponding to the gastric pathological image to obtain the enhanced text features corresponding to the gastric pathological image, including:
[0119] Word vector embedding is performed on the text features corresponding to the gastric pathological image to obtain the word vector sequence corresponding to the text features.
[0120] The method provided in this invention extracts text features using a hierarchical discriminator based on semantic part-of-speech tags and a hierarchical attention mechanism, and then performs text enhancement according to task requirements. First, for each word in the text features corresponding to a gastric pathology image, a word embedding layer is used to map it to a word vector space, thereby obtaining the word vectors corresponding to each word in the text features. The word vectors corresponding to each word form a sequence of word vectors corresponding to the text feature. For example, the text features corresponding to a gastric pathology image can be represented as: X={x1,x2,x3,…,x…} n}, where x i Representing words. The sequence of word vectors obtained by word vector embedding processing of text features can be represented as: E={e1,e2,…,e...} n}, e i ∈R d Among them, e i =f embed (x i ), representing the word x i The corresponding word vector, where d represents the word vector dimension.
[0121] The word vector sequence is layered to obtain the word vectors corresponding to each layer.
[0122] In the method provided by this invention, semantic units in the descriptive text corresponding to a gastric pathological image are modeled using a preset hierarchical discriminator, that is, the text is hierarchically divided according to the semantic properties of language. This invention divides the text into three levels: the first level is local location information, the second level is core entities, and the third level is main feature descriptions. For example, for the descriptive text "There is a noticeable dark bulge in the upper left corner," the local location information "upper left corner" can be classified as the semantics of the first level, "a bulge" as the semantics of the second level, and "dark bulge" as the semantics of the third level. In the specific data processing, semantic hierarchical division is performed using word vector sequences to obtain word vectors for each level. The core part-of-speech function f for hierarchical modeling can be predefined. core This is used to extract key semantic levels. For each semantic level, the core part-of-speech function f is used. core Semantic extraction can be formally represented as follows:
[0123] (Equation 16).
[0124] Among them, H k w represents the semantics of the k-th layer. k This represents the word vector in the k-th layer. For example, for the text "There is a noticeable dark bulge in the upper left corner", the layer order is: H1=f core (Top left corner), H2=f core (One piece), H3=f core(Dark raised areas).
[0125] Determine the attention weight for each word vector at each level.
[0126] The method provided in this invention introduces a hierarchical attention mechanism, assigning higher weights to key features at specific levels for feature extraction. Therefore, for each word vector in each level, the attention weight of that word vector in that level is calculated according to a pre-defined attention weight calculation method. In calculating the attention weight, it is not only dependent on the local representation of the word, but also needs to consider its relationship with other words and the influence of the context. The attention weight calculation method is as follows:
[0127] (Equation 17).
[0128] in, This represents the attention weight of the i-th word vector in the k-th layer, which measures the relative importance of the word to other words in the current layer. This represents the weighted representation of the i-th word in the k-th layer, combining word vectors. and context information The meaning of the parameters carried by z with other subscripts and superscripts is the same as... Similar to the above, I will not elaborate further here. W represents the learning weights of the k-th layer, used to adjust the weighted contributions of word vectors and context representations. It determines the importance of the k-th layer features in the attention calculation. k W represents the dynamic context-aware weights of the k-th layer, used to measure the influence of contextual information on word importance at that layer. It determines the weight allocation of context in different contexts. j and W m These are also context-aware weights for the corresponding levels. The context information of the i-th word can be other words related to its meaning, disease types, pathological descriptions, etc. Indicates the weight of inter-layer interactions. This indicates the mutual influence between different layers, for example This represents the relationship between the k-th layer and the j-th layer, which can be a semantic connection between different layers. This indicates the relationship between the k-th layer and other layers (such as the layer above or the layer below), that is, the mutual influence between a layer and its adjacent layers. The calculation method is as follows:
[0129] (Equation 18).
[0130] in, This represents the word vector of the i-th word in the k-th layer. The context information of the i-th word can be represented, such as the surrounding vocabulary, sentence structure, or the context of the entire paragraph. f(·) represents the function used to fuse word vectors and context information; common operations include weighted summation, concatenation, or neural networks.
[0131] Based on a pre-defined bidirectional long short-term memory network and attention weights corresponding to each word vector, contextual features of each word vector are fused, and the fusion result is used as the enhanced text feature corresponding to the gastric pathological image.
[0132] The method provided in this invention can encode word vectors based on a hierarchical bidirectional long short-term memory network (BiLSTM) to extract and fuse text features. An attention mechanism is introduced during feature extraction, where each word vector is extracted and fused according to its corresponding attention weight. The fusion result is used as the enhanced text feature corresponding to the gastric pathological image. In the process of feature extraction and fusion using BiLSTM, the input of each layer comes from the hidden state of the previous layer. The expression for this process is as follows:
[0133] (Equation 19).
[0134] (Equation 20).
[0135] (Equation 21).
[0136] in, This represents the hidden state of the k-th layer.
[0137] exist Figure 1 Based on the method shown, in the method provided by this embodiment of the invention, the process of performing cross-modal feature optimization on the enhanced image features corresponding to the gastric pathological image based on its corresponding enhanced text features, as mentioned in step S104, to obtain the cross-modal optimized image features corresponding to the gastric pathological image, includes:
[0138] The pathological location of the gastric pathological image is obtained by indexing the enhanced text features corresponding to the image using a preset relative position discriminator.
[0139] In the method provided by this invention, a relative position discriminator is predefined. This discriminator can identify the semantic representation of location in the text through the basic positional semantic features of the text, thereby locating the pathological location in a gastric pathological image. Specifically, the relative position discriminator can identify the positional information in the enhanced text features corresponding to the gastric pathological image, output the region index corresponding to that location, and use the location corresponding to that region index as the pathological location corresponding to the gastric pathological image. For example, assuming the gastric pathological image is R, and its center point coordinates are C=(x... c ,y c The image is divided into four regions: R1, R2, R3, and R4.
[0140] (Equation 22).
[0141] Where R1 represents the features of the upper left region, R2 represents the features of the upper right region, R3 represents the features of the lower left region, and R4 represents the features of the lower right region. Relative position discriminator f pos Location information T acting on text features pos Output the region index r corresponding to this position:
[0142] (Equation 23).
[0143] For example, for the description "there is a noticeable dark bulge in the upper left corner", the index corresponding to "upper left corner" is identified, and "upper left corner" is the pathological location corresponding to the gastric pathological image.
[0144] In the enhanced image features corresponding to the pathological image of the stomach, the feature region corresponding to the pathological location is determined.
[0145] In the method provided by the embodiments of the present invention, the image features of the corresponding region can be located in the corresponding enhanced image features according to the pathological location corresponding to the gastric pathological image, and the feature region in the enhanced image features that matches the pathological location can be taken as the feature region corresponding to the pathological location.
[0146] Edge sensing processing is performed on the feature region to obtain the edge sensing features corresponding to the feature region.
[0147] In the method provided by this invention, a centroid extraction module (GCEmodule) can be pre-configured to perform a round of deep image feature extraction and fusion on the feature regions corresponding to the pathological location. The centroid extraction module is configured with three parallel processing paths, focusing on edge details, texture response, and high-level semantic information, respectively. Each path is jointly modeled in the fusion stage to improve the discriminativeness and robustness of the overall feature representation. These three processing paths are the Edge Path, the Texture Path, and the Semantic Path.
[0148] In the method provided by this embodiment of the invention, edge sensing is performed on the feature region corresponding to the pathological location in the edge-aware path, and edge-sensitive features are extracted. The extracted edge-sensitive features are then used as the edge-aware features corresponding to the feature region. Specifically, the feature region corresponding to the pathological location is first processed by a channel compression layer. This channel compression layer uses a 1×1 convolution kernel with a stride of 1, mapping the number of channels C of the input features to C / 4. The expression for the channel compression process is as follows:
[0149] (Equation 24).
[0150] in, This represents the original input feature map. That is, the dimensions are height H, width W, and number of channels C. Indicates the convolution kernel weights, The corresponding operation is to perform a 3×3 convolution on the features with C input channels, and output C' (e.g., C / 4) channels. This represents the bias term of the convolution. Add bias according to the output channel. This represents the standard convolution operation. This represents a non-linear activation function that retains positive numbers and suppresses negative numbers. This represents the output feature map after convolution and activation. .
[0151] For features that have undergone channel compression, edge filtering is performed using a non-trainable Sobel edge filter (kernel size 3×3) to enhance local edge contour responses, while maintaining the same number of output channels, C / 4. The edge-filtered features are then spatially downsampled using a depthwise convolution, and the resulting features are used as edge-aware features. Specifically, the depthwise convolution uses a 3×3 kernel with a stride of 2, and the output size is H / 2×W / 2. This process can be formally represented as:
[0152] (Equation 25).
[0153] in, Represents the input feature map, The space dimensions are H×W, and the number of channels C'=C / 4. This represents a depthwise separable convolution (DepthwiseConv) with a kernel size of 3×3 and a stride of 2, achieving spatial downsampling. Output This means that the spatial size of the output feature map is reduced by half, while the number of channels remains the same.
[0154] Edge-aware paths can extract edge-sensitive features with extremely low parameter overhead, making them suitable for boundary modeling under high-resolution input.
[0155] Texture-aware processing is performed on the feature regions to obtain the texture-aware features corresponding to the feature regions.
[0156] In the method provided by this invention embodiment, the texture-aware path enhances the texture features of the feature regions corresponding to the pathological locations, and the enhanced features are used as texture-aware features. Specifically, firstly, a standard 3×3 convolution reduces the number of input feature channels from C to C / 4 while maintaining the original spatial resolution. Then, a directionally selective Gabor convolution kernel (5×5) is used for feature enhancement, emphasizing the texture response at different directions and frequencies. This process can be formally represented as:
[0157] (Equation 26).
[0158] Then, a 3×3 depthwise separable convolution (with a stride of 2) is used to spatially compress the feature-enhanced features to obtain output features (i.e., texture-aware features) with a size of H / 2×W / 2, thereby capturing local texture differences.
[0159] Semantic aggregation processing is performed on the feature regions to obtain the semantic aggregated features corresponding to the feature regions.
[0160] In the method provided by this embodiment of the invention, the semantic aggregation path involves semantic aggregation processing of the feature regions corresponding to the pathological locations to fuse the semantics within the features and obtain semantically aggregated features. Specifically, the features of the feature regions are used as input features, and the number of channels of the input features is compressed from C to C / 2 using a standard 3×3 convolution to preserve spatial dimensions. Then, the SE-Block (Squeeze-and-Excitation) module is introduced, which recalibrates the channel dimensions using global average pooling (GAP) and two fully connected layers. This module can be formally described as follows:
[0161] (Equation 27).
[0162] Where SE(X) represents the output of the SE-Block module, and X represents the characteristics of the input SE-Block module. Represents the ReLU activation function. It is the Sigmoid activation function. and These represent the weights of the fully connected layer, and GAP() represents the global average pooling operation.
[0163] The channel attention weights of the SE-Block module's output are multiplied channel-by-channel with the original feature map to enhance key channels. Finally, a 3×3 convolution (with a stride of 2) is used for downsampling, outputting features of size H / 2×W / 2 (i.e., semantic aggregation features), with the number of channels remaining at C / 2.
[0164] Edge-aware features, texture-aware features, and semantic aggregation features are fused, and the fusion result is used as the cross-modal optimized image feature corresponding to the gastric pathological image.
[0165] In the method provided by this invention, after obtaining the edge-aware features, texture-aware features, and semantic aggregation features corresponding to the feature region, these three types of features can be fused and stitched together in the channel dimension through a preset fusion module. The stitched features (i.e., the fusion result) are then used as the cross-modal optimized image features corresponding to the gastric pathological image. Specifically, firstly, the edge-aware features, texture-aware features, and semantic aggregation features are stitched together in the channel dimension to obtain the fused features:
[0166] (Equation 28).
[0167] in, This represents the output features from the edge path. . This represents the output features from the texture path. . The output features represent the semantic path. . This indicates a concatenation operation along the channel dimension. This indicates the blending characteristics after splicing. The number of channels is restored to the original number of channels C.
[0168] Then, a gated attention mechanism with a bottleneck structure is used to learn the channel importance weights of different paths. This process is implemented by two fully connected layers and can be formally represented as follows:
[0169] (Equation 29).
[0170] in, This represents the input feature map, which is typically the result of concatenating features from multiple paths. . This indicates that a global average pooling operation is performed on F, that is, the average is calculated for each channel to obtain a global description of the channel. . , representing the weights of the first fully connected layer, used for channel compression. , represents the weight of the second fully connected layer, used to restore the channel dimension, and r is the intermediate compression dimension, usually r=C / 8 or 16. Represents the ReLU activation function. This represents the Sigmoid activation function, used to generate the weight coefficients (0~1) for each channel. , representing the attention weight for each channel. The ultimate effect of the attention weights: the channel attention broadcasts the weights and multiplies them by F, i.e., F' = F.Gate(F).
[0171] Finally, channel compression is performed using a 1×1 convolution to obtain C channels. out The output features are used as the fusion result of edge-aware features, texture-aware features, and semantic aggregation features (i.e., cross-modal optimized image features F). out This can be formally represented as:
[0172] (Formula 30).
[0173] Where Conv represents the convolution operation, F gate This represents the fused feature output after processing edge-aware features, texture-aware features, and semantic aggregation features based on a gating attention mechanism.
[0174] exist Figure 1 Based on the method shown, in the method provided by this embodiment of the invention, the process mentioned in step S105 of performing manifold space mapping on the cross-modal optimized image features and enhanced text features corresponding to the gastric pathological image to obtain the mapped image features and mapped text features corresponding to the gastric pathological image includes:
[0175] The cross-modal optimized image features corresponding to the gastric pathological image are processed by manifold space mapping using a preset symmetric positive definite matrix, and the processing result is used as the mapped image features corresponding to the gastric pathological image.
[0176] In the method provided by the embodiments of the present invention, optimized image features and text features are respectively mapped to the Riemannian manifold space to obtain manifold embedding features. By adopting bidirectional geodesic contrastive learning loss to optimize the relationship between image and text features, introducing a curvature-aware manifold attention mechanism, introducing a pseudo-negative sample mechanism to generate pseudo-negative samples and manifold dynamic weighted multi-task planning, cross-modal semantic consistency is further enhanced.
[0177] Specifically, in the feature mapping process, cross-modal optimized features are mapped using a symmetric positive definite matrix (SPD) to a Riemannian manifold space. The image features mapped to the Riemannian manifold space are then used as the mapped image features corresponding to the gastric pathological image. The process of manifold space mapping of cross-modal optimized image features can be formally represented as follows:
[0178] (Equation 31).
[0179] Among them, F img I represents the mapped image features, and I represents the cross-modal optimized image features. This represents the feature extraction mapping operation. This represents the symmetry operation. This indicates a matrix exponential mapping operation.
[0180] The enhanced text features corresponding to the gastric pathological image are processed by manifold space mapping using a preset covariance matrix, and the processing result is used as the mapped text features corresponding to the gastric pathological image.
[0181] In the method provided by this invention, the enhanced text features corresponding to the gastric pathological image are mapped to a manifold space using the covariance matrix, and the text features mapped to the Riemannian manifold space are used as the mapped text features corresponding to the gastric pathological image. The process of manifold space mapping of the enhanced text features can be formally represented as follows:
[0182] (Equation 32).
[0183] Among them, F text Let t represent the mapped text features, N represent the number of word vectors in the enhanced text features, and t represent the number of word vectors in the enhanced text features. i The vector representation of the i-th word (i.e., the word vector). This represents the mean of all word vectors.
[0184] exist Figure 1 Based on the method shown, in the method provided by this embodiment of the invention, the process of fusing the mapped image features and mapped text features corresponding to the gastric pathological image based on a preset active spatial focusing mechanism in step S106 to obtain the image-text fusion features corresponding to the gastric pathological image includes:
[0185] Two-way distance measurement is performed between the mapped image features and mapped text features corresponding to the gastric pathological image to obtain the first geodesic distance and the second geodesic distance corresponding to the gastric pathological image.
[0186] In the method provided by this invention, geodesic distance between image and text features is defined in the SPD manifold space. For example, the geodesic distance from mapped image features to mapped text features can be calculated as follows:
[0187] (Equation 33).
[0188] d geodesic (F) img F text The distance from the mapped image feature to the mapped text feature is represented by , and the calculation method for the distance from the mapped text feature to the mapped image feature is similar. In this embodiment of the invention, bidirectional distance measurement is performed on the mapped image feature and the mapped text feature, that is, the geodesic distance from the mapped image feature to the mapped text feature is calculated and used as the first geodesic distance, and the geodesic distance from the mapped text feature to the mapped image feature is calculated and used as the second geodesic distance.
[0189] To ensure alignment consistency in both the image-to-text and text-to-image directions, a bidirectional geodesic contrast loss L is defined. bi-geodesic for:
[0190] (Equation 34).
[0191] This loss comprehensively reflects the degree of bidirectional consistency of text and image features in the manifold space.
[0192] The weighted attention corresponding to the gastric pathological image is determined based on the first geodesic distance and the second geodesic distance.
[0193] The method provided in this embodiment of the invention configures an active spatial focusing mechanism, introduces a curvature-aware manifold attention mechanism (to solve the problem of modeling the importance of local regions), and defines weighted attention during fusion:
[0194] (Equation 35).
[0195] Where, α i This represents the weighted attention corresponding to the i-th local unit in the mapped text features, where γ is the curvature adjustment weight parameter, and F... text,i F represents the i-th local unit in the mapped text features. text,j The meaning is the same. K(F) text,iThe ) represents the local curvature index, which measures the degree of manifold curvature at the feature, i.e., the degree of local nonlinear change of the feature in the manifold space. This index can be obtained by performing second-order geometric analysis on the mapped text features.
[0196] In the method provided by the embodiments of the present invention, the weighted attention corresponding to each unit in the mapped text features can be calculated based on the principle shown in Equation 35, and the weighted attention of each unit can be used as the weighted attention corresponding to the gastric pathological image.
[0197] Based on the weighted attention corresponding to the gastric pathological image, the mapped image features and mapped text features corresponding to the gastric pathological image are fused, and the fusion result is used as the image-text fusion feature corresponding to the gastric pathological image.
[0198] In the method provided by this invention, the mapped image features and mapped text features can be fused based on the weighted attention corresponding to the gastric pathological image according to a pre-set cross-modal feature fusion method to obtain image-text fused features. For example, the image-text fused feature F fusion This can be formally represented as:
[0199] (Equation 36).
[0200] exist Figure 1 Based on the method shown, the method provided in this embodiment of the invention further includes:
[0201] For each gastric pathology image, pseudo-negative samples are generated based on the mapped text features corresponding to the gastric pathology image.
[0202] The visual language model is trained based on pseudo-negative samples corresponding to various gastric pathological images.
[0203] The method provided in this embodiment of the invention establishes an active discrimination enhancement mechanism, employing a method for optimizing mutual information of negative samples and introducing a pseudo-negative sample mechanism to generate pseudo-negative samples, thereby addressing the problem of insufficient semantic consistency. Based on the mapped text features of gastric pathological images, corresponding pseudo-negative samples can be generated using the appropriate mechanism for model training. Specifically, pseudo-negative samples can be formally represented as follows:
[0204] (Equation 37).
[0205] in, Indicates a false negative sample. This represents the small-amplitude noise perturbation matrix.
[0206] Furthermore, in the method provided in this embodiment of the invention, the negative sample mutual information optimization loss is defined to actively suppress surface matching phenomena and strengthen cross-modal deep semantic connections. The calculation method of this loss is as follows:
[0207] (Equation 38).
[0208] in, For mutual information estimation, This is the adjustment coefficient for the difference in mutual information between positive and negative samples.
[0209] exist Figure 1 Based on the method shown, the method provided in this embodiment of the invention further includes:
[0210] Based on the image-text fusion features corresponding to each gastric pathological image, the complexity index corresponding to each gastric pathological image is determined.
[0211] The method provided in this embodiment of the invention establishes an active training attention transfer mechanism, introducing manifold dynamic weighted multi-task programming to resolve inter-task conflicts and training adaptability issues. First, for each gastric pathological image corresponding to an image-text fusion feature, the complexity of each image-text fusion feature is calculated, and the complexity of each feature is used as a complexity index for the corresponding gastric pathological image.
[0212] Based on the complexity index corresponding to each gastric pathological image, the preset complexity sensitivity adjustment coefficient, and the preset complexity threshold corresponding to each training task, the task weight corresponding to each training task is determined.
[0213] In the method provided by this invention, a complexity-sensitive adjustment coefficient and a complexity threshold (i.e., a preset complexity threshold) corresponding to each training task can be pre-set according to the training requirements of each training task. Specifically, the task weight corresponding to each training task can be calculated according to the following task weight calculation method, based on the complexity index corresponding to each gastric pathological image, the complexity-sensitive adjustment coefficient, and the preset complexity threshold corresponding to each training task:
[0214] (Equation 39).
[0215] Where, λ i Let C(F) represent the training weights corresponding to the i-th training task, β represent the complexity-sensitive adjustment coefficient, and C(F) represent the training weights corresponding to the i-th training task. fusion θ represents the complexity index corresponding to a gastric pathological image. i This represents the complexity threshold corresponding to the i-th training task.
[0216] Based on the task weights corresponding to each training task, the weights of each training task are adaptively adjusted.
[0217] In the method provided by this invention, the training weights of each training task can be adaptively adjusted according to the task weights corresponding to each training task. That is, the task weights corresponding to each training task are used as the adjusted training weights for model training. In this invention embodiment, the total loss function L of each training task is used... total Defined as:
[0218] (Formula 40).
[0219] Where M represents the total number of training tasks, L i (·) represents the loss of the i-th subtask (such as segmentation, classification, or generating diagnostics).
[0220] and Figure 1 Corresponding to the manifold alignment-based model training method shown, this embodiment of the invention also provides a manifold alignment-based model training apparatus for training models. Figure 1 The specific implementation of the method shown is illustrated in the following diagram. Figure 4 As shown, it includes:
[0221] The first determining unit 401 is used to determine the training dataset; the training dataset includes multiple gastric pathological images and the corresponding descriptive text and annotation labels of the gastric pathological images.
[0222] The feature extraction unit 402 is used to iteratively train the pre-built visual language model based on the training dataset and according to the preset number of training times. When entering the current iterative training cycle, the visual language model is used to extract image features from each gastric pathological image to obtain the image features corresponding to each gastric pathological image, and to extract text features from the descriptive text corresponding to each gastric pathological image to obtain the text features corresponding to each gastric pathological image.
[0223] The feature enhancement unit 403 is used to apply a visual language model to each gastric pathological image to enhance the image features corresponding to the gastric pathological image, thereby obtaining the enhanced image features corresponding to the gastric pathological image, and to enhance the text features corresponding to the gastric pathological image, thereby obtaining the enhanced text features corresponding to the gastric pathological image.
[0224] The feature optimization unit 404 is used to perform cross-modal feature optimization on the enhanced image features corresponding to each gastric pathology image based on its corresponding enhanced text features, so as to obtain the cross-modal optimized image features corresponding to the gastric pathology image.
[0225] The feature mapping unit 405 is used to perform manifold space mapping on the cross-modal optimized image features and enhanced text features corresponding to each gastric pathological image to obtain the mapped image features and mapped text features corresponding to the gastric pathological image.
[0226] The feature fusion unit 406 is used to perform feature fusion on the mapped image features and mapped text features corresponding to each gastric pathological image based on a preset active spatial focusing mechanism, so as to obtain the image-text fusion features corresponding to the gastric pathological image.
[0227] The feature training unit 407 is used to train the visual language model based on the image-text fusion features corresponding to each gastric pathology image to obtain a gastric pathology detection model.
[0228] Using the apparatus provided in this embodiment of the invention, during the training process of the gastric pathology detection model, cross-modal feature optimization of image features can be performed first through text features. Then, for the text features and image features mapped to the manifold space, feature fusion can be further performed on the text features and image features based on the active spatial focusing mechanism. This is beneficial for integrating the cross-modal information of text features and image features, improving the training effect of the model, and thereby improving the detection accuracy of the gastric pathology detection model.
[0229] exist Figure 4 Based on the device shown, the device provided in this embodiment of the invention can be further extended to include multiple units. The functions of each unit can be found in the descriptions of the various embodiments of the manifold alignment-based model training method provided above, and will not be further illustrated here.
[0230] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on its differences from other embodiments. In particular, for system or system embodiments, since they are fundamentally similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. The systems and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. Components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0231] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0232] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A model training method based on manifold alignment, characterized in that, include: Determine the training dataset; The training dataset includes multiple gastric pathology images, as well as corresponding descriptive text and annotation labels for each gastric pathology image; Based on the training dataset, the pre-built visual language model is iteratively trained according to a preset number of training iterations. When entering the current iterative training cycle, the visual language model is applied to extract image features from each of the gastric pathological images to obtain the image features corresponding to each of the gastric pathological images. Text features are also extracted from the descriptive text corresponding to each of the gastric pathological images to obtain the text features corresponding to each of the gastric pathological images. For each of the gastric pathological images, the visual language model is applied to enhance the image features corresponding to the gastric pathological image to obtain the enhanced image features corresponding to the gastric pathological image, and the text features corresponding to the gastric pathological image are enhanced to obtain the enhanced text features corresponding to the gastric pathological image. For each of the gastric pathology images, based on its corresponding enhanced text features, cross-modal feature optimization is performed on the enhanced image features corresponding to the gastric pathology image to obtain the cross-modal optimized image features corresponding to the gastric pathology image; For each of the gastric pathological images, manifold space mapping is performed on the cross-modal optimized image features and enhanced text features corresponding to the gastric pathological image to obtain the mapped image features and mapped text features corresponding to the gastric pathological image; For each of the gastric pathological images, based on a preset active spatial focusing mechanism, the mapping image features and mapping text features corresponding to the gastric pathological image are fused to obtain the image-text fusion features corresponding to the gastric pathological image. Based on the image-text fusion features corresponding to each of the gastric pathology images, the visual language model is trained to obtain a gastric pathology detection model. Specifically, the step of performing cross-modal feature optimization on the enhanced image features corresponding to the gastric pathological image based on its corresponding enhanced text features to obtain the cross-modal optimized image features corresponding to the gastric pathological image includes: The pathological location of the gastric pathological image is obtained by indexing the enhanced text features corresponding to the gastric pathological image using a preset relative position discriminator. In the enhanced image features corresponding to the pathological image of the stomach, the feature region corresponding to the pathological location is determined; Edge sensing processing is performed on the feature region to obtain the edge sensing features corresponding to the feature region; The feature region is subjected to texture-aware processing to obtain the texture-aware features corresponding to the feature region; The feature regions are subjected to semantic aggregation processing to obtain the semantic aggregated features corresponding to the feature regions; The edge-aware features, texture-aware features, and semantic aggregation features are fused, and the fusion result is used as the cross-modal optimized image feature corresponding to the gastric pathological image. The method, based on a preset active spatial focusing mechanism, fuses the mapped image features and mapped text features corresponding to the gastric pathological image to obtain the image-text fusion features corresponding to the gastric pathological image, including: Two-way distance measurement is performed between the mapped image features and mapped text features corresponding to the gastric pathological image to obtain the first geodesic distance and the second geodesic distance corresponding to the gastric pathological image; Based on the first geodesic distance and the second geodesic distance, the weighted attention corresponding to the gastric pathological image is determined; Based on the weighted attention corresponding to the gastric pathological image, the mapped image features and mapped text features corresponding to the gastric pathological image are fused, and the fusion result is used as the image-text fusion feature corresponding to the gastric pathological image.
2. The model training method based on manifold alignment according to claim 1, characterized in that, The step of enhancing the image features corresponding to the gastric pathological image to obtain the enhanced image features corresponding to the gastric pathological image includes: Determine whether the current iterative training cycle meets the preset dynamic adjustment conditions; If the current iterative training cycle meets the dynamic adjustment conditions, then the number of pixels to be filtered corresponding to the current iterative training cycle is determined. The comprehensive score corresponding to each pixel in the gastric pathological image is determined. Based on the preset loss function, the number of pixels screened, and the comprehensive score corresponding to each pixel in the gastric pathological image, the difficult region features corresponding to the gastric pathological image are determined from the image features corresponding to the gastric pathological image. Feature enhancement processing is performed on the difficult regions corresponding to the gastric pathological image, and the processed features are used as the enhanced image features corresponding to the gastric pathological image, so that the visual language model focuses on training the difficult regions.
3. The model training method based on manifold alignment according to claim 2, characterized in that, The step of determining whether the current iterative training cycle meets the preset dynamic adjustment conditions includes: Determine the training loss value and the validation set intersection-union ratio (MIU) corresponding to the current iterative training cycle; Based on the training loss value and the validation set mean intersection and union ratio (MIRR) index, the index trend is predicted to obtain the predicted training loss value and the predicted MIRR value corresponding to the current iterative training cycle. Based on the predicted training loss value and the predicted validation set intersection-union ratio value, determine whether the current iterative training cycle meets the preset first triggering condition, and obtain the first triggering condition judgment result; Determine the curvature anomaly ratio corresponding to the current iterative training cycle; Based on the curvature anomaly ratio, determine whether the current iterative training cycle meets the preset second triggering condition, and obtain the second triggering condition judgment result; Determine the average loss conflict degree corresponding to the current iterative training cycle; Based on the average loss conflict degree, it is determined whether the current iterative training cycle meets the preset third trigger condition, and the third trigger condition judgment result is obtained; Based on the first trigger condition judgment result, the second trigger condition judgment result, and the third trigger condition judgment result, the dynamically adjusted activation probability corresponding to the current iterative training cycle is determined; Determine whether the dynamically adjusted activation probability is greater than a preset probability threshold; If the dynamically adjusted activation probability is greater than the probability threshold, then the current iterative training cycle is determined to meet the dynamic adjustment condition.
4. The model training method based on manifold alignment according to claim 1, characterized in that, The step of enhancing the text features corresponding to the gastric pathological image to obtain enhanced text features corresponding to the gastric pathological image includes: Word vector embedding is performed on the text features corresponding to the gastric pathological image to obtain the word vector sequence corresponding to the text features; The word vector sequence is layered to obtain the word vectors corresponding to each layer; Determine the attention weight corresponding to each word vector for each of the aforementioned levels; Based on a pre-defined bidirectional long short-term memory network and attention weights corresponding to each word vector, contextual features are fused to each word vector, and the fusion result is used as the enhanced text feature corresponding to the gastric pathological image.
5. The model training method based on manifold alignment according to claim 1, characterized in that, The step of performing manifold space mapping on the cross-modal optimized image features and enhanced text features corresponding to the gastric pathological image to obtain the mapped image features and mapped text features corresponding to the gastric pathological image includes: The cross-modal optimized image features corresponding to the gastric pathological image are processed by manifold space mapping using a preset symmetric positive definite matrix, and the processing result is used as the mapped image features corresponding to the gastric pathological image. The enhanced text features corresponding to the gastric pathological image are processed by manifold space mapping using a preset covariance matrix, and the processing result is used as the mapped text features corresponding to the gastric pathological image.
6. The model training method based on manifold alignment according to claim 1, characterized in that, Also includes: For each of the gastric pathological images, a pseudo-negative sample generation process is performed based on the mapped text features corresponding to the gastric pathological image to obtain the pseudo-negative sample corresponding to the gastric pathological image. The visual language model is trained based on the pseudo-negative samples corresponding to each of the aforementioned gastric pathological images.
7. The model training method based on manifold alignment according to claim 1, characterized in that, Also includes: Based on the image-text fusion features corresponding to each of the gastric pathological images, the complexity index corresponding to each of the gastric pathological images is determined. Based on the complexity index corresponding to each of the gastric pathological images, the preset complexity sensitivity adjustment coefficient, and the preset complexity threshold corresponding to each training task, the task weight corresponding to each training task is determined. Based on the task weights corresponding to each training task, the weights of each training task are adaptively adjusted.
8. A model training device based on manifold alignment, characterized in that, include: The first determining unit is used to determine the training dataset; The training dataset includes multiple gastric pathology images, as well as corresponding descriptive text and annotation labels for each gastric pathology image; The feature extraction unit is used to iteratively train the pre-built visual language model based on the training dataset according to a preset number of training iterations; when entering the current iterative training cycle, the visual language model is used to extract image features from each of the gastric pathological images to obtain the image features corresponding to each of the gastric pathological images, and to extract text features from the descriptive text corresponding to each of the gastric pathological images to obtain the text features corresponding to each of the gastric pathological images. The feature enhancement unit is used to apply the visual language model to each of the gastric pathological images to enhance the image features corresponding to the gastric pathological image, thereby obtaining the enhanced image features corresponding to the gastric pathological image, and to enhance the text features corresponding to the gastric pathological image, thereby obtaining the enhanced text features corresponding to the gastric pathological image. The feature optimization unit is used to perform cross-modal feature optimization on the enhanced image features corresponding to each gastric pathological image based on its corresponding enhanced text features, so as to obtain the cross-modal optimized image features corresponding to the gastric pathological image. The feature mapping unit is used to perform manifold space mapping on the cross-modal optimized image features and enhanced text features corresponding to each gastric pathological image to obtain the mapped image features and mapped text features corresponding to the gastric pathological image. The feature fusion unit is used to perform feature fusion on the mapped image features and mapped text features corresponding to each gastric pathological image based on a preset active spatial focusing mechanism, so as to obtain the image-text fusion features corresponding to the gastric pathological image. The feature training unit is used to train the visual language model based on the image-text fusion features corresponding to each of the gastric pathology images to obtain a gastric pathology detection model. Specifically, the step of performing cross-modal feature optimization on the enhanced image features corresponding to the gastric pathological image based on its corresponding enhanced text features to obtain the cross-modal optimized image features corresponding to the gastric pathological image includes: The pathological location of the gastric pathological image is obtained by indexing the enhanced text features corresponding to the gastric pathological image using a preset relative position discriminator. In the enhanced image features corresponding to the pathological image of the stomach, the feature region corresponding to the pathological location is determined; Edge sensing processing is performed on the feature region to obtain the edge sensing features corresponding to the feature region; The feature region is subjected to texture-aware processing to obtain the texture-aware features corresponding to the feature region; The feature regions are subjected to semantic aggregation processing to obtain the semantic aggregated features corresponding to the feature regions; The edge-aware features, texture-aware features, and semantic aggregation features are fused, and the fusion result is used as the cross-modal optimized image feature corresponding to the gastric pathological image. The method, based on a preset active spatial focusing mechanism, fuses the mapped image features and mapped text features corresponding to the gastric pathological image to obtain the image-text fusion features corresponding to the gastric pathological image, including: Two-way distance measurement is performed between the mapped image features and mapped text features corresponding to the gastric pathological image to obtain the first geodesic distance and the second geodesic distance corresponding to the gastric pathological image; Based on the first geodesic distance and the second geodesic distance, the weighted attention corresponding to the gastric pathological image is determined; Based on the weighted attention corresponding to the gastric pathological image, the mapped image features and mapped text features corresponding to the gastric pathological image are fused, and the fusion result is used as the image-text fusion feature corresponding to the gastric pathological image.
Citation Information
Patent Citations
Interactive image editing method and device, readable storage medium and electronic equipment
CN113448477A
Intelligent query semantic understanding method based on information geometry and Riemannian manifold
CN119829740A