Model training method and device based on manifold alignment

Through the manifold alignment model training method, the visual language model is used to extract, enhance and fuse gastric pathology images and text features, which solves the problem of insufficient cross-modal information integration in the existing model and improves the detection accuracy.

CN120805067AActive Publication Date: 2025-10-17XIANGJIANG LAB
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511254799.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2025-10-17
Estimated Expiration
2045-09-04

AI Technical Summary

Technical Problem

During the training process, existing gastric pathology detection models are unable to effectively capture the complex nonlinear relationship between images and text in their respective latent manifold spaces, resulting in limited cross-modal information integration and poor detection accuracy.

Method used

A model training method based on manifold alignment is adopted to extract, enhance and optimize the features of gastric pathology images and text through a visual language model, map them to the manifold space, and use the active spatial focusing mechanism to perform feature fusion to improve the cross-modal information integration effect.

Benefits of technology

The detection accuracy of the gastric pathology detection model is improved, the training effect of the model is improved, and the ability to integrate cross-modal information of image and text features is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120805067A_ABST
    Figure CN120805067A_ABST
Patent Text Reader

Abstract

The invention provides a manifold alignment-based model training method and device, and belongs to the field of machine learning, and the method comprises the steps: carrying out the image feature and text feature extraction of a stomach pathological image through a visual language model in a training process, and obtaining the image features and text features; performing feature enhancement on the features to obtain enhanced image features and enhanced text features; performing cross-modal feature optimization on the enhanced image features based on the enhanced text features to obtain cross-modal optimized image features; manifold space mapping is carried out on the cross-modal optimization image features and the enhanced text features, and mapping image features and mapping text features are obtained; and based on an active space focusing mechanism, carrying out feature fusion on the mapping image features and the mapping text features to obtain image-text fusion features, and carrying out model training on the image-text fusion features. By applying the method, cross-modal fusion can be performed on the image features and the text features, the model training effect can be improved, and the model detection precision can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of machine learning, in particular to a model training method and device based on manifold alignment. BACKGROUND

[0002] With the development of machine learning technology, in the medical field, machine learning technology is also gradually applied to assist in diagnosis. One of the common applications is to train a stomach pathology detection model to detect stomach pathology images and obtain diagnosis results about pathology types, lesion locations, etc.

[0003] In the stomach pathology detection scenario, multiple modal data processing is usually involved, such as image data and text data. In the training process of the existing stomach pathology detection model, the multiple modal data is usually directly concatenated for feature, and the model training is completed based on the concatenated features.

[0004] The direct feature concatenation of multiple modal data usually cannot effectively capture the complex nonlinear relationship of images and texts in their respective latent manifold spaces. Therefore, the stomach pathology detection model trained based on the existing method has limited integration effect on cross-modal information, resulting in poor detection accuracy. SUMMARY

[0005] Therefore, the embodiments of the present application provide a model training method based on manifold alignment to solve the problem that the existing model training method has limited integration effect on cross-modal information, resulting in poor model detection accuracy.

[0006] The embodiments of the present application also provide a model training device based on manifold alignment to ensure the implementation and application of the above method in practice.

[0007] To achieve the above object, the embodiments of the present application provide the following technical solutions:

[0008] The first aspect of the embodiments of the present application provides a model training method based on manifold alignment, comprising:

[0009] determining a training data set; the training data set includes a plurality of stomach pathology images, and description texts and label tags corresponding to the stomach pathology images;

[0010] based on the training data set, iteratively training a pre-constructed visual language model according to a preset training number; when entering the current iteration training period, applying the visual language model to respectively extract image features of each stomach pathology image to obtain image features corresponding to each stomach pathology image, and respectively extract text features of description texts corresponding to each stomach pathology image to obtain text features corresponding to each stomach pathology image;

[0011] For each of the stomach pathology images, the visual language model is applied to perform feature enhancement on the image features corresponding to the stomach pathology image to obtain enhanced image features corresponding to the stomach pathology image, and to perform feature enhancement on the text features corresponding to the stomach pathology image to obtain enhanced text features corresponding to the stomach pathology image;

[0012] For each of the stomach pathology images, the enhanced image features corresponding to the stomach pathology image are optimized based on the enhanced text features corresponding to the stomach pathology image to obtain cross-modal optimized image features corresponding to the stomach pathology image;

[0013] For each of the stomach pathology images, the cross-modal optimized image features and the enhanced text features corresponding to the stomach pathology image are respectively mapped to a manifold space to obtain mapped image features and mapped text features corresponding to the stomach pathology image;

[0014] For each of the stomach pathology images, the mapped image features and the mapped text features corresponding to the stomach pathology image are fused based on a preset active space focusing mechanism to obtain image-text fusion features corresponding to the stomach pathology image;

[0015] Based on the image-text fusion features corresponding to each of the stomach pathology images, the visual language model is trained to obtain a stomach pathology detection model.

[0016] The second aspect of the embodiment of the present application provides a model training device based on manifold alignment, comprising:

[0017] A first determination unit is configured to determine a training data set; the training data set comprises a plurality of stomach pathology images, description texts corresponding to the stomach pathology images, and annotation labels corresponding to the stomach pathology images;

[0018] A feature extraction unit is configured to perform iterative training on a pre-constructed visual language model based on the training data set according to a preset training number; when entering a current iteration training period, the visual language model is applied to perform image feature extraction on each of the stomach pathology images to obtain image features corresponding to each of the stomach pathology images, and to perform text feature extraction on the description texts corresponding to each of the stomach pathology images to obtain text features corresponding to each of the stomach pathology images;

[0019] A feature enhancement unit is configured to, for each of the stomach pathology images, apply the visual language model to perform feature enhancement on the image features corresponding to the stomach pathology image to obtain enhanced image features corresponding to the stomach pathology image, and to perform feature enhancement on the text features corresponding to the stomach pathology image to obtain enhanced text features corresponding to the stomach pathology image;

[0020] a feature optimization unit, configured to, for each of the gastric pathology images, perform cross-modal feature optimization on the enhanced image feature corresponding to the gastric pathology image based on the enhanced text feature corresponding to the gastric pathology image, to obtain a cross-modal optimized image feature corresponding to the gastric pathology image;

[0021] a feature mapping unit, configured to, for each of the gastric pathology images, respectively perform manifold space mapping on the cross-modal optimized image feature and the enhanced text feature corresponding to the gastric pathology image, to obtain a mapped image feature and a mapped text feature corresponding to the gastric pathology image;

[0022] a feature fusion unit, configured to, for each of the gastric pathology images, perform feature fusion on the mapped image feature and the mapped text feature corresponding to the gastric pathology image based on a preset active space focusing mechanism, to obtain a graphic-text fusion feature corresponding to the gastric pathology image;

[0023] a feature training unit, configured to train the visual language model based on the graphic-text fusion feature corresponding to each of the gastric pathology images, to obtain a gastric pathology detection model.

[0024] Based on the manifold alignment-based model training method provided in the above embodiments of the present application, the method comprises the following steps: determining a training data set; the training data set comprises a plurality of stomach pathological images, and description texts and annotation labels corresponding to the stomach pathological images; based on the training data set, a pre-constructed visual language model is iteratively trained according to a preset training number; when entering a current iteration training cycle, the visual language model is applied to respectively extract image features of each stomach pathological image, to obtain image features corresponding to each stomach pathological image, and to respectively extract text features of description texts corresponding to each stomach pathological image, to obtain text features corresponding to each stomach pathological image; for each stomach pathological image, the visual language model is applied to enhance features of the image features corresponding to the stomach pathological image, to obtain enhanced image features corresponding to the stomach pathological image, and to enhance features of the text features corresponding to the stomach pathological image, to obtain enhanced text features corresponding to the stomach pathological image; for each stomach pathological image, based on the enhanced text features corresponding thereto, the enhanced image features corresponding to the stomach pathological image are optimized in a cross-modal manner, to obtain cross-modal optimized image features corresponding to the stomach pathological image; for each stomach pathological image, the cross-modal optimized image features corresponding thereto and the enhanced text features corresponding thereto are respectively mapped in a manifold space, to obtain mapped image features and mapped text features corresponding to the stomach pathological image; for each stomach pathological image, based on a preset active space focusing mechanism, the mapped image features corresponding thereto and the mapped text features corresponding thereto are fused, to obtain image-text fused features corresponding to the stomach pathological image; based on the image-text fused features corresponding to each stomach pathological image, the visual language model is trained, to obtain a stomach pathological detection model. By applying the method provided in the embodiments of the present application, in the training process of the stomach pathological detection model, the image features can be first optimized in a cross-modal manner by the text features, then the text features and the image features mapped to the manifold space can be further fused based on the active space focusing mechanism, which is beneficial to integrating the cross-modal information of the text features and the image features, improving the training effect of the model, and then improving the detection accuracy of the stomach pathological detection model. BRIEF DESCRIPTION OF DRAWINGS

[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are only embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor based on the provided drawings.

[0026] Figure 1 The method flow chart of the manifold alignment-based model training method provided in the embodiments of the present application;

[0027] Figure 2 Another method flowchart of the model training method based on manifold alignment provided by an embodiment of the present application;

[0028] Figure 3 Another method flowchart of the model training method based on manifold alignment provided by an embodiment of the present application;

[0029] Figure 4 A structural schematic diagram of the model training device based on manifold alignment provided by an embodiment of the present application. DETAILED DESCRIPTION

[0030] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all the other embodiments obtained by those skilled in the art without creative work fall within the protection scope of the present application.

[0031] In the present application, the term “comprising” or “containing” or any other variant thereof is intended to cover the non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or includes the elements inherent to such process, method, article or device. Without more limitations, the element defined by the sentence “including a…” does not exclude the existence of other same elements in the process, method, article or device including the element.

[0032] The embodiment of the present application provides a model training method based on manifold alignment, which can be applied to a model training platform, and the execution subject of the method can be a processor of the model training platform. The flowchart of the method is as shown in Figure 1 The method comprises the following steps.

[0033] S101: determining a training data set; the training data set comprises a plurality of stomach pathological images, and description text and a label corresponding to the stomach pathological images.

[0034] In the method provided by the embodiment of the present application, when it is necessary to train the stomach pathology detection model, a stomach pathology image dataset, i.e., stomach pathology images containing different stomach pathologies, can be acquired first. A corresponding description text can be configured for each stomach pathology image through a large model configuration or manual configuration, to describe the stomach pathology features. Meanwhile, each stomach pathology image can be labeled through a labeling tool or manual configuration, i.e., the pathology type and lesion position of the stomach pathology image are labeled, and the labeled information is used as the labeling label of the stomach pathology image, which can also be understood as a sample label in model training. The pre-acquired stomach pathology images, the pre-configured description texts corresponding to the stomach pathology images, and the labeling labels corresponding to the stomach pathology images are used as the training dataset, to train the stomach pathology model.

[0035] In the method provided by the embodiment of the present application, a visual language model can be pre-configured according to actual needs, as an initial model of the stomach pathology model, and the visual language model can be configured based on an existing visual language algorithm. In addition, the training times can be pre-configured according to actual needs.

[0036] In the method provided by the embodiment of the present application, a visual language model can be pre-configured according to actual needs, as an initial model of the stomach pathology model, and the visual language model can be configured based on an existing visual language algorithm. In addition, the training times can be pre-configured according to actual needs.

[0037] In the model training process, the pre-configured visual language model is iteratively trained based on the training dataset and according to the pre-configured training times. When each iteration training period is entered, the visual language model can be applied to detect each stomach pathology image and the description text corresponding to the stomach pathology image, to train the visual language model.

[0038] When the visual language model is applied to detect the stomach pathology image and the description text corresponding to the stomach pathology image, the image encoder in the visual language model can be used to extract image features of the stomach pathology image, and the extracted image features are used as the image features corresponding to the stomach pathology image. The text encoder in the visual language model can be used to extract text features of the description text corresponding to the stomach pathology image, and the extracted text features are used as the text features corresponding to the stomach pathology image.

[0039] S103: For each stomach pathology image, apply the visual language model to perform feature enhancement on image features corresponding to the stomach pathology image to obtain enhanced image features corresponding to the stomach pathology image, and perform feature enhancement on text features corresponding to the stomach pathology image to obtain enhanced text features corresponding to the stomach pathology image.

[0040] In the method provided by the embodiment of the application, when the visual language model is applied to detect the stomach pathology image and the description text corresponding to the stomach pathology image, the image feature extraction branch structure in the visual language model is used to perform feature enhancement processing on the stomach pathology image object, and the image features output by the image feature extraction branch structure are used as the enhanced image features corresponding to the stomach pathology image. The text feature extraction branch structure in the visual language model is used to perform feature enhancement processing on the text features corresponding to the stomach pathology image, and the text features output by the text feature extraction branch structure are used as the enhanced text features corresponding to the stomach pathology image.

[0041] S104: For each stomach pathology image, based on the enhanced text features corresponding to the stomach pathology image, cross-modal feature optimization is performed on the enhanced image features corresponding to the stomach pathology image to obtain cross-modal optimized image features corresponding to the stomach pathology image.

[0042] In the method provided by the embodiment of the application, when the visual language model is applied to detect the stomach pathology image and the description text corresponding to the stomach pathology image, the enhanced text features corresponding to the stomach pathology image can be used to perform cross-modal feature optimization on the enhanced image features corresponding to the stomach pathology image, and the optimized enhanced image features are used as the cross-modal optimized image features corresponding to the stomach pathology image, so that the image features are strengthened based on the text features, and the text features and the image features are fused.

[0043] S105: For each stomach pathology image, manifold space mapping is performed on the cross-modal optimized image features and the enhanced text features corresponding to the stomach pathology image respectively to obtain mapped image features and mapped text features corresponding to the stomach pathology image.

[0044] In the method provided by the embodiment of the application, when the visual language model is applied to detect the stomach pathology image and the description text corresponding to the stomach pathology image, the cross-modal optimized image features corresponding to the stomach pathology image can be subjected to manifold space mapping, and the cross-modal optimized image features are mapped to the Riemann manifold space, and the cross-modal optimized image features mapped to the Riemann manifold space are used as the mapped image features corresponding to the stomach pathology image. Meanwhile, the enhanced text features corresponding to the stomach pathology image can be subjected to manifold space mapping, and the enhanced text features are mapped to the Riemann manifold space, and the enhanced text features mapped to the Riemann manifold space are used as the mapped text features corresponding to the stomach pathology image.

[0045] S106: For each stomach pathology image, based on the preset active space focusing mechanism, the mapping image features and the mapping text features corresponding to the stomach pathology image are fused to obtain image-text fusion features corresponding to the stomach pathology image.

[0046] In the method provided by the embodiment of the application, when the visual language model is applied to detect the stomach pathology image and the description text corresponding to the stomach pathology image, the mapping image features corresponding to the stomach pathology image and the mapping text features corresponding to the stomach pathology image are aligned and fused in depth based on the preset active space focusing mechanism, and the fused features are taken as image fusion features corresponding to the stomach pathology image.

[0047] S107: Based on the image-text fusion features corresponding to each stomach pathology image, the visual language model is trained to obtain a stomach pathology detection model.

[0048] In the method provided by the embodiment of the application, the visual language model is used to make a pathological diagnosis based on the image-text fusion features corresponding to each stomach pathology image, the visual language model is trained based on the diagnosis result of the visual language model and the annotation label corresponding to each stomach pathology image, and then the next training iteration cycle is entered until the training effect of the model reaches the expectation or the number of training iterations reaches the preset number of training, thereby completing the training process of the visual language model, and the trained visual language model is taken as the stomach pathology detection model.

[0049] In the training process of the stomach pathology detection model, the method provided by the embodiment of the application can first optimize the cross-modal features of the image features based on the text features, and then further fuse the text features and the image features based on the active space focusing mechanism, which is beneficial to integrating the cross-modal information of the text features and the image features, improving the training effect of the model, and then improving the detection accuracy of the stomach pathology detection model.

[0050] In the method provided by the embodiment of the application, the process of enhancing the image features corresponding to the stomach pathology image to obtain enhanced image features corresponding to the stomach pathology image in step S103 includes: Figure 1 Figure 2 In the method provided by the embodiment of the application, the process of enhancing the image features corresponding to the stomach pathology image to obtain enhanced image features corresponding to the stomach pathology image in step S103 includes:

[0051] S201: Determine whether the current iteration training cycle meets the preset dynamic adjustment condition.

[0052] ​In the method provided by the embodiment of the present application, the image region to which the model needs to pay attention can be dynamically adjusted according to the model iteration process, so that the model focuses on detecting difficult image regions. The dynamic adjustment condition can be set in advance according to actual needs, and the dynamic adjustment condition is used to judge whether the image feature enhancement mechanism is adjusted in the current iteration training period.

[0053] In S202, if the current iteration training period meets the dynamic adjustment condition, the pixel screening number corresponding to the current iteration training period is determined.

[0054] In the method provided by the embodiment of the present application, if the current iteration training period meets the dynamic adjustment condition, the corresponding pixel screening number can be calculated based on the current training progress and the training number corresponding to the current iteration training period. The pixel screening number can be calculated based on a proportional function of the total image pixel value. The value of the proportional function is dynamically changed. As the training progresses, the proportion calculated in different iteration training periods may be different. Generally, the pixel screening number is small in the early training stage, so that the model pays attention to a small number of representative pixels. In the later training stage, the pixel screening number is large, so that the model pays attention to more difficult samples.

[0055] Specifically, for example, let k represent the pixel screening number, let N represent the total number of image pixels, and let γ(t) represent the screening proportional function. The calculation method of γ(t) can be as follows:

[0056] (Formula 1).

[0057] Wherein, t represents the period number of the current iteration training period, that is, the current training epoch, T max represents the preset training number, that is, the total training epoch. γ start represents the pixel proportion reserved in the early training stage. It can be set according to actual needs. Generally, it is set to a low value, for example, 0.2, to avoid being sensitive to noise too early. γ end represents the pixel proportion reserved in the later training stage. It can be set according to actual needs. Generally, it is set to a high value, for example, 0.5, to accelerate the learning of the model for difficult regions.

[0058] The calculation method of the pixel screening number k can be as follows:

[0059] k=[γ(t)·N] (Formula 2).

[0060] In the method provided by the embodiment of the present application, if the current iteration training period does not meet the dynamic adjustment condition, the image features of the gastric pathological image can be globally enhanced, and the image features subjected to the global enhancement are used as the enhanced image features corresponding to the gastric pathological image.

[0061] S203: Determine the comprehensive score corresponding to each pixel point in the stomach pathological image, and determine the difficult region feature corresponding to the stomach pathological image in the image feature based on the preset loss function, the pixel screening quantity and the comprehensive score corresponding to each pixel point in the stomach pathological image.

[0062] In the method provided by the embodiment of the application, in the process of image coding of the stomach pathological image by the visual language model to obtain the image feature, a corresponding score is also obtained for each pixel point in the stomach pathological image, and the score corresponding to each pixel point is taken as the comprehensive score corresponding to the pixel point in the embodiment of the application. The calculation method of the comprehensive score corresponding to each pixel point can be as follows:

[0063] (Formula 3).

[0064] wherein j represents a certain pixel point in the image, s j represents the comprehensive score corresponding to the pixel point j in the image. represents the normalized pixel-level loss value of the pixel point j, and is used to measure the difficulty degree of the model in predicting the pixel. represents the boundary perception intensity of the pixel point j, and represents whether the pixel is close to the edge of the lesion. represents the prediction confidence of the pixel point j, and is used to measure the confidence degree of the model for the pixel. and is the evaluation value of the confidence or other evaluation standard of the model for the pixel, which can be calculated based on the prediction of the model for the pixel, and is usually calculated based on the difference between the prediction result of the model at the current stage and the real label. The three weight hyperparameters α, β and δ represent the comprehensive importance of the above three factors, and the three weight hyperparameters satisfy the following relationship: α+β+δ=1. The comprehensive score comprehensively considers the difficult sample, the boundary pixel and the uncertain region, and the pixel with a high score is considered to be more critical for the model training.

[0065] In the embodiment of the application, a loss function for difficult sample feature extraction is preset according to actual needs, and the loss function can be referred to as an LSL loss function.

[0066] In the method provided by the embodiment of the application, the corresponding difficult region feature can be located in the image feature based on the comprehensive score of each pixel point in the stomach pathological image, the LSL loss function and the pixel screening quantity, that is, the model detects the relatively difficult region. Let T represent the pixel set of the difficult region feature, and the calculation method of the pixel set can be as follows:

[0067] (Formula 4).

[0068] ​Wherein, Select represents a mechanism for constructing the pixel set T, which is defined as selecting the top k pixels with the highest scores from all pixels in the image, and the meanings of other parameters can be referred to the foregoing description. The pixel set T is selected from the image according to the comprehensive score s j The top k pixels are screened out in descending order. These pixels are usually the most representative or most difficult regions, which may involve small lesions, edge regions or other difficult-to-distinguish regions.

[0069] In the method provided by the embodiment of the present application, the loss value for feature extraction of the difficult region can be calculated based on the following manner:

[0070] (Formula 5).

[0071] Wherein, Loss represents a loss function, p j represents the predicted value of the model for the pixel j in the feature extraction stage, y j represents the true label of the pixel j.

[0072] S204: Perform feature enhancement processing on the difficult region features corresponding to the gastric pathology image, and take the processed features as the enhanced image features corresponding to the gastric pathology image, so as to focus the training of the visual language model on the difficult region.

[0073] In the method provided by the embodiment of the present application, the existing image feature enhancement manner can be used to perform feature enhancement processing on the difficult region features corresponding to the gastric pathology image, and the image features obtained after the processing are taken as the enhanced image features corresponding to the gastric pathology image.

[0074] Based on the method provided by the embodiment of the present application, the model can be focused on the global structure in the early training stage to prevent noise interference, and gradually shift the attention to more difficult regions in the later training stage, especially the places with fuzzy boundaries and small lesions. This strategy can help the model better adapt to the needs of different training stages, thereby improving the learning effect of difficult samples.

[0075] In Figure 2 Based on the method shown in Figure 3 In the method provided by the embodiment of the present application, the process of determining whether the current iteration training period meets the preset dynamic adjustment condition in step S201 includes:

[0076] S301: Determine the training loss value and the average precision index of the validation set corresponding to the current iteration training period.

[0077] In the method provided by the embodiment of the present application, a future trend perception module can be defined, through which the next round of index trend of model training is predicted, and then whether the current iteration training period meets the dynamic adjustment mechanism is evaluated. First, the future trend perception module can calculate the training loss value of the current iteration training period and the mean intersection over union (mIoU) index of the validation set by using the predicted value of the model for each pixel in the stomach pathological image and the labeled label (i.e., the real label) in the current iteration training period.

[0078] S302: Based on the training loss value and the mIoU index of the validation set, the index trend is predicted to obtain the training loss prediction value and the mIoU prediction value of the validation set corresponding to the current iteration training period.

[0079] In the method provided by the embodiment of the present application, a sliding window model or a lightweight sequence prediction model can be applied to predict the training loss value and the mIoU index of the next iteration training period, to obtain the prediction value of the training loss value and the prediction value of the mIoU index of the next iteration training period. The prediction value of the training loss value of the next period is taken as the training loss prediction value of the current iteration training period, and the prediction value of the mIoU index of the validation set of the next period is taken as the mIoU prediction value of the validation set of the current iteration training period.

[0080] S303: Based on the training loss prediction value and the mIoU prediction value of the validation set, whether the current iteration training period meets the preset first trigger condition is judged to obtain a first trigger condition judgment result.

[0081] In the method provided by the embodiment of the present application, the trigger condition of the training loss and the mIoU of the validation set can be set according to actual needs, that is, when the development trend of the training loss value and the mIoU index of the validation set meets certain conditions, it is considered that the mechanism of dynamic adjustment needs to be triggered. For example, based on the variation amplitude of the training loss value and the mIoU index of the validation set, the first trigger condition can be set, that is, when the variation amplitude of the training loss value and the mIoU index of the validation set meets certain conditions, it is considered that the first trigger condition is met. Specifically, TraninLoss t represents the training loss value of the current iteration training period, mIoU t represents the mIoU index of the validation set of the current iteration training period. represents the training loss prediction value, represents the mIoU prediction value of the validation set. The training loss difference value is obtained by performing difference operation on the training loss prediction value and the training loss value, and the mIoU difference value of the validation set is obtained by performing difference operation on the mIoU prediction value of the validation set and the mIoU index of the validation set. The specific calculation method can be as follows:

[0082] (Formula 6).

[0083] (Formula 7).

[0084] wherein, represents a training loss difference value, represents a validation set mean intersection over union difference value.

[0085] In the embodiment of the present application, the first trigger condition is that the training loss difference value is greater than or equal to zero and the validation set mean intersection over union difference value is less than zero. That is, if the training loss difference value and the validation set mean intersection over union difference value satisfy the following formula, it is considered that the current iteration training period satisfies the preset first trigger condition.

[0086] (Formula 8).

[0087] In the method provided by the embodiment of the present application, the corresponding parameter can be set to represent the first trigger condition judgment result, and the first trigger condition judgment result indicates that the current iteration training period satisfies the first trigger condition or does not satisfy the first trigger condition. For example, the parameter S FATE represents the first trigger condition judgment result. If the current iteration training period satisfies the first trigger condition, S FATE = 1 is set, and if it does not satisfy the first trigger condition, S FATE = 0 is set.

[0088] Further, the selection of the pixel screening ratio in the LSL loss function can also be selected in the following manner:

[0089] (Formula 9).

[0090] wherein, Loγ(t) represents the pixel screening ratio, γ low and γ high can be set according to actual needs. In the embodiment of the present application, γ low = 0.2 and γ high = 0.5.

[0091] S304: Determine the curvature anomaly ratio corresponding to the current iteration training period.

[0092] In the method provided by the embodiment of the present application, a training anomaly detection module can be defined to detect whether the model training is abnormal, and then evaluate whether the current iteration training period satisfies the dynamic adjustment mechanism. The training anomaly detection module can evaluate the curvature anomaly index of the current iteration training period through the gradient and derivative of the current training loss, and obtain the curvature anomaly ratio. Specifically, the calculation method of the curvature anomaly ratio A t is as follows:

[0093] (Formula 10).

[0094] wherein, denotes the first order gradient of the total loss of the current iteration training period, denotes the second order derivative of the total loss of the current iteration training period.

[0095] S305: determining whether the current iteration training period meets the preset second trigger condition based on the curvature anomaly ratio, to obtain a second trigger condition determination result.

[0096] In the method provided by the embodiments of the present application, the trigger condition for the curvature anomaly ratio can be set according to actual needs, that is, when the curvature anomaly ratio meets certain conditions, it is considered that the mechanism of dynamic adjustment needs to be triggered. For example, the second trigger condition is set based on the size relationship between the current curvature anomaly ratio and the historical curvature anomaly ratio. Whether the current iteration training period meets the second trigger condition can be determined by the following condition:

[0097] (Formula 11).

[0098] wherein, δ denotes the preset curvature anomaly detection threshold, EMA(A) denotes the exponential moving average of the historical curvature, and the historical curvature refers to the curvature anomaly ratio corresponding to the iteration training period before the current period.

[0099] In the method provided by the embodiments of the present application, the corresponding parameter can be set to represent the second trigger condition determination result. The second trigger condition determination result indicates that the current iteration training period meets the second trigger condition or does not meet the second trigger condition. For example, the parameter S ASL denotes the second trigger condition determination result. If the current iteration training period meets the second trigger condition, S ASL =1 is set, and if it does not meet the second trigger condition, S ASL =0 is set.

[0100] S306: determining the average loss conflict degree corresponding to the current iteration training period.

[0101] In the method provided by the embodiments of the present application, a multi-task conflict detection module (MOCL) can be defined. Through the module, whether there is an optimization direction conflict between the multi-loss tasks in the model training is monitored, and then whether the current iteration training period meets the dynamic adjustment mechanism is evaluated. Specifically, the calculation method of the average loss conflict degree can be as follows:

[0102] (Formula 12).

[0103] wherein, Conf trepresents the average loss conflict degree, M represents the number of loss functions currently participating in optimization, that is, each loss function involved in the current task. represents the gradient of the i-th loss function, represents the cosine similarity of the gradient direction.

[0104] S307: Based on the average loss conflict degree, it is judged whether the current iteration training period meets the preset third trigger condition, and a third trigger condition judgment result is obtained.

[0105] In the method provided by the embodiments of the application, the trigger condition for the average loss conflict degree can be set according to actual needs, that is, when the average loss conflict degree meets certain conditions, it is considered that the mechanism of dynamic adjustment needs to be triggered. For example, based on the size relationship between the average loss conflict degree and the preset threshold, the third trigger condition is set, such as taking the average loss conflict degree greater than the preset threshold as the third trigger condition. Whether the current iteration training period meets the third trigger condition can be judged by the condition shown in the following formula:

[0106] Conf t > 0.8 (formula 13).

[0107] In the method provided by the embodiments of the application, the corresponding parameter can be set to represent the third trigger condition judgment result, and the third trigger condition judgment result represents that the current iteration training period meets the third trigger condition or does not meet the third trigger condition. For example, the conflict trigger flag S MOCL represents the third trigger condition judgment result, if the current iteration training period meets the third trigger condition, S MOCL = 1 is set, and if the third trigger condition is not met, S MOCL = 0 is set.

[0108] S308: Based on the first trigger condition judgment result, the second trigger condition judgment result and the third trigger condition judgment result, the dynamic adjustment activation probability corresponding to the current iteration training period is determined.

[0109] In the method provided by the embodiments of the application, the activation probability fusion module can be defined, which evaluates the probability of the dynamic adjustment mechanism based on the first trigger condition judgment result, the second trigger condition judgment result and the third trigger condition judgment result, and obtains the dynamic adjustment activation probability corresponding to the current iteration training period. Specifically, the calculation method of the dynamic adjustment activation probability can be as shown in the following formula:

[0110] P t = σ (γ1·S FATE + γ2·S ASL + γ3·S MOCL ) (formula 14).

[0111] Wherein, P tis the dynamically adjusted activation probability corresponding to the current iterative training cycle, and σ(·) is the Sigmoid function used to smoothly map the data to the (0, 1) interval. γ1, γ2, and γ3 represent the weights corresponding to the respective judgment results, satisfying γ1+γ2+γ3=1. In this embodiment of the present invention, the following values ​​are used: γ1=0.5, γ2=0.3, and γ3=0.2.

[0112] S309: Determine whether the dynamically adjusted activation probability is greater than a preset probability threshold.

[0113] In the method provided by the embodiment of the present invention, a probability threshold may be set in advance according to actual needs, and the dynamically adjusted activation probability is compared with the preset probability threshold. For example, the probability threshold may be set to 0.5.

[0114] S310: If the dynamically adjusted activation probability is greater than the probability threshold, it is determined that the current iterative training cycle meets the dynamic adjustment condition.

[0115] In the method provided by the embodiment of the present invention, if the dynamically adjusted activation probability is greater than the preset probability threshold, the current iterative training cycle is determined to meet the dynamic adjustment conditions. If the dynamically adjusted activation probability is less than the preset probability threshold, the current iterative training cycle is determined to not meet the dynamic adjustment conditions. When the dynamically adjusted activation probability is equal to the preset probability threshold, the current iterative training cycle can be configured to meet the dynamic adjustment conditions. Specifically, the loss function selection during the training process can be determined based on the following formula:

[0116] (Equation 15).

[0117] Among them, Loss t Represents the loss function applied in the current iterative training cycle, and P(t) is the dynamically adjusted activation probability P t , CE represents the CE (cross entropy) loss function, and Dice represents the Dice loss function. LSL γ(t) represents the LSL loss function.

[0118] exist Figure 1 On the basis of the method shown in FIG. 1 , in the method provided in an embodiment of the present invention, the process of enhancing the text features corresponding to the gastric pathology image in step S103 to obtain the enhanced text features corresponding to the gastric pathology image includes:

[0119] The text features corresponding to the gastric pathology image are embedded in word vectors to obtain a word vector sequence corresponding to the text features.

[0120] In the method provided by the embodiment of the application, text features are extracted through a hierarchical discriminator based on semantic word features and a hierarchical attention mechanism, and text enhancement is performed according to task requirements. First, for each word in the text features corresponding to the stomach pathological image, the word embedding layer is used to map to the word vector space, so as to obtain the word vector corresponding to each word in the text features, and the word vector sequence corresponding to the text features is composed of the word vectors corresponding to each word. For example, the text features corresponding to the stomach pathological image can be represented as: X={x1,x2,x3,…,x n}, wherein x i represents a word. The word vector sequence obtained by performing word vector embedding processing on the text features can be represented as: E={e1,e2,…,e n}, e i ∈R d . Wherein, e i =f embed (x i ), represents the word vector corresponding to the word x i , and d represents the dimension of the word vector.

[0121] The word vector sequence is layered to obtain the word vectors corresponding to each level.

[0122] In the method provided by the embodiment of the application, the semantic units in the description text corresponding to the stomach pathological image are modeled through a preset hierarchical discriminator, that is, the text is layered according to the semantic property features of the language. In the embodiment of the application, three levels are divided, the first level is local position information, the second level is core entity, and the third level is main feature description. For example, for the description text “there is an obvious dark protrusion in the upper left corner”, the local position information “upper left corner” can be divided into the first level semantics, “a piece” can be divided into the second level semantics, and “dark protrusion” can be divided into the third level semantics. In the specific data processing process, the word vector sequence is used for semantic layering to obtain the word vectors of each level. The core word function f core of the hierarchical modeling can be defined in advance to extract the key semantic levels. For the semantics of each level, the core word function f core is used for semantic extraction, which can be specifically represented as:

[0123] (Formula 16).

[0124] Wherein, H k represents the semantics of the kth level, and w k represents the word vector in the kth level. For example, for the text “there is an obvious dark protrusion in the upper left corner”, the hierarchical order is: H1=f core (upper left corner), H2=f core (a piece), and H3=f core(dark color protrusions).

[0125] The attention weight corresponding to each word vector corresponding to each level is determined.

[0126] In the method provided by the embodiment of the application, a hierarchical attention mechanism is introduced, and a key feature is given a higher weight in a specific level to perform feature extraction. Therefore, for each word vector in each level, the attention weight of the word vector in the level is calculated according to a pre-set attention weight calculation manner. In the process of calculating the attention weight, not only the local representation of the word is relied on, but also the relationship with other words and the influence of the context need to be considered. The calculation manner of the attention weight can be as follows:

[0127] (Formula 17).

[0128] wherein, represents the attention weight of the i th word vector in the k th layer, which measures the relative importance of the word to other words in the current level. represents the weighted representation of the i th word in the k th layer, which combines the word vector and the context information , and the parameter meanings of other subscripts and superscripts carried by z are similar to those of , which will not be described here. represents the learning weight of the k th layer, which is used to adjust the weighted contribution of the word vector and the context representation, and determines the importance of the k th layer feature in calculating the attention. k represents the dynamic context awareness weight of the k th layer, which is used to measure the influence of the context information on the importance of the word in the layer, and determines the weight distribution of the context in different contexts. j and W m are respectively the context awareness weights of the corresponding levels. represents the context information of the i th word, which can be other words, disease types, pathological descriptions and the like related to the semantics thereof. represents the interlayer interaction weight, represents the mutual influence between different layers, for example represents the mutual relationship between the k th layer and the j th layer, which can be the semantic connection between different levels. represents the relationship between the k th layer and other layers (such as the previous layer or the next layer), that is, the mutual influence between a level and its adjacent layers. The calculation manner of W

[0129] (Formula 18).

[0130] wherein, represents the word vector of the i th word in the k th layer, Context information representing the i-th word, for example, it can include the context of the surrounding vocabulary, sentence structure or entire paragraph. f(·) represents a function for fusing word vectors and context information, common operations are weighted summation, concatenation or neural network, etc.

[0131] Based on the preset bidirectional long short-term memory network and the attention weight corresponding to each word vector, the context feature fusion is performed on each word vector, and the fusion result is taken as the enhanced text feature corresponding to the gastric pathological image.

[0132] In the method provided by the embodiment of the application, the word vector can be encoded based on a hierarchical bidirectional long short-term memory network (BiLSTM) to extract and fuse the text feature, and the attention mechanism is introduced in the feature extraction process to extract and fuse each word vector according to the corresponding attention weight, and the fusion result is taken as the enhanced text feature corresponding to the gastric pathological image. In the process of applying BiLSTM to feature extraction and fusion, the input of each layer comes from the hidden state of the previous layer, and the expression of the process can be as follows:

[0133] (Formula 19).

[0134] (Formula 20).

[0135] (Formula 21).

[0136] wherein, represents the hidden state of the k-th layer.

[0137] In the method shown in the method, Figure 1 On the basis of the method shown in the method, in the method provided by the embodiment of the application, the process of optimizing the enhanced image feature corresponding to the gastric pathological image based on the enhanced text feature corresponding thereto in step S104 to obtain the cross-modal optimized image feature corresponding to the gastric pathological image includes:

[0138] The relative position discriminator is used to index the position of the enhanced text feature corresponding to the gastric pathological image to obtain the pathological position corresponding to the gastric pathological image.

[0139] In the method provided by the embodiment of the present application, a relative position discriminator is defined in advance, which can recognize the semantic of the position in the text through the basic position semantic feature of the text, so as to locate the pathological position in the stomach pathological image. Specifically, the relative position discriminator can recognize the position information in the enhanced text feature corresponding to the stomach pathological image, output the region index corresponding to the position, and take the position corresponding to the region index as the pathological position corresponding to the stomach pathological image. For example, assuming that the stomach pathological image is R, the center point coordinates of which are C=(x c ,y c ), the image is divided into four regions R1, R2, R3 and R4:

[0140] (formula 22).

[0141] wherein R1 represents the feature of the upper left region, R2 represents the feature of the upper right region, R3 represents the feature of the lower left region, and R4 represents the feature of the lower right region. The relative position discriminator f pos acts on the position information T pos in the text feature, and outputs the region index r corresponding to the position:

[0142] (formula 23).

[0143] For example, for the description “there is an obvious dark color protrusion in the upper left corner”, the index corresponding to “the upper left corner” is recognized, and the “upper left corner” is the pathological position corresponding to the stomach pathological image.

[0144] In the enhanced image feature corresponding to the stomach pathological image, the feature region corresponding to the pathological position is determined.

[0145] In the method provided by the embodiment of the present application, according to the pathological position corresponding to the stomach pathological image, the image feature of the corresponding region in the enhanced image feature is located, and the feature region matched with the pathological position in the enhanced image feature is taken as the feature region corresponding to the pathological position.

[0146] The edge perception processing is performed on the feature region, and the edge perception feature corresponding to the feature region is obtained.

[0147] In the method provided by the embodiment of the application, a gravity center extraction module (GCE module) can be pre-configured, and a feature region corresponding to a pathological position is subjected to one round of depth image feature extraction and fusion through the gravity center extraction module. Three parallel processing paths are configured in the gravity center extraction module, and are respectively focused on edge details, texture response and high-level semantic information. Each path is finally subjected to joint modeling in a fusion stage to improve the discriminability and robustness of overall feature expression. The three processing paths are an edge perception path (Edge Path), a texture perception path (Texture Path) and a semantic aggregation path (Semantic Path).

[0148] In the method provided by the embodiment of the application, in the edge perception path, edge perception is performed on the feature region corresponding to the pathological position, and edge-sensitive features are extracted, and the extracted edge-sensitive features are taken as edge perception features corresponding to the feature region. Specifically, first, a channel compression layer is used to process the feature region corresponding to the pathological position, the channel compression layer uses a 1x1 convolution kernel and a stride of 1, maps the channel number C of the input feature to C / 4, and the expression of the channel compression process can be as follows:

[0149] (Formula 24).

[0150] wherein, represents an original input feature map, , that is, the size is H in height, W in width and C in channel number. represents a convolution kernel weight, , the corresponding operation is 3x3 convolution on the feature with the input channel number C, and the output channel is C' (such as C / 4). represents a convolution bias term, , and the bias is added according to the output channel. represents a standard convolution operation. represents a nonlinear activation function, which retains positive numbers and suppresses negative numbers. represents an output feature map after convolution and activation, .

[0151] For the feature subjected to channel compression, an untrainable Sobel edge filter (kernel size 3x3) is used for edge filtering processing to enhance local edge contour response, and the output channel number remains unchanged, which is C / 4. For the feature subjected to edge filtering, a depthwise convolution is used for spatial down-sampling, and the feature obtained after processing is taken as the edge perception feature. Specifically, the depthwise convolution uses a 3x3 convolution kernel and a stride of 2, and the output size is H / 2xW / 2, and this process can be formalized as:

[0152] (Formula 25).

[0153] wherein, represents an input feature map, , the spatial size is HxW, and the number of channels C' = C / 4. represents a depthwise convolution (DepthwiseConv), the kernel size is 3x3, the stride is 2, and spatial down-sampling is realized. , that is, the spatial size of the output feature map is reduced by half, and the number of channels remains unchanged.

[0154] The edge-aware path can extract edge-sensitive features with extremely low parameter overhead, and is suitable for boundary modeling under high-resolution input.

[0155] The feature region is subjected to texture-aware processing to obtain texture-aware features corresponding to the feature region.

[0156] In the method provided by the embodiment of the application, in the texture-aware path, the feature region corresponding to the pathological position is subjected to texture feature enhancement, and the enhanced features are used as texture-aware features. Specifically, first, a 3x3 standard convolution is used to reduce the number of channels of the input features from C to C / 4, while maintaining the original spatial resolution. Then, a Gabor convolution kernel (5x5) with direction selectivity is used for feature enhancement to emphasize texture responses in different directions and frequencies. This process can be formally represented as:

[0157] (Formula 26).

[0158] Then, a 3x3 depthwise convolution (with a stride of 2) is used to spatially compress the feature-enhanced features to obtain output features (i.e., texture-aware features) with a size of H / 2xW / 2, so as to capture local texture differences.

[0159] The feature region is subjected to semantic aggregation processing to obtain semantic aggregation features corresponding to the feature region.

[0160] In the method provided by the embodiment of the application, in the semantic aggregation path, the feature region corresponding to the pathological position is subjected to semantic aggregation processing to fuse the semantics in the features and obtain semantic aggregation features. Specifically, the features of the feature region are used as input features, a 3x3 standard convolution is used to compress the number of channels of the input features from C to C / 2, while maintaining the spatial size, and then an SE-Block (Squeeze-and-Excitation) module is introduced to re-label the channel dimension with the help of global average pooling (GAP) and two fully connected layers. This module can be formally represented as:

[0161] (Formula 27).

[0162] where SE(X) represents the output of the SE-Block module, X represents the features input into the SE-Block module, represents a ReLU activation function, is a Sigmoid activation function, and are full connection layer weights, and GAP() represents a global average pooling operation.

[0163] The channel attention weights of the output of the SE-Block module are multiplied with the original feature map channel by channel to realize key channel enhancement. Finally, a 3x3 convolution (with a stride of 2) is performed for downsampling, and the output size is H / 2xW / 2 features (i.e. semantic aggregation features), and the number of channels is still C / 2.

[0164] The edge perception features, texture perception features, and semantic aggregation features are fused, and the fusion result is taken as the cross-modal optimized image feature corresponding to the gastric pathological image.

[0165] In the method provided by the embodiment of the application, after the edge perception features, texture perception features, and semantic aggregation features corresponding to the feature region are obtained, the three types of features can be fused and spliced in the channel dimension by a preset fusion module (Fusion Module), and the spliced features (i.e. the fusion result) are taken as the cross-modal optimized image feature corresponding to the gastric pathological image. Specifically, first, the edge perception features, texture perception features, and semantic aggregation features are spliced in the channel dimension to obtain fusion features:

[0166] (Formula 28).

[0167] wherein, represents the output features from the edge path, . represents the output features from the texture path, . represents the output features from the semantic path, . represents a splicing operation in the channel dimension, represents the spliced fusion features, , and the number of channels is restored to the original number of channels C.

[0168] Then, a gate attention mechanism (Gate Attention) with a bottleneck structure is used to learn the channel importance weights of different paths, and the process is realized by two full connection layers. The process can be formally represented as:

[0169] (Formula 29).

[0170] wherein, represents an input feature map, which is generally the result of multi-path feature splicing, . represents a global average pooling operation on F, that is, an average is taken for each channel to obtain a channel global description, . , represents the first fully connected layer weight, used for channel compression, , represents the second fully connected layer weight, used for restoring the channel dimension, and r is an intermediate compression dimension, which is usually C / 8 or 16. represents a ReLU activation function, represents a Sigmoid activation function, used to generate a weight coefficient (0~1) for each channel. , represents the attention weight of each channel. The final effect of the attention weight is that the channel attention broadcasts and multiplies the weight of F, that is, F'=F.Gate(F).

[0171] Finally, a 1x1 convolution is performed for channel compression to obtain an output feature with a channel number of C out , which is taken as the fusion result of the edge perception feature, the texture perception feature and the semantic aggregation feature (that is, the cross-modal optimized image feature), and the cross-modal optimized image feature F out can be formally represented as:

[0172] (Formula 30).

[0173] wherein, Conv represents a convolution operation, and F gate represents the fusion feature output after processing the edge perception feature, the texture perception feature and the semantic aggregation feature based on the gated attention mechanism.

[0174] On the basis of the method shown in Figure 1 , the method provided in the embodiment of the present application comprises the following steps:

[0175] The cross-modal optimized image feature corresponding to the gastric pathological image is mapped in a manifold space by using a preset symmetric positive definite matrix, and the processing result is taken as the mapped image feature corresponding to the gastric pathological image.

[0176] In the method provided by the embodiment of the application, the optimized image features and text features are respectively mapped to the Riemann manifold space to obtain manifold embedding features, the cross-modal semantic consistency is further strengthened by adopting bidirectional geodesic contrast learning loss to optimize the image-text feature relationship, introducing a curvature perception manifold attention mechanism, and introducing a pseudo-negative sample mechanism to generate pseudo-negative samples and manifold dynamic weighted multi-task planning.

[0177] Specifically, in the feature mapping process, the cross-modal optimized features are mapped to the Riemann manifold space through a symmetric positive definite matrix (SPD), and the image features mapped to the Riemann manifold space are taken as the mapping image features corresponding to the stomach pathological image. The process of manifold space mapping processing of the cross-modal optimized image features can be formally represented as:

[0178] (Formula 31).

[0179] wherein F img represents the mapping image features, I represents the cross-modal optimized image features, represents a feature extraction mapping operation, represents a symmetrization operation, represents a matrix exponential mapping operation.

[0180] The enhanced text features corresponding to the stomach pathological image are processed in the manifold space through a preset covariance matrix, and the processing result is taken as the mapping text features corresponding to the stomach pathological image.

[0181] In the method provided by the embodiment of the application, the enhanced text features corresponding to the stomach pathological image are processed in the manifold space through a covariance matrix, and the text features mapped to the Riemann manifold space are taken as the mapping text features corresponding to the stomach pathological image. The process of manifold space mapping processing of the enhanced text features can be formally represented as:

[0182] (Formula 32).

[0183] wherein F text represents the mapping text features, N represents the number of word vectors in the enhanced text features, t i represents the vector representation (i.e., the word vector) of the i th word, represents the mean of all word vectors.

[0184] Based on the method shown in Figure 1 , in the method provided by the embodiment of the application, the process of feature fusion of the mapping image features and the mapping text features corresponding to the stomach pathological image to obtain the image-text fusion features corresponding to the stomach pathological image based on the preset active space focusing mechanism in step S106 includes:

[0185] The bidirectional ranging between the mapping image features and the mapping text features corresponding to the stomach pathological image is performed to obtain the first geodesic distance and the second geodesic distance corresponding to the stomach pathological image.

[0186] In the method provided by the embodiment of the application, the geodesic distance between the image features and the text features is defined in the SPD manifold space, and for example, the calculation manner of the geodesic distance from the mapping image features to the mapping text features can be as follows:

[0187] (Formula 33).

[0188] d geodesic (F img , F text ) represents the geodesic distance from the mapping image features to the mapping text features, and the calculation manner of the geodesic distance from the mapping text features to the mapping image features is similar to that. In the embodiment of the application, the bidirectional ranging between the mapping image features and the mapping text features is performed, that is, the geodesic distance from the mapping image features to the mapping text features is calculated, the geodesic distance is taken as the first geodesic distance, and the geodesic distance from the mapping text features to the mapping image features is calculated, and the geodesic distance is taken as the second geodesic distance.

[0189] To ensure the alignment consistency in both directions of image to text and text to image, a bidirectional geodesic contrast loss L bi-geodesic is defined as follows:

[0190] (Formula 34).

[0191] The loss comprehensively reflects the bidirectional consistency degree of the image features and the text features in the manifold space.

[0192] Based on the first geodesic distance and the second geodesic distance, the weighted attention corresponding to the stomach pathological image is determined.

[0193] In the method provided by the embodiment of the application, an active space focusing mechanism is configured, a curvature-aware manifold attention mechanism (solving the problem of modeling the importance of a local region) is introduced, and the weighted attention in the fusion is defined as follows:

[0194] (Formula 35).

[0195] Wherein, α i represents the weighted attention corresponding to the i-th local unit in the mapping text features, γ is a curvature adjustment weight parameter, F text,i represents the i-th local unit in the mapping text features, and the meaning of F text,j is the same. K(F text,i) represents the local curvature index, which is used to measure the degree of manifold curvature at the feature, that is, the degree of local nonlinear change of the feature in the manifold space. This index can be obtained by performing a second-order geometric analysis on the mapped text features.

[0196] In the method provided by the embodiment of the present invention, the weighted attention corresponding to each unit in the mapped text feature can be calculated based on the principle shown in Formula 35, and the weighted attention of each unit can be used as the weighted attention corresponding to the gastric pathology image.

[0197] Based on the weighted attention corresponding to the gastric pathology image, feature fusion is performed on the mapped image features and mapped text features corresponding to the gastric pathology image, and the fusion result is used as the image-text fusion feature corresponding to the gastric pathology image.

[0198] In the method provided by the embodiment of the present invention, the mapped image features and the mapped text features can be fused according to a pre-set cross-modal feature fusion method based on the weighted attention corresponding to the gastric pathology image to obtain the image-text fusion feature. For example, the image-text fusion feature F fusion It can be formally expressed as:

[0199] (Equation 36).

[0200] exist Figure 1 On the basis of the method shown, the method provided in the embodiment of the present invention further includes:

[0201] For each gastric pathology image, a pseudo negative sample generation process is performed based on the mapped text features corresponding to the gastric pathology image to obtain a pseudo negative sample corresponding to the gastric pathology image.

[0202] The visual language model is trained based on the pseudo negative samples corresponding to each gastric pathology image.

[0203] The method provided in the embodiment of the present invention sets up an active discrimination reinforcement mechanism, adopts the method of optimizing the mutual information of negative samples, and introduces a pseudo-negative sample mechanism to generate pseudo-negative samples to solve the problem of insufficient semantic consistency. Based on the mapped text features of gastric pathology images, corresponding pseudo-negative samples can be generated based on the corresponding mechanism for model training. Specifically, the pseudo-negative samples can be formally expressed as:

[0204] (Equation 37).

[0205] in, represents a pseudo negative sample, represents the small-amplitude noise perturbation matrix.

[0206] Further, the method provided by the embodiment of the present application defines the negative sample mutual information optimization loss, actively suppresses the surface matching phenomenon, and strengthens the cross-modal deep semantic connection, and the calculation method of the loss is as follows:

[0207] (Formula 38).

[0208] Wherein, is the mutual information estimation, is the positive and negative sample mutual information difference adjustment coefficient.

[0209] On the basis of the method shown in Figure 1 , the method provided by the embodiment of the present application further comprises the following steps:

[0210] Based on the corresponding graphic-text fusion features of each gastric pathology image, the complexity index corresponding to each gastric pathology image is determined.

[0211] In the method provided by the embodiment of the present application, an active training attention transfer mechanism is provided, and a manifold dynamic weighted multi-task planning is introduced to solve the task conflict and training adaptability problem. First, the complexity of each graphic-text fusion feature can be calculated according to the graphic-text fusion features corresponding to each gastric pathology image, and the complexity of each graphic-text fusion feature is taken as the complexity index corresponding to the corresponding gastric pathology image.

[0212] Based on the complexity index corresponding to each gastric pathology image, the preset complexity sensitive adjustment coefficient and the preset complexity threshold corresponding to each training task, the task weight corresponding to each training task is determined.

[0213] In the method provided by the embodiment of the present application, the complexity sensitive adjustment coefficient and the complexity threshold corresponding to each training task (i.e. the preset complexity threshold) can be set in advance according to the training requirements of each training task. Specifically, the task weight corresponding to each training task can be calculated based on the complexity index corresponding to each gastric pathology image, the complexity sensitive adjustment coefficient and the preset complexity threshold corresponding to each training task according to the following task weight calculation method:

[0214] (Formula 39).

[0215] Wherein, λ i represents the training weight corresponding to the i th training task, β represents the complexity sensitive adjustment coefficient, C (F fusion ) represents the complexity index corresponding to the gastric pathology image. θ i represents the complexity threshold corresponding to the i th training task.

[0216] Based on the task weight corresponding to each training task, the weight of each training task is adaptively adjusted.

[0217] In the method provided by the embodiment of the present application, the training weights of the current training tasks can be adaptively adjusted according to the task weights corresponding to the respective training tasks, that is, the task weights corresponding to the training tasks are taken as the adjusted training weights, so as to perform model training. total is defined as:

[0218] (Formula 40).

[0219] wherein M represents the total number of training tasks, L i represents the loss of the i th subtask (such as segmentation, classification, and diagnosis generation).

[0220] and Figure 1 Corresponding to the model training method based on manifold alignment shown in the above formula (40), the embodiment of the present application further provides a model training device based on manifold alignment, which is used for implementing the method shown in the above formula (40). Figure 1 The structural diagram of the model training device based on manifold alignment is shown in Figure 4 The model training device based on manifold alignment comprises:

[0221] A first determination unit 401 is configured to determine a training data set, wherein the training data set comprises a plurality of stomach pathological images, and description texts and annotation labels corresponding to the stomach pathological images.

[0222] A feature extraction unit 402 is configured to perform iterative training on a pre-constructed visual language model according to a preset training number based on the training data set; when entering a current iteration training period, the visual language model is applied to perform image feature extraction on each stomach pathological image respectively to obtain image features corresponding to each stomach pathological image, and text feature extraction on the description texts corresponding to each stomach pathological image respectively to obtain text features corresponding to each stomach pathological image.

[0223] A feature enhancement unit 403 is configured to, for each stomach pathological image, apply the visual language model to perform feature enhancement on the image features corresponding to the stomach pathological image to obtain enhanced image features corresponding to the stomach pathological image, and perform feature enhancement on the text features corresponding to the stomach pathological image to obtain enhanced text features corresponding to the stomach pathological image.

[0224] A feature optimization unit 404 is configured to, for each stomach pathological image, perform cross-modal feature optimization on the enhanced image features corresponding to the stomach pathological image based on the enhanced text features corresponding to the stomach pathological image to obtain cross-modal optimized image features corresponding to the stomach pathological image.

[0225] The feature mapping unit 405 is configured to, for each gastric pathology image, respectively perform manifold space mapping on the cross-modal optimized image feature and the enhanced text feature corresponding to the gastric pathology image, to obtain a mapped image feature and a mapped text feature corresponding to the gastric pathology image.

[0226] The feature fusion unit 406 is configured to, for each gastric pathology image, perform feature fusion on the mapped image feature and the mapped text feature corresponding to the gastric pathology image based on a preset active space focusing mechanism, to obtain a graphic-text fusion feature corresponding to the gastric pathology image.

[0227] The feature training unit 407 is configured to train a visual language model based on the graphic-text fusion feature corresponding to each gastric pathology image, to obtain a gastric pathology detection model.

[0228] In the training process of the gastric pathology detection model, the device provided by the embodiment of the present application can first optimize the image feature through the text feature in a cross-modal manner, then can further fuse the text feature and the image feature based on the active space focusing mechanism, after the text feature and the image feature are mapped to the manifold space, which is beneficial to integrating the cross-modal information of the text feature and the image feature, improves the training effect of the model, and then improves the detection accuracy of the gastric pathology detection model.

[0229] In Figure 4 Based on the device shown in the figure, the device provided by the embodiment of the present application can further extend a plurality of units, and the functions of the units can be referred to the descriptions of the various embodiments provided in the method for training the model based on manifold alignment, which will not be further illustrated.

[0230] Each of the embodiments in the specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the difference from other embodiments. Especially, for the system or system embodiment, since it is basically similar to the method embodiment, it is described more simply, and the related parts can be referred to the part of the method embodiment. The system and system embodiment described above are only illustrative, and the units described as separate components can be or can not be physically separated, and the components displayed as units can be or can not be physical units, that is, they can be located in one place or distributed on multiple network units. According to the actual needs, part or all of the modules can be selected to achieve the purpose of the embodiment. Those skilled in the art can understand and implement without creative labor.

[0231] Those skilled in the art will further realize that the mechanisms of the various examples described herein are capable of being implemented using any number of combinations of the described features. Accordingly, these examples are not limited to the mechanisms described herein, but rather, the intent is to cover all modifications and alternatives equivalent thereto. The preceding description of the examples is illustrative, and not restrictive. Many other examples will be apparent to those of skill in the art upon reviewing the above description. The scope of the examples should, therefore, be determined not with reference to the above description, but instead should be given to the appended claims, along with their full scope of equivalents.

[0232] The above description of disclosed examples is intended to be illustrative, and not restrictive. Many other examples will be apparent to those of skill in the art upon reviewing the above description. The scope of the examples should, therefore, be determined not with reference to the above description, but instead should be given to the appended claims, along with their full scope of equivalents.

Claims

1. A model training method based on manifold alignment, characterized in that: include: Determine the training dataset; The training data set includes a plurality of gastric pathology images and description texts and annotation labels corresponding to the gastric pathology images; Based on the training data set, the pre-constructed visual language model is iteratively trained according to a preset number of training times; when entering the current iterative training cycle, the visual language model is applied to perform image feature extraction on each of the gastric pathology images to obtain image features corresponding to each of the gastric pathology images, and text feature extraction is performed on the description text corresponding to each of the gastric pathology images to obtain text features corresponding to each of the gastric pathology images; For each of the gastric pathology images, applying the visual language model, performing feature enhancement on image features corresponding to the gastric pathology image to obtain enhanced image features corresponding to the gastric pathology image, and performing feature enhancement on text features corresponding to the gastric pathology image to obtain enhanced text features corresponding to the gastric pathology image; For each of the gastric pathology images, based on the corresponding enhanced text features, cross-modal feature optimization is performed on the enhanced image features corresponding to the gastric pathology image to obtain cross-modal optimized image features corresponding to the gastric pathology image; For each of the gastric pathology images, performing manifold space mapping on the cross-modal optimized image features and enhanced text features corresponding to the gastric pathology image to obtain mapped image features and mapped text features corresponding to the gastric pathology image; For each of the gastric pathology images, based on a preset active spatial focusing mechanism, feature fusion is performed on the mapped image features and the mapped text features corresponding to the gastric pathology image to obtain an image-text fusion feature corresponding to the gastric pathology image; Based on the image-text fusion features corresponding to each of the gastric pathology images, the visual language model is trained to obtain a gastric pathology detection model.

2. The model training method based on manifold alignment according to claim 1, characterized in that: The step of performing feature enhancement on the image features corresponding to the gastric pathological image to obtain enhanced image features corresponding to the gastric pathological image includes: Determining whether the current iterative training cycle meets the preset dynamic adjustment conditions; If the current iterative training cycle meets the dynamic adjustment condition, determining the number of pixels to be screened corresponding to the current iterative training cycle; Determining a comprehensive score corresponding to each pixel in the gastric pathology image, and determining a difficult area feature corresponding to the gastric pathology image from image features corresponding to the gastric pathology image based on a preset loss function, the number of pixels screened, and the comprehensive score corresponding to each pixel in the gastric pathology image; Feature enhancement processing is performed on the difficult area features corresponding to the gastric pathological image, and the processed features are used as enhanced image features corresponding to the gastric pathological image, so that the visual language model focuses on the training of difficult areas.

3. The model training method based on manifold alignment according to claim 2, characterized in that: The determining whether the current iterative training cycle meets the preset dynamic adjustment condition includes: Determine the training loss value and the validation set mean intersection-over-union (IoU) index corresponding to the current iterative training cycle; Based on the training loss value and the validation set mean intersection-over-union (IoU) index, an indicator trend forecast is performed to obtain a training loss prediction value and a validation set mean intersection-over-union (IoU) prediction value corresponding to the current iterative training cycle; Based on the training loss prediction value and the validation set mean intersection-union ratio prediction value, determining whether the current iterative training cycle meets a preset first trigger condition, and obtaining a first trigger condition judgment result; Determining a curvature anomaly ratio corresponding to the current iterative training cycle; Based on the curvature anomaly ratio, determining whether the current iterative training cycle meets a preset second trigger condition, and obtaining a second trigger condition determination result; Determine the average loss conflict degree corresponding to the current iterative training cycle; Based on the average loss conflict degree, determining whether the current iterative training cycle meets a preset third trigger condition, and obtaining a third trigger condition determination result; Determining a dynamically adjusted activation probability corresponding to the current iterative training cycle based on the first trigger condition judgment result, the second trigger condition judgment result, and the third trigger condition judgment result; Determining whether the dynamically adjusted activation probability is greater than a preset probability threshold; If the dynamically adjusted activation probability is greater than the probability threshold, it is determined that the current iterative training cycle meets the dynamic adjustment condition.

4. The model training method based on manifold alignment according to claim 1, characterized in that The step of enhancing the text features corresponding to the gastric pathological image to obtain the enhanced text features corresponding to the gastric pathological image includes: Perform word vector embedding on the text features corresponding to the gastric pathology image to obtain a word vector sequence corresponding to the text features; Layering the word vector sequence to obtain word vectors corresponding to each layer; Determine the attention weight corresponding to each word vector corresponding to each of the layers; Based on a preset bidirectional long short-term memory network and the attention weights corresponding to each of the word vectors, contextual feature fusion is performed on each of the word vectors, and the fusion result is used as the enhanced text feature corresponding to the gastric pathology image.

5. The model training method based on manifold alignment according to claim 1, characterized in that: The cross-modal feature optimization of the enhanced image features corresponding to the gastric pathology image based on the corresponding enhanced text features to obtain the cross-modal optimized image features corresponding to the gastric pathology image includes: Performing position indexing on the enhanced text features corresponding to the gastric pathology image through a preset relative position discriminator to obtain the pathological position corresponding to the gastric pathology image; Determining a feature region corresponding to the pathological position in the enhanced image features corresponding to the stomach pathological image; Performing edge-aware processing on the feature region to obtain edge-aware features corresponding to the feature region; Performing texture perception processing on the feature area to obtain texture perception features corresponding to the feature area; Performing semantic aggregation processing on the feature area to obtain semantic aggregation features corresponding to the feature area; The edge perception feature, the texture perception feature and the semantic aggregation feature are fused, and the fusion result is used as the cross-modal optimized image feature corresponding to the gastric pathology image.

6. The model training method based on manifold alignment according to claim 1, characterized in that: The performing manifold space mapping on the cross-modal optimized image features and enhanced text features corresponding to the gastric pathology image to obtain the mapped image features and mapped text features corresponding to the gastric pathology image includes: Performing manifold space mapping processing on the cross-modal optimized image features corresponding to the gastric pathology image through a preset symmetric positive definite matrix, and using the processing results as the mapped image features corresponding to the gastric pathology image; The enhanced text features corresponding to the gastric pathology image are subjected to manifold space mapping processing through a preset covariance matrix, and the processing results are used as the mapped text features corresponding to the gastric pathology image.

7. The model training method based on manifold alignment according to claim 1, characterized in that: The feature fusion of the mapped image features and the mapped text features corresponding to the gastric pathology image based on the preset active spatial focusing mechanism to obtain the image-text fusion features corresponding to the gastric pathology image includes: Performing bidirectional distance measurement between features on the mapped image features and the mapped text features corresponding to the gastric pathology image to obtain a first geodesic distance and a second geodesic distance corresponding to the gastric pathology image; determining a weighted attention corresponding to the gastric pathological image based on the first geodesic distance and the second geodesic distance; Based on the weighted attention corresponding to the gastric pathology image, feature fusion is performed on the mapped image features and mapped text features corresponding to the gastric pathology image, and the fusion result is used as the image-text fusion feature corresponding to the gastric pathology image.

8. The model training method based on manifold alignment according to claim 1, characterized in that: Also includes: For each of the gastric pathology images, performing pseudo negative sample generation processing based on the mapped text features corresponding to the gastric pathology image to obtain a pseudo negative sample corresponding to the gastric pathology image; The visual language model is trained based on pseudo negative samples corresponding to each of the gastric pathology images.

9. The model training method based on manifold alignment according to claim 1, characterized in that: Also includes: determining a complexity index corresponding to each of the gastric pathology images based on the image-text fusion features corresponding to each of the gastric pathology images; Determining a task weight corresponding to each of the training tasks based on a complexity index corresponding to each of the gastric pathology images, a preset complexity sensitivity adjustment coefficient, and a preset complexity threshold corresponding to each of the training tasks; Based on the task weights corresponding to the respective training tasks, the weights of the respective training tasks are adaptively adjusted.

10. A model training device based on manifold alignment, characterized in that: include: A first determining unit, configured to determine a training data set; The training data set includes a plurality of gastric pathology images and description texts and annotation labels corresponding to the gastric pathology images; a feature extraction unit, configured to iteratively train a pre-constructed visual language model based on the training data set according to a preset number of training times; and when entering a current iterative training cycle, applying the visual language model to perform image feature extraction on each of the gastric pathology images to obtain image features corresponding to each of the gastric pathology images, and performing text feature extraction on the descriptive text corresponding to each of the gastric pathology images to obtain text features corresponding to each of the gastric pathology images; a feature enhancement unit configured to apply the visual language model to each of the gastric pathology images, perform feature enhancement on image features corresponding to the gastric pathology image to obtain enhanced image features corresponding to the gastric pathology image, and perform feature enhancement on text features corresponding to the gastric pathology image to obtain enhanced text features corresponding to the gastric pathology image; a feature optimization unit configured to perform cross-modal feature optimization on the enhanced image features corresponding to each of the gastric pathology images based on the enhanced text features corresponding to the gastric pathology images, so as to obtain cross-modal optimized image features corresponding to the gastric pathology images; a feature mapping unit, configured to perform manifold space mapping on the cross-modal optimized image features and enhanced text features corresponding to each of the gastric pathology images, to obtain mapped image features and mapped text features corresponding to the gastric pathology image; a feature fusion unit configured to, for each of the gastric pathology images, fuse the mapped image features and the mapped text features corresponding to the gastric pathology image based on a preset active spatial focusing mechanism to obtain an image-text fusion feature corresponding to the gastric pathology image; A feature training unit is used to train the visual language model based on the image-text fusion feature corresponding to each of the gastric pathology images to obtain a gastric pathology detection model.

Citation Information

Patent Citations

  • Interactive image editing method and device, readable storage medium and electronic equipment

    CN113448477A

  • Image recognition method and data processing method for image recognition

    CN116109896A

  • Intelligent query semantic understanding method based on information geometry and Riemannian manifold

    CN119829740A

  • Methods, systems and software for improved diagnosis of a medical condition

    WO2021044431A1