Method and system for generating elastic-like ultrasonic image through TDM-GAN (Time Division Multiplexing-Generation Area Network) with AttConv-Transform detection introduced

The generation of elastic ultrasound-like images through the TDM-GAN network detected by AttConv-Transformer solves the problems of high cost and inconsistent results of traditional ultrasound elastologic imaging, and improves diagnostic accuracy and equipment applicability.

CN120148778APending Publication Date: 2025-06-13ZHEJIANG CANCER HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510056520.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Traditional ultrasonic elastic imaging equipment is costly, has low scanning efficiency and inconsistent examination results. The existing artificial intelligence model ignores influencing factors, which leads to low trust in the generation of images, and is unable to effectively assist in disease diagnosis.

Method used

The TDM-GAN network detected by AttConv-Transformer is adopted, combined with Transformer, self-attention mechanism and convolutional operations, and generates an adversarial network through multimodal data to generate elastic ultrasound images, integrate ultrasound B-mode images, elastic images and diagnostic reports to improve diagnostic accuracy.

Benefits of technology

It improves the accuracy and efficiency of ultrasound imaging generation, overcomes the influence of tissue characteristics, operating technology and instrument differences, enhances physicians' ability to diagnose lesions, and improves the applicability of the model in different brands of equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148778A_ABST
    Figure CN120148778A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a system for generating an elastic-like ultrasonic image by introducing a TDM-GAN (Time Division Multiplexing-Generation Area Network) for AttConv-Transform detection. The generation method comprises the following steps: S1, integrating data, and screening data of a patient subjected to organ ultrasonic examination; s2, feature engineering: acquiring a focus area and annotation information existing in an ultrasonic B-mode image delineated by a physician, performing coding according to different data characteristics, and constructing a multi-mode data feature group; s3, model construction: designing a detection model in which AttConv-Transform is introduced, and designing an applicable loss function in combination with a TDM-GAN network model controlled by multi-modal data; s4, model training: using transfer learning to train a detection model, using a multi-modal data feature group to train a TDM-GAN network model, and combining with fining to optimize joint model performance; and S5, model application. According to the method, the elasticity-like ultrasonic image with relatively high accuracy can be generated based on the common ultrasonic B-mode image, and the problems of high equipment cost, low scanning efficiency and great influence of subjective and objective factors in the traditional ultrasonic elasticity inspection are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical imaging multi-modal data fusion and medical imaging generation, and particularly relates to a method and system for generating elastogram-like images by introducing a TDM-GAN network with AttConv-Transformer detection. Background Art

[0002] Tumor imaging examinations mainly include X-ray examination, computed tomography (CT), magnetic resonance imaging (MRI),

[0003] ultrasonic examination (US), positron emission computed tomography (PET-CT), and radionuclide imaging. Compared with radiological examinations, ultrasonic examination has the advantages of being radiation-free, repeatable, real-time imaging, flexible operation, low cost, and suitable for special populations. At the same time, when performing an ultrasonic examination on a suspected lesion, elastography plays an important role in differentiating the benign and malignant nature, composition, and evaluating the invasiveness of tumors in the medical field.

[0004] Traditional ultrasonic elastography not only has a high equipment cost and low scanning efficiency, but may also show inconsistent examination results for the same lesion under different circumstances due to factors such as tissue characteristics (composition, structure, and depth), operation techniques (pressure, position, and angle), patient individuals (age, underlying diseases, respiration, and heartbeat), and instrument performance (frequency, brand, and model). Existing scholars use artificial intelligence models to generate ultrasonic elastogram images through conventional B-mode ultrasonic examinations, but only use the patient's image data for modeling, ignoring the above four factors that may affect elastography, resulting in a low trust level in the generated images and no help in assisting disease diagnosis when physicians diagnose patients' diseases by combining the generated elastogram images. Summary of the Invention

[0005] Aiming at the above problems, the purpose of the present invention is to provide a method and system for generating elastogram-like images by introducing a TDM-GAN network with AttConv-Transformer detection. This method constructs a method for generating elastogram-like images based on ultrasonic B-mode images with higher accuracy and efficiency by introducing a detection network model that combines Transformer, self-attention mechanism, and convolutional operation structure (AttConv-Transformer) into a Transformer-Driven Multimodal Generative Adversarial Model (TDM-GAM).

[0006] To achieve the above purpose, the present invention adopts the following technical solutions:

[0007] On the one hand, the present invention proposes a method for generating elastogram-like ultrasound images using a TDM-GAN network with AttConv-Transformer detection, comprising the following steps:

[0008] S1. Data integration: Screening patient data undergoing organ ultrasound examinations, where the patient data includes at least ultrasound B-mode images, corresponding ultrasound elastography images, and diagnostic reports. Among them, both the ultrasound B-mode images and the corresponding ultrasound elastography images are Dicom format images, containing at least the patient's age, gender, examined organ, ultrasound instrument brand and model, and probe model information. The diagnostic report contains at least the XX-Rads score of the corresponding organ, lesion location, and lesion section information;

[0009] S2. Feature engineering: Obtaining the lesion areas existing in the ultrasound B-mode images outlined by physicians, as well as the annotations of the examined organ and lesion information based on the diagnostic report; Encoding continuous data according to data characteristics and performing one-hot encoding on discrete data to construct a multi-modal data feature group;

[0010] S3. Model construction: Designing a detection network model incorporating the AttConv-Transformer structure and a TDM-GAN network model controlled by multi-modal data, and simultaneously designing loss functions suitable for the detection network model and the TDM-GAN network model respectively. Among them, the TDM-GAN network model includes a generation network and a discriminator network;

[0011] S4. Model training: Using transfer learning to train the detection network model incorporating the AttConv-Transformer structure, and training the TDM-GAN network model using ultrasound B-mode images and the multi-modal data feature group; After independently training the two models, perform serial splicing and optimize the performance of the combined model through fine-tuning;

[0012] S5. Model application: Generating corresponding elastogram-like ultrasound images for target ultrasound B-mode images based on the trained model.

[0013] Further, the above generation method further includes the step of:

[0014] S6. Model conversion and deployment: Deploying the TDM-GAN network model with AttConv-Transformer detection into a C++ callable type for efficient inference operations when integrated into ultrasound instruments of different brands.

[0015] Further, in the above step S1, the screening of patient data for organ ultrasound examination further includes including the ultrasonic B-mode image data of the patient's ultrasound screening, physical examination, routine examination, and reexamination, as well as the corresponding elastography data. The included organs include one or more of the thyroid gland, breast, liver, ovary, prostate gland, and kidney; the lesion location is the orientation of the current organ to which the lesion belongs; the lesion section information includes transverse section, longitudinal section, oblique section, or coronal section; the probe models include convex array, linear array, phased array, or micro-convex array probes.

[0016] Further, in the above step S2, obtaining the lesion area existing in the ultrasonic B-mode image outlined by the physician, as well as the annotation of the examined organ and lesion information according to the diagnostic report, includes:

[0017] 1) Preprocessing the image:

[0018] a. Desensitizing the image data: Removing the identity recognition information of the patient's name, medical record number, ID number, contact information, and image number from the Dicom format image; at least retaining the information of age, gender, examined organ, ultrasonic instrument brand and model, and probe model;

[0019] b. Converting the image format: Saving the Dicom format image as a high-quality TIFF image through a lossless conversion method;

[0020] 2) Obtaining the physician's outlined information: Obtaining the outlined information of the examined organ and lesion area by the physician under the TIFF format ultrasonic B-mode image. The types of labels after outlining include four categories, namely whether there is a lesion, organ, lesion location, and XX-Rads score of the corresponding organ.

[0021] Further, in the above step S3, the detection network model introducing the AttConv-Transformer structure uses YoloV10 as the base model, adding a Transformer structure to the Backbone part of YoloV10 and adding a self-attention mechanism structure to the Neck part of YoloV10.

[0022] Further, the loss function used by the detection network model introducing the AttConv-Transformer structure is L Att-Trs , which is composed of classification loss L Focal , bounding box loss L GIoU , similarity L1 penalty L Dice_L1 loss functions. The expression of its total loss function is:

[0023] L Att-Trs = L Focal + L GIoU + LDice_L1

[0024] Among them, L Focal is a multi-class loss function for distinguishing benign and malignant lesions and different organ positions. C_num is the number of detection categories, and α Focal is a balance factor hyperparameter used to adjust the weights between different categories. P c is the predicted probability belonging to category c. γ is a tempering factor hyperparameter used to adjust the detection weights between micro-lesions, normal lesions, and non-lesions. Its expression is:

[0025]

[0026] Among them, L GIoU is the intersection over union area of the detected bounding boxes of different categories. G 1 ∩G 2 is the intersection area of two detection boxes of a certain category. G 1 ∪G 2 is the union area between the two. G min is G 1 and G 2 The minimum bounding rectangle area of the two rectangular boxes. Its expression is:

[0027]

[0028] Among them, L Dice_L1 is used to measure the similarity between the true target region box and the predicted target region box, improve the detection rate of micro-lesions, and enable the model to better learn the position and shape information of the target. G 1 is the true label box of the lesion or organ annotation information outlined by the physician. G 2 is the predicted box detected during the model training stage. X and Y correspond to the coordinate values of the center coordinates of the rectangular box respectively. W and H are the width and height of the rectangular box respectively. The subscripts true and pred are the rectangular box outlined by the physician and the rectangular box predicted by the model respectively. γ and δ are the weight coefficients automatically adjusted during training. Its expression is:

[0029]

[0030] Furthermore, in the above step S3, a Transformer structure is added to the encoding part of the generation network of the TDM-GAN network model. Based on the traditional binary cross-entropy loss of the GAN network, a joint three-channel color loss function L Color is added. L consistency is the color consistency loss. L distribution is the color distribution loss. L DiscriminationTo distinguish the discriminative loss between the generated pseudo-elastic and real ultrasound elastograms, its expression is as follows:

[0031] L Color = L Consistency + L Distribution + L Discrimination

[0032] Among them, L Donsistency The color consistency loss function ensures the color consistency of the generated pseudo-elastic ultrasound images. The Euclidean distance is used to calculate the difference between two adjacent pixels. G Fake is the pseudo-elastic ultrasound image generated by TDM-GAN, and G Fake (i, j) is the pixel value at the point with coordinates i, j in the pseudo-elastic ultrasound image. Its expression is as follows:

[0033]

[0034] Among them, L Distribution The color distribution loss function calculates the distribution difference between the pseudo-elastic and real ultrasound elastograms, which can help the generation network generate pseudo-elastic ultrasound images with a color distribution consistent with that of real ultrasound elastograms. KL is the Kullback-Leibler divergence, which is used to calculate the probability distribution difference between two images. P(G Fake ) is the color probability distribution function of the pseudo-elastic ultrasound image, and P(G Real ) is the color distribution function of the real ultrasound elastogram. Its expression is as follows:

[0035] L Distribution = KL(P(G Fake ), P(G Real ))

[0036] Among them, L Discriminaction is the discriminative loss function, which is used to judge the pseudo-elastic and real elastogram data generated by the generation network. Y G is the model's judgment on whether it is a real elastogram. Y G ∈{0, 1}, and P is the probability predicted to be a real elastogram. Its expression is as follows:

[0037] L Discrimination = -[Y G LogP + (1 - Y G )Log(1 - P)].

[0038] Furthermore, in the above step S4, the performance of the combined model is optimized by combining fine-tuning, including:

[0039] Generate elastography-like ultrasound images in the test set, evaluate the generated elastography-like ultrasound images according to evaluation metrics, and fine-tune and optimize the hyperparameters of the TDM-GAN network with AttConv-Transformer detection introduced based on the levels of the evaluation metrics. Train repeatedly until the model converges. The evaluation metrics include peak signal-to-noise ratio, structural similarity index, and physician evaluation metric.

[0040] Further, in the above step S6, deploying the TDM-GAN network model with AttConv-Transformer detection introduced into a C++-callable type includes converting the weight file in Pth format obtained after training and optimization through the Pytorch framework into a weight file in Pt format, enabling the C++ language to implement the call and inference of the model and improving the portability of the model.

[0041] On the other hand, the present invention proposes a system for generating elastography-like ultrasound images using a TDM-GAN network with AttConv-Transformer detection introduced, for implementing the above generation method, including:

[0042] A data integration module, configured to screen patient data of patients undergoing organ ultrasound examinations. The patient data at least includes ultrasound B-mode images, corresponding ultrasound elastography images, and diagnostic reports. Among them, both the ultrasound B-mode images and the corresponding ultrasound elastography images are Dicom format images, at least including the patient's age, gender, examined organ, ultrasound instrument brand and model, and probe model information. The diagnostic report at least includes the XX-Rads score of the corresponding organ, lesion location, and lesion section information;

[0043] A feature engineering module, configured to obtain the lesion areas existing in the ultrasound B-mode images outlined by a physician, and the annotations of the examined organ and lesion information according to the diagnostic report; encode continuous data according to data characteristics, perform one-hot encoding on discrete data, and construct a multi-modal data feature group;

[0044] A model construction module, configured to design a detection network model introducing an AttConv-Transformer structure, and a TDM-GAN network model controlled by multi-modal data, and simultaneously design loss functions respectively applicable to the detection network model and the TDM-GAN network model. Among them, the TDM-GAN network model includes a generation network and a discriminant network;

[0045] The model training module is configured to train a detection network model introducing the AttConv-Transformer structure using the transfer learning method, and train the TDM-GAN network model using ultrasonic B-mode images and multi-modal data feature groups. After independently training the two models, they are concatenated in series, and the performance of the combined model is optimized by fine-tuning.

[0046] The model application module is configured to generate corresponding elastographic ultrasound images from the target ultrasonic B-mode images based on the trained model.

[0047] The model conversion and deployment module is configured to deploy the TDM-GAN network model introducing AttConv-Transformer detection into a C++ callable type, facilitating integration into ultrasonic instruments of different brands for efficient inference operations.

[0048] The present invention has the following beneficial effects:

[0049] The present invention extracts features from ultrasonic B-mode images through Transformer, self-attention mechanism, object detection, and adversarial network generation models in the deep learning convolutional neural network, and combines the patient's clinical information and the information in the Dicom format ultrasonic B-mode images to generate elastographic ultrasound images. This deep learning detection and generation network introducing Transformer and self-attention mechanism can automatically judge whether there are lesions in the ultrasonic images. If there are lesions, the designed model can extract the scores of organs and lesions, overcoming the subjective and objective factors brought by tissue characteristics, physician operation techniques, patient individual differences, and instrument imaging differences in traditional ultrasonic elastography examinations. The TDM-GAN network detected by AttConv-Transformer can objectively give the ultrasonic elastography corresponding to the ultrasonic B-mode images, which is more suitable for assisting ultrasonic physicians in diagnosing organs with lesions, further improving the diagnostic accuracy of physicians for organ lesions. And by deploying the model into a C++ callable type, the portability of this method in devices of different brands is improved, greatly enhancing the general applicability of the deep learning model. Description of the Drawings

[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for description in this embodiment. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0051] Figure 1Flow chart designed for the generation method of elastography-like ultrasound images by the TDM-GAN network with AttConv-Transformer detection in the embodiments of the present invention;

[0052] Figure 2 Backbone structure diagram of the Transformer structure and the detection network model added in the generation method of elastography-like ultrasound images by the TDM-GAN network with AttConv-Transformer detection in the embodiments of the present invention;

[0053] Figure 3 Neck structure diagram of the self-attention mechanism structure and the detection network model added in the generation method of elastography-like ultrasound images by the TDM-GAN network with AttConv-Transformer detection in the embodiments of the present invention;

[0054] Figure 4 Structure diagram of the network model of the generation method of elastography-like ultrasound images by the TDM-GAN network with AttConv-Transformer detection in the embodiments of the present invention;

[0055] Figure 5 Detection and generation graph display in the generation method of elastography-like ultrasound images by the TDM-GAN network with AttConv-Transformer detection in the embodiments of the present invention. Detailed implementation manners

[0056] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0057] Term explanation:

[0058] Transformer: A deep learning model based on the self-attention mechanism;

[0059] AttConv (Attentive Convolution): A neural network module that combines the attention mechanism and convolutional operation;

[0060] TDM-GAN (Transformer-Driven Multimodal Generative Adversarial Model): A multi-modal generative adversarial network model based on the Transformer architecture;

[0061] XX-Rads Score: A scoring method used in the Radiology Reporting and Data System;

[0062] Transfer learning: A technique in machine learning that allows a model to apply the knowledge learned in one task (source task) to another different but related task (target task);

[0063] YoloV10 (You Only Look Once version 10): The latest generation of object detection algorithm developed by the research team at Tsinghua University;

[0064] Backbone: Refers to the core part of a deep neural network for feature extraction;

[0065] Neck: Refers to the connecting part between the backbone network (Backbone) and the head (Head). Its main function is to further process and integrate the features extracted by the backbone network, helping to better transfer the features to the head for the final prediction or classification task;

[0066] Fine-tuning: A technique in machine learning and deep learning used to adjust a model that has been trained on certain data in order to optimize it for a new, related task

[0067] C2f (Channel-to-Pixel): An important component in YOLOv10, mainly used for feature fusion;

[0068] SCDown (Spatial-Channel Decoupled Downsampling): An important component in YOLOv10. It decouples the downsampling in the spatial dimension and the downsampling in the channel dimension to achieve more effective feature extraction and dimensionality reduction operations;

[0069] C2fCIB (Convolution to Feature Map Fusion and Inverted Bottleneck): A key component in YOLOv10 for improving the efficiency and performance of feature extraction. By combining the advantages of the C2f and CIB modules, it achieves more effective feature fusion and extraction;

[0070] SPPF (Spatial Pyramid Pooling Fast): An improved spatial pyramid pooling technology mainly used for multi-scale feature extraction to enhance the model's detection ability for targets of different sizes;

[0071] PSA (Partial Self-Attention): An efficient self-attention mechanism that reduces computational complexity by applying self-attention only to a part of the features while enhancing the model's global modeling ability;

[0072] Feature Map Scale: In image processing and computer vision, it refers to the size and dimension of the feature map output by each layer in a convolutional neural network (CNN).

[0073] Upsampling: Also known as resampling or interpolation, it is a technique in signal processing and image processing used to increase the sampling rate or resolution of a signal or image, thus making the image or signal larger;

[0074] Downsampling: Refers to the technique of reducing the sampling rate or resolution of a signal or image in signal processing and image processing;

[0075] Softmax: An activation function widely used in machine learning and deep learning, especially when dealing with multi-class classification problems.

[0076] Embodiment 1

[0077] Figure 1 Shows the flowchart of the design of the generation method of the TDM-GAN network introducing AttConv-Transformer detection for class elastic ultrasound images in the embodiments of the present invention. As Figure 1 shown, this embodiment mainly includes six steps S1 to S6, including:

[0078] Step S1, data integration, screening patient data for organ ultrasound examinations. The patient data includes at least ultrasound B-mode images, corresponding ultrasound elastograms, and diagnostic reports. Among them, both the ultrasound B-mode images and the corresponding ultrasound elastograms are Dicom format images, and at least include the patient's age, gender, examined organ, ultrasound instrument brand and model, and probe model information. The diagnostic report at least includes the XX-Rads score of the corresponding organ, lesion location, and lesion section information.

[0079] Specifically, when screening patient data, in addition to the patient's clinical information and basic information such as gender and age, the ultrasonic B-mode data and the corresponding elastography data of ultrasonic screening, physical examination, routine examination, and reexamination need to be included. The organs included are not limited to the thyroid, breast, liver, ovary, prostate, and kidney. If there are lesions in the collected ultrasonic image data, the diagnosis report needs to include the corresponding lesion nature, the orientation of the current organ to which the lesion belongs (such as the left lobe, right lobe, and isthmus of the thyroid, the upper outer, lower outer, upper inner, lower inner quadrants and central area of the breast, the left lobe, right lobe, quadrate lobe, and caudate lobe of the liver), and the section information (transverse section, longitudinal section, oblique section, and coronal section) of the current position, as well as the complete pre-examination, in-examination, and post-examination result data of the ultrasonic probe model (convex array, linear array, phased array, micro-convex array probe).

[0080] Among them, the initial formats of the included ultrasonic B-mode and the corresponding ultrasonic elastography images need to be in the Dicom medical image storage standard data format. One image contains the ultrasonic B-mode image and the corresponding ultrasonic elastography image (the left and right positions are indistinguishable). This format contains the patient's basic information (name, gender, age, and date of birth), examination information (examination date, time, examination type, and examination modality), basic parameters of the image data (image size, resolution, bit depth, and pixel format), and image annotation information (anatomical markers, positioning, annotations, and markings). Due to the requirements of medical data ethics, subsequent desensitization processing must be carried out in the data preprocessing stage to remove the identity recognition information such as the patient's name, ID number, and contact information in the Dicom format, and retain the multi-modal information of age, gender, XX-Rads score, lesion location, lesion section, brand of the ultrasonic examination instrument, model, and probe type required in this embodiment.

[0081] Step S2, feature engineering, obtaining the lesion areas existing in the ultrasonic B-mode images outlined by the physician, and the annotations of the examined organs and lesion information according to the diagnosis report; encoding continuous data and performing one-hot encoding on discrete data according to the data characteristics to construct a multi-modal data feature group.

[0082] Specifically, the ultrasonic B-mode and elastography ultrasonic image data in the original DICOM format are parsed by using the Pydicom and SimpleITK scientific databases in the Python language through a lossless conversion method. The patient's basic information (identity identification information such as name, medical record number, and image number), examination information (examination type, scanning method, and parameters), and other information (physician annotations, marks, image orientation, coordinate system, and storage information) are removed. After parsing, the image pixel data (original resolution and window width and window level), examination site, and equipment information (equipment manufacturer, model, and acquisition parameters) are saved. To minimize image loss during the conversion process, the DICOM-format ultrasonic image data is converted into a high-quality TIFF image format through a lossless conversion method.

[0083] Subsequently, the physician uses the Labelme drawing software written in the Python language and annotates using the rectangular box annotation method in combination with the benign and malignant regions of the lesions diagnosed by the pathological gold standard and the scanning sites marked by the probe in the ultrasonic image. When using the rectangular box method for annotation, the physician needs to determine the position and size of the rectangular box according to the lesion and the scanning site, making the rectangular box fit the contours of the lesion and the scanning site as much as possible, and avoiding including too many non-target regions. The types of labels after drawing are four categories, namely whether there is a lesion (present, absent), organ (thyroid, breast, liver, ovary, prostate, kidney), lesion location, and XX-Rads score (e.g., Ti-Rads 1-6 levels for the thyroid, where grade 4 is not sub-classified in the present invention).

[0084] Furthermore, one-hot encoding is performed on discrete data such as XX-Rads score, gender, organ type, whether there is a lesion, lesion location, lesion section information, and examination instrument information. For example, the discrete Ti-Rads score has 6 categories C = {1, 2, 3, 4, 5, 6}. For the score C i (i = 1, 2, 3, 4, 5, 6), its encoding rule is as follows:

[0085]

[0086] where, v i,j is the j-th element of the vector v i After encoding, the length of the v i,j vector is the number of categories of the discrete variable.

[0087] Furthermore, normalization encoding is performed on the continuous data of age information, and its encoding rule is as follows:

[0088]

[0089] where, x age0-1is the normalized age data, whose range is within the interval of 0 - 1 (both ends are closed intervals), x is the original age data, x min is the minimum value in the age data, x max is the maximum value in the age data.

[0090] Step S3, model construction. Design a detection network model introducing the AttConv-Transformer structure and a TDM-GAN network model controlled by multi-modal data. At the same time, design loss functions respectively applicable to the detection network model and the TDM-GAN network model. Among them, the TDM-GAN network model includes a generation network and a discriminant network.

[0091] Specifically, the combined model includes a detection network model introducing the AttConv-Transformer structure and a TDM-GAN network model for generating elastography-like ultrasound images. The purpose of the detection network model introducing the AttConv-Transformer structure is to detect the organ type, the presence of lesions, the lesion location, and the XX-Rads grade score of the corresponding organ in the ultrasound B-mode image. Specifically, the detection network model detects the organ to be examined. When the detection accuracy of the organ is 90% or above, one-hot encoding is performed on the organ category; the detection network model detects whether there is a lesion. If no lesion is detected, the generation of elastography-like ultrasound image data stops. If a lesion is detected, the lesion location and XX-Rads score are judged in combination with its organ type. After obtaining the lesion location and XX-Rads score, one-hot encoding is performed on them; when a lesion is present, information such as the patient's age, gender, lesion section, ultrasound instrument model, and probe signal during scanning in the Dicom format is read, and corresponding data type encoding is performed on it.

[0092] Furthermore, the detection network model introducing the AttConv-Transformer structure uses the object detection model of the tenth generation of Yolo (abbreviated as YoloV10) as the base model. Since there may be tiny lesions (diameter < 5 mm) in the organs, a Transformer structure is added to the backbone part of YoloV10, enabling the detection model part to effectively integrate global information, realize the fusion of features at different scales, and at the same time utilize the dynamic learning characteristics of the Transformer structure to adaptively adjust the weights, thereby automatically learning the feature representations of nodules of different sizes under different organs; a self-attention mechanism is added to the neck of the detection model to enhance the detailed information of lesions of different sizes, thereby enhancing the feature expression of tiny targets. On the other hand, the self-attention mechanism has the ability to adaptively adjust the focal length. When tiny lesions are under occlusion or incomplete scanning, this structure automatically adjusts the convolutional kernel weights through the content of the whole image, making the convolutional operation focus on the tiny target area, thereby enhancing the feature expression of tiny targets.

[0093] Furthermore, in the AttConv-Transformer detection part, due to the unbalanced data ratios of the cases of no lesions, presence of lesions, and tiny lesions, a loss function L suitable for the AttConv-Transformer detection part is designed. Att-Trs , which consists of a classification loss L Focal , a bounding box loss L GIoU , and a similarity L1 penalty L Dice_L1 loss function. The expression of its total loss function is:

[0094] L Att-Trs = L Eocal + L GIoU + L Dice_L1

[0095] Furthermore, the TDM-GAN network model controlled by multi-modal information is used to generate elastography-like ultrasound images, which includes a generator network and a discriminator network. A Transformer structure is added to the encoding part of the generator network. Different from traditional generation models, in the generated feature maps, three single-channel elastography-like ultrasound images are generated, and their channels are R, G, and B respectively. Subsequently, three hyperparameters θ, σ, and ε that control the weights of the R, G, and B channels are updated through the backpropagation algorithm during the training process. Then, the three single-channel images are synthesized to obtain a three-channel elastography-like ultrasound image.

[0096] Since the generated elastography-like ultrasound image is an RGB three-channel image, on the basis of the traditional binary cross-entropy loss of the GAN network, a joint three-channel color loss function L Color , Lconsistency is the color consistency loss, L distribution is the color distribution loss, L Discrimination is the discriminant loss for distinguishing between generated and real ultrasound elastography data, and its expression is:

[0097] L Color = L Consistency + L Distribution + L Discrimination

[0098] Step S4, model training. Use the transfer learning method to train the detection network model introducing the AttConv-Transformer structure, and use the ultrasound B-mode images and multi-modal data feature groups to train the TDM-GAN network model; after independently training the two models, perform series splicing, and combine fine-tuning to optimize the performance of the joint model.

[0099] Specifically, to ensure the accuracy of its detection and generation, the AttConv-Transformer lesion and organ detection network model and the TDM-GAN class elastographic ultrasound image generation network model are trained independently. Use the YoloV10-Medium model combined with the transfer learning method to achieve a better balance between data resources and generalization performance for the detection network model. When loading the pre-trained weight parameters, for the Transformer and self-attention mechanism parts added in the Backbone and Neck of the model, use random initialization parameters and do not load the pre-trained model parameters. The weight parameters of the TDM-GAN network model are randomly initialized without transfer learning.

[0100] Furthermore, after independently training the two models, perform series splicing, and combine fine-turning to optimize the trained joint model. Generate elastographic ultrasound-like images in the test set, and evaluate the generated elastographic ultrasound-like images according to objective evaluation indicators such as peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), and subjective evaluation indicators of physicians. Fine-tune and optimize the hyperparameters of the model according to the level of the evaluation indicators, and repeatedly train until the model converges.

[0101] Step S5, model application. Based on the trained model, generate the corresponding elastographic ultrasound-like image for the target ultrasound B-mode image. This will not be elaborated here.

[0102] Step S6, model conversion and deployment. To optimize the model inference speed and improve portability and security, deploy the TDM-GAN network model introducing AttConv-Transformer detection as a C++ callable type, which is convenient for integration into ultrasound instruments of different brands for efficient inference operations.

[0103] Specifically, in this embodiment, during the training process, the TDM-GAN network model introduced with AttConv-Transformer detection is trained using the open-source deep learning framework Pytorch (2.0.1 GPU version). The weight file in Pth format obtained after training and optimization by the Pytorch framework is converted and deployed as a weight file in Pt format. Since the model weight parameters are saved in Pth format after training the model using the Python language, this format only contains the weight parameters of the model and does not include the defined model structure. However, using Pytorch, the saved Pth format model can be deployed as a Pt format file that contains both the constructed model structure and the optimal weight parameters obtained during training. The model in this format can directly perform inference and prediction on the TDM-GAN network model introduced with AttConv-Transformer detection by loading the Pt file, without the need to redefine the model structure in the ultrasonic device to be transplanted. Finally, the open-source library libtorch can be used to call the Pt file in the C++ development environment to perform the inference process of the detection and generation tasks of elastography-like ultrasound images.

[0104] Embodiment 2

[0105] Figure 2 It shows the structural diagram of the Transformer structure added and the Backbone of the detection network model in the method for generating elastography-like ultrasound images by the TDM-GAN network introduced with AttConv-Transformer detection in the embodiments of the present invention.

[0106] After the ultrasonic B-mode image enters the AttConv-Transformer detection network model, it first enters the Backbone part of the model. This part first adjusts the size of the image and automatically learns the low-level and high-level features of the labeled target area through convolutional neural network (CNN) layers with different structures. For example, the low-level features such as the labeled position of the organ to be examined, the shape of the organ, and the edge structure. As the number and complexity of CNNs increase, more complex feature information of the organ part structure and overall shape will be extracted, enabling the detection model to better distinguish the background (non-organ and non-lesion areas) from the target area (lesion area).

[0107] Specifically, the Backbone part of the detection network model introduced with the AttConv-Transformer structure is composed of a convolutional layer, Transformer, C2f, SCDown, C2fCIB, SPPF, and PSA.

[0108] Among them, C2f is a feature fusion module, whose purpose is to fuse features at different levels, enhance the model's understanding of features, and at the same time reduce the computational amount and model parameters. After the features extracted by the convolutional layer enter the C2f module, they first pass through a convolutional layer to double the number of channels of the feature map, and then it is split into two parts. One part of the feature map enters the bottleneck structure (formed by a depthwise convolution and a pointwise convolution) to extract the feature channel information, and then it is concatenated with the features of the other part. The concatenated features pass through a convolutional operation to obtain the output feature map.

[0109] Among them, SCDown is a spatial and channel decoupled downsampling operation, whose purpose is to extract features from feature maps of different scales during the downsampling process and at the same time reduce the loss of feature information. SCDown adjusts the number of channels of the input feature map through a 1x1 pointwise convolution, and then uses a 3x3 depthwise convolution for spatial downsampling to obtain the processed feature map.

[0110] Among them, C2fCIB replaces the original first convolutional operation in the C2f structure with a depthwise convolution and a pointwise convolution, enabling effective fusion of spatial and channel features of the feature map, and at the same time greatly reducing the computational cost.

[0111] Among them, SPPF is a fast spatial pyramid pooling operation. This structure enhances the richness of features of the input feature map through a pooling kernel of size 5x5, then concatenates it with the original input feature map, and finally passes through a 3x3 convolutional operation to obtain the final input. This operation increases the spatial features of different scales.

[0112] Among them, PSA is a self-attention mechanism. This structure divides the input feature map into two parts. One part extracts the feature information of different subspaces through the multi-head self-attention module and the feed-forward network, and at the same time performs non-linear transformation and integration of feature information on the features. Then it is concatenated with the other part of the input features, and a convolutional operation is performed for feature fusion. This operation enhances the model's understanding of global information.

[0113] Furthermore, after Transformer is added to the second and third convolutional layers in the Backbone part, its structure is to divide the input feature map into F 1 and F 2 , F 1 enters the multi-head attention mechanism after batch normalization (Batch Normalization, BN), and then is concatenated with the feature F 2 . The concatenated feature F 3 , F 3 is also divided into F 3-1 and F 3-2 , F 3-1After passing through the BN layer, it is connected to a Multilayer Perceptron (MLP) structure composed of multiple neurons, and its output is concatenated with F 3-2 to obtain the output feature map by concatenation.

[0114] Among them, the role of adding the Transformer structure after the convolutional layer is as follows: after the convolutional layer captures the local features of the feature edges and textures, the Transformer structure is used to further extract the relationship between the long-distance feature regions and the background, and at the same time fuse the extracted features, so that the model can have both the detailed features in the local features and the coarse features in the global features, thereby improving the richness of the features.

[0115] Among them, the Backbone part saves three output features, namely the feature output by the first Transformer structure Figure 1 (F Backbone-1 ), the feature after the second SCDown operation Figure 2 (F Backbone-2 ) and the output feature of the Backbone Figure 3 (F Backbone-output ). The above three features are used for upsampling in the Neck structure, feature fusion with the C2f structure, and input features of the Neck respectively.

[0116] Embodiment 3

[0117] Figure 3 Shows the self-attention mechanism structure added in the generation method of the TDM-GAN network for class elastic ultrasound images introducing AttConv-Transformer detection and the Neck structure diagram of the detection network model.

[0118] Specifically, the feature Figure 3 (F Backbone-output ) obtained after passing the image through the Backbone structure is used as the input of the Neck structure. The role of this part is to adjust the scale and dimension of the input feature map through upsampling, C2f, convolution, self-attention mechanism, SCDown and C2fCIB and channel concatenation operations on the feature information of different levels and scales extracted by the Backbone part, so as to perform feature fusion and enhancement, so that the feature map passing through the Neck structure has both resolution feature maps of different sizes and rich characteristic information, enabling the detection model to accurately detect target regions of different sizes.

[0119] Among them, the self-attention mechanism is added after the convolutional operation of the Neck structure, and the feature output by the convolutional layer is divided into four feature maps F Conv2Att-1 、F Conv2Att-2 、FConv2Att-3 and F Conv2Att-4 , where the three feature maps are multiplied by the α, β, and γ matrices that control the importance of the attention regions, and then simultaneously pass through a 1x1 convolutional kernel to obtain the feature map F Att-α 、F Att-β and F Att-γ , F Att-α and F Att-β After splicing, the Softmax activation function is used to enhance the non-linear expression ability of the feature map, and then it is spliced and fused with the F Att-γ feature map to obtain F Att-s , this operation retains the global features of the original feature map and also enhances the model's attention to the target region. Finally, F Att-s passes through a 1x1 convolutional kernel and is then spliced with the input feature F Conv2Att-4 to complete one self-attention mechanism operation.

[0120] Furthermore, the features output by the self-attention mechanism pass through C2f and SCDown and are then channel-spliced with the input F Backbone-output feature to complete the feature fusion and enhancement operation of the Neck structure.

[0121] Example 4

[0122] Figure 4 shows the structural diagram of the network model of the method for generating class elastography ultrasound images by the TDM-GAN network with AttConv-Transformer detection introduced in the embodiments of the present invention.

[0123] Specifically, after the AttConv-Transformer detection network model and the TDM-GAN class elastography ultrasound image generation network model are independently trained, the detection model and the generation model are spliced into an end-to-end network model. After inputting an ultrasound B-mode image to be detected, it first enters the AttConv-Transformer detection network model. If no lesion area is detected, it does not enter the generation network of the TDM-GAN network model, and the character "no lesion detected" is output to end the prediction; if the detection model detects a lesion area in the image, after determining the organ type of the lesion in the image, the XX-Rads score and the lesion location of the organ are detected. At the same time, the age, gender, ultrasound instrument brand and model, probe model of the patient in the Dicom format image, and section information during the examination, etc. are combined for feature encoding to obtain the feature vector f 1 .

[0124] Furthermore, after detecting a lesion in the AttConv-Transformer detection network model, an image img containing only the lesion area and the surrounding tissues is obtained through cropping Us , imgUs It serves as the input part of the generation network of the TDM-GAN network model.

[0125] Among them, img Us First, it passes through a convolutional layer for primary image semantic feature extraction, and then enters the Transformer structure with the same structure as in the Backbone to process img Us The features of the lesion and surrounding tissue regions in the image are enhanced, and at the same time, it is concatenated with the original feature map to improve the feature richness in the feature map. Then, it passes through another convolutional layer and Transformer structure, and combines with the flattening operation to obtain the feature map vector f Us ,

[0126] Subsequently, for the feature map vector f Us and the multi-modal information vector f of organs, XX-Rads, device information, etc. 1 are concatenated to obtain the feature vector after multi-modal information fusion. Then, the feature map is restored to the feature map img Us of the same size as img through convolutional layers and Transformer operations Gen-Us-F .

[0127] Among them, during the encoding and decoding process of img Us The dense connection method is used to directly connect each layer to all the previous layers. This connection method strengthens the transmission and utilization of features, enabling feature information to fully propagate to the next layer, and each layer will use the feature information of all the previous layers.

[0128] Among them, img Gen-Us-F Controls the pixel values of the R, G, and B channel images through three hyperparameters θ, σ, and ε, and then obtains the pseudo-elastic ultrasound image img Gen-Elast through the way of channel concatenation.

[0129] Among them, in the TDM-GAN network model, the discriminant network is used to judge the real ultrasound elasticity image img Real-Elast and the pseudo-elastic ultrasound image img Gen-Elast , and judges the authenticity, color space, and color consistency of the two through the proposed L Color loss function. The weight parameters in the network are iteratively updated through backpropagation to complete the training and prediction process.

[0130] Example 5

[0131] Figure 5 This is the detection and generation graph display in the method for generating pseudo-elastic ultrasound images by the TDM-GAN network with AttConv-Transformer detection introduced in the embodiments of the present invention.

[0132] In the effect display diagram, the examination site of patient 1A is the thyroid gland. The patient is female, 52 years old. The diagnosis result of the lesion is in the upper middle part of the left lobe, heterogeneous hypoecho, thyroid bilateral lobe nodules (ACR Ti-Rads classification category 3), and the elastography score is 3 points. The examination site of patient 2A is the ovary. The patient is female, 35 years old. The diagnosis result of the lesion is a cystic-solid mass, about 57mm * 47mm in size, with unclear boundary, uneven internal echo, poor sound transmission in the cystic part, and CDFI shows a blood flow signal level 3.

[0133] Further, in the effect diagram of the generation model, 1B is the thyroid lesion, 2B is the real ultrasound elastogram, and 3B is the elastogram-like ultrasound image generated by the TDM-GAN network with AttConv-Transformer detection introduced; 1C is the ovarian lesion, 2C is the real ultrasound elastogram, and 3C is the elastogram-like ultrasound image generated by the TDM-GAN network with AttConv-Transformer detection introduced. By comparing the real ultrasound elastogram and the elastogram-like ultrasound image, it is found that in the elastic region part of the three-channel color, the TDM-GAN network with AttConv-Transformer detection introduced can effectively generate elastic images. However, due to the imbalance of organ data during the training process, there is a small part of blurring at the boundary of different elastic regions in the generated elastogram-like ultrasound images. Subsequently, the model will be continuously optimized in multi-center, multi-instrument, and large cohort data to make the generated images more similar to the real ultrasound elastograms.

[0134] For the specific calculation of a method for generating elastogram-like ultrasound images by a TDM-GAN network with AttConv-Transformer detection introduced, reference can be made to the description of data integration, feature engineering, model construction, model training, model application, model conversion, and deployment in the invention method in the above text, which will not be elaborated here. All or part of the above method for generating elastogram-like ultrasound images by the TDM-GAN network with AttConv-Transformer detection introduced can be implemented through software to facilitate the calculation of the corresponding operations in each step.

[0135] The process embodiments described above are merely illustrative. The organs, lesions, scoring detection models, or elastogram-like ultrasound image generation networks can also be of other different structures. The loss function, training method, and deployment method during the training process can or may not be the L Color loss, independent training, and C++ deployment methods proposed in this article. One can choose some of them according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.

[0136] Through the description of the above embodiments, those skilled in the art can clearly understand each embodiment, and the proposed method can be implemented by means of different languages and deep learning frameworks. Based on such an understanding, the essence of the above technical solution, or rather the contribution made to the prior art, can be presented in the form of software or a Web full stack. The model can be embedded in the software or deployed in a Web network database, such as MySQL, SQL Server, Oracle, DB2, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0137] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that it is still possible to modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for generating elastic ultrasound images using a TDM-GAN network with AttConv-Transformer detection, characterized in that: The steps include: S1. Data integration: Screening patient data that undergo organ ultrasound examination, wherein the patient data at least includes ultrasound B-mode images, corresponding ultrasound elasticity images, and diagnostic reports, wherein the ultrasound B-mode images and corresponding ultrasound elasticity images are both Dicom format images, and at least include the patient's age, gender, examined organ, ultrasound instrument brand and model, and probe model information; the diagnostic report at least includes the XX-Rads score of the corresponding organ, lesion location, and lesion section information; S2. Feature Engineering: Obtain the lesion area under the ultrasound B-mode image drawn by the physician, and annotate the examined organs and lesion information according to the diagnosis report; encode continuous data according to data characteristics, perform one-hot encoding on discrete data, and construct multimodal data feature groups; S3. Model construction: Design a detection network model that introduces the AttConv-Transformer structure, and combine it with a TDM-GAN network model controlled by multimodal data, and design loss functions that are respectively applicable to the detection network model and the TDM-GAN network model, wherein the TDM-GAN network model includes a generation network and a discrimination network; S4. Model training: Use the transfer learning method to train the detection network model that introduces the AttConv-Transformer structure, and use the ultrasound B-mode image and multimodal data feature group to train the TDM-GAN network model; After independently training the two models, they are connected in series and combined with fine-tuning to optimize the performance of the joint model; S5. Model application: Generate corresponding elastic-like ultrasound images from the target ultrasound B-model image based on the trained model.

2. The generation method according to claim 1, characterized in that: Also includes the steps: S6. Model conversion and deployment: The TDM-GAN network model introduced with AttConv-Transformer detection is deployed as a C++ callable type to facilitate integration into ultrasound instruments of different brands for efficient reasoning operations.

3. The generation method according to claim 1, characterized in that: In step S1, the screening of patient data undergoing organ ultrasound examination also includes ultrasound B-model image data and corresponding elastic imaging data included in the patient's ultrasound screening, physical examination, routine examination, and reexamination, and the included organs include one or more of the thyroid, breast, liver, ovary, prostate, and kidney; the lesion location is the orientation of the current organ to which the lesion belongs; the lesion section information includes a cross section, a longitudinal section, an oblique section or a coronal section; the probe model includes a convex array, a linear array, a phased array or a micro-convex array probe.

4. The generation method according to claim 1, characterized in that: In step S2, the step of obtaining the lesion area existing in the ultrasound B-model image outlined by the physician, and marking the examined organ and lesion information according to the diagnosis report, includes: 1) Preprocess the image: a. Desensitize the image data: remove the patient's name, medical record number, ID number, contact information, and image number in the Dicom format image; at least retain the age, gender, examination organ, ultrasound instrument brand and model, and probe model information; b. Image format conversion: Save Dicom format images as high-quality TIFF images through lossless conversion; 2) Obtain the physician's outlining information: Obtain the physician's outlining information of the examined organs and lesion areas under the ultrasound B-model image in TIFF format. The outlining label types include four categories, namely, whether there is a lesion, organ, lesion location, and XX-Rads score of the corresponding organ.

5. The generation method according to claim 1, characterized in that: In step S3, the detection network model with the AttConv-Transformer structure is introduced with YoloV10 as the base model, the Transformer structure is added to the Backbone part of YoloV10, and the self-attention mechanism structure is added to the Neck part of YoloV10.

6. The generation method according to claim 5, characterized in that: The loss function used by the detection network model introducing the AttConv-Transformer structure is L Att-Trs , by the classification loss L Focal , bounding box loss L GIoU , similarity L1 penalty L Dice_L1 The loss function is composed of the following expressions: L Att-Trs =L Focal +L GIoU +L Dice_L1 Among them, L Focal To distinguish benign and malignant lesions and different organ locations, C_num is the number of detection categories, α Focal is the balancing factor hyperparameter, which is used to adjust the weights between different categories. c is the predicted probability of belonging to category c, γ is the tuning factor hyperparameter, which is used to adjust the detection weights between micro-lesions, normal lesions and no lesions. Its expression is: Among them, L GIoU is the intersection area of ​​the detected position boxes of different categories, G1∩G2 is the intersection area of ​​two detection boxes of a certain category, and G1∪G2 is the union area between the two. min is the minimum circumscribed rectangular area of ​​the two rectangular frames G1 and G2, and its expression is: Among them, L Dice_L1 In order to compare and measure the similarity between the real target area box and the predicted target area box, improve the detection rate of small lesions, and enable the model to better learn the location and shape information of the target, G1 is the real label box of the lesion or organ annotation information drawn by the physician, G2 is the predicted box detected in the model training stage, X and Y correspond to the coordinate values ​​of the center coordinates of the rectangular box, W and H are the width and height of the rectangular box, the subscripts true and pred are the rectangular box drawn by the physician and the rectangular box predicted by the model, γ and δ are the weight coefficients automatically adjusted during training, and their expressions are:

7. The generation method according to claim 1, characterized in that: In step S3, a Transformer structure is added to the encoding part of the generation network of the TDM-GAN network model. The TDM-GAN network model adds a joint three-channel color loss function L based on the traditional binary cross entropy loss of the GAN network. Color , L consistency is the color consistency loss, L distribution is the color distribution loss, L Discrimination The discriminant loss to distinguish the generated elastic-like and real ultrasound elastic images is expressed as: L Color =L Consistency +L Distribution +L Discrimination Among them, L Donsistency The color consistency loss function is to ensure that the generated elastic ultrasound image has color consistency. The Euclidean distance is used to calculate the difference between two adjacent pixels. Fake Elastic ultrasound image generated by TDM-GAN, G Fake (i, j) is the pixel value at the point i, j in the quasi-elastic ultrasound image, and its expression is: Among them, L Distribution The color distribution loss function is to calculate the distribution difference between the quasi-elastic and real elastic ultrasound images, which can help the generative network generate quasi-elastic ultrasound images with the same color distribution as the real ultrasound elastic images. KL is the Kullback-Leibler divergence, which is used to calculate the probability distribution difference between the two images. P(G Fake ) is the color probability distribution function of the elastic ultrasound image, P(G Real ) is the color distribution function of the real ultrasound elastic image, and its expression is: L Distribution =KL(P(G Fake ),P(G Real )) Among them, L Discrimination is the discriminant loss function, which is used to judge the quasi-elastic and real elastic image data generated by the generative network, Y G For the model to determine whether it is a real elastic image, Y G ∈{0,1}, P is the probability of being predicted as a true elastic image, and its expression is: L Discrimination =-[Y G LogP+(1-Y G )Log(1-P)]。 8. The generation method according to claim 1, characterized in that: In step S4, the performance of the joint model optimized by fine-tuning includes: Elastic-like ultrasound images were generated in the test set and evaluated according to evaluation indicators. The hyperparameters of the TDM-GAN network introduced into the AttConv-Transformer detection were fine-tuned and optimized according to the evaluation indicators, and the training was repeated until the model converged. The evaluation indicators included peak signal-to-noise ratio, structural similarity index and physician evaluation index.

9. The generation method according to claim 2, characterized in that: In step S6, the TDM-GAN network model introduced with AttConv-Transformer detection is deployed as a C++ callable type, including converting the Pth format weight file obtained after training and optimization with the Pytorch framework and deploying it as a Pt format weight file, so that the C++ language can call and reason about the model, thereby improving the portability of the model.

10. A system for generating elastic ultrasound images using a TDM-GAN network with AttConv-Transformer detection, used to implement the generation method according to any one of claims 1 to 9, characterized in that: include: A data integration module is configured to screen patient data undergoing organ ultrasound examination, wherein the patient data at least includes an ultrasound B-mode image, a corresponding ultrasound elasticity image, and a diagnosis report, wherein the ultrasound B-mode image and the corresponding ultrasound elasticity image are both Dicom format images, and at least include the patient's age, gender, examined organ, ultrasound instrument brand and model, and probe model information, and the diagnosis report at least includes the XX-Rads score of the corresponding organ, the location of the lesion, and the section information of the lesion; The feature engineering module is configured to obtain the lesion area existing in the ultrasound B-mode image outlined by the physician, and to annotate the examined organ and lesion information according to the diagnosis report; encode the continuous data according to the data characteristics, perform one-hot encoding on the discrete data, and construct a multimodal data feature group; A model building module is configured to design a detection network model that introduces an AttConv-Transformer structure and a TDM-GAN network model controlled by multimodal data, and to design loss functions respectively applicable to the detection network model and the TDM-GAN network model, wherein the TDM-GAN network model includes a generation network and a discrimination network; The model training module is configured to use the transfer learning method to train the detection network model that introduces the AttConv-Transformer structure, and to train the TDM-GAN network model using ultrasound B-mode images and multimodal data feature groups; after independently training the two models, they are concatenated in series and combined with fine-tuning to optimize the performance of the joint model; A model application module is configured to generate a corresponding elastic-like ultrasound image from a target ultrasound B-model image based on the trained model; The model conversion and deployment module is configured to deploy the TDM-GAN network model introduced with AttConv-Transformer detection as a C++ callable type, making it easier to integrate it into ultrasound instruments of different brands for efficient reasoning operations.