Ultrasonic image artificial intelligence auxiliary judgment method and system based on TimeSform
By using the feature fusion and visualization assistance of the TimeSformer model, the problems of low efficiency and high subjectivity in the ultrasound imaging diagnosis of gout are solved, and highly accurate and interpretable automatic auxiliary judgment is achieved.
Patent Information
- Application Number
- CN202511178958.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-11-14
AI Technical Summary
Existing technologies for the ultrasound imaging diagnosis of gout suffer from problems such as low efficiency, high subjectivity, lack of visualization aids, and failure to effectively integrate features from multiple frames of images.
The TimeSformer model is adopted, and an auxiliary gout diagnosis model is constructed by segmenting image frames and image blocks, spatial attention layers and temporal attention layers, fully connected layers and Softmax function layers, so as to realize feature fusion of multi-frame images and visual auxiliary diagnosis.
It improves the accuracy and interpretability of automated ultrasound imaging for gout diagnosis, enhances the reliability and consistency of doctors' diagnoses, and reduces the misdiagnosis rate.
Smart Images

Figure CN120953619A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing of ultrasound images, and more specifically, to an AI-assisted judgment method and system for ultrasound images based on TimeSformer. Background Technology
[0002] Gouty arthritis is a common type of osteoarthritis, primarily caused by the deposition of monosodium urate (MSU) crystals in and around the joints. Currently, the global incidence of gouty arthritis is rapidly increasing, making the development of reliable diagnostic methods particularly urgent.
[0003] In clinical diagnostic practice, the diagnosis of gout mainly relies on the following criteria: first, detecting the serum uric acid level in suspected patients; second, performing joint aspiration and synovial fluid microscopy; and third, utilizing imaging techniques such as ultrasound. Since many clinicians do not routinely perform synovial fluid analysis, ultrasound examination, due to its convenience, has become a widely used diagnostic tool. Related studies have shown that ultrasound can effectively identify characteristic manifestations of gout, such as the double-track sign and tophi, and these markers are widely used in clinical diagnosis.
[0004] However, with the increasing number of gout patients requiring early screening and diagnosis, several problems have emerged in clinical diagnosis. Manual analysis and interpretation of ultrasound images by physicians is not only inefficient but also inherently subjective. Although a few studies have explored the application of deep learning technology to the automated diagnostic assessment of gout using ultrasound images, these methods still face challenges: firstly, how to utilize sequences composed of a large number of ultrasound images; and secondly, how to reliably detect subtle pathological features.
[0005] A search revealed Chinese patent number 202110314834.X, which discloses an ultrasound-based identification device. The device acquires ultrasound images of the user's knee, ankle, and first metatarsophalangeal joint via an acquisition module. The identification module uses N training models, each corresponding to a specific body part, to identify characteristic symptoms. A scoring module calculates a score based on the characteristic symptoms and preset corresponding scores. An output module outputs the detection result of gout or asymptomatic hyperuricemia based on the score and a preset score threshold, and can also output the problem level. The shortcomings of this approach are: it lacks an image preprocessing step, making it susceptible to interference from noise in the original image; it uses multiple independent training models to identify features without feature fusion, resulting in low efficiency; and it lacks visualization aids, leading to weak interpretability. Summary of the Invention
[0006] In view of the deficiencies in the prior art, the purpose of this application is to provide an artificial intelligence-assisted judgment method and system for ultrasound images based on TimeSformer.
[0007] The first aspect of this application provides an AI-assisted judgment method for ultrasound images based on TimeSformer, comprising:
[0008] A dataset of sample ultrasound images;
[0009] The sample image dataset is read, denoised, segmented, and normalized.
[0010] An auxiliary gout diagnosis model is established, including a segmentation image frame and image block layer, a parallel spatial attention layer and temporal attention layer, a fully connected layer, and a Softmax function layer. The segmentation image frame and image block layer divides the input multi-frame video into multi-frame images and segments each frame image into image blocks. The spatial attention layer and temporal attention layer focus on the relationship between image blocks in each frame image and the relationship between image blocks at the same position in different frames, respectively. The fully connected layer and Softmax function layer convert the relationship between image blocks in each frame image and the relationship between image blocks at the same position in different frames into a probability distribution for auxiliary diagnosis.
[0011] The auxiliary gout diagnosis model is trained based on the sample image dataset.
[0012] The ultrasound images to be tested are read, denoised, segmented, and normalized before being input into the trained auxiliary gout diagnosis model to obtain the auxiliary gout diagnosis results.
[0013] Optionally, the sample image dataset for acquiring ultrasound images includes:
[0014] Ultrasound images of the first metatarsophalangeal joint were acquired from gout patients and healthy individuals, including three locations: the dorsum of the foot, the medial side of the foot, and the sole of the foot. Five images were saved for each location.
[0015] All collected ultrasound images of the first metatarsophalangeal joint were labeled and divided into two categories: gout patients and healthy individuals.
[0016] All of the above-mentioned ultrasound images of the first metatarsophalangeal joint and their corresponding annotations constitute the sample image dataset.
[0017] Optionally, the process of reading, denoising, segmenting, and normalizing the sample image dataset includes:
[0018] The CV algorithm is used to read the ultrasound images in the sample image dataset into an m*n*15 matrix data, where m*n represents the size of each image as m*n pixels, m and n are both integers greater than 1, and 15 represents a 15-frame video consisting of 5 images of each of the 3 parts.
[0019] Gaussian filtering is used to remove noise generated during ultrasound image acquisition and instrument signal processing from the matrix data.
[0020] For the noise-removed matrix data, perform region segmentation, execute continuous "close-open-close" morphological operations to generate a binary mask, locate the strong echo region, and extract a central rectangular region of 700×320 pixels from the strong echo region.
[0021] The pixel values of the central rectangular region are normalized to obtain 15-channel images of the three regions.
[0022] Optionally, the segmented image frame and image block layer divides the 15-channel image into 15 frames, and each frame is divided into image blocks of a fixed size.
[0023] Optionally, the spatial attention layer focuses on the interior of a single frame image and analyzes the relationship between different image blocks in the same frame; the temporal attention layer analyzes the relationship between image blocks at the same position in different frames.
[0024] Optionally, the fully connected layer fuses the relationships between image blocks in each frame and the relationships between image blocks at the same position in different frames to obtain a pathological feature map;
[0025] The Softmax function layer obtains auxiliary judgment results based on the pathological feature map.
[0026] Optionally, it also includes:
[0027] The PyTorch framework was used to extract features from the spatial attention layer and temporal attention layer in the auxiliary gout diagnosis model.
[0028] The features are input into the GradCAM function to generate a gradcam image;
[0029] The gradcam image is overlaid with the ultrasound image to be tested to obtain a heat map, which is used to assist in the judgment.
[0030] A second aspect of this application provides an AI-assisted judgment system for ultrasound images based on TimeSformer, comprising:
[0031] Acquisition module: Acquires sample image datasets of ultrasound images;
[0032] Preprocessing module: Reads, denoises, segments, and normalizes the sample image dataset;
[0033] Model building module: A model for assisting gout diagnosis is established, including image frame and image patch segmentation layers, parallel spatial attention layers and temporal attention layers, fully connected layers and softmax function layers. The image frame and image patch segmentation layers divide the input multi-frame video into multi-frame images and segment each frame image into image patches. The spatial attention layer and temporal attention layer focus on the relationship between image patches in each frame image and the relationship between image patches at the same position in different frames, respectively. The fully connected layer and softmax function layer convert the relationship between image patches in each frame image and the relationship between image patches at the same position in different frames into a probability distribution for assisting diagnosis.
[0034] Training module: Based on the sample image dataset, train the auxiliary gout diagnosis model;
[0035] Test module: The ultrasound images to be tested are read, denoised, segmented and normalized before being input into the training module.
[0036] A third aspect of this application provides a terminal, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it can be used to execute the aforementioned AI-assisted judgment method for ultrasound images based on multi-scale features, or to run the aforementioned AI-assisted judgment system for ultrasound images based on multi-scale features.
[0037] A fourth aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, can be used to perform the aforementioned AI-assisted judgment method for ultrasound images based on multi-scale features, or to run the aforementioned AI-assisted judgment system for ultrasound images based on multi-scale features.
[0038] The ultrasound imaging AI-assisted gout diagnosis method provided in this application takes into account the characteristics of automatic gout ultrasound diagnosis: it needs to mine the spatial correlation between image blocks and the temporal correlation between different frames from multiple ultrasound images. It constructs an auxiliary gout diagnosis model, whose image frame and image block segmentation layer can divide the input multi-frame video into multi-frame images and segment them into image blocks; its parallel spatial attention layer and temporal attention layer can respectively focus on the relationship between image blocks in each frame and the relationship between image blocks in the same position in different frames; the fully connected layer and the Softmax function layer convert these relationships into probability distributions for auxiliary diagnosis. The entire model can realize automatic auxiliary diagnosis of gout ultrasound images.
[0039] Other technical effects resulting from the additional features will be further illustrated in the corresponding embodiments. Attached Figure Description
[0040] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0041] Figure 1 This is a flowchart illustrating an ultrasound imaging AI-assisted gout diagnosis method according to an exemplary embodiment;
[0042] Figure 2 This is a schematic diagram of the structure of an auxiliary gout diagnosis model according to an exemplary embodiment;
[0043] Figure 3 This is a structural diagram of an ultrasound imaging AI-assisted gout diagnosis system according to an exemplary embodiment;
[0044] Figure 4 This is a visualization heatmap (Grad-Cam diagram) illustrated according to an exemplary embodiment. Detailed Implementation
[0045] The present application will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present application, but do not limit the present application in any way. It should be noted that those skilled in the art can make several modifications and improvements without departing from the concept of the present application, and these all fall within the protection scope of the present application. Parts not described in detail in the following embodiments can be implemented using existing technology.
[0046] Given the large number of existing ultrasound images, there is a need to utilize sequences composed of these images to reliably detect subtle pathological features. Based on this need, embodiments of this application provide an AI-assisted ultrasound image analysis method based on TimeSformer to address the aforementioned problems.
[0047] Reference Figure 1 As shown in one embodiment of this application, an AI-assisted judgment method for ultrasound images based on TimeSformer includes the following steps:
[0048] S100, a sample image dataset for acquiring ultrasound images;
[0049] S200 performs reading, denoising, segmentation, and normalization on the sample image dataset;
[0050] S300, establish an auxiliary gout diagnosis model, which includes, in sequence, a segmented image frame and image block layer, a parallel spatial attention layer and a temporal attention layer, a fully connected layer, and a Softmax function layer. The segmented image frame and image block layer divides the input ultrasound image into multiple frames and divides each frame into image blocks. The spatial attention and temporal attention focus on the relationship between image blocks in each frame and the relationship between image blocks at the same position in different frames, respectively. The fully connected layer and the Softmax function layer convert the relationship between image blocks in each frame and the relationship between image blocks at the same position in different frames into a probability distribution for auxiliary diagnosis.
[0051] S400, based on a sample image dataset, trains an auxiliary gout diagnosis model;
[0052] The S500 reads, denoises, segments, and normalizes the ultrasound images to be tested, then inputs them into a trained auxiliary gout diagnosis model to obtain the auxiliary gout diagnosis results.
[0053] The embodiments described above in this application construct a deep learning model based on segmented image frame and image block layers, spatial attention and temporal attention layers, fully connected layers, and Softmax function layers. This model enables automatic auxiliary judgment of gout ultrasound images, achieving high classification accuracy and strong generalization to different center data, thus realizing automatic auxiliary judgment of gout diseases based on ultrasound images.
[0054] To provide raw data for all subsequent steps, a sample image dataset is needed. To make the obtained image dataset more comprehensive, in some specific embodiments of this application, step S100 may employ the following steps:
[0055] S101, acquire multi-section ultrasound image data of the first metatarsophalangeal joint.
[0056] Specifically, the data is collected by the hospital and includes two categories: ultrasound images of the first metatarsophalangeal joint from patients with gout and ultrasound images of the first metatarsophalangeal joint from healthy individuals. Furthermore, the multi-sectional ultrasound images of the first metatarsophalangeal joint should be saved as multi-sectional images, divided into three parts: the dorsum of the foot, the medial side of the foot, and the sole of the foot.
[0057] S102 involves a professional ultrasound physician diagnosing and labeling the acquired multi-sectional ultrasound images of the first metatarsophalangeal joint.
[0058] Specifically, data on people with gout can be placed in one folder, and data on healthy people can be placed in another folder.
[0059] The embodiments described above in this application acquire multi-sectional ultrasound images of the first metatarsophalangeal joint, which are then labeled by a professional physician. This provides direct and crucial raw data for all subsequent steps, including sample image preprocessing, model algorithm construction, and model training. The acquired images include those from both gout patients and healthy individuals, and cover multi-sectional images of the dorsum, medial, and plantar surfaces of the foot. This comprehensive data coverage allows the model to be exposed to a rich variety of samples during training, learning the differences in images of different locations between gout patients and healthy individuals. This improves the model's adaptability to different situations and the universality of its judgments, reducing model bias caused by limited data.
[0060] The ultrasound images acquired in step S100, compared to conventional images, exhibit various noises and artifacts due to inherent limitations in the acquisition instrument and ultrasound imaging itself. Furthermore, foot ultrasound images contain numerous elements irrelevant to gout diagnosis, such as text and blank areas. Given these characteristics, to extract more accurate key regions from the ultrasound images, Gaussian filtering can be used to remove noise, and morphological methods can be employed to obtain key regions relevant to gout identification. In some specific embodiments of this application, step S200, which involves reading, denoising, segmenting, and normalizing the sample image dataset, can employ the following steps:
[0061] S201, Image Input: Using a Computer Vision algorithm, the image data obtained in step S100 is read into an m×n×15 matrix. Here, m×n represents the pixel size of a single image, and 15 corresponds to the image composition of each case (3 body parts, each containing 5 images).
[0062] S202, Noise Removal: Ultrasonic images are prone to noise during acquisition and instrument signal processing. This noise can interfere with subsequent region segmentation and feature extraction. If noise is not removed beforehand, it may be misidentified as valid information, thus affecting the accuracy of key region segmentation. Therefore, Gaussian filtering is used to suppress noise to obtain clearer, cleaner images, providing a reliable foundation for subsequent processing.
[0063] S203, Effective Identification Region Segmentation: On the clear image after noise removal, perform continuous "close-open-close" morphological operations (using a 6×6 kernel matrix). This operation can accurately locate strong echo regions with identification value and extract a central rectangular region of 700×320 pixels from them.
[0064] Specifically, morphological opening operation refers to erosion followed by dilation, morphological closing operation refers to dilation followed by erosion, and "closing-opening-closing" morphological operation refers to performing closing operation, opening operation, and a second closing operation in sequence (using 6×6 cores).
[0065] S204, Image pixel value normalization: Adjust the pixel values of the segmented effective recognition region image.
[0066] Specifically: calculate the average value of the image pixels, subtract the average value from each pixel value, and then divide the result by the maximum value of the calculation.
[0067] In the above embodiments of this application, S201 converts the image into a computer-processable matrix form, providing a structured data foundation for subsequent processing; S202 reduces noise interference, improves image quality, and ensures the accuracy of subsequent region segmentation; S203 accurately locates key recognition regions, eliminates invalid information, and provides the model with high-quality analysis objects (central rectangular region); S204 unifies the pixel value range, eliminates interference from invalid regions, enhances data consistency, and helps the model learn efficiently.
[0068] To obtain more accurate judgment results, in some specific embodiments of this application, in step S300, an auxiliary gout judgment model is established, such as... Figure 2 As shown, the process sequentially includes image frame and image patch segmentation layers, parallel spatial attention layers and temporal attention layers, fully connected layers, and a softmax function layer. The steps involved are as follows:
[0069] S301 divides the input video image into a 15-frame image sequence by segmenting the image frame and image block layers, and divides each frame image into image blocks.
[0070] Specifically, after preprocessing, each patient's image data exists in 15-channel format (corresponding to the original information of 3 sites × 5 images).
[0071] The image frame and image patch layer divides the 15-channel image into 15 frames, and each frame is further divided into 20×20 image patches. This segmentation allows the subsequent model to learn the relationship between image patches to extract gout image features.
[0072] S302, the spatial attention and temporal attention layers, respectively focus on the relationships between image patches in each frame and the relationships between image patches at the same position in different frames.
[0073] Specifically, the original ultrasound image is first segmented into frames and image patches. The image patches are converted into low-dimensional feature vectors through linear embedding, and then spatial (the coordinates of the patch within the frame) and temporal (the sequence number of the frame) information is injected through positional encoding to form a feature sequence with positional information. The preprocessed feature sequence is input into the Transformer basic structure, whose core self-attention mechanism supports two specialized sub-layers:
[0074] Spatial attention layer: Under the Transformer framework, it focuses on the correlation between different image patches within a single frame, such as the spatial distribution of lesion areas;
[0075] Temporal attention layer: Also based on the Transformer mechanism, it focuses on the temporal changes of image patches at the same location in different frames, such as the spatial location of lesion areas.
[0076] In step S303, the relationships between image patches in each frame obtained in step S302 and the relationships between image patches at the same position in different frames are input into a fully connected layer for fusion to obtain fused features. These fused features are the subtle pathological features.
[0077] S304, the feature input function Softmax layer is fused to obtain the judgment result.
[0078] Specifically, the result is a judgment to assist doctors in diagnosing whether someone has gout, and it can also output a probability of having gout.
[0079] The embodiments described above in this application first segment the multi-channel image by dividing the image frame and image block, then focus on key information through spatial attention and temporal attention, and finally fuse comprehensive features through a fully connected layer, gradually building a complete algorithm framework for the model from "processing raw data" to "generating high-quality judgment criteria".
[0080] To optimize the parameters of the above model, in some specific embodiments of this application, S400, training the auxiliary gout diagnosis model based on the sample image dataset can be carried out by the following steps:
[0081] S401, set model parameters.
[0082] Specifically, image data of two cases were input in each batch, with the number of categories set to 2 (gout patients and healthy individuals), and the Adam optimizer was used for training, with the initial learning rate set to 0.001.
[0083] The 2D matrix output by the model is converted into probability values in the 0-1 interval using the softmax function. The sum of the two probability values is 1, which corresponds to the prediction probabilities of the two categories respectively.
[0084] S402 sets the cross-entropy loss function as the loss function. This function calculates a quantitative index by taking the two probability values of the output; the smaller the index value, the better the model's prediction performance.
[0085] During training, the S403, in conjunction with the Adam optimizer, automatically adjusts the model parameters after each batch of input data to reduce the loss function value and optimize the model. After 50 rounds of training, the trained model is obtained.
[0086] After training according to the above embodiments of this application, the final trained model has a more stable and accurate ability to assist in the diagnosis of gout, which can effectively assist in the clinical diagnosis of gouty arthritis.
[0087] The above embodiments enable automatic assisted diagnosis of gout using ultrasound images. In addition, a thermogram can be provided to assist doctors in diagnosis. In some specific embodiments of this application, step S600 further includes establishing a thermogram, with the specific steps as follows:
[0088] S601, using the PyTorch framework to extract features of the spatial attention layer and temporal attention layer in the auxiliary gout diagnosis model;
[0089] S602, Input the features into the GradCAM function to generate the gradcam image;
[0090] S603 overlays the gradcam image with the ultrasound image to be tested to obtain a heat map, which is used to assist in the judgment.
[0091] In the embodiments described above, the feature maps of the spatial and temporal attention layers in the gout diagnosis model are extracted using the PyTorch framework. These feature maps are then generated using the GradCAM function and compared with a heatmap obtained by overlaying the original ultrasound image. This allows for a direct visualization of the key areas the model focuses on during diagnosis (such as areas of strong echoes and urate deposition). This enables doctors to clearly see the "focus" of the model's diagnosis, transforming the model's abstract decision-making process into visualized image information and enhancing the interpretability of the model's diagnosis results.
[0092] Meanwhile, heat maps provide doctors with additional diagnostic references, especially in the identification of subtle pathological features, complementing their professional judgment. Doctors can combine the key areas marked on the heat maps to more efficiently verify diagnostic results, reduce misjudgments caused by subjective experience differences, improve the accuracy and consistency of gout ultrasound imaging diagnosis, and further enhance the clinical practical value of the automated auxiliary judgment system.
[0093] Based on the same technical concept, some other embodiments of this application provide an ultrasound imaging AI-assisted gout diagnosis system 100, such as... Figure 3 As shown, it includes:
[0094] Acquisition module 110: Acquires a sample image dataset of ultrasound images;
[0095] Preprocessing module 120: Reads, denoises, segments, and normalizes the sample image dataset;
[0096] Model Module 130: Establishes an auxiliary gout diagnosis model, including segmenting image frames and image blocks, parallel spatial attention layer and temporal attention layer, fully connected layer and Softmax function layer. The segmentation of image frames and image blocks divides the input multi-frame video into multi-frame images, and each frame image is segmented into image blocks. The spatial attention and temporal attention focus on the relationship between image blocks in each frame image and the relationship between image blocks at the same position in different frames, respectively. The fully connected layer and Softmax function layer convert the relationship between image blocks in each frame image and the relationship between image blocks at the same position in different frames into a probability distribution for auxiliary diagnosis.
[0097] Training Module 140: Train an auxiliary gout diagnosis model based on a sample image dataset;
[0098] Test module 150: After reading, denoising, segmenting and normalizing the ultrasound images to be tested, the data is input into the trained auxiliary gout judgment model to obtain the auxiliary gout judgment result.
[0099] The specific implementation techniques of each module / unit in the above examples of this application can be referred to the steps of the ultrasound imaging artificial intelligence-assisted gout diagnosis method in the above embodiments, and will not be repeated here.
[0100] The preferred features in the above embodiments can be used individually in any embodiment, or in any combination thereof, provided they do not conflict with each other. Furthermore, parts not described in detail in the embodiments can be implemented using existing technologies.
[0101] The following examples and comparative examples will be used to further illustrate this application in order to better understand the above-mentioned technical solutions. It should be understood that the following are only some examples and are not intended to limit this application.
[0102] Application Example 1
[0103] The method described in the above embodiments of this application is executed as follows: A Linux 4.4.0 operating system with an x86_64 processor is built, using the PyTorch 1.7.1 framework, Python 3.8 as the programming language, Visual Studio 1.75.1 as the editor, an Intel(R) Xeon(R) CPU E5-2620v4 @ 2.10GHz, an NVIDIA GeForce RTX 3090 GPU, and 24GB of RAM. All programs are implemented using the PyTorch open-source framework. Figure 4 As shown, the heatmap obtained using the Grad-CAM method further enhances the technical effect of assisting doctors in diagnosis by clearly marking areas in ultrasound images related to gout identification (such as areas with strong echoes and urate deposition areas).
[0104] This clear regional labeling allows doctors to quickly locate the core areas on which the model makes its judgments, eliminating the need to search the entire map one by one, significantly improving diagnostic efficiency. At the same time, the accuracy of the labeling provides doctors with more specific references, especially for less experienced physicians, helping them better identify key diagnostic features, reducing the possibility of missed or misdiagnosed cases, effectively enhancing the practical application value of assisted judgment, and enabling a closer collaboration between the automated assisted judgment system and the doctor's diagnosis.
[0105] Based on the same technical concept, in other embodiments of this application, a terminal is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it can be used to execute the above-mentioned ultrasound imaging artificial intelligence-assisted gout diagnosis method, or to run the above-mentioned ultrasound imaging artificial intelligence-assisted gout diagnosis system.
[0106] Based on the same technical concept, in other embodiments of this application, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, it can be used to execute the above-described ultrasound imaging AI-assisted gout diagnosis method, or to run the above-described ultrasound imaging AI-assisted gout diagnosis system.
[0107] Optionally, the memory is used to store programs; the memory may include volatile memory, such as random-access memory (RAM), such as static random-access memory (SRAM), double data rate synchronous dynamic random-access memory (DDRSDRAM), etc.; the memory may also include non-volatile memory, such as flash memory. The memory is used to store computer programs (such as application programs, functional modules, etc. that implement the above methods), computer instructions, etc., and the aforementioned computer programs, computer instructions, etc., can be partitioned and stored in one or more memories. Furthermore, the aforementioned computer programs, computer instructions, data, etc., can be accessed by the processor.
[0108] The aforementioned computer programs, computer instructions, etc., can be stored in partitions within one or more memory locations. Furthermore, the aforementioned computer programs, computer instructions, data, etc., can be accessed by a processor.
[0109] A processor is used to execute a computer program stored in memory to implement the various steps of the methods involved in the above embodiments. For details, please refer to the relevant descriptions in the preceding method embodiments.
[0110] The processor and memory can be separate structures or integrated structures. When the processor and memory are separate structures, they can be coupled together via a bus.
[0111] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0112] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0113] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0114] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0115] The foregoing has described some specific embodiments of this application. It should be understood that this application is not limited to the specific embodiments described above, and those skilled in the art can make various modifications or variations within the scope of the claims, which do not affect the substantive content of this application. The above-described preferred features can be used in any combination without conflict.
Claims
1. A time-former-based artificial intelligence-assisted judgment method for ultrasound images, characterized in that, include: A dataset of sample ultrasound images; The sample image dataset is read, denoised, segmented, and normalized. An auxiliary gout diagnosis model is established, including a segmentation image frame and image block layer, a parallel spatial attention layer and temporal attention layer, a fully connected layer, and a Softmax function layer. The segmentation image frame and image block layer divides the input multi-frame video into multi-frame images and segments each frame image into image blocks. The spatial attention layer and temporal attention layer focus on the relationship between image blocks in each frame image and the relationship between image blocks at the same position in different frames, respectively. The fully connected layer and Softmax function layer convert the relationship between image blocks in each frame image and the relationship between image blocks at the same position in different frames into a probability distribution for auxiliary diagnosis. The auxiliary gout diagnosis model is trained based on the sample image dataset. The ultrasound images to be tested are read, denoised, segmented, and normalized before being input into the trained auxiliary gout diagnosis model to obtain the auxiliary gout diagnosis results.
2. The AI-assisted judgment method for ultrasound images based on TimeSformer according to claim 1, characterized in that, The sample image dataset for acquiring ultrasound images includes: Ultrasound images of the first metatarsophalangeal joint were acquired from gout patients and healthy individuals, including three locations: the dorsum of the foot, the medial side of the foot, and the sole of the foot. Five images were saved for each location. All collected ultrasound images of the first metatarsophalangeal joint were labeled and divided into two categories: gout patients and healthy individuals. All of the above-mentioned ultrasound images of the first metatarsophalangeal joint and their corresponding annotations constitute the sample image dataset.
3. The AI-assisted judgment method for ultrasound images based on TimeSformer according to claim 1, characterized in that, The process of reading, denoising, segmenting, and normalizing the sample image dataset includes: The CV algorithm is used to read the ultrasound images in the sample image dataset into an m*n*15 matrix data, where m*n represents the size of each image as m*n pixels, m and n are both integers greater than 1, and 15 represents a 15-frame video consisting of 5 images of each of the 3 parts. Gaussian filtering is used to remove noise generated during ultrasound image acquisition and instrument signal processing from the matrix data. For the noise-removed matrix data, perform region segmentation, execute continuous "close-open-close" morphological operations to generate a binary mask, locate the strong echo region, and extract a central rectangular region of 700×320 pixels from the strong echo region. The pixel values of the central rectangular region are normalized to obtain 15-channel images of the three regions.
4. The AI-assisted judgment method for ultrasound images based on TimeSformer according to claim 3, characterized in that, The segmented image frame and image block layer divides the 15-channel image into 15 frames, and then divides each frame into image blocks of a fixed size.
5. The AI-assisted judgment method for ultrasound images based on TimeSformer according to claim 4, characterized in that, The spatial attention layer focuses on the interior of a single frame image and analyzes the relationship between different image blocks within the same frame; the temporal attention layer analyzes the relationship between image blocks at the same position in different frames.
6. The AI-assisted judgment method for ultrasound images based on TimeSformer according to claim 5, characterized in that, The fully connected layer fuses the relationships between image blocks in each frame and the relationships between image blocks at the same position in different frames to obtain a pathological feature map. The Softmax function layer obtains auxiliary judgment results based on the pathological feature map.
7. The AI-assisted judgment method for ultrasound images based on TimeSformer according to claim 1, characterized in that, Also includes: The PyTorch framework was used to extract features from the spatial attention layer and temporal attention layer in the auxiliary gout diagnosis model. The features are input into the GradCAM function to generate a gradcam image; The gradcam image is overlaid with the ultrasound image to be tested to obtain a heat map, which is used to assist in the judgment.
8. An AI-assisted judgment system for ultrasound images based on TimeSformer, characterized in that, include: Acquisition module: Acquires sample image datasets of ultrasound images; Preprocessing module: Reads, denoises, segments, and normalizes the sample image dataset; Model building module: A model for assisting gout diagnosis is established, including image frame and image patch segmentation layers, parallel spatial attention layers and temporal attention layers, fully connected layers and softmax function layers. The image frame and image patch segmentation layers divide the input multi-frame video into multi-frame images and segment each frame image into image patches. The spatial attention layer and temporal attention layer focus on the relationship between image patches in each frame image and the relationship between image patches at the same position in different frames, respectively. The fully connected layer and softmax function layer convert the relationship between image patches in each frame image and the relationship between image patches at the same position in different frames into a probability distribution for assisting diagnosis. Training module: Based on the sample image dataset, train the auxiliary gout diagnosis model; The testing module reads, denoises, segments, and normalizes the ultrasound images to be tested, then inputs them into the trained auxiliary gout diagnosis model to obtain the auxiliary gout diagnosis results.
9. A terminal, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it can be used to perform the method of any one of claims 1-7, or to run the system of claim 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program can be used to perform the methods described in claims 1-7, or to run the system described in claim 8.
Citation Information
Patent Citations
Recognition device based on ultrasonic image
CN115120262A
Cited By
A chronic kidney disease information processing system and method based on a morphological characteristic parameter of uric acid crystals
CN122474306A
A chronic kidney disease information processing system and method based on a morphological characteristic parameter of uric acid crystals
CN122474306B