PDA (Personal Digital Assistant) key area enhanced display method based on deep learning
The TimeSformer model is used to automatically identify and enhance the display of PDA key areas in Doppler ultrasound videos, solving the problem of insufficient utilization of video information in echocardiography examinations and achieving efficient and accurate PDA diagnostic assistance.
Patent Information
- Application Number
- CN202511300977.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-12
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-09-12
AI Technical Summary
Existing echocardiographic examinations have difficulty effectively utilizing the temporal dimension information in video when processing patent ductus arteriosus (PDA) images, resulting in diagnostic accuracy being highly dependent on the operator's experience and image interpretation ability.
The TimeSformer model based on the Transformer architecture is used to process Doppler ultrasound videos. Through section recognition, contrast learning, and enhanced display modules, it automatically identifies key areas of PDA and generates dynamic overlay attention heat maps to assist doctors in diagnosis.
It significantly improves the efficiency and accuracy of PDA diagnosis, provides objective visual guidance, reduces subjective errors, and improves the reliability of clinical diagnosis.
Smart Images

Figure CN120807511A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of medical image analysis and artificial intelligence, in particular to a PDA key area enhanced display method based on deep learning. BACKGROUND
[0002] Patent Ductus Arteriosus (PDA) refers to a pathological state in which the ductus arteriosus fails to close as expected after birth and remains open. PDA is a common cardiovascular malformation in newborns and infants. Currently, echocardiography is the main non-invasive examination method for PDA, and one of the most commonly used sections to observe PDA is the parasternal short-axis view at the level of great vessels. The accuracy is highly dependent on the experience and image interpretation ability of the operator.
[0003] In recent years, with the rapid development of deep learning technology, its application in the field of medical image analysis has become increasingly widespread and in-depth. Traditional computer-aided diagnosis systems often rely on hand-designed features, while deep learning models, especially convolutional neural networks (CNN), can automatically learn high-dimensional and discriminative features from image data. However, echocardiography is usually recorded in video form, containing rich temporal dimension information. Static image analysis methods will ignore these key dynamic information. To solve this problem, models capable of processing video or continuous time series data have emerged. TimeSformer is a time series model based on Transformer architecture, which can effectively capture long-range dependencies in video data through self-attention mechanism. This feature makes it very suitable for analyzing dynamic echocardiography videos. Based on the above, the present application proposes a PDA key area enhanced display method based on deep learning. SUMMARY
[0004] The purpose of the present application is to provide a PDA key area enhanced display method based on deep learning, which can automatically identify Doppler ultrasound videos and enhance the display of PDA key areas to assist doctors in judgment.
[0005] To achieve the above purpose, the present application provides the following scheme:
[0006] A PDA key area enhanced display method based on deep learning, comprising:
[0007] obtaining a target Doppler ultrasound video;
[0008] inputting the target Doppler ultrasound video into a pre-constructed enhanced display model, and outputting an enhanced video segment dynamically superimposed with an attention heat map, wherein the enhanced display model adopts a TimeSformer architecture to identify a parasternal short-axis section in each frame of the video and locate a region, and to generate an enhanced display by weighted fusion.
[0009] Optionally, the enhanced display model comprises:
[0010] a section identification module configured to analyze the target Doppler ultrasound video by using a first model, identify a parasternal short-axis section in each frame, and determine a key frame based on a preset section identification confidence threshold;
[0011] a contrast learning module configured to extract features by calculating similarity scores of a key frame sequence with preset positive samples and negative samples by using a second model, and locate a key region by assigning an attention score;
[0012] an enhanced display module configured to superimpose the key frame and the corresponding key region, and output an enhanced video segment dynamically superimposed with an attention heat map.
[0013] Optionally, the analysis of the target Doppler ultrasound video by using the first model to identify a parasternal short-axis section in each frame comprises:
[0014] ;
[0015] wherein, represents a parameter of the first model , is a classification result output by the first model , to determine whether it is a parasternal short-axis section, is an input video.
[0016] Optionally, the determination of the key frame based on the preset section identification confidence threshold comprises:
[0017] When the first model identifies the parasternal short-axis section, it outputs a confidence score for each frame of the target Doppler ultrasound video. When the confidence score corresponding to a video frame is greater than or equal to the preset section identification confidence threshold, the video frame is determined to be a key frame, and is added to a key frame sequence for enhanced display.
[0018] Optionally, the calculation of the similarity scores of the key frame sequence with the preset positive samples and negative samples comprises:
[0019] The second model maps the key frame sequence into a feature vector, calculates distances between the feature vector and preset positive sample feature centers and negative sample feature centers, and obtains the similarity score, wherein the positive sample is a set of labeled positive patent ductus arteriosus sample videos, and the negative sample is a set of normal sample videos.
[0020] Optionally, the second model adopts a contrast learning strategy for training, in the training process, an anchor sample is set, a set of labeled positive patent ductus arteriosus sample videos is set as positive samples, and a set of normal sample videos is set as negative samples, and the feature vector distances of the positive sample pairs are pulled closer and the feature vector distances of the negative sample pairs are pushed farther apart in a high-dimensional feature space through a contrast learning loss function for optimization.
[0021] Optionally, the key region is located by assigning an attention score, comprising:
[0022] When the second model extracts features by calculating the similarity, an attention score is assigned to each image block in each frame, the key region is determined through the attention score, and an attention map is formed by collecting the attention scores of all image blocks, and an attention heat map is generated by pseudo-color rendering.
[0023] Optionally, the first model and the second model adopt a weight sharing mechanism in the training process, and the same weight parameters are shared in a preset part of the network.
[0024] Optionally, the enhanced video segment dynamically superimposes the attention heat map by setting a joint display control parameter to adjust the heat map display intensity and spatial focus range.
[0025] The beneficial effects of the present application are:
[0026] The present application can automatically identify Doppler ultrasound videos, enhance and display PDA key frames and key regions, and provide precise visual guidance, thereby significantly improving the efficiency and accuracy of reading; the present application provides a quantitative similarity between the video to be detected and typical diseases through contrast learning, changes subjective impression into objective data, and provides a new and reliable reference dimension for clinical diagnosis, assisting doctors in making more accurate judgments. The weight sharing mechanism designed in the present application is a core advantage, by sharing weights in the backbone part of the two models, not only the overall performance of the model is significantly improved, but also the number of parameters to be trained is effectively reduced, and the convergence process of the model is accelerated. BRIEF DESCRIPTION OF DRAWINGS
[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed in the embodiments. Obviously, the drawings described below are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.
[0028] Figure 1 A flow chart of a PDA key area enhanced display method based on deep learning according to an embodiment of the present application. DETAILED DESCRIPTION
[0029] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0030] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application will be further described in detail below with reference to the drawings and specific embodiments.
[0031] The present embodiment provides a PDA key area enhanced display method based on deep learning, as shown in Figure 1 , which comprises:
[0032] Obtaining a target Doppler ultrasound video;
[0033] Inputting the target Doppler ultrasound video into a pre-constructed enhanced display model to output an enhanced video segment dynamic superimposed attention heat map, wherein the enhanced display model adopts a TimeSformer architecture, is used to identify the parasternal short-axis section of each frame in the video and locate the region, and generates an enhanced display through weighted fusion.
[0034] Specifically, the present embodiment can automatically identify the Doppler ultrasound video, enhance the display of PDA key frames and key areas, and provide accurate visual guidance, thereby significantly improving the efficiency and accuracy of reading. The present embodiment provides a quantitative similarity between the video to be detected and the typical disease through contrast learning, changes the subjective impression into objective data, provides a new and reliable reference dimension for clinical diagnosis, and assists doctors in making more accurate judgments. The weight sharing mechanism designed in the present embodiment is a core advantage. By sharing the weights in the two model backbone parts, not only the overall performance of the model is significantly improved, but also the number of parameters that need to be trained is effectively reduced, and the convergence process of the model is accelerated.
[0035] Further, the enhanced display model comprises:
[0036] a section recognition module configured to analyze the target Doppler ultrasound video using a first model to identify a parasternal short-axis view in each frame and determine a key frame based on a preset section recognition confidence threshold;
[0037] a contrast learning module configured to extract features by calculating similarity scores of the key frame sequence and preset positive and negative samples using a second model, and locate key regions by assigning attention scores;
[0038] an enhanced display module configured to superimpose the key frame and the corresponding key region, and output an enhanced video segment dynamic attention heat map.
[0039] Further, the section recognition module comprises:
[0040] The section recognition module is configured to automatically identify a parasternal short-axis view (PSA) from an inputted original echocardiogram video.
[0041] Input: an original Doppler echocardiogram video, denoted as The video may contain multiple different heart sections.
[0042] Objective: accurately identify the parasternal short-axis view (PSA).
[0043] Process: the first model analyzes the input video . The model is based on the TimeSformer architecture, and its task is to classify the section in the video dimension. The processing process can be formally represented as:
[0044] ;
[0045] wherein, represents the parameters of the first model , is the classification result output by the first model , which determines whether the video is a PSA section.
[0046] Output: only when is determined to be a “PSA section”, the video will be passed to the next stage.
[0047] In addition, in order to ensure that the enhanced display is only performed on high-quality and high-confidence section images, a threshold-based key frame selection strategy is introduced.
[0048] When the first model identifies the section, it will assign a confidence score to each frame output a confidence score , which represents the degree of confidence that the model judges the frame to be a standard PSA section (the range is usually [0, 1]). Set a section recognition confidence threshold , only when its confidence score is greater than or equal to the threshold , the frame will be selected into the final selected key frame sequence for enhanced display , denoted as:
[0049] ;
[0050] wherein, is the first model identified in the video all video frames that are PSA sections.
[0051] Further, the contrast learning module:
[0052] uses a contrast learning scheme to train the model to effectively distinguish the features of positive samples (containing PDA) and negative samples (normal).
[0053] Input: receives the video frame data filtered by the section recognition module and judged as a PSA section , and a small number of positive (PDA) sample video sets and negative (normal) sample video sets labeled by doctors.
[0054] Objective: Through the same encoder processing of positive and negative sample pairs, extract feature representations, and pull apart the distance of paired samples in the feature space.
[0055] Processing:
[0056] a) Model training stage (contrast learning): a feature extraction model based on TimeSformer architecture is used as the second model. In the training process, given an anchor sample (Anchor), such as a PDA video, other PDA videos are considered as positive samples, and all normal videos are considered as negative samples. The model is optimized by a contrast learning loss function (such as InfoNCELoss), the goal is to pull the feature vector (Embedding) distance of positive sample pairs closer in a high-dimensional feature space, while pushing the feature vector distance of negative sample pairs further apart.
[0057] ;
[0058] wherein, is the contrast learning loss, is the query sample, is the positive sample, is a negative sample, N is the total number of negative samples, is the cosine similarity function, is the temperature hyperparameter.
[0059] b) Model application stage (similarity calculation): When the video data processed by the section recognition module is input, the trained second model It is mapped into a feature vector; by calculating the distance between the vector and the preset positive sample feature center and the negative sample feature center, a quantitative similarity score is obtained.
[0060] The calculation process of the quantitative similarity score is as follows: first, the new video to be tested is mapped into a high-dimensional feature vector through the model trained in the above process; then, the distance between this vector and two preset "feature centers" is calculated. These two center points are pre-constructed by averaging the feature vectors of a large number of known positive (PDA) and negative (normal) samples; finally, based on the principle that the closer the distance, the higher the similarity, the two distance values are converted into an intuitive score between 0 and 1 through normalization.
[0061] Output: A quantitative similarity score or distance relationship with a typical sample.
[0062] Using the second model Attention mechanism to locate key areas related to lesions: Contains ,Model In the pair When extracting features, each key frame Each image block (patch) in each key frame Each pixel in Assign an attention score ), the higher the attention score, the greater the contribution of the area to the model's judgment. The set of attention scores of all image blocks constitutes an attention map, and an intuitive attention heat map can be generated through pseudo-color rendering.
[0063] Furthermore, the display module is enhanced:
[0064] The enhanced video clips are dynamically superimposed with attention heatmaps, and the heatmap display intensity and spatial focus range are adjusted by setting joint display control parameters.
[0065] Specifically, set a joint display control parameter , while adjusting the heat map display intensity and spatial focus range, the final display intensity of the heat map overlay Calculated by the following formula:
[0066] ;
[0067] where, is a normalization function, is the attention score of pixel point in the superimposed attention heat map, is the maximum attention score, is a threshold parameter, is an exponential function used to construct a Gaussian distribution, is the physical distance between pixel point and the attention peak pixel point in the superimposed attention heat map, is the standard deviation of the Gaussian distribution, the larger the value of σ, the wider the range of the heat map, and the smaller the value of σ, the narrower the range; is a proportionality coefficient, is a small correction term to ensure that still has a display area when it is set to 1.
[0068] By adjusting a single parameter , the overall "focus intensity" of the heat map can be controlled. A lower value will display the heat map with a wider range and lower threshold, suitable for global overview; a higher value will narrow the display range and increase the display threshold, highlighting only the core signal area, suitable for precise focus.
[0069] Further, the first model and the second model employ a weight sharing mechanism during training, sharing the same weight parameters in a pre-set part of the network.
[0070] Specifically, to improve the overall performance and training efficiency of the model, a weight sharing strategy is adopted between the two-stage models.
[0071] The backbone network of the two models ( and ) shares the same set of weight parameters, denoted as .
[0072] The two models ( and ) have independent classification heads for specific tasks, with their parameters denoted as and .
[0073] Therefore, the total parameters of the two-stage model can be represented as:
[0074] Stage 1 model parameters: =[
[0075] Stage 2 model parameters: [ ];
[0076] This design significantly improves the performance of the model by sharing learned features between different but related tasks, while effectively reducing the number of trainable parameters and accelerating the model convergence process.
[0077] The above-described embodiments are merely intended to describe the preferred modes of the present application, and are not intended to limit the scope of the present application. Various modifications and improvements to the technical solutions of the present application made by those of ordinary skill in the art without departing from the design spirit of the present application shall fall within the scope of protection of the present application as defined by the claims.
Claims
1. A PDA key area enhanced display method based on deep learning, characterized in that: include: Acquire target Doppler ultrasound video; The target Doppler ultrasound video is input into a pre-built enhanced display model, and the enhanced video clip is output with a dynamic superimposed attention heat map. The enhanced display model adopts the TimeSformer architecture to identify the parasternal short axis section of each frame in the video and locate the area, and generates an enhanced display through weighted fusion.
2. The PDA key area enhanced display method based on deep learning according to claim 1, characterized in that: The enhanced display model includes: a section recognition module, configured to analyze the target Doppler ultrasound video using the first model, identify the parasternal short-axis section in each frame, and determine a key frame based on a preset section recognition confidence threshold; a contrastive learning module for extracting features by calculating similarity scores between the key frame sequence and preset positive and negative samples using the second model, and locating key areas by assigning attention scores; The enhanced display module is used to superimpose the key frames and the corresponding key areas, and output the enhanced video clips with dynamic superimposed attention heat maps.
3. The PDA key area enhanced display method based on deep learning according to claim 2, characterized in that: Analyzing the target Doppler ultrasound video using the first model to identify the parasternal short-axis section in each frame includes: ; in, Represents the first model Parameters, The first model The output classification result determines whether it is a parasternal short axis section. For input video.
4. The method for enhancing display of key areas of a PDA based on deep learning according to claim 2, characterized in that: Determining key frames based on a preset slice recognition confidence threshold includes: When identifying the parasternal short-axis section, the first model outputs a confidence score for each frame in the target Doppler ultrasound video. When the confidence score corresponding to the video frame is greater than or equal to a preset section identification confidence threshold, the video frame is determined to be a key frame, and a key frame sequence for enhanced display is added.
5. The PDA key area enhanced display method based on deep learning according to claim 2, characterized in that: Calculate the similarity scores between the key frame sequence and the preset positive and negative samples, including: The second model maps the key frame sequence into a feature vector, calculates the distance between the feature vector and the preset positive sample feature center and the negative sample feature center, and obtains the similarity score, wherein the positive sample is a video set of labeled positive patent ductus arteriosus samples, and the negative sample is a video set of normal samples.
6. The method for enhancing display of key areas of a PDA based on deep learning according to claim 5, characterized in that: The second model is trained using a contrastive learning strategy. During the training process, anchor samples are set, and the video set of labeled positive patent ductus arteriosus samples is set as positive samples, and the normal sample video set is set as negative samples. Through the contrastive learning loss function, optimization is performed to shorten the feature vector distance of the positive sample pair in the high-dimensional feature space, while pushing the feature vector distance of the negative sample pair further away.
7. The method for enhancing the display of key areas of a PDA based on deep learning according to claim 2, characterized in that: Key areas are located by assigning attention scores, including: When the second model performs feature extraction by calculating the similarity score, an attention score is assigned to each image block in each frame, the key area is determined by the attention score, and the attention score set of all image blocks is used to form an attention map, and an attention heat map is generated using pseudo-color rendering.
8. The method for enhancing display of key areas of a PDA based on deep learning according to claim 2, characterized in that: The first model and the second model adopt a weight sharing mechanism during the training process and share the same weight parameters in a preset partial network.
9. The method for enhancing display of key areas of a PDA based on deep learning according to claim 2, characterized in that: The enhanced video clip dynamically overlays the attention heat map and adjusts the heat map display intensity and spatial focus range by setting joint display control parameters.
Citation Information
Patent Citations
Hysteromyoma diagnosis method and device based on deep learning
CN113951866A
Intelligent image diagnosis system, method, equipment and medium for hospital radiology department
CN118553385A
Medical ultrasound field knowledge base construction method, related equipment and program product
CN120338078A
Energy based devices and methods for treatment of patent foramen ovale
US20040193147A1
Method and system for PDA-based ultrasound
US20180185009A1
Cited By
Feature fusion method for intelligent identification of pulmonary valve stenosis echocardiogram
CN121661455A