A method and device for cardiac feature point localization based on a multi-task framework
By performing image preprocessing and task branch training in a multi-task framework for cardiac feature point positioning, the problem of view classification tasks increasing model complexity is solved, and faster and more efficient cardiac feature point positioning is achieved.
Patent Information
- Application Number
- CN202410509492.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-25
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2044-04-25
AI Technical Summary
During the training process of cardiac feature point positioning, the training task of adding view classification results in an increase in the amount of model parameters and calculations, and a decrease in the training speed and inference speed.
The cardiac feature point positioning method based on a multi-task framework is adopted, and the target image data is preprocessed, including image scaling, normalization and standardization processing, and only the cardiac feature point positioning branch or view classification branch is trained in the multi-task framework, or both are trained at the same time, and the two tasks are connected through the attention module.
The number of parameters and calculations of the training model is reduced, the training speed and inference speed are improved, and the loss impact between feature point positioning tasks and view classification tasks is avoided, and the application practicality and stability of the multi-task framework are improved.
Smart Images

Figure CN118470104B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to a method and device for cardiac feature point localization based on a multi-task framework. Background Art
[0002] During the process of cardiac quantitative analysis based on echocardiograms, the localization of cardiac anatomical feature points is a very crucial step. A large number of research algorithms have been proposed for the localization of cardiac anatomical feature points in the prior art. Among them, in order to achieve the localization of cardiac feature points, researchers have proposed to combine the method of conventional cardiac feature point localization with the view classification of echocardiograms to improve the accuracy of cardiac feature point localization.
[0003] However, practice shows that if the training task of view classification is added during the training process of cardiac feature point localization, it will lead to an increase in the number of parameters and computational complexity of the overall training model. At the same time, the training speed and inference speed of the overall training model will also decrease accordingly. Therefore, it is particularly important to provide a method that can increase the training task of view classification in the training model while reducing the number of parameters and computational complexity of the training model and improving the training speed and inference speed of the training model. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a method and device for cardiac feature point localization based on a multi-task framework, which can increase the training task of view classification in the training model while reducing the number of parameters and computational complexity of the training model and improving the training speed and inference speed of the training model.
[0005] To solve the above technical problem, in the first aspect of the present invention, a method for cardiac feature point localization based on a multi-task framework is disclosed. The method includes:
[0006] Performing a preset preprocessing operation on the acquired target image data to obtain a preprocessing result corresponding to the target image data. The preprocessing operation at least includes image scaling, normalization, and standardization processing. The preprocessing result is used to perform different training tasks during framework training.
[0007] According to a preset model training strategy, inputting the preprocessing result into a multi-task framework to be trained to obtain a first target data corresponding to the preprocessing result. The multi-task framework at least includes two training branches, and these two training branches include a first branch for cardiac feature point localization and a second branch for echocardiogram view classification. Among them, each time the framework training is performed in the multi-task framework, only one of the two training branches is trained or both of the two training branches are trained simultaneously. The first target data includes prediction heat map data corresponding to the first branch and view classification data corresponding to the second branch.
[0008] Perform preset heatmap post - processing on the first target data to obtain second target data corresponding to the first target data. The second target data includes multiple feature coordinates; all the feature coordinates are positioning coordinates corresponding to heart feature points; the heatmap post - processing operation includes non - maximum suppression calculation, image scaling processing, and coordinate selection;
[0009] After determining that the multi - task framework has completed training, determine the second target data and multiple view categories corresponding to the view classification data as target output data.
[0010] As an alternative implementation, in the first aspect of the present invention, the step of inputting the pre - processing result into a multi - task framework to be trained according to a preset model training strategy to obtain first target data corresponding to the pre - processing result includes:
[0011] According to a preset model training strategy, determine the training type of the current framework training executed by the multi - task framework. The training type includes an early - stage training type or a late - stage training type;
[0012] According to the model training probability set in the model training strategy, determine the target training branch corresponding to the training type. When the training type is the early - stage training type, the target training branch is the first branch or the second branch; when the training type is the late - stage training type, the target training branch is the third branch. The third branch is used to train the first branch, the second branch, and a preset attention module as a whole, where the attention module is used to connect the training tasks corresponding to the first branch and the second branch;
[0013] Input the pre - processing result into the target training branch to obtain branch output data corresponding to the pre - processing result;
[0014] Among them, when the training type is the late - stage training type and it is determined that the multi - task framework has been trained to convergence, determine the branch output data as the first target data corresponding to the pre - processing result.
[0015] As an alternative implementation, in the first aspect of the present invention, the attention module is used to transfer the target information output by the second branch to the first branch. The target information includes channel attention and spatial attention;
[0016] When the training type is the late training type, the attention module is further configured to perform a first operation corresponding to the channel attention and a second operation corresponding to the spatial attention on the predicted heatmap data corresponding to the first branch and the view classification data corresponding to the second branch in sequence.
[0017] As an alternative implementation, in the first aspect of the present invention, the first branch and the second branch share the same preset residual network in the encoding layer;
[0018] Before determining that the task framework is trained to convergence, inputting the preprocessing result into the target training branch to obtain branch output data corresponding to the preprocessing result includes:
[0019] When the multi-task framework executes the first training task corresponding to the first branch, input the preprocessing result into the residual network to obtain a residual processing result corresponding to the preprocessing result, and the residual processing result corresponds to the predicted heatmap data;
[0020] When the multi-task framework executes the second training task corresponding to the second branch, input the preprocessing result into a preset view classification network to obtain a view classification result corresponding to the preprocessing result, and the view classification result corresponds to the view classification data;
[0021] Wherein, when the training type is the early training type, determine the residual processing result and / or the view classification result as the branch output data corresponding to the preprocessing result;
[0022] When the training type is the late training type, the residual processing result and the view classification result are used to be transmitted to the attention module; at the same time, determine the attention processing result output by the attention module via the residual processing result and the view classification result as the branch output data corresponding to the preprocessing result;
[0023] Wherein, the view classification network includes a feature extraction layer, a global pooling layer, a multi-layer perceptron MLP, and a normalization layer; the feature extraction layer is the residual network.
[0024] As an alternative implementation, in the first aspect of the present invention, during the process of performing the framework training on the multi-task framework, when the training type is the early training type, only calculate the loss value corresponding to the target training branch, and update the network parameters adopted by the target training branch through the loss value corresponding to the target training branch;
[0025] When the training type is the late training type, calculate the first loss value corresponding to the first branch and the second loss value corresponding to the second branch simultaneously; then, in combination with the set gradient calculation formula, perform gradient calculation on the first loss value and the second loss value, and then update the network parameters adopted by the target training branch through the obtained gradient calculation results.
[0026] As an optional implementation manner, in the first aspect of the present invention, during the process of performing the framework training on the multi-task framework, and when the training type is the late training type, the loss functions used to perform the loss value calculation include the DWA function and the DTP function;
[0027] Among them, the calculation formula corresponding to the DWA function is:
[0028]
[0029] Among them, w t (n) represents the weight occupied by the Loss of the training task t in the nth round of training; represents the task set corresponding to all training tasks when performing the framework training; the Loss of the task t in the nth round of training is Loss t (n), then the training speed of the task t in the nth round of training is r t (n) = Loss t (n) / Loss t (n - 1); T is the temperature constant;
[0030] The calculation formula corresponding to the DTP function is:
[0031]
[0032] Among them, k t is used to represent the KPI of each training task t; and k t ∈ [0, 1]; γ t is used to update the network weights for a certain training task.
[0033] As an optional implementation manner, in the first aspect of the present invention, the view classification data includes multiple view category probability vectors;
[0034] Before performing the preset heatmap post-processing on the first target data to obtain the second target data corresponding to the first target data, the method further includes:
[0035] Perform preset category post-processing on all the obtained view category probability vectors to obtain at least one view category corresponding to the view classification data;
[0036] Update the view classification data according to all the view categories, and trigger the execution of the preset heatmap post - processing on the first target data to obtain second target data corresponding to the first target data;
[0037] Among them, the execution of preset category post - processing on all the obtained view category probability vectors to obtain at least one view category corresponding to the view classification data includes:
[0038] Execute maximum value indexing on all the obtained view category probability vectors according to a preset indexing function to obtain an indexing result corresponding to all the view category probability vectors; the indexing function includes the Argmax function;
[0039] Determine at least one view category corresponding to the indexing result according to the indexing result and a preset category encoding; the category encoding is a pre - determined encoding for the target image data.
[0040] The second aspect of the present invention discloses a cardiac feature point positioning device based on a multi - task framework, and the device includes:
[0041] A pre - processing module, configured to perform a preset pre - processing operation on the obtained target image data to obtain a pre - processing result corresponding to the target image data, and the pre - processing operation at least includes image scaling, normalization, and standardization processing; the pre - processing result is used to perform different training tasks during framework training;
[0042] A multi - task training module, configured to input the pre - processing result into a multi - task framework to be trained according to a preset model training strategy to obtain first target data corresponding to the pre - processing result; the multi - task framework at least includes two training branches, and the two training branches include a first branch for cardiac feature point positioning and a second branch for echocardiogram view classification; among them, each time the multi - task framework performs the framework training, only one of the two training branches is trained or the two training branches are trained simultaneously; the first target data includes prediction heatmap data corresponding to the first branch and view classification data corresponding to the second branch;
[0043] A heatmap post - processing module, configured to perform a preset heatmap post - processing on the first target data to obtain second target data corresponding to the first target data, and the second target data includes multiple feature coordinates; all the feature coordinates are positioning coordinates corresponding to cardiac feature points; the heatmap post - processing operation includes non - maximum suppression calculation, image scaling processing, and coordinate selection;
[0044] A determination module, configured to, after determining that the multi-task framework has completed training, determine the second target data and multiple view categories corresponding to the view classification data as target output data.
[0045] As an optional implementation manner, in the second aspect of the present invention, the manner in which the multi-task training module inputs the preprocessing result into a multi-task framework to be trained according to a preset model training strategy to obtain first target data corresponding to the preprocessing result specifically includes:
[0046] According to a preset model training strategy, determine the training type of the multi-task framework currently performing framework training, where the training type includes an early training type or a late training type;
[0047] According to the model training probability set in the model training strategy, determine the target training branch corresponding to the training type; when the training type is the early training type, the target training branch is the first branch or the second branch; when the training type is the late training type, the target training branch is the third branch; the third branch is used to train the first branch, the second branch, and a preset attention module as a whole, where the attention module is used to connect the training tasks corresponding to the first branch and the second branch;
[0048] Input the preprocessing result into the target training branch to obtain branch output data corresponding to the preprocessing result;
[0049] Wherein, when the training type is the late training type and it is determined that the multi-task framework has been trained to convergence, determine the branch output data as the first target data corresponding to the preprocessing result.
[0050] As an optional implementation manner, in the second aspect of the present invention, the attention module is used to transmit the target information output by the second branch to the first branch, and the target information includes channel attention and spatial attention;
[0051] When the training type is the late training type, the attention module is further used to sequentially perform a first operation operation corresponding to the channel attention and a second operation operation corresponding to the spatial attention on the predicted heat map data corresponding to the first branch and the view classification data corresponding to the second branch.
[0052] As an optional implementation manner, in the second aspect of the present invention, the first branch and the second branch share the same preset residual network in the encoding layer;
[0053] The multi-task training module inputs the pre-processing result into the target training branch to obtain branch output data corresponding to the pre-processing result, and the specific method includes:
[0054] Before determining that the multi-task framework converges during training, when the multi-task framework executes the first training task corresponding to the first branch, the pre-processing result is input into the residual network to obtain a residual processing result corresponding to the pre-processing result, and the residual processing result corresponds to the predicted heat map data;
[0055] When the multi-task framework executes the second training task corresponding to the second branch, the pre-processing result is input into a preset view classification network to obtain a view classification result corresponding to the pre-processing result, and the view classification result corresponds to the view classification data;
[0056] Among them, when the training type is the early training type, the residual processing result and / or the view classification result is determined as the branch output data corresponding to the pre-processing result;
[0057] When the training type is the late training type, the residual processing result and the view classification result are used to be transmitted to the attention module; meanwhile, the attention processing result output by the attention module through the residual processing result and the view classification result is determined as the branch output data corresponding to the pre-processing result;
[0058] Among them, the view classification network includes a feature extraction layer, a global pooling layer, a multi-layer perceptron MLP, and a normalization layer; the feature extraction layer is the residual network.
[0059] As an optional implementation manner, in the second aspect of the present invention, during the process of performing the framework training on the multi-task framework, when the training type is the early training type, only the loss value corresponding to the target training branch is calculated, and the network parameters adopted by the target training branch are updated through the loss value corresponding to the target training branch;
[0060] When the training type is the late training type, the first loss value corresponding to the first branch and the second loss value corresponding to the second branch are calculated simultaneously; then, in combination with the set gradient calculation formula, gradient calculation is performed on the first loss value and the second loss value, and then the network parameters adopted by the target training branch are updated through the obtained gradient calculation result.
[0061] As an alternative implementation, in the second aspect of the present invention, during the process of performing the framework training on the multi-task framework, and when the training type is the late training type, the loss function used to perform the loss value calculation includes the DWA function and the DTP function;
[0062] Among them, the calculation formula corresponding to the DWA function is:
[0063]
[0064] Among them, w t (n) represents the weight of the Loss of the training task t in the nth round of training; represents the task set corresponding to all training tasks when performing the framework training; the Loss of task t in the nth round of training is Loss t (n), then the training speed of task t in the nth round of training is r t (n) = Loss t (n) / Loss t (n - 1); T is the temperature constant;
[0065] The calculation formula corresponding to the DTP function is:
[0066]
[0067] Among them, k t is used to represent the KPI of each training task t; and k t ∈ [0, 1]; γ t is used to update the network weights for a certain training task.
[0068] As an alternative implementation, in the second aspect of the present invention, the view classification data includes multiple view category probability vectors; the device further includes:
[0069] A category post-processing module, configured to perform a preset category post-processing on all the obtained view category probability vectors to obtain at least one view category corresponding to the view classification data before the heatmap post-processing module performs a preset heatmap post-processing on the first target data to obtain second target data corresponding to the first target data;
[0070] An update module, configured to update the view classification data according to all the view categories, and trigger the heatmap post-processing module to perform the preset heatmap post-processing on the first target data to obtain second target data corresponding to the first target data;
[0071] Among them, the specific way that the category post - processing module performs preset category post - processing on all the obtained view category probability vectors to obtain at least one view category corresponding to the view classification data includes:
[0072] Perform maximum value indexing on all the obtained view category probability vectors according to a preset indexing function to obtain the indexing results corresponding to all the view category probability vectors; the indexing function includes the Argmax function;
[0073] Determine at least one view category corresponding to the indexing results according to the indexing results and a preset category encoding; the category encoding is the encoding determined in advance for the target image data.
[0074] The third aspect of the present invention discloses another cardiac feature point localization device based on a multi - task framework. The device includes:
[0075] A memory storing executable program code;
[0076] A processor coupled to the memory;
[0077] The processor calls the executable program code stored in the memory and executes the cardiac feature point localization method based on the multi - task framework disclosed in the first aspect of the present invention.
[0078] The fourth aspect of the present invention discloses a computer storage medium. The computer storage medium stores computer instructions, which are used to execute the cardiac feature point localization method based on the multi - task framework disclosed in the first aspect of the present invention when called.
[0079] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0080] In an embodiment of the present invention, a method for localizing cardiac feature points based on a multi-task framework is provided. The method includes: performing a preset preprocessing operation on the acquired target image data to obtain a preprocessing result corresponding to the target image data, where the preprocessing operation at least includes frame image extraction, image scaling, normalization, and standardization processing; the preprocessing result includes two pieces of preprocessing data; different preprocessing data are used to perform different training tasks when performing framework training; according to a preset model training strategy, input the preprocessing result into the multi-task framework to be trained to obtain first target data corresponding to the preprocessing result; the multi-task framework includes two training branches, and the two training branches include a first branch for localizing cardiac feature points and a second branch for classifying echocardiogram views; the two pieces of preprocessing data respectively correspond to the first branch and the second branch; wherein, each time the multi-task framework performs framework training, only one of the two training branches is trained; the first target data includes predicted heat map data corresponding to the first branch and view classification data corresponding to the second branch; performing a preset post-processing on the heat map of the first target data to obtain second target data corresponding to the first target data, and the second target data includes multiple feature coordinates, and all the feature coordinates are positioning coordinates corresponding to the cardiac feature points; the heat map post-processing operation includes non-maximum suppression calculation, image scaling processing, and coordinate selection. It can be seen that by implementing the present invention, by performing preset preprocessing on the acquired target image data, including image scaling, normalization, and standardization processing, it helps to improve the stability, convergence speed, generalization ability, and training efficiency of the model; after obtaining the required preprocessing result, it is possible to adjust the training branches of the multi-task framework according to the set model training strategy during different training periods, thereby avoiding the situation where the feature point localization task (the first branch) can be normally trained due to sharing the feature point localization data, while the view classification algorithm (the second branch) cannot converge normally, that is, improving the application practicability and stability of the multi-task framework; among them, corresponding different training branches to different training periods of the multi-task framework is beneficial to reducing the influence of the gradient of the feature point localization (the first branch) on the backpropagation of the view classification branch (the second branch) when the multi-task framework is not stable enough in the early stage of training of the multi-task framework, that is, the setting of this model training strategy is beneficial to reducing the loss influence between the first branch and the second branch in the early stage of training; in the later stage of training, training the two branches simultaneously through the model training strategy setting is beneficial to improving the training accuracy and reliability of the multi-task framework. BRIEF DESCRIPTION OF THE DRAWINGS
[0081] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0082] Figure 1 It is a schematic flowchart of a method for locating cardiac feature points based on a multi-task framework disclosed in an embodiment of the present invention;
[0083] Figure 2 It is a schematic flowchart of another method for locating cardiac feature points based on a multi-task framework disclosed in an embodiment of the present invention;
[0084] Figure 3 It is a schematic structural diagram of a device for locating cardiac feature points based on a multi-task framework disclosed in an embodiment of the present invention;
[0085] Figure 4 It is a schematic structural diagram of another device for locating cardiac feature points based on a multi-task framework disclosed in an embodiment of the present invention;
[0086] Figure 5 It is a schematic structural diagram of yet another device for locating cardiac feature points based on a multi-task framework disclosed in an embodiment of the present invention;
[0087] Figure 6 It is a schematic structural diagram of a multi-task framework network disclosed in an embodiment of the present invention;
[0088] Figure 7 It is a schematic structural diagram of an attention module disclosed in an embodiment of the present invention;
[0089] Figure 8 It is a schematic overall inference flowchart of a multi-task framework disclosed in an embodiment of the present invention. Detailed implementation manners
[0090] In order to enable those skilled in the art to better understand the solutions of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0091] In the description, claims and above-mentioned drawings of the present invention, terms such as "first" and "second" are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or terminal comprising a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or terminals.
[0092] Reference to "embodiment" herein means that a specific feature, structure or characteristic described in connection with the embodiment can be included in at least one embodiment of the present invention. The phrase appears in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0093] The present invention discloses a method and device for cardiac feature point localization based on a multi-task framework. By performing preset preprocessing on the acquired target image data, including image scaling, normalization, and standardization processing, it helps to improve the stability, convergence speed, generalization ability, and training efficiency of the model; after obtaining the required preprocessing results, it is possible to adjust the training branches of the multi-task box according to the set model training strategy during different training epochs, thereby avoiding the situation where the feature point localization task (the first branch) can be normally trained due to sharing feature point localization data, while the view classification algorithm (the second branch) cannot converge normally, that is, improving the application practicability and stability of the multi-task framework; among them, corresponding different training branches are trained during different training epochs of the multi-task framework, which is beneficial to reducing the influence of the gradient of the feature point localization (the first branch) on the backpropagation of the view classification branch (the second branch) when the multi-task framework is not stable enough in the early stage of training of the multi-task framework, that is, the setting of this model training strategy is beneficial to reducing the loss influence between the first branch and the second branch in the early stage of training; in the later stage of training, training the two branches simultaneously through the model training strategy setting is beneficial to improving the training accuracy and reliability of the multi-task framework.
[0094] Embodiment 1
[0095] Please refer to Figure 1 , Figure 1 is a schematic flowchart of a method for cardiac feature point localization based on a multi-task framework disclosed in an embodiment of the present invention. Among them, Figure 1 The described method for cardiac feature point localization based on a multi-task framework can be applied to a device for cardiac feature point localization based on a multi-task framework, which is not limited in the embodiments of the present invention. AsFigure 1 As shown in Figure 1 , the method for locating cardiac feature points based on a multi-task framework may include the following operations:
[0096] 101. Perform a preset preprocessing operation on the acquired target image data to obtain a preprocessing result corresponding to the target image data.
[0097] In an embodiment of the present invention, the preprocessing operation at least includes image scaling, normalization, and standardization processing; the preprocessing result is used to perform different training tasks during framework training.
[0098] In an embodiment of the present invention, by performing an image scaling operation on the acquired target image data, a large image can be scaled to an appropriate size, which is beneficial to significantly reducing the computational amount of the neural network (the neural network refers to the subsequent multi-task framework to be trained) when processing the image, thereby accelerating the training and inference speeds of the neural network. At the same time, this image scaling operation can also remove some unnecessary details and noises in the image, helping the neural network to focus more on the main features of the image.
[0099] In an embodiment of the present invention, the set normalization processing can scale the pixel values of the image to a specific interval (such as [0, 1] or [-1, 1]), making the neural network more stable during the training process and reducing training problems caused by uneven data distribution; at the same time, the normalization processing also helps the neural network converge to the optimal solution faster and improves the training efficiency; in addition, by normalizing the input data, the neural network can better adapt to different data distributions, thereby improving its generalization ability on unknown data.
[0100] In an embodiment of the present invention, the standardization operation can make the pixel value distribution of the image more conform to the normal distribution, helping the neural network better learn the internal laws of the data; at the same time, standardization can also avoid the problem of too large gradient values caused by unstandardized data, making the neural network more stable during the training process and reducing the risk of gradient disappearance or gradient explosion; in addition, it can also accelerate the training speed of the neural network and reduce the training time required.
[0101] In an embodiment of the present invention, the target image data may be a conventional echocardiogram. Since a conventional echocardiogram is a video in segments and the actual algorithm processes an image of a fixed size, when actually inputting the target image data into the subsequent multi-task framework, it is necessary to combine the training strategy to extract frames of images from this video in segments (that is, a frame image extraction operation will be performed), then scale the extracted image to a unified fixed size of 512×512 through bilinear interpolation, and finally perform normalization and standardization processing again to obtain the required preprocessing result. At the same time, considering the subsequent training of the multi-task framework, the preprocessing result will be evenly divided into two portions of data.
[0102] In an embodiment of the present invention, optionally, after obtaining the preprocessing result, the method further includes:
[0103] Performing a preset data augmentation operation on the preprocessing result to obtain a data augmentation result corresponding to the preprocessing result. The data augmentation operation includes affine transformation and pixel adjustment operations; the image parameters corresponding to the pixel adjustment operations include at least one of brightness, contrast, and saturation;
[0104] Performing a preset coordinate transformation operation on the data augmentation result to obtain a coordinate transformation result corresponding to the data augmentation result. The coordinate transformation operation includes adding coordinate information based on CoordConv;
[0105] Updating the preprocessing result according to the coordinate transformation result.
[0106] In an embodiment of the present invention, the above affine transformation includes image rotation, translation, and scaling operations; secondly, random changes in brightness, contrast, and saturation on each pixel are also considered to be added.
[0107] In an embodiment of the present invention, in order to enhance the detection ability of the subsequent multi-task framework, it can be achieved by setting corresponding data augmentation means.
[0108] 102. According to the preset model training strategy, input the preprocessing result into the multi-task framework to be trained to obtain the first target data corresponding to the preprocessing result.
[0109] In an embodiment of the present invention, the multi-task framework includes at least two training branches, and the two training branches include a first branch for cardiac feature point localization and a second branch for echocardiogram view classification.
[0110] In an embodiment of the present invention, it should be noted that when performing framework training on the multi-task framework each time, only one of the two training branches is trained or the two training branches are trained simultaneously; the first target data includes the predicted heat map data corresponding to the first branch and the view classification data corresponding to the second branch.
[0111] In an embodiment of the present invention, it should be noted that for the training strategy of the multi-task framework, the following three points are considered:
[0112] (1) Image extraction strategy: Compared with taking out all the images in the echocardiogram in each round, randomly taking one frame of the image and sending it into the training model corresponding to the multi-task framework. This training method speeds up the convergence speed of the network, theoretically makes full use of all the images in each echocardiogram, and avoids some model biases that may be caused by unequal numbers of frames in the echocardiogram.
[0113] (2) Sampling strategy: Since the number of samples in different categories varies by several times, category-balanced sampling is adopted here, that is, the overall sampling probability of samples in each category is equal. Compared with sample-balanced sampling, using category-balanced sampling not only makes the training more stable, but also alleviates the problem that the recognition effect of relatively small-number categories is relatively poor.
[0114] (3) Pre-training strategy: Load the weights pre-trained on ImageNet to accelerate the convergence speed. Although the images in ImageNet are quite different from those in the view classification dataset of this paper, pre-training in this task still significantly accelerates the convergence speed of the model.
[0115] In terms of the ensemble optimization strategy, each time the trained model makes an inference, K frames are randomly selected from one echocardiogram, and the voting is performed on the prediction results of these K frames to determine the final predicted category of the echocardiogram. Not adopting the ensemble optimization strategy corresponds to K = 1.
[0116] 103. Perform preset heatmap post-processing on the first target data to obtain second target data corresponding to the first target data.
[0117] In the embodiments of the present invention, the second target data includes multiple feature coordinates, and all feature coordinates are positioning coordinates corresponding to cardiac feature points; the heatmap post-processing operations include non-maximum suppression calculation, image scaling processing, and coordinate selection.
[0118] 104. After determining that the multi-task framework has completed training, determine the second target data and multiple view categories corresponding to the view classification data as the target output data.
[0119] It can be seen that the implementation Figure 1The described heart feature point localization method based on a multi-task framework performs preset preprocessing on the acquired target image data, including image scaling, normalization, and standardization processing, which helps to improve the stability, convergence speed, generalization ability, and training efficiency of the model; after obtaining the required preprocessing results, it can adjust the training branches of the multi-task box according to the set model training strategy during different training epochs, thus avoiding the situation where the feature point localization task (the first branch) can be normally trained while the view classification algorithm (the second branch) cannot converge normally, that is, improving the application practicability and stability of the multi-task framework; among them, setting different training branches for different training epochs of the multi-task framework is beneficial to reducing the influence of the gradient of the feature point localization (the first branch) on the backpropagation of the view classification branch (the second branch) when the multi-task framework is not stable enough in the early stage of training of the multi-task framework, that is, the setting of this model training strategy is beneficial to reducing the loss influence between the first branch and the second branch in the early stage of training; in the later stage of training, training the two branches simultaneously through the model training strategy setting is beneficial to improving the training accuracy and reliability of the multi-task framework.
[0120] In an optional embodiment, the manner of inputting the preprocessing result into the multi-task framework to be trained according to the preset model training strategy to obtain the first target data corresponding to the preprocessing result in step 102 specifically includes:
[0121] According to the preset model training strategy, determine the training type of the current execution of the framework training of the multi-task framework, and the training type includes the early training type or the late training type;
[0122] According to the model training probability set in the model training strategy, determine the target training branch corresponding to the training type; when the training type is the early training type, the target training branch is the first branch or the second branch; when the training type is the late training type, the target training branch is the third branch; the third branch is used to train the first branch, the second branch, and the preset attention module as a whole, where the attention module is used to connect the training tasks corresponding to the first branch and the second branch;
[0123] Input the preprocessing result into the target training branch to obtain the branch output data corresponding to the preprocessing result;
[0124] Among them, when the training type is the late training type and it is determined that the multi-task framework has been trained to convergence, determine that the branch output data is the first target data corresponding to the preprocessing result.
[0125] It can be seen that in this alternative embodiment, different training branches are correspondingly set for different training epochs of the multi-task framework, which helps to reduce the impact of the backpropagation of the gradient of feature point localization (the first branch) on the view classification branch (the second branch) during the early stage of training of the multi-task framework when the multi-task framework is not stable enough and both training branches are trained simultaneously. That is to say, the setting of this model training strategy helps to reduce the loss impact between the first branch and the second branch in the early stage of training; in the later stage of training, training both branches simultaneously through the model training strategy setting helps to improve the training accuracy and reliability of the multi-task framework.
[0126] In this alternative embodiment, for the specific data flow after the pre-processing result is input into the multi-task framework, please refer to Figure 6 , Figure 6 which is a schematic structural diagram of a multi-task framework network disclosed in an embodiment of the present invention. As Figure 6 shown, the multi-task framework integrates the first branch for right ventricular anatomical feature point localization and the second branch for echocardiogram view classification through the following two ways, where:
[0127] (1) The first branch and the second branch share the ResNet part in the encoding layer; further, the shared residual network can be a RestNet50 network. On the one hand, the view classification algorithm (corresponding to the second branch) can be trained using each frame in the echocardiogram, while only the ED / ES frames in the echocardiogram are used in the training process of the feature point localization algorithm (corresponding to the first branch). Therefore, the view classification model actually has a larger dataset. By sharing the ResNet part in the encoding layer, the feature point localization branch can benefit from the larger dataset scale of view classification; on the other hand, sharing part of the network can reduce the total number of parameters of the two tasks and reduce the total inference time of the two tasks.
[0128] (2) When the multi-task framework enters the later stage of training, a new attention module is introduced. In this multi-task framework, the view classification branch enhances the feature map extracted by the feature point localization branch through the attention module. More specifically, the view classification branch conveys two types of information to the feature point localization branch through the attention module: which types of features should be focused on (channel attention), and which positions on the feature map should be noted for features (spatial attention).
[0129] It can be seen that in this alternative embodiment, through the sharing of the RestNet network set in the multi-task framework, the feature point localization branch can benefit from the view classification branch, while reducing the total number of parameters and the total inference time of the two tasks, which is beneficial to improving the operation efficiency and speed of the multi-task framework; at the same time, the setting of the attention module is beneficial to enhancing the feature map extracted by the feature point localization branch, that is, it is beneficial to improving the accuracy of the subsequent adjusted data.
[0130] In this alternative embodiment, for Figure 6 the description is as follows:
[0131] After performing a preset preprocessing operation on the target image data to obtain a preprocessing result, the preprocessing result is input into the multi-task framework. In the vertical direction, after the first layer of processing, the first layer includes at least one convolutional layer, one batch normalization layer, and one ReLu layer; the next layer consists of a max pooling layer and three residual blocks; the next layer consists of a (2x downsampling) residual block and three residual blocks; and so on, each subsequent layer consists of one (2x downsampling) residual block and multiple residual blocks. At the same time, at the second-to-last layer in the vertical direction, a global pooling layer is set, and the next layer below it is set with an MLP perception layer and a Softmax normalization layer. And there are two output directions in the vertical direction. One output direction is the same as the original vertical direction, and the corresponding output result corresponds to Figure 8 the class probability vector output after passing through the multi-task framework in Figure 7 ; the other output direction in the vertical direction corresponds to the view classification branch (the second branch) in
[0132] That is, in this other direction, it is input into the newly added attention module to convey two types of information, channel attention and spatial attention, to the feature point localization branch. Figure 7 In the horizontal direction, each processing layer consists of at least one convolutional layer, one batch normalization layer, and one ReLu layer, and they are stacked, and at the same time, one transposed convolutional layer is concatenated; when finally output, that is, the first layer in the vertical direction corresponds to Figure 8 the output of the feature point localization branch in
[0133] Embodiment 2
[0134] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of another method for cardiac feature point localization based on a multi-task framework disclosed in an embodiment of the present invention. Among them, Figure 2The described method for locating cardiac feature points based on a multi-task framework can be applied to a device for locating cardiac feature points based on a multi-task framework, which is not limited in the embodiments of the present invention. As Figure 2 shown, the method for locating cardiac feature points based on a multi-task framework may include the following operations:
[0135] 201. Perform a preset preprocessing operation on the obtained target image data to obtain a preprocessing result corresponding to the target image data.
[0136] 202. According to a preset model training strategy, determine the training type for the current execution framework training of the multi-task framework, where the training type includes an early training type or a late training type.
[0137] 203. According to the model training probability set in the model training strategy, determine the target training branch corresponding to the current training type.
[0138] In the embodiments of the present invention, when the training type is the early training type, the target training branch is the first branch or the second branch; when the training type is the late training type, the target training branch is the third branch; the third branch is used to train the first branch, the second branch, and a preset attention module as a whole, where the attention module is used to connect the training tasks corresponding to the first branch and the second branch.
[0139] In the embodiments of the present invention, the attention module is used to transfer the target information output by the second branch to the first branch, and the target information includes channel attention and spatial attention;
[0140] When the training type is the late training type, the attention module is further used to perform a first operation corresponding to channel attention and a second operation corresponding to spatial attention on the predicted heatmap data corresponding to the first branch and the view classification data corresponding to the second branch in sequence.
[0141] 204. Input the preprocessing result into the target training branch to obtain branch output data corresponding to the preprocessing result.
[0142] 205. When the training type is the late training type and it is determined that the multi-task framework training has converged, determine the branch output data as the first target data corresponding to the preprocessing result.
[0143] 206. Perform a preset heatmap postprocessing on the first target data to obtain second target data corresponding to the first target data.
[0144] 207. After determining that the multi-task framework has completed training, determine the second target data and multiple view categories corresponding to the view classification data as the target output data.
[0145] In the embodiments of the present invention, for other descriptions of step 201 and steps 206-207, please refer to the other specific descriptions of step 101 and steps 103-104 in the first embodiment, and the embodiments of the present invention will not be elaborated herein.
[0146] It can be seen that in the Figure 2 described heart feature point positioning method based on a multi-task framework, when specifically setting up the training, it can also combine the training type and training strategy of the currently executed framework training. Specifically, based on the early training stage and the late training stage, corresponding target training operations are executed. That is, the training of the multi-task framework training model takes into account different training periods, which is beneficial to improving the training stability of the multi-task framework in the early training stage and is also beneficial to improving the accuracy of training the multi-task framework.
[0147] In an optional embodiment, the first branch and the second branch share the same preset residual network in the encoding layer;
[0148] Before determining that the task framework training converges, the manner of inputting the preprocessing result into the target training branch in step 204 to obtain the branch output data corresponding to the preprocessing result specifically includes:
[0149] When the multi-task framework executes the first training task corresponding to the first branch, input the preprocessing result into the residual network to obtain the residual processing result corresponding to the preprocessing result, and the residual processing result corresponds to the predicted heat map data;
[0150] When the multi-task framework executes the second training task corresponding to the second branch, input the preprocessing result into the preset view classification network to obtain the view classification result corresponding to the preprocessing result, and the view classification result corresponds to the view classification data;
[0151] Among them, when the training type is the early training type, determine the residual processing result and / or the view classification result as the branch output data corresponding to the preprocessing result;
[0152] When the training type is the late training type, the residual processing result and the view classification result are used to be transmitted to the attention module; at the same time, determine the attention processing result output by the attention module from the residual processing result and the view classification result as the branch output data corresponding to the preprocessing result;
[0153] Among them, the view classification network includes a feature extraction layer, a global pooling layer, a multi-layer perceptron MLP, and a normalization layer; the feature extraction layer is a residual network.
[0154] In this optional embodiment, it should be noted that during the process of performing framework training on the multi-task framework, when the training type is the early training type, only the loss value corresponding to the target training branch is calculated, and the network parameters adopted by the target training branch are updated through the loss value corresponding to the target training branch;
[0155] When the training type is the late training type, the first loss value corresponding to the first branch and the second loss value corresponding to the second branch are calculated simultaneously; then, in combination with the set gradient calculation formula, gradient calculation is performed on the first loss value and the second loss value, and then the network parameters adopted by the target training branch are updated through the obtained gradient calculation result.
[0156] In this optional embodiment, when the training type is the early training type, and it is determined that the loss values corresponding to the two training branches (the first branch and the second branch) are the smallest, or it is determined that the two training branches are trained to convergence, the training type is switched to the late training type, and at the same time, the attention module is added to the update object that performs parameter update. At this time, the two branches are trained simultaneously.
[0157] In this optional embodiment, please refer to Figure 7 , Figure 7 which is a schematic structural diagram of an attention module disclosed in an embodiment of the present invention; as Figure 7 shown, the attention module connects the decoding branches of the view classification task (the second branch) and the feature point localization task (the first branch), and information flows from the view classification branch to the feature point localization branch.
[0158] In this optional embodiment, in the early stage of training the multi-task framework, if the entire multi-task framework is directly trained, the gradient of feature point localization will be backpropagated to the view classification branch along the attention module, that is, the view classification branch will be affected not only by the view classification loss but also by the feature point localization loss, which will affect the training of the view classification branch. Therefore, when actually training the training model, it is necessary to divide the training process of the training model into the early training stage and the late training stage based on whether the attention module participates in parameter update. This setting of dividing the training period / training type is to improve the training accuracy of the overall training model.
[0159] In this optional embodiment, specifically, as Figure 7 shown, denote the class prediction vector activated by the view classification branch as p view , then the attention module first makes p view pass through multiple layers of MLP and activate to obtain channel attention, and then makes p view pass through another group of multiple layers of ML to obtain a vector v spatial with a length of 64, v spatialIt will perform inner products with the multi-channel vectors at each spatial position on the original feature map in sequence and activate them to obtain spatial attention. Finally, the original feature map will enhance the appropriate feature channels and feature spaces according to the channel attention and spatial attention in sequence.
[0160] In this alternative embodiment, it should be noted that the number of parameters and the amount of computation introduced by this attention module are not large. Under the experimental settings of this alternative embodiment, adding this attention module only requires an additional 9.22K parameters and 8.96K FLOPs of computational volume. The increased number of parameters and computational volume are not in the same order of magnitude as the original multi-task framework.
[0161] In this alternative embodiment, since the attention module does not participate in the update of network parameters in the early stage of training the model, at this time, only the first branch and the second branch are trained and the parameters are updated; at the same time, the first branch and the second branch use a unified multi-task loss function to calculate the overall framework loss value. When it is determined that the framework loss value reaches the preset loss value range, that is, the preset condition is met, at this time, it is determined that the training progress of the training model reaches a certain level, so that the training model as a whole reaches a relatively stable state, and the attention module can be added to the ranks of parameter update, that is, it is correspondingly determined that the training type can be switched to the later training type.
[0162] In this alternative embodiment, during the process of performing framework training on the multi-task framework and when the training type is the later training type, the loss function used to perform loss value calculation includes the DWA function and the DTP function;
[0163] Among them, the calculation formula corresponding to the DWA function is:
[0164]
[0165] Among them, w t (n) represents the weight of the Loss of training task t in the nth round of training; represents the task set corresponding to all training tasks when performing framework training; the Loss of task t in the nth round of training is Loss t (n), then the training speed of task t in the nth round of training is r t (n) = Loss t (n) / Loss t (n - 1); T is the temperature constant; when T = 1, this expression degenerates to: w t (n) = N · Softmax t (r i (n)).
[0166] The calculation formula corresponding to the DTP function is:
[0167]
[0168] where k t is used to represent the KPI for each training task t; and k t ∈ [0, 1]; γ t is used to update the network weights for a certain training task.
[0169] In this alternative embodiment, compared with DWA, DTP is slightly more complex in terms of both concept and computational complexity. First, it is necessary to define KPIs (Key Performance Indicators) for each task; and the KPI for the echocardiogram view classification task is defined as the classification accuracy, and the KPI for the feature point localization task is the feature point detection rate SDR.
[0170] It can be seen that in this alternative embodiment, when training the training model, an update process for network parameters and network weights is set; at the same time, based on the training situation of the framework loss value, monitoring of the training type is achieved, corresponding to the monitoring of the early-stage training type and the late-stage training type; which is beneficial to improving the applicability of the training model of the overall multi-task framework.
[0171] In another alternative embodiment, before the above step 205 performs preset heatmap post-processing on the first target data to obtain the second target data corresponding to the first target data, the method further includes:
[0172] Performing preset class post-processing on all obtained view category probability vectors to obtain at least one view category corresponding to the view classification data;
[0173] Updating the view classification data according to all view categories, and triggering the execution of the above-mentioned preset heatmap post-processing on the first target data to obtain the second target data corresponding to the first target data;
[0174] Among them, the manner of performing preset class post-processing on all obtained view category probability vectors to obtain at least one view category corresponding to the view classification data specifically includes:
[0175] Performing maximum value indexing on all obtained view category probability vectors according to a preset indexing function to obtain the indexing results corresponding to all view category probability vectors; the indexing function includes the Argmax function;
[0176] Determining at least one view category corresponding to the indexing result according to the indexing result and a preset class encoding; the class encoding is a previously determined encoding for the target image data.
[0177] Specifically, the implementation process of this class post-processing can refer toFigure 8 , Figure 8 It is a schematic diagram of the overall inference process of a multi-task framework disclosed in an embodiment of the present invention.
[0178] It can be seen that in this optional embodiment, a post-processing process for the second branch is set. When the target image data input into the multi-task framework is output as view classification data (including class probability vectors) via the second branch, through the set class post-processing process, it is converted into specific view classes, further improving the applicability of the multi-task framework.
[0179] Embodiment Three
[0180] Please refer to Figure 3 , Figure 3 which is a schematic structural diagram of a heart feature point positioning device based on a multi-task framework disclosed in an embodiment of the present invention. Among them, the heart feature point positioning device based on the multi-task framework can be a heart feature point positioning terminal, device, system or server based on the multi-task framework. The server can be a local server, a remote server, or a cloud server (also known as a cloud server). When the server is a non-cloud server, the non-cloud server can communicate with the cloud server. The embodiments of the present invention do not make limitations. As Figure 3 shown, the heart feature point positioning device based on the multi-task framework may include a pre-processing module 301, a multi-task training module 302, a heat map post-processing module 303, and a determination module 304, where:
[0181] The pre-processing module 301 is configured to perform a preset pre-processing operation on the acquired target image data to obtain a pre-processing result corresponding to the target image data. The pre-processing operation at least includes image scaling, normalization, and standardization processing; the pre-processing result is used to perform different training tasks during framework training;
[0182] The multi-task training module 302 is configured to input the pre-processing result into a multi-task framework to be trained according to a preset model training strategy to obtain first target data corresponding to the pre-processing result; the multi-task framework at least includes two training branches, and the two training branches include a first branch for heart feature point positioning and a second branch for echocardiogram view classification; among them, each time the multi-task framework performs framework training, only one of the two training branches is trained or the two training branches are trained simultaneously; the first target data includes predicted heat map data corresponding to the first branch and view classification data corresponding to the second branch;
[0183] The heat map post - processing module 303 is used to perform preset heat map post - processing on the first target data to obtain second target data corresponding to the first target data. The second target data includes multiple feature coordinates; all the feature coordinates are positioning coordinates corresponding to the heart feature points; the heat map post - processing operations include non - maximum suppression calculation, image scaling processing, and coordinate selection;
[0184] The determination module 304 is used to determine the second target data and multiple view categories corresponding to the view classification data as the target output data after determining that the multi - task framework has completed training.
[0185] It can be seen that implementing Figure 3 the described heart feature point positioning device based on the multi - task framework, by performing preset pre - processing on the acquired target image data, including image scaling, normalization, and standardization processing, helps to improve the stability, convergence speed, generalization ability, and training efficiency of the model; after obtaining the required pre - processing results, it can adjust the training branches of the multi - task framework according to the set model training strategy during different training epochs, thus avoiding the situation where the feature point positioning task (the first branch) can be normally trained while the view classification algorithm (the second branch) cannot converge normally due to sharing the feature point positioning data, that is, improving the application practicability and stability of the multi - task framework; among them, setting different training branches for different training epochs of the multi - task framework is beneficial to reducing the impact of the gradient of the feature point positioning (the first branch) on the backpropagation of the view classification branch (the second branch) during the early stage of the multi - task framework training when the multi - task framework is not stable enough, that is, the setting of this model training strategy is beneficial to reducing the loss impact between the first branch and the second branch during the early training stage; in the later stage of training, training the two branches simultaneously through the model training strategy setting is beneficial to improving the training accuracy and reliability of the multi - task framework.
[0186] In an optional embodiment, the specific manner in which the multi - task training module 302 inputs the pre - processing result into the multi - task framework to be trained according to the preset model training strategy to obtain the first target data corresponding to the pre - processing result includes:
[0187] According to the preset model training strategy, determine the training type of the current execution of the framework training of the multi - task framework. The training type includes the early - stage training type or the late - stage training type;
[0188] Determine the target training branch corresponding to the current training type according to the model training probability set in the model training strategy; when the training type is the early training type, the target training branch is the first branch or the second branch; when the training type is the late training type, the target training branch is the third branch; the third branch is used to train the first branch, the second branch, and a preset attention module as a whole, where the attention module is used to connect the training tasks corresponding to the first branch and the second branch;
[0189] Input the preprocessing result into the target training branch to obtain branch output data corresponding to the preprocessing result;
[0190] Among them, when the training type is the late training type and it is determined that the multi-task framework is trained to convergence, determine that the branch output data is the first target data corresponding to the preprocessing result.
[0191] In this optional embodiment, the attention module is used to transfer the target information output by the second branch to the first branch, and the target information includes channel attention and spatial attention;
[0192] When the training type is the late training type, the attention module is also used to sequentially perform a first operation operation corresponding to channel attention and a second operation operation corresponding to spatial attention on the predicted heat map data corresponding to the first branch and the view classification data corresponding to the second branch.
[0193] It can be seen that in this optional embodiment, different training branches are set for different training periods of the multi-task framework, which is beneficial to reducing the influence of the gradient of feature point localization (the first branch) on the backpropagation of the view classification branch (the second branch) when the multi-task framework is not stable enough during the early training stage of the multi-task framework. That is, the setting of this model training strategy is beneficial to reducing the loss influence between the first branch and the second branch during the early training stage; in the late training stage, training the two branches simultaneously through the model training strategy setting is beneficial to improving the training accuracy and reliability of the multi-task framework.
[0194] In another optional embodiment, the first branch and the second branch share the same preset residual network in the encoding layer;
[0195] The specific way for the multi-task training module 302 to input the preprocessing result into the target training branch to obtain branch output data corresponding to the preprocessing result includes:
[0196] Before determining that the multi-task framework is trained to convergence, when the multi-task framework executes the first training task corresponding to the first branch, input the preprocessing result into the residual network to obtain a residual processing result corresponding to the preprocessing result, and the residual processing result corresponds to the predicted heat map data;
[0197] When the multi-task framework executes the second training task corresponding to the second branch, the pre-processing result is input into a preset view classification network to obtain a view classification result corresponding to the pre-processing result, and the view classification result corresponds to view classification data;
[0198] Among them, when the training type is the early training type, the residual processing result and / or the view classification result are determined as the branch output data corresponding to the pre-processing result;
[0199] When the training type is the late training type, the residual processing result and the view classification result are used to be transmitted to the attention module; meanwhile, the attention processing result output by the attention module through the residual processing result and the view classification result is determined as the branch output data corresponding to the pre-processing result;
[0200] Among them, the view classification network includes a feature extraction layer, a global pooling layer, a multi-layer perceptron MLP, and a normalization layer; the feature extraction layer is a residual network.
[0201] In this optional embodiment, it should be noted that during the process of performing framework training on the multi-task framework, when the training type is the early training type, only the loss value corresponding to the target training branch is calculated, and the network parameters adopted by the target training branch are updated through the loss value corresponding to the target training branch;
[0202] When the training type is the late training type, the first loss value corresponding to the first branch and the second loss value corresponding to the second branch are calculated simultaneously; then, in combination with the set gradient calculation formula, gradient calculation is performed on the first loss value and the second loss value, and then the network parameters adopted by the target training branch are updated through the obtained gradient calculation result.
[0203] In this optional embodiment, during the process of performing framework training on the multi-task framework, and when the training type is the late training type, the loss functions used to perform loss value calculation include the DWA function and the DTP function;
[0204] Among them, the calculation formula corresponding to the DWA function is:
[0205]
[0206] Among them, w t (n) represents the weight occupied by the Loss of the training task t in the nth round of training; represents the task set corresponding to all training tasks when performing framework training; the Loss of the task t in the nth round of training is loss t (n), then the training speed of the task t in the nth round of training is r t (n) = Loss t(n) / Loss t (n - 1); T is the temperature constant;
[0207] The calculation formula corresponding to this DTP function is:
[0208]
[0209] where k t is used to represent the KPI of each training task t; and k t ∈ [0, 1]; γ t is used to update the network weights for a certain training task.
[0210] It can be seen that in this optional embodiment, when training the training model, an update process for network parameters and network weights is set; at the same time, based on the training situation of the framework loss value, monitoring of the training type is achieved, corresponding to the monitoring of the early training type and the late training type; it is beneficial to improve the applicability of the training model of the overall multi - task framework.
[0211] In another optional embodiment, as Figure 4 shown, the device further includes a category post - processing module 305 and an update module 306, where:
[0212] The category post - processing module 305 is used to perform a preset category post - processing on all obtained view category probability vectors to obtain at least one view category corresponding to the view classification data before the heat - map post - processing module 303 performs a preset heat - map post - processing on the first target data to obtain the second target data corresponding to the first target data;
[0213] The update module 306 is used to update the view classification data according to all view categories, and trigger the heat - map post - processing module 303 to perform the above - mentioned preset heat - map post - processing on the first target data to obtain the second target data corresponding to the first target data;
[0214] Among them, the way that the category post - processing module 305 performs a preset category post - processing on all obtained view category probability vectors to obtain at least one view category corresponding to the view classification data specifically includes:
[0215] Performing a maximum value index on all obtained view category probability vectors according to a preset index function to obtain an index result corresponding to all view category probability vectors; the index function includes the Argmax function;
[0216] Determining at least one view category corresponding to the index result according to the index result and a preset category encoding; the category encoding is an encoding determined in advance for the target image data.
[0217] It can be seen that in this alternative embodiment, a post - processing process for the second branch is set up. When the target image data input into the multitask framework is output as view classification data (including the class probability vector) via the second branch, through the set class post - processing process, it is converted into specific view classes, further improving the applicability of this multitask framework.
[0218] Embodiment Four
[0219] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of another heart feature point positioning device based on a multitask framework disclosed in the embodiments of the present invention. As Figure 5 shown, the heart feature point positioning device based on the multitask framework may include:
[0220] A memory 401 storing executable program code;
[0221] A processor 402 coupled to the memory 401;
[0222] The processor 402 calls the executable program code stored in the memory 401 and executes the steps in the heart feature point positioning method based on the multitask framework described in Embodiment One or Embodiment Two of the present invention.
[0223] Embodiment Five
[0224] The embodiments of the present invention disclose a computer storage medium. The computer storage medium stores computer instructions, which are used to execute the steps in the heart feature point positioning method based on the multitask framework described in Embodiment One or Embodiment Two of the present invention when the computer instructions are called.
[0225] Embodiment Six
[0226] The embodiments of the present invention disclose a computer program product. The computer program product includes a non - transitory computer storage medium storing a computer program, and the computer program is operable to cause a computer to execute the steps in the heart feature point positioning method based on the multitask framework described in Embodiment One or Embodiment Two.
[0227] The device embodiments described above are only illustrative. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0228] Through the specific descriptions of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solutions, in essence, or the parts that contribute to the prior art can be embodied in the form of a software product. This computer software product can be stored in a computer storage medium, and the storage medium includes read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc memories, magnetic disk memories, tape memories, or any other computer-readable medium that can be used to carry or store data.
[0229] Finally, it should be noted that: The method and device for locating cardiac feature points based on a multi-task framework disclosed in the embodiments of the present invention only disclose the preferred embodiments of the present invention, and are only used to illustrate the technical solutions of the present invention, rather than limiting them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A cardiac feature point positioning method based on a multi-task framework, characterized in that: The method comprises: Performing a preset pre-processing operation on the acquired target image data to obtain a pre-processing result corresponding to the target image data, wherein the pre-processing operation at least includes image scaling, normalization and standardization processing; the pre-processing result is used to perform different training tasks during framework training; According to a preset model training strategy, the pre-processing result is input into a multi-task framework to be trained to obtain first target data corresponding to the pre-processing result; the multi-task framework includes at least two training branches, the two training branches include a first branch for locating cardiac feature points and a second branch for classifying echocardiographic views; wherein, each time the multi-task framework performs the framework training, only one of the two training branches is trained or both training branches are trained simultaneously; the first target data includes predicted thermal map data corresponding to the first branch and view classification data corresponding to the second branch; Performing a preset heat map post-processing on the first target data to obtain second target data corresponding to the first target data, wherein the second target data includes a plurality of feature coordinates; all of the feature coordinates are positioning coordinates corresponding to cardiac feature points; the heat map post-processing operation includes non-maximum suppression calculation, image scaling processing, and coordinate selection; After determining that the multi-task framework has completed training, determining the second target data and a plurality of view categories corresponding to the view classification data as target output data; The method of inputting the pre-processing result into the multi-task framework to be trained according to the preset model training strategy to obtain the first target data corresponding to the pre-processing result includes: According to a preset model training strategy, determining a training type of the framework training currently executed by the multi-task framework, wherein the training type includes an early training type or a late training type; According to the model training probability set in the model training strategy, determine the target training branch currently corresponding to the training type; when the training type is the early training type, the target training branch is the first branch or the second branch; when the training type is the late training type, the target training branch is the third branch; the third branch is used to train the first branch, the second branch and the preset attention module as a whole, wherein the attention module is used to connect the training tasks corresponding to the first branch and the second branch; Inputting the pre-processing result into the target training branch to obtain branch output data corresponding to the pre-processing result; Wherein, when the training type is the late training type and it is determined that the multi-task framework is trained to convergence, the branch output data is determined to be the first target data corresponding to the pre-processing result.
2. The method for locating cardiac feature points based on a multi-task framework according to claim 1, characterized in that: The attention module is used to transfer the target information output by the second branch to the first branch, wherein the target information includes channel attention and spatial attention; When the training type is the late training type, the attention module is also used to sequentially perform a first operation corresponding to the channel attention and a second operation corresponding to the spatial attention on the predicted heat map data corresponding to the first branch and the view classification data corresponding to the second branch.
3. The cardiac feature point positioning method based on a multi-task framework according to claim 2, characterized in that: The first branch and the second branch share the same preset residual network at the coding layer; Before determining that the task framework is trained to convergence, inputting the pre-processing result into the target training branch to obtain branch output data corresponding to the pre-processing result includes: When the multi-task framework executes the first training task corresponding to the first branch, the pre-processing result is input into the residual network to obtain a residual processing result corresponding to the pre-processing result, and the residual processing result corresponds to the predicted heat map data; When the multi-task framework executes the second training task corresponding to the second branch, the pre-processing result is input into a preset view classification network to obtain a view classification result corresponding to the pre-processing result, wherein the view classification result corresponds to the view classification data; Wherein, when the training type is the pre-training type, the residual processing result and / or the view classification result is determined as branch output data corresponding to the pre-processing result; When the training type is the post-training type, the residual processing result and the view classification result are used to be transmitted to the attention module; and the residual processing result and the view classification result are determined as branch output data corresponding to the pre-processing result through the attention processing result output by the attention module; The view classification network includes a feature extraction layer, a global pooling layer, a multi-layer perceptron MLP and a normalization layer; the feature extraction layer is the residual network.
4. The method for locating cardiac feature points based on a multi-task framework according to claim 3, characterized in that: In the process of performing the framework training on the multi-task framework, when the training type is the early training type, only the loss value corresponding to the target training branch is calculated, and the network parameters adopted by the target training branch are updated according to the loss value corresponding to the target training branch; When the training type is the late training type, a first loss value corresponding to the first branch and a second loss value corresponding to the second branch are calculated simultaneously; then, in combination with the set gradient calculation formula, gradient calculation is performed on the first loss value and the second loss value, and then the network parameters adopted by the target training branch are updated through the obtained gradient calculation results.
5. The cardiac feature point positioning method based on a multi-task framework according to claim 1, 2 or 4, characterized in that: In the process of performing the framework training on the multi-task framework, when the training type is the late training type, the loss function corresponding to the loss value calculation includes a DWA function and a DTP function; The calculation formula corresponding to the DWA function is: Among them, w t (n) represents the weight of the Loss of training task t in the nth round of training; represents the task set corresponding to all training tasks when executing the framework training; the Loss of task t in the nth round of training is Loss t (n), then the training speed of task t in the nth round of training is r t (n) = Loss t (n) / Loss t (n-1); T is the temperature constant; The calculation formula corresponding to the DTP function is: Among them, k t (n) is used to represent the KPI of the training task t in the nth round of training; and k t (n)∈[0,1]; γ t Used to perform network weight updates for a certain training task.
6. The method for locating cardiac feature points based on a multi-task framework according to claim 5, characterized in that: The view classification data includes a plurality of view category probability vectors; Before performing a preset heat map post-processing on the first target data to obtain second target data corresponding to the first target data, the method further includes: performing preset category post-processing on all the acquired view category probability vectors to obtain at least one view category corresponding to the view classification data; According to all the view categories, the view classification data is updated, and the preset heat map post-processing is triggered to be performed on the first target data to obtain second target data corresponding to the first target data; The performing of preset category post-processing on all the acquired view category probability vectors to obtain at least one view category corresponding to the view classification data includes: According to a preset indexing function, performing maximum indexing on all the obtained view category probability vectors to obtain index results corresponding to all the view category probability vectors; the indexing function includes an Argmax function; At least one view category corresponding to the index result is determined according to the index result and a preset category code; the category code is a predetermined code for the target image data.
7. A cardiac feature point positioning device based on a multi-task framework, characterized in that: The device comprises: A pre-processing module is used to perform a preset pre-processing operation on the acquired target image data to obtain a pre-processing result corresponding to the target image data, wherein the pre-processing operation at least includes image scaling, normalization and standardization processing; the pre-processing result is used to perform different training tasks during framework training; A multi-task training module, used for inputting the pre-processing result into the multi-task framework to be trained according to a preset model training strategy, and obtaining first target data corresponding to the pre-processing result; the multi-task framework includes at least two training branches, the two training branches include a first branch for locating cardiac feature points and a second branch for classifying echocardiographic views; wherein, each time the multi-task framework executes the framework training, only one of the two training branches is trained or both training branches are trained simultaneously; the first target data includes predicted heat map data corresponding to the first branch and view classification data corresponding to the second branch; A heat map post-processing module, configured to perform a preset heat map post-processing on the first target data to obtain second target data corresponding to the first target data, wherein the second target data includes a plurality of feature coordinates, and all of the feature coordinates are positioning coordinates corresponding to cardiac feature points; the heat map post-processing operation includes non-maximum suppression calculation, image scaling processing, and coordinate selection; a determination module, configured to determine the second target data and a plurality of view categories corresponding to the view classification data as target output data after determining that the multi-task framework has completed training; The multi-task training module inputs the pre-processing result into the multi-task framework to be trained according to the preset model training strategy, and the method of obtaining the first target data corresponding to the pre-processing result specifically includes: According to a preset model training strategy, determining a training type of the framework training currently executed by the multi-task framework, wherein the training type includes an early training type or a late training type; According to the model training probability set in the model training strategy, determine the target training branch currently corresponding to the training type; when the training type is the early training type, the target training branch is the first branch or the second branch; when the training type is the late training type, the target training branch is the third branch; the third branch is used to train the first branch, the second branch and the preset attention module as a whole, wherein the attention module is used to connect the training tasks corresponding to the first branch and the second branch; Inputting the pre-processing result into the target training branch to obtain branch output data corresponding to the pre-processing result; Wherein, when the training type is the late training type and it is determined that the multi-task framework is trained to convergence, the branch output data is determined to be the first target data corresponding to the pre-processing result.
8. A cardiac feature point positioning device based on a multi-task framework, characterized in that: The device comprises: A memory storing executable program code; a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the cardiac feature point positioning method based on a multi-task framework as described in any one of claims 1-6.
9. A computer storage medium, characterized in that: The computer storage medium stores computer instructions, which, when called, are used to execute the cardiac feature point positioning method based on a multi-task framework as described in any one of claims 1-6.