Semi-supervised key point positioning method and semi-supervised key point positioning equipment for swallowing contrast analysis
By combining the combined training of marked and unlabeled data in swallowing contrast analysis, and introducing semantic guidance module and Kalman filtering algorithm, the problem of insufficient accuracy and robustness of key point positioning in swallowing contrast analysis in the prior art is solved, and a more efficient and accurate key point positioning effect is achieved.
Patent Information
- Application Number
- CN202411922583.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-25
AI Technical Summary
The existing semi-supervised key point positioning technology has many limitations in the field of swallowing contrast analysis, including the lack of semantic guidance strategies for specific areas of swallowing contrast, the difficulty in making full use of background knowledge to improve model robustness, and the failure to effectively consider timing consistency between video frames.
A semi-supervised key point positioning method for swallowing contrast analysis is adopted, and an efficient and accurate key point positioning framework is constructed by combining the combined training of marked data and unlabeled data. Specific steps include: obtaining swallowing contrast images and annotating anatomical key points, designing joint optimization strategies to integrate supervision losses and self-supervising consistency losses, building a semantic guidance module to enhance the model's ability to capture key point regional features, and introducing Kalman filtering algorithms in the video processing stage to optimize timing consistency.
Through this method, the accuracy and robustness of key point positioning are significantly improved, especially in complex backgrounds and fuzzy targets, which can provide more accurate key point kinematic parameters and support efficient analysis of clinical scenarios such as swallowing disorder diagnosis.
Smart Images

Figure CN119942049A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of image processing, computer vision, medical image analysis, and in particular to a semi-supervised key point positioning method and device for swallowing radiography analysis. Background Art
[0002] Dysphagia is a major health problem that affects human eating function. Its accurate diagnosis and treatment are of great significance to improving the quality of life of patients. As the gold standard for clinical evaluation of swallowing function, swallowing contrast examination can dynamically collect videos of the oropharyngeal-esophageal swallowing process through X-rays, comprehensively record the structural and functional changes of the laryngeal and esophageal regions, and provide a reliable basis for the diagnosis of dysphagia. However, due to the complexity of the anatomical structure and the fuzzy target boundaries and severe noise interference in the contrast video, there are significant challenges in accurately measuring the temporal and kinematic parameters involved in the swallowing process. As the core feature of the anatomical structure, the key point can not only focus and locate the moving target, but also provide a basis for the calculation of kinematic parameters (such as hyoid displacement) and key indicators. Therefore, achieving accurate positioning of key points is of great value for the quantitative evaluation of swallowing function and the optimization of clinical diagnosis and treatment. Furthermore, the automated key point positioning method will provide an efficient solution for the objective analysis and large-scale clinical application of dysphagia.
[0003] At present, key point localization technology is widely used in the field of computer vision. It extracts structural features in images or videos through deep learning models to achieve the purpose of accurate positioning. In fully supervised key point localization, the model relies on a large number of manually annotated high-quality key point coordinates as training data. Such methods usually extract features from images or video frames through convolutional neural networks (CNN) or models based on self-attention mechanisms (such as Transformer), and predict the position of key points through coordinate regression or heat map regression. These methods can achieve high positioning accuracy when there is sufficient annotated data. However, the anatomical structure in swallowing angiography videos is complex, the annotation cost is high, and the subjectivity is high, which makes it difficult to obtain training data on a large scale, thus limiting the applicability of fully supervised methods.
[0004] In order to solve the problem of insufficient annotation, semi-supervised key point localization methods have emerged. This type of method combines a small amount of labeled data with a large amount of unlabeled data, and uses the potential structural information in the unlabeled data to improve model performance. Semi-supervised methods usually include strategies such as generating pseudo-labels, contrastive learning, or consistency constraints, and reduce dependence on manual annotation by mining the characteristics of unlabeled data. Especially in swallowing angiography videos, it is easier to obtain unlabeled data. Semi-supervised methods can improve the accuracy and robustness of key point localization under the premise of controllable annotation costs. In addition, when dealing with the fuzzy boundaries and complex dynamic changes of specific anatomical structures in videos, semi-supervised methods can better utilize time series information and complement multimodal features, thereby enhancing the model's ability to capture key points. These characteristics make semi-supervised key point localization a very promising research direction in swallowing angiography analysis.
[0005] However, existing semi-supervised key point localization technologies have many limitations in the field of swallowing radiography analysis. Current methods lack semantic guidance strategies for specific areas of swallowing radiography, and it is difficult to fully utilize background knowledge to improve the robustness of the model to complex backgrounds and small target areas, especially when the swallowing action target is blurred or the background is complex. At the same time, these technologies fail to effectively consider the temporal consistency between video frames, which may cause the key point prediction results to fluctuate greatly in the time dimension, making it difficult to meet the strict clinical requirements for high precision and temporal stability of the action trajectory. Summary of the invention
[0006] In order to solve at least one of the technical problems existing in the prior art to a certain extent, the object of the present invention is to provide a semi-supervised key point positioning method, device and medium for swallowing radiography analysis.
[0007] The first technical solution adopted by the present invention is:
[0008] A semi-supervised key point positioning method for swallowing radiography analysis includes the following steps:
[0009] Acquire a swallowing angiography image, annotate anatomical key points in the swallowing angiography image, and obtain annotated data;
[0010] Design a joint optimization strategy to integrate the supervised loss of labeled data and the self-supervised consistency loss of unlabeled data in a unified training framework; generate pseudo labels using unlabeled data and dynamically update pseudo labels during model training, and jointly train the model using supervised and unsupervised data;
[0011] Build a semantic guidance module to assist the model in capturing the features of key point areas more accurately, thereby improving the positioning performance of complex backgrounds and blurred key points;
[0012] The Kalman filter algorithm is introduced in the video processing stage to achieve key point timing calibration by fusing the key point prediction results of multiple frames.
[0013] Furthermore, the step of acquiring the swallowing contrast image and annotating the anatomical key points in the swallowing contrast image to obtain the annotated data includes:
[0014] Based on professional medical anatomical knowledge and the characteristics of swallowing angiography images, key parts that can accurately reflect the changes in anatomical structures during swallowing are selected as anatomical key points that need to be marked;
[0015] Using professional image annotation tools, medical imaging professionals accurately annotate the selected key points on the swallowing angiography images to form a small amount of annotated data sets, providing a reliable supervised data foundation for subsequent model training.
[0016] Furthermore, the joint optimization strategy is designed to integrate the supervised loss of labeled data and the self-supervised consistency loss of unlabeled data in a unified training framework; pseudo labels are generated using unlabeled data, and the pseudo labels are dynamically updated during the model training process, and the model is jointly trained using supervised data and unsupervised data, including:
[0017] Construct a unified training framework that includes a labeled data supervision branch and an unlabeled data self-supervision branch. In the labeled data supervision branch, a regression loss function is used to measure the difference between the model prediction results and the true labels of the labeled data to guide the model to learn the key features in the labeled data.
[0018] In the unlabeled data self-supervision branch, a consistency loss function based on data enhancement is designed to mine the potential features in the unlabeled data;
[0019] During the model training process, the prediction results of unlabeled data are used to generate pseudo labels. By setting the confidence threshold, the prediction results with high confidence are selected as pseudo labels, and these pseudo labels are used together with the labeled data in the next round of model training. At the same time, according to the iterative training of the model, the confidence evaluation and selection strategy of the pseudo labels are continuously updated to dynamically improve the reliability of the pseudo labels, thereby realizing the joint effective training of supervised data and unlabeled data, so as to improve the generalization ability of the model and its ability to adapt to complex situations.
[0020] Furthermore, for the labeled data, the heat map is used as the supervision signal of the key points. Given a labeled image x s and its corresponding annotation key point coordinate set Where K is the number of key points, and a Gaussian heat map is generated for each key point as a supervision signal:
[0021]
[0022] Among them, (u, v) is the image pixel coordinate, σ is the standard deviation of the Gaussian distribution, which controls the diffusion range of the heat map;
[0023] The regression loss function is the mean square error loss function; the key point heat map predicted by the model is The supervised loss of a single sample is calculated using the mean squared error (MSE) loss function:
[0024]
[0025] In the formula, H k Indicates a key point.
[0026] Furthermore, for unlabeled data, a teacher-student training paradigm is adopted to construct a teacher model f T and student model f S ; For unlabeled images x u Generate input data using weak data enhancement and strong data enhancement respectively and
[0027] The data Input to the teacher model f T In the output, the corresponding heat map prediction Then the heat map prediction Perform weak to strong conversion and generate pseudo labels The data Input student model f S The output heat map is
[0028] Only when the maximum response value of the pseudo label exceeds the confidence threshold τ, it is used for student model training; otherwise, the pseudo label does not participate in the training, and the expression is:
[0029]
[0030] in, is the indicator function, when When , the pseudo label is set to zero and does not participate in training; τ(t) is the dynamic confidence threshold.
[0031] Furthermore, the dynamic confidence threshold τ(t) increases linearly with the training process, and the expression is:
[0032]
[0033] In the formula, t is the current training round, T is the total training round, τ init and τ end represents the initial and final values of the confidence threshold;
[0034] The loss function for training the student model using unlabeled data is defined as:
[0035]
[0036] Furthermore, the semantic guidance module is constructed to assist the model in capturing the features of the key point area more accurately, including:
[0037] The pre-trained DINOv2 model was used to extract semantic features from swallowing radiography images;
[0038] A semantic guidance module is designed, and the features extracted from swallowing radiography images by the key point localization model are used to align semantic features, so that the model can focus more on the semantic information of the key point area, thereby accurately locating key anatomical structures in complex backgrounds and improving the positioning accuracy and robustness of the model.
[0039] Furthermore, the expression of the extracted semantic features is:
[0040] F sec =f DINOv2 (x)
[0041] In the formula, f DINOv2 Represents the DINOv2 model;
[0042] The key point positioning model is implemented by its own encoder f enc Extract the features F of swallowing radiography images model :
[0043] F model =f enc (x)
[0044] The semantic guidance module calculates F sec and F model The cosine similarity of is used to align the semantic features, so as to guide the model to focus more on the semantic information of the key point area; the calculation formula of cosine similarity is:
[0045]
[0046] In the formula, ‖·‖ represents the bi-norm of the vector.
[0047] Furthermore, the Kalman filter algorithm is introduced in the video processing stage to achieve key point timing calibration by fusing key point prediction results of multiple frames, including:
[0048] After locating the key points of the swallowing angiography video frame by frame, the key point prediction results of each frame are used as the input of the Kalman filter algorithm;
[0049] The Kalman filter algorithm estimates and predicts the state of key points by establishing a state space model. It combines the observations of the current frame (i.e., the key point prediction results of the model) and the state estimation values of the previous frame to calculate a more accurate and smooth key point position estimate for the current frame, thereby effectively reducing the jitter and discontinuity problems of the positioning trajectory and improving the timing consistency and accuracy of key point detection in video sequences.
[0050] The second technical solution adopted by the present invention is:
[0051] An electronic device comprises a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement a semi-supervised key point positioning method for swallowing radiography analysis as described above.
[0052] The third technical solution adopted by the present invention is:
[0053] A computer-readable storage medium, wherein at least one instruction, at least one program, a code set or an instruction set is stored in the storage medium, wherein the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement a semi-supervised key point positioning method for swallowing radiography analysis as described above.
[0054] The fourth technical solution adopted by the present invention is:
[0055] A computer program product or a computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above method.
[0056] The beneficial effects of the present invention are as follows: the present invention constructs an efficient and accurate key point positioning framework by combining joint training of labeled data and unlabeled data. In addition, the present invention enhances the model's ability to capture key point regional features by introducing a semantic guidance module, and optimizes temporal consistency through Kalman filtering, effectively improving the accuracy and robustness of key point positioning, especially in the case of complex backgrounds and blurred targets. The present invention will provide more accurate key point kinematic parameters for clinical scenarios such as dysphagia diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the embodiments of the present invention or the drawings of related technical solutions in the prior art are introduced below. It should be understood that the drawings introduced below are only for the convenience of clearly describing some embodiments of the technical solutions of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0058] Figure 1 is a flowchart of the steps of a semi-supervised key point positioning method for swallowing angiography analysis in an embodiment of the present invention;
[0059] Figure 2 Schematic diagram of anatomical key points in an embodiment of the present invention. DETAILED DESCRIPTION
[0060] The embodiments of the present invention are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and are not to be construed as limitations of the present invention. For the step numbers in the following embodiments, they are only provided for the convenience of explanation, and the order between the steps is not limited in any way, and the execution order of each step in the embodiment can be adaptively adjusted according to the understanding of those skilled in the art.
[0061] In the description of the present invention, it should be understood that descriptions involving orientations, such as up, down, front, back, left, right, etc., and orientations or positional relationships indicated are based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present invention.
[0062] In the description of the present invention, "several" means one or more, "more" means more than two, "greater than", "less than", "exceed" etc. are understood as not including the number itself, and "above", "below", "within" etc. are understood as including the number itself. If there is a description of "first" or "second", it is only used for the purpose of distinguishing the technical features, and cannot be understood as indicating or implying the relative importance or implicitly indicating the number of the indicated technical features or implicitly indicating the order of the indicated technical features.
[0063] In the description of the present invention, unless otherwise clearly defined, terms such as setting, installing, connecting, etc. should be understood in a broad sense, and technicians in the relevant technical field can reasonably determine the specific meanings of the above terms in the present invention based on the specific content of the technical solution.
[0064] In view of the existing technical problems, the present invention proposes a semi-supervised key point positioning scheme for swallowing angiography analysis. By combining the joint training of supervised data and unsupervised data, an efficient and accurate key point positioning framework is constructed, aiming to solve the technical problems such as complex anatomical structure and scarce key point annotation in swallowing angiography videos. Specifically, firstly, the anatomical key points in the swallowing angiography images are annotated based on clinical knowledge to form a small amount of high-quality annotated data sets; secondly, a joint optimization strategy is designed to integrate the supervised loss of the annotated data and the self-supervised consistency loss of the unlabeled data in a unified training framework, and pseudo labels are generated using unlabeled data, and the pseudo labels are dynamically updated during the model training process to improve their reliability; then, with the help of the semantic feature extraction capability of the DINOv2 model, a semantic guidance module is constructed to assist the model in more accurately capturing the features of the key point area and improving the positioning performance of complex backgrounds and blurred key points; finally, the Kalman filter algorithm is introduced in the video processing stage to smooth the positioning trajectory by fusing the key point prediction results of multiple frames, effectively improving the temporal consistency and accuracy of key point detection.
[0065] Example 1
[0066] like Figure 1 As shown, this embodiment provides a semi-supervised key point positioning method for swallowing radiography analysis, comprising the following steps:
[0067] S1. Obtain a swallowing angiography image, annotate anatomical key points in the swallowing angiography image, and obtain annotated data.
[0068] Based on clinical knowledge, the anatomical key points in the swallowing radiography images are annotated to form a small amount of high-quality annotated data sets, which provide a basis for subsequent model training. Specifically, step S1 includes the following steps:
[0069] S11. Based on professional medical anatomical knowledge and the characteristics of swallowing radiography images, select key parts that can accurately reflect the changes in anatomical structures during swallowing, such as the soft palate and hyoid bone, as anatomical key points that need to be marked.
[0070] S12. Using professional image annotation tools, experienced medical imaging professionals accurately annotate the selected key points on the swallowing radiography images and record their coordinate positions and other information. After strict quality review, a small but highly accurate and representative annotation dataset is formed, providing a reliable supervised data foundation for subsequent model training.
[0071] S2. Design a joint optimization strategy to integrate the supervised loss of labeled data and the self-supervised consistency loss of unlabeled data in a unified training framework; use unlabeled data to generate pseudo labels, and dynamically update the pseudo labels during the model training process, and use supervised data and unsupervised data to jointly train the model.
[0072] In some embodiments, step S2 specifically includes the following steps:
[0073] S21. Construct a unified training framework that includes a labeled data supervision branch and an unlabeled data self-supervision branch. In the labeled data supervision branch, use a common regression loss function (such as mean square error loss) to measure the difference between the model prediction result and the true label of the labeled data, and guide the model to learn the key features in the labeled data.
[0074] S22. In the unlabeled data self-supervision branch, a consistency loss function based on data enhancement is designed. For example, after performing random transformations (such as rotation, flipping, etc.) on the unlabeled data, the model is required to maintain consistency in the prediction results of the data before and after the transformation, so as to explore the potential features in the unlabeled data.
[0075] S23. During the model training process, pseudo labels are generated using the prediction results of unlabeled data. By setting a certain confidence threshold, high-confidence prediction results are selected as pseudo labels, and these pseudo labels are used together with labeled data in the next round of model training. At the same time, according to the iterative training of the model, the confidence evaluation and selection strategy of the pseudo labels are continuously updated to dynamically improve the reliability of the pseudo labels, thereby achieving effective joint training of supervised data and unlabeled data, and improving the model's generalization ability and adaptability to complex situations.
[0076] S3. Build a semantic guidance module to assist the model in capturing the features of key point areas more accurately, thereby improving the positioning performance of complex backgrounds and blurred key points.
[0077] Specifically, with the help of the semantic feature extraction capability of the DINOv2 model, a semantic guidance module is constructed to assist the model in more accurately capturing the features of the key point area, thereby improving the positioning performance of complex backgrounds and blurred key points. Step S3 specifically includes the following steps:
[0078] S31. The DINOv2 model is introduced into the key point localization framework as a pre-trained model, and its excellent semantic feature extraction capability is used to extract semantic features from swallowing angiography images.
[0079] S32. Based on the extracted semantic features, a semantic guidance module is designed. The features extracted by the key point positioning model encoder are aligned with the semantic features extracted by DINOv2 using the cosine similarity loss, so that the model can focus more on the semantic information of the key point area, thereby accurately locating the key anatomical structure in a complex background and improving the positioning accuracy and robustness of the model.
[0080] S4. Introduce the Kalman filter algorithm in the video processing stage to achieve key point timing calibration by fusing the key point prediction results of multiple frames.
[0081] The Kalman filter algorithm is introduced in the video processing stage to smooth the positioning trajectory by fusing the key point prediction results of multiple frames, thereby effectively improving the temporal consistency and accuracy of key point detection. This method constructs an efficient and accurate key point positioning framework, solving technical problems such as complex anatomical structures and scarce key point annotations in swallowing angiography videos. In some embodiments, step S4 specifically includes the following steps:
[0082] S41. After locating the key points of the swallowing radiography video frame by frame, the key point prediction result of each frame is used as the input of the Kalman filter algorithm.
[0083] S42. The Kalman filter algorithm estimates and predicts the position, speed and other states of key points by establishing a state-space model. It combines the observation value of the current frame (i.e., the key point prediction result of the model) and the state estimation value of the previous frame to calculate a more accurate and smooth key point position estimate for the current frame, thereby effectively reducing the jitter and discontinuity of the positioning trajectory caused by factors such as noise between video frames and key point detection errors, improving the timing consistency and accuracy of key point detection in the video sequence, and making the key point positioning results of the entire swallowing angiography video more stable and reliable, providing more accurate data support for subsequent clinical analysis.
[0084] The above method is explained in detail below with reference to the accompanying drawings and specific embodiments.
[0085] This embodiment provides a semi-supervised key point positioning method for swallowing radiography analysis, which specifically includes the following steps:
[0086] Step 1: Identify and annotate anatomical key points in swallowing images.
[0087] In this step, based on medical expertise and clinical experience, the present invention determines anatomical parts closely related to the swallowing process as key points. These parts usually include the soft palate, hyoid bone, etc. in the swallowing video. Specific key point categories are: suprahyoid convex point, subhyoid convex point, left hyoid convex point, left end point of soft palate, right end point of soft palate, soft palate peak point, left lower vertex of the second vertebra, left lower vertex of the fourth vertebra, such as Figure 2 shown.
[0088] Using professional image annotation tools, experienced medical imaging professionals accurately annotated the selected key points on the swallowing radiography images and recorded their coordinate positions and other information. After strict quality review, 13,895 annotated images were obtained, providing a reliable supervised data basis for subsequent model training. In addition, 31,536 unannotated images were collected, providing a richer feature learning basis for the model.
[0089] Step 2: Design a joint optimization strategy for labeled and unlabeled data.
[0090] First, for the labeled data, we use the heat map as the supervision signal of the key points. Given a labeled image x s and its corresponding annotation key point coordinate set Where K is the number of key points. In this embodiment, K=8. A Gaussian heat map is generated for each key point as a supervision signal:
[0091]
[0092] Among them, (u,v) is the image pixel coordinate, σ is the standard deviation of the Gaussian distribution, which controls the diffusion range of the heat map. The key point heat map predicted by the model is The supervised loss of a single sample is calculated using the mean squared error (MSE) loss function:
[0093]
[0094] This loss guides the model to learn key point features in the labeled data and optimize the accuracy of heat map prediction.
[0095] For unlabeled data, a teacher-student training paradigm is adopted to construct a teacher model f T and student model f S For unlabeled images x u Generate input data using weak data enhancement (such as simple preprocessing) and strong data enhancement (such as affine transformation such as rotation) and Bundle Input to the teacher model f T In the output, the corresponding heat map prediction Then Perform weak to strong transformation, for example, generate pseudo labels through the same affine transformation Correspondingly, the input is strongly enhanced Give the student model f S , the heat map output by the student model is
[0096] To ensure the quality of the pseudo-label, only when the maximum response value of the pseudo-label exceeds the confidence threshold τ, it will be used for student model training. Otherwise, the pseudo-label does not participate in the training, which can be expressed mathematically as:
[0097]
[0098] Here, is the indicator function, when When , the pseudo label is set to zero and does not participate in training. τ(t) is a dynamic confidence threshold, which can be continuously increased with the number of training iterations. Optionally, τ(t) here increases linearly with the training process, such as:
[0099]
[0100] Among them, t is the current training round, T is the total training round, τ init and τ end represents the initial and final values of the confidence threshold. Then the loss function of training the student model using unlabeled data is defined as:
[0101]
[0102] Pseudo labels can not only reflect the complex characteristics of unlabeled data, but also ensure high confidence reliability, thereby effectively supervising the student model to learn key point features under complex transformations.
[0103] Step 3: Construct a semantic guidance module to assist the model in capturing key point area features.
[0104] In order to improve the model's ability to capture key point regional features in swallowing radiography images, this embodiment proposes a semantic guidance module, which uses the semantic feature extraction capability of the DINOv2 model to guide the key point positioning model to focus on the key area more accurately. Use the pre-trained DINOv2 model swallowing radiography image x (both annotated and unannotated images are applicable) to extract high-dimensional semantic features F sec :
[0105] F sec =f DINOv2 (x)
[0106] Among them, f DINOv2 represents the DINOv2 model. At the same time, the key point localization model is implemented by its own encoder f enc Extract the features F of swallowing radiography images model :
[0107] F model =f enc (x)
[0108] The semantic guidance module calculates F sec and F model The cosine similarity is used to align the semantic features, thereby guiding the model to focus more on the semantic information of the key point area. The calculation formula of cosine similarity is:
[0109]
[0110] Among them, ‖·‖ represents the bi-norm of the vector. Then the semantic-assisted loss function L sec Defined as:
[0111]
[0112] Where N represents the number of pixels on the feature map, and They represent the feature vectors at the i-th position respectively.
[0113] Combining the two loss functions in step 2, the total loss function L is:
[0114]
[0115] Among them, λ1 and λ2 are hyperparameters that adjust the weight of the loss function. The stochastic gradient descent algorithm (SGD) is used to optimize the loss function, so as to train an accurate key point positioning model.
[0116] Step 4: Use the Kalman filter-guided inter-frame key point estimation method to achieve key point timing calibration.
[0117] In order to improve the timing consistency and accuracy of video key point positioning, this embodiment introduces a Kalman filter algorithm in the key point detection process, and effectively reduces noise interference and positioning errors through multi-frame information fusion and trajectory smoothing.
[0118] First, the key point positioning model is used to predict each frame of the swallowing angiography video frame by frame to obtain the two-dimensional coordinate prediction results of n key points in each frame. Where t is the frame number and n is the number of key points. t As the observation value input of the Kalman filter algorithm. Define the state vector of the key point The state of each key point includes two-dimensional position and velocity:
[0119]
[0120] Among them, p 1x ,p 1y and v 1x ,v 1yRepresent the two-dimensional position and velocity of the i-th key point respectively. The state space model consists of a state transition model and an observation model. State transition model:
[0121] x t =Fx t-1 +w t-1
[0122] Among them, F is the state transfer matrix, which describes the relationship between position and velocity, for example:
[0123]
[0124] Δt represents the time interval between adjacent frames, I 2n and 0 2n They represent the identity matrix and zero matrix of dimension 2n×2n respectively, and w t-1 is the process noise, satisfying Q is the covariance matrix of process noise. Observation model:
[0125] z t =Hx t +v t
[0126] Where H is the observation matrix, which is used to extract position information from the state vector, for example, H = [I 2n 0 2n ],v t is the observation noise, satisfying R is the covariance matrix of the observation noise.
[0127] At each frame, the keypoint states are estimated and calibrated using the following steps. Prediction step:
[0128]
[0129]
[0130] Among them, among them, is the state estimate of the previous frame, P t-1|t-1 is its covariance matrix. Update steps:
[0131]
[0132] P t|t =(IK t H)P t|t-1
[0133] Among them, K t is the Kalman gain, which is used to balance the weights of observed values and predicted values.
[0134] The state estimate output by the Kalman filter Location information in This is the key point position after smooth calibration. Through the temporal smoothing effect of Kalman filtering, inter-frame jitter and discontinuity problems can be effectively reduced, and the temporal consistency and accuracy of key point detection can be significantly improved under complex background and noise conditions, providing high-quality key point kinematic parameters for subsequent clinical analysis.
[0135] In summary, the present invention proposes a semi-supervised key point localization method for swallowing radiography analysis, and constructs an efficient and accurate key point localization framework by combining joint training of labeled data and unlabeled data. Compared with existing methods, the present invention enhances the model's ability to capture key point regional features by introducing a semantic guidance module based on the DINOv2 model, and optimizes temporal consistency through Kalman filtering, which effectively improves the accuracy and robustness of key point localization, especially in the case of complex backgrounds and blurred targets. The present invention will provide more accurate key point kinematic parameters for clinical scenarios such as swallowing disorder diagnosis.
[0136] Example 2
[0137] An embodiment of the present invention further provides an electronic device, the electronic device comprising a processor and a memory, the memory storing at least one instruction, at least one program, a code set or an instruction set, the at least one instruction, the at least one program, the code set or the instruction set being loaded and executed by the processor to implement the following Figure 1 A semi-supervised key point localization method for swallowing angiography analysis is shown.
[0138] It is understood that the memory may include a random access memory (RAM) or a read-only memory (ROM). Optionally, the memory includes a non-transitory computer-readable storage medium. The memory may be used to store instructions, programs, codes, code sets, or instruction sets. The memory may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function, instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data created according to the use of the server, etc.
[0139] The processor may include one or more processing cores. The processor uses various interfaces and lines to connect the various parts of the entire server, and executes various functions of the server and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory, and calling data stored in the memory. Optionally, the processor can be implemented in at least one hardware form of digital signal processing (DSP), field programmable gate array (FPGA), and programmable logic array (PLA). The processor can integrate one or a combination of a central processing unit (CPU) and a modem. Among them, the CPU mainly processes the operating system and application programs; the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor, but implemented separately through a chip.
[0140] Since the electronic device is an electronic device corresponding to a semi-supervised key point positioning method for swallowing radiography analysis in an embodiment of the present invention, and the principle of solving the problem by the electronic device is similar to that of the method, the implementation of the electronic device can refer to the implementation process of the above-mentioned method embodiment, and the repeated parts will not be repeated.
[0141] Example 3
[0142] The embodiment of the present invention further provides a computer-readable storage medium, wherein the storage medium stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, the at least one program, the code set or instruction set is loaded and executed by a processor to implement the following Figure 1 A semi-supervised key point localization method for swallowing angiography analysis is shown.
[0143] Those skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, and the storage medium includes a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable rewritable read-only memory (EEPROM), a compact disc (CD-ROM) or other optical disc storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium that can be used to carry or store data.
[0144] Since the storage medium is a storage medium corresponding to a semi-supervised key point positioning method for swallowing radiography analysis in an embodiment of the present invention, and the principle of solving the problem by the storage medium is similar to that of the method, the implementation of the storage medium can refer to the implementation process of the above-mentioned method embodiment, and the repeated parts will not be repeated.
[0145] Example 4
[0146] In some possible implementations, various aspects of the method of the embodiment of the present invention can also be implemented in the form of a program product, which includes a program code. When the program product is run on a computer device, the program code is used to enable the computer device to execute the steps of a semi-supervised key point positioning method for swallowing imaging analysis according to various exemplary embodiments of the present application described above in this specification. Among them, the executable computer program code or "code" for executing each embodiment can be written in a high-level programming language such as C, C++, C#, Smalltalk, Java, JavaScript, Visual Basic, structured query language (e.g., Transact-SQL), Perl, or in various other programming languages.
[0147] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0148] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0149] The above embodiments are only for illustrating the technical concept and features of the present invention, and their purpose is to enable ordinary technicians in the field to understand the content of the present invention and implement it accordingly, and they cannot be used to limit the protection scope of the present invention. Any equivalent changes or modifications made based on the essence of the content of the present invention should be included in the protection scope of the present invention.
Claims
1. A semi-supervised key point localization method for swallowing radiography analysis, characterized in that: The following steps are involved: Acquire a swallowing angiography image, annotate anatomical key points in the swallowing angiography image, and obtain annotated data; Design a joint optimization strategy to integrate the supervised loss of labeled data and the self-supervised consistency loss of unlabeled data in a unified training framework; generate pseudo labels using unlabeled data and dynamically update pseudo labels during model training, and jointly train the model using supervised and unsupervised data; Construct a semantic guidance module to assist the model in capturing the features of key point areas more accurately; The Kalman filter algorithm is introduced in the video processing stage to achieve key point timing calibration by fusing the key point prediction results of multiple frames.
2. A semi-supervised key point positioning method for swallowing imaging analysis according to claim 1, characterized in that: The step of acquiring a swallowing radiography image and annotating anatomical key points in the swallowing radiography image to obtain annotated data includes: Based on professional medical anatomical knowledge and the characteristics of swallowing angiography images, key parts that can accurately reflect the changes in anatomical structures during swallowing are selected as anatomical key points that need to be marked; Using professional image annotation tools, medical imaging professionals accurately annotate the selected key points on the swallowing angiography images to form a small amount of annotated data sets, providing a reliable supervised data foundation for subsequent model training.
3. A semi-supervised key point positioning method for swallowing imaging analysis according to claim 1, characterized in that: The joint optimization strategy is designed to integrate the supervised loss of labeled data and the self-supervised consistency loss of unlabeled data in a unified training framework; pseudo labels are generated using unlabeled data, and the pseudo labels are dynamically updated during the model training process, and the model is jointly trained using supervised data and unsupervised data, including: Construct a unified training framework that includes a labeled data supervision branch and an unlabeled data self-supervision branch. In the labeled data supervision branch, a regression loss function is used to measure the difference between the model prediction results and the true labels of the labeled data to guide the model to learn the key features in the labeled data. In the unlabeled data self-supervision branch, a consistency loss function based on data enhancement is designed to mine the potential features in the unlabeled data; During the model training process, the prediction results of unlabeled data are used to generate pseudo labels. By setting the confidence threshold, the prediction results with high confidence are selected as pseudo labels, and these pseudo labels are used together with the labeled data in the next round of model training. At the same time, according to the iterative training of the model, the confidence evaluation and selection strategy of the pseudo labels are continuously updated to dynamically improve the reliability of the pseudo labels, thereby realizing the joint effective training of supervised data and unlabeled data, so as to improve the generalization ability of the model and its ability to adapt to complex situations.
4. A semi-supervised key point positioning method for swallowing imaging analysis according to claim 3, characterized in that: For the labeled data, the heat map is used as the supervision signal of the key points. Given a labeled image x s and its corresponding annotation key point coordinate set Where K is the number of key points, and a Gaussian heat map is generated for each key point as a supervision signal: Among them, (u, v) is the image pixel coordinate, σ is the standard deviation of the Gaussian distribution, which controls the diffusion range of the heat map; The regression loss function is the mean square error loss function; the key point heat map predicted by the model is The supervised loss of a single sample is calculated using the mean square error loss function: In the formula, H k Indicates a key point.
5. The semi-supervised key point positioning method for swallowing imaging analysis according to claim 3, characterized in that: For unlabeled data, a teacher-student training paradigm is adopted to construct a teacher model f T and student model f S ; For unlabeled images x u Generate input data using weak data enhancement and strong data enhancement respectively and The data Input to the teacher model f T In the output, the corresponding heat map prediction Then the heat map prediction Perform weak to strong conversion and generate pseudo labels The data Input student model f S The output heat map is Only when the maximum response value of the pseudo label exceeds the confidence threshold τ, it is used for student model training; otherwise, the pseudo label does not participate in the training, and the expression is: in, is the indicator function, when When , the pseudo label is set to zero and does not participate in training; τ(t) is the dynamic confidence threshold.
6. A semi-supervised key point positioning method for swallowing imaging analysis according to claim 5, characterized in that: The dynamic confidence threshold τ(t) increases linearly with the training process, and the expression is: In the formula, t is the current training round, T is the total training round, τ init and τ end represents the initial and final values of the confidence threshold; The loss function for training the student model using unlabeled data is defined as:
7. The semi-supervised key point positioning method for swallowing imaging analysis according to claim 1, characterized in that: The semantic guidance module is constructed to assist the model in capturing the features of key point areas more accurately, including: The pre-trained DINOv2 model was used to extract semantic features from swallowing radiography images; A semantic guidance module is designed, and the features extracted from swallowing radiography images by the key point localization model are used to align semantic features, so that the model can focus more on the semantic information of the key point area, thereby accurately locating key anatomical structures in complex backgrounds and improving the positioning accuracy and robustness of the model.
8. The semi-supervised key point positioning method for swallowing imaging analysis according to claim 8, characterized in that: The expression of the extracted semantic features is: F sec =f DINOv2 (x) In the formula, f DINOv2 Represents the DINOv2 model; The key point positioning model is implemented by its own encoder f enc Extract the features F of swallowing radiography images model : F model =f enc (x) The semantic guidance module calculates F sec and F model The cosine similarity of is used to align the semantic features, so as to guide the model to focus more on the semantic information of the key point area; the calculation formula of cosine similarity is: In the formula, ‖·‖ represents the bi-norm of the vector.
9. The semi-supervised key point positioning method for swallowing imaging analysis according to claim 1, characterized in that: The Kalman filter algorithm is introduced in the video processing stage to achieve key point timing calibration by fusing key point prediction results of multiple frames, including: After locating the key points of the swallowing angiography video frame by frame, the key point prediction results of each frame are used as the input of the Kalman filter algorithm; The Kalman filter algorithm estimates and predicts the state of key points by establishing a state space model, and calculates a more accurate and smooth estimate of the key point position for the current frame by combining the observation value of the current frame and the state estimate of the previous frame, thereby effectively reducing the jitter and discontinuity of the positioning trajectory and improving the timing consistency and accuracy of key point detection in video sequences.
10. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method described in any one of claims 1 to 9.
Citation Information
Patent Citations
Semi-supervised nasopharyngeal carcinoma segmentation method based on image text contrast learning
CN118297960A
DETR-based semi-supervised medical image target detection method
CN118840331A
Method, apparatus, electronic device and readable storage medium for constructing key-point learning model
US20210201161A1
METHOD AND SYSTEM FOR SEMANTIC APPEARANCE TRANSFER USING SPLICING ViT FEATURES
US20240419382A1
Cited By
Single sample learning path construction method for medical image key point detection
CN121616845A
Single sample learning path construction method for medical image key point detection
CN121616845B