Swallowing radiography video micro-action identification and positioning method and electronic equipment

By extracting multimodal information fusion of anatomical key point sequences and appearance features in swallowing contrast videos, a joint optimization strategy is designed to solve the problem that the existing technology is difficult to accurately capture swallowing movement characteristics, and the accurate identification and positioning of swallowing micromoves is achieved, and more accurate clinical analysis support is provided.

CN119942635AActive Publication Date: 2025-05-06SOUTH CHINA UNIV OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411922577.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-05-06
Estimated Expiration
2044-12-25

AI Technical Summary

Technical Problem

The prior art is difficult to accurately capture the key dynamic features and action boundaries of swallowing movements in swallowing contrast videos, and faces the problems of small size of the moving target, large background/noise interference, and blurred contours of the contrast targets such as organs.

Method used

By determining and extracting the anatomical key point sequence in swallowing contrast video, obtaining the spatial position and motion trend of the throat, combining the appearance characteristics of the video (RGB and optical flow characteristics), a multimodal information complementary method is used to integrate it, and a joint optimization strategy for action classification and positioning is designed to achieve accurate identification and positioning of swallowing micromoves.

Benefits of technology

It realizes the accurate identification and positioning of swallowing micromoves, effectively deals with the problems of small size of the moving target and blurred contour of the contrast target in swallowing contrast video, and provides more accurate time-based parameter analysis to meet clinical needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942635A_ABST
    Figure CN119942635A_ABST
Patent Text Reader

Abstract

The invention discloses a swallowing angiography video micro-motion identification and positioning method and electronic equipment, and the method comprises the steps: determining anatomical key points in a swallowing angiography video, extracting a key point sequence, obtaining the spatial position and motion trend of the throat, and providing structural guidance for micro-motion identification and positioning; extracting appearance characteristics of the swallowing contrast video; the key point sequence and the appearance features are fused, and the distinguishing capability of a micro-action recognition and positioning model on fine-grained actions is improved through multi-modal information complementation; and designing a joint optimization strategy of action classification and positioning, predicting a micro-action category and a time sequence boundary thereof, and realizing accurate identification and positioning of the swallowing micro-action. According to the method, a unified analysis framework is constructed by fusing video time sequence features and anatomical key point sequence information, a key point sequence guide model is used for focusing key areas and time periods, and then accurate micro-action recognition and positioning are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of image processing, computer vision, medical video analysis, and in particular to a swallowing radiography video micro-motion recognition and positioning method and electronic equipment. Background Art

[0002] Content understanding of medical videos has a wide range of application scenarios and important application value in real life. With the development of artificial intelligence technology, micro-motion recognition and positioning have important research value and application prospects in the medical field, especially in the diagnosis and treatment of swallowing disorders. As an important tool for clinical evaluation of swallowing function, swallowing angiography video can intuitively record and analyze the patient's swallowing process. However, due to the rapidity and complexity of swallowing movements, manual annotation of micro-movements in swallowing videos is time-consuming and easily interfered by subjective factors, which is difficult to meet clinical needs. Therefore, it is an inevitable trend to develop automated micro-motion recognition and positioning methods for swallowing angiography videos.

[0003] At present, video-based action recognition technology usually adopts deep learning models to extract spatiotemporal features in videos by combining convolutional neural networks (CNN) and recurrent neural networks (RNN) or models based on self-attention mechanisms (such as Transformer) to achieve action classification and location. Such methods usually decompose videos into continuous frame sequences, extract spatial features from them to describe the key information of static images, and capture the dynamic changes of actions by modeling the temporal series relationship between frames. In addition, action localization usually uses time series regression or bounding box regression technology to accurately locate the time range of action occurrence. These methods have achieved remarkable results in standard action recognition and localization tasks, but still face specific technical difficulties in swallowing angiography videos. On the one hand, swallowing actions are accompanied by complex anatomical changes, the moving target is small in size, the background / noise interference is large, and the contours of imaging targets such as organs are blurred, making it difficult for existing technologies to accurately capture key dynamic features and action boundaries. On the other hand, the temporal boundaries of swallowing micro-movements are fuzzy, the duration is short, and there are multiple actions occurring simultaneously, which further increases the difficulty of modeling temporal information and fine-grained features. Summary of the invention

[0004] In order to solve at least one of the technical problems existing in the prior art to a certain extent, the purpose of the present invention is to provide a swallowing angiography video micro-motion recognition and positioning method based on key point sequence, electronic equipment and medium.

[0005] The first technical solution adopted by the present invention is:

[0006] A swallowing radiography video micro-motion recognition and positioning method comprises the following steps:

[0007] Determine the anatomical key points in the swallowing video and extract the key point sequence to obtain the spatial position and movement trend of the larynx, providing structured guidance for micro-motion recognition and positioning;

[0008] Extract the appearance features (RGB and optical flow features) of swallowing radiography videos;

[0009] The key point sequence is integrated with the appearance features to enhance the ability of the micro-motion recognition and localization model to distinguish fine-grained motions through complementary multi-modal information.

[0010] Design a joint optimization strategy for motion classification and positioning, predict micro-motion categories and their temporal boundaries, and achieve accurate identification and positioning of swallowing micro-motions.

[0011] Furthermore, the step of determining the anatomical key points in the swallowing radiography video and extracting the key point sequence to obtain the spatial position and movement trend of the larynx includes:

[0012] Based on professional medical anatomical atlases and clinical swallowing function assessment standards, key anatomical parts closely related to swallowing movements are selected as key points;

[0013] Through the image recognition algorithm, the pixel coordinates of the key points in each frame of the swallowing angiography video image are located, and then these coordinate information are arranged in sequence according to the time sequence of the video to obtain the key point sequence;

[0014] Analyze the coordinate changes reflected by the key point sequence, obtain the spatial position changes and movement trend characteristics of key parts during swallowing, and provide detailed and organized guidance for subsequent identification and positioning operations.

[0015] Furthermore, the extraction of appearance features of the swallowing radiography video includes:

[0016] The swallowing angiography video is input into the convolutional neural network frame by frame, and the RGB features corresponding to each frame of the image are extracted to capture the image's color, texture and other appearance information;

[0017] Use the optical flow algorithm to calculate the pixel motion information between adjacent frames, construct the optical flow field, and then use the deep neural network to extract the optical flow features to characterize the motion of objects in the video;

[0018] The extracted RGB features and optical flow features are concatenated in feature dimensions to obtain fused appearance features.

[0019] Furthermore, for swallowing radiography videos, a sliding window method was used to construct a dataset;

[0020] Assume that the number of frames corresponding to the total duration of the swallowing video is N, the length of the sliding window (the number of frames included) is L, the step size (the number of frames per slide) is S, and the number of video segments n divided by the sliding window is calculated according to the following formula:

[0021]

[0022] In the formula, Indicates a round-down operation;

[0023] Each divided video segment will be used as a data unit for subsequent RGB feature extraction, and then the RGB features and optical flow features are obtained from each frame image in each segment.

[0024] Furthermore, the key point sequence is fused with the appearance feature to enhance the ability of the micro-motion recognition and positioning model to distinguish fine-grained motions through complementary multimodal information, including:

[0025] The attention mechanism is used to assign different weights to different parts of the appearance features based on the key area guidance information provided by the key point sequence, so that the micro-movement recognition and positioning model can focus more on the features related to the key areas of the swallowing movement during the analysis process, thus achieving deep fusion of multimodal information.

[0026] By constructing a training sample set containing fused feature vectors and using supervised learning to train the micro-motion recognition and localization model, the micro-motion recognition and localization model can fully learn the advantages brought by the complementarity of multimodal information during the training process, thereby improving the ability to distinguish the detailed features of micro-motions.

[0027] Furthermore, the attention mechanism is used to assign different weights to different parts of the appearance features based on the key area guidance information provided by the key point sequence, including:

[0028] For key point sequences, graph convolutional neural network (GCN) or convolutional neural network (CNN) is used to extract key point sequence features.

[0029] Using cross attention mechanism to fuse key point sequence features With appearance features F i , and obtain enhanced key point sequence features and appearance features

[0030] Furthermore, the joint optimization strategy of designing motion classification and positioning, predicting micro-motion categories and their temporal boundaries, and realizing accurate recognition and positioning of swallowing micro-motions includes:

[0031] Construct a joint loss function that includes a classification loss function and a localization loss function. The classification loss function is used to measure the difference between the predicted micro-motion category and the true category, and the localization loss function is used to measure the error between the predicted action timing boundary and the actual timing boundary.

[0032] The constructed joint loss function is optimized using an optimization algorithm. During the model training process, the model parameters are continuously adjusted according to the feedback of the loss value, so that the model gradually converges, improving the accuracy of micro-motion category prediction and the accuracy of temporal boundary positioning;

[0033] During the testing phase, the swallowing angiography video to be analyzed is input into the trained model, and the category of micro-movements and the corresponding timing boundary information are directly obtained through the model output, thereby achieving accurate recognition and positioning of swallowing micro-movements.

[0034] Furthermore, the joint loss function is constructed in the following way:

[0035] Enhanced appearance features and key point sequence features They are respectively sent into the action recognition and positioning backbone networks with the same structure, and the expressions are:

[0036]

[0037] Where Y i , is the final output of the model, including action classification results and positioning results, Y i , They correspond to the outputs obtained by the enhanced appearance features and the enhanced key point sequence features respectively; φ TAL Represents the ActionMamba model;

[0038] Focal Loss is used as the classification loss function L cls , the expression is:

[0039]

[0040] In the formula, p i is the predicted micro-action category probability distribution, y i is the true category, γ is the focus parameter, α is the balance factor, and C is the total number of micro-motion categories:

[0041] The distance intersection loss function is used to measure the error between the predicted action timing boundary and the actual timing boundary. The expression of the positioning loss function is:

[0042]

[0043] Where DIoU is the distance intersection over union ratio; s p is the predicted action start time, e p is the end time, s g is the actual action start time, e g is the end time, ρ is the distance from the center point, and c is the duration;

[0044] The joint loss function is:

[0045] L=L cls +λL loc

[0046] Where λ is the weight coefficient for balancing classification loss and localization loss.

[0047] Furthermore, the training steps of the micro-motion recognition and positioning model include:

[0048] In each round of training, the model parameters are updated using the stochastic gradient descent algorithm according to the loss function value; at the same time, the exponential moving average and gradient clipping techniques are used to stabilize the training process and prevent gradient explosion. The expression is:

[0049]

[0050] θ ema,t =γθ ema,t-1 +(1-γ)θ t

[0051] In the formula, θ t is the model parameter at the tth iteration, is the loss function L with respect to parameter θ t The gradient of G is the gradient threshold, θ ema,t is the exponential moving average parameter at the tth iteration, and γ is the decay rate.

[0052] The second technical solution adopted by the present invention is:

[0053] An electronic device comprises a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, at least one program, the code set or the instruction set is loaded and executed by the processor to implement a swallowing angiography video micro-motion recognition and positioning method as described above.

[0054] The third technical solution adopted by the present invention is:

[0055] A computer-readable storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement a swallowing angiography video micro-motion recognition and positioning method as described above.

[0056] The fourth technical solution adopted by the present invention is:

[0057] A computer program product or a computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above method.

[0058] The beneficial effects of the present invention are as follows: the present invention constructs a unified analysis framework by integrating video timing features and anatomical key point sequence information, and uses the key point sequence to guide the model to focus on key areas and time periods, thereby achieving accurate micro-motion recognition and positioning. In addition, the present invention not only focuses on appearance information, but also fully explores anatomical key points and kinematic information during swallowing, effectively addressing the problems of small size of moving targets and blurred contours of imaging targets such as organs in swallowing angiography videos. The present invention will provide more accurate temporal parameter analysis for clinical scenarios such as dysphagia diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the embodiments of the present invention or the drawings of related technical solutions in the prior art are introduced below. It should be understood that the drawings introduced below are only for the convenience of clearly describing some embodiments of the technical solutions of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0060] Figure 1 This is a flowchart of the steps of a swallowing angiography video micro-motion recognition and positioning method in an embodiment of the present invention;

[0061] Figure 2 Schematic diagram of anatomical key points in an embodiment of the present invention. DETAILED DESCRIPTION

[0062] The embodiments of the present invention are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and are not to be construed as limitations of the present invention. For the step numbers in the following embodiments, they are only provided for the convenience of explanation, and the order between the steps is not limited in any way, and the execution order of each step in the embodiment can be adaptively adjusted according to the understanding of those skilled in the art.

[0063] In the description of the present invention, it should be understood that descriptions involving orientations, such as up, down, front, back, left, right, etc., and orientations or positional relationships indicated are based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present invention.

[0064] In the description of the present invention, "several" means one or more, "more" means more than two, "greater than", "less than", "exceed" etc. are understood as not including the number itself, and "above", "below", "within" etc. are understood as including the number itself. If there is a description of "first" or "second", it is only used for the purpose of distinguishing the technical features, and cannot be understood as indicating or implying the relative importance or implicitly indicating the number of the indicated technical features or implicitly indicating the order of the indicated technical features.

[0065] In the description of the present invention, unless otherwise clearly defined, terms such as setting, installing, connecting, etc. should be understood in a broad sense, and technicians in the relevant technical field can reasonably determine the specific meanings of the above terms in the present invention based on the specific content of the technical solution.

[0066] At present, swallowing radiography video analysis faces specific technical difficulties. On the one hand, swallowing movements are accompanied by complex anatomical changes, the moving target is small in size, the background / noise interference is large, and the contours of imaging targets such as organs are blurred, making it difficult for existing technologies to accurately capture key dynamic features and action boundaries. On the other hand, the temporal boundaries of swallowing micro-movements are fuzzy, the duration is short, and there are multiple actions occurring simultaneously, further increasing the difficulty of modeling temporal information and fine-grained features. Existing methods focus on utilizing the RGB modality of the video and only focus on appearance information. They fail to fully explore the anatomical key points and their kinematic information that are crucial in the swallowing process, and are difficult to meet the strict requirements of clinical practice for accurate identification and positioning of micro-movements in swallowing radiography videos.

[0067] Based on this, the present invention proposes a swallowing angiography video micro-motion recognition and positioning scheme based on key point sequence. By integrating the video temporal features and anatomical key point sequence information to build a unified analysis framework, the key point sequence is used to guide the model to focus on key areas and time periods, thereby achieving accurate micro-motion recognition and positioning. Specifically, firstly, the anatomical key points in the swallowing angiography video are determined based on clinical knowledge and the key point sequence is extracted to obtain the spatial position and movement trend of the larynx, providing structured guidance for micro-motion recognition and positioning; secondly, the appearance features (RGB and optical flow features) of the swallowing angiography video are extracted using a deep learning model; then, the key point sequence is fused with the appearance features, and the model's ability to distinguish fine-grained movements is enhanced through multimodal information complementarity; finally, a joint optimization strategy for movement classification and positioning is designed to predict the micro-motion category and its temporal boundary, so as to achieve accurate recognition and positioning of swallowing micro-motions.

[0068] Example 1

[0069] like Figure 1 As shown, this embodiment provides a method for identifying and locating micro-movements in swallowing angiography videos, so as to achieve accurate identification and location of micro-movements in swallowing angiography videos, overcome the deficiencies of the prior art in processing such videos, and improve the diagnosis and analysis effects in clinical applications. The method includes the following steps:

[0070] S1. Based on clinical knowledge, determine the anatomical key points in the swallowing video and extract the key point sequence to obtain the spatial position and movement trend of the larynx, providing structured guidance for micro-motion recognition and positioning.

[0071] In some embodiments, step S1 specifically includes the following steps:

[0072] S11. Based on professional medical anatomical atlases and relevant standards for clinical swallowing function assessment, key anatomical parts closely related to swallowing movements are selected as key points. These key points include the throat, esophageal entrance and other parts.

[0073] S12. Using an image recognition algorithm, the pixel coordinates of the key points are located in each frame of the swallowing angiography video image, and then the coordinate information is arranged in sequence according to the time series of the video to extract a key point sequence.

[0074] S13. Analyze the coordinate changes reflected by the key point sequence to obtain the spatial position changes and movement trend characteristics of key parts such as the larynx during the swallowing process, so as to provide detailed and organized guidance for subsequent identification and positioning operations.

[0075] S2. Use deep learning models to extract the appearance features (RGB and optical flow features) of swallowing radiography videos.

[0076] In some embodiments, step S2 specifically includes the following steps:

[0077] S21. Select a suitable convolutional neural network (CNN) architecture, input the swallowing angiography video into the network frame by frame, and extract the RGB feature map corresponding to each frame image to capture the image's color, texture and other appearance information.

[0078] S22. Use the optical flow algorithm to calculate the pixel motion information between adjacent frames, construct an optical flow field, and then extract the optical flow features based on another deep neural network to characterize the motion of objects in the video.

[0079] S23, concatenating the extracted RGB features and optical flow features in feature dimensions to obtain fused appearance features.

[0080] S3. Fusion of key point sequences and appearance features to enhance the ability of micro-motion recognition and localization models to discern fine-grained motions through complementary multimodal information.

[0081] In some embodiments, step S3 specifically includes the following steps:

[0082] S31. The attention mechanism is used to assign different weights to different parts of the appearance features based on the key area guidance information provided by the key point sequence, so that the model can focus more on the features related to the key areas of the swallowing action during the analysis process, thereby achieving deep fusion of multimodal information.

[0083] S32. By constructing a training sample set containing fused feature vectors, the classification and positioning model is trained using supervised learning, so that it can fully learn the advantages brought by the complementarity of multimodal information during the training process, thereby improving the ability to distinguish the detailed features of micro-motions.

[0084] S4. Design a joint optimization strategy for motion classification and positioning to predict micro-motion categories and their temporal boundaries, and achieve accurate recognition and positioning of swallowing micro-motions.

[0085] In some embodiments, step S4 specifically includes the following steps:

[0086] S41. Constructing a joint loss function including a classification loss function and a positioning loss function, wherein the classification loss function is used to measure the difference between the predicted micro-motion category and the true category, and the positioning loss function is used to measure the error between the predicted action timing boundary and the actual timing boundary;

[0087] S42, using an optimization algorithm (such as stochastic gradient descent, etc.) to optimize the constructed joint loss function. During the model training process, the parameters of the model are continuously adjusted according to the feedback of the loss value, so that the model gradually converges, thereby improving the accuracy of micro-motion category prediction and the accuracy of temporal boundary positioning;

[0088] S43. During the testing phase, the swallowing angiography video to be analyzed is input into the trained model, and the category of micro-movements and the corresponding temporal boundary information are directly obtained through the model output, so as to achieve accurate recognition and positioning of swallowing micro-movements.

[0089] The above method is explained in detail below with reference to the accompanying drawings and specific implementation methods.

[0090] This embodiment provides a method for identifying and locating micro-movements in swallowing angiography videos based on key point sequences, which specifically includes the following steps:

[0091] Step 1: Determine and extract anatomical key point sequences in swallowing angiography videos.

[0092] In this step, based on medical expertise and clinical experience, the present invention determines anatomical parts closely related to the swallowing process as key points. These parts usually include the soft palate, hyoid bone, etc. in the swallowing video. Specific key point categories are: suprahyoid convex point, subhyoid convex point, left hyoid convex point, left end point of soft palate, right end point of soft palate, soft palate peak point, left lower vertex of the second vertebra, left lower vertex of the fourth vertebra, such as Figure 2 shown.

[0093] For each frame of the swallowing angiography video, the key point positioning model is used to locate the position of the key points. Specifically, the present invention constructs a key point dataset of the swallowing angiography dataset based on the labeled data, which is used to train the existing key point positioning model HRNet. Then, the model is used to extract the key points of each frame of the swallowing angiography video. Let the i-th key point in the j-th frame be the key point p ij =(x ij ,y ij ), where x and y represent the horizontal and vertical coordinates in the image respectively. All key points of the jth frame can be expressed as:

[0094] k j =[P 1j ,P 2j ,…,P 8j ]

[0095] The key point sequence of N frames of swallowing angiography video can be expressed as:

[0096] k=[k 1 ,k 2 ,…,k N ]

[0097] Step 2: Extract the appearance features (RGB and optical flow features) of the swallowing radiography video.

[0098] First, for the available swallowing videos, the sliding window method is used to construct a data set. Assume that the number of frames corresponding to the total duration of the swallowing video is N, the length of the sliding window (the number of frames included) is L, and the step size (the number of frames per slide) is S. The number of video segments n divided by the sliding window can be calculated according to the following formula:

[0099]

[0100] in Each segmented video segment will be used as a data unit for subsequent RGB feature extraction, and then its RGB features and optical flow features can be obtained from each frame image in each segment. Optionally, an I3D model is used as a feature extractor to extract video features. Among them, the pre-trained RGB I3D model φ RGB and the pre-trained optical flow I3D model φ Flow When processing video segment C i In the process, the RGB feature g can be extracted accordingly i and optical flow features fi i , the specific formula is as follows:

[0101] g i =φ RGB (C i ),f i =φ Flow (C i )

[0102] In order to better integrate these two features, the present invention adopts the common feature splicing operation to fuse the RGB and optical flow features to obtain the fused appearance feature F i , the specific formula is as follows:

[0103] F i =[C i ;f i ]

[0104] Step 3: Design a multimodal information fusion mechanism for fusing key point sequences and appearance features.

[0105] This embodiment uses a graph convolutional neural network (GCN) or a convolutional neural network (CNN) to extract key point sequence features.

[0106] 1) If GCN is used, the key point sequence must first be constructed as a graph structure data, with each anatomical key point as a node of the graph, and edges are constructed based on the anatomical association and the mutual influence relationship in the swallowing action. Optionally, the ST-GCN model is used to extract key point sequence features. Before extracting features, it is also necessary to perform the sliding window operation mentioned in step 2 on the key point sequence. The sliding window is divided on the key point sequence data according to the set window length and step size, and the original continuous key point sequence is divided into multiple subsequence fragments. Then the key point subsequence fragment X i Through the ST-GCN model, features are extracted The specific calculation formula is as follows:

[0107]

[0108] 2) If CNN is used, the key point sequence needs to be organized into a format suitable for CNN input, such as converting the key point coordinates into a heat map form, and then stacking them in chronological order into a video-like format. Then, the features are extracted through the 3D-CNN network. Optionally, the I3D model φ HM To extract X i The corresponding heat map sequence feature H i , the specific calculation formula is as follows:

[0109]

[0110] After obtaining the key point sequence features Then, the cross attention mechanism is used to fuse it with the previously obtained appearance feature F i Specifically, the key point sequence features can be used as the query (Query, denoted as Q), the appearance features as the key (Key, denoted as K) and the value (Value, denoted as V), or the reverse setting, to construct a bidirectional cross-attention path to mine bimodal correlation information. By using the Transformer module based on cross-attention, the enhanced key point sequence features can be expressed as The specific calculation formula is as follows:

[0111]

[0112] Among them, F Q ,W K ,W V are all learnable linear projection layer parameters. Similarly, enhanced appearance features can be obtained The specific formula is as follows:

[0113]

[0114] Step 4: Design a joint optimization strategy for micro-motion classification and localization.

[0115] Enhanced appearance features and key point sequence features They are fed into the action recognition and positioning backbone networks of the same structure respectively. Optionally, the ActionMamba model φ is used. TAL To obtain the prediction results, the specific calculation formula is as follows:

[0116]

[0117] Among them, Y i , is the final output of the model, including action classification results and positioning results, Y i , They correspond to the outputs obtained by the enhanced appearance features and the enhanced key point sequence features respectively.

[0118] In order to alleviate the problem of class imbalance and pay more attention to difficult-to-classify samples, Focal Loss is used as the classification loss function. Suppose the predicted micro-action category probability distribution is p i (i represents the category), the true category is y i (using one-hot encoding), γ is the focus parameter (usually 2), α is the balance factor, then the classification loss is as follows, where C is the total number of micro-motion categories:

[0119]

[0120] The distance intersection over union (DIoU) loss function is used to measure the error between the predicted action timing boundary and the actual timing boundary. Assume that the predicted action start time is s p , the end time is e p , the actual action start time is s g , the end time is e g , center point distance ρ, duration c, then:

[0121]

[0122] The localization loss is expressed as:

[0123] L loc =1-DIoU

[0124] The joint loss function used in the final training is L = L cls +λL loc , where λ is the weight coefficient for balancing classification loss and positioning loss, which can be determined through experimental adjustment.

[0125] Step 5: Train to obtain the micro-motion recognition and positioning model.

[0126] In each round of training, the model parameters are updated using the stochastic gradient descent algorithm according to the loss function value in step 4. At the same time, the exponential moving average (EMA) and gradient clipping techniques are used to stabilize the training process and prevent gradient explosion. The specific formula is as follows:

[0127]

[0128] θ ema,t =γθ ema,t-1 +(1-γ)θ t

[0129] where θ t is the model parameter at the tth iteration, is the loss function L with respect to parameter θ t The gradient of G is the gradient threshold, θ ema,t is the exponential moving average parameter at the tth iteration, and γ is the decay rate. In the test phase, the swallowing angiography video to be analyzed is input into the trained model, and the category of micro-movements and the corresponding temporal boundary information are directly obtained through the model output, so as to realize the accurate recognition and positioning of swallowing micro-movements.

[0130] In summary, the present invention proposes a method for identifying and locating micro-movements in swallowing angiography videos based on key point sequences. By integrating video timing features with anatomical key point sequence information, a unified analysis framework is constructed, and the key point sequence is used to guide the model to focus on key areas and time periods, thereby achieving accurate micro-movement identification and positioning. Compared with existing methods, this invention not only focuses on appearance information, but also fully explores anatomical key points and kinematic information in the swallowing process, effectively addressing the problems of small size of moving targets and blurred contours of imaging targets such as organs in swallowing angiography videos. The present invention will provide more accurate temporal parameter analysis for clinical scenarios such as swallowing disorder diagnosis.

[0131] Example 2

[0132] An embodiment of the present invention further provides an electronic device, the electronic device comprising a processor and a memory, the memory storing at least one instruction, at least one program, a code set or an instruction set, the at least one instruction, the at least one program, the code set or the instruction set being loaded and executed by the processor to implement the following Figure 1 A method for identifying and locating micro-movements in swallowing angiography videos is shown.

[0133] It is understood that the memory may include a random access memory (RAM) or a read-only memory (ROM). Optionally, the memory includes a non-transitory computer-readable storage medium. The memory may be used to store instructions, programs, codes, code sets, or instruction sets. The memory may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function, instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data created according to the use of the server, etc.

[0134] The processor may include one or more processing cores. The processor uses various interfaces and lines to connect the various parts of the entire server, and executes various functions of the server and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory, and calling data stored in the memory. Optionally, the processor can be implemented in at least one hardware form of digital signal processing (DSP), field programmable gate array (FPGA), and programmable logic array (PLA). The processor can integrate one or a combination of a central processing unit (CPU) and a modem. Among them, the CPU mainly processes the operating system and application programs; the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor, but implemented separately through a chip.

[0135] Since the electronic device is an electronic device corresponding to a swallowing angiography video micro-motion recognition and positioning method of an embodiment of the present invention, and the principle of solving the problem by the electronic device is similar to that of the method, the implementation of the electronic device can refer to the implementation process of the above-mentioned method embodiment, and the repeated parts will not be repeated.

[0136] Example 3

[0137] The embodiment of the present invention further provides a computer-readable storage medium, wherein the storage medium stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, the at least one program, the code set or instruction set is loaded and executed by a processor to implement the following Figure 1 A method for identifying and locating micro-movements in swallowing angiography videos is shown.

[0138] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, the storage medium including a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable programmable read-only memory (EEPROM), a compact disc (CD-ROM) or other optical disc storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium that can be used to carry or store data.

[0139] Since the storage medium is a storage medium corresponding to a swallowing angiography video micro-motion recognition and positioning method in an embodiment of the present invention, and the principle of solving the problem by the storage medium is similar to that of the method, the implementation of the storage medium can refer to the implementation process of the above-mentioned method embodiment, and the repeated parts will not be repeated.

[0140] Example 4

[0141] In some possible implementations, various aspects of the method of the embodiment of the present invention can also be implemented in the form of a program product, which includes a program code. When the program product is run on a computer device, the program code is used to enable the computer device to execute the steps of a swallowing angiography video micro-motion recognition and positioning method according to various exemplary embodiments of the present application described above in this specification. Among them, the executable computer program code or "code" for executing each embodiment can be written in a high-level programming language such as Python, C, C++, C#, Smalltalk, Java, JavaScript, Visual Basic, structured query language (e.g., Transact-SQL), Perl, or in various other programming languages.

[0142] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0143] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.

[0144] The above embodiments are only for illustrating the technical concept and features of the present invention, and their purpose is to enable ordinary technicians in the field to understand the content of the present invention and implement it accordingly, and they cannot be used to limit the protection scope of the present invention. Any equivalent changes or modifications made based on the essence of the content of the present invention should be included in the protection scope of the present invention.

Claims

1. A swallowing angiography video micro-motion recognition and positioning method, characterized in that: The following steps are involved: Determine the anatomical key points in the swallowing video and extract the key point sequence to obtain the spatial position and movement trend of the larynx, providing structured guidance for micro-motion recognition and positioning; Extract appearance features of swallowing radiography videos; The key point sequence is integrated with the appearance features to enhance the ability of the micro-motion recognition and localization model to distinguish fine-grained motions through complementary multi-modal information. Design a joint optimization strategy for motion classification and positioning, predict micro-motion categories and their temporal boundaries, and achieve accurate identification and positioning of swallowing micro-motions.

2. A swallowing imaging video micro-motion recognition and positioning method according to claim 1, characterized in that: The method of determining the anatomical key points in the swallowing radiography video and extracting the key point sequence to obtain the spatial position and movement trend of the larynx includes: Based on professional medical anatomical atlases and clinical swallowing function assessment standards, key anatomical parts closely related to swallowing movements are selected as key points; Through the image recognition algorithm, the pixel coordinates of the key points in each frame of the swallowing angiography video image are located, and then these coordinate information are arranged in sequence according to the time sequence of the video to obtain the key point sequence; Analyze the coordinate changes reflected by the key point sequence, obtain the spatial position changes and movement trend characteristics of key parts during swallowing, and provide detailed and organized guidance for subsequent identification and positioning operations.

3. The swallowing imaging video micro-motion recognition and positioning method according to claim 1, characterized in that: The method of extracting the appearance features of the swallowing radiography video includes: The swallowing angiography video is input into the convolutional neural network frame by frame to extract the RGB features corresponding to each frame image; Use the optical flow algorithm to calculate the pixel motion information between adjacent frames, construct the optical flow field, and then use the deep neural network to extract the optical flow features to characterize the motion of objects in the video; The extracted RGB features and optical flow features are concatenated in feature dimensions to obtain fused appearance features.

4. The swallowing imaging video micro-motion recognition and positioning method according to claim 3, characterized in that: For swallowing angiography videos, a sliding window method is used to construct a data set; Assume that the number of frames corresponding to the total duration of the swallowing video is N, the length of the sliding window is L, and the step size is S. The number of video segments n divided by the sliding window is calculated according to the following formula: In the formula, Indicates a round-down operation; Each divided video segment will be used as a data unit for subsequent RGB feature extraction, and then the RGB features and optical flow features are obtained from each frame image in each segment.

5. The swallowing imaging video micro-motion recognition and positioning method according to claim 1, characterized in that: The key point sequence is fused with the appearance feature to enhance the ability of the micro-motion recognition and positioning model to distinguish fine-grained motions through complementary multi-modal information, including: The attention mechanism is used to assign different weights to different parts of the appearance features based on the key area guidance information provided by the key point sequence, so that the micro-movement recognition and positioning model can focus more on the features related to the key areas of the swallowing movement during the analysis process, thus achieving deep fusion of multimodal information. By constructing a training sample set containing fused feature vectors and using supervised learning to train the micro-motion recognition and localization model, the micro-motion recognition and localization model can fully learn the advantages brought by the complementarity of multimodal information during the training process, thereby improving the ability to distinguish the detailed features of micro-motions.

6. A swallowing imaging video micro-motion recognition and positioning method according to claim 5, characterized in that: The attention mechanism is used to assign different weights to different parts of the appearance features based on the key area guidance information provided by the key point sequence, including: For key point sequences, graph convolutional neural networks or convolutional neural networks are used to extract key point sequence features. Using cross attention mechanism to fuse key point sequence features With appearance features F i , and obtain enhanced key point sequence features and appearance features 7. The swallowing imaging video micro-motion recognition and positioning method according to claim 1, characterized in that: The joint optimization strategy of designing motion classification and positioning, predicting micro-motion categories and their temporal boundaries, and realizing accurate recognition and positioning of swallowing micro-motions includes: Construct a joint loss function including classification loss function and positioning loss function; The constructed joint loss function is optimized using an optimization algorithm. During the model training process, the model parameters are continuously adjusted according to the feedback of the loss value, so that the model gradually converges, improving the accuracy of micro-motion category prediction and the accuracy of temporal boundary positioning; During the testing phase, the swallowing angiography video to be analyzed is input into the trained model, and the category of micro-movements and the corresponding timing boundary information are directly obtained through the model output, thereby achieving accurate recognition and positioning of swallowing micro-movements.

8. The swallowing imaging video micro-motion recognition and positioning method according to claim 7, characterized in that: The joint loss function is constructed in the following way: Enhanced appearance features and key point sequence features They are respectively sent into the action recognition and positioning backbone networks with the same structure, and the expressions are: Where Y i , is the final output of the model, including action classification results and positioning results, Y i , They correspond to the outputs obtained by the enhanced appearance features and the enhanced key point sequence features respectively; φ TAL Represents the ActionMamba model; Focal Loss is used as the classification loss function L cls , the expression is: In the formula, p i is the predicted micro-action category probability distribution, y i is the true category, γ is the focus parameter, α is the balance factor, and C is the total number of micro-motion categories: The distance intersection loss function is used to measure the error between the predicted action timing boundary and the actual timing boundary. The expression of the positioning loss function is: L loc =1-DIoU Where DIoU is the distance intersection over union ratio; s p is the predicted action start time, e p is the end time, s g is the actual action start time, e g is the end time, ρ is the distance from the center point, and c is the duration; The joint loss function is: L=L cls +λL loc Where λ is the weight coefficient for balancing classification loss and localization loss.

9. The swallowing imaging video micro-motion recognition and positioning method according to claim 7, characterized in that: The training steps of the micro-motion recognition and positioning model include: In each round of training, the model parameters are updated using the stochastic gradient descent algorithm according to the loss function value; at the same time, the exponential moving average and gradient clipping techniques are used to stabilize the training process and prevent gradient explosion. The expression is: i ema,t =γθ ema,t-1 +(1-γ)θ t In the formula, θ t is the model parameter at the tth iteration, is the loss function L with respect to parameter θ t The gradient of G is the gradient threshold, θ ema,t is the exponential moving average parameter at the tth iteration, and γ is the decay rate.

10. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Self-adaptive perception video time sequence action positioning system and method thereof

    CN116052034A

  • Micro-action chronological parameter acquisition method and device and medium

    CN116863367A

  • Identification and de-identification within a video sequence

    US20170249500A1