A pharyngeal phase swallowing reconstruction method and system based on multi-dimensional image fusion
By using multimodal image acquisition and fusion technology, a dynamic model of the pharynx is constructed, which solves the problem that single-view images cannot fully capture the three-dimensional dynamic structure of the pharynx, and achieves high-fidelity restoration and accurate simulation of the swallowing process, thereby improving the reliability and accuracy of the swallowing process.
Patent Information
- Application Number
- CN202511860710.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-11
- Publication Date
- 2026-03-17
- Estimated Expiration
- 2045-12-11
AI Technical Summary
In existing technologies, two-dimensional images from a single perspective cannot fully capture the complex three-dimensional dynamic structure of the pharynx, resulting in information loss. The manual analysis process is highly dependent on physician experience, is subjective and difficult to quantify, and lacks accurate reconstruction and dynamic simulation of the entire swallowing process, leading to insufficient accuracy and reliability in the reconstruction of the swallowing process.
Multimodal image acquisition equipment is used to simultaneously acquire swallowing dynamic data from multiple perspectives, extract image feature parameters, configure image fusion weights to perform multidimensional image fusion, construct a pharyngeal dynamic model, use optimization algorithms to iteratively correct the motion trajectory, construct a swallowing restoration simulator, and output pharyngeal swallowing restoration results.
It achieves high-fidelity reconstruction of the swallowing process, improving the accuracy and reliability of the reconstruction. Through multi-dimensional image fusion and trajectory optimization correction, it improves the accuracy of motion trajectory and outputs highly realistic dynamic results of the swallowing process.
Smart Images

Figure CN121304945B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image data processing technology, specifically to a method and system for restoring pharyngeal swallowing based on multidimensional image fusion. Background Technology
[0002] With the acceleration of population aging, the demand for diagnosis and treatment of dysphagia continues to grow, and the application of multimodal imaging technology in medical scenarios is becoming increasingly widespread. Existing technologies mostly acquire swallowing process data through single or limited-view images, combined with basic image processing methods to assist in observing the dynamics of swallowing during the pharyngeal phase.
[0003] However, existing technologies have obvious limitations: two-dimensional images from a single perspective cannot fully capture the complex three-dimensional dynamic structure of the pharynx, resulting in information loss; the manual analysis process is highly dependent on physician experience, is subjective, and difficult to quantify; in addition, there is a lack of effective means to accurately reconstruct and dynamically simulate the entire swallowing process, ultimately leading to insufficient accuracy and reliability in reconstructing the swallowing process. Summary of the Invention
[0004] This invention provides a method and system for pharyngeal swallowing reconstruction based on multidimensional image fusion, aiming to solve the technical problem of insufficient accuracy and reliability in the reconstruction of the swallowing process in the prior art.
[0005] In view of the above problems, the present invention provides a method and system for pharyngeal swallowing restoration based on multidimensional image fusion.
[0006] In a first aspect, the present invention provides a method for restoring pharyngeal swallowing based on multidimensional image fusion, comprising:
[0007] Multimodal imaging acquisition equipment is used to acquire multidimensional image sequences of the pharyngeal swallowing process, obtain dynamic swallowing data from multiple imaging perspectives, and extract multiple image feature parameters.
[0008] Based on the multiple image feature parameters, image fusion weights are configured to perform multidimensional image fusion and construct a dynamic model of the pharynx.
[0009] Based on the pharyngeal dynamic model, pharyngeal motion trajectory analysis is performed to obtain the initial motion trajectory. The comprehensive trajectory similarity between the initial motion trajectory and the reference motion trajectory is calculated. An optimization objective function is constructed based on the comprehensive trajectory similarity. An optimization algorithm is used to iteratively optimize the initial motion trajectory to obtain the corrected motion trajectory.
[0010] Based on the corrected motion trajectory, a swallowing restoration simulator is constructed to simulate swallowing restoration and output the swallowing restoration results during the pharyngeal phase.
[0011] Secondly, the present invention provides a pharyngeal swallowing reconstruction system based on multidimensional image fusion, comprising:
[0012] The multi-dimensional image acquisition module is used to acquire multi-dimensional image sequences of the pharyngeal swallowing process through multi-modal image acquisition equipment, obtain swallowing dynamic data from multiple image perspectives, and extract multiple image feature parameters.
[0013] The image fusion modeling module is used to configure image fusion weights based on the multiple image feature parameters, perform multi-dimensional image fusion, and construct a dynamic model of the pharynx.
[0014] The trajectory optimization and correction module is used to perform pharyngeal motion trajectory analysis based on the pharyngeal dynamic model, obtain an initial motion trajectory, calculate the comprehensive trajectory similarity between the initial motion trajectory and the reference motion trajectory, construct an optimization objective function based on the comprehensive trajectory similarity, and use an optimization algorithm to iteratively optimize the initial motion trajectory to obtain a corrected motion trajectory.
[0015] The swallowing restoration simulation module is used to construct a swallowing restoration simulator based on the corrected motion trajectory, perform swallowing restoration simulation, and output the swallowing restoration results during the pharyngeal phase.
[0016] One or more technical solutions provided in this invention have at least the following technical effects or advantages:
[0017] This invention provides a method and system for pharyngeal swallowing reconstruction based on multidimensional image fusion. It comprehensively acquires dynamic swallowing data through multidimensional image acquisition and integrates multi-source information through dynamic fusion modeling to construct a high-fidelity dynamic pharyngeal model. Then, through trajectory optimization and correction, the initial motion trajectory is compared with the standard trajectory and iteratively optimized to improve the accuracy of the motion trajectory. Finally, the swallowing reconstruction simulation outputs highly realistic dynamic results of the swallowing process, effectively improving the accuracy and reliability of swallowing process reconstruction. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 A flowchart illustrating a pharyngeal swallowing reconstruction method based on multidimensional image fusion provided in an embodiment of the present invention;
[0020] Figure 2 A schematic diagram of a pharyngeal swallowing reconstruction system based on multidimensional image fusion provided in an embodiment of the present invention;
[0021] The components represented by each number in the attached diagram are explained below:
[0022] Multidimensional image acquisition module 11, image fusion modeling module 12, trajectory optimization and correction module 13, swallowing restoration simulation module 14. Detailed Implementation
[0023] This invention provides a method and system for pharyngeal swallowing reconstruction based on multidimensional image fusion, which addresses the technical problem of insufficient accuracy and reliability of existing technologies in reconstructing the swallowing process.
[0024] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0025] It should be noted that the terms "comprising" and "having" are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or modules that are not explicitly listed or that are inherent to these processes, methods, products, or devices.
[0026] Example 1, as Figure 1 As shown, this invention provides a method for pharyngeal swallowing reconstruction based on multidimensional image fusion, the method comprising:
[0027] S100: Through multimodal image acquisition equipment, it acquires multidimensional image sequences of the pharyngeal swallowing process, obtains swallowing dynamic data from multiple image perspectives, and extracts multiple image feature parameters.
[0028] In this embodiment of the invention, a multimodal imaging acquisition device is used to acquire multidimensional image sequences of the pharyngeal swallowing process, obtain swallowing dynamic data from multiple imaging perspectives, and extract multiple image feature parameters. The swallowing process involves the coordinated movement of complex pharyngeal muscle groups. Traditional single-view imaging can only capture local dynamics and cannot fully reflect the spatial displacement and movement patterns of the pharyngeal structure. Furthermore, single-image modalities lack the ability to quantify three-dimensional motion trajectories, leading to subjective and inaccurate assessments of swallowing disorders. Therefore, multimodal imaging acquisition and feature extraction techniques are needed to achieve high-resolution, multidimensional data acquisition of pharyngeal swallowing dynamics, providing a reliable foundation for subsequent modeling and analysis.
[0029] Step S100 in the method provided in this embodiment of the invention includes:
[0030] The dynamic image sequence of the pharyngeal swallowing process is simultaneously acquired by multiple high-speed cameras arranged around the pharynx, wherein the multiple high-speed cameras include at least one frontal view camera, one side view camera and one oblique view camera.
[0031] Image feature parameters for each time point are extracted from the dynamic image sequence, wherein the image feature parameters include the coordinates of the pharyngeal contour points, the motion velocity value, and the motion direction angle.
[0032] First, multiple high-speed cameras deployed around the pharynx synchronously acquire dynamic image sequences of the pharyngeal swallowing process. These high-speed cameras include at least one frontal view camera, one side view camera, and one oblique view camera. High-speed cameras refer to imaging devices with a frame rate of at least 200 frames per second, capable of accurately capturing the high-speed movements of structures during pharyngeal swallowing, avoiding the motion blur problem of ordinary cameras. The dynamic image sequence refers to a series of ordered image frames generated by continuous shooting from the high-speed cameras, each frame corresponding to the pharyngeal state at a specific time point. Combining these frames reconstructs the dynamic changes during swallowing. Three sets of high-speed cameras are deployed around the patient's pharynx, each corresponding to the observation needs of different anatomical regions: a frontal view camera facing the patient's oropharynx to capture the rising and falling movements of the epiglottis and the retraction trajectory of the tongue root; a side view camera shooting from the patient's ear side, focusing on recording the degree of pharyngeal expansion and the path of the food bolus; and an oblique view camera covering the posterior pharyngeal wall and laryngeal vestibule region at a 45° angle to supplement the dynamic contraction of the pharyngeal constrictor muscles. By calibrating timestamps and matching frame rates, it is ensured that the three sets of cameras record the image sequence on the same timeline when the swallowing action occurs, thus avoiding spatial misalignment caused by time shift.
[0033] For example, using a 65-year-old patient with dysphagia as the subject, a frontal camera was fixed 50cm directly in front of the patient, with the lens aimed at the midline of the pharynx; a side camera was fixed 30cm to the left of the patient, with the lens horizontally aimed at the side of the pharynx; and an oblique camera was fixed 60cm away at a 45° angle to the right front of the patient, with the lens obliquely aimed at the pharynx. After activating the synchronization controller, the patient was asked to swallow 10ml of warm water, and the three cameras simultaneously captured images at a frame rate of 300 frames per second, obtaining a total of 600 frames of images. This formed a dynamic image sequence containing three perspectives: frontal, side, and oblique. The frontal view clearly shows the left and right contraction range of the pharyngeal constrictor muscles, the side view captures the up and down rotation angle of the epiglottis, and the oblique view observes the opening and closing state of the laryngeal vestibule.
[0034] Secondly, image feature parameters for each time point are extracted from the dynamic image sequence. These image feature parameters include pharyngeal contour point coordinates, motion velocity values, and motion direction angles. Image feature parameters refer to key data extracted from dynamic image frames that quantify the motion state of the pharyngeal structures, including three dimensions: position, velocity, and direction. They serve as inputs for subsequent image fusion and model construction. Pharyngeal contour point coordinates refer to the three-dimensional coordinates (x, y, z) of points on the contours of key pharyngeal structures marked in the image frame, allowing for precise structural location. The x-axis represents horizontal left and right, the y-axis represents vertical up and down, and the z-axis represents anterior and posterior depth; the origin is 2 cm directly above the Adam's apple. The motion velocity value refers to the displacement distance of the pharyngeal contour point per unit time, measured in mm / s, representing the speed of structural motion. The motion direction angle refers to the angle between the displacement vector and each axis of the three-dimensional coordinate system, ranging from 0° to 360°, representing the orientation of structural motion. The above dynamic image sequence was loaded using medical image processing software. First, each frame of the image was preprocessed by grayscale conversion and noise reduction to remove ambient light interference. Then, in the preprocessed image, three key contour points, namely the epiglottis tip, the upper edge of the laryngeal vestibule, and the midpoint of the pharyngeal constrictor muscle, were manually marked, and the x, y, and z coordinates of each point in each frame were recorded. Finally, based on the coordinate changes of adjacent frames, the motion velocity value and motion direction angle of each contour point were calculated.
[0035] For example, in the dynamic image sequence of a patient, frames 100 and 101 are selected for feature extraction: Pharyngeal contour coordinates: In frame 100 (0.33s), the coordinates of the epiglottis tip are marked as (150, 200, 45), the coordinates of the upper edge of the laryngeal vestibule are marked as (150, 180, 44), and the coordinates of the midpoint of the pharyngeal constrictor muscle are marked as (120, 200, 46); in frame 101, the coordinates of the three points become (152, 198, 43), (150, 178, 44), and (118, 200, 46), respectively. Motion velocity value: Taking the epiglottis tip as an example, the displacement distance between the two frames is... Pixel, actual distance 0.346mm, adjacent frame time interval = 1 / 300s = 0.00333s, speed = 0.283mm / 0.00333s ≈ 85mm / s. Motion direction angle: Calculated using the direction cosine, the angle with each positive axis is: angle with the positive x-axis ≈ 45° (to the right); angle with the positive y-axis ≈ 135° (downward); angle with the positive z-axis ≈ 135° (backward). This direction angle reflects the physiological movement pattern of the epiglottis during swallowing: it flips downward and to the right and shifts backward to close the laryngeal opening, providing a quantitative basis for subsequent judgment of whether the movement is abnormal.
[0036] In this embodiment of the invention, multi-view high-speed synchronous acquisition solves the problems of missing information from a single viewpoint and the inability of ordinary frame rates to capture high-speed motion. The coordinated acquisition of three views enables comprehensive recording of pharyngeal structural motion. Through three-dimensional feature extraction of pharyngeal contour point coordinates, motion velocity values, and motion direction angles, visual information in the images is transformed into quantifiable and precise data. Furthermore, the extraction process, combined with specific anatomical structures, ensures the medical relevance of the data. Ultimately, this provides comprehensive, accurate, and reliable basic data for subsequent image fusion and pharyngeal dynamic modeling, avoiding subsequent reconstruction deviations caused by initial data defects and laying the foundation for accurate reconstruction of the swallowing process.
[0037] S200: Based on the multiple image feature parameters, configure image fusion weights, perform multi-dimensional image fusion, and construct a dynamic model of the pharynx.
[0038] In this embodiment of the invention, image fusion weights are configured based on the multiple image feature parameters to perform multi-dimensional image fusion and construct a pharyngeal dynamic model. S100 has acquired image feature parameters from three perspectives: frontal, lateral, and oblique. However, the multi-view data exhibits differences: the frontal view shows high clarity of the pharyngeal contour but insufficient depth information; the lateral view shows good continuity of the motion trajectory but lower detail resolution; and the oblique view balances both but is susceptible to lighting interference. Directly stitching multi-view data can lead to subsequent analysis biases due to information redundancy or conflicting key features. Furthermore, pharyngeal swallowing is a coupled process of geometric structure and temporal motion, and traditional static modeling methods cannot capture the coordinated dynamics of epiglottic inversion and pharyngeal constrictor muscle contraction. Therefore, this step requires multi-dimensional image fusion and the construction of a pharyngeal dynamic model to transform the multi-view discrete data into a unified and complete pharyngeal dynamic representation, providing reliable model support for subsequent trajectory analysis.
[0039] Step S200 in the method provided in this embodiment of the invention includes:
[0040] Specifically, based on the multiple image feature parameters, image fusion weights are configured to perform multi-dimensional image fusion, including:
[0041] A standard swallowing sample dataset is constructed, and reference image feature parameters of the standard swallowing process are extracted from the standard swallowing sample dataset as reference feature parameters;
[0042] Calculate the similarity between each feature parameter of the image to be fused and the reference feature parameter to obtain multiple feature similarities;
[0043] Based on the multiple feature similarities, the image fusion weight for each image viewpoint is calculated, wherein the feature similarities are positively correlated with the image fusion weight;
[0044] Based on the image fusion weights, swallowing dynamic data from multiple image perspectives are weighted and fused to obtain fused multidimensional image data.
[0045] First, a standard swallowing sample dataset is constructed, from which baseline image feature parameters of the standard swallowing process are extracted as reference feature parameters. The standard swallowing sample dataset refers to a clinically validated database of swallowing images and features, containing data from subjects of different ages, genders, and swallowing functional states. Each data point is labeled with baseline image feature parameters of the standard swallowing process, such as the average speed of epiglottic inversion and the coordinate variation range of pharyngeal constrictor muscle contraction in healthy adults, serving as a similarity comparison standard.
[0046] For example, 800 clinical swallowing angiography data were collected, including 500 healthy individuals and 300 patients with disabilities. The S100 feature extraction method was used to obtain the pharyngeal contour point coordinate sequence, motion velocity sequence, and orientation angle sequence for each data point. Three ENT physicians jointly annotated the standard swallowing baseline features. Taking the movement of the epiglottis tip as an example, the baseline velocity range of the epiglottis tip during swallowing in healthy adults is 80-120 mm / s, and the baseline coordinate change trajectory is anterior-superior, posterior-inferior, and repositioning. For the patients' swallowing scenarios, 10 samples of healthy males aged 60-70 years swallowing warm water were selected from the standard set. The average coordinates of their epiglottis tip at 0.33 seconds of swallowing were calculated to be (151, 199, 44), and the average velocity was 95 mm / s, which were used as reference feature parameters for that time point.
[0047] Next, the similarity between each image feature parameter to be fused and the reference feature parameter is calculated to obtain multiple feature similarities. Feature similarity is used to quantify the degree of matching between the feature to be fused and the baseline feature, with a value range of 0-1, where 1 is a complete match and 0 is a complete mismatch. In this embodiment, the cosine similarity algorithm is used to calculate and measure the cosine value of the angle between two feature vectors. For example, the epiglottic tip feature of the patient's three perspectives at 0.33 seconds is extracted, and the cosine similarity is calculated with the reference feature parameter respectively: Frontal view feature: coordinates (150, 200, 45), velocity 85 mm / s, the similarity with the reference feature is calculated to be 0.85; Lateral view feature: coordinates (152, 198, 43), velocity 90 mm / s, due to the high matching degree of depth information, the similarity is 0.92; Oblique view feature: coordinates (149, 201, 46), velocity 88 mm / s, affected by light and shadow interference, the similarity is 0.78.
[0048] Further, based on the multiple feature similarities, an image fusion weight is calculated for each image viewpoint, wherein the feature similarity is positively correlated with the image fusion weight. The image fusion weight refers to the confidence coefficient assigned to each viewpoint data, ranging from 0 to 1, with the sum of all viewpoint weights being 1. Higher feature similarity results in a larger weight, ensuring that high-confidence viewpoint data dominates the fusion result. Normalization is used to calculate the weights, ensuring that the weights are positively correlated with the similarity and that their sum is 1. For example, total similarity = 0.85 + 0.92 + 0.78 = 2.55; frontal viewpoint weight = 0.85 / 2.55 ≈ 0.33; side viewpoint weight = 0.92 / 2.55 ≈ 0.36; oblique viewpoint weight = 0.78 / 2.55 ≈ 0.31.
[0049] Finally, based on the image fusion weights, the swallowing dynamic data from multiple image perspectives are weighted and fused to obtain fused multidimensional image data. Weighted fusion refers to the weighted calculation of similar feature parameters at the same time point based on the weights of each perspective. The coordinates x, y, and z values and velocities of the three perspectives are weighted and calculated separately, and the direction angle is taken from the side perspective with the highest similarity. For example, taking the fusion of the three-dimensional coordinates and velocity of the epiglottis tip at 0.33 seconds as an example: fused coordinates x=150×0.33+152×0.36+149×0.31≈150.35; fused coordinates y=200×0.33+198×0.36+201×0.31≈199.51; fused coordinates z=45×0.33+43×0.36+46×0.31≈44.21; fused velocity=85×0.33+90×0.36+88×0.31≈87.53mm / s; the final output of the fused features at this time point is (150.35,199.51,44.21), velocity 87.53mm / s, and orientation angle 123.7°.
[0050] The steps for constructing the pharyngeal dynamic model include:
[0051] During the training phase, based on historical clinical swallowing image data, sample swallowing dynamic datasets were collected, and the pharyngeal structures in each swallowing dynamic dataset were three-dimensionally annotated to obtain a sample pharyngeal structure annotation set. Each sample pharyngeal structure annotation includes the spatial coordinate sequence and motion trajectory of key pharyngeal contour points.
[0052] A network architecture for constructing a dynamic model of the pharynx based on a three-dimensional convolutional neural network is provided, which includes a feature extraction module, a spatiotemporal fusion module, and a dynamic reconstruction module.
[0053] The pharyngeal dynamic model is trained under supervision using the sample swallowing dynamic dataset and the sample pharyngeal structure annotation set to obtain the trained pharyngeal dynamic model.
[0054] In the application phase, the fused multidimensional image data is input into the trained pharyngeal dynamic model, and the pharyngeal dynamic model outputs a complete swallowing dynamic result containing the pharyngeal geometry and motion sequence.
[0055] First, during the training phase, based on historical clinical swallowing image data, a sample swallowing dynamic dataset was collected, and the pharyngeal structures in each swallowing dynamic dataset were 3D annotated to obtain a sample pharyngeal structure annotation set. Each sample pharyngeal structure annotation includes the spatial coordinate sequence and motion trajectory of key pharyngeal contour points.
[0056] Specifically, the pharyngeal structures in each swallowing dynamic data point are 3D annotated to obtain a sample pharyngeal structure annotation set, including:
[0057] Collect historical clinical swallowing imaging data from clinical swallowing contrast examinations to construct an original dynamic swallowing dataset;
[0058] Based on a medical image annotation platform, key pharyngeal structures in each swallowing dynamic data are annotated, including the epiglottis, laryngeal vestibule, and pharyngeal constrictor muscle contours.
[0059] A three-dimensional spatial registration algorithm is used to fuse the annotation results and generate standardized three-dimensional annotations of the pharyngeal structure.
[0060] The consistency of the three-dimensional annotations of the pharyngeal structure is verified. When the distance error between annotations is less than a preset error threshold, it is included in the sample pharyngeal structure annotation set. Each sample pharyngeal structure annotation contains the spatial coordinate sequence of key contour points of the pharynx and the corresponding motion trajectory parameters.
[0061] Through iterative optimization, the pharyngeal structure annotation set of the sample is continuously expanded and optimized until it covers all typical swallowing stages and abnormal swallowing patterns.
[0062] First, historical clinical swallowing imaging data from clinical swallowing contrast studies were collected to construct the original dynamic swallowing dataset. Clinical swallowing contrast studies, also known as video-assisted fluoroscopy (VFSS), is a clinical assessment of swallowing function. By having subjects swallow food containing contrast agent, the swallowing process is dynamically captured on X-ray during the pharyngeal phase, clearly showing key information such as epiglottic inversion, laryngeal vestibule closure, and the food's path. The imaging data combines medical authority with complete motor detail. The original dynamic swallowing dataset refers to a collection of unlabeled clinical swallowing contrast image sequences, requiring coverage of both temporal and population dimensions. Data source selection: Archived clinical swallowing contrast examination data from the ENT / rehabilitation departments of three tertiary hospitals over the past five years were selected, prioritizing cases with complete examination reports to ensure a one-to-one correspondence between imaging data and clinical diagnosis. Data inclusion criteria: ① Coverage of the entire population: ages 18-85, balanced gender ratio, encompassing healthy individuals, those with mild, moderate, and severe dysphagia, and other functional states; ② Coverage of all swallowing scenarios: including swallowing images of different food textures such as thin water, thick porridge, and soft rice, covering the complete swallowing stages including oral preparation, pharyngeal phase, and esophageal ingress phase; ③ Achieving satisfactory image quality: frame rate ≥ 50 frames / second, clear contrast agent imaging, and no obvious motion blur or artifacts. Data preprocessing: The selected image data underwent format standardization, frame rate normalization, and invalid frame cropping to ultimately form a structured raw swallowing dynamic dataset.
[0063] For example, a total of 1200 cases of raw data were collected, including 400 healthy individuals: 100 young people aged 20-30, 100 middle-aged people aged 40-50, and 200 elderly people over 60 years old; and 800 patients with dysphagia: 300 mild, 300 moderate, and 200 severe. Each case of data included swallowing image sequences of 2-3 food textures. Taking a similar case of a patient, namely a 65-year-old patient with mild dysphagia, the contrast images of swallowing thin water and thick porridge were collected. The raw data was 1000 frames / case of DICOM format images, covering the complete pharyngeal process of ingesting contrast agent, posterior displacement of the tongue base, epiglottis inversion, and food entering the esophagus.
[0064] Secondly, based on a medical image annotation platform, key pharyngeal structures in each swallowing dynamic data point are annotated. These key pharyngeal structures include the epiglottis, laryngeal vestibule, and pharyngeal constrictor muscle contour. The medical image annotation platform is a professional tool supporting multimodal image visualization and 3D interactive annotation. It features image frame jumping, contour point marking, real-time coordinate display, and annotation result export, adapting to the dynamic sequence annotation needs of swallowing angiography images. The epiglottis is a structure that flips to close the laryngeal opening during swallowing, preventing aspiration of food; the laryngeal vestibule is the uppermost cavity of the larynx, which closes during swallowing to prevent food from entering the airway; the pharyngeal constrictor muscle is a group of muscles surrounding the pharyngeal wall that contracts to propel food towards the esophagus. The movement of these three structures directly reflects the safety and efficiency of the patient's swallowing. The original image sequence was imported into a medical image annotation platform and independently annotated by at least three experts in dysphagia diagnosis. In each frame, ten contour points were marked: the epiglottis (two key contour points at the tip and root), the laryngeal vestibule (three contour points at the upper left, middle, and right edges), and the pharyngeal constrictor muscles (five contour points on the anterior, lateral, and posterior walls). Sequence recording: Annotations were performed sequentially from frame 1 to frame 2 according to the swallowing timeline, automatically generating a frame number-two-dimensional coordinate correspondence table for each contour point, forming a single-view structural contour annotation sequence. If the original data contained multi-view contrast images, annotation was performed separately for each view, maintaining consistent naming and number of annotation points to reserve a matching benchmark for subsequent three-dimensional fusion.
[0065] For example, taking a patient's water swallowing contrast imaging image (1000-frame sequence) as an example, after importing the image into the annotation platform, experts annotated it according to the following standards: Annotated the 3D coordinates of key structures in frame 100 (0.33s): epiglottis tip (150,200,45), midpoint of the upper border of the laryngeal vestibule (150,180,44), midpoint of the anterior wall of the pharyngeal constrictor muscle (120,200,46); Frame 101 (0.34s): epiglottis tip (152,198,43), midpoint of the upper border of the laryngeal vestibule (150,178,44), midpoint of the anterior wall of the pharyngeal constrictor muscle (118,200,46), completely recording the changes in 3D position. After completing frame by frame, a 3D coordinate sequence of 10 contour points was generated.
[0066] Furthermore, a 3D spatial registration algorithm is used to fuse the annotation results, generating standardized 3D annotations of the pharyngeal structure. The 3D spatial registration algorithm refers to an algorithm that maps the 2D coordinates of each viewpoint to a unified 3D coordinate system by finding common reference points between images from different perspectives, achieving spatial alignment and fusion of multi-view data. For example, the Iterative Closest Point (ICP) algorithm based on feature points can achieve a registration accuracy of 0.1 pixels. The standardized 3D annotations of the pharyngeal structure are fused 3D data results that conform to clinical anatomical standards, containing the 3D coordinate sequence of each key contour point, which can directly reflect the change in the spatial position of the structure with swallowing time. Unified coordinate system establishment: The origin (0,0,0) is set at the midpoint of the upper edge of the laryngeal vestibule in the first frame image, defining the x-axis horizontally to the right, the y-axis vertically upward, and the z-axis forward and backward. Registration fusion: The 3D coordinates of the multi-view annotations are input, and the transformation matrix is calculated using the ICP algorithm. All contour point coordinates are spatially aligned to eliminate viewpoint deviations; the coordinate sequence is smoothed to remove abnormal jitter points. Motion parameter calculation: Based on the coordinate difference between adjacent frames, calculate the three-dimensional displacement, velocity, and three-dimensional orientation angle of each contour point, and integrate them into complete annotation data.
[0067] For example, taking the epiglottis tip as the calculation object, based on the above-mentioned labeled coordinates: the coordinate difference between frame 101 and frame 100 is (152-150, 198-200, 43-45) = (+2, -2, -2). The actual distance is 0.3464mm, the time interval is 0.01s, and the speed is 0.3464mm / 0.01s≈34.64mm / s; the angle with the x-axis is ≈45° (to the right), the angle with the y-axis is ≈135° (downward), and the angle with the z-axis is ≈135° (backward), accurately reflecting the three-dimensional flipping trajectory of the epiglottis to the lower right and backward.
[0068] Furthermore, consistency verification was performed on the 3D annotations of the pharyngeal structures. When the distance error between annotations was less than a preset error threshold, the annotations were included in the sample pharyngeal structure annotation set. Each sample pharyngeal structure annotation included the spatial coordinate sequence of key pharyngeal contour points and the corresponding motion trajectory parameters. Consistency verification refers to a verification method in which multiple annotators independently annotate the same image data, and the consistency of annotations is judged by calculating the error between the annotation results. Distance error is used as the core indicator, that is, the Euclidean distance between the 3D coordinate differences of different annotators for the same contour point in the same frame. The preset error threshold is an upper limit of error set based on the accuracy requirements of clinical diagnosis. Combined with the resolution of 0.1 mm / pixel for swallowing contrast imaging, the preset error threshold is set to 0.3 pixels. A value below this threshold indicates that the annotation consistency meets the standard. Three experts independently completed the annotation and fusion process for the same original image data, generating three independent 3D annotation results. The three annotation results were matched according to the same contour points and the same frame number, and the Euclidean distance error of each set of matching coordinates was calculated. The average error and maximum error of the individual data were statistically analyzed. If the average error was less than 0.3 pixels and the maximum error was less than 0.5 pixels, the 3D annotation was directly included in the sample set. If the error exceeded the standard, the three experts jointly reviewed the frames with larger errors, analyzed the reasons for the deviation, re-annotated and verified again until the error met the standard. The format of the finally qualified annotation data was standardized.
[0069] For example, consistency verification was performed on the three-dimensional annotation data of a patient: the three-dimensional coordinates of the epiglottis tip (frame number 100) annotated by three experts were (150.00, 200.00, 52.00), (150.05, 200.10, 52.08), and (149.98, 199.95, 51.95), respectively. The pairwise distance errors were calculated to be 0.14 pixels, 0.11 pixels, and 0.16 pixels, respectively, with an average error of 0.14 pixels < 0.3 pixels; the maximum error of the entire sequence was 0.42 pixels < 0.5 pixels at the posterior wall point of the pharyngeal constrictor muscle (frame number 200), which was deemed to meet the standard; after format standardization, it was included in the sample pharyngeal structure annotation set, and the annotation file was named: Patient Name-Water Swallowing-Three-Dimensional Annotation.
[0070] Subsequently, through iterative optimization, the pharyngeal structure annotation set of the samples was continuously expanded and optimized until it covered all typical swallowing stages and abnormal swallowing patterns. Typical swallowing stages refer to the complete physiological sequence of pharyngeal swallowing, including the swallowing initiation phase, laryngeal closure phase, food passage phase, and recovery phase. Abnormal swallowing patterns refer to clinically common swallowing dysfunction manifestations, such as delayed epiglottic inversion, pharyngeal constrictor muscle weakness, and incomplete laryngeal vestibule closure. The number of samples in the set for typical swallowing stages and abnormal swallowing patterns was periodically counted to identify weak points; for these weak points, corresponding case images were collected, and after annotation verification, they were included in the set; experts were invited quarterly to randomly review 10% of the samples to correct biases; annotation standards were updated in accordance with new clinical guidelines; the final set was formed when the combined samples of all typical swallowing stages and abnormal swallowing patterns reached ≥20 cases and the review accuracy rate was ≥98%.
[0071] For example, the initial sample set contained 600 cases. Coverage assessment revealed only 8 cases of food passage phase + pharyngeal constrictor muscle weakness and only 5 cases of recovery phase + epiglottic repositioning incomplete repositioning. Imaging data of 15 patients with pharyngeal constrictor muscle weakness and 10 patients with epiglottic repositioning incomplete repositioning were collected and included in the set after annotation and verification according to the procedure. During clinical expert review, it was found that the pharyngeal constrictor muscle contour annotation of 12 samples was too wide. After readjusting the annotation boundary, the error met the standard. The final set contained 800 samples. The sample size of all typical swallowing stages and abnormal swallowing pattern combinations was ≥20 cases, and the review accuracy rate was 98.5%. The iteration was terminated.
[0072] Secondly, based on a 3D convolutional neural network, a network architecture for a dynamic pharyngeal model is constructed. This architecture includes a feature extraction module, a spatiotemporal fusion module, and a dynamic reconstruction module. The 3D convolutional neural network (3DCNN) simultaneously captures the spatial geometric features and temporal motion features of data through 3D convolutional kernels, making it a core algorithm for processing 3D dynamic image sequences. The spatiotemporal fusion module strengthens the weights of key stage / structural features, increasing the proportion of core information expressed. The dynamic reconstruction module maps high-dimensional features into interpretable dynamic results. First, a feature extraction module is constructed, stacking three 3D convolutional layers (3×3×3 kernel size, covering 3×3 spatial dimensions and 3 temporal frames), with ReLU activation. A 2×2×2 max-pooling layer is then added after the first and second convolutional layers. Next, a spatiotemporal fusion module is constructed, embedding a self-attention mechanism to calculate feature association weights, followed by a batch normalization layer. Finally, a dynamic reconstruction module is constructed, connecting two fully connected layers and one output layer, with Sigmoid activation, completing the module concatenation. For example, using a patient's swallowing fusion data as input, the feature extraction module extracts three-dimensional contour features from the epiglottic region in frame 100; the spatiotemporal fusion module increases the feature weight of the key stage of epiglottic flipping (frames 100-125) to 0.8; and the dynamic reconstruction module outputs the coordinate temporal sequence of the epiglottic tip in this stage.
[0073] Furthermore, the pharyngeal dynamic model is trained under supervised supervision using the sample swallowing dynamic dataset and the sample pharyngeal structure annotation set to obtain a trained pharyngeal dynamic model. Supervised training is a learning process that uses the labeled data as the standard answer and optimizes the model parameters through error feedback. The validation set is a sample set used to adjust the training parameters to avoid model overfitting. The early stopping mechanism is a strategy to terminate training when the error in the validation set does not decrease. For example, 1000 cases of 3D fusion data and 800 cases of labeled data are divided into training, validation, and test sets in a 7:2:1 ratio; features such as coordinates and velocities are normalized to the 0-1 interval; the Adam optimizer is used to iterate training at 16 cases / batch using mean squared error as the loss function; and the early stopping mechanism is introduced to stop training when the coordinate prediction accuracy of the test set is ≥96% and the velocity error is ≤3mm / s. For example, when training a patient with a similar delayed epiglottic flipping sample, the error between the predicted epiglottic tip velocity and the labeled value in the 80th round of the test set is only 2.1mm / s, and the model is saved after reaching the target.
[0074] Finally, in the application phase, the fused multidimensional image data is input into the trained pharyngeal dynamic model, which outputs a complete swallowing dynamic result including pharyngeal geometry and motion sequences. Model inference is the process of generating prediction results by calling the trained weights. Denormalization is the operation of restoring the normalized values of the model output to the actual physiological parameters. The three-dimensional geometric model is a visually viewable spatial morphology file of the pharyngeal structure. After verifying the format of the input fused data, it is normalized according to the training standard; the model weights are called to complete the inference; the output result is denormalized, and then Gaussian filtering is used to smooth the jitter; the three-dimensional geometric model in STL format and the motion sequence report containing coordinates / velocities are output. For example, after inputting the fused data of a patient, the output three-dimensional model shows that the contraction amplitude of the pharyngeal constrictor muscle is 30% lower; the velocity of the epiglottis tip in frames 100-125 in the motion sequence report is 45-55 mm / s, matching the diagnosis: mild dysphagia.
[0075] In this embodiment of the invention, a fusion strategy adapted to multi-view data is used to eliminate biases and redundancies in information from different perspectives, integrating complete dynamic data of the pharynx and providing a reliable foundation for subsequent analysis. Based on a three-dimensional convolutional neural network model architecture, the spatial morphology and temporal motion characteristics of the pharyngeal structure are captured simultaneously, achieving precise dynamic reconstruction of the swallowing process. The structured output results can directly support clinical assessment of swallowing function, improving the accuracy and reliability of swallowing process reconstruction.
[0076] S300: Based on the pharyngeal dynamic model, perform pharyngeal motion trajectory analysis to obtain an initial motion trajectory, calculate the comprehensive trajectory similarity between the initial motion trajectory and the reference motion trajectory, construct an optimization objective function based on the comprehensive trajectory similarity, and use an optimization algorithm to iteratively optimize the initial motion trajectory to obtain a corrected motion trajectory.
[0077] In this embodiment of the invention, based on the pharyngeal dynamic model, pharyngeal movement trajectory analysis is performed to obtain an initial movement trajectory. The comprehensive trajectory similarity between the initial movement trajectory and a reference movement trajectory is calculated. An optimization objective function is constructed based on the comprehensive trajectory similarity, and an optimization algorithm is used to iteratively optimize the initial movement trajectory to obtain a corrected movement trajectory. Although the pharyngeal dynamic model output by S200 has restored the basic movement state, the initial movement trajectory may be affected by image noise and individual anatomical differences, resulting in local offsets or temporal misalignments, making it unsuitable for direct clinical abnormality assessment. This step, through trajectory extraction, similarity comparison, and iterative optimization, combined with a standard reference trajectory to calibrate the initial trajectory, eliminates deviations and enhances the clinical comparability of the trajectory, providing accurate trajectory evidence for the identification of swallowing dysfunction.
[0078] Step S300 in the method provided in this embodiment of the invention includes:
[0079] Specifically, based on the pharyngeal dynamic model, pharyngeal motion trajectory analysis is performed to obtain an initial motion trajectory, and the comprehensive trajectory similarity between the initial motion trajectory and a reference motion trajectory is calculated, including:
[0080] Motion path data of key points in the pharynx are extracted from the dynamic model of the pharynx to construct an initial set of motion trajectories;
[0081] A reference motion trajectory is obtained from a standard swallowing sample dataset, wherein the reference motion trajectory is obtained by statistical analysis of the motion trajectory of the standard swallowing process;
[0082] The trajectory similarity between the initial motion trajectory and the reference motion trajectory is calculated, including trajectory shape similarity and motion temporal similarity, wherein the trajectory shape similarity is calculated based on the dynamic time warping algorithm, and the motion temporal similarity is calculated based on the correlation coefficient method;
[0083] The trajectory shape similarity and motion temporal similarity are weighted and fused to obtain a comprehensive trajectory similarity.
[0084] First, motion path data of key pharyngeal points are extracted from the pharyngeal dynamic model to construct an initial motion trajectory set. Key pharyngeal points reflect the core structural locations of swallowing function during the pharyngeal phase, and their movement directly relates to swallowing safety and efficiency. The initial motion trajectory set refers to the complete path sequence combination of the spatial coordinates of each key pharyngeal point changing over swallowing time, serving as the foundational data for trajectory analysis. Based on the output of the pharyngeal dynamic model, preset key points such as the epiglottis apex and the midpoint of the laryngeal vestibule are selected. The spatial coordinate data of each point per frame is exported, categorized and organized into continuous path sequences, and all path sequences are combined to form the initial motion trajectory set. For example, from a patient's pharyngeal dynamic model, the coordinate sequences of the epiglottis apex and the midpoint of the laryngeal vestibule are exported, organized into epiglottis apex path sequences and laryngeal vestibule midpoint path sequences respectively, collectively forming the initial motion trajectory set.
[0085] Secondly, reference motion trajectories are obtained from a standard swallowing sample dataset. These reference motion trajectories are obtained through statistical analysis of the motion trajectories during a standard swallowing process. The reference motion trajectory refers to the statistical mean of the swallowing trajectories of healthy individuals in the standard swallowing sample dataset who are consistent with the target subject's age and food characteristics; it serves as a physiological benchmark for determining whether the trajectory is normal. From the standard swallowing sample dataset, healthy samples matching the target subject's age and swallowed food characteristics are selected. The motion trajectories of corresponding key points in the pharynx are extracted, and the mean sequence of coordinates for each frame is calculated to obtain the reference motion trajectory for that point. For example, from the standard dataset, samples of 10 healthy 65-year-old men swallowing warm water are selected, the motion trajectory of the epiglottis tip of each sample is extracted, and the mean of coordinates for each frame is calculated to generate the reference motion trajectory of the epiglottis tip.
[0086] Further, the trajectory similarity between the initial trajectory and the reference trajectory is calculated, including trajectory shape similarity and motion temporal similarity. The trajectory shape similarity is calculated based on the Dynamic Time Warping (DTW) algorithm, and the motion temporal similarity is calculated based on the correlation coefficient method. Trajectory shape similarity is an indicator of the degree of matching in the spatial path shape of two trajectories. The Dynamic Time Warping (DTW) algorithm is an algorithm that eliminates the influence of trajectory length differences by dynamically adjusting the temporal alignment method to calculate spatial path shape differences. Motion temporal similarity is an indicator of the consistency of the changing trends of two trajectories over time. The correlation coefficient method is a method that quantifies the degree of matching in temporal changing trends by calculating the Pearson correlation coefficient between sequences. The coordinate sequences of the initial trajectory and the reference trajectory are imported into the trajectory analysis tool, and the DTW algorithm is called to calculate the cumulative distance between the two sequences. The distance value is normalized to the 0-1 interval to obtain the trajectory shape similarity; the closer the value is to 1, the higher the shape matching degree. For the coordinate sequences of the initial trajectory and the reference trajectory, the Pearson correlation coefficient is calculated along the time dimension; the coefficient value is the motion temporal similarity; the closer the value is to 1, the more consistent the temporal changing trends. For example, by importing the initial trajectory and reference trajectory of a patient's epiglottis, and after calculating the cumulative distance using the DTW algorithm and normalizing it, the trajectory shape similarity is found to be 0.78; by calculating the Pearson correlation coefficient between the initial trajectory and reference trajectory of a patient's epiglottis, the motion temporal similarity is found to be 0.72.
[0087] Finally, the trajectory shape similarity and motion temporal similarity are weighted and fused to obtain the comprehensive trajectory similarity. The comprehensive trajectory similarity is a comprehensive index of the degree of trajectory matching by fusing the trajectory shape and motion temporal dimensions. Weighted fusion refers to a fusion method that assigns weights to the sub-similarities according to clinical priority and performs linear calculation. Based on the characteristic that clinical attention to trajectory morphology is higher than that to temporal trend, a higher weight is assigned to trajectory shape similarity and a lower weight to motion temporal similarity. The two similarities are multiplied by their respective weights and then summed to obtain the comprehensive trajectory similarity. For example, setting the weight of trajectory shape similarity to 0.6 and the weight of motion temporal similarity to 0.4, the comprehensive trajectory similarity of a patient's epiglottic tip is calculated as: 0.78 × 0.6 + 0.72 × 0.4 = 0.756.
[0088] Specifically, an optimization objective function is constructed based on the comprehensive trajectory similarity, and an optimization algorithm is used to iteratively optimize the initial motion trajectory to obtain a corrected motion trajectory, including:
[0089] To maximize the comprehensive trajectory similarity, an optimization objective function is constructed, wherein the optimization objective function considers both trajectory shape similarity and motion temporal similarity.
[0090] The initial trajectory is iteratively optimized using a particle swarm optimization algorithm, with particle swarm size parameters and iteration control parameters set.
[0091] In each iteration, new candidate trajectories are generated by updating particle positions and velocities, the corresponding comprehensive trajectory similarity is calculated, and the individual optimal solution and the global optimal solution are recorded.
[0092] Based on the comprehensive trajectory similarity, the particle search strategy is dynamically adjusted, and when the iteration termination condition is met, the globally optimal motion trajectory is output as the corrected motion trajectory.
[0093] First, an optimization objective function is constructed with the goal of maximizing the comprehensive trajectory similarity. This objective function considers both trajectory shape similarity and motion temporal similarity. The objective function is a mathematical expression centered on maximizing the comprehensive trajectory similarity, used to quantify the optimization direction while constraining the matching accuracy of trajectory shape and motion temporal sequence. Comprehensive trajectory similarity is a weighted index that integrates trajectory shape similarity and motion temporal similarity, with weights set based on clinical attention. Using the goal of maximizing comprehensive trajectory similarity, and combining the weight allocation of trajectory shape similarity and motion temporal similarity (e.g., shape weight 0.6, temporal weight 0.4), the objective function is constructed as follows: Objective function value = Shape weight × Trajectory shape similarity + Temporal weight × Motion temporal similarity. This clarifies that the function needs to simultaneously optimize the similarity of both dimensions. For example, for trajectory optimization of a patient's epiglottis tip, the objective function is constructed as: Comprehensive trajectory similarity = max(0.6 × Trajectory shape similarity + 0.4 × Motion temporal similarity), ensuring that the optimization process both conforms to the normal trajectory shape and matches the physiological movement rhythm.
[0094] Secondly, the initial trajectory is iteratively optimized using a particle swarm optimization (PSO) algorithm, with particle swarm size parameters and iteration control parameters set. PSO is a swarm intelligence optimization algorithm where each particle searches for the optimal solution in the solution space by mimicking the cooperative behavior of a flock of birds; these particles are the candidate trajectories. The particle swarm size parameter refers to the number of candidate trajectories participating in the search, affecting search diversity and computational efficiency. The iteration control parameters are the core parameters for regulating the optimization process, including the maximum number of iterations and the convergence threshold. First, the particle swarm size is set based on the complexity of the initial trajectory: 50-100 particles for trajectories with many nodes and high dimensionality, and 30-50 particles for simple trajectories. Then, the iteration control parameters are set based on the initial comprehensive trajectory similarity: 100-200 iterations for similarity < 0.6, and 50-100 iterations for similarity between 0.6 and 0.8. Simultaneously, a convergence threshold is set: a similarity change of <1e-5 or a similarity ≥ 0.95 after 5-10 consecutive iterations. For example, the trajectory of a patient's epiglottis tip is a 600-frame sequence of three-dimensional coordinates, which is highly complex. The particle swarm size is set to 50. The initial comprehensive trajectory similarity is 0.756. The maximum number of iterations is set to 100. The convergence threshold is: the similarity change is <1e-5 or the similarity is ≥0.95 after 8 consecutive iterations.
[0095] Furthermore, in each iteration, new candidate trajectories are generated by updating particle positions and velocities, the corresponding comprehensive trajectory similarity is calculated, and individual optimal solutions and global optimal solutions are recorded. A particle corresponds to a candidate trajectory, its position being the coordinate sequence of the trajectory, and its velocity representing the magnitude and direction of coordinate adjustment. The individual optimal solution refers to the candidate trajectory with the highest comprehensive trajectory similarity among all iterations for a single particle. The global optimal solution refers to the candidate trajectory with the highest comprehensive trajectory similarity among the individual optimal solutions of all particles. Particles are initialized, and 50 candidate trajectories with structures similar to the initial trajectory are randomly generated. In each iteration, particle positions and velocities are updated according to the particle swarm optimization algorithm rules, generating new candidate trajectories. The comprehensive similarity of each new trajectory is calculated, and the individual optimal solution is updated by comparing it with the particle's historical best value. Finally, the global optimal solution is updated by comparing it with all individual optimal solutions. For example, 50 candidate trajectories for the tip of a patient's epiglottis are initialized. After the first iteration, the overall similarity of a particle's trajectory is 0.81, and it is updated to be the best individual trajectory for that particle. After the 50th iteration, the best individual trajectory similarity of that particle increases to 0.90, which is better than other particles, and it becomes the global optimal solution.
[0096] Finally, based on the comprehensive trajectory similarity, the particle search strategy is dynamically adjusted. When the iteration termination condition is met, the globally optimal motion trajectory is output as the corrected motion trajectory. The dynamic search strategy refers to adjusting the particle search range according to the comprehensive similarity of the globally optimal solution. The higher the similarity, the smaller the search range, focusing on the optimal region for refined optimization. The iteration termination condition is the basis for determining when optimization stops. It can be terminated if any one of the following conditions is met: reaching the maximum number of iterations, the globally optimal solution being stable for multiple consecutive rounds, or the similarity reaching the target threshold. After each iteration, if the comprehensive similarity of the globally optimal solution is high, the search range of the corresponding particle is reduced; the termination condition is checked in real time: reaching the maximum number of iterations, or the global optimal similarity change being <1e-5 for 5-10 consecutive iterations, or the similarity being ≥0.95; after the condition is met, the trajectory corresponding to the globally optimal solution is output as the corrected motion trajectory. For example, after the 80th iteration, the comprehensive similarity of the globally optimal trajectory of a patient's epiglottis reaches 0.92, and the similarity change is <1e-5 for 8 consecutive iterations, triggering the termination condition; after dynamic adjustment, this trajectory is output as the corrected motion trajectory of the epiglottis. The obtained corrected motion trajectory is verified for swallowing coordination to ensure that it conforms to the physiological characteristics of pharyngeal movement and meets the preset trajectory accuracy requirements.
[0097] In this embodiment of the invention, an initial trajectory is constructed by extracting the motion paths of key points in the pharynx, providing structured foundational data for subsequent analysis and ensuring the relevance of the trajectory analysis. A comprehensive approach is adopted, integrating trajectory similarity with both shape and temporal dimensions. Dynamic time warping and correlation coefficient methods address the limitations of single-dimensional comparisons, enhancing the comprehensiveness of trajectory matching assessment. A particle swarm optimization algorithm, combined with a dynamic search strategy, efficiently iteratively optimizes the initial trajectory, balancing search diversity and computational efficiency through reasonable parameter settings while avoiding local optima, thus improving the matching degree between the trajectory and the reference benchmark. The optimized corrected motion trajectory is verified for consistency, conforming to the physiological characteristics of pharyngeal movement and meeting accuracy requirements, effectively correcting noise and deviations in the initial trajectory. This further strengthens the reliability of the trajectory data, providing accurate trajectory evidence for the identification of swallowing dysfunction and contributing to improved accuracy in clinical diagnosis.
[0098] S400: Based on the corrected motion trajectory, construct a swallowing restoration simulator, perform swallowing restoration simulation, and output the pharyngeal swallowing restoration results.
[0099] In this embodiment of the invention, a swallowing reconstruction simulator is constructed based on the corrected motion trajectory to simulate swallowing and output the pharyngeal swallowing reconstruction results. Although the corrected motion trajectory output by S300 has eliminated deviations and possesses accurate quantitative characteristics, this trajectory is essentially a three-dimensional coordinate time sequence of key points in the pharynx, presented only in data form, and cannot intuitively reflect the dynamic coordination relationship of pharyngeal structures and the complete scenario of food transmission during swallowing. Clinicians need to reconstruct the swallowing process concretely to intuitively judge abnormalities such as delayed epiglottic inversion and food residue. To this end, S400, through the process of constructing a professional simulator, configuring clinical parameters, dynamic simulation, and verifying the output, transforms the abstract corrected trajectory into a visualized swallowing process that fits clinical reality, providing an intuitive and reliable basis for clinical diagnosis.
[0100] Step S400 in the method provided in this embodiment of the invention includes:
[0101] Based on the corrected motion trajectory, a computational framework for a swallowing reconstruction simulator is constructed, wherein the computational framework includes a motion trajectory analysis unit, a pharyngeal dynamics calculation unit, and a swallowing process reconstruction unit;
[0102] Based on historical clinical swallowing process data, the simulation parameters of the swallowing simulation simulator are configured, wherein the simulation parameters include pharyngeal structure motion parameters and food passage path parameters;
[0103] The corrected motion trajectory is input into the swallowing simulation simulator. The motion sequence is analyzed by the motion trajectory analysis unit, the interaction between the pharyngeal geometric structures is calculated by the pharyngeal dynamics calculation unit, and a complete swallowing process simulation is generated by the swallowing process restoration unit.
[0104] The simulation results are validated for completeness. When the simulation results include the complete pharyngeal structure movement process and food passage path, the pharyngeal swallowing reconstruction results are output.
[0105] First, based on the corrected motion trajectory, a computational framework for a swallowing reconstruction simulator is constructed. This framework includes a motion trajectory analysis unit, a pharyngeal dynamics calculation unit, and a swallowing process reconstruction unit. The swallowing reconstruction simulator is a specialized tool that integrates motion analysis, mechanical calculation, and dynamic reconstruction functions based on the corrected motion trajectory, transforming the quantified trajectory into a visualized swallowing process. The motion trajectory analysis unit is responsible for decomposing the coordinate time-series data of the corrected motion trajectory into executable structural motion commands, clarifying the direction, speed, and time point of each key point. The pharyngeal dynamics calculation unit, based on pharyngeal anatomy and biomechanical models, calculates the interactions between structures during swallowing, such as the closing pressure of the epiglottis and laryngeal vestibule, and the frictional force between the pharyngeal constrictor muscles and food. The swallowing process reconstruction unit integrates the output data from the first two units along the time axis to generate a continuous simulation scene containing pharyngeal structural dynamics and the food transport path. Framework architecture design: A serial architecture is adopted, connecting the motion trajectory analysis unit, pharyngeal dynamics calculation unit, and swallowing process reconstruction unit in the logical order of data input, analysis, calculation, and reconstruction to ensure continuous data flow; based on the physiological characteristics of pharyngeal swallowing, basic rules are preset for each unit: the motion trajectory analysis unit needs to identify key timing nodes such as the swallowing initiation period and the laryngeal closure period; the pharyngeal dynamics calculation unit embeds a timing synchronization model of epiglottic inversion and laryngeal vestibule closure; the swallowing process reconstruction unit supports real-time overlay rendering of structural motion and food flow.
[0106] For example, a computational framework is constructed for the corrected motion trajectory of a patient: the motion trajectory analysis unit presets the temporal node recognition rules: frames 1-50 are the swallowing initiation period, and frames 51-150 are the laryngeal closure period; the pharyngeal dynamics calculation unit loads the biomechanical model: when the epiglottis inversion amplitude is ≥30°, the laryngeal vestibule begins to close; the swallowing process reconstruction unit configures the rendering rules: the pharyngeal structure is semi-transparent and the food is blue fluid and visualized, to ensure that the structure and food can be intuitively distinguished in subsequent simulations.
[0107] Secondly, based on historical clinical swallowing process data, the simulation parameters of the swallowing simulation simulator are configured. These simulation parameters include pharyngeal structure motion parameters and food passage path parameters. These simulation parameters are key quantitative indicators supporting the swallowing simulation's close resemblance to clinical reality, and include pharyngeal structure motion parameters and food passage path parameters. Pharyngeal structure motion parameters describe the movement of the pharyngeal structure itself, such as the range of movement speed and the threshold of contraction amplitude; food passage path parameters describe the transmission of food in the pharynx, such as food flow rate and the range of path boundaries. Historical clinical swallowing process data refers to a multi-center collection of swallowing imaging data containing complete clinical diagnoses, covering cases of different ages, food characteristics, and swallowing function states, and serves as the objective basis for parameter configuration. Historical clinical data were selected and matched based on the target subjects' age, swallowed food characteristics, and basic physiological indicators, with an age error of ≤5 years, consistent food characteristics, and similar swallowing function grades. From the matched historical data, motion parameters such as the range of motion speed and contraction amplitude of key pharyngeal structures, as well as path parameters such as the flow rate and transmission path boundaries of corresponding food, were extracted, and the statistical mean was taken as the initial parameters. The initial parameters were then fine-tuned based on the corrected motion trajectory characteristics of the target subjects to ensure that the parameters and trajectory were well matched.
[0108] For example, consider a 65-year-old patient with mild dysphagia who swallows 10ml of warm water. The simulation parameters are configured as follows: Historical data screening: Select 20 cases of patients aged 60-70 with mild dysphagia who swallowed warm water from historical clinical data; Initial parameter extraction: Calculate the statistical mean of the 20 cases to obtain the initial parameters: Structural motion parameters: Epiglottic inversion speed 45-55mm / s, pharyngeal constrictor muscle contraction amplitude 35-45 pixels, laryngeal vestibule closure time 0.15-0.25s; Food path parameters: Warm water flow rate 18-22mm / frame, path boundary within a 5-pixel range of the pharyngeal wall anatomical contour; Adaptation adjustment: The average epiglottic inversion speed in the patient's calibration trajectory is 50mm / s, consistent with the initial parameters, requiring no adjustment; Finally, determine the simulation parameters to ensure they match the patient's physiological characteristics.
[0109] Then, the corrected motion trajectory is input into the swallowing simulation simulator. The motion sequence is analyzed by the motion trajectory analysis unit, the interaction between pharyngeal geometric structures is calculated by the pharyngeal dynamics calculation unit, and a complete swallowing process simulation is generated by the swallowing process reconstruction unit. Motion sequence analysis refers to the motion trajectory analysis unit converting the three-dimensional coordinate time-series data of the corrected motion trajectory into structured motion instructions, clarifying the motion direction, velocity, and displacement of each key point in each frame. The interaction between pharyngeal geometric structures refers to the coordinated motion relationship and mechanical action of various pharyngeal structures during swallowing, such as the epiglottis flipping to cover the laryngeal vestibule to prevent aspiration, the contraction of the pharyngeal constrictor muscles to propel food into the esophagus, and the closing pressure of the epiglottis and laryngeal vestibule. Trajectory Input and Analysis: The corrected motion trajectory is imported into the swallowing simulation simulator. The motion trajectory analysis unit breaks down the coordinate data frame by frame, generating structured motion instructions with frame number, key points, motion direction, velocity, and displacement. Dynamics Calculation: The pharyngeal dynamics calculation unit reads the motion instructions and, in conjunction with the configured simulation parameters, calculates the cooperative relationship and mechanical parameters between structures in each frame. Dynamic Reconstruction and Rendering: The swallowing process reconstruction unit integrates the motion instructions and mechanical calculation results along the time axis, rendering the dynamic movement of the pharyngeal structure and the flow of food in real time, generating a continuous simulation animation.
[0110] For example, the patient's corrected motion trajectory (600-frame coordinate sequence of the epiglottis tip: frame 100 (150,200,45), frame 101 (152,198,43) ... frame 125 (118,88,43)) is input into the simulator: Motion sequence parsing: Output structured motion instructions, such as: frames 100-125 (0.33-0.42s): the epiglottis tip moves downward and to the right at a speed of 50mm / s (x-axis +2, y-axis -2, z-axis -2). The epiglottis flips downwards and to the right, with a cumulative displacement of 40 pixels. Dynamic calculation: At frame 110 (epiglottic flip amplitude of 30°), the laryngeal vestibule begins to close, with a closing pressure of 0.8 kPa. At frame 130, the thrust is 1.2 N, pushing the warm water flow rate to 20 mm / frame. Dynamic rendering: An animation is generated to visually present the complete process of the epiglottis flipping downwards and to the right from frame 100 to 125, the laryngeal vestibule gradually closing, and the pharyngeal constrictor muscle contracting in frame 130 to push the blue warm water along the pharyngeal wall towards the esophagus.
[0111] Furthermore, the simulation results are validated for completeness. When the simulation results include the complete pharyngeal structure movement process and food passage path, the pharyngeal swallowing reconstruction results are output.
[0112] The integrity verification of the simulation results includes:
[0113] Check the integrity of the pharyngeal structure movement process in the simulation results and verify whether the movement trajectories of each key structure are continuous and complete;
[0114] Verify the continuity of the swallowing sequence and check whether the movement sequence of each pharyngeal structure conforms to the physiological laws of swallowing;
[0115] When the verification result does not meet the preset integrity standard, the comprehensive trajectory similarity between the simulation result and the reference motion trajectory is calculated.
[0116] Based on the comprehensive trajectory similarity, the adjustment direction and adjustment range of the simulation parameters are determined. The adjustment direction is determined according to the deviation between the trajectory shape similarity and motion timing similarity and the target value. The adjustment range is calculated proportionally based on the difference between the current comprehensive trajectory similarity and the target similarity.
[0117] Based on the adjusted simulation parameters, a new swallowing simulation is performed to generate new simulation results. The verification and adjustment process is repeated until the simulation results meet the preset integrity standards.
[0118] First, the integrity of the pharyngeal structure movement process in the simulation results is checked to verify whether the motion trajectories of each key structure are continuous and complete. Pharyngeal structure movement integrity means that the motion trajectories of all preset key structures in the simulation results must continuously cover the entire swallowing process without breaks, jumps, or missing key structures, ensuring the continuity of structural movement. The motion trajectory sequences of each preset key structure are extracted from the simulation results and compared with the corrected motion trajectory output by the S300, checking two dimensions: ① Coverage: Does it include all key structures? ② Continuity: Are the motion coordinates of each structure continuous in each frame, without abrupt changes in coordinates between frames? For example, based on the patient's simulation results, the trajectory sequences of three key structures—the epiglottis tip, the midpoint of the laryngeal vestibule, and the midpoint of the anterior wall of the pharyngeal constrictor muscle—were extracted: ① Coverage check: All three key structures were present without any missing structures; ② Continuity check: It was found that the coordinates of the epiglottis tip in frames 110-111 abruptly changed from (130,150,44) to (120,130,44), with an inter-frame displacement of 14.14 pixels. This far exceeded the maximum inter-frame displacement of 50 mm / s / 30 frames / s × 10 pixels / mm ≈ 1.67 pixels corresponding to the epiglottis flipping speed in the simulation parameters, indicating that the structural motion process was incomplete.
[0119] Secondly, the continuity of the swallowing sequence is verified, checking whether the movement sequence of each pharyngeal structure conforms to the physiological laws of swallowing. The continuity of the swallowing sequence means that the order and time interval of the movements of each pharyngeal structure in the simulation results must conform to the physiological laws of swallowing. For example, the laryngeal vestibule should close within 0.1-0.2 seconds after the epiglottis initiates movement, and the contraction of the pharyngeal constrictor muscle should be later than the epiglottis initiation, ensuring that the timing logic is consistent with the human physiological mechanism. Based on historical clinical data, a swallowing physiological timing rule base is compiled, and the initiation frame number of the movement of each key structure in the simulation results is extracted. The time interval between the initiation of movement between structures is calculated, and it is checked whether it meets the threshold of the rule base: ① Sequence: For example, the initiation frame number of epiglottis initiation must be earlier than the initiation frame number of laryngeal vestibule closure; ② Time interval: For example, the difference between the initiation frames of epiglottis initiation and laryngeal vestibule closure should correspond to 3-6 frames. For example, calling the rule in the rule base: the laryngeal vestibule closes within 3-6 frames after the epiglottis initiates. Checking the patient's simulation results, the epiglottis initiates at frame 100, and the laryngeal vestibule closes at frame 108, with a frame difference of 8 frames, which exceeds the physiological threshold of 3-6 frames, the swallowing sequence is determined to be discontinuous.
[0120] Furthermore, when the verification result does not meet the preset integrity standard, the comprehensive trajectory similarity between the simulation result and the reference motion trajectory is calculated. The target comprehensive trajectory similarity refers to the benchmark index used to measure the degree of matching between the current simulation result and the reference trajectory when the verification fails. The preset target value is ≥0.9, and parameters need to be adjusted if it is lower than this value. If the above verification fails, the motion trajectory of each key structure in the simulation result is extracted. Using the reference motion trajectory in the standard swallowing sample dataset as the benchmark, the calculation method of S300 is followed: ① The trajectory shape similarity is calculated using the dynamic time warping algorithm; ② The motion temporal similarity is calculated using the correlation coefficient method; ③ The current comprehensive trajectory similarity is obtained by weighting and fusing the data with a shape weight of 0.6 and a temporal weight of 0.4. For example, the simulation results of this patient failed the validation due to a sudden change in the trajectory of the epiglottis tip and abnormal temporal intervals. The current simulated trajectory of the epiglottis tip was extracted and compared with the reference trajectory to calculate: ① Shape similarity 0.80; ② Temporal similarity 0.85; ③ Comprehensive trajectory similarity = 0.80×0.6+0.85×0.4=0.82, which is lower than the target value of 0.9, and the parameters need to be adjusted.
[0121] Subsequently, based on the comprehensive trajectory similarity, the adjustment direction and magnitude of the simulation parameters are determined. The adjustment direction is determined by the deviation between the trajectory shape similarity and motion temporal similarity and the target value. The adjustment magnitude is calculated proportionally based on the difference between the current comprehensive trajectory similarity and the target similarity. The adjustment direction refers to clarifying the type of simulation parameter that needs correction based on the deviation between the shape / temporal similarity and the target value. The adjustment magnitude refers to calculating the specific adjustment amount of the parameters proportionally based on the difference between the current comprehensive similarity and the target value, ensuring adjustment accuracy. Determining the adjustment direction: If the trajectory shape similarity is low, focus on adjusting the pharyngeal structure motion parameters; if the motion temporal similarity is low, focus on adjusting the temporal correlation parameters of the structure motion. Calculating the adjustment magnitude: Adjustment magnitude = Adjustment coefficient × (Target similarity - Current similarity), with the adjustment magnitude controlled within ±20%. The adjustment coefficient is a constant pre-set based on historical data. For example, by analyzing the effect of multiple parameter adjustments in the sample swallowing process record set, the adjustment coefficient is determined to be in the range of 0.1 to 0.3. The specific value is determined according to the correspondence between the parameter adjustment range and the similarity improvement effect in the historical swallowing process data.
[0122] For example, the patient's overall trajectory similarity is 0.82, shape similarity is 0.80, and temporal similarity is 0.85: ① Adjusting direction: Adjusting structural motion parameters (epiglottic tumbling speed, current value 50 mm / s) and temporal correlation parameters (motion interval threshold, current corresponding frame difference 8 frames); ② Determining adjustment coefficients: The adjustment coefficient for epiglottic tumbling speed is 0.28, and the adjustment coefficient for motion interval threshold is 0.25; ③ Calculating adjustment magnitude: Epiglottic tumbling speed adjustment magnitude = 0.28 × (0.9 - 0.82) = 0.0224 (unit: mm / s × 100, adapted according to parameter magnitude), the actual adjustment is 50 + 2.24 ≈ 52.24 mm / s; Motion interval threshold adjustment magnitude = 0.25 × (0.9 - 0.82) = 0.02 (unit: frames × 100), the actual adjustment is 8 - 2 = 6 frames (fitting the physiological threshold of 3-6 frames).
[0123] Based on this, the swallowing reconstruction simulation is re-performed according to the adjusted simulation parameters to generate new simulation results. The verification and adjustment process is repeated until the simulation results meet the preset integrity standards. Iterative verification optimizes the simulation results step by step through a cycle of adjusting parameters, re-simulating, and re-verifying until the integrity standards are met, forming a closed-loop optimization mechanism. Re-simulation: The adjusted parameters are imported into the swallowing reconstruction simulator, the corrected motion trajectory is input, and the simulation results are regenerated; Repeated verification: The structural motion integrity and temporal continuity are checked again according to the above standards. If the verification is passed, the result is output; if it is not passed, the adjustment is repeated until the standards are met. For example, after adjusting the epiglottic turning speed of this patient to 54.4 mm / s and the motion interval threshold to 5 frames, the simulation is re-performed: ① Structural motion integrity check: The maximum inter-frame displacement of the epiglottic tip is 1.8 pixels, which meets the speed threshold, and the trajectory is continuous; ② Temporal continuity check: Epiglottic turning starts at frame 100, and laryngeal vestibule closure starts at frame 105, with a frame difference of 5 frames, corresponding to 0.17 seconds, which meets the physiological threshold; The verification is passed, and the final pharyngeal swallowing reconstruction result is output.
[0124] In this embodiment of the invention, following the corrected motion trajectory of S300, a computational framework comprising three units—motion trajectory analysis, pharyngeal dynamics calculation, and swallowing process reconstruction—is constructed to transform the abstract quantitative trajectory into a concrete simulation of the swallowing process. Simulation parameters are configured based on historical clinical data matching the target object, improving the alignment of simulation results with clinical reality. Through multi-dimensional integrity verification and iterative parameter adjustment, the simulation results effectively ensure coverage of the entire swallowing process and conformity to physiological laws. The output of visual simulation animations and quantitative parameter reports not only addresses the pain point of the difficulty in intuitively interpreting quantified trajectories but also provides doctors with observable and traceable diagnostic evidence, building a crucial bridge between data processing and clinical diagnosis, and improving the accuracy of swallowing dysfunction identification and the efficiency of clinical diagnosis.
[0125] Through the specific implementation methods described above, the embodiments of the present invention achieve the following technical effects:
[0126] This invention provides a method and system for pharyngeal swallowing reconstruction based on multi-dimensional image fusion. Through multi-stage collaborative operation, it achieves accurate reconstruction of the swallowing process from all angles. First, multi-view data fusion eliminates information bias, providing a complete and reliable dynamic basis for subsequent analysis of the pharynx. Then, a network architecture based on 3DCNN synchronously captures the spatial morphology and temporal motion characteristics of the pharyngeal structure to generate an initial dynamic model. Subsequently, the initial trajectory is iteratively optimized through comprehensive trajectory similarity calculation and particle swarm optimization algorithm, effectively correcting noise and offset to obtain a corrected motion trajectory that conforms to physiological laws. Finally, relying on a three-unit collaborative swallowing reconstruction simulator, combined with matched clinical data and configured parameters, and after integrity verification and iterative parameter adjustment, a visual simulation animation and quantitative report are output. The entire process forms a closed loop of data fusion, model construction, trajectory optimization, and simulation reconstruction. Through precise control of each stage and adaptation to clinical data, the accuracy, completeness, and clinical fit of swallowing process reconstruction are improved, providing a comprehensive and reliable basis for the identification of swallowing dysfunction, thereby effectively improving the efficiency and accuracy of clinical diagnosis.
[0127] Example 2, as Figure 2 As shown, this invention provides a pharyngeal swallowing reconstruction system based on multidimensional image fusion, the system comprising:
[0128] The multi-dimensional image acquisition module 11 is used to acquire multi-dimensional image sequences of the pharyngeal swallowing process through a multi-modal image acquisition device, obtain swallowing dynamic data from multiple image perspectives, and extract multiple image feature parameters.
[0129] The image fusion modeling module 12 is used to configure image fusion weights according to the multiple image feature parameters, perform multi-dimensional image fusion, and construct a dynamic model of the pharynx.
[0130] The trajectory optimization and correction module 13 is used to perform pharyngeal motion trajectory analysis based on the pharyngeal dynamic model, obtain an initial motion trajectory, calculate the comprehensive trajectory similarity between the initial motion trajectory and the reference motion trajectory, construct an optimization objective function based on the comprehensive trajectory similarity, and use an optimization algorithm to iteratively optimize the initial motion trajectory to obtain a corrected motion trajectory.
[0131] The swallowing restoration simulation module 14 is used to construct a swallowing restoration simulator based on the corrected motion trajectory, perform swallowing restoration simulation, and output the swallowing restoration results during the pharyngeal phase.
[0132] In one embodiment, the multidimensional image acquisition module 11 is further configured to:
[0133] The dynamic image sequence of the pharyngeal swallowing process is simultaneously acquired by multiple high-speed cameras arranged around the pharynx, wherein the multiple high-speed cameras include at least one frontal view camera, one side view camera and one oblique view camera.
[0134] Image feature parameters for each time point are extracted from the dynamic image sequence, wherein the image feature parameters include the coordinates of the pharyngeal contour points, the motion velocity value, and the motion direction angle.
[0135] In one embodiment, the image fusion modeling module 12 is further configured to:
[0136] Specifically, based on the multiple image feature parameters, image fusion weights are configured to perform multi-dimensional image fusion, including:
[0137] A standard swallowing sample dataset is constructed, and reference image feature parameters of the standard swallowing process are extracted from the standard swallowing sample dataset as reference feature parameters;
[0138] Calculate the similarity between each feature parameter of the image to be fused and the reference feature parameter to obtain multiple feature similarities;
[0139] Based on the multiple feature similarities, the image fusion weight for each image viewpoint is calculated, wherein the feature similarities are positively correlated with the image fusion weight;
[0140] Based on the image fusion weights, swallowing dynamic data from multiple image perspectives are weighted and fused to obtain fused multidimensional image data.
[0141] The steps for constructing the pharyngeal dynamic model include:
[0142] During the training phase, based on historical clinical swallowing image data, sample swallowing dynamic datasets were collected, and the pharyngeal structures in each swallowing dynamic dataset were three-dimensionally annotated to obtain a sample pharyngeal structure annotation set. Each sample pharyngeal structure annotation includes the spatial coordinate sequence and motion trajectory of key pharyngeal contour points.
[0143] A network architecture for constructing a dynamic model of the pharynx based on a three-dimensional convolutional neural network is provided, which includes a feature extraction module, a spatiotemporal fusion module, and a dynamic reconstruction module.
[0144] The pharyngeal dynamic model is trained under supervision using the sample swallowing dynamic dataset and the sample pharyngeal structure annotation set to obtain the trained pharyngeal dynamic model.
[0145] In the application phase, the fused multidimensional image data is input into the trained pharyngeal dynamic model, and the pharyngeal dynamic model outputs a complete swallowing dynamic result containing the pharyngeal geometry and motion sequence.
[0146] Specifically, the pharyngeal structures in each swallowing dynamic data point are 3D annotated to obtain a sample pharyngeal structure annotation set, including:
[0147] Collect historical clinical swallowing imaging data from clinical swallowing contrast examinations to construct an original dynamic swallowing dataset;
[0148] Based on a medical image annotation platform, key pharyngeal structures in each swallowing dynamic data are annotated, including the epiglottis, laryngeal vestibule, and pharyngeal constrictor muscle contours.
[0149] A three-dimensional spatial registration algorithm is used to fuse the annotation results and generate standardized three-dimensional annotations of the pharyngeal structure.
[0150] The consistency of the three-dimensional annotations of the pharyngeal structure is verified. When the distance error between annotations is less than a preset error threshold, it is included in the sample pharyngeal structure annotation set. Each sample pharyngeal structure annotation contains the spatial coordinate sequence of key contour points of the pharynx and the corresponding motion trajectory parameters.
[0151] Through iterative optimization, the pharyngeal structure annotation set of the sample is continuously expanded and optimized until it covers all typical swallowing stages and abnormal swallowing patterns.
[0152] In one embodiment, the trajectory optimization and correction module 13 is further configured to:
[0153] Specifically, based on the pharyngeal dynamic model, pharyngeal motion trajectory analysis is performed to obtain an initial motion trajectory, and the comprehensive trajectory similarity between the initial motion trajectory and a reference motion trajectory is calculated, including:
[0154] Motion path data of key points in the pharynx are extracted from the dynamic model of the pharynx to construct an initial set of motion trajectories;
[0155] A reference motion trajectory is obtained from a standard swallowing sample dataset, wherein the reference motion trajectory is obtained by statistical analysis of the motion trajectory of the standard swallowing process;
[0156] The trajectory similarity between the initial motion trajectory and the reference motion trajectory is calculated, including trajectory shape similarity and motion temporal similarity, wherein the trajectory shape similarity is calculated based on the dynamic time warping algorithm, and the motion temporal similarity is calculated based on the correlation coefficient method;
[0157] The trajectory shape similarity and motion temporal similarity are weighted and fused to obtain a comprehensive trajectory similarity.
[0158] Specifically, an optimization objective function is constructed based on the comprehensive trajectory similarity, and an optimization algorithm is used to iteratively optimize the initial motion trajectory to obtain a corrected motion trajectory, including:
[0159] To maximize the comprehensive trajectory similarity, an optimization objective function is constructed, wherein the optimization objective function considers both trajectory shape similarity and motion temporal similarity.
[0160] The initial trajectory is iteratively optimized using a particle swarm optimization algorithm, with particle swarm size parameters and iteration control parameters set.
[0161] In each iteration, new candidate trajectories are generated by updating particle positions and velocities, the corresponding comprehensive trajectory similarity is calculated, and the individual optimal solution and the global optimal solution are recorded.
[0162] Based on the comprehensive trajectory similarity, the particle search strategy is dynamically adjusted, and when the iteration termination condition is met, the globally optimal motion trajectory is output as the corrected motion trajectory.
[0163] In one embodiment, the swallowing simulation module 14 is further configured to:
[0164] Based on the corrected motion trajectory, a computational framework for a swallowing reconstruction simulator is constructed, wherein the computational framework includes a motion trajectory analysis unit, a pharyngeal dynamics calculation unit, and a swallowing process reconstruction unit;
[0165] Based on historical clinical swallowing process data, the simulation parameters of the swallowing simulation simulator are configured, wherein the simulation parameters include pharyngeal structure motion parameters and food passage path parameters;
[0166] The corrected motion trajectory is input into the swallowing simulation simulator. The motion sequence is analyzed by the motion trajectory analysis unit, the interaction between the pharyngeal geometric structures is calculated by the pharyngeal dynamics calculation unit, and a complete swallowing process simulation is generated by the swallowing process restoration unit.
[0167] The simulation results are validated for completeness. When the simulation results include the complete pharyngeal structure movement process and food passage path, the pharyngeal swallowing reconstruction results are output.
[0168] The integrity verification of the simulation results includes:
[0169] Check the integrity of the pharyngeal structure movement process in the simulation results and verify whether the movement trajectories of each key structure are continuous and complete;
[0170] Verify the continuity of the swallowing sequence and check whether the movement sequence of each pharyngeal structure conforms to the physiological laws of swallowing;
[0171] When the verification result does not meet the preset integrity standard, the comprehensive trajectory similarity between the simulation result and the reference motion trajectory is calculated.
[0172] Based on the comprehensive trajectory similarity, the adjustment direction and adjustment range of the simulation parameters are determined. The adjustment direction is determined according to the deviation between the trajectory shape similarity and motion timing similarity and the target value. The adjustment range is calculated proportionally based on the difference between the current comprehensive trajectory similarity and the target similarity.
[0173] Based on the adjusted simulation parameters, a new swallowing simulation is performed to generate new simulation results. The verification and adjustment process is repeated until the simulation results meet the preset integrity standards.
[0174] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0175] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
[0176] This specification and accompanying drawings are merely illustrative examples of the invention and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of the invention. Clearly, those skilled in the art can make various alterations and modifications to the invention without departing from its scope. Therefore, if such modifications and modifications fall within the scope of the invention and its equivalents, the invention is intended to include these modifications and modifications.
Claims
1. A method for pharyngeal phase swallowing reconstruction based on multi-dimensional video fusion, characterized in that, The method comprises: acquiring a multi-dimensional image sequence of a pharyngeal stage swallowing process by a multi-modal image acquisition device, obtaining swallowing dynamic data under multiple image perspectives, and extracting multiple image feature parameters; configuring image fusion weights according to the multiple image feature parameters, performing multi-dimensional image fusion, and constructing a pharyngeal dynamic model; based on the pharyngeal dynamic model, performing pharyngeal motion trajectory analysis to obtain an initial motion trajectory, calculating the comprehensive trajectory similarity between the initial motion trajectory and a reference motion trajectory, and constructing an optimization objective function based on the comprehensive trajectory similarity, using an optimization algorithm to iteratively optimize the initial motion trajectory to obtain a corrected motion trajectory; based on the corrected motion trajectory, constructing a swallowing restoration simulator, performing swallowing restoration simulation, and outputting a pharyngeal stage swallowing restoration result, comprising: based on the corrected motion trajectory, constructing a computational framework of the swallowing restoration simulator, wherein the computational framework comprises a motion trajectory analysis unit, a pharyngeal dynamics calculation unit, and a swallowing process restoration unit; configuring simulation parameters of the swallowing restoration simulator according to historical clinical swallowing process data, wherein the simulation parameters include pharyngeal structure motion parameters and food passage path parameters; inputting the corrected motion trajectory into the swallowing restoration simulator, analyzing the motion sequence through the motion trajectory analysis unit, calculating the interaction between the pharyngeal geometric structures through the pharyngeal dynamics calculation unit, and generating a complete swallowing process simulation through the swallowing process restoration unit; performing integrity verification on the simulation result, and outputting the pharyngeal stage swallowing restoration result when the simulation result contains a complete pharyngeal structure motion process and a food passage path, wherein the integrity verification on the simulation result comprises: checking the integrity of the pharyngeal structure motion process in the simulation result, and verifying whether the motion trajectories of each key structure are continuous and complete; verifying the continuity of the swallowing timing, and checking whether the motion timing of each pharyngeal structure conforms to the swallowing physiological law; when the verification result does not meet the preset integrity standard, calculating the comprehensive trajectory similarity between the simulation result and the reference motion trajectory; based on the comprehensive trajectory similarity, determining the adjustment direction and adjustment amplitude of the simulation parameters, wherein the adjustment direction is determined according to the deviation of the trajectory shape similarity and the motion timing similarity from the target value, and the adjustment amplitude is calculated in proportion to the difference between the current comprehensive trajectory similarity and the target similarity; re-performing the swallowing restoration simulation according to the adjusted simulation parameters to generate a new simulation result, and repeating the verification and adjustment process until the simulation result meets the preset integrity standard.
2. The method of claim 1, wherein the method is based on multi-dimensional image fusion. acquiring a multi-dimensional image sequence of a pharyngeal stage swallowing, obtaining swallowing dynamic data under multiple image perspectives, comprising: synchronously acquiring dynamic image sequences of the pharyngeal stage swallowing process by multiple high-speed cameras arranged around the pharynx, wherein the multiple high-speed cameras include at least one front perspective camera, one side perspective camera, and one oblique perspective camera; extracting image feature parameters at each time point from the dynamic image sequences, wherein the image feature parameters include pharyngeal contour point coordinates, motion velocity values, and motion direction angles. 3.The method of claim 1, wherein, According to the plurality of image feature parameters, an image fusion weight is configured, and multi-dimensional image fusion is performed, comprising: Constructing a standard swallowing sample data set, and extracting reference feature parameters of a standard swallowing process from the standard swallowing sample data set; Calculating the similarity of each to-be-fused image feature parameter and the reference feature parameter to obtain a plurality of feature similarities; According to the plurality of feature similarities, the image fusion weight of each image view is calculated, wherein the feature similarity is positively correlated with the image fusion weight; Based on the image fusion weight, swallowing dynamic data under a plurality of image views are weighted and fused to obtain fused multi-dimensional image data.
4. The method of claim 3, wherein the method is a multi-dimensional image fusion-based pharyngeal phase swallowing reconstruction method. The construction steps of the pharyngeal dynamic model include: In the training stage, sample swallowing dynamic data sets are collected according to historical clinical swallowing image data, and the pharyngeal structure in each swallowing dynamic data is three-dimensionally labeled to obtain a sample pharyngeal structure labeling set, wherein each sample pharyngeal structure labeling includes a sequence of spatial coordinates of pharyngeal key contour points and a motion trajectory; Based on a three-dimensional convolutional neural network, a network architecture of the pharyngeal dynamic model is constructed, and the network architecture includes a feature extraction module, a space-time fusion module and a dynamic reconstruction module; The sample swallowing dynamic data sets and the sample pharyngeal structure labeling set are used to supervise the training of the pharyngeal dynamic model to obtain a trained pharyngeal dynamic model; In the application stage, the fused multi-dimensional image data is input into the trained pharyngeal dynamic model, and a complete swallowing dynamic result containing pharyngeal geometric structure and motion sequence is output by the pharyngeal dynamic model.
5. The method of claim 4, wherein the method is based on multi-dimensional image fusion. The pharyngeal structure in each swallowing dynamic data is three-dimensionally labeled to obtain a sample pharyngeal structure labeling set, comprising: The historical clinical swallowing image data of clinical swallowing radiography is collected to construct an original swallowing dynamic data set; Based on a medical image labeling platform, the pharyngeal key structure in each swallowing dynamic data is labeled, wherein the pharyngeal key structure includes the epiglottis, the laryngeal vestibule and the pharyngeal constrictor contour; A three-dimensional space registration algorithm is used to fuse and process the labeling results to generate standardized pharyngeal structure three-dimensional labeling; The pharyngeal structure three-dimensional labeling is verified for consistency, and when the distance error between the labelings is less than a preset error threshold, it is included in the sample pharyngeal structure labeling set, wherein each sample pharyngeal structure labeling contains a sequence of spatial coordinates of pharyngeal key contour points and corresponding motion trajectory parameters; Through an iterative optimization process, the sample pharyngeal structure labeling set is continuously expanded and optimized until all typical swallowing stages and abnormal swallowing modes are covered.
6. The method of claim 1, wherein the method is a multi-dimensional image fusion-based pharyngeal phase swallowing reconstruction method. Based on the pharyngeal dynamic model, pharyngeal motion trajectory analysis is performed to obtain an initial motion trajectory, and the comprehensive trajectory similarity of the initial motion trajectory and a reference motion trajectory is calculated, comprising: Motion path data of pharyngeal key points are extracted from the pharyngeal dynamic model to construct an initial motion trajectory set; A reference motion trajectory is obtained from a standard swallowing sample data set, wherein the reference motion trajectory is obtained by statistically analyzing the motion trajectory of a standard swallowing process; calculating a trajectory similarity between the initial motion trajectory and a reference motion trajectory, including a trajectory shape similarity and a motion timing similarity, wherein the trajectory shape similarity is calculated based on a dynamic time warping algorithm, and the motion timing similarity is calculated based on a correlation coefficient method; performing weighted fusion on the trajectory shape similarity and the motion timing similarity to obtain a comprehensive trajectory similarity.
7. The method of claim 1, wherein the method is a multi-dimensional image fusion-based pharyngeal phase swallowing reconstruction method. constructing an optimization objective function based on the comprehensive trajectory similarity, and iteratively optimizing the initial motion trajectory using an optimization algorithm to obtain a corrected motion trajectory, including: maximizing the comprehensive trajectory similarity as an optimization objective to construct an optimization objective function, wherein the optimization objective function simultaneously considers the trajectory shape similarity and the motion timing similarity; iteratively optimizing the initial motion trajectory using a particle swarm optimization algorithm, and setting particle swarm size parameters and iteration control parameters; in each iteration process, generating a new candidate motion trajectory by updating particle positions and velocities, calculating the corresponding comprehensive trajectory similarity, and recording individual optimal solutions and global optimal solutions; based on the comprehensive trajectory similarity, dynamically adjusting the particle search strategy, and outputting the global optimal motion trajectory as the corrected motion trajectory when the iteration termination condition is met.
8. A pharyngeal phase swallowing reconstruction system based on multi-dimensional video fusion, characterized in that, A multi-dimensional image fusion-based pharyngeal phase swallowing reconstruction method for implementing any one of claims 1-7, the system comprising: a multi-dimensional image acquisition module for acquiring a multi-dimensional image sequence of a pharyngeal phase swallowing process through a multi-modal image acquisition device, obtaining swallowing dynamic data under multiple image perspectives, and extracting multiple image feature parameters; an image fusion modeling module for configuring image fusion weights according to the multiple image feature parameters, performing multi-dimensional image fusion, and constructing a pharyngeal dynamic model; a trajectory optimization and correction module for performing pharyngeal motion trajectory analysis based on the pharyngeal dynamic model, obtaining an initial motion trajectory, calculating a comprehensive trajectory similarity between the initial motion trajectory and a reference motion trajectory, and constructing an optimization objective function based on the comprehensive trajectory similarity, iteratively optimizing the initial motion trajectory using an optimization algorithm to obtain a corrected motion trajectory; a swallowing reconstruction simulation module for constructing a swallowing reconstruction simulator based on the corrected motion trajectory, performing swallowing reconstruction simulation, and outputting pharyngeal phase swallowing reconstruction results, including: constructing a computational framework of the swallowing reconstruction simulator based on the corrected motion trajectory, wherein the computational framework includes a motion trajectory analysis unit, a pharyngeal dynamics calculation unit, and a swallowing process reconstruction unit; configuring simulation parameters of the swallowing reconstruction simulator based on historical clinical swallowing process data, wherein the simulation parameters include pharyngeal structure motion parameters and food passage path parameters; inputting the corrected motion trajectory into the swallowing reconstruction simulator, analyzing the motion sequence through the motion trajectory analysis unit, calculating the interaction between pharyngeal geometric structures through the pharyngeal dynamics calculation unit, and generating a complete swallowing process simulation through the swallowing process reconstruction unit; The integrity of the simulation result is verified, and when the simulation result contains complete pharyngeal structure movement process and food passing path, the pharyngeal swallowing reconstruction result is output, wherein the integrity of the simulation result is verified, including: checking the integrity of the pharyngeal structure movement process in the simulation result, verifying whether the movement trajectory of each key structure is continuous and complete; verifying the continuity of the swallowing timing, checking whether the movement timing of each pharyngeal structure conforms to the physiological law of swallowing; when the verification result does not satisfy the preset integrity standard, calculating the comprehensive trajectory similarity between the simulation result and the reference movement trajectory; based on the comprehensive trajectory similarity, determining the adjustment direction and adjustment amplitude of the simulation parameters, wherein the adjustment direction is determined according to the deviation of the trajectory shape similarity and the movement timing similarity from the target value, and the adjustment amplitude is calculated in proportion according to the difference between the current comprehensive trajectory similarity and the target similarity; according to the adjusted simulation parameters, the swallowing reconstruction simulation is performed again to generate a new simulation result, and the verification and adjustment process is repeated until the simulation result satisfies the preset integrity standard.
Citation Information
Patent Citations
Virtual simulation system for preventing and controlling dysphagia of old people in pension institution
CN115797121A
Swallowing dysfunction improvement evaluation system based on head hooking and throat shrinking swallowing actions
CN118762845A