Image-enhanced head and neck radiotherapy patient dysphagia assessment system

By employing a sensitivity-weighted sampling strategy and multi-scale feature extraction, combined with adaptive modulation parameter generation and feature encoding units, the accuracy and edge blurring issues of the assessment system for difficulty in opening the mouth in head and neck radiotherapy patients were resolved, achieving more accurate assessment and segmentation results.

CN120318226BActive Publication Date: 2025-11-18THE FIFTH MEDICAL CENT OF CHINESE PLA GENERAL HOSPITAL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510780744.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-11-18
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

Existing systems for assessing the degree of difficulty in opening the mouth in patients undergoing head and neck radiotherapy suffer from poor accuracy due to uneven distribution of foreground and background samples and a lack of adaptive correction mechanisms for different facial regions. Post-radiotherapy tissue degradation leads to blurred edge information, affecting the clarity of segmentation results.

Method used

A sensitivity-weighted sampling strategy, multi-scale degradation feature extraction, and adaptive modulation parameter generation are adopted. By combining feature extraction coding units and decoding upsampling units, image segmentation and evaluation are optimized through region penalty loss and boundary distance weighting.

Benefits of technology

It improves the accuracy of assessing the degree of difficulty in opening the mouth and the precision of segmentation results, adapts to the complex morphological changes in the oral cavity after radiotherapy, and enhances the ability to handle edges.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318226B_ABST
    Figure CN120318226B_ABST
Patent Text Reader

Abstract

The application discloses a head and neck radiotherapy patient mouth opening difficulty degree evaluation system based on image enhancement, which comprises an image acquisition module, a mouth opening image enhancement module, a mouth opening image segmentation module and a mouth opening difficulty degree evaluation module. The application belongs to the field of image processing, and particularly relates to a head and neck radiotherapy patient mouth opening difficulty degree evaluation system based on image enhancement. The scheme adopts a sampling strategy based on sensitivity weight, alleviates the problem of subsequent evaluation influenced by class imbalance, introduces a local attenuation compensation term, and corrects based on a multi-scale degradation feature extraction unit and an adaptive modulation parameter generation unit. Based on a constructed regional penalty loss and a boundary distance weighted boundary numerical supervision loss, global and local information are taken into account, and the features of edge blur and local morphology difference commonly existing in the images of patients after radiotherapy are adapted, so that the subsequent head and neck radiotherapy patient mouth opening difficulty degree evaluation effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing, specifically to an image enhancement-based system for assessing the degree of difficulty in opening the mouth in patients undergoing head and neck radiotherapy. Background Technology

[0002] A system for assessing the degree of mouth opening difficulty in patients undergoing head and neck radiotherapy generally refers to a comprehensive system that utilizes medical imaging technology to quantitatively assess the patient's mouth opening limitation by acquiring images of the patient's open mouth, performing fine image segmentation, and extracting relevant morphological features. However, such systems typically suffer from uneven distribution of foreground and background samples, lack of adaptive correction mechanisms for different facial regions, and inability to adjust compensation levels based on local conditions, leading to poor accuracy in subsequent mouth opening difficulty assessments. Furthermore, these systems often lack dedicated processing mechanisms for edge regions. After radiotherapy, due to tissue degradation and significant changes in local morphology, edge information is easily blurred, resulting in unclear segmentation edges and affecting the overall assessment effectiveness. Summary of the Invention

[0003] To address the above issues and overcome the shortcomings of existing technologies, this invention provides an image-enhanced assessment system for mouth-opening difficulty in head and neck radiotherapy patients. Common head and neck radiotherapy patients' mouth-opening difficulty assessment systems suffer from uneven distribution of foreground and background samples, lack of adaptive correction mechanisms for different facial regions, and inability to adjust compensation intensity based on local conditions, leading to poor accuracy in subsequent mouth-opening difficulty assessments. This solution employs a sensitivity-weighted sampling strategy to alleviate the impact of class imbalance on subsequent assessments. For the complex local degradation phenomena in radiotherapy patients, a local attenuation compensation term is introduced to simulate additional local attenuation, and correction is performed based on a multi-scale degradation feature extraction unit and an adaptive modulation parameter generation unit. This corrects global illumination reflection and noise while restoring tooth edges and soft tissue details, thereby improving the accuracy of subsequent mouth-opening difficulty assessments. Furthermore, common head and neck radiotherapy patients' mouth-opening difficulty assessment systems lack specialized processing mechanisms for edge regions, resulting in poor accuracy after radiotherapy. Due to tissue degradation and significant local morphological changes, edge information is easily blurred, leading to unclear segmentation results and affecting the overall evaluation effect. This scheme addresses this issue by incorporating learnable positional encoding into the feature extraction encoding unit to capture spatial distribution information, thereby more accurately locating subtle local regions in oral images. Multi-scale feature extraction ensures that it can capture details while maintaining global contours even when dealing with blurred edges and local degradation, adapting to the complex morphological changes in the oral region after radiotherapy. Skip connections in the decoding upsampling unit effectively preserve low-level details, which is particularly crucial for addressing the loss of local details caused by radiotherapy, resulting in more accurate segmentation results in terms of both detail and overall morphology. Finally, based on the construction of a region penalty loss and the introduction of boundary distance-weighted boundary numerical supervision loss, both global and local information are considered, adapting to the common features of blurred edges and inconsistent local morphology in post-radiotherapy patient images. This improves the evaluation effect of mouth-opening difficulty in subsequent head and neck radiotherapy patients.

[0004] The technical solution adopted by the present invention is as follows: The image enhancement-based assessment system for the degree of difficulty in opening the mouth of patients undergoing head and neck radiotherapy provided by the present invention includes an image acquisition module, an opening image enhancement module, an opening image segmentation module, and an opening difficulty assessment module;

[0005] The image acquisition module acquires oral cavity opening images of historical head and neck radiotherapy patients and filters images using unbalanced ratios and sensitivity weights.

[0006] The mouth opening image enhancement module uses degradation feature extraction and adaptive modulation parameter generation to correct the acquired oral cavity opening image, thereby achieving mouth opening image enhancement.

[0007] The mouth opening image segmentation module achieves fine pixel-level segmentation of the oral cavity opening region after radiotherapy by using multi-scale feature extraction, position encoding, and Transformer global modeling, combined with stepwise upsampling and dual loss optimization.

[0008] The mouth opening difficulty assessment module uses morphological processing to extract the oral cavity opening boundary and contour, and uses a pre-trained regression model to generate a quantitative score to assess the mouth opening difficulty.

[0009] Furthermore, the image acquisition module annotates the oral opening region in historical head and neck radiotherapy patients' oral opening images; defines an imbalance ratio and introduces a sampling method based on the imbalance ratio; treats the oral opening region as the foreground and the remaining facial regions as the background; and introduces sensitivity weights. The final imbalance ratio IR is expressed as: ; ; Where x and y are pixel coordinate indices; It is the weighted area of ​​the oral cavity opening region; It is the area of ​​the oral cavity opening region; It is the area of ​​the remaining facial region; It is a regulatory factor; and These are the gradient values ​​in the horizontal and vertical directions at pixel (x,y), respectively; images of oral openings of historical head and neck radiotherapy patients are sorted in ascending order according to their imbalance ratio, and N images are selected as the final image set.

[0010] Furthermore, the mouth-opening image enhancement module specifically includes the following:

[0011] Comprehensive simulation unit for oral cavity opening image degradation; introduction of local attenuation compensation term. The collected images of the patient's oral cavity opening are simulated as follows: ; ;in, These are images of the patient's oral cavity opening. It is an image of an ideal oral cavity structure; It is additional noise; L represents the visibility inside the oral cavity; L represents the intensity of light reflected from the dental mirror. It is a convolution operation; It is the degenerate feature map extracted at (x,y);

[0012] Oral cavity opening image rearrangement unit; rearranged oral cavity opening image Represented as: ; ; ; and These are modulation parameters;

[0013] The degradation feature extraction unit takes the patient's oral cavity opening image, along with gradient and texture information, as input. The unit consists of T convolutional layers, each followed by a ReLU activation function and a normalization layer to form a series of local feature representations. The overall representation is as follows: ;in, It is feature fusion; It is a feature concatenation operation; , and Features are extracted using 3×3, 5×5, and 7×7 convolutional kernels;

[0014] Adaptive modulation parameter generation unit; generates modulation parameters in the channel and spatial dimensions using Conv1DNet and embedding blocks. and Three parallel 1D convolutional branches are used on the embedded features, with kernel sizes of 3, 5, and 7, respectively, to capture contextual information at different scales. The outputs at different scales are fused to obtain the final modulation parameters, expressed as follows: ; ;in, , and It is the local modulation parameter extracted at position (x,y) by 1D convolution branches with different kernel sizes. ; , and It is the local modulation parameter extracted at position (x,y) by 1D convolution branches with different kernel sizes. ; It uses fully connected layers to fuse multi-scale features;

[0015] Reconstruction error minimization unit; defines a reconstruction loss function, which optimizes modulation parameters by minimizing the error between the reconstructed image and the ideal image, expressed as: ;in, The modulation parameters that minimize the loss function L are selected; the Conv1DNet parameters are adjusted based on the loss gradient, and the mouth-opening image enhancement module is trained based on the original image set.

[0016] Furthermore, the mouth-opening image segmentation module performs pixel-level segmentation on the image output by the mouth-opening image enhancement module; the mouth-opening image segmentation module includes a feature extraction encoding unit B(·), a decoding upsampling unit D(·), and a dual loss optimization unit; the image processed by the mouth-opening image enhancement module is then segmented at the pixel level. As input to the segmentation network; the output is a pixel-level probability map P. F(x), where the value of each pixel represents the probability that it belongs to the oral cavity opening region, and the final segmentation result is obtained; the entire oral cavity image segmentation module is represented as: Specifically, it includes the following:

[0017] Feature extraction coding unit: extracts low-level features from the input image to obtain a feature map. : is represented as: Add a positional encoding P(x,y), which is used as a learnable parameter and initialized with random values. During training, it is automatically adjusted through backpropagation to obtain the embedded features. , is represented as: To address the scale difference between global and local information in oral cavity opening images, three parallel 1D convolutional branches are employed. , and , is represented as: ; ; Each branch focuses on a different receptive field, capturing details of the labial margins and interdental spaces, as well as features of the overall oral cavity contour. The outputs of the multi-scale branches are fused to generate locally modulated features. , is represented as: Adding Transformer encoding captures long-range dependencies, resulting in a high-dimensional feature map. , is represented as: ; and will As the final output of the high-dimensional feature map; among which... This is the initial convolution operation; It is a low-level feature map; It is a one-dimensional convolution operation using a kernel size of k, where k=3, 5, or 7 is the kernel size; It is the ReLU activation function; , and It is a fusion weight; It is Transformer encoding, which uses a self-attention mechanism to globally reorganize the fused features, highlighting key structures and edge information;

[0018] Decoding upsampling unit; the decoder converts the high-dimensional feature map into a pixel-level prediction map of the same size as the input image; it uses stepwise upsampling to restore spatial resolution, while combining skip connections to preserve low-level details, represented as: Finally, pixel-level prediction probabilities are obtained through a single convolutional layer and activation function. , is represented as: ;in, It is a high-dimensional feature at layer l in the encoding stage; It is an upsampling operation; These are low-level features from the corresponding layer of the encoder; It is the decoded upsampled output of layer l; It is the Sigmoid function; This is the last convolutional operation, which converts the upsampled features into predicted values. This is the total decoded upsampled output;

[0019] Dual loss optimization unit; including:

[0020] Define region penalty; design region penalty loss; define pixel-level labels Y∈{0,1}, where 0 represents background and 1 represents the oral cavity opening region; the prediction probability is P. F (x), and power-modulate the predicted probability; region penalty loss Represented as: ;in, and It is a weighting factor; It is the modulation index;

[0021] Define boundary numerical supervision loss Introducing boundary distance weighting The boundary numerical supervision loss is expressed as: ; ;in, It is the probability of matching between the predicted boundary numerical labels and the true boundary labels; It is an image region; and It refers to adjusting parameters; It is the distance from the pixel to the actual boundary;

[0022] The final loss function L is expressed as: ;in, It is the weighting factor for the boundary loss.

[0023] Furthermore, the mouth opening difficulty assessment module acquires real-time images of the mouth opening of patients undergoing head and neck radiotherapy. After processing by the mouth opening image enhancement module and the mouth opening image segmentation module, the mouth opening boundary and regional contour are extracted using morphological processing. The morphological processing indicators are then input into a pre-trained regression model to generate a quantitative mouth opening difficulty score for assessing the degree of mouth opening difficulty.

[0024] The beneficial effects achieved by the present invention using the above solution are as follows:

[0025] (1) In view of the problem that the general head and neck radiotherapy patients’ mouth opening difficulty assessment system has uneven distribution of foreground and background samples, lacks an adaptive correction mechanism for different facial regions, and cannot adjust the compensation intensity according to local conditions, resulting in poor accuracy of subsequent mouth opening difficulty assessment, this solution adopts a sampling strategy based on sensitivity weight to alleviate the problem of class imbalance affecting subsequent assessment. In view of the complex local degradation phenomenon of radiotherapy patients, a local attenuation compensation term is introduced to simulate local additional attenuation, and correction is performed based on multi-scale degradation feature extraction unit and adaptive modulation parameter generation unit. This corrects global illumination reflection and noise, and restores tooth edge and soft tissue details, thereby improving the accuracy of subsequent mouth opening difficulty assessment.

[0026] (2) In view of the lack of a special processing mechanism for edge regions in the general head and neck radiotherapy patients' mouth opening difficulty assessment system, the edge information is easily blurred due to tissue degradation and obvious local morphological changes after radiotherapy, resulting in unclear edge segments and affecting the overall assessment effect. This scheme adds learnable positional coding to the feature extraction coding unit to capture spatial distribution information, thereby more accurately locating local fine regions in the oral cavity image; based on multi-scale feature extraction, it ensures that it can capture details and maintain the global contour when dealing with edge blurring and local degradation, adapting to the complex situation of oral cavity morphological changes after radiotherapy; by using the skip connection of the decoding upsampling unit, low-level details are effectively preserved, which is particularly important for the problem of local detail loss caused by radiotherapy, thus making the final segmentation result more accurate in terms of details and overall morphology; finally, based on the construction of region penalty loss and the introduction of boundary distance weighted boundary numerical supervision loss, both global and local information are taken into account, adapting to the edge blurring and local morphological differences that are common in patient images after radiotherapy; thereby improving the assessment effect of mouth opening difficulty in subsequent head and neck radiotherapy patients. Attached Figure Description

[0027] Figure 1 A schematic diagram of the process for the image-enhanced assessment system for mouth-opening difficulty in head and neck radiotherapy patients provided by the present invention;

[0028] Figure 2 This is a flowchart illustrating the open-mouth image segmentation module.

[0029] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof. Detailed Implementation

[0030] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0031] In the description of this invention, it should be understood that the terms "upper", "lower", "front", "rear", "left", "right", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.

[0032] Example 1, see Figure 1 The present invention provides an image enhancement-based system for assessing the degree of difficulty in opening the mouth in patients undergoing head and neck radiotherapy, comprising an image acquisition module, an opening image enhancement module, an opening image segmentation module, and an opening difficulty assessment module.

[0033] The image acquisition module acquires oral cavity images of patients who have undergone historical head and neck radiotherapy, and filters the images using an imbalance ratio and sensitivity weight; and sends the data to the oral cavity image enhancement module.

[0034] The mouth opening image enhancement module uses degradation feature extraction and adaptive modulation parameter generation to correct the acquired oral cavity opening image, thereby achieving mouth opening image enhancement; and sends the data to the mouth opening image segmentation module.

[0035] The mouth opening image segmentation module achieves fine pixel-level segmentation of the oral cavity opening region after radiotherapy by using multi-scale feature extraction, position encoding, and Transformer global modeling, combined with stepwise upsampling and dual loss optimization; and sends the data to the mouth opening difficulty assessment module.

[0036] The mouth opening difficulty assessment module uses morphological processing to extract the oral cavity opening boundary and contour, and uses a pre-trained regression model to generate a quantitative score to assess the mouth opening difficulty.

[0037] Example 2, see Figure 1This embodiment is based on the above embodiment. The image acquisition module annotates the oral opening region in historical head and neck radiotherapy patients' oral opening images; defines an imbalance ratio, since the oral opening region usually occupies a small proportion in the overall facial image, direct training may face serious class imbalance problems; introduces a sampling method based on the imbalance ratio; treats the oral opening region as the foreground and the remaining facial regions as the background; and introduces sensitivity weights. By enlarging the proportion of key structures, even if the overall area does not change significantly, changes in key areas can be fully reflected, thus more accurately measuring the degree of mouth opening restriction in patients. The resulting imbalance ratio (IR) is expressed as: ; ; Where x and y are pixel coordinate indices; It is the weighted area of ​​the oral cavity opening region; It is the area of ​​the oral cavity opening region; It is the area of ​​the remaining facial region; It is a regulatory factor; and These are the gradient values ​​in the horizontal and vertical directions at pixel (x, y), respectively; the area of ​​the oral cavity opening region when the patient has difficulty opening their mouth. It will become relatively smaller, while the area of ​​the remaining facial region will be smaller. The imbalance ratio remains basically unchanged, which reduces the overall imbalance ratio. Therefore, a low imbalance ratio can intuitively reflect the reduced mouth opening of the patient and accurately reflect the actual situation of the patient's limited mouth opening. The oral opening images of historical head and neck radiotherapy patients are arranged in ascending order according to the imbalance ratio, and N images are selected as the final image set.

[0038] Example 3, see Figure 1 This embodiment is based on the above embodiment, and the mouth opening image enhancement module specifically includes the following:

[0039] A comprehensive simulation unit for oral opening image degradation is included; it takes into account the degradation of oral opening images caused by factors such as soft tissue fibrosis, mucosal damage, and muscle stiffness in radiotherapy patients; and a local attenuation compensation term is introduced. This is used to describe the additional attenuation effect caused by fibrosis or tissue stiffness in a local area. The acquired images of the patient's oral cavity opening are simulated as follows: ; ;in, These are images of the patient's oral cavity opening. It is an ideal image of the oral cavity structure that can accurately measure the intraoral opening and the spacing between the teeth; This includes additional noise, including ambient light interference and sensor noise; L represents the visibility inside the oral cavity; L represents the intensity of light reflected from the dental mirror. It is a convolution operation; It is the degenerate feature map extracted at (x,y);

[0040] The oral cavity opening image rearrangement unit simplifies the complex degradation process into a local multiply-accumulate operation, using local visibility to correct the acquired image and compensate for the effects of illumination reflection and noise; enabling the restoration of key tooth edges and soft tissue details in the image; rearranged oral cavity opening images. Represented as: ; ; ; and These are modulation parameters, representing the combined effects of local imaging conditions and occlusion reflection noise, respectively;

[0041] The degradation feature extraction unit takes the patient's oral cavity opening image, along with gradient and texture information, as input. The unit consists of multiple convolutional layers, each followed by a ReLU activation function and a normalization layer, forming a series of local feature representations. The overall representation is as follows: ;in, It is feature fusion; It is a feature concatenation operation; , and Features are extracted using 3×3, 5×5, and 7×7 convolutional kernels;

[0042] An adaptive modulation parameter generation unit is used to adaptively adjust to different regions such as the alveolar region, the vicinity of the joint, and the soft tissue boundary, thereby correcting local feature distortion caused by insufficient mouth opening. Modulation parameters are generated in both channel and spatial dimensions using Conv1DNet and embedding blocks. and To address the scale differences between the overall width of the oral cavity opening and the local details of the lip contour and interdental spaces in the image, three parallel 1D convolutional branches with kernel sizes of 3, 5, and 7 are used on the embedded features to capture contextual information at different scales. The outputs from different scales are then fused to obtain the final modulation parameters, expressed as follows: ; ;in, , and It is the local modulation parameter extracted at position (x,y) by 1D convolution branches with different kernel sizes. ; , and It is the local modulation parameter extracted at position (x,y) by 1D convolution branches with different kernel sizes. ; It uses fully connected layers to fuse multi-scale features;

[0043] Reconstruction error minimization unit; defines a reconstruction loss function, which optimizes modulation parameters by minimizing the error between the reconstructed image and the ideal image, expressed as: ;in, The modulation parameters that minimize the loss function L are selected; the Conv1DNet parameters are adjusted based on the loss gradient, and the mouth-opening image enhancement module is trained based on the original image set.

[0044] By performing the above operations, this solution addresses the problem that general head and neck radiotherapy patients' mouth-opening difficulty assessment systems suffer from uneven distribution of foreground and background samples, lack of adaptive correction mechanisms for different facial regions, and inability to adjust compensation intensity according to local conditions, leading to poor accuracy in subsequent mouth-opening difficulty assessments. This solution employs a sensitivity-weighted sampling strategy to alleviate the impact of class imbalance on subsequent assessments. For the complex local degradation phenomena in radiotherapy patients, a local attenuation compensation term is introduced to simulate additional local attenuation. Correction is then performed based on a multi-scale degradation feature extraction unit and an adaptive modulation parameter generation unit, correcting both global illumination reflection and noise while restoring tooth margins and soft tissue details; thereby improving the accuracy of subsequent mouth-opening difficulty assessments.

[0045] Example 4, see Figure 1 and Figure 2 This embodiment is based on the above embodiment. The mouth opening image segmentation module performs pixel-level segmentation on the image output by the mouth opening image enhancement module, aiming to locate the oral opening region of patients undergoing head and neck radiotherapy. The mouth opening image segmentation module includes a feature extraction encoding unit B(·), a decoding upsampling unit D(·), and a dual loss optimization unit. The image processed by the mouth opening image enhancement module is... As input to the segmentation network; the output is a pixel-level probability map P. F (x), where the value of each pixel represents the probability that it belongs to the oral cavity opening region, and the final segmentation result is obtained; the entire oral cavity image segmentation module is represented as: Specifically, it includes the following:

[0046] Feature extraction coding unit: extracts low-level features from the input image to obtain a feature map. : is represented as: To enable the mouth opening image segmentation module to perceive information about the spatial positions of each part of the oral cavity opening image, a positional encoding P(x,y) is added. This positional encoding is used as a learnable parameter, initialized with random values, and automatically adjusted through backpropagation during training to obtain embedded features. , is represented as: To address the scale difference between global and local information in oral cavity opening images, three parallel 1D convolutional branches are employed. , and , is represented as: ; ; Each branch focuses on a different receptive field, capturing details of the labial margins and interdental spaces, as well as features of the overall oral cavity contour. The outputs of the multi-scale branches are fused to generate locally modulated features. , is represented as: Adding Transformer encoding captures long-range dependencies, resulting in a high-dimensional feature map. , is represented as: ; and will As the final output of the high-dimensional feature map; among which... This is the initial convolution operation; It is a low-level feature map; It is a one-dimensional convolution operation using a kernel size of k, where k=3, 5, or 7 is the kernel size; It is the ReLU activation function; , and It is a fusion weight; It is Transformer encoding, which uses a self-attention mechanism to globally reorganize the fused features, highlighting key structures and edge information;

[0047] Decoding upsampling unit; the decoder converts the high-dimensional feature map into a pixel-level prediction map of the same size as the input image; it uses stepwise upsampling to restore spatial resolution, while combining skip connections to preserve low-level details, represented as: Finally, pixel-level prediction probabilities are obtained through a single convolutional layer and activation function. , is represented as: ;in, It is a high-dimensional feature at layer l in the encoding stage; It is an upsampling operation; These are low-level features from the corresponding layer of the encoder; It is the decoded upsampled output of layer l; It is the Sigmoid function; This is the last convolutional operation, which converts the upsampled features into predicted values. This is the total decoded upsampled output;

[0048] Dual-loss optimization unit: In order to improve segmentation accuracy, especially in dealing with edge blurring and local morphological changes in oral opening images after radiotherapy, two loss functions were designed.

[0049] Define a region penalty; to accurately segment the oral cavity opening region at the pixel level, a region penalty loss is designed to impose a higher penalty on missegmented regions; define pixel-level labels Y∈{0,1}, where 0 represents the background and 1 represents the oral cavity opening region; the prediction probability is P. F (x), and power-law modulation is applied to the predicted probability to reduce the loss contribution of easily classifiable pixels; region penalty loss Represented as: ;in, and It is a weighting factor used to balance the penalty for positive and negative samples; It is the modulation index;

[0050] Define boundary numerical supervision loss It is used to optimize the complex boundaries of the oral opening area, and is more sensitive to the oral opening area under the influence of radiotherapy; and introduces boundary distance weighting. Pixels closer to the boundary are assigned higher weights, making them more sensitive to complex boundaries in the oral cavity after radiotherapy; the boundary numerical supervision loss is represented as: ; ;in, It is the probability of matching between the predicted boundary numerical labels and the true boundary labels; It is an image region; and It refers to adjusting parameters; It is the distance from the pixel to the actual boundary;

[0051] The final loss function L is expressed as: ;in, It is the weighting factor for the boundary loss.

[0052] By performing the above operations, this solution addresses the problem that general head and neck radiotherapy patients' mouth-opening difficulty assessment systems lack a dedicated mechanism for processing edge regions. After radiotherapy, due to tissue degradation and significant local morphological changes, edge information is easily blurred, leading to unclear segmentation results and affecting the overall assessment effect. This solution incorporates learnable positional encoding into the feature extraction encoding unit to capture spatial distribution information, thereby more accurately locating subtle local regions in oral images. Multi-scale feature extraction ensures that while handling edge blurring and local degradation, it can capture details while maintaining the global contour, adapting to the complex morphological changes in the oral region after radiotherapy. Skip connections in the decoding upsampling unit effectively preserve low-level details, which is particularly crucial for addressing the loss of local details caused by radiotherapy, resulting in more accurate segmentation results in terms of both detail and overall morphology. Finally, by constructing a region penalty loss and introducing a boundary distance-weighted boundary numerical supervision loss, it takes into account both global and local information, adapting to the common edge blurring and inconsistent local morphology features in post-radiotherapy patient images, thus improving the assessment effect of mouth-opening difficulty in subsequent head and neck radiotherapy patients.

[0053] Example 6, see Figure 1 This embodiment is based on the above embodiment. The mouth opening difficulty assessment module acquires real-time images of the oral opening of patients undergoing head and neck radiotherapy. After processing by the mouth opening image enhancement module and the mouth opening image segmentation module, the oral opening boundary and regional contour are extracted using morphological processing. The oral opening width is obtained by calculating the farthest distance between the two boundaries of the segmented region. The pixels of the segmented region are counted and converted into actual area to obtain the oral opening area. The edge smoothness and curvature are calculated. The above indicators are input into a pre-trained regression model to generate a quantitative mouth opening difficulty score for assessing the degree of mouth opening difficulty. If the assessment result is mouth opening difficulty, an early warning is issued to the relevant personnel.

[0054] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0055] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention.

[0056] The present invention and its embodiments have been described above. This description is not restrictive, and the accompanying drawings are only one embodiment of the present invention; the actual structure is not limited thereto. In conclusion, if those skilled in the art are inspired by this description and design similar structures and embodiments without departing from the spirit of the invention, such designs should fall within the protection scope of the present invention.

Claims

1. An image-enhanced assessment system for mouth-opening difficulty in head and neck radiotherapy patients, characterized in that: The system includes an image acquisition module, an open-mouth image enhancement module, an open-mouth image segmentation module, and an open-mouth difficulty assessment module; The image acquisition module acquires oral cavity opening images of historical head and neck radiotherapy patients and filters images using unbalanced ratios and sensitivity weights. The mouth opening image enhancement module uses degradation feature extraction and adaptive modulation parameter generation to correct the acquired oral cavity opening image, thereby achieving mouth opening image enhancement. The mouth opening image segmentation module achieves fine pixel-level segmentation of the oral cavity opening region after radiotherapy by using multi-scale feature extraction, position encoding, and Transformer global modeling, combined with stepwise upsampling and dual loss optimization. The mouth opening difficulty assessment module uses morphological processing to extract the oral cavity opening boundary and contour, and uses a pre-trained regression model to generate a quantitative score to assess the mouth opening difficulty. The image acquisition module marks the oral opening area in historical images of head and neck radiotherapy patients. An imbalance ratio is defined, and a sampling method based on the imbalance ratio is introduced; the oral cavity opening area is considered as the foreground, and the remaining facial areas are considered as the background; sensitivity weights are introduced. The final imbalance ratio IR is expressed as: ; ; Where x and y are pixel coordinate indices; It is the weighted area of ​​the oral cavity opening region; It is the area of ​​the oral cavity opening region; It is the area of ​​the remaining facial region; It is a regulatory factor; and These are the gradient values ​​in the horizontal and vertical directions at pixel (x,y), respectively; images of oral openings of historical head and neck radiotherapy patients are sorted in ascending order according to their imbalance ratio, and N images are selected as the final image set.

2. The image-enhanced assessment system for mouth-opening difficulty in head and neck radiotherapy patients according to claim 1, characterized in that: The mouth-opening image enhancement module specifically includes the following: Comprehensive simulation unit for oral cavity opening image degradation; introduction of local attenuation compensation term. The collected images of the patient's oral cavity opening are simulated as follows: ; ;in, These are images of the patient's oral cavity opening. It is an image of an ideal oral cavity structure; It is additional noise; L represents the visibility inside the oral cavity; L represents the intensity of light reflected from the dental mirror. It is a convolution operation; It is the degenerate feature map extracted at (x,y); Oral cavity opening image rearrangement unit; rearranged oral cavity opening image Represented as: ; ; ; and These are modulation parameters; The degradation feature extraction unit takes the patient's oral cavity opening image, along with gradient and texture information, as input. The unit consists of T convolutional layers, each followed by a ReLU activation function and a normalization layer to form a series of local feature representations. The overall representation is as follows: ;in, It is feature fusion; It is a feature concatenation operation; , and Features are extracted using 3×3, 5×5, and 7×7 convolutional kernels; Adaptive modulation parameter generation unit; generates modulation parameters in the channel and spatial dimensions using Conv1DNet and embedding blocks. and Three parallel 1D convolutional branches are used on the embedded features, with kernel sizes of 3, 5, and 7, respectively, to capture contextual information at different scales. The outputs at different scales are fused to obtain the final modulation parameters, expressed as follows: ; ;in, , and It is the local modulation parameter extracted at position (x,y) by 1D convolution branches with different kernel sizes. ; , and It is the local modulation parameter extracted at position (x,y) by 1D convolution branches with different kernel sizes. ; It uses fully connected layers to fuse multi-scale features; Reconstruction error minimization unit; defines a reconstruction loss function, which optimizes modulation parameters by minimizing the error between the reconstructed image and the ideal image, expressed as: ;in, The modulation parameters that minimize the loss function L are selected; the Conv1DNet parameters are adjusted based on the loss gradient, and the mouth-opening image enhancement module is trained based on the original image set.

3. The image-enhanced assessment system for mouth-opening difficulty in head and neck radiotherapy patients according to claim 2, characterized in that: The mouth-opening image segmentation module performs pixel-level segmentation on the image output by the mouth-opening image enhancement module; the mouth-opening image segmentation module includes a feature extraction encoding unit B(·), a decoding upsampling unit D(·), and a dual loss optimization unit; the image processed by the mouth-opening image enhancement module... As input to the segmentation network; the output is a pixel-level probability map P. F (x), where the value of each pixel represents the probability that it belongs to the oral cavity opening region, and the final segmentation result is obtained; the entire oral cavity image segmentation module is represented as: ;specific Includes the following: Feature extraction coding unit; extracts low-level features from the input image to obtain a feature map, represented as: ; A positional encoding P(x,y) is added, which is used as a learnable parameter and initialized with random values. During training, it is automatically adjusted through backpropagation to obtain the embedded features. , is represented as: To address the scale difference between global and local information in oral cavity opening images, three parallel 1D convolutional branches are employed. , and , is represented as: ; ; Each branch focuses on a different receptive field, capturing details of the labial margins and interdental spaces, as well as features of the overall oral cavity contour. The outputs of the multi-scale branches are fused to generate locally modulated features. , is represented as: ; By incorporating Transformer encoding to capture long-range dependencies, a high-dimensional feature map is obtained. , is represented as: ; and will As the final output of the high-dimensional feature map; among which... This is the initial convolution operation; It is a low-level feature map; It is a one-dimensional convolution operation using a kernel size of k, where k=3, 5, or 7 is the kernel size; It is the ReLU activation function; , and It is a fusion weight; It is Transformer encoding, which uses a self-attention mechanism to globally reorganize the fused features, highlighting key structures and edge information; Decoding upsampling unit; the decoder converts the high-dimensional feature map into a pixel-level prediction map of the same size as the input image; it uses stepwise upsampling to restore spatial resolution, while combining skip connections to preserve low-level details, represented as: Finally, pixel-level prediction probabilities are obtained through a single convolutional layer and activation function. , is represented as: ;in, It is a high-dimensional feature at layer l in the encoding stage; It is an upsampling operation; These are low-level features from the corresponding layer of the encoder; It is the decoded upsampled output of layer l; It is the Sigmoid function; This is the last convolutional operation, which converts the upsampled features into predicted values. This is the total decoded upsampled output; Dual loss optimization unit; including: Define region penalty; design region penalty loss; define pixel-level labels Y∈{0,1}, where 0 represents background and 1 represents the oral cavity opening region; the prediction probability is P. F (x), and power-modulate the predicted probability; region penalty loss Represented as: ;in, and It is a weighting factor; It is the modulation index; Define boundary numerical supervision loss Introducing boundary distance weighting The boundary numerical supervision loss is expressed as: ; ;in, It is the probability of matching between the predicted boundary numerical labels and the true boundary labels; It is an image region; and It refers to adjusting parameters; It is the distance from the pixel to the actual boundary; The final loss function L is expressed as: ;in, It is the weighting factor for the boundary loss.

Citation Information

Patent Citations

  • Method and apparatus for generating object detection model

    KR102379855B1

  • KR20220075713A