Radiotherapy target volume identification method based on large language model

By collecting and processing radiotherapy target images and historical data, and using a large language model for feature similarity calculation and texture profiling, the problems of inaccurate boundary extraction and insufficient texture features in radiotherapy target volume recognition in existing technologies are solved, and efficient radiotherapy target volume recognition and pathological state inference are achieved.

CN121033461APending Publication Date: 2025-11-28NINGBO INST OF TECH ZHEJIANG UNIV ZHEJIANG
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511138597.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-14
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing radiotherapy target volume recognition methods based on large language models suffer from inaccurate boundary extraction and insufficient texture features when dealing with complex pathological conditions. This makes it difficult to achieve multimodal medical image fusion and dynamic tumor change tracking, thus limiting the accuracy and effectiveness of radiotherapy planning.

Method used

By collecting radiotherapy target images and historical data, converting them into target grayscale images and extracting planar contours, identifying target plaque images for texture dissection, combining large language models for feature similarity calculation, matching similar pathological states, querying target parameters with consistent contours in historical data, and extracting the volume of radiotherapy target parameters.

Benefits of technology

It improves the accuracy and stability of radiotherapy target volume identification, enhances the system's adaptability and robustness to complex image features, supports data integration and joint analysis of multiple imaging modalities, and improves identification efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121033461A_ABST
    Figure CN121033461A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image recognition, in particular to a radiotherapy target volume recognition method based on a large language model, and the method comprises the following steps: collecting a radiotherapy target image and historical data, converting the target image into a grayscale image, separating a background, extracting a target plane contour, and cutting a target plaque image based on the contour. Texture analysis is carried out to generate target texture features, texture and contour information is combined, the pathological state of the target is speculated, a preset large language model is utilized to calculate the feature similarity between the pathological state of the target and historical data, a historical radiotherapy target set with similar pathological states is matched, and data with consistent historical contours are screened through contour matching. And corresponding radiotherapy parameters and volume information are extracted, and identification and analysis of precise radiotherapy target parameters are realized. According to the method, the automation level and the intelligent degree of target volume recognition are improved, the manual intervention requirement is reduced, and the recognition efficiency and accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image recognition, and in particular to a radiotherapy target volume recognition method based on a large language model. BACKGROUND

[0002] Current radiotherapy target volume recognition techniques rely heavily on traditional image processing and shallow machine learning methods. These methods often exhibit inaccurate boundary extraction and insufficient texture features when dealing with fuzzy target patches and complex tissue structures, which directly limits the accuracy and effectiveness of radiotherapy plan development, especially in the aspects of multi-modal medical image fusion and dynamic tumor change tracking. Traditional methods are difficult to achieve accurate inference and real-time adjustment of complex pathological states.

[0003] With the rapid development of large language models in the field of natural language processing, attempts have been made to introduce them into medical image feature semantic understanding and pathological state inference, which has become a new research hotspot. However, existing methods based on large language models mostly focus on the mining of single text information and simple feature matching, ignoring the deep fusion of image texture and structure multi-scale, multi-modal information, and lacking an efficient feature similarity calculation mechanism, which limits the model's ability to distinguish complex clinical pathological states and makes it difficult to meet the needs of precise radiotherapy.

[0004] Existing methods have significant shortcomings in matching similar pathological states in historical case libraries and updating dynamic pathological states, lacking a comprehensive recognition framework that takes into account image texture, structural features, and semantic information. This not only affects the efficiency of radiotherapy plan development, but also reduces the predictability and safety of treatment effectiveness, thus there is an urgent need for an innovative and highly generalizable radiotherapy target volume recognition method based on a large language model to break through existing bottlenecks. SUMMARY

[0005] To solve the above technical problems, the present application proposes a radiotherapy target volume recognition method based on a large language model to solve at least one of the above technical problems.

[0006] To achieve the above purpose, the present application provides a radiotherapy target volume recognition method based on a large language model, comprising the following steps:

[0007] Step S1: Collecting radiotherapy target images and historical radiotherapy target data; converting the radiotherapy target images into target grayscale images, and performing background separation according to the target grayscale images to extract target plane contours;

[0008] Step S2: Identifying target patch images in the radiotherapy target images according to the target plane contours, and performing texture dissection on the target patch images to generate radiotherapy target textures;

[0009] Step S3: Pathological state inference is performed according to the radiotherapy target texture and the target plane contour, so as to obtain a target pathological state;

[0010] Step S4: Feature similarity calculation is performed on the target pathological state and historical radiotherapy target data based on a preset large language model, and a historical radiotherapy target collection of similar pathological states is matched according to the feature similarity;

[0011] Step S5: Radiotherapy target parameters with consistent contours in the historical radiotherapy target collection are queried according to the target plane contour, and the volume of the radiotherapy target parameters is extracted.

[0012] The present application realizes accurate extraction of the target plane contour, enhances the precision of background separation, improves the quality of the basic data for subsequent target region recognition, guarantees the standardization and consistency of the input data, realizes recognition and texture analysis of the target patch image based on the target plane contour, enhances the multidimensional expression ability of the texture features, improves the fine-grained resolution of the texture information, realizes deep capture of complex texture structures, refines the texture segmentation and feature extraction process, enhances the discriminability and stability of the texture features, infers the pathological state in combination with the texture features and the plane contour, strengthens the fusion of spatial structures and texture information, improves the accuracy and stability of the pathological state inference, effectively reduces the influence of noise interference on the inference result, improves the generalization ability and adaptability of the inference model, performs high-dimensional feature similarity calculation on the pathological state and historical data by using a preset large language model, realizes deep matching at the semantic level, promotes efficient fusion and information complementation of multi-source heterogeneous data, improves the reliability and generalization ability of similar pathological state recognition, optimizes the matching precision through multidimensional mapping of the feature space, enhances the adaptability of the model to different data distributions, queries target parameters with consistent contours in the historical data according to the target contour and accurately extracts the volume information, realizes accurate mapping and quantization of spatial geometric parameters, provides accurate numerical support for volume recognition, ensures the stability and consistency of volume measurement, and overall improves the automation level and intelligent degree of target volume recognition, reduces the need for manual intervention, improves the recognition efficiency and accuracy, enhances the adaptability and robustness of the system to complex image features, and simultaneously supports data integration and joint analysis of multiple imaging modalities. BRIEF DESCRIPTION OF DRAWINGS

[0013] Figure 1 A step flowchart of the present application is shown in the figure.

[0014] Figure 2 A detailed implementation step flowchart of step S1 is shown in the figure.

[0015] Figure 3 A detailed implementation step flowchart of step S2 is shown in the figure. DETAILED DESCRIPTION

[0016] It should be understood that the specific embodiments described herein are merely exemplary and do not limit the application.

[0017] The application example provides a radiotherapy target volume identification method based on a large language model. The execution subject of the radiotherapy target volume identification method based on the large language model includes but is not limited to mechanical equipment, a data processing platform, a cloud server node, a network upload device and the like which can be regarded as a general computing node of the application, and the data processing platform includes but is not limited to at least one of an audio image management system, an information management system and a cloud data management system.

[0018] Please refer to Figures 1 to 3 The application provides a radiotherapy target volume identification method based on a large language model, which includes the following steps:

[0019] Step S1: collecting radiotherapy target images and historical radiotherapy target data; converting the radiotherapy target images into target gray scale images, and performing background separation according to the target gray scale images to extract target plane contours;

[0020] Step S2: identifying target patch images in the radiotherapy target images according to the target plane contours, and performing texture profiling on the target patch images to generate radiotherapy target textures;

[0021] Step S3: performing pathological state speculation according to the radiotherapy target textures and the target plane contours to obtain a target pathological state;

[0022] Step S4: performing feature similarity calculation on the target pathological state and the historical radiotherapy target data based on a preset large language model, and matching a historical radiotherapy target collection of a similar pathological state according to the feature similarity;

[0023] Step S5: querying radiotherapy target parameters with consistent contours in the historical radiotherapy target collection according to the target plane contours, and extracting the volume of the radiotherapy target parameters.

[0024] The application realizes accurate extraction of target plane contour, enhances the precision of background separation, improves the quality of basic data for subsequent target region recognition, guarantees the standardization and consistency of input data, identifies and texture profiling of target patch image based on target plane contour, enhances the multi-dimensional expression ability of texture features, improves the fine-grained resolution of texture information, realizes deep capture of complex texture structure, refines texture segmentation and feature extraction process, enhances the discriminability and stability of texture features, combines texture features and plane contour to infer pathological state, strengthens the fusion of spatial structure and texture information, improves the accuracy and stability of pathological state inference, effectively reduces the influence of noise interference on inference results, improves the generalization ability and adaptability of the inference model, uses a preset large language model to calculate the high-dimensional feature similarity of the pathological state and historical data, realizes deep matching at the semantic level, promotes efficient fusion and information complementation of multi-source heterogeneous data, improves the reliability and generalization ability of similar pathological state recognition, optimizes matching accuracy through multi-dimensional mapping of feature space, enhances the adaptability of the model to different data distribution, queries the target parameters with consistent contour in the historical data according to the target contour and accurately extracts the volume information, realizes accurate mapping and quantization of spatial geometric parameters, provides accurate numerical support for volume recognition, ensures the stability and consistency of volume measurement, and overall improves the automation level and intelligent degree of target volume recognition, reduces the demand for manual intervention, improves the recognition efficiency and accuracy, enhances the adaptability and robustness of the system to complex image features, and supports data integration and joint analysis of multiple imaging modalities.

[0025] In the embodiment of the application, reference is made to Figure 1 The steps of the radiotherapy target volume recognition method based on a large language model according to the application are shown in the flowchart, and in this example, the steps of the radiotherapy target volume recognition method based on a large language model include:

[0026] Step S1: Collect radiotherapy target images and historical radiotherapy target data; convert the radiotherapy target images into target gray images, and perform background separation according to the target gray images to extract target plane contour;

[0027] In this embodiment, when collecting the radiotherapy target image and the historical radiotherapy target data, a multi-modal medical image collection system is used, including but not limited to magnetic resonance imaging (MRI) and computed tomography (CT), to perform three-dimensional high-resolution scanning on the radiotherapy target area of the patient. During the image collection process, a fixed dose and a unified scanning protocol are used to ensure image quality and data consistency. After the collection is completed, an image preprocessing algorithm is used to convert color or multi-channel images into single-channel grayscale images. The background and target area in the grayscale image are separated by an adaptive threshold segmentation technique. A region growing algorithm or morphological closing operation is used to remove noise and smooth the boundary. The edge contour of the largest connected domain is extracted from the segmentation result. The contour extraction is based on the Canny edge detection algorithm, and a closed target plane contour is formed by combining the contour tracking technology. The entire contour extraction process controls the edge details by adjusting the kernel size of the Gaussian filter. The commonly used Gaussian kernel radius is 1.5 pixels.

[0028] Step S2: identifying the target patch image in the radiotherapy target image according to the target plane contour, and performing texture profiling on the target patch image to generate radiotherapy target texture;

[0029] In this embodiment, according to the extracted target plane contour, a cropping algorithm is used to cut the target area from the original radiotherapy image. During the cropping process, a polygon clipping algorithm is used to achieve accurate positioning. The cropping boundary completely matches the target contour. Then, the target patch image obtained by cropping is subjected to texture profiling. The texture profiling first extracts statistical features based on the gray level co-occurrence matrix (GLCM) method, including contrast, correlation, energy, and homogeneity. Further, local binary pattern (LBP) encoding is applied to image texture information. A 3x3 neighborhood window is used for LBP encoding to extract local texture microstructure. A Gabor filter bank is used to extract multi-scale and multi-direction texture responses. The Gabor filter parameters are set to 4 scales and 6 directions. The filter kernel size is fixed at 31x31. The texture feature vector is integrated into a high-dimensional texture descriptor by feature splicing. Finally, the radiotherapy target texture data is generated.

[0030] Step S3: performing pathological state inference according to the radiotherapy target texture and the target plane contour to obtain the target pathological state;

[0031] In this embodiment, principal component analysis (PCA) is performed on the texture data for dimensionality reduction, and the top 10 principal components are selected to represent the main variation characteristics of the texture. Then, geometric features are calculated based on the two-dimensional target contour data, including perimeter, compactness, and shape factor. The shape factor is calculated by dividing the target area by the square of the contour perimeter, and is used to describe the regularity of the shape. Subsequently, the principal components of the texture and the geometric features are jointly input into a multivariate regression model. The regression model weights are obtained by training historical pathology labeled samples. The output is a pathology state estimation value. The estimation process uses a multi-layer perception (MLP) structure, with 64 hidden layer nodes and a ReLU activation function. The batch size is set to 32 during training, and the number of iterations is 100. The mean squared error (MSE) is used as the loss function for optimization. Finally, the target pathology state is obtained by combining the model output and the contour area feature.

[0032] Step S4: Based on the pre-set large language model, the target pathology state and the historical radiotherapy target data are calculated for feature similarity, and the historical radiotherapy target set of similar pathology state is matched according to the feature similarity;

[0033] In this embodiment, the pathology state features in the target pathology state data and the historical radiotherapy target data are mapped in the feature space. A pre-trained large language model BERT (Bidirectional Encoder Representations from Transformers) is used to encode the text form of the pathology state information into a 768-dimensional embedding vector. The maximum sequence length is 128 in the vectorization process, and the input that does not meet the length is truncated or padded. After encoding, the cosine similarity between the matching target and the historical target is calculated as the feature similarity index. The cosine similarity calculation formula is similarity = vector A · vector B / (||vector A||*||vector B||). Based on the similarity value, a threshold of 0.85 is set to filter out a set of highly similar historical radiotherapy targets. The screening results are subdivided into similar sample categories through clustering analysis based on the k-nearest neighbor algorithm (KNN). In the clustering parameters, k is set to 5, and the clustering distance uses the Euclidean distance. Finally, the historical radiotherapy target set data containing similar pathology states is output.

[0034] Step S5: According to the target plane contour, the radiotherapy target parameters with consistent contours in the historical radiotherapy target set are queried, and the volume of the radiotherapy target parameters is extracted.

[0035] In this embodiment, according to the target plane profile data, profile matching query is performed from the historical radiotherapy target database, the profile matching adopts Hausdorff distance algorithm to evaluate the maximum distance error between two profile point sets, the matching threshold is set to 5 pixels, the historical target highly consistent with the current target profile is screened out, and the corresponding radiotherapy target parameters including volume, dose distribution and boundary shape are extracted according to the matching result, the volume calculation is based on the three-dimensional reconstruction model, the voxel counting method is adopted for three-dimensional reconstruction, the voxel side length is set to 0.5mm, the volume is the product of the voxel number and the voxel volume, and the extracted volume parameter is fused with the historical multiple treatment data through weighted average processing, and finally the parameter data set for radiotherapy target volume identification is generated.

[0036] In this embodiment, referring to Figure 2 The detailed implementation steps of step S1 include:

[0037] Step S11: collect radiotherapy target image and historical radiotherapy target data; perform gray scale conversion on the radiotherapy target image to generate a target gray scale image;

[0038] Step S12: perform preset pseudo-color mapping on the radiotherapy target image based on the target gray scale image to obtain a target pseudo-color mapping image;

[0039] Step S13: overlay the target pseudo-color mapping image into the radiotherapy target image to generate a color fusion image;

[0040] Step S14: separate the background color according to the color fusion image, and identify the target region image in the radiotherapy target image through the background separation data;

[0041] Step S15: extract the region boundary in the target region image to obtain a target plane profile.

[0042] In this embodiment, magnetic resonance imaging (MRI) and computed tomography (CT) equipment are used, and the acquisition parameters are set to a slice thickness of 1 millimeter and an image resolution of 512x512 pixels, ensuring the spatial resolution and detail integrity of the data. All image data is saved in DICOM format and imported into the processing system through a secure transmission protocol. After data acquisition, grayscale conversion is performed on the color or multi-channel images. The conversion process calls the cvtColor function of the OpenCV library with the parameter set to COLOR_BGR2GRAY to generate a single-channel 8-bit grayscale image. The size and spatial resolution of the converted image remain unchanged. The grayscale conversion uses a linear mapping formula to ensure that the pixel value range is 0 to 255. Based on the generated target grayscale image, a pseudo-color lookup table (LUT) is constructed. The lookup table includes 256 entries, each mapping a grayscale level to a corresponding RGB color. The color gradient transitions from deep blue to bright red. The mapping process calls the LUT function of OpenCV to map each pixel value of the grayscale image as an index to a color value, outputting a three-channel color image. The image size is the same as the original grayscale image Figure 1The mapping parameters are debugged to ensure that the organizational structure of different gray levels is represented by different colors, the pseudo-color image and the original radiotherapy image are fused, a weighted linear fusion formula is used, the pixel value of the fused image is alpha times the corresponding pixel value of the original image plus (1-alpha) times the corresponding pixel value of the pseudo-color image, the weight alpha is set to 0.6, the OpenCV addWeighted function is used to perform the fusion operation, the size, channel and data type of the two images are matched, the fusion process preserves the structural information of the original image and improves the color distinction between the target and the background, and the fusion result is output as a three-channel color image. The fused color image is converted to the HSV color space, the saturation (S) and lightness (V) channels are extracted to analyze the color features to distinguish the background and foreground regions, the Otsu threshold segmentation algorithm is applied to the lightness channel, the automatic threshold is calculated based on the image gray histogram, the image is segmented into foreground (potential target region) and background two parts, the segmentation result is a binary image, then the morphological opening operation is applied, the structure element is a 3*3 square, the first corrosion and then inflation processing is performed to remove small noise points and keep the integrity of the target region. After the opening operation, the connected component analysis is performed, the largest connected region is selected as the final target region mask, the Canny edge detection operator is called to extract the edge according to the target region mask, the low threshold is set to 50 and the high threshold is set to 150, the double threshold mechanism is used to determine the edge of the image gradient, and a binary edge image is generated. The white pixels in the edge image correspond to the detected edge position, then the Suzuki contour extraction algorithm is used to traverse the edge image, the closed contour is located, and the contour is stored in the form of a point set. The coordinate values in the point set are floating-point numbers to ensure spatial accuracy. Then, the polygon approximation processing is performed on the contour point set, the Douglas-Peucker algorithm is used, the approximation error is set to 2 pixels, the number of contour points is reduced, and the contour simplicity is improved. The point set of the contour represents the target plane boundary.

[0043] In this embodiment, refer to Figure 3 For the detailed implementation step flow diagram of step S2, in this embodiment, the detailed implementation step of step S2 includes:

[0044] Step S21: performing edge structure line enhancement processing on the radiotherapy target image according to the target plane contour to generate a target region edge line enhancement image;

[0045] Step S22: locating a contour region image in the radiotherapy target image through the target region edge line enhancement image, and performing patch cutting according to the contour region image to generate a target patch image;

[0046] Step S23: performing texture double-scale staggered filtering on the target patch image to generate a structure texture image and a detail texture image;

[0047] Step S24: performing texture block replacement on the structure texture block of the corresponding texture region in the structure texture image according to the detail texture image, to generate a radiotherapy target texture.

[0048] In this embodiment, the original image is subjected to Gaussian blur processing to reduce noise interference, a convolution kernel size of 5x5 is used, the standard deviation is set to 1.0, and the blurred image is used for subsequent edge enhancement operations. Then, the multi-scale Laplacian of Gaussian (LoG) is used to detect the edge features in the image, the spatial scale parameter σ is selected as 1.5, the significance of the structural lines in the image is enhanced, the detection results are normalized, and then the Non-Maximum Suppression is applied to refine the edge position and remove redundant responses. Further, a double-threshold edge detection algorithm is used, the threshold is set to a low threshold of 30 and a high threshold of 90, the edge breakpoints are connected in combination with the hysteresis threshold, a delicate and clear edge line enhancement image is generated, the edge line enhancement image is binarized, the threshold is set to 100, and the white pixels in the binary image represent the edge line region. The connected component analysis algorithm is used to extract all closed edge contours, the area threshold of the connected component is set to not less than 500 pixels to filter out noise contours, the extracted contour point set is stored in order, the original radiotherapy target image is cropped based on the contour bounding box, the cropping region coordinates are determined according to the four extreme points of the contour bounding box, and the cropped image is the target plaque image. The original resolution and pixel depth of the image are kept unchanged during the cropping process. Two sets of filter groups are constructed, the first group is a large-scale guided filter, the guide image is the target plaque image itself, the filter radius r is set to 15, and the regularization parameter ε is set to 1e-6, which is used to smooth the large-scale texture and preserve the main organizational hierarchy. The second group is a small-scale edge enhancement filter, a high-pass filter kernel is used, and the filter radius is set to 3 to enhance the fine texture details. The two-scale filter results are fused according to the time interleaving strategy, and the weighted average is used in the fusion process, with weights of 0.6 and 0.4 respectively. The fused image contains structural texture information and detailed texture information, and the image size and resolution are kept unchanged during the fusion process. Two separate channel images are output, which are the structural texture image and the detailed texture image, both in grayscale format. The structural texture image is divided into texture blocks with a size of 16x16 pixels, all texture blocks are traversed, and the texture block corresponding to the same position in the detailed texture image is calculated for each block. The local contrast and edge intensity difference of the two texture blocks are calculated, and the local variance and Sobel operator response are used as the measurement indicators. The local variance threshold is set to 15, and the Sobel operator response threshold is set to 20. The texture block that shows strong details in the detailed texture image is used to replace the corresponding block in the structural texture image. The replacement process is realized by pixel-by-pixel coverage to ensure the spatial consistency and edge continuity of the replacement area. After all the texture blocks are replaced, a complete radiotherapy target texture image is generated.

[0049] In this embodiment, the specific steps of the texture double-scale interleaved filtering of the target plaque image are as follows:

[0050] Large-scale guided filtering is applied to the target patch image to generate an initial structure-guided image;

[0051] By compressing the brightness and contrast fluctuations within the structural layers of the initial structure-guided image, a texture structure-adapted image is obtained.

[0052] Simultaneous small-scale edge enhancement filtering is applied to the target patch image to generate a detailed texture response image;

[0053] The local texture change rate is obtained by performing local texture change statistics on the detail texture response image;

[0054] Based on the local texture change rate, the texture structure adaptation image and the detail texture response image are fused to construct a texture interlacing fused image;

[0055] Large kernel guided filtering is applied to the texture interlacing and fusion image to extract the low-frequency response region and generate a structured texture image;

[0056] The remaining texture regions in the texture interleaving and fusion image are selected from the structural texture image to generate a detailed texture image.

[0057] In this embodiment, a guided filter is used, with the target patch image itself as the guide image. The filtering radius r is set to 15 pixels, and the regularization parameter ε is set to 1e-6. The filtering process smooths the image by calculating the local mean and variance, while preserving significant structural boundaries. The calculation formula is that pixel i in the output image q is equal to the local mean a multiplied by pixel i in the input image p plus the bias b, i.e., q_i = a_k × p_i + b_k. a_k and b_k are obtained by statistically calculating the pixels within the window. The window size is (2r+1) × (2r+1), dividing the image into 8×8 pixel blocks. The maximum gray value within each block is calculated. The minimum grayscale difference is defined as the local brightness contrast. After multiplying the contrast value by a compression factor of 0.5, the grayscale values ​​of all pixels in the small block are linearly scaled and adjusted. The adjustment formula is: new grayscale value = mean + (original grayscale value - mean) × 0.5. The mean is the grayscale mean of the small block. This ensures that the brightness fluctuation inside the structure layer is reduced, forming a texture structure that adapts to the image. The Sobel operator in the high-pass filter is selected as the edge enhancement operator. The filter kernel size is 3×3. The horizontal and vertical gradients are calculated separately. The gradient magnitude is calculated by the square root of the square. The specific calculation formula is: G = sqrt(Gx^2 + Gy^2).

[0058] Where Gx is the horizontal gradient and Gy is the vertical gradient, after filtering, the gradient magnitude image is normalized so that the pixel values ​​are mapped to the range of 0 to 255, generating a detailed texture response image. The sliding window method is adopted, with the window size set to 16×16 pixels and the step size set to 8 pixels. The entire image is traversed, and the standard deviation of the pixel gray level in each window is calculated as the local texture change rate. The standard deviation calculation formula is: σ=sqrt((1 / N)Σ(x_i-μ)^2);

[0059] Where N is the number of pixels in the window, x_i is the pixel gray value, and μ is the average gray value of the pixels in the window. The statistical results form a local texture change rate map, and the value reflects the intensity of texture change in the region. The fusion value of each pixel is calculated as follows: F(i,j)=α(i,j)×S(i,j)+(1-α(i,j))×D(i,j);

[0060] Where S(i,j) is the gray value of the texture structure adaptation image, D(i,j) is the gray value of the detail texture response image, and the fusion weight α(i,j) is obtained by mapping the local texture change rate L(i,j). The specific mapping function is α(i,j) = 1 - L(i,j). When the local texture change rate is low, the weight of the structure image is emphasized, and when it is high, the weight of the detail image is enhanced. The fusion process keeps the image size and gray depth unchanged and outputs a texture interleaved fused image. A guided filter is used, with the filter radius r set to 25 pixels and the regularization parameter ε set to 5e-6. The filtering process is also based on the local linear model to calculate the filter coefficients a and bias b. Local statistics are used to smooth the low-frequency region. After filtering, a structure texture image is generated. The structure texture image is then thresholded, with the threshold set to the local mean minus 0.1 times the standard deviation. After segmentation, a structure texture mask is obtained. In the mask, a pixel value of 1 represents a structure texture region, and 0 represents a non-structure texture region. This mask is applied to the texture interleaved fused image, and the regions with a mask value of 0 are extracted, which are the detail texture regions.

[0061] In this embodiment, the detailed implementation steps of step S3 include:

[0062] Extract the local texture direction distribution of the radiotherapy target texture;

[0063] The difference between the dominant and secondary dominant directions is statistically analyzed based on the local texture direction distribution, and the texture direction dispersion is calculated based on the difference.

[0064] The degree of texture dispersion is mapped to the roughness of the target region;

[0065] The radiotherapy target area is spatially defined based on the target plane contour, and the area of ​​the defined target space is calculated to generate the radiotherapy target area.

[0066] By integrating the target area of ​​radiotherapy and the roughness of the target region, the target pathological state can be obtained.

[0067] In this embodiment, the radiotherapy target texture image is divided into multiple non-overlapping basic sub-regions, each 32×32 pixels in size. Gradient direction extraction is performed within each sub-region, using the Sobel operator to extract the horizontal and vertical gradients Gx and Gy. The principal direction θ = arctangent(Gy / Gx) is calculated for each pixel. The calculated direction angle values ​​are limited to the range of 0 to 180 degrees. Histogram statistics are performed on the direction angles of all pixels within each sub-region, with an angle binning interval of 10 degrees, resulting in 18 discrete angle intervals. The number of pixels within each angle interval is counted to generate the direction distribution vector for that sub-region. This process is completed for the entire image. A local texture direction distribution map is then generated. The data format of this map is a two-dimensional array, with the horizontal axis being the sub-region index and the vertical axis being an 18-dimensional direction distribution vector. In each sub-region direction distribution vector, the angle corresponding to the largest count value is found as the dominant direction, and the angle with the second largest count value is the secondary dominant direction. The difference between the two is defined as the direction difference value Δθ=|θ_main-θ_sub|. The difference value is limited to the range of 0 to 90 degrees. The Δθ values ​​of all sub-regions are normalized. The largest direction difference is uniformly normalized to 1, and the smallest direction difference is 0, resulting in a local standardized direction difference map. For example, the direction distribution vector of sub-region [4,6] is [6,12,23,31,18,7,1,0,...,0].The angle interval with the largest count value (e.g., 40°) is extracted as the dominant direction, and the angle interval with the second largest count value (e.g., 90°) is extracted as the secondary dominant direction. The direction difference is 50°. This difference value is mapped to the normalized range [0,1], i.e., 50 / 90≈0.556, indicating that the texture direction of the current region has a moderate degree of dispersion. The sliding window mean method is used to process the standardized direction difference map. The window size is set to 3×3 sub-regions, and the step size is 1 sub-region. The average of the 9 direction differences within each sliding window is taken as the texture direction dispersion value of the central sub-region, generating a texture direction dispersion map. This map is a two-dimensional array with the same number of local sub-regions, and the value range is between 0 and 1, representing the dispersion intensity of the texture direction within each sub-region. The value of the region with uniformly arranged directions approaches 0, and the value of the region with variable direction distribution approaches 1. Binarization is used. The planar contour mask image spatially filters the texture direction discrete map, retaining only the regions with a mask value of 1, i.e., the texture discrete values ​​within the target region. Then, the mean of all discrete values ​​within this region is calculated, denoted as R_texture, where R_texture∈[0,1], and is defined as the roughness of the target region. If the target region area is divided into N effective sub-regions, edge contour data is called, and a closed polygon is constructed through the sequence of contour boundary points. The Green's formula (i.e., the polygon area formula) is used to sum the areas of the contour regions. All points are iteratively calculated sequentially to obtain the pixel area value of the target region on the image plane. Then, it is converted according to the actual spatial resolution of the image. If the image resolution is 0.25mm / pixel, the area conversion value is the original pixel area multiplied by (0.25×0.25)mm. 2 The output is the target area data, in mm. 2 Construct a two-dimensional feature vector F = [A, R], where A is the actual area of ​​the target region in mm. 2 R is the normalized value of texture roughness, in the form of a dimensionless real number. This feature vector is then input into the feature embedding layer derived from the large language model for encoding. The embedding process quantizes the real value into a fixed-position embedding vector representation by looking up a table.

[0068] In this embodiment, the specific steps for calculating the difference between the dominant and secondary dominant directions based on the local texture direction distribution, and for calculating the texture direction dispersion based on the difference, are as follows:

[0069] The directional angles of the local texture distribution are statistically analyzed, and the frequencies corresponding to all directional angles are identified to construct directional frequency data.

[0070] The dominant direction is obtained by filtering the direction angles with the highest frequency among all direction angles based on the direction frequency data.

[0071] Eliminate the dominant direction from all directional angles to determine the remaining directions;

[0072] Based on a preset directional angle difference threshold, texture directions in the remaining directions whose angular difference from the dominant direction is greater than the preset directional angle difference threshold are identified and marked as secondary dominant directions;

[0073] Calculate the angular difference between the dominant direction and the secondary dominant direction;

[0074] The degree of uniformity of texture direction is determined based on the angle difference, and then quantified as the degree of dispersion of texture direction.

[0075] In this embodiment, a sliding window is used to divide the radiotherapy target texture image into regions. Each sliding window is set to a size of 32×32 pixels and a window step size of 16 pixels. The gradient orientation angle θ of each pixel within the window region is obtained by calculating the gradient components Gx and Gy using the Sobel operator. The orientation angle calculation formula is θ = arctangent(Gy / Gx). All orientation angle data are limited to the range of 0° to 180°. Statistical processing is performed on the orientation angle data of all pixels within each sliding window. The angle interval is set to 10° increments, for a total of 18 orientation intervals. The orientation value of each pixel is assigned to the corresponding angle interval, and the number of pixels in each angle interval is accumulated to generate the orientation frequency. According to the data, this is represented as a one-dimensional vector of length 18. Each vector element corresponds to the frequency of occurrence of a direction angle interval, and the total number of pixels in each direction interval is represented as an integer. The 18 elements of the direction frequency vector are traversed, and the index of the highest frequency value is recorded. The angle interval corresponding to this index is the dominant direction angle. The center angle of the direction interval is determined by the angle binning setting. If the highest frequency occurs in the 5th interval, the dominant direction angle is (5-0.5)×10=45 degrees. This dominant direction angle represents the maximum trend direction of the texture direction in the current local area. The angle value is represented as an integer. The dominant direction data is stored in a dominant direction array, the length of which is equal to the number of sliding windows. The data is obtained from the original direction frequency vector... The element containing the dominant direction angle is cleared to zero. The index of the dominant direction in the direction frequency vector is set to i, and freq[i] is assigned a value of 0. The frequencies of other directions remain unchanged to ensure that the same direction is not repeatedly identified as a secondary dominant direction in subsequent processing. A direction difference threshold of 30 degrees is set. The center angles of all angle intervals are traversed from the direction frequency vector, and the difference between each angle and the dominant direction angle is calculated. When Δθ ≥ 30 degrees and the frequency value of that angle interval is greater than zero, it is marked as a candidate for secondary dominant direction. Among all directions that meet the condition, the direction angle with the highest frequency value is taken as the final secondary dominant direction angle and recorded as θ_sub. After processing, the angle pairs of the dominant and secondary dominant directions are obtained. The angle difference is calculated. The angle values ​​are limited to the range of 0 to 180 degrees, and the final Δθ value is limited to the range of 0 to 90 degrees to avoid amplifying angle errors caused by angular symmetry. The obtained angle differences are stored as real numbers in degrees, with decimal precision maintained to one decimal place. The angle differences are stored in the angle difference map as a direct representation of texture direction diversity. All angle differences are normalized as follows: the maximum difference of 90 degrees is normalized to 1, and the minimum difference of 0 degrees is normalized to 0. The normalized texture dispersion is then D = Δθ / 90. The normalization result is retained to four decimal places. The resulting D value constitutes a texture direction dispersion map, representing the texture consistency distribution of various local regions in the entire image. This dispersion map format is the same as the angle difference. Figure 1A two-dimensional array is used as the direct input data source for calculating texture roughness in subsequent processing. All discrete values ​​are unitless floating-point numbers with a value range limited to 0.0000 to 1.0000. The final output is an image data structure for texture direction discreteness.

[0076] In this embodiment, the detailed implementation steps of step S4 include:

[0077] Based on a pre-defined large language model, semantic fingerprint encoding is performed on the target pathological state to obtain pathological semantic vector features;

[0078] By using the same large language model, each pathological state in the historical radiotherapy target data is structurally semantically transformed to construct a set of historical semantic features;

[0079] The similarity is calculated for each feature vector in the pathological semantic vector feature and the historical semantic feature set to obtain the feature similarity.

[0080] By matching the target pathological state with historical radiotherapy target data based on feature similarity, a set of historical radiotherapy targets with similar pathological states is obtained.

[0081] In this embodiment, the obtained target pathological state is organized into standardized pathological description text in natural language, with a unified format of "target region area X, roughness Y", where X is the area value of the target planar contour in square millimeters, and Y is the texture dispersion in dimensionless numerical value. The input text undergoes semantic processing through a large language model built on a transformer structure. The large language model used employs the GPT-4 architecture, and the model's pre-training corpus is a joint integration of medical image structural language data and publicly available medical image diagnostic report corpus. The maximum input length of the model is 4096 tokens, and the model encoding dimension is 1536. The input text is mapped to a token sequence through a word embedding layer and processed through a multi-head self-attention mechanism. After extracting semantic context relevance using self-attention, the model's pooling layer outputs the final sentence vector representation, which serves as the semantic fingerprint encoding. The output is a 1536-dimensional floating-point vector, with each dimension ranging from -1.0 to +1.0. All floating-point numbers are rounded to four decimal places. All pathological states recorded in the historical radiotherapy dataset are extracted into structured data. Each record includes an area field and a dispersion field, with the field format uniformly set to "Area value: X mm". 2"Dispersion: Y", where X comes from the contour area analysis results of historical images, and Y is the dispersion value obtained from historical texture analysis. Each pathological state is converted into a natural language input with a unified semantic expression format consistent with the target pathological state. A set of standard input sequences is constructed, with the number of sequences equal to the number of historical pathological state records. Each input statement is independently semantically encoded using a pre-set large language model, with the processing method being completely consistent with the target semantic vector extraction process. The model output is the feature of each historical semantic vector, generating a total of N historical semantic vectors, where N is the total number of historical pathological data. Each vector has a dimension of 1536, and the vector set forms a two-dimensional array with shape [N, 1536], resulting in a collection of historical semantic features. Cosine... Similarity (cosine similarity) is used as the metric function. Let the target pathological semantic vector be V, the historical semantic feature set be matrix H, and the vector in the i-th row of H be h_i. All similarity values ​​are rounded to four decimal places. The final output is a one-dimensional similarity array of length N, where each value corresponds to the similarity index of a historical data point. The calculated similarity array is sorted in descending order, and the top K historical record indices are obtained. K is set to 10, indicating that the 10 records with the highest similarity are matched. Based on the sorting results, the historical radiotherapy target data corresponding to the index is extracted, including complete records such as historical image ID, target region contour, pathological structure description and parameters. Finally, the matched historical pathological state entries are constructed into a list structure and named the similar pathological state historical radiotherapy target collection. The output format is a JSON structured data object, where each record contains fields such as semantic vector, area, dispersion, original image path, and analysis timestamp.

[0082] In this embodiment, the detailed implementation steps of step S5 include:

[0083] Extract the historical contour of each radiotherapy target data in the historical radiotherapy target set, and encode and transform it to generate historical contour features;

[0084] Based on the target plane contour, match the historical contour features that are consistent with the historical contour features, and query the corresponding radiotherapy target parameters in the historical radiotherapy target set based on the historical contour consistent features.

[0085] Extract the volume of the target parameters for radiotherapy and determine the target volume parameters for radiotherapy.

[0086] In this embodiment, the two-dimensional contour data corresponding to each radiotherapy record in the historical radiotherapy target set of similar pathological states is read as a boundary coordinate sequence. Each contour data consists of continuous pixel pairs (x, y). The contour originates from the segmentation boundary of the target region in the original medical image. The segmentation boundary adopts the segmentation result based on the fusion of region growing algorithm and edge detection. The boundary coordinates are normalized and resampled using an equally spaced contour sampling method. Each contour is uniformly sampled into 256 coordinate points. All point coordinates are normalized so that the x and y values ​​are mapped to the [0, 1] interval. The normalized coordinate sequence is input into a Long Short-Term Memory (LSTM) network. Feature encoding is performed in a contour encoder model constructed using memory. The LSTM network structure is set to a unidirectional 3-layer stacked structure, with 512 hidden units in each layer. The output is a contour feature vector with a dimension of 512. All historical radiotherapy contour codes constitute a historical contour feature set of shape N×512, where N represents the number of contours in the historical pathology matching set. The target contour image obtained after background separation and plaque recognition of the target radiotherapy image also undergoes the same 256-point resampling, normalization, and LSTM encoding process to finally generate a 512-dimensional feature vector V of the target contour. t , for V t The encoding H of each historical contour in the historical contour feature set i Matching was performed, and similarity was calculated using Euclidean distance. A threshold of θ = 3.0 was set. When the Euclidean distance between two contours was less than θ, they were considered structurally similar. The volume parameter field in the historical radiotherapy data was named "Volume_mm3", with units of cubic millimeters (mm). 3 This field stores data in floating-point format. Volume data matching all historical contours are extracted into a one-dimensional array, such as: [3231.4,3368.9,3402.7,3275.5,3310.8]. The average of the radiotherapy volume parameters corresponding to the above five historical contours is calculated as follows:

[0087] Predicted volume = (3231.4 + 3368.9 + 3402.7 + 3275.5 + 3310.8) / 5 = 3317.86 mm 3 ;

[0088] The final target radiotherapy volume was 3317.86 mm². 3 The value is rounded to two decimal places and is used as the final radiotherapy volume identification result.

[0089] Therefore, the embodiments should be considered as exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of the equivalents of the application are intended to be included within the invention.

[0090] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.

Claims

1. A method for radiotherapy target volume recognition based on a large language model, characterized in that, Includes the following steps: Step S1: Acquire images of the radiotherapy target and historical radiotherapy target data; The radiotherapy target image is converted into a target grayscale image, and background separation is performed based on the target grayscale image to extract the target planar contour; Step S2: Identify the target patch image in the radiotherapy target image based on the target plane contour, and perform texture dissection on the target patch image to generate the radiotherapy target texture; Step S3: Based on the texture and planar contour of the radiotherapy target, the pathological state is inferred to obtain the target pathological state; Step S4: Based on the preset large language model, calculate the feature similarity of the target pathological state and historical radiotherapy target data, and match the set of historical radiotherapy targets with similar pathological states according to the feature similarity. Step S5: Query the radiotherapy target parameters with consistent contours in the historical radiotherapy target collection based on the target plane contour, and extract the volume of the radiotherapy target parameters.

2. The method for radiotherapy target volume recognition based on a large language model according to claim 1, characterized in that, The specific steps of step S1 are as follows: Acquire images of radiotherapy targets and historical radiotherapy target data; perform grayscale conversion on the radiotherapy target images to generate target grayscale images; A pseudo-color mapping of the radiotherapy target image is performed on the target grayscale image to obtain the target pseudo-color mapped image; The target pseudo-color mapping image is overlaid onto the radiotherapy target image to generate a color fusion image; Background color separation is performed based on the color fusion image, and the target region image in the radiotherapy target image is identified through the background separation data. Extract the region boundaries from the target region image to obtain the target planar contour.

3. The method for radiotherapy target volume recognition based on a large language model according to claim 1, characterized in that, The specific steps of step S2 are as follows: The edge structure line enhancement processing of the radiotherapy target image is performed based on the target plane contour to generate an enhanced edge line image of the target region. The contour region image in the radiotherapy target image is located by enhancing the image of the target region edge line, and the target patch image is generated by patch cropping based on the contour region image; The target patch image is subjected to dual-scale interleaved texture filtering to generate structural texture image and detail texture image; Based on the detailed texture image, the structural texture blocks in the corresponding texture region of the structural texture image are replaced to generate the radiotherapy target texture.

4. The method for radiotherapy target volume recognition based on a large language model according to claim 1, characterized in that, The specific steps of performing texture dual-scale interleaved filtering on the target patch image are as follows: Large-scale guided filtering is applied to the target patch image to generate an initial structure-guided image; By compressing the brightness and contrast fluctuations within the structural layers of the initial structure-guided image, a texture structure-adapted image is obtained. Simultaneous small-scale edge enhancement filtering is applied to the target patch image to generate a detailed texture response image; The local texture change rate is obtained by performing local texture change statistics on the detail texture response image; Based on the local texture change rate, the texture structure adaptation image and the detail texture response image are fused to construct a texture interlacing fused image; Large kernel guided filtering is applied to the texture interlacing and fusion image to extract the low-frequency response region and generate a structured texture image; The remaining texture regions in the texture interleaving and fusion image are selected from the structural texture image to generate a detailed texture image.

5. The method for radiotherapy target volume recognition based on a large language model according to claim 1, characterized in that, Step S3 is as follows: Extract the local texture direction distribution of the radiotherapy target texture; The difference between the dominant and secondary dominant directions is statistically analyzed based on the local texture direction distribution, and the texture direction dispersion is calculated based on the difference. The degree of texture dispersion is mapped to the roughness of the target region; The radiotherapy target area is spatially defined based on the target plane contour, and the area of ​​the defined target space is calculated to generate the radiotherapy target area. By integrating the target area of ​​radiotherapy and the roughness of the target region, the target pathological state can be obtained.

6. The method for radiotherapy target volume recognition based on a large language model according to claim 1, characterized in that, The step of calculating the difference between the dominant and secondary dominant directions based on the local texture direction distribution, and then calculating the texture direction dispersion based on the difference, specifically involves: The directional angles of the local texture distribution are statistically analyzed, and the frequencies corresponding to all directional angles are identified to construct directional frequency data. The dominant direction is obtained by filtering the direction angles with the highest frequency among all direction angles based on the direction frequency data. Eliminate the dominant direction from all directional angles to determine the remaining directions; Based on a preset directional angle difference threshold, texture directions in the remaining directions whose angular difference from the dominant direction is greater than the preset directional angle difference threshold are identified and marked as secondary dominant directions; Calculate the angular difference between the dominant direction and the secondary dominant direction; The degree of uniformity of texture direction is determined based on the angle difference, and then quantified as the degree of dispersion of texture direction.

7. The method for radiotherapy target volume recognition based on a large language model according to claim 1, characterized in that, Step S4 is as follows: Based on a pre-defined large language model, semantic fingerprint encoding is performed on the target pathological state to obtain pathological semantic vector features; By using the same large language model, each pathological state in the historical radiotherapy target data is structurally semantically transformed to construct a set of historical semantic features; The similarity is calculated for each feature vector in the pathological semantic vector feature and the historical semantic feature set to obtain the feature similarity. By matching the target pathological state with historical radiotherapy target data based on feature similarity, a set of historical radiotherapy targets with similar pathological states is obtained.

8. The method for radiotherapy target volume recognition based on a large language model according to claim 1, characterized in that, The specific steps of step S5 are as follows: Extract the historical contour of each radiotherapy target data in the historical radiotherapy target set, and encode and transform it to generate historical contour features; Based on the target plane contour, match the historical contour features that are consistent with the historical contour features, and query the corresponding radiotherapy target parameters in the historical radiotherapy target set based on the historical contour consistent features. Extract the volume of the target parameters for radiotherapy and determine the target volume parameters for radiotherapy.

Citation Information

Cited By

  • Skin cancer risk assessment method and system combining vision and language model

    CN121639704A

  • Intelligent medical care interaction platform system based on image processing technology

    CN122067752A