Image processing method for oral mucosa squamous carcinoma

Through the dual-model synergistic segmentation and boundary-guided diffusion map of saliva mask and lesion primary segmentation mask, the problem of lesion area identification under complex background in oral mucosal squamous cell carcinoma images is solved, and more accurate lesion segmentation and detection are achieved.

CN120298403AActive Publication Date: 2025-07-11HUNAN RUICHENG BIOTECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510774347.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-07-11
Estimated Expiration
2045-06-11

AI Technical Summary

Technical Problem

The existing oral mucosal squamous cell carcinoma image processing methods are difficult to accurately identify lesion areas with blurred boundaries under the treatment of reflective interference and complex backgrounds, resulting in high leakage detection rate and lack of joint modeling of multidimensional structural features of color, texture and spatial boundaries.

Method used

A dual-model synergistic segmentation mechanism of saliva mask and lesion primary segmentation mask is adopted, combined with the boundary-guided diffusion map of boundary mutation intensity, and through multi-dimensional fusion of color features, texture features and boundary structure features, an image processing method is constructed to improve the segmentation accuracy and robustness of the lesion area.

Benefits of technology

Effectively distinguishing highly reflective interference areas from real lesion areas, improving the recovery and detection capabilities of lesion structures in fuzzy areas, reducing the missed detection rate, and improving the boundary coherence and regional integrity of lesion segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298403A_ABST
    Figure CN120298403A_ABST
Patent Text Reader

Abstract

The invention discloses an image processing method for oral mucosa squamous carcinoma, and relates to the technical field of image processing, and the method comprises the steps: inputting # imgabs0 # into a first semantic segmentation model, inputting # imgabs1 # into a second semantic segmentation model, and obtaining a saliva mask and a local positive focus mask # imgabs2 #; a transition mask # imgabs3 # is obtained; calculating boundary mutation intensity, constructing a boundary guide diffusion diagram # imgabs4 # according to the boundary mutation intensity, segmenting a local positive focus mask # imgabs7 # from # imgabs6 # by using # imgabs5 #, and merging # imgabs8 # and # imgabs9 # to obtain an overall focus mask; according to the method, the oral mucosa squamous cell carcinoma lesion area can be segmented more completely and accurately under the conditions of strong reflective shielding and complex background.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and specifically to an image processing method for oral squamous cell carcinoma of the mucosa. Background Art

[0002] Oral squamous cell carcinoma of the mucosa is a common oral malignant lesion. Its early stage is characterized by irregular tissue structure, blurred boundaries, and small area. With the development of digital medical imaging technology, in more and more clinical application scenarios, oral endoscopy images or digital scan images are used to analyze the lesion area to achieve the auxiliary detection and judgment of positive lesions.

[0003] Existing oral image processing methods mostly perform image segmentation based on color features or shallow semantic information. However, in actual images, due to interference factors such as saliva reflection, imaging angle, and illumination change, there are a large number of highlighted areas in the images, and these areas usually block the visible information of the actual lesion tissue, thus significantly reducing the segmentation accuracy. In addition, the color differences and texture features of the early lesion areas are highly similar to those of normal tissues, and it is difficult to accurately identify them through traditional methods.

[0004] On the other hand, existing methods usually process based on a single feature dimension and lack the ability to jointly model multi-dimensional structural features such as color information, texture features, and spatial boundaries, resulting in insufficient recognition robustness in complex backgrounds. Especially in the lesion areas with blurred edges and being occluded, the missed detection rate is relatively high.

[0005] Therefore, there is an urgent need for an image processing method that can enhance the recognition ability of the blurred boundary lesion area while dealing with reflection interference, and can have a strong boundary understanding ability to improve the automatic processing quality and segmentation accuracy of oral squamous cell carcinoma of the mucosa images. Summary of the Invention

[0006] The present invention aims to solve at least one of the technical problems existing in the prior art. To this end, the present invention proposes an image processing method for oral squamous cell carcinoma of the mucosa.

[0007] To achieve the above object, the present invention provides an image processing method for oral squamous cell carcinoma of the mucosa, including: Obtaining a first oral image , and a replicated second oral image ; Inputting into a pre-trained first semantic segmentation model to obtain a saliva mask ; and inputting into a pre-trained second semantic segmentation model for initial segmentation to obtain a first local positive lesion mask ; Obtain the and to perform an AND operation to generate an overlapping transition mask ; Calculate the boundary mutation intensity, construct a boundary-guided diffusion map based on the boundary mutation intensity , and use to segment the second local positive lesion mask from , and merge with to obtain the overall lesion mask .

[0008] Preferably, the training method of the first semantic segmentation model is as follows: Obtain a historical oral image sample set, divide the historical oral image sample set into a first semantic segmentation training set and a first semantic segmentation test set. The historical oral image sample set includes multiple grayscale sample oral images and corresponding first annotation labels; Among them, the first annotation labels include a first positive label and a first negative label. The first positive label represents a saliva mask , and the first negative label represents a non-saliva mask; Initialize the first convolutional neural network, use the in the first semantic segmentation training set as the input of the first convolutional neural network, and use the first annotation labels as the output, and train the first convolutional neural network to obtain a first semantic segmentation network to be verified; Use the first semantic segmentation test set to verify the first semantic segmentation network, and output the first semantic segmentation network with an error less than or equal to the preset test error threshold as the trained first semantic segmentation model.

[0009] Preferably, the acquisition method of the first annotation labels is as follows: Obtain a standard lesion image and a sample oral image , and obtain the first color feature from the saliva area in , and obtain the second color feature of each pixel in ; Among them, the first color feature includes a target saturation interval and a target lightness interval, and the second color feature includes a saturation value and a lightness value; Compare the second color feature of each pixel in with the first color feature to screen out the first target pixels and non-first target pixels in : in ; In the formula: represents the sample oral cavity image at the color features of the pixel, including the saturation value and the lightness value; is the first target pixel, represents a non-first target pixel, represents the target saturation interval, represents the target lightness interval; All pixel regions formed by the connection of the first target pixels are used as the saliva mask , and labeled as the first positive label, and all pixel regions formed by the connection of non-first target pixels are used as non-saliva masks and labeled as the first negative label; Repeat the above steps until all in the historical oral cavity image sample set are labeled, and all first annotation labels are obtained.

[0010] Preferably, the method for obtaining the target saturation interval is as follows: The standard lesion image is converted from the RGB color space to the HSV color space to obtain the saturation image , where the saturation range of each pixel is [0,1]; From the saturation values corresponding to the pixels belonging to the saliva region and the lesion region are respectively extracted to form the first saturation set and the second saturation set ; The saturation range [0,1] is evenly divided into K saturation sub-intervals, and each saturation sub-interval is defined as: , is the i-th saturation sub-interval; Count the number of pixels in the first saturation set falling into each saturation sub-interval, denoted as , and count the number of pixels in the second saturation set falling into each saturation sub-interval, denoted as ; According to and respectively construct the first saturation histogram and the second saturation histogram ; Arbitrarily select two consecutive saturation sub-intervals as the saturation sub-interval combination , where 0 ≤ m ≤ n < K, and define the first error term and the second error term : Among them, the first error term is calculated as: ; Among them, the second error term is calculated as: ; According to the first error term and the second error term construct the discriminant error total loss function : ; In the formula: and are the weighting coefficients of the first error term and the second error term respectively, ; Traverse all combinations of saturation sub-intervals, and calculate the corresponding discriminant error total loss function for each combination of saturation sub-intervals; For all combinations of saturation sub-intervals corresponding perform ascending sorting, and take the combination of saturation sub-intervals corresponding to the first as the target saturation interval .

[0011] Preferably, the training method of the second semantic segmentation model is as follows: Obtain a historical oral image sample set, divide the historical oral image sample set into a second semantic segmentation training set and a second semantic segmentation test set, and the historical oral image sample set also includes a second annotation label corresponding to the sample oral image ; Among them, the second annotation label includes a second positive label and a second negative label. The second positive label is represented as a first local positive lesion mask , and the second negative label represents a non-first local positive lesion mask; Initialize the second convolutional neural network, use the in the second semantic segmentation training set as the input of the second convolutional neural network, and use the second annotation label as the output, and train the second convolutional neural network to obtain a second semantic segmentation network to be verified; Use the second semantic segmentation test set to verify the second semantic segmentation network, and output the second semantic segmentation network with an error less than or equal to the preset test error threshold as the trained second semantic segmentation model.

[0012] Preferably, the acquisition method of the second annotation label is as follows: The grayscale sample oral image is divided into local window regions in a sliding window manner to obtain a set of local window region blocks , where the sliding step is 1, the window size is W×W, and Z and W are positive integers; Obtain the first texture features of the lesion area in the standard lesion image , and the second texture features of each local window region in; wherein, the first texture feature and the second texture feature are both statistical feature sets of the gray-level co-occurrence matrix, including multiple statistical features, including energy, average contrast, and entropy; Calculate the texture similarity between the first texture feature of the lesion area and the second texture feature of the local window region , and its calculation formula is as follows: ; In the formula: is the eigenvalue of the r-th statistical feature in the first texture feature of the lesion area, is the eigenvalue of the r-th statistical feature in the second texture feature of the local window region, is the total number of statistical features, is a positive real number; Compare the texture similarity with a preset texture similarity threshold. If the texture similarity is greater than or equal to the texture similarity threshold, the corresponding local window region is used as the first local positive lesion mask , and is labeled as the second positive label. Otherwise, the corresponding local window region is used as a non-first local positive lesion mask and is labeled as the second negative label; Repeat the above steps until all in the historical oral image sample set are labeled, and all second annotation labels are obtained.

[0013] Preferably, the calculation of the boundary mutation intensity includes: For each pixel point in , perform the following operations: For the pixel point calculate the component of the Sobel gradient in the x direction and the component in the y direction : ; ; In the formula: represents the partial derivative of the image in the horizontal x direction, represents the image Partial derivative in the vertical y direction; Combine the component with the component into a gradient vector , and calculate its counterclockwise perpendicular vector according to the gradient vector as the boundary normal direction vector: ; Normalize the boundary normal direction vector to obtain the unit normal direction ; Starting from , sample e pixel points in sequence along the unit normal direction to form a sampling sequence: ; For each sampling point , calculate the Sobel gradient magnitude on the original first oral cavity image : ; Combine the gradient magnitudes of e pixel points into a gradient sequence , and calculate the maximum value of the first-order difference of this gradient sequence, which is defined as the boundary mutation intensity : .

[0014] Preferably, constructing the boundary-guided diffusion map according to the boundary mutation intensity , includes: Construct a two-dimensional floating-point map with the same size as the original first oral cavity image , and initialize all its pixel positions to zero values: ; Traverse all pixel points in the transition mask , and judge whether the boundary mutation intensity of each pixel point is greater than the preset structure mutation discrimination threshold ; If not, mark the corresponding pixel point as a non-diffusion source point; If so, mark the corresponding pixel point as a diffusion source point, and for each diffusion source point, start again from and sample pixel points in sequence along the unit normal direction vector to form a new sampling sequence: ; For each sampling point , calculate its diffusion weight: ; In the formula: is the standard deviation of the diffusion kernel, is the exponential function; The boundary mutation intensity , according to the weight is weighted and accumulated onto the boundary-guided diffusion map , and the update formula is: .

[0015] Preferably, the second local positive lesion mask is segmented from , including: Set the lesion possibility threshold ; For each pixel in, perform the following pixel discrimination: ; In the formula: , indicating that the pixel is marked as 1 in , and its value in the boundary-guided diffusion map is greater than or equal to the lesion possibility threshold ; indicates that the pixel does not satisfy or does not satisfy , 1 represents the second target pixel, and 0 represents the non-second target pixel; The region formed by the second target pixels obtained after discrimination is segmented from , and is marked as the second local positive lesion mask .

[0016] An image processing system for oral squamous cell carcinoma, implemented based on the above-mentioned image processing method for oral squamous cell carcinoma, includes: An acquisition module for acquiring a first oral image , and a second oral image after copying ; A primary segmentation module for inputting into a pre-trained first semantic segmentation model to obtain a saliva mask ; and inputting into a pre-trained second semantic segmentation model for primary segmentation to obtain a first local positive lesion mask ; A merging module for obtaining when and A transition mask that produces an overlapping part after performing an AND operation ; A secondary segmentation module for calculating the boundary mutation intensity and constructing a boundary-guided diffusion map based on the boundary mutation intensity , using to segment out the second local positive lesion mask , and merge with to obtain the overall lesion mask

[0017] Compared with the prior art, the beneficial effects of the present invention are as follows: The oral mucosal squamous cell carcinoma image processing method provided by the present invention can effectively distinguish the highly reflective interference region and the real lesion region in the image and reduce the influence of the saliva region on the accuracy of lesion recognition by introducing a dual-model collaborative segmentation mechanism of the saliva mask and the initial lesion segmentation mask; at the same time, by constructing a boundary-guided diffusion map based on the boundary mutation intensity, potential occluded lesion edges are mined in the transition region, effectively improving the ability to recover and detect the lesion structure in the blurred region.

[0018] Compared with the existing methods based on single feature or single model processing, the present invention utilizes the multi-dimensional fusion of color features, texture features and boundary structure features, enhances the detection robustness for boundary blur and small-scale lesions, and realizes a more complete and accurate segmentation of the oral mucosal squamous cell carcinoma lesion region under strong specular occlusion and complex background conditions; overall, the present invention can improve the boundary coherence and regional integrity of lesion segmentation, reduce the missed detection rate caused by specular occlusion, and has good application prospects and practical promotion value. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0020] Figure 1 is a flowchart of the method of the present invention; Figure 2 is a structural diagram of the system of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] The technical solution of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative work belong to the scope of protection of the present invention.

[0022] Please refer to Figure 1 , an embodiment of the first aspect of the present invention provides an image processing method for oral squamous cell carcinoma, including: Step 1: Obtain a first oral image , and a second oral image after replication ; It should be understood that: the first oral image is obtained by photographing the oral cavity of a patient with oral squamous cell carcinoma using a device such as an oral endoscope (such as Olympus ENF-VH) or a digital oral scanner, with a resolution of 1024×1024, an output image format of PNG, and a channel number of RGB; the second oral image is obtained by directly replicating the first oral image ; it should be noted that the first oral image is exactly the same as the second oral image in terms of size, resolution, direction, etc.

[0023] Step 2: Input into a pre-trained first semantic segmentation model to obtain a saliva mask ; and input into a pre-trained second semantic segmentation model for initial segmentation to obtain a first local positive lesion mask ; Specifically, the training method of the first semantic segmentation model is as follows: Obtain a historical oral image sample set, divide the historical oral image sample set into a first semantic segmentation training set and a first semantic segmentation test set. The historical oral image sample set includes multiple grayscale sample oral images and corresponding first annotation labels; Among them, the first annotation label includes a first positive label and a first negative label. The first positive label represents a saliva mask , and the first negative label represents a non-saliva mask; In practice, the acquisition method of the first annotation label is as follows: Obtain a standard lesion image and a sample oral image , and from the saliva area in Obtain a first color feature and obtain the second color feature of each pixel in Among them, the standard lesion image includes a pre-annotated saliva region and a lesion region ; Among them, the first color feature includes a target saturation interval and a target lightness interval, and the second color feature includes a saturation value and a lightness value; Specifically, the acquisition method of the target saturation interval is as follows: Convert the standard lesion image from the RGB color space to the HSV color space to obtain a saturation image , where the saturation range of each pixel is [0,1]; It should be noted that: the standard lesion image is pre-stored in the system database and includes a saliva region and a lesion region. Among them, the saliva region and the lesion region in the standard lesion image are manually annotated by oral squamous cell carcinoma doctors and experts; Among them, the conversion from the RGB color space to the HSV color space includes: 1) Calculate the hue, specifically as follows: ; 2) Calculate the saturation, specifically as follows: ; 3) Calculate the lightness, specifically as follows: ; In the formula: is the color difference, , is the maximum value of the color channel, is the minimum value of the color channel, is the hue, is the saturation, is the lightness, is the pixel value of the red channel, G is the pixel value of the green channel, the pixel value of the blue channel; Extract the saturation values corresponding to the pixels belonging to the saliva region and the lesion region from respectively to form a first saturation set and a second saturation set ; Among them, , represents the saturation values of all pixels belonging to the saliva region; , representing the saturation values of all pixels belonging to the lesion area; The saturation range [0, 1] is evenly divided into K saturation sub - intervals, and each saturation sub - interval is defined as: , is the i - th saturation sub - interval; Count the number of pixels in the first saturation set that fall into each saturation sub - interval, denoted as , and count the number of pixels in the second saturation set that fall into each saturation sub - interval, denoted as ; According to and respectively construct the first saturation histogram and the second saturation histogram ; Arbitrarily select two consecutive saturation sub - intervals as the saturation sub - interval combination , where 0 ≤ m ≤ n < K, and define the first error term and the second error term : Among them, the first error term , representing the number of pixels in the selected saturation sub - interval combination that are mis - identified as pixels in the saliva area from the lesion area pixels, and the calculation method is: ; Among them, the second error term , representing the number of saliva area pixels that are missed as non - saliva area pixels outside the interval , and the calculation method is: ; According to the first error term and the second error term construct the discriminant error total loss function : ; In the formula: and are the weighting coefficients of the first error term and the second error term respectively, used to adjust the influence intensity of mis - judging the lesion and missing the reflection in the objective function, , which is set by the technical staff according to historical data; Traverse all saturation sub - interval combinations, and calculate the corresponding discriminant error total loss function for each saturation sub - interval combination; For all saturation sub - interval combinations corresponding Perform ascending sorting and take the one ranked first The corresponding saturation sub-interval combination As the target saturation interval ; It should be noted that the acquisition logic of the target lightness interval is the same as that of the above-mentioned target saturation interval. The difference is that when determining the target saturation interval, the lightness range is first obtained as [0,1], and the first lightness set (representing the lightness values of all pixels belonging to the saliva region) and the second lightness set (representing the lightness values of all pixels belonging to the lesion region) are extracted. K lightness sub-intervals are divided, and the number of pixels falling into each lightness sub-interval is counted And the number of pixels falling into each lightness sub-interval is used to construct the first lightness histogram and the second lightness histogram to define the third error term and the fourth error term to construct the discriminant error total loss function , and are the weighting coefficients of the third error term and the fourth error term respectively, , and finally all lightness sub-interval combinations are traversed, and through ascending sorting and selection of all lightness sub-interval combinations corresponding , the target lightness interval is obtained. For details, refer to the relevant part of the above-mentioned target saturation interval, and no more elaboration will be made here; Compare the second color feature of each pixel in with the color feature of the first color feature to screen out the first target pixels and non-first target pixels in ; In the formula: represents the color feature of the pixel of the sample oral image at , including the saturation value and the lightness value; is the first target pixel, represents the non-first target pixel, represents the target saturation interval, represents the target lightness interval; Take all the pixel regions formed by the connection of the first target pixels as the saliva mask , and label it as the first positive label, and take all the pixel regions formed by the connection of the non-first target pixels as the non-saliva mask and label it as the first negative label; Repeat the above steps until all in the historical oral image sample set are labeled, obtaining all first annotation labels; Initialize the first convolutional neural network, use the in the first semantic segmentation training set as the input of the first convolutional neural network, and use the first annotation label as the output, and train the first convolutional neural network to obtain the to-be-verified first semantic segmentation network; Use the first semantic segmentation test set to verify the model of the first semantic segmentation network, and output the first semantic segmentation network with a value less than or equal to the preset test error threshold as the trained first semantic segmentation model; Specifically, the training method of the second semantic segmentation model is as follows: Obtain the historical oral image sample set, divide the historical oral image sample set into a second semantic segmentation training set and a second semantic segmentation test set, and the historical oral image sample set also includes the corresponding to the sample oral image second annotation labels; It should be noted that: the sample oral images in the historical oral image sample set are obtained by technicians taking medical images of the oral cavities of several patients or non-patients in advance and stored in the system database; Among them, the second annotation labels include a second positive label and a second negative label, and the second positive label is represented as a first local positive lesion mask and the second negative label represents a non-first local positive lesion mask; In practice, the acquisition method of the second annotation label is as follows: Perform window division on the grayscale sample oral image in a sliding window manner to obtain a set of local window region blocks , where the sliding step is 1, the window size is W×W (such as W = 32), and Z and W are positive integers; It should be noted that: when the sliding window exceeds the image boundary, mirror filling or zero filling is used for processing to ensure the integrity of the window block; Obtain the first texture feature of the lesion area in the standard lesion image , and the second texture feature of each local window region in; Among them, both the first texture feature and the second texture feature are statistical feature sets of the gray-level co-occurrence matrix, including multiple statistical features, including but not limited to energy, average contrast, and entropy, etc.; Calculate the texture similarity between the first texture feature of the lesion area and the second texture feature of the local window region , and its calculation formula is as follows: ; In the formula: is the eigenvalue of the r-th statistical feature in the first texture feature of the lesion area, is the eigenvalue of the r-th statistical feature in the second texture feature of the local window area, is the total number of statistical features, is a (very small) positive real number used to avoid division by zero error; Compare the texture similarity with a preset texture similarity threshold. If the texture similarity is greater than or equal to the texture similarity threshold, then the corresponding local window area is used as the first local positive lesion mask , and mark it as the second positive label. Otherwise, the corresponding local window area is used as a non-first local positive lesion mask and marked as the second negative label; Repeat the above steps until all in the historical oral image sample set are labeled, and all second annotation labels are obtained; Initialize the second convolutional neural network, use the in the second semantic segmentation training set as the input of the second convolutional neural network, and use the second annotation label as the output to train the second convolutional neural network to obtain the second semantic segmentation network to be verified; Use the second semantic segmentation test set to verify the model of the second semantic segmentation network, and output the second semantic segmentation network with an error less than or equal to the preset test error threshold as the trained second semantic segmentation model; Drive the dual-model training through the three features of saturation, lightness, and texture to improve the accuracy of the initial segmentation and enhance the processing ability for local small lesions and specular interference.

[0024] Step 3: Obtain the transition mask and of the overlapping part generated after performing the AND operation on ; Among them, the expression of the AND operation is as follows: ; where represents the overlapping part that is recognized as both the saliva area and the lesion area in the image, called the transition area; the transition area represents the area where some lesions are blurred due to saliva visually, but there is still a possibility of lesions; It should be understood that the image will be binarized before the operation. A fixed threshold (0.5) or an empirical threshold based on the probability distribution output by the model is used for binarization. Among them, the AND operation is a common logical operation in computer image processing, and its purpose is to find the transition region. Further explanation is that for two segmented images (binary images where each pixel value is 0 or 1), if the pixel values at the same position in both images are 1, then 1 is output (indicating the overlapping part), and if any one of the pixel values is 0, then 0 is output (indicating that this part is not the overlapping part).

[0025] Step 4: Calculate the boundary mutation intensity and construct a boundary-guided diffusion map based on the boundary mutation intensity , using to segment out the second local positive lesion mask , and merge it with to obtain the overall lesion mask ; In implementation, the calculation of the boundary mutation intensity includes: For each pixel point in , the following operations are performed: Calculate the component of the Sobel gradient in the x direction and the component in the y direction : ; ; In the formula: represents the partial derivative of the image in the horizontal x direction, represents the partial derivative of the image in the vertical y direction, which is actually calculated by the Sobel filter; Combine the component with the component to form a gradient vector , and calculate its counterclockwise vertical vector according to the gradient vector as the boundary normal direction vector: ; Normalize the boundary normal direction vector to obtain the unit normal direction ; It should be noted that the normalization formula is as follows: ; where represents the Euclidean norm of the normal vector; Taking Starting from sample e pixel points in sequence along the unit normal direction to form a sampling sequence: ; For each sampling point , calculate the Sobel gradient magnitude on the original first oral cavity image : ; In an alternative embodiment, if the sampling point is at a non-integer coordinate position, then use bilinear interpolation on the original first oral cavity image to estimate the Sobel gradient components from the gray values of the four neighboring pixels; Form a gradient sequence from the gradient magnitudes of the e pixel points, and calculate the maximum first-order difference of this gradient sequence, which is defined as the boundary mutation intensity :

[0026] In the formula: represents the maximum degree of mutation of the image gradient intensity along the normal direction at the boundary ; the larger the value, the more obvious the boundary interruption, and the more likely it is the position of the lesion structure occluded by saliva; in practice, constructing the boundary-guided diffusion map according to the boundary mutation intensity includes: Construct a two-dimensional floating-point map with the same size as the original first oral cavity image , and initialize all its pixel positions to zero values: ; In the formula: represents the universal quantifier, indicating that the initialization operation is performed on all pixel coordinates in the image; Traverse all pixel points in the transition mask , and judge whether the boundary mutation intensity of each pixel point is greater than the preset structural mutation discrimination threshold ; If not, mark the corresponding pixel point as a non-diffusion source point; If so, mark the corresponding pixel point as a diffusion source point, and for each diffusion source point, start again from and sample pixel points in sequence along the unit normal direction vector to form a new sampling sequence: ; For each sampling point , calculate its diffusion weight: ; In the formula: is the standard deviation of the diffusion kernel, is the exponential function; Multiply the boundary mutation intensity by the weight and accumulate it into the boundary-guided diffusion map . The update formula is: ; Exemplarily, assume that the pixel point , its normal vector is [0.6, 0.8]. After sampling e = 5 points outward, the calculated gradient sequence is {20, 25, 40, 30, 15}, then , if , then it is marked as a diffusion source point.

[0027] In implementation, the process of segmenting the second local positive lesion mask from includes: Set the lesion possibility threshold ; For each pixel in , perform the following pixel discrimination: ; In the formula: means that the pixel is marked as 1 in and its value in the boundary-guided diffusion map is greater than or equal to the lesion possibility threshold ; means that the pixel does not meet or does not meet . 1 represents the second target pixel, 0 represents the non-second target pixel, The typical value range is 5 - 20, which is set according to the experience of the blurring degree of the lesion boundary in the image; Segment the area formed by the second target pixels obtained after discrimination from and mark it as the second local positive lesion mask ; By performing information diffusion along the normal direction of the boundary through the image structure mutation points, the blurred boundary information blocked by saliva can be restored, thereby achieving more refined lesion restoration, making up for the segmentation omission, and significantly improving the detection rate, segmentation coherence, and structural integrity of early blurred lesions in oral squamous cell carcinoma images.

[0028] Please refer toFigure 2 , based on the same inventive concept, the second aspect of the present invention provides an image processing system for oral squamous cell carcinoma. For the details not described in this embodiment, please refer to the relevant parts in Embodiment 1. The system includes: An acquisition module for acquiring a first oral image , and a second oral image after copying ; A primary segmentation module for inputting into a pre-trained first semantic segmentation model to obtain a saliva mask ; and inputting into a pre-trained second semantic segmentation model for primary segmentation to obtain a first local positive lesion mask ; A merging module for obtaining a transition mask of the overlapping part generated after performing an AND operation on and ; ; A secondary segmentation module for calculating the boundary mutation intensity, constructing a boundary-guided diffusion map according to the boundary mutation intensity , using to segment a second local positive lesion mask from , and merging with to obtain an overall lesion mask . .

[0029] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired or wireless network. The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that contains one or more collections of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0030] In several embodiments provided by the present invention, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only one way, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0031] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0032] In addition, in each embodiment of the present invention, the functional units can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0033] Some of the data in the above formula are calculated by removing the dimension and taking their numerical values. The formula is a formula that is closest to the actual situation obtained through software simulation of a large amount of collected data. The preset parameters and preset thresholds in the formula are set by those skilled in the art according to the actual situation or obtained through simulation of a large amount of data.

[0034] The above embodiments are only used to illustrate the technical method of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical method of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical method of the present invention.

Claims

1. An image processing method for oral mucosal squamous cell carcinoma, characterized in that Including: Obtain a first oral cavity image , and copy the second oral cavity image after copying ; Input into a pre-trained first semantic segmentation model to obtain a saliva mask ; and input into a pre-trained second semantic segmentation model for initial segmentation to obtain a first local positive lesion mask ; Obtain the transitional mask that generates the overlapping part after performing the AND operation on and ; ; Calculate the boundary mutation intensity and construct a boundary-guided diffusion map based on the boundary mutation intensity , and use to split out the second local positive lesion mask from , and merge with to obtain the overall lesion mask . .

2. The image processing method for oral mucosal squamous cell carcinoma according to claim 1, wherein, The training method of the first semantic segmentation model is as follows: Obtain a historical oral image sample set, and divide the historical oral image sample set into a first semantic segmentation training set and a first semantic segmentation test set. The historical oral image sample set includes multiple grayscale sample oral images and corresponding first annotation labels; Among them, the first annotation label includes a first positive label and a first negative label. The first positive label is represented as a saliva mask , and the first negative label represents a non-saliva mask; Initialize the first convolutional neural network, and use the in the first semantic segmentation training set as the input of the first convolutional neural network, and use the first annotation label as the output to train the first convolutional neural network to obtain the first semantic segmentation network to be verified; Using the first semantic segmentation test set to verify the first semantic segmentation network, and outputting the first semantic segmentation network with a test error less than or equal to the preset test error threshold as the trained first semantic segmentation model.

3. The image processing method for oral mucosal squamous cell carcinoma according to claim 2, wherein The acquisition method of the first annotation label is as follows: Obtain the standard lesion image and the sample oral image , and from the saliva region obtain the first color feature, and obtain the second color feature of each pixel in; Wherein, the first color feature includes a target saturation interval and a target lightness interval, and the second color feature includes a saturation value and a lightness value; Compare the second color feature of each pixel in with the first color feature to screen out the first target pixels and non-first target pixels in ; Wherein: represents a sample oral image at the color features of the pixel, including the saturation value and the lightness value; is the first target pixel, represents a non-first target pixel, represents a target saturation interval, represents a target lightness interval; All pixel regions formed by the connection of the first target pixels are used as the saliva mask , and are labeled as the first positive label. In addition, all pixel regions formed by the connection of non-first target pixels are used as non-saliva masks and are labeled as the first negative label; Repeat the above steps until all in the historical oral image sample set are labeled, and all first annotation labels are obtained.

4. An image processing method for oral mucosal squamous cell carcinoma according to claim 3, characterized in that, The acquisition method of the target saturation interval is as follows: Convert the standard lesion image from the RGB color space to the HSV color space to obtain a saturation image where the saturation range of each pixel is [0, 1]; Extract from the saturation values corresponding to the pixels belonging to the saliva region and the lesion region respectively to form a first saturation set and a second saturation set ; The saturation range [0, 1] is evenly divided into K saturation sub-intervals, and each saturation sub-interval is defined as: , is the i-th saturation sub-interval; Statistically analyze the first saturation set the number of pixels falling into each saturation sub-interval, denoted as , and statistically analyze the second saturation set the number of pixels falling into each saturation sub-interval, denoted as ; According to and construct a first saturation histogram and a second saturation histogram ; Arbitrarily select two consecutive saturation sub - intervals as the saturation sub - interval combination , where \(0\leq m\leq n < K\), and define the first error term and the second error term : Among them, the first error term The calculation method is as follows: ; Among them, the second error term The calculation method is as follows: ; According to the first error term and the second error term construct the total discriminant error loss function : ; In the formula: and are the weighting coefficients of the first error term and the second error term respectively, ; Traverse all combinations of saturation sub - intervals, and for each combination of saturation sub - intervals, calculate its corresponding total discriminant error loss function as well. ; For all combinations of saturation sub-intervals corresponding perform ascending sorting, and take the saturation sub-interval combination corresponding to the first one in the sorting as the target saturation interval .​ 5. The image processing method for oral mucosal squamous cell carcinoma according to claim 4, wherein The training method of the second semantic segmentation model is as follows: Obtain a historical oral image sample set, and divide the historical oral image sample set into a second semantic segmentation training set and a second semantic segmentation test set. The historical oral image sample set further includes a second annotation label corresponding to the sample oral image ; Among them, the second annotation label includes a second positive label and a second negative label, and the second positive label is represented as a first local positive lesion mask , and the second negative label represents a non-first local positive lesion mask; Initialize the second convolutional neural network, and use the in the second semantic segmentation training set as the input of the second convolutional neural network, and use the second annotation label as the output to train the second convolutional neural network to obtain the second semantic segmentation network to be verified; Using the second semantic segmentation test set to verify the second semantic segmentation network, and outputting the second semantic segmentation network with a test error less than or equal to the preset test error threshold as the trained second semantic segmentation model.

6. The image processing method for oral squamous cell carcinoma according to claim 5, wherein The acquisition method of the second annotation label is as follows: Perform window partitioning on the grayscale sample oral cavity image in a sliding window manner to obtain a set of local window region blocks , where the sliding step size is 1, the window size is W×W, and Z and W are positive integers; Obtain a standard lesion image The first texture feature of the lesion area, and The second texture feature of each local window area in Wherein, both the first texture feature and the second texture feature are statistical feature sets of the gray-level co-occurrence matrix, including multiple statistical features, including energy, average contrast, and entropy; Calculate the texture similarity between the first texture feature of the lesion area and the second texture feature of the local window area , and its calculation formula is as follows: ; Wherein: is the eigenvalue of the r-th statistical feature in the first texture feature of the lesion area, is the eigenvalue of the r-th statistical feature in the second texture feature of the local window area, is the total number of statistical features, is a positive real number; Compare the texture similarity with a preset texture similarity threshold. If the texture similarity is greater than or equal to the texture similarity threshold, use the corresponding local window area as the first local positive lesion mask and mark it as the second positive label. Otherwise, use the corresponding local window area as a non-first local positive lesion mask and mark it as the second negative label; Repeat the above steps until all in the historical oral image sample set are labeled, and all second annotation labels are obtained.

7. A method for processing an image of oral mucosal squamous cell carcinoma according to claim 6, characterized in that Calculating the boundary mutation intensity includes: For each pixel in , perform the following operations: For a pixel point Calculate the component of the Sobel gradient in the x - direction and the component in the y - direction : ; ; In the formula: represents the partial derivative of the image in the horizontal x direction, represents the partial derivative of the image in the vertical y direction; Combine the component with the component to form a gradient vector , and calculate its counterclockwise perpendicular vector according to the gradient vector as the boundary normal direction vector: ; Unitize the boundary normal direction vector to obtain the unit normal direction ; Starting from as the starting point, sample e pixel points in sequence along the unit normal direction to form a sampling sequence: ; For each sampling point , calculate the Sobel gradient magnitude on the original first oral image : ; Form a gradient sequence from the gradient magnitudes of e pixel points , and calculate the maximum value of the first-order difference of this gradient sequence, which is defined as the boundary mutation intensity : .

8. A method for processing an image of oral mucosal squamous cell carcinoma according to claim 7, characterized in that, Constructing a boundary-guided diffusion map according to the boundary mutation intensity , including: Construct a two-dimensional floating-point image with exactly the same size as the original first oral cavity image , and initialize all its pixel positions to zero values: ; Traverse the transition mask for all pixel points and determine each pixel point to see if the boundary mutation intensity is greater than the preset structural mutation discrimination threshold ; Otherwise, mark the corresponding pixel point as a non-diffusion source point; If so, mark the corresponding pixel point as a diffusion source point, and for each diffusion source point, re-start with as the starting point and sample successively pixel points to form a new sampling sequence: ; For each sampling point , calculate its diffusion weight: ; In the formula: is the standard deviation of the diffusion kernel, is the exponential function; The boundary mutation intensity , according to the weight is weighted and accumulated onto the boundary-guided diffusion map . The update formula is as follows: .

9. The image processing method for oral mucosal squamous cell carcinoma according to claim 8, wherein The said extraction from to split out the second partial positive lesion mask , including: Set the lesion possibility threshold ; For each pixel in perform the following pixel discrimination: ; Wherein: represents a pixel is marked as 1 in and has a value greater than or equal to the lesion likelihood threshold in the boundary-guided diffusion map ; ; represents a pixel that does not satisfy or does not satisfy , where 1 represents a second target pixel and 0 represents a non-second target pixel; The region formed by the second target pixels obtained after differentiation is separated from and marked as the second local positive lesion mask .

10. An image processing system for oral mucosal squamous cell carcinoma, which is implemented based on the image processing method for oral mucosal squamous cell carcinoma described in any one of claims 1-9, characterized in that Including: An acquisition module for acquiring a first oral cavity image , and a second oral cavity image after copying ; ; An initial segmentation module for inputting into a pre-trained first semantic segmentation model to obtain a saliva mask ; and input it into a pre-trained second semantic segmentation model for initial segmentation to obtain a first local positive lesion mask ;​ A merging module for obtaining and a transitional mask of the overlapping part generated after performing an AND operation ; The secondary segmentation module is used to calculate the boundary mutation intensity and construct a boundary-guided diffusion map according to the boundary mutation intensity. , and use to segment out the second local positive lesion mask , and merge with to obtain the overall lesion mask.

Citation Information

Patent Citations

  • Medical ultrasonic image recognition method and device and storage medium

    CN114757953A

  • Image diagnosis system for mutual recognition of medical examination results

    CN117635616A

  • Image processing apparatus, image processing method, and non-transitory computer-readable storage medium

    US20180089529A1

  • Image segmentation method and apparatus, electronic device, and storage medium

    WO2021184972A1