Image segmentation method, device, electronic device, computer equipment and storage medium

By training the CE-Net model, and using a fully trained tongue segmentation model to segment the tongue images from different data sources, the problem of lack of generalization in tongue image segmentation in the prior art is solved, and higher segmentation accuracy and generalization are achieved.

CN113674282BActive Publication Date: 2025-06-27ZHEJIANG YISHAN SMART MEDICAL RES CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110775412.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-07-08
Publication Date
2025-06-27
Estimated Expiration
2041-07-08

AI Technical Summary

Technical Problem

The prior art lacks generalization when segmenting the original tongue image of different data sources, resulting in low accuracy of tongue diagnosis results.

Method used

By obtaining training samples, including images containing tongue objects of different angles and binary images obtained by pre-notation, the CE-Net model is used for training to obtain a complete tongue segmentation model. Then, the tongue image to be detected is input to the model to be detected, and the tongue object is segmented.

Benefits of technology

Accurate segmentation of tongue images acquired under different acquisition environments is achieved, and the generalization and accuracy of tongue image segmentation is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113674282B_ABST
    Figure CN113674282B_ABST
Patent Text Reader

Abstract

The present application relates to an image segmentation method, apparatus, electronic device, computer device, and storage medium. By obtaining training samples, where the training samples include training tongue images and binary images obtained by pre-labeling the training tongue images, training a CE-Net model using the training samples to obtain a fully trained tongue segmentation model, then obtaining a tongue image to be detected, and inputting the tongue image to be detected into the tongue segmentation model to obtain the tongue image in the tongue image to be detected. The image segmentation method provided by the present application trains the CE-Net model to obtain a fully trained tongue segmentation model, and segments the tongue object in the tongue image to be detected through the fully trained tongue segmentation model, which can achieve accurate segmentation of tongue images obtained in different acquisition environments, thereby improving the generalization and accuracy of tongue image segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing, and particularly to an image segmentation method, apparatus, electronic device, computer device, and storage medium. Background Art

[0002] As one of the important indicators of traditional Chinese medicine tongue diagnosis, tongue diagnosis can evaluate the physical condition of the patient by observing and detecting the shape, color, and moisture of the tongue coating of the patient, and then judge the nature and severity of the disease. Therefore, tongue diagnosis can be performed by analyzing the collected original tongue image. Due to the diversity of the original tongue images, there are often many interference factors in the obtained original tongue images that affect the accuracy of the tongue diagnosis results. For example, the original tongue images also contain other facial regions such as the lips and oral cavity, there are large differences in the illumination of different original tongue images, and the shooting angles of the tongues in different original tongue images are different. Therefore, it is necessary to extract the tongue region from the original tongue image to improve the accuracy of tongue diagnosis.

[0003] Current tongue image segmentation technologies often obtain original tongue images from a closed acquisition environment, which have many restrictions on the data source and can only perform image segmentation on the front of the tongue coating. Therefore, there is a lack of generalization in tongue image segmentation for original tongue images from different data sources. For the problem of lack of generalization in tongue image segmentation for original tongue images from different data sources in the related art, no effective solution has been proposed yet. Summary of the Invention

[0004] In this embodiment, an image segmentation method, apparatus, electronic device, computer device, and storage medium are provided to solve the problem of lack of generalization in tongue image segmentation for original tongue images from different data sources in the related art.

[0005] In the first aspect, in this embodiment, an image segmentation method is provided for segmenting a tongue object in an image. The method includes:

[0006] Obtain training samples, where the training samples include training tongue images and binary images obtained by pre-annotating the training tongue images;

[0007] Train a CE-Net model using the training samples to obtain a trained tongue segmentation model;

[0008] Obtain a tongue image to be detected;

[0009] Input the tongue image to be detected into the tongue segmentation model to obtain the tongue image in the tongue image to be detected.

[0010] In some of these embodiments, the training tongue images are images containing tongue objects at different angles.

[0011] In some of these embodiments, if the output result of the CE-Net model meets the preset iteration criteria, the CE-Net model is trained using the training samples to obtain a fully trained tongue segmentation model, including:

[0012] The CE-Net model is iteratively trained using the training samples to obtain CE-Net models at multiple training stages;

[0013] The CE-Net model that meets the preset iteration error condition among the CE-Net models at the multiple training stages is used as the fully trained tongue segmentation model.

[0014] In some of these embodiments, training the CE-Net model using the training samples to obtain a fully trained tongue segmentation model further includes:

[0015] If the output result of the CE-Net model does not meet the preset iteration criteria, the parameters of the CE-Net are adjusted according to the preset parameter interval value until the output result of the adjusted CE-Net model meets the preset iteration criteria, and then the CE-Net model is iteratively trained using the training samples.

[0016] In some of these embodiments, the tongue segmentation model includes a feature encoder, a context extractor, and a feature decoder. The tongue image to be detected is input into the fully trained tongue segmentation model, and the tongue object in the tongue image to be detected is segmented to obtain the tongue image of the tongue image to be detected, including:

[0017] The tongue image to be detected is input into the tongue segmentation model, and after being processed by the feature encoder, the context extractor, and the feature decoder, a binary image containing the tongue edge region of the tongue image to be detected is obtained;

[0018] The tongue image is obtained according to the binary image containing the tongue edge region of the tongue image to be detected.

[0019] In some of these embodiments, the feature encoder is a Res-Net34 encoder.

[0020] Second, in this embodiment, an image segmentation device is provided for segmenting tongue objects in an image, including: a first acquisition module, a training module, a second acquisition module, and a segmentation module, where:

[0021] The first acquisition module is configured to acquire training samples, where the training samples include training tongue images and binary images obtained by pre-annotating the training tongue images;

[0022] The training module is configured to train the CE-Net model using the training samples to obtain a trained and complete tongue segmentation model;

[0023] The second acquisition module is configured to acquire a tongue image to be detected;

[0024] The segmentation module is configured to input the tongue image to be detected into the tongue segmentation model to obtain the tongue object in the tongue image to be detected.

[0025] In a third aspect, an electronic device is provided in this embodiment, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the image segmentation method described in the first aspect above is implemented.

[0026] In a fourth aspect, a computer device is provided in this embodiment, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the image segmentation method described in the first aspect above is implemented.

[0027] In a fifth aspect, a storage medium is provided in this embodiment, on which a computer program is stored. When the program is executed by a processor, the image segmentation method described in the first aspect above is implemented.

[0028] Compared with the related art, in the image segmentation method, device, electronic device, computer device, and storage medium provided in this embodiment, by acquiring training samples, where the training samples include training tongue images and binary images obtained by pre-annotating the training tongue images, training the CE-Net model using the training samples to obtain a trained and complete tongue segmentation model, then acquiring a tongue image to be detected, and inputting the tongue image to be detected into the tongue segmentation model to obtain the tongue image in the tongue image to be detected. The image segmentation method provided in this application trains the CE-Net model to obtain a trained and complete tongue segmentation model, and segments the tongue object in the tongue image to be detected through the trained and complete tongue segmentation model, which can achieve accurate segmentation of tongue images obtained in different acquisition environments, thereby improving the generalization and accuracy of tongue image segmentation.

[0029] Details of one or more embodiments of this application are set forth in the following drawings and description to make other features, objects, and advantages of this application more concise and understandable. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] The accompanying drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0031] Figure 1 is a hardware structure block diagram of a terminal of an image segmentation method in the related art;

[0032] Figure 2 is a flowchart of the image segmentation method of this embodiment;

[0033] Figure 3 is a flowchart of another image segmentation method of this embodiment;

[0034] Figure 4 is a schematic structural diagram of the image segmentation device of this embodiment;

[0035] Figure 5 is a schematic structural diagram of the electronic device of this embodiment;

[0036] Figure 6 is a schematic structural diagram of the computer device of this embodiment. Detailed implementation manners

[0037] To understand the purpose, technical solution and advantages of the present application more clearly, the present application will be described and explained below with reference to the accompanying drawings and embodiments.

[0038] Unless otherwise defined, technical terms or scientific terms involved in this application shall have the ordinary meanings understood by those with ordinary skills in the technical field to which this application belongs. In this application, words such as "a", "an", "one kind", "the", "these", etc. do not indicate a limitation in quantity, and they can be singular or plural. The terms "include", "comprise", "have" and any variants thereof involved in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product or device including a series of steps or modules (units) is not limited to the listed steps or modules (units), but may include unlisted steps or modules (units), or may include other steps or modules (units) inherent in these processes, methods, products or devices. The terms "connect", "be connected", "couple" and other similar words involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "plurality" involved in this application means two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, and B exists alone. Usually, the character " / " means that the objects associated before and after are in an "or" relationship. The terms "first", "second", "third", etc. involved in this application are only used to distinguish similar objects and do not represent a specific sorting of the objects.

[0039] The method embodiment provided in this embodiment can be executed on a terminal, a computer or a similar computing device. For example, when running on a terminal, Figure 1 is the hardware structure block diagram of the terminal of the image segmentation method in this embodiment. As Figure 1 shown, the terminal may include one or more ( Figure 1 only one is shown in the figure) processors 102 and a memory 104 for storing data. Among them, the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA. The above terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above terminal. For example, the terminal may further include more or fewer components than Figure 1 shown in the figure, or have a different configuration from Figure 1 shown.

[0040] The memory 104 can be used to store computer programs, such as software programs and modules of application software, such as the computer program corresponding to the image segmentation method in this embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the terminal through a network. Examples of the above network include but are not limited to the Internet, intranet, local area network, mobile communication network, and combinations thereof.

[0041] The transmission device 106 is used to receive or send data via a network. The above network includes a wireless network provided by a communication provider of the terminal. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0042] In this embodiment, an image segmentation method is provided. Figure 2 is a flowchart of the image segmentation method of this embodiment, as Figure 2 shown, the process includes the following steps:

[0043] Step S210, obtain training samples, where the training samples include training tongue images and binary images obtained by pre-annotating the training tongue images.

[0044] Among them, the training tongue image can specifically be an original image of a tongue object containing different angles, volumes, morphologies, and colors provided by different data sources. For example, an image containing the side of the tongue, an image containing the front of the tongue, and an image containing the back of the tongue. The training tongue image can specifically be provided by the hospital or obtained by manual collection using a camera. In addition, due to differences in the shooting environment, the training tongue images from different data sources also have inconsistent lighting and contrast. Therefore, using multiple training tongue images with large differences provided by different data sources as training samples can improve the versatility of the model, thereby better meeting the actual clinical tongue diagnosis requirements.

[0045] Additionally, the training samples also include binary images, which are obtained by manually or algorithmically annotating the tongue edge regions of the corresponding training tongue images. Among them, in the binary image, the original tongue region has a gray value of 255, and other regions in the image are regarded as background regions with a gray value of 0. Specifically, the tongue edge regions in the training tongue images can be pre-annotated through the deep learning image annotation tool labelme to obtain the corresponding binary images. After obtaining the binary image corresponding to the training tongue image, the binary image can be used as the label of its corresponding training tongue image, and together with the training tongue image, it can be used as a training sample to train the tongue segmentation model.

[0046] Step S220: Use the training samples to train the CE-Net model to obtain a well-trained tongue segmentation model.

[0047] Among them, the Context Encoder-Network (CE-Net) model is specifically composed of three parts: a feature encoder module, a context extractor, and a feature decoder module. When processing medical images, a convolutional neural network can be used for image segmentation, blood vessel detection, lung image segmentation, and cell image segmentation of medical images. Compared with the U-net used in current image processing using convolutional neural networks, which will cause the loss of image spatial information due to continuous pooling and a series of convolutional operations, the CE-Net model uses the Deep Residual Network (ResNet) instead of the fully convolutional network U-net as the feature encoder, which can retain the spatial information of the image while extracting the high-level semantic features of the image, thereby improving the accuracy of image segmentation. Additionally, the influence of interference factors such as light spots and shadows brought by light in different shooting environments on CE-Net is relatively low, and directly using the CE-Net model to process the original tongue image can also avoid the segmentation failure caused by the close skin color of the lips and face when using traditional image processing methods for tongue image processing. In addition, the residual multi-core pooling adopted by the CE-Net model can also adapt to the differences in the volume, shape, and color of the tongue in the original tongue image. Therefore, training the CE-Net model to obtain a well-trained tongue segmentation model and using it for tongue image segmentation can improve the accuracy and robustness of tongue image segmentation in different scenarios.

[0048] Specifically, for the CE-Net model whose output results meet the preset iteration criteria, the CE-Net model can be iteratively trained until the CE-Net model meets the iteration error condition, at which point the iterative training is terminated, and the CE-Net model of the current training stage that meets the iteration error condition is used as the trained tongue segmentation model. Among them, the iteration error condition can be an error threshold set in advance according to empirical values. To determine whether the CE-Net model meets the iteration error condition, specifically, during the iterative training process, it can be judged whether the error of the output results of the CE-Net model no longer decreases. If so, it is confirmed that the CE-Net model of the current training stage meets the iteration error condition.

[0049] Additionally, if the CE-Net model of the current training stage does not meet the iteration error condition, the label of the training tongue image in the training sample is refined according to the binary image output by the CE-Net model of the current training stage to form a new binary image, and the new binary image is used as the new label in the training sample to train the CE-Net model in the next training stage. In addition, before training the CE-Net model, if the output results of the CE-Net model do not meet the preset iteration criteria, the parameters of the CE-Net model can be adjusted to improve its ability to segment the tongue edge area from the image. Specifically, the parameters of the model can be dynamically adjusted according to a preset interval value. The parameters of the CE-Net model are dynamically adjusted using the preset interval value until the output results of the CE-Net model meet the preset iteration criteria. Among them, the preset iteration criteria can be that the accuracy of the output results of the CE-Net model is within the range of the accuracy of the manually annotated binary image, or other iteration criteria determined in advance according to actual scenario requirements.

[0050] In addition, to test the generalization of the CE-Net model, the CE-Net model can also be tested using test images outside the training samples after training.

[0051] Step S230, obtain the tongue image to be detected.

[0052] Specifically, the tongue image to be detected can be the original image obtained by the hospital after photographing the tongue of the patient. The original image contains the tongue object of the patient, and may also contain areas such as the patient's oral cavity and lips. Among them, the tongue object contained in the tongue image to be detected may have differences in volume, shape, and color due to factors such as shooting angle and lighting. The tongue image to be detected can be used for subsequent tongue diagnosis using machine learning. Therefore, after obtaining the tongue image to be detected, the tongue object contained in the tongue image to be detected needs to be segmented to obtain a separate tongue object.

[0053] Step S240: Input the tongue image to be detected into the tongue segmentation model to obtain the tongue object in the tongue image to be detected.

[0054] Specifically, the tongue segmentation model is the trained and complete tongue segmentation model obtained by training the CE-Net model in the above step S220. After inputting the tongue image to be detected into the CE-Net model, a binary image output by the CE-Net can be obtained. According to the binary image corresponding to the tongue image to be detected, the tongue object in the tongue image to be detected can be obtained. Further, since the gray values of the background regions of non-tongue objects in the binary image are all set to 0, the binary image corresponding to the tongue image to be detected can be intersected with the tongue image to be detected to obtain the tongue image.

[0055] Compared with the current methods that can only perform tongue image segmentation on the original tongue image based on the front of the tongue coating by color, gray threshold, etc., the tongue segmentation model based on the CE-Net model can be applied to the tongue images to be detected in different acquisition environments, has stronger applicability to different illuminations, and has generalization for the segmentation of tongue objects with different shapes, volumes, and colors obtained by shooting at different angles.

[0056] In the above steps S210 to S240, by obtaining training samples, where the training samples include training tongue images and binary images obtained by pre-labeling the training tongue images, the CE-Net model is trained using the training samples to obtain a trained and complete tongue segmentation model. Then, the tongue image to be detected is obtained, and the tongue image to be detected is input into the tongue segmentation model to obtain the tongue image in the tongue image to be detected. The image segmentation method provided in this application trains the CE-Net model to obtain a trained and complete tongue segmentation model, and segments the tongue object in the tongue image to be detected through the trained and complete tongue segmentation model, which can achieve accurate segmentation of tongue images obtained in different acquisition environments, thereby improving the generalization and accuracy of tongue image segmentation.

[0057] In one embodiment, the training tongue image is an image including tongue objects at different angles.

[0058] Among them, the volumes and shapes of the tongue objects at different angles in the training tongue image are different. Therefore, using an image including tongue objects at different angles to train the CE-Net model can improve the generalization of the model, thus better meeting the needs of actual clinical tongue diagnosis.

[0059] In one embodiment, based on the above step S220, if the output result of the CE-Net model meets the preset iteration criteria, training the CE-Net model using the training samples to obtain a trained and complete tongue segmentation model specifically includes the following steps:

[0060] Step S221: Iteratively train the CE-Net model using training samples to obtain CE-Net models at multiple training stages.

[0061] Specifically, the iterative training process may generate multiple training stages. For each training stage, a corresponding CE-Net model can be obtained, and CE-Net models with different parameters are obtained at different training stages.

[0062] Step S222: Use the CE-Net models at multiple training stages that meet the preset iterative error condition as the trained tongue segmentation model.

[0063] Specifically, if the error of the output result of the CE-Net model at the current training stage no longer decreases compared to the error of the output result of the CE-Net model at the previous training stage, then use the CE-Net model at the current training stage as the trained tongue segmentation model and end the training. In addition, the CE-Net model with the smallest error in the output results among the CE-Net models at multiple training stages can also be used as the tongue segmentation model for the Internet cafe.

[0064] In one embodiment, based on the above step S220, when training the CE-Net model using training samples to obtain the trained tongue segmentation model, the following steps may further be included:

[0065] Step S223: If the output result of the CE-Net model does not meet the preset iterative standard, adjust the parameters of the CE-Net according to the preset parameter interval value until the output result of the adjusted CE-Net model meets the preset iterative standard, and then iteratively train the CE-Net model using training samples.

[0066] In addition, in one embodiment, based on the above step S240, the tongue segmentation model includes a feature encoder, a context extractor, and a feature decoder. Input the tongue image to be detected into the trained tongue segmentation model to segment the tongue object in the tongue image to be detected, and obtain the tongue image of the tongue image to be detected. Specifically, the following steps are included:

[0067] Step S241: Input the tongue image to be detected into the tongue segmentation model. After being processed by the feature encoder, the context extractor, and the feature decoder, obtain a binary image including the tongue edge region of the tongue image to be detected.

[0068] Step S242: Obtain the tongue image according to the binary image including the tongue edge region of the tongue image to be detected.

[0069] Further, in one embodiment, based on the above steps, the feature encoder is a Res-Net34 encoder.

[0070] Specifically, in the Res-Net34 encoder, the first four feature extraction modules of the U-Net encoder are retained, and the average pooling layer and the fully connected layer are removed. Compared with the U-Net encoder, the connection mechanism of Res-Net34 is faster, which can avoid the network degradation problem caused by the increase in the depth of the network layer, and can accelerate the network convergence speed, thereby improving the efficiency of the CE-Net model in processing images. In addition, the dense atrous convolution module DAC (Dense Atrous Convolution) in the context extractor is composed of atrous convolutions in a cascaded manner. When processing tongue images, it has a larger receptive field. By combining atrous convolutions with different dilation rates, more information can be extracted from images of different sizes. The residual multi-kernel pooling RMP (residual multi-kernel pooling) based on spatial pyramid pooling can further encode the multi-scale context features of the objects extracted in the DAC module by using pooling operations of various sizes, so that richer context information can be further obtained from the multi-scale context features output by the DAC module.

[0071] Therefore, by training the CE-Net model to obtain a tongue segmentation model, due to its strong generalization ability, it can recognize tongue images at various angles, and has relatively fewer restrictions on the shooting image environment compared with the current tongue image segmentation methods, so that it can segment tongue images collected in an open environment.

[0072] The following describes and illustrates this embodiment through preferred embodiments.

[0073] Figure 3 is a flowchart of the image segmentation method of this preferred embodiment. As Figure 3 shown, it includes the following steps:

[0074] Step S310, obtain the original image dataset from the hospital side, artificial camera collection, and network open-source image channels;

[0075] Step S320, pre-annotate the tongue edge region in the original image to obtain a label dataset, where the label in the label dataset is a binary image with a gray value of 255 for the tongue region and a gray value of 0 for the background region;

[0076] Step S330, randomly split the original image dataset and the label dataset into training data and prediction data, where the labels in the training data and prediction data correspond to the original images;

[0077] Step S340: Input the above training data into the CE-Net model for training, and use the prediction data to test it;

[0078] Step S350: Use the output results obtained in the training phase to improve the label data set;

[0079] Step S360: Repeat the above steps S330 to S350 until a CE-Net model that meets the iterative error condition is obtained as the trained tongue segmentation model.

[0080] In the above steps, by using the images containing tongue objects at different angles as the training tongue images, the generalization of the tongue segmentation model for image segmentation of tongue objects with different shapes can be improved. By iteratively training the CE-Net model using the training samples, CE-Net models at multiple training stages are obtained. The CE-Net model that meets the preset iterative error condition among the CE-Net models at multiple training stages is used as the trained tongue segmentation model, thereby improving the accuracy of the tongue segmentation model for tongue image segmentation. Finally, the segmentation of tongue images at different angles is achieved, the limitation on the tongue acquisition environment is reduced, and the generalization and accuracy of tongue image segmentation for different data sources are improved.

[0081] In this embodiment, an image segmentation device is further provided. This device is used to implement the above embodiment and the preferred implementation manners, and those that have been described will not be repeated here. The following terms such as "module", "unit", "sub-unit", etc. can be a combination of software and / or hardware that can achieve a predetermined function. Although the device described in the following embodiments is preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0082] Figure 4 is the structural block diagram of the image segmentation device 40 of this embodiment, as Figure 4 shown, this device includes: a first acquisition module 42, a training module 44, a second acquisition module 46, and a segmentation module 48, where:

[0083] The first acquisition module 42 is used to acquire training samples, where the training samples include training tongue images and binary images obtained by pre-labeling the training tongue images;

[0084] The training module 44 is used to train the CE-Net model using the training samples to obtain a trained tongue segmentation model;

[0085] The second acquisition module 46 is used to acquire the tongue image to be detected;

[0086] The segmentation module 48 is used to input the tongue image to be detected into the tongue segmentation model to obtain the tongue image in the tongue image to be detected.

[0087] In one embodiment, the training tongue image is an image containing tongue objects at different angles.

[0088] In one embodiment, the training module 44 is further configured to iteratively train the CE-Net model using the training samples to obtain CE-Net models at multiple training stages, and use the CE-Net model that meets the preset iteration error condition among the CE-Net models at multiple training stages as the trained tongue segmentation model.

[0089] In one embodiment, if the output result of the CE-Net model does not meet the preset iteration standard, the training module 44 is further configured to adjust the parameters of the CE-Net at preset parameter interval values until the output result of the adjusted CE-Net model meets the preset iteration standard, and then iteratively train the CE-Net model using the training samples.

[0090] In one embodiment, the segmentation module 48 is further configured to input the tongue image to be detected into the tongue segmentation model, and after being processed by the feature encoder, the context extractor, and the feature decoder, obtain a binary image of the tongue edge region including the tongue image to be detected, and obtain the tongue image according to the binary image of the tongue edge region including the tongue image to be detected.

[0091] In one embodiment, the feature encoder is a Res-Net34 encoder.

[0092] The above image segmentation device 40 obtains training samples, where the training samples include training tongue images and binary images obtained by pre-labeling the tongue edge regions of the training tongue images. The training tongue images contain tongue objects, and uses the training samples to train the CE-Net model to obtain a trained tongue segmentation model. Then, it obtains the tongue image to be detected, where the tongue image to be detected contains tongue objects, and inputs the tongue image to be detected into the tongue segmentation model to segment the tongue objects in the tongue image to be detected, and obtain the tongue image in the tongue image to be detected. The image segmentation method provided in this application trains the CE-Net model to obtain a trained tongue segmentation model, and segments the tongue objects in the tongue image to be detected through the trained tongue segmentation model, which can achieve accurate segmentation of tongue images obtained in different acquisition environments, thereby improving the generalization and accuracy of tongue image segmentation.

[0093] It should be noted that each of the above modules can be a functional module or a program module, and can be implemented either by software or by hardware. For the modules implemented by hardware, each of the above modules can be located in the same processor; or each of the above modules can be located in different processors in any combined form.

[0094] In this embodiment, an electronic device is also provided. As Figure 5 shown, it includes a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0095] Optionally, the above electronic device may further include a transmission device and an input / output device. Among them, the transmission device is connected to the above processor, and the input / output device is connected to the above processor.

[0096] Optionally, in this embodiment, the above processor may be configured to execute the following steps through a computer program:

[0097] Obtain training samples, where the training samples include training tongue images and binary images obtained by pre-annotating the training tongue images;

[0098] Use the training samples to train the CE-Net model to obtain a trained and complete tongue segmentation model;

[0099] Obtain the tongue image to be detected;

[0100] Input the tongue image to be detected into the tongue segmentation model to obtain the tongue image in the tongue image to be detected.

[0101] In one embodiment, the training tongue image is an image containing tongue objects at different angles.

[0102] In one embodiment, when the processor executes the computer program, the following steps are also implemented:

[0103] Iteratively train the CE-Net model using the training samples to obtain CE-Net models at multiple training stages;

[0104] Use the CE-Net model that meets the preset iteration error condition among the CE-Net models at multiple training stages as the trained and complete tongue segmentation model.

[0105] In one embodiment, when the processor executes the computer program, the following steps are also implemented:

[0106] If the output result of the CE-Net model does not meet the preset iteration standard, the parameters of the CE-Net are adjusted according to the preset parameter interval value until the output result of the adjusted CE-Net model meets the preset iteration standard, and then the CE-Net model is iteratively trained using the training samples.

[0107] In one embodiment, when the processor executes the computer program, the following steps are further implemented:

[0108] Input the tongue image to be detected into the tongue segmentation model, and after being processed by the feature encoder, the context extractor, and the feature decoder, obtain a binary image including the tongue edge region of the tongue image to be detected;

[0109] Obtain the tongue image according to the binary image including the tongue edge region of the tongue image to be detected.

[0110] In one embodiment, the feature encoder is a Res-Net34 encoder.

[0111] The above electronic device obtains training samples, where the training samples include training tongue images and binary images obtained by pre-labeling the tongue edge regions of the training tongue images. The training tongue images include tongue objects. The CE-Net model is trained using the training samples to obtain a trained and complete tongue segmentation model. Then, a tongue image to be detected is obtained, where the tongue image to be detected includes a tongue object. The tongue image to be detected is input into the tongue segmentation model, and the tongue object in the tongue image to be detected is segmented to obtain the tongue image in the tongue image to be detected. The image segmentation method provided in this application trains the CE-Net model to obtain a trained and complete tongue segmentation model, and segments the tongue object in the tongue image to be detected through the trained and complete tongue segmentation model, which can achieve accurate segmentation of tongue images obtained in different acquisition environments, thereby improving the generalization and accuracy of tongue image segmentation.

[0112] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and alternative embodiments, and will not be repeated in this embodiment.

[0113] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 6As shown. The computer device includes a processor, a memory, a network interface, and a database connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store a set of preset configuration information. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, the above-mentioned image segmentation method is implemented.

[0114] In one embodiment, a computer device is provided. The computer device may be a terminal. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, an image segmentation method is implemented. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device may be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the computer device housing, or an external keyboard, touchpad, or mouse, etc.

[0115] In addition, in combination with the image segmentation method provided in the above embodiments, a storage medium may also be provided in this embodiment to implement. A computer program is stored on the storage medium; when the computer program is executed by the processor, any one of the above-mentioned image segmentation methods in the embodiments is implemented.

[0116] It should be understood that the specific embodiments described here are only used to explain this application, rather than to limit it. According to the embodiments provided in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of this application.

[0117] Obviously, the accompanying drawings are only some examples or embodiments of the present application. For those of ordinary skill in the art, the present application can also be applied to other similar situations based on these drawings without creative work. Additionally, it can be understood that although the work done during this development process may be complex and time-consuming, for those of ordinary skill in the art, certain design, manufacturing, or production changes based on the technical content disclosed in the present application are only conventional technical means and should not be regarded as insufficient disclosure of the present application.

[0118] The term "embodiment" in this application means that the specific features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of the present application. The phrase appears in various positions in the specification and does not necessarily mean the same embodiment, nor does it mean being independent or alternative to other embodiments and mutually exclusive. Those of ordinary skill in the art can clearly or implicitly understand that the embodiments described in this application can be combined with other embodiments without conflict.

[0119] The above-described embodiments merely represent several implementation manners of the present application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of patent protection. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. An image segmentation method for segmenting the tongue object in an image, characterized in that, Including: Obtain training samples, where the training samples include training tongue images and binary images obtained by pre-annotating the training tongue images. The training tongue images include original images of tongue objects of different colors, and the training tongue images are captured under various lighting conditions; Use the training samples to train the CE-Net model to obtain a trained and complete tongue segmentation model; Obtain a tongue image to be detected; Input the tongue image to be detected into the tongue segmentation model. After processing by a deep residual network, a dense dilated convolution module, and a residual multi-kernel pooling, obtain a binary image of the tongue edge region including the tongue image to be detected. Among them, the tongue segmentation model includes a deep residual network, a dense dilated convolution module, and a residual multi-kernel pooling. The deep residual network is used to extract high-level semantic features of the tongue image to be detected and retain the spatial information of the tongue image to be detected. The dense dilated convolution module is used to extract multi-scale context features, and the residual multi-kernel pooling is used to encode the multi-scale context features; Perform an intersection process on the binary image and the tongue image to be detected to obtain the tongue image in the tongue image to be detected.

2. The image segmentation method according to claim 1, wherein The training tongue image is an image including tongue objects at different angles.

3. The image segmentation method according to claim 1, characterized in that If the output result of the CE-Net model meets the preset iteration standard, use the training samples to train the CE-Net model to obtain a trained and complete tongue segmentation model, including: Use the training samples to perform iterative training on the CE-Net model to obtain CE-Net models at multiple training stages; Use the CE-Net model that meets the preset iteration error condition among the CE-Net models at the multiple training stages as the trained and complete tongue segmentation model.

4. The image segmentation method according to claim 1, characterized in that Using the training samples to train the CE-Net model to obtain a trained and complete tongue segmentation model further includes: If the output result of the CE-Net model does not meet the preset iteration standard, adjust the parameters of the CE-Net according to the preset parameter interval value until the output result of the adjusted CE-Net model meets the preset iteration standard, and then use the training samples to perform iterative training on the CE-Net model.

5. The image segmentation method according to claim 1, wherein The tongue segmentation model includes a feature encoder, a context extractor, and a feature decoder; the feature encoder is a Res-Net34 encoder.

6. An image segmentation device for segmenting the tongue object in an image, characterized in that, Including: A first acquisition module, a training module, a second acquisition module, and a segmentation module, where: The first acquisition module is used to obtain training samples, where the training samples include training tongue images and binary images obtained by pre-annotating the training tongue images. The training tongue images include original images of tongue objects of different colors, and the training tongue images are captured under various lighting conditions; The training module is used to use the training samples to train the CE-Net model to obtain a trained and complete tongue segmentation model; The second acquisition module is used to obtain a tongue image to be detected; The segmentation module is configured to input the tongue image to be detected into the tongue segmentation model. After processing by a deep residual network, a dense dilated convolution module, and a residual multi-kernel pooling, a binary image of the tongue edge region containing the tongue image to be detected is obtained. The tongue segmentation model includes a deep residual network, a dense dilated convolution module, and a residual multi-kernel pooling. The deep residual network is used to extract high-level semantic features of the tongue image to be detected and retain the spatial information of the tongue image to be detected. The dense dilated convolution module is used to extract multi-scale context features. The residual multi-kernel pooling is used to encode the multi-scale context features. The binary image is intersected with the tongue image to be detected to obtain the tongue image in the tongue image to be detected.

7. The image segmentation device according to claim 6, wherein The training module is further configured to: Iteratively train the CE-Net model using the training samples to obtain CE-Net models at multiple training stages; Use the CE-Net model that meets the preset iterative error condition among the CE-Net models at the multiple training stages as the trained tongue segmentation model.

8. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to run the computer program to execute the image segmentation method according to any one of claims 1 to 5.

9. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor is configured to run the computer program to execute the image segmentation method according to any one of claims 1 to 5.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the image segmentation method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Tongue image segmentation method based on context-aware residual network

    CN110729045A

  • Tongue body automatic segmentation method based on U-net model

    CN111260619A