An auricle reference segmentation method and system
By introducing text guidance and angle transformation techniques in auricle image segmentation, multi-stage fusion and angle transformation of text features and visual features are achieved, solving the problem of mask overlap and poor segmentation in auricle segmentation, and improving the segmentation effect and robustness.
Patent Information
- Application Number
- CN202510140742.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-08
AI Technical Summary
In the process of segmenting auricle image, the prior art results are poor in segmentation effects, and masks overlap each other, making it difficult to achieve accurate segmentation due to factors such as dense auricle parts, similar textures, and changes in light and postures.
The auricle reference segmentation method based on text guidance and angle transformation is adopted. Through the combination of text encoding module, visual encoding module, visual decoding module and angle transformation module, multi-stage fusion and angle transformation of text features and visual features are realized to generate high-quality segmentation masks.
The robustness and accuracy of the auricle segmentation model in complex scenarios is improved, the segmentation ability of the fine-grained structure of the ear is enhanced, the mask overlap phenomenon is reduced, and the precise segmentation of key parts such as the throbbing and tragus is achieved.
Smart Images

Figure CN119579905B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of auricle reference segmentation, and particularly to an auricle reference segmentation method and system. Background Art
[0002] The human ear, as a biometric feature, is considered a highly reliable way of identity recognition because of its relatively stable shape, significant individual differences, and insensitivity to changes in expression and age. It can be used as an auxiliary means of identity recognition when face biometric features are restricted and inconvenient for face recognition. In addition, in the medical field, it is believed that by observing the color, shape, texture, etc. of the auricle and combining specific changes in reflection points, it is possible to assist in judging the functions of internal organs or the health status of the body. In the process of intelligentization combined with AI technology, instance segmentation based on auricle images and image feature analysis of regions of interest are very important links. In the process of image segmentation, there is a great demand for a reference segmentation method that incorporates language or text information to guide the segmentation results in actual application scenarios of human-computer interaction.
[0003] However, the "auricle inclination" (i.e., the rotation angle of the outer ear on the side of the head) of different individuals shows a variety from 0 degrees to 70 degrees, and the auricle includes roughly divided parts such as the helix, antihelix, tragus, antitragus, earlobe, cymba conchae, and cavum conchae. When divided in detail, there can be up to dozens of parts. These parts are visually characterized by dense distribution, that is, each structure is closely adjacent, which will lead to the phenomenon of overlapping of each region mask during image segmentation, resulting in poor segmentation effect. Moreover, the auricle parts have similar texture features, which further exacerbates the difficulty of the segmentation task. In addition, in actual segmentation tasks, it may be affected by various interference factors. For example, changes in illumination will cause the ear features to be blurred, and changes in the pose of the human ear may lead to deformation of the auricle shape. Summary of the Invention
[0004] In order to solve the above technical problems existing in the prior art, the present invention provides an auricle reference segmentation method, and the technical solution is as follows:
[0005] On the one hand, an auricle reference segmentation method is provided, and the method includes:
[0006] S1. Obtain the ear image to be segmented and the corresponding text description;
[0007] S2. Input the ear image to be segmented and the corresponding text description into an auricle reference segmentation model based on text guidance and angle transformation. The auricle reference segmentation model includes a text encoding module, a text-guided visual encoding module, a text-guided visual decoding module, and an angle transformation module;
[0008] The text encoding module embeds the text description into a high-dimensional word vector to obtain text features ;
[0009] The text-guided visual encoding module realizes the fusion of text features and image features through a four-stage structure. These four stages are connected in series, with the output of the previous stage serving as the input of the next stage to achieve multi-stage feature fusion. Each stage includes a visual encoder, a cross-modal perception module, and an attention gating module. The visual encoder of each stage generates visual features , and the cross-modal perception module cross-modally aligns and fuses the visual features with text features to obtain multi-modal features . Each element in the multi-modal features is weighted by the attention gating module to obtain weighted multi-modal features . The weighted multi-modal features are added element-wise to the original visual features to generate a set of enhanced visual features embedding language information . The enhanced visual features with the smallest scale generated by the last stage are input into the text-guided visual decoding module;
[0010] The text-guided visual decoding module gradually restores the spatial resolution of the image, while further fusing text features and visual features, provides high-quality feature representations for the final segmentation task, and outputs multi-scale features;
[0011] The angle transformation module performs angle transformation on the multi-scale features and outputs a segmentation mask for the region related to the text description.
[0012] Optionally, the text encoding module uses BERT for text encoding, embeds the text description into high-dimensional word vectors, and obtains text features as , where and N represent the number of word vector channels and the number of words contained in the text description, respectively.
[0013] Optionally, the visual encoder adopts the encoder of Swin-Transformer. The input feature map is divided into windows of a fixed size, and the features within each window are modeled through the multi-head self-attention mechanism to capture the relationships in the local region. After the modeling at the window level, the positions of the windows are periodically shifted to achieve global context association across windows. In the specific encoding process, the feature map first passes through a linear embedding module to be mapped into a high-dimensional space, and then layer by layer, feature extraction and non-linear transformation are performed through the sliding window attention module and the multi-layer perceptron to enhance the expression ability of high-level semantic features, and the visual features are output.
[0014] Optionally, the cross-modal perception module uses the visual features as queries, and the text features as keys and values to calculate scaled dot-product attention, realizing the process of supervising visual features with text features, promoting the fusion of multi-modal information. After completing the calculation of scaled dot-product attention, the obtained features and the input visual features are projected again to obtain and , and then a dot product operation is performed to finally obtain the multi-modal features . The process formula is as follows:
[0015] (1)
[0016] (2)
[0017] (3)
[0018] (4)
[0019] (5)
[0020] (6)
[0021] (7)
[0022] where , , , , are projection functions, represents the dot product operation, the key function and the value function are 1×1 convolutions with output channels, the query function is also a 1×1 convolution, and the number of output channels is the same as that of the text features. When calculating , the two-dimensional visual feature map is cut and arranged into a one-dimensional sequence data to adapt to the subsequent attention calculation operation. After completing the attention calculation, the calculated one-dimensional sequence data is restored to the form of a two-dimensional feature map.
[0023] Optionally, the attention gating module learns a set of weight maps based on the multi-modal features to rescale each element in to prevent the multi-modal features from overwhelming the visual features The visual signals in it and allow an adaptive amount of text semantic information flow to the next stage. The specific process is as follows:
[0024] Perform 1×1 convolution and ReLu activation on the multi-modal features Then perform 1×1 convolution and Tanh activation again. In this process, the 1×1 convolution adjusts the number of channels to achieve the recombination and weight allocation of channel information without changing the spatial resolution of the input feature map, while the ReLu and Tanh activation functions can increase the non-linear ability of the model and generate weights for the multi-modal features Multiply the weights with the original feature map through a dot product operation to obtain the weighted multi-modal features .
[0025] Optionally, the text-guided visual decoding module consists of three stages. The input of each stage is the fused features output by the previous stage and the multi-modal features of the corresponding stage in the visual encoding module ;
[0026] In each stage, first, the visual decoder upsamples the feature map of the previous stage through bilinear interpolation, gradually expanding the spatial resolution of the feature map, restoring the spatial information of the original image, and laying a foundation for subsequent image segmentation;
[0027] The upsampled features are added to the multi-modal features of the corresponding stage in the visual encoding module in the channel dimension to further achieve the deep fusion of visual features and text features, make full use of the semantic information in the text features to guide the restoration of visual features, enhance the sensitivity of the segmentation task to details and semantic boundaries, and through the layer-by-layer fusion and upsampling operation of the multi-modal features The features output by the visual decoding module gradually approach the original resolution of the input image and provide a more delicate and semantically clear feature representation for the subsequent angle transformation module;
[0028] During the entire decoding process, the visual decoding module takes the enhanced visual features of the smallest scale output by the visual encoding module as the input, and through three levels of progressive decoding, gradually generates three larger-scale feature maps, which together with the initial input feature map form four-scale feature maps. Finally, the sizes of the four-scale feature maps are consistent with the corresponding features of the visual encoding module Figure 1 to provide structural support for the fusion and segmentation of features, and finally output multi-scale features .
[0029] Optionally, the angle transformation module includes three angle transformation layers, three non-linear projection layers, and one linear projection layer, and receives the multi-scale features output by the text-guided visual decoding module As the input, the final mask prediction is output, and the whole process is shown in the following formula:
[0030] (8)
[0031] (9)
[0032] (10)
[0033] Among them, represents the feature map with the smallest scale of the input, represents the input of each angle transformation layer, represents the concat operation to achieve feature fusion in the channel dimension, and the feature maps of different scales are unified in scale through bilinear interpolation, The process includes a 3×3 convolutional layer, a batch normalization layer and a ReLU activation function, corresponding to the non-linear projection layer, represents the angle transformation layer, represents the final segmentation mask, is a linear projection function that projects the feature onto the final segmentation mask.
[0034] Optionally, the angle transformation layer captures angle information from the input features for angle transformation and dynamically re-parameterizes the convolutional kernel weights, filtering out redundant features through the updated convolutional kernel weights. The process is as follows:
[0035] The feature map is fed into a depthwise separable convolutional layer and an average pooling layer to further extract features, and then input into a linear prediction layer. The linear prediction layer includes a fully connected layer and an activation function, and predicts n angle parameters and the corresponding weight parameters , , according to the coordinates of the points on the feature map, coordinate transformation is performed to obtain , and after angle transformation, the coordinates will make the weights fall on new positions. Bilinear interpolation is used to ensure the continuity of weight changes and prevent weight information loss caused by angle changes. Finally, the updated convolutional kernel weights, weight parameters and the original visual features are fed into a convolutional layer to filter out redundant features and generate new features. The formula is as follows:
[0036] (11)
[0037] (12)
[0038] (13)
[0039] Among them, is the coordinate of a point on the original feature map, is the coordinate after angle transformation, is the inverse matrix of the rotation matrix of the angle affine transformation, is bilinear interpolation, represents the weight of the original static convolution kernel, represents according to the transformed for the original weight to be updated, is the original feature map, is the weighted feature map.
[0040] Optionally, the training process of the auricle reference segmentation model uses instance contrast learning to improve the discriminability of the model for different instances and its robustness to different languages describing the same instance;
[0041] For the instance contrast learning, b text descriptions of an image are sampled. These text descriptions may point to the same instance or different instances in the image. The model will generate b mask predictions based on the image and the b text descriptions, making the masks of the sentences describing the same instance similar and suppressing the overlap of the masks of different instances. The overlap score is obtained by calculating the overlap between every two masks. The overlap score between the i-th and j-th final masks and is calculated by the following formula:
[0042] (14)
[0043] Where, , n is the pixel size of the feature map S, is the normalized value of the i-th segmentation mask at position n. According to the overlap score, the instance contrast loss is calculated and used to enhance the discrimination ability of the model for different instances. The formula is as follows:
[0044] (15)
[0045] Where, represents and whether they correspond to the same instance, is the intersection over union between the i-th mask and the actual annotation, which is used to prevent incorrect mask predictions;
[0046] The loss function of the auricle reference segmentation model uses a joint optimization loss function, combining binary cross-entropy loss, Dice loss and the instance contrast loss Combined, it improves the segmentation accuracy and boundary discrimination ability of the model, expressed as:
[0047] (16)
[0048] Among them, represents the binary cross-entropy loss, represents the Dice loss, represents the instance contrast loss mentioned above.
[0049] On the other hand, an auricle reference segmentation system is provided, and the system includes:
[0050] An acquisition module for acquiring a human ear image to be segmented and a corresponding text description;
[0051] A segmentation module for inputting the human ear image to be segmented and the corresponding text description into an auricle reference segmentation model based on text guidance and angle transformation, and the auricle reference segmentation model includes a text encoding module, a text-guided visual encoding module, a text-guided visual decoding module, and an angle transformation module;
[0052] The text encoding module embeds the text description into a high-dimensional word vector to obtain text features ;
[0053] The text-guided visual encoding module realizes the fusion of text features and image features through a four-stage structure. These four stages are connected in series, and the output of the previous stage is used as the input of the next stage to achieve multi-stage feature fusion. Each stage includes a visual encoder, a cross-modal perception module, and an attention gating module. The visual encoder of each stage generates visual features , and the cross-modal perception module cross-modally aligns and fuses the visual features with the text features to obtain multi-modal features . Each element in the multi-modal features is weighted by the attention gating module to obtain weighted multi-modal features , and added element-wise to the original visual features to generate a set of enhanced visual features embedding language information . The enhanced visual features with the smallest scale generated by the last stage
[0054] are input into the text-guided visual decoding module;
[0055] The angle transformation module performs angle transformation on the multi-scale features and outputs a segmentation mask for the region related to the text description.
[0056] On the other hand, an electronic device is provided. The electronic device includes a processor and a memory. At least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the above-mentioned auricle reference segmentation method.
[0057] On the other hand, a computer-readable storage medium is provided. At least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement the above-mentioned auricle reference segmentation method.
[0058] The beneficial effects brought by the technical solution provided by the present invention at least include:
[0059] 1) By introducing a text guidance mechanism, the present invention combines the high-level semantic information contained in natural language descriptions with the color and texture features of RGB images, realizing a more comprehensive understanding of the ear region. This method is particularly suitable for scenarios where ear features are not obvious or the background is complex, effectively improving the robustness and accuracy of the segmentation model in complex scenarios. Specifically, by using a high-resolution imaging device to collect ear images as input, performing fine-grained mask annotation on the collected images, while using text descriptions to provide high-level semantic information for the ear region and features, extracting visual features from the ear images through a visual encoding module, and encoding the text descriptions through a text encoding module to obtain corresponding semantic embedding representations. Then, using a cross-modal perception module, combining multi-scale image features with text features to obtain multi-modal features, establishing the association between pixel-level features and high-level semantic information, strengthening the segmentation ability of the model. Then, sending the multi-modal features into an attention gating module, performing attention weighting on the multi-modal features, and performing element-wise accumulation with the original visual features. Then, the visual decoding module upsamples the features sent by the visual encoding module, gradually increasing the scale of the feature map. During the upsampling process, multi-scale channel dimension fusion is performed with the multi-modal features of the encoding module, realizing a high degree of alignment between text semantic information and visual semantic information, which is helpful for subsequent fine-grained segmentation.
[0060] 2) The angle transformation module designed by the present invention corrects the deformation characteristics of the ear part of the input image at different tilts and postures, thereby enhancing the pose adaptation ability of the segmentation model. Combining multi-view features with text guidance, the model can more accurately segment diverse auricle regions.
[0061] 3) By designing a text-guided segmentation strategy, the present invention performs fine-grained annotation on the fine-grained structure of the ear, and at the same time introduces an instance contrast learning method to improve the boundary discrimination ability of the model, significantly reducing the mask overlap phenomenon between the dense structures of the ear, thereby achieving precise segmentation of key parts such as the helix and tragus.
[0062] 4) By designing a text-guided human ear image segmentation method, the present invention provides convenience for human-computer interaction and is more suitable for task scenarios with high requirements for human-computer interaction. For example, in the medical field, the positioning of ear acupoints is a skill that requires experience accumulation. The text-guided segmentation method provides clear segmentation results and visual semantic annotations, and can dynamically adjust the segmented target area according to text descriptions. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the following-described accompanying drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts.
[0064] Figure 1 It is a flowchart of a method for auricle reference segmentation provided by an embodiment of the present invention;
[0065] Figure 2 It is a general block diagram of a cross-modal outer ear region segmentation and key point localization method provided by an embodiment of the present invention;
[0066] Figure 3 It is a block diagram of the structure of a cross-modal perception module provided by an embodiment of the present invention;
[0067] Figure 4 It is a block diagram of the structure of an attention gating module provided by an embodiment of the present invention;
[0068] Figure 5 It is a block diagram of the structure of an angle transformation module provided by an embodiment of the present invention;
[0069] Figure 6 It is a block diagram of the structure of an angle transformation layer provided by an embodiment of the present invention;
[0070] Figure 7 It is a schematic diagram of instance contrast learning provided by an embodiment of the present invention;
[0071] Figure 8 It is a block diagram of a system for auricle reference segmentation provided by an embodiment of the present invention;
[0072] Figure 9 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0073] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.
[0074] An embodiment of the present invention provides a method for auricle reference segmentation. This method can be implemented by an electronic device, which can be a terminal or a server. Figure 1 The flowchart of this method is shown as follows, Figure 2 The overall block diagram of this method is shown as follows. The processing flow may include the following steps:
[0075] S1. Obtain the ear image to be segmented and the corresponding text description;
[0076] In an embodiment of the present invention, the ear image to be segmented is collected by a high-definition camera, and the corresponding text description is obtained. The text description of the ear acupoint area has corresponding organ attributes, position attributes or shape attributes, etc. For example, the text description of the concha is: hemispherical acupoint related to the heart, acupoint related to the heart, acupoints under the ear nail boat.
[0077] S2. Input the ear image to be segmented and the corresponding text description into an auricle reference segmentation model based on text guidance and angle transformation. The auricle reference segmentation model includes a text encoding module, a text-guided visual encoding module, a text-guided visual decoding module, and an angle transformation module, as Figure 2 shown;
[0078] The text encoding module embeds the text description into a high-dimensional word vector to obtain text features ;
[0079] The text-guided visual encoding module realizes the fusion of text features and image features through a four-stage structure. These four stages are connected in series, and the output of the previous stage is used as the input of the next stage to achieve multi-stage feature fusion. Each stage includes a visual encoder, a cross-modal perception module, and an attention gating module. The visual encoder of each stage generates visual features (the visual encoders of the four stages will generate visual features of four scales ), and the cross-modal perception module cross-modally aligns and fuses the visual features with the text features To improve the model accuracy, multi-modal feature alignment and fusion are performed in four stages, enabling the text semantic information to effectively supervise the feature extraction process of the visual coding module, and multi-modal features are obtained. , and each element in the multi-modal features is weighted by the attention gating module to obtain weighted multi-modal features , which are added element-wise to the original visual features to generate a set of enhanced visual features embedding language information . The enhanced visual features at the smallest scale generated in the last stage are input into the text-guided visual decoding module;
[0080] The text-guided visual decoding module gradually restores the spatial resolution of the image, further fuses text features and visual features simultaneously, provides high-quality feature representations for the final segmentation task, and outputs multi-scale features;
[0081] The angle transformation module performs angle transformation on the multi-scale features and outputs a segmentation mask for the region related to the text description.
[0082] Optionally, the text encoding module uses BERT for text encoding, embeds the text description into high-dimensional word vectors, and obtains text features as , where and N represent the number of word vector channels and the number of words in the text description respectively.
[0083] Optionally, the visual encoder adopts the encoder of Swin-Transformer. The input feature map is divided into windows of a fixed size, and the features within each window are modeled through the multi-head self-attention mechanism to capture the relationships in local regions. After modeling at the window level, the positions of the windows are periodically shifted to achieve global context association across windows (solving the problem of information isolation in traditional windowing methods). In the specific encoding process, the feature map first passes through a linear embedding module to be mapped into a high-dimensional space, and then layer by layer, feature extraction and non-linear transformation are performed through the sliding window attention module and the multi-layer perceptron to enhance the expression ability of high-level semantic features, and the visual features are output.
[0084] Optionally, as Figure 3 shown, the cross-modal perception module uses the visual features as queries, the text features as keys and values to calculate the scaled dot-product attention, realizes the process of supervising visual features with text features, promotes the fusion of multi-modal information, and after calculating the scaled dot-product attention, the obtained features are combined with the input visual features Perform the projection operation again to obtain and , and then perform a dot product operation to finally obtain the multi-modal feature . The process formula is as follows:
[0085] (1)
[0086] (2)
[0087] (3)
[0088] (4)
[0089] (5)
[0090] (6)
[0091] (7)
[0092] Among them, , , , , are projection functions, represents the dot product operation, the key function and the value function are 1×1 convolutions with output channel numbers, and the query function is also a 1×1 convolution, and the number of output channels is the same as that of the text feature. When calculating , the two-dimensional visual feature map is cut and arranged into a one-dimensional sequence data to adapt to the subsequent attention calculation operation. After the attention calculation is completed, the calculated one-dimensional sequence data is restored to the form of a two-dimensional feature map.
[0093] Optionally, as shown in Figure 4 , the attention gating module learns a set of weight maps based on the multi-modal feature to rescale each element in to prevent the multi-modal feature from overwhelming the visual signal in the visual feature and allowing an adaptive amount of text semantic information to flow to the next stage. The specific process is as follows:
[0094] The multi-modal feature Perform 1×1 convolution and ReLu activation, and then perform 1×1 convolution and Tanh activation again. During this process, the 1×1 convolution adjusts the number of channels to achieve the recombination and weight distribution of channel information without changing the spatial resolution of the input feature map, while the ReLu and Tanh activation functions can increase the nonlinear ability of the model and generate multi-modal features The weights of are obtained, and the weighted multi-modal features are obtained by performing a dot product operation on the original feature map with the weights .
[0095] Optionally, as Figure 2 shown, the text-guided visual decoding module consists of three stages, and the input of each stage is the fused features output by the previous stage and the multi-modal features of the corresponding stage in the visual encoding module ;
[0096] In each stage, first, the visual decoder upsamples the feature map of the previous stage through bilinear interpolation, gradually expanding the spatial resolution of the feature map, restoring the spatial information of the original image, and laying a foundation for subsequent image segmentation;
[0097] The upsampled features are added to the multi-modal features of the corresponding stage in the visual encoding module in the channel dimension to further achieve the deep fusion of visual features and text features, fully utilize the semantic information in the text features to guide the restoration of visual features, enhance the sensitivity of the segmentation task to details and semantic boundaries, and through the multi-modal features layer-by-layer fusion and upsampling operations, the features output by the visual decoding module gradually approach the original resolution of the input image, and provide a more delicate and semantically clear feature representation for the subsequent angle transformation module;
[0098] During the entire decoding process, the visual decoding module uses the enhanced visual features of the smallest scale output by the visual encoding module as the input, and after three levels of decoding, gradually generates three larger-scale feature maps, which together with the initial input feature map form four-scale feature maps. Finally, the sizes of the four-scale feature maps are consistent with the corresponding features of the visual encoding module Figure 1 to provide structural support for the fusion and segmentation of features, and finally output multi-scale features .
[0099] Optionally, as Figure 5 shown, the angle transformation module includes three angle transformation layers, three non-linear projection layers and one linear projection layer, receives the multi-scale features output by the text-guided visual decoding module as the input, and outputs the final mask prediction. The whole process is shown in the following formula:
[0100] (8)
[0101] (9)
[0102] (10)
[0103] Among them, represents the feature map with the smallest input scale, represents the input of each angle transformation layer, represents the concat operation, which realizes feature fusion in the channel dimension. The feature maps of different scales are unified in scale through bilinear interpolation, The process includes a 3×3 convolutional layer, a batch normalization layer, and a ReLU activation function, corresponding to the non-linear projection layer, represents the angle transformation layer, represents the final segmentation mask, is a linear projection function that projects the feature onto the final segmentation mask.
[0104] Optionally, as Figure 6 shown, the angle transformation layer captures angle information from the input features for angle transformation and dynamically re-parameterizes the convolutional kernel weights, filtering out redundant features through the updated convolutional kernel weights. The process is as follows:
[0105] The feature map is sent to a depthwise separable convolutional layer and an average pooling layer to further extract features, and then input into a linear prediction layer. The linear prediction layer includes a fully connected layer and an activation function, and predicts n angle parameters and the corresponding weight parameters , , and performs coordinate transformation according to the coordinates of the points on the feature map to obtain , the coordinates after angle transformation will make the weights fall on new positions. Bilinear interpolation is used to ensure the continuity of weight changes and prevent weight information loss caused by angle changes. Finally, the updated convolutional kernel weights, weight parameters, and the original visual features are sent into a convolutional layer to filter out redundant features and generate new features. The formula is as follows:
[0106] (11)
[0107] (12)
[0108] (13)
[0109] Among them, is the coordinate of the point on the original feature map, is the coordinate after angle transformation, is the inverse matrix of the rotation matrix of the angle affine transformation, is bilinear interpolation, represents the original static convolution kernel weight, represents according to the transformed for the original weight to be updated, is the original feature map, is the weighted feature map.
[0110] Optionally, the training process of the auricle reference segmentation model uses instance contrast learning to improve the discriminability of the model for different instances and its robustness to different languages describing the same instance;
[0111] As Figure 7 shown, in the instance contrast learning, b text descriptions of an image are sampled, and these text descriptions may point to the same instance or different instances in the image. The model will generate b mask predictions based on the image and the b text descriptions, making the masks of the sentences describing the same instance similar and suppressing the overlap of the masks of different instances. The overlap score is obtained by calculating the overlap between every two masks. The overlap score between the i-th and j-th final masks and is calculated by the following formula:
[0112] (14)
[0113] where, , n is the pixel size of the feature map S, is the normalized value of the i-th segmentation mask at position n. According to the overlap score, the instance contrast loss is calculated and used to enhance the discrimination ability of the model for different instances. The formula is as follows:
[0114] (15)
[0115] where, represents and whether they correspond to the same instance, is the intersection over union between the i-th mask and the actual annotation, which is used to prevent incorrect mask predictions;
[0116] The loss function of the auricle reference segmentation model uses a joint optimization loss function, which combines binary cross-entropy loss, Dice loss and the instance contrast loss to improve the segmentation accuracy and boundary discrimination ability of the model, expressed as:
[0117] (16)
[0118] Among them, represents the binary cross - entropy loss, represents the Dice loss, represents the instance contrast loss described above.
[0119] The binary cross - entropy loss in the embodiments of the present invention is used for pixel - level classification, determining whether each pixel belongs to the target area, punishing the accuracy of the model prediction probability, being able to provide a fine - grained supervision signal, and improving the segmentation accuracy. The Dice loss is a loss function based on the degree of regional overlap, used to evaluate the similarity between the predicted area and the real area, being more sensitive to small target areas, being able to effectively alleviate the problem of unbalanced target areas, and significantly enhancing the model's perception ability of the boundary area. The instance contrast loss is used to enhance the model's ability to distinguish different instances, prevent the phenomenon that there may be partial mask overlap in different ear regions, and improve the model's boundary discrimination ability.
[0120] The training process of the auricle reference segmentation model based on text guidance and angle transformation in the embodiments of the present invention is as follows:
[0121] 1) Data collection.
[0122] Collect a certain number of human ear images through a high - definition camera, manually annotate the acupoint areas of the human ear, and associate them with text descriptions.
[0123] 2) Dataset division.
[0124] Divide all the collected data samples into a training set, a validation set, and a test set according to the ratio of 8:1:1 to form a complete dataset. The training set is used for the training of the algorithm, the validation set is used for parameter tuning and validating the training effect, and the test set is used for evaluating the final performance of the model. Finally, select the model with the best performance on the test set as the final model.
[0125] 3) Model training.
[0126] The model is trained on the training set and the model effect is verified on the validation set. Select the model with the best performance on the validation set as the final model.
[0127] As Figure 8 shown, the embodiments of the present invention also provide an auricle reference segmentation system, and the system includes:
[0128] An acquisition module 810, configured to acquire the human ear image to be segmented and the corresponding text description;
[0129] A segmentation module 820, configured to input the ear image to be segmented and the corresponding text description into an auricle reference segmentation model based on text guidance and angle transformation, where the auricle reference segmentation model includes a text encoding module, a text-guided visual encoding module, a text-guided visual decoding module, and an angle transformation module;
[0130] The text encoding module embeds the text description into a high-dimensional word vector to obtain text features ;
[0131] The text-guided visual encoding module realizes the fusion of text features and image features through a four-stage structure. These four stages are connected in series, and the output of the previous stage is used as the input of the next stage to achieve multi-stage feature fusion. Each stage includes a visual encoder, a cross-modal perception module, and an attention gating module. The visual encoder of each stage generates visual features , the cross-modal perception module cross-modally aligns and fuses the visual features with the text features to obtain multi-modal features , and each element in the multi-modal features is weighted by the attention gating module to obtain weighted multi-modal features , and added element-wise to the original visual features to generate a set of enhanced visual features embedding language information , and the enhanced visual features with the smallest scale generated by the last stage are input into the text-guided visual decoding module;
[0132] The text-guided visual decoding module gradually restores the spatial resolution of the image, and further fuses text features and visual features to provide high-quality feature representations for the final segmentation task, and outputs multi-scale features;
[0133] The angle transformation module performs angle transformation on the multi-scale features and outputs a segmentation mask for the region related to the text description.
[0134] An auricle reference segmentation system provided by an embodiment of the present invention has a functional structure corresponding to an auricle reference segmentation method provided by an embodiment of the present invention, and will not be elaborated here.
[0135] Figure 9FIG. 0 is a schematic structural diagram of an electronic device 900 provided by an embodiment of the present invention. The electronic device 900 may vary greatly due to different configurations or performances, and may include one or more processors (central processing units, CPUs) 901 and one or more memories 902. Among them, at least one instruction is stored in the memory 902, and the at least one instruction is loaded and executed by the processor 901 to implement the steps of the above-mentioned auricle reference segmentation method.
[0136] In an exemplary embodiment, a computer-readable storage medium is also provided. For example, a memory including instructions, and the above instructions can be executed by a processor in a terminal to complete the above auricle reference segmentation method. For example, the computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0137] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiments can be completed by hardware, or can be completed by a program instructing related hardware. The program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a magnetic disk or an optical disc, etc.
[0138] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for segmenting auricle reference, characterized in that: The method comprises: S1, obtaining the human ear image to be segmented and the corresponding text description; S2, inputting the human ear image to be segmented and the corresponding text description into an auricle reference segmentation model based on text guidance and angle transformation, wherein the auricle reference segmentation model includes a text encoding module, a text-guided visual encoding module, a text-guided visual decoding module and an angle transformation module; The text encoding module embeds the text description into a high-dimensional word vector to obtain text features ; The text-guided visual encoding module realizes the fusion of text features and image features by organizing into a four-stage structure. The four stages are connected in series, and the output of the previous stage is used as the input of the next stage to realize multi-stage feature fusion. Each stage includes a visual encoder, a cross-modal perception module and an attention gating module. The visual encoder of each stage generates visual features. The cross-modal perception module aligns and fuses the visual features across modalities. With text features , get multimodal features , the multimodal features Each element in is weighted by the attention gating module to obtain weighted multimodal features , by element and original visual features Added together, they produce a set of enhanced visual features that embed linguistic information , the smallest scale enhanced visual features produced by the last stage inputting the text-guided visual decoding module; The text-guided visual decoding module gradually restores the spatial resolution of the image, and further integrates text features and visual features to provide high-quality feature representation for the final segmentation task and output multi-scale features; The angle transformation module performs angle transformation on the multi-scale features and outputs a segmentation mask of an area related to the text description; The text-guided visual decoding module consists of three stages, and the input of each stage is the fusion feature output of the previous stage. And the multimodal features of the corresponding stages in the visual encoding module ; In each stage, the visual decoder first upsamples the feature map of the previous stage through bilinear interpolation, gradually expands the spatial resolution of the feature map, restores the spatial information of the original image, and lays the foundation for subsequent image segmentation; The upsampled features and the multimodal features of the corresponding stage of the visual encoding module Add in the channel dimension to further achieve deep fusion of visual features and text features; During the entire decoding process, the visual decoding module uses the smallest scale enhanced visual features output by the visual encoding module. As input, after three levels of decoding, three larger-scale feature maps are gradually generated, which together with the initial input feature map constitute four-scale feature maps. Finally, the sizes of the feature maps of the four scales are consistent with the feature maps corresponding to the visual encoding module, which provides structural support for feature fusion and segmentation, and finally outputs multi-scale features. .
2. The method according to claim 1, characterized in that: The text encoding module uses BERT to perform text encoding and embeds the text description into a high-dimensional word vector to obtain the text feature: in and and N represent the number of word embedding channels and the number of words contained in the text description, respectively.
3. The method according to claim 1, characterized in that The visual encoder adopts the Swin-Transformer encoder. The input feature map is divided into windows of fixed size. The features in each window are modeled through a multi-head self-attention mechanism to capture the relationship between local areas. After the modeling is completed at the window level, the position of the window will be periodically translated to achieve global context association across windows. In the specific encoding process, the feature map is first mapped to a high-dimensional space through a linear embedding module, and then feature extraction and nonlinear transformation are performed layer by layer through a sliding window attention module and a multi-layer perceptron to enhance the expression ability of high-level semantic features and output the visual feature. .
4. The method according to claim 1, characterized in that: The cross-modal perception module uses the visual features As a query, the text feature As the key and value, we use the scaled dot product attention to calculate the process of supervising visual features with text features, and promote the fusion of multimodal information. After calculating the scaled dot product attention, we will obtain the features With input visual features Perform the projection operation again to obtain and , and then perform a dot product operation to finally obtain the multimodal feature , the process formula is as follows: (1) (2) (3) (4) (5) (6) (7) in, , , , , is the projection function, Represents the dot product operation, key function Sum value function Is has 1×1 convolution with the number of output channels, query function It is also a 1×1 convolution, and the number of output channels remains the same as the text feature. When the attention is calculated, the two-dimensional visual feature map is cut and arranged into one-dimensional sequence data to adapt to the subsequent attention calculation operation. After the attention calculation is completed, the one-dimensional sequence data is restored to the form of a two-dimensional feature map.
5. The method according to claim 1, characterized in that The attention gating module learns a set of The weight map is used to rescale Each element in , to prevent multimodal features Overwhelming visual features The visual signal in the image is used to generate the image and allow an adaptive amount of textual semantic information to flow to the next stage. The specific process is as follows: The multimodal features Perform 1×1 convolution and ReLu activation, and then perform 1×1 convolution and Tanh activation again. In this process, 1×1 convolution adjusts the number of channels to achieve the reorganization and weight distribution of channel information without changing the spatial resolution of the input feature map. The ReLu and Tanh activation functions can increase the nonlinear ability of the model and generate multimodal features. The weight is multiplied by the original feature map to obtain the weighted multimodal feature .
6. The method according to claim 1, characterized in that The angle transformation module includes three angle transformation layers, three nonlinear projection layers and one linear projection layer, and receives the multi-scale features output by the text-guided visual decoding module. As input, the final mask prediction is output. The whole process is shown in the following formula: (8) (9) (10) in, Represents the smallest feature map of the input, represents the input of each angle transformation layer, Represents the concat operation, which realizes the feature fusion in the channel dimension. The scale of feature maps of different scales is unified through bilinear interpolation. The process includes 3×3 convolutional layers, batch normalization layers, and ReLU activation functions, corresponding to nonlinear projection layers. represents the angle transformation layer, represents the final segmentation mask, is the linear projection function, and the feature Mapped to the final segmentation mask.
7. The method according to claim 6, characterized in that The angle transformation layer captures angle information from the input features to transform the angle, and dynamically re-parameterizes the convolution kernel weights to filter out redundant features through the updated convolution kernel weights. The process is as follows: The feature map is sent to the depth-separable convolution layer and the average pooling layer to further extract features, and then input to the linear prediction layer, which includes a fully connected layer and an activation function. n Angle parameters and the corresponding weight parameters , , according to the coordinates of the points on the feature map Perform coordinate transformation to obtain , coordinates after angle transformation The weights will fall to the new position, and the bilinear interpolation will ensure the continuity of weight changes to prevent the loss of weight information due to angle changes. Finally, the updated convolution kernel weights, weight parameters and original visual features are sent to the convolution layer to filter redundant features and generate new features. The formula is as follows: (11) (12) (13) in, is the coordinate of the point on the original feature map, is the coordinate after angle transformation, is the inverse matrix of the angular affine transformation rotation matrix, is bilinear interpolation, represents the original static convolution kernel weight, According to the transformed The original weight To update, is the original feature map, is the weighted feature map.
8. The method according to claim 1, characterized in that The training process of the ear pinna segmentation model uses instance contrast learning to improve the model's discriminability for different instances and its robustness to different languages describing the same instance; The instance contrastive learning is used to compare an image b These text descriptions may refer to the same instance or different instances in the image. The model will b Text description generation b mask prediction, making the masks of sentences describing the same instance similar, suppressing the overlap of masks of different instances, and obtaining the overlap score by calculating the overlap between every two masks. i and j Final mask and The overlap fraction between them is calculated by the following formula: (14) in, , m The feature map S The number of pixels, , It is i Segmentation mask In Location n The normalized value is used to calculate the instance contrast loss based on the overlap score. , which is used to enhance the model's ability to distinguish different instances. The formula is as follows: (15) in, express and Whether they correspond to the same instance, It is i Mask The intersection-over-union ratio between the actual annotation and the original annotation is used to prevent incorrect mask prediction; The loss function of the ear pinna segmentation model uses a joint optimization loss function to combine the binary cross entropy loss, Dice loss and the instance contrast loss Combined with the above, the segmentation accuracy and boundary distinction ability of the model are improved, which can be expressed as: (16) in, represents the binary cross entropy loss, represents the Dice loss, Denotes the instance contrast loss.
9. An auricle reference segmentation system, characterized in that: The system comprises: An acquisition module, used for acquiring a human ear image to be segmented and a corresponding text description; A segmentation module, used for inputting the human ear image to be segmented and the corresponding text description into an auricle reference segmentation model based on text guidance and angle transformation, wherein the auricle reference segmentation model includes a text encoding module, a text-guided visual encoding module, a text-guided visual decoding module and an angle transformation module; The text encoding module embeds the text description into a high-dimensional word vector to obtain text features ; The text-guided visual encoding module realizes the fusion of text features and image features by organizing into a four-stage structure. The four stages are connected in series, and the output of the previous stage is used as the input of the next stage to realize multi-stage feature fusion. Each stage includes a visual encoder, a cross-modal perception module and an attention gating module. The visual encoder of each stage generates visual features. The cross-modal perception module aligns and fuses the visual features across modalities. With text features , get multimodal features , the multimodal features Each element in is weighted by the attention gating module to obtain weighted multimodal features , by element and original visual features Added together, they produce a set of enhanced visual features that embed linguistic information , the smallest scale enhanced visual features produced by the last stage inputting the text-guided visual decoding module; The text-guided visual decoding module gradually restores the spatial resolution of the image, and further integrates text features and visual features to provide high-quality feature representation for the final segmentation task and output multi-scale features; The angle transformation module performs angle transformation on the multi-scale features and outputs a segmentation mask of an area related to the text description; The text-guided visual decoding module consists of three stages, and the input of each stage is the fusion feature output of the previous stage. And the multimodal features of the corresponding stages in the visual encoding module ; In each stage, the visual decoder first upsamples the feature map of the previous stage through bilinear interpolation, gradually expands the spatial resolution of the feature map, restores the spatial information of the original image, and lays the foundation for subsequent image segmentation; The upsampled features and the multimodal features of the corresponding stage of the visual encoding module Add in the channel dimension to further achieve deep fusion of visual features and text features; During the entire decoding process, the visual decoding module uses the smallest scale enhanced visual features output by the visual encoding module. As input, after three levels of decoding, three larger-scale feature maps are gradually generated, which together with the initial input feature map constitute four-scale feature maps. Finally, the sizes of the feature maps of the four scales are consistent with the feature maps corresponding to the visual encoding module, which provides structural support for feature fusion and segmentation, and finally outputs multi-scale features. .
Citation Information
Patent Citations
Picture referentiality segmentation method and device, computer equipment and storage medium
CN113592881A
Anaphora segmentation method based on multi-scale feature selective fusion
CN116152265A