A caenorhabditis elegans segmentation method and device based on local high-order semantic aggregation
By combining a local high-order semantic aggregation and hierarchical feature filtering aggregation module with a regression-guided localization information calibration module, the segmentation problem of overlapping and complex backgrounds in *C. elegans* segmentation was solved, achieving high-precision automated segmentation results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGDONG UNIV OF TECH
- Filing Date
- 2026-02-12
- Publication Date
- 2026-05-29
AI Technical Summary
Existing methods for segmenting *C. elegans* rely on manual observation, which is time-consuming and labor-intensive. Furthermore, it is difficult to accurately segment overlapping nematodes and tiny eggs in complex backgrounds. Existing deep learning models do not perform well when dealing with complex background interference.
The Local High-Order Semantic Aggregation (LHSA) module is used to enhance features, combined with the Hierarchical Feature Selection Aggregation (HFSA) module and the Regression-Guided Locality Correction (RGLCC) module. Through LHSA, hierarchical feature selection and regression feature correction, accurate segmentation of *C. elegans* is achieved.
It achieves high-precision segmentation of *C. elegans*, effectively addressing nematode overlap, tiny eggs, and complex background interference, thus improving segmentation accuracy and efficiency.
Smart Images

Figure CN122116353A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing, and specifically to a method and apparatus for segmenting *C. elegans* based on local high-order semantic aggregation. Background Technology
[0002] Caenorhabditis elegans, as an important model organism, is widely used in environmental toxicology, genetics, pharmacology, and other fields. Its biological characteristics give it unique advantages in pollutant toxicity assessment, ecological risk assessment, and mechanistic research. Especially in soil and aquatic environments, Caenorhabditis elegans is highly sensitive to a variety of pollutants and can rapidly respond to the effects of pharmaceuticals, personal care products, heavy metals, microplastics, and other pollutants. Calculating its growth and reproduction indicators is very helpful in determining the toxicity of environmental pollutants, which requires precise segmentation of Caenorhabditis elegans and its eggs.
[0003] However, traditional methods for segmenting *C. elegans* rely primarily on manual observation and counting, such as indicators of movement and egg production. This is not only time-consuming and labor-intensive but also limited by the operator's experience and has a low throughput. Especially in complex environments, background interference factors such as contaminants, food scraps, and excrement often increase the difficulty of counting. Existing manual counting methods struggle to address these challenges, resulting in compromised accuracy and efficiency.
[0004] With the development of computer technology, deep learning methods have made significant progress in the field of computer vision, especially in the increasingly widespread application of biomedical image processing. Deep learning can automatically learn deeper feature representations from large-scale data, effectively improving the performance of image segmentation tasks. However, for the segmentation problem of *C. elegans*, although existing deep learning methods have made some progress, the performance of current models is still unsatisfactory when dealing with overlapping nematodes, tiny eggs, and complex background interference. This is mainly because existing methods are insufficient in balancing the capture of global information and detailed features, especially when dealing with high-density, overlapping nematode groups, which easily leads to blurred boundaries and missed detections.
[0005] Therefore, developing an automated method for segmenting *C. elegans* that can effectively address issues such as nematode overlap, tiny eggs, and complex background interference is of great significance for assisting in the monitoring of environmental pollutants. Summary of the Invention
[0006] To address the aforementioned technical problems, this invention provides a method and apparatus for segmenting *C. elegans* based on local high-order semantic aggregation, which can provide users with a more accurate, efficient, and highly adaptable *C. elegans* segmentation solution.
[0007] Specifically, the method includes the following steps: S1: Acquire microscopic images of *C. elegans* and preprocess them. Input the images into the backbone network to extract shallow, middle and deep features. S2: Construct a Local High-Order Semantic Aggregation Module (LHSA) to perform local high-order semantic aggregation on deep features to obtain enhanced features; S3: Construct a hierarchical feature filtering and aggregation module HFSA to fuse and enhance the features of the backbone network and the neck network to obtain fused features; specifically, this includes the following steps: S3.1: Perform hierarchical feature filtering on the features of the backbone network and the neck network; S3.2: The selected backbone network and neck network features are fused to obtain the fused features. ; S4: Input the fused features into the segmentation head to generate bounding box regression features and classification features, and construct the regression-guided localization accuracy calibration module RGLCC to perform quality correction on the classification features using the regression features; specifically, this includes the following steps: S4.1: Generate bounding box regression features and classification features; S4.2: Construct a regression-guided localization calibration module RGLCC to perform quality correction on classification features using regression features; S5: Decode and output the detection bounding box, instance mask, and category results.
[0008] Preferably, S1 includes the following steps: S1.1: Acquire high-resolution microscopic images of *C. elegans* and crop the images into four equal parts (length and width); S1.2: Divide the preprocessed data into training set, validation set and test set; S1.3: Define the input image data as... The input is processed by a feature extraction network to extract shallow features. Mid-layer characteristics with deep features .
[0009] Preferably, S2 includes the following steps: deep features Input the local high-order semantic aggregation module to obtain enhanced features. The operation is described as follows:
[0010]
[0011]
[0012]
[0013] In the above formula, represent The output after convolution represent Perform twice The output of the operation Indicates to and The output after splicing express The output after convolution Represents convolution calculation, This represents the splicing operation along the channel dimension. The operation represents the local high-order semantic aggregation of features; Local higher-order semantic aggregation operations Described as:
[0014]
[0015]
[0016]
[0017] In the above formula, represent Output after multi-head attention operation and The representation is the result after convolution calculation Features obtained by splitting by channel dimension Representative to The guiding features are obtained after performing localized and effective information aggregation. Representative use right Modulate and The output is the sum of the residuals. This represents bullish attention-based trading. This represents the operation of splitting features by channel dimension. This represents a depthwise separable convolution operation. This represents an element-wise addition operation. This represents element-wise multiplication. This represents the activation function.
[0018] Preferably, S3 includes the following steps: S3.1: Perform hierarchical feature filtering on the features of the backbone network and the neck network; S3.2: The selected backbone network and neck network features are fused to obtain the fused features. .
[0019] Preferably, S3.1 includes the following steps: Define the output features of the neck network as , and , The set representing the output features of the backbone network. The set representing the output features of the neck network, performing hierarchical feature filtering on the features of the backbone network and the neck network, is described as follows:
[0020]
[0021]
[0022]
[0023] In the above formula, represent The output after convolution represent The output after convolution represent The output after hierarchical feature filtering represent The output after hierarchical feature filtering Represents step length, Representative hierarchical feature filtering operation; Hierarchical feature filtering operation The definition is as follows: Let the input be a feature tensor. , , The feature map is divided by step size The space dimension is divided into non-overlapping patches, and each patch is denoted as . , , , each Average along the channel dimension and expand into a one-dimensional vector. :
[0024]
[0025] In the above formula, Non-overlapping patches representing feature maps. The representative will The output after unfolding into one dimension The representative is responsible for unfolding the operation; For each vector go through The transformation is then applied, followed by activation and spatial normalization to obtain the weights at each position within the vector. Finally, element-wise multiplication is performed to obtain the final weights. :
[0026]
[0027]
[0028] In the above formula, represent go through The transformed output, represent The weights obtained after spatial dimension normalization represent and The output after element-wise multiplication This represents the feature transformation operation of the feedforward network. Represents normalization operation; A feature selection mechanism is introduced to highlight effective features relevant to the task and suppress interference from irrelevant or redundant features. A channel gating signal is generated for each token, and this signal is used to filter tokens in the channel dimension. Subsequently, a linear transformation is applied to the filtered tokens to complete the channel selection.
[0029]
[0030] In the above formula, represent The output after passing through a multilayer perceptron and the sigmoid function represent and The output after element-wise multiplication and linear transformation Represents the Sigmoid function. Represents multilayer sensor operation. Represents a linear transformation operation; Each Reshaping and interpolation operations are performed to ultimately generate features. ;
[0031]
[0032]
[0033]
[0034] In the above formula, Representative to Do =2 operate, Representative to Do =4 operate, Representative to Do =2 operate, Representative to Do =4 operate; Preferably, S3.2 includes the following steps: The selected backbone network and neck network features are fused to obtain fused features. The operation is described as follows:
[0035]
[0036]
[0037]
[0038]
[0039] In the above formula, represent and The output after concatenation along the channel dimension represent and The output after concatenation along the channel dimension represent and The output after concatenation along the channel dimension represent , and The output after concatenation along the channel dimension represent The fused features after feature refinement This represents reparameterized convolution.
[0040] Preferably, S4 includes the following steps: S4.1: Generate bounding box regression features and classification features; S4.2: Construct a regression-guided localization calibration module RGLCC to perform quality correction on classification features using regression features.
[0041] Preferably, S4.1 includes the following steps: The operation description for generating bounding box regression features and classification features is as follows:
[0042]
[0043] In the above formula, represent The bounding box regression features are obtained through a series of convolution operations. represent The classification features are obtained through a series of convolution operations. Represents a two-dimensional convolution operation; Preferably, S4.2 includes the following steps: The operation description of constructing the Regression-Guided Positional Information Calibration Module (RGLCC) to perform quality correction on classification features using regression features is as follows:
[0044]
[0045]
[0046]
[0047] In the above formula, represent The output obtained by normalizing the first dimension of the channel. Representative from Before selecting the first dimension of the channel The output of the maximum value feature, represent and The output is obtained by concatenating the channels based on the mean of the first dimension of the channel. Representative to The output after performing a multilayer perceptron operation. Representative at Before selecting the channel The maximum value feature operation, Representative of features in Average operation across multiple channels; The adjusted quality score is added to the initial classification score to obtain the final classification score. :
[0048] In the above formula, This represents the final classification score.
[0049] Preferably, S5 includes the following steps: S5.1: Regression features of the bounding box With category features Decode the candidate detection boxes to obtain their category probabilities and confidence scores; S5.2: Perform threshold filtering and redundancy removal on the candidate results to obtain the final detection box and its corresponding category; S5.3: Generate and restore instance masks based on the final detection boxes, and output the detection boxes, instance masks and category prediction results.
[0050] The second technical solution adopted in this invention is: a segmentation device for *C. elegans* based on local high-order semantic aggregation, comprising: LHSA module: Utilizes multi-scale depthwise separable convolution operations to focus on and enhance local detail features, thereby effectively improving the representation ability of deep features and optimizing the recognition accuracy of nematode boundaries and morphology; HFSA module: Through hierarchical feature selection and fusion, it effectively aggregates features from the backbone network and the neck network, highlighting key information relevant to the task, suppressing irrelevant features and background noise, and improving the model's sensitivity to detailed information. The RGLCC module effectively improves the accuracy of classification features through regression feature-guided localization calibration, further optimizing the regression and classification accuracy of the detection box, ensuring accurate identification and localization of *C. elegans* and its eggs in complex backgrounds.
[0051] The beneficial effects of this invention are as follows: First, in the data preprocessing stage, cropping the high-resolution microscopic images of *C. elegans* makes the images easier to train, laying a good foundation for subsequent model training and testing. During model construction, a Local High-Order Semantic Aggregation Module (LHSA) is proposed, effectively enhancing the local detail representation of features. Next, a Hierarchical Feature Selection Aggregation Module (HFSA) is proposed. This module employs a hierarchical feature selection mechanism, enabling hierarchical attention-based fusion of features at different levels, suppressing irrelevant background and noise, and strengthening discriminative cues related to the target. Subsequently, Regression-Guided Locality Calibration (RGLCC) is proposed, using regression features to correct the quality of classification features, further improving detection accuracy. Finally, by decoding to generate detection boxes, instance masks, and category results, accurate segmentation of *C. elegans* is achieved. Through these innovative designs, this invention achieves high-precision segmentation of *C. elegans*, providing an efficient and reliable solution for the *C. elegans* segmentation task. Attached Figure Description
[0052] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following is a brief introduction to the drawings used in the prior art and embodiments. The following drawings are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0053] Figure 1 is a flowchart illustrating the nematode segmentation method for *C. elegans* based on local high-order semantic aggregation according to the present invention. Figure 2 is a model architecture diagram of the *C. elegans* segmentation method based on local high-order semantic aggregation according to the present invention. Figure 3 is a schematic diagram of the Local High-Order Semantic Aggregation (LHSA) module of the Caenorhabditis elegans segmentation device based on local high-order semantic aggregation according to the present invention. Figure 4 is a schematic diagram of the hierarchical feature filtering and aggregation module HFSA of the nematode segmentation device based on local high-order semantic aggregation according to the present invention. Figure 5 is a schematic diagram of the regression-guided localization information calibration module RGLCC of the nematode segmentation device based on local high-order semantic aggregation according to the present invention. Detailed Implementation Plan To make the objectives, features, and advantages of this invention more apparent and understandable, the technical solutions of the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of this application. Unless otherwise specified, all other embodiments obtained by those skilled in the art based on the embodiments of this invention without inventive effort are within the scope of protection of this invention.
[0054] This invention provides a method and apparatus for segmenting *C. elegans* based on local high-order semantic aggregation, which can be used to achieve accurate segmentation of *C. elegans*.
[0055] Reference Figure 1 The method includes the following steps: S1: Acquire microscopic images of *C. elegans* and preprocess them. Input the images into the backbone network to extract shallow, middle and deep features. S2: Construct a Local High-Order Semantic Aggregation Module (LHSA) to perform local high-order semantic aggregation on deep features to obtain enhanced features; S3: Construct the hierarchical feature filtering and aggregation module HFSA to fuse and enhance the features of the backbone network and the neck network to obtain fused features; S4: Input the fused features into the segmentation head to generate bounding box regression features and classification features, and construct the regression-guided localization information calibration module RGLCC to perform quality correction on the classification features using the regression features; S5: Decode and output the detection bounding box, instance mask, and category results.
[0056] Furthermore, S1 includes the following steps: S1.1: Acquire high-resolution microscopic images of *C. elegans* and crop the images into four equal parts (length and width); S1.2: Divide the preprocessed data into training set, validation set and test set; S1.3: Define the input image data as... The input is processed by a feature extraction network to extract shallow features. Mid-layer characteristics with deep features .
[0057] Furthermore, refer to Figure 3 S2 includes the following steps: deep features Input the local high-order semantic aggregation module to obtain enhanced features. The operation is described as follows:
[0058]
[0059]
[0060]
[0061] In the above formula, represent The output after convolution represent Perform twice The output of the operation Indicates to and The output after splicing express The output after convolution Represents convolution calculation, This represents the splicing operation along the channel dimension. The operation represents the local high-order semantic aggregation of features; Local higher-order semantic aggregation operations Described as:
[0062]
[0063]
[0064]
[0065] In the above formula, represent Output after multi-head attention operation and The representation is the result after convolution calculation Features obtained by splitting by channel dimension Representative to The guiding features are obtained after performing localized and effective information aggregation. Representative use right Modulate and The output is the sum of the residuals. This represents bullish attention-based trading. This represents the operation of splitting features by channel dimension. This represents a depthwise separable convolution operation. This represents an element-wise addition operation. This represents element-wise multiplication. This represents the activation function.
[0066] Furthermore, refer to Figure 4 S3 includes the following steps: S3.1: Perform hierarchical feature filtering on the features of the backbone network and the neck network; S3.2: The selected backbone network and neck network features are fused to obtain the fused features. .
[0067] Furthermore, S3.1 includes the following steps: Define the output features of the neck network as , and , The set representing the output features of the backbone network. The set representing the output features of the neck network, performing hierarchical feature filtering on the features of the backbone network and the neck network, is described as follows:
[0068]
[0069]
[0070]
[0071] In the above formula, represent The output after convolution represent The output after convolution represent The output after hierarchical feature filtering represent The output after hierarchical feature filtering Represents step length, Representative hierarchical feature filtering operation; Hierarchical feature filtering operation The definition is as follows: Let the input be a feature tensor. , , The feature map is divided by step size The space dimension is divided into non-overlapping patches, and each patch is denoted as . , , , each Average along the channel dimension and expand into a one-dimensional vector. :
[0072]
[0073] In the above formula, Non-overlapping patches representing feature maps. The representative will The output after unfolding into one dimension The representative is responsible for unfolding the operation; For each vector go through The transformation is then applied, followed by activation and spatial normalization to obtain the weights at each position within the vector. Finally, element-wise multiplication is performed to obtain the final weights. :
[0074]
[0075]
[0076] In the above formula, represent go through The transformed output, represent The weights obtained after spatial dimension normalization represent and The output after element-wise multiplication This represents the feature transformation operation of the feedforward network. Represents normalization operation; A feature selection mechanism is introduced to highlight effective features relevant to the task and suppress interference from irrelevant or redundant features. A channel gating signal is generated for each token, and this signal is used to filter tokens in the channel dimension. Subsequently, a linear transformation is applied to the filtered tokens to complete the channel selection.
[0077]
[0078] In the above formula, represent The output after passing through a multilayer perceptron and the sigmoid function represent and The output after element-wise multiplication and linear transformation Represents the Sigmoid function. Represents multilayer sensor operation. Represents a linear transformation operation; Each Reshaping and interpolation operations are performed to ultimately generate features. ;
[0079]
[0080]
[0081]
[0082] In the above formula, Representative to Do =2 operate, Representative to Do =4 operate, Representative to Do =2 operate, Representative to Do =4 operate; Furthermore, S3.2 includes the following steps: The selected backbone network and neck network features are fused to obtain fused features. The operation is described as follows:
[0083]
[0084]
[0085]
[0086]
[0087] In the above formula, represent and The output after concatenation along the channel dimension represent and The output after concatenation along the channel dimension represent and The output after concatenation along the channel dimension represent , and The output after concatenation along the channel dimension represent The fused features after feature refinement This represents reparameterized convolution.
[0088] Furthermore, refer to Figure 5 S4 includes the following steps: S4.1: Generate bounding box regression features and classification features; S4.2: Construct a regression-guided localization calibration module RGLCC to perform quality correction on classification features using regression features.
[0089] Furthermore, S4.1 includes the following steps: The operation description for generating bounding box regression features and classification features is as follows:
[0090]
[0091] In the above formula, represent The bounding box regression features are obtained through a series of convolution operations. represent The classification features are obtained through a series of convolution operations. Represents a two-dimensional convolution operation; Furthermore, S4.2 includes the following steps: The operation description of constructing the Regression-Guided Positional Information Calibration Module (RGLCC) to perform quality correction on classification features using regression features is as follows:
[0092]
[0093]
[0094]
[0095] In the above formula, represent The output obtained by normalizing the first dimension of the channel. Representative from Before selecting the first dimension of the channel The output of the maximum value feature, represent and The output is obtained by concatenating the channels based on the mean of the first dimension of the channel. Representative to The output after performing a multilayer perceptron operation. Representative at Before selecting the channel The maximum value feature operation, Representative of features in Average operation across multiple channels; The adjusted quality score is added to the initial classification score to obtain the final classification score. :
[0096] In the above formula, This represents the final classification score.
[0097] Furthermore, refer to Figure 2 S5 includes the following steps: S5.1: Regression features of the bounding box With category features Decode the candidate detection boxes to obtain their category probabilities and confidence scores; S5.2: Perform threshold filtering and redundancy removal on the candidate results to obtain the final detection box and its corresponding category; S5.3: Generate and restore instance masks based on the final detection boxes, and output the detection boxes, instance masks and category prediction results.
[0098] The above is a detailed description of the preferred embodiments of the present invention. However, the present invention is not limited to the embodiments described. Those skilled in the art can make other equivalent modifications or substitutions without departing from the spirit of the invention. These equivalent modifications or substitutions are included within the scope defined by the claims.
Claims
1. A segmentation method for *C. elegans* based on local high-order semantic aggregation, characterized in that, The method includes the following steps: S1: Acquire microscopic images of *C. elegans* and preprocess them. Input the images into the backbone network to extract shallow, middle and deep features. S2: Construct a Local High-Order Semantic Aggregation Module (LHSA) to perform local high-order semantic aggregation on deep features to obtain enhanced features; S3: Construct a hierarchical feature filtering and aggregation module HFSA to fuse and enhance the features of the backbone network and the neck network to obtain fused features; specifically, this includes the following steps: S3.1: Perform hierarchical feature filtering on the features of the backbone network and the neck network; S3.2: The selected backbone network and neck network features are fused to obtain the fused features. ; S4: Input the fused features into the segmentation head to generate bounding box regression features and classification features, and construct the regression-guided localization accuracy calibration module RGLCC to perform quality correction on the classification features using the regression features; specifically, this includes the following steps: S4.1: Generate bounding box regression features and classification features; S4.2: Construct a regression-guided localization calibration module RGLCC to perform quality correction on classification features using regression features; S5: Decode and output the detection bounding box, instance mask, and category results.
2. The method for segmenting *C. elegans* based on local high-order semantic aggregation according to claim 1, characterized in that, S1 includes the following steps: S1.1: Acquire high-resolution microscopic images of *C. elegans* and crop the images into four equal parts (length and width); S1.2: Divide the preprocessed data into training set, validation set and test set; S1.3: Define the input image data as... The input is processed by a feature extraction network to extract shallow features. Mid-layer characteristics with deep features .
3. The method for segmenting *C. elegans* based on local high-order semantic aggregation according to claim 2, characterized in that, S2 includes the following steps: deep features Input the local high-order semantic aggregation module to obtain enhanced features. The operation is described as follows: In the above formula, represent The output after convolution represent Perform twice The output of the operation Indicates to and The output after splicing express The output after convolution Represents convolution calculation, This represents the splicing operation along the channel dimension. The operation represents the local high-order semantic aggregation of features; Local higher-order semantic aggregation operations Described as: In the above formula, represent Output after multi-head attention operation and The representation is the result after convolution calculation Features obtained by splitting by channel dimension Representative to The guiding features are obtained after performing localized and effective information aggregation. Representative use right Modulate and The output is the sum of the residuals. This represents bullish attention-based trading. This represents the operation of splitting features by channel dimension. This represents a depthwise separable convolution operation. This represents an element-wise addition operation. This represents element-wise multiplication. This represents the activation function.
4. The method for segmenting *C. elegans* based on local high-order semantic aggregation according to claim 3, characterized in that, The sub-steps in S3 further include: S3.1: Define the output features of the neck network as follows , and , The set representing the output features of the backbone network. The set representing the output features of the neck network, performing hierarchical feature filtering on the features of the backbone network and the neck network, is described as follows: In the above formula, represent The output after convolution represent The output after convolution represent The output after hierarchical feature filtering represent The output after hierarchical feature filtering Represents step length, Representative hierarchical feature filtering operation; Hierarchical feature filtering operation The definition is as follows: Let the input be a feature tensor. , , The feature map is divided by step size The space dimension is divided into non-overlapping patches, and each patch is denoted as . , , , each Average along the channel dimension and expand into a one-dimensional vector. : In the above formula, Non-overlapping patches representing feature maps. The representative will The output after unfolding into one dimension The representative is responsible for unfolding the operation; For each vector go through The transformation is then applied, followed by activation and spatial normalization to obtain the weights at each position within the vector. Finally, element-wise multiplication is performed to obtain the final weights. : In the above formula, represent go through The transformed output, represent The weights obtained after spatial dimension normalization represent and The output after element-wise multiplication This represents the feature transformation operation of the feedforward network. Represents normalization operation; A feature selection mechanism is introduced to highlight effective features relevant to the task and suppress interference from irrelevant or redundant features. A channel gating signal is generated for each token, and this signal is used to filter tokens in the channel dimension. Subsequently, a linear transformation is applied to the filtered tokens to complete the channel selection. In the above formula, represent The output after passing through a multilayer perceptron and the sigmoid function represent and The output after element-wise multiplication and linear transformation Represents the Sigmoid function. Represents multilayer sensor operation. Represents a linear transformation operation; Each Reshaping and interpolation operations are performed to ultimately generate features. ; In the above formula, Representative to Do =2 operate, Representative to Do =4 operate, Representative to Do =2 operate, Representative to Do =4 operate; S3.2: The selected backbone network and neck network features are fused to obtain the fused features. The operation is described as follows: In the above formula, represent and The output after concatenation along the channel dimension represent and The output after concatenation along the channel dimension represent and The output after concatenation along the channel dimension represent , and The output after concatenation along the channel dimension represent The fused features after feature refinement This represents reparameterized convolution.
5. The method for segmenting *C. elegans* based on local high-order semantic aggregation according to claim 4, characterized in that, The sub-steps in S4 further include: S4.1: The operation description for generating bounding box regression features and classification features is as follows: In the above formula, represent The bounding box regression features are obtained through a series of convolution operations. represent The classification features are obtained through a series of convolution operations. Represents a two-dimensional convolution operation; S4.2: The operation description of constructing the regression-guided localization information calibration module RGLCC, which uses regression features to perform quality correction on classification features, is as follows: In the above formula, represent The output obtained by normalizing the first dimension of the channel. Representative from Before selecting the first dimension of the channel The output of the maximum value feature represent and The output is obtained by concatenating the channels based on the mean of the first dimension of the channel. Representative to The output after performing a multilayer perceptron operation. Representative at Before selecting the channel The maximum value feature operation, Representative of features in Average operation across multiple channels; The adjusted quality score is added to the initial classification score to obtain the final classification score. : In the above formula, This represents the final classification score.
6. The method for segmenting *C. elegans* based on local high-order semantic aggregation according to claim 5, characterized in that, S5 includes the following steps: S5.1: Regression features of the bounding box With category features Decode the candidate detection boxes to obtain their category probabilities and confidence scores; S5.2: Perform threshold filtering and redundancy removal on the candidate results to obtain the final detection box and its corresponding category; S5.3: Generate and restore instance masks based on the final detection boxes, and output the detection boxes, instance masks and category prediction results.
7. A segmentation device for *C. elegans* based on local high-order semantic aggregation, characterized in that, include: LHSA module: Utilizes multi-scale depthwise separable convolution operations to focus on and enhance local detail features, thereby effectively improving the representation ability of deep features and optimizing the recognition accuracy of nematode boundaries and morphology; HFSA module: Through hierarchical feature selection and fusion, it effectively aggregates features from the backbone network and the neck network, highlighting key information relevant to the task, suppressing irrelevant features and background noise, and improving the model's sensitivity to detailed information. The RGLCC module effectively improves the accuracy of classification features through regression feature-guided localization calibration, further optimizing the regression and classification accuracy of the detection box, ensuring accurate identification and localization of *C. elegans* and its eggs in complex backgrounds.