Biological Tissue Recognition Method Based on Axial Free Attention Network

By using an axial free attention network method based on an axial free attention network in tissue recognition during embryonic blastocyst period, an axial free attention transformer with an encoder-decoder structure is solved, and a high-accurate biological tissue recognition is achieved.

CN119723577BActive Publication Date: 2025-06-13ANHUI UNIVERSITY OF TRADITIONAL CHINESE MEDICINE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510218542.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-13
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

The prior art is difficult to effectively capture non-local dependencies and alleviate local inconsistencies in tissue recognition during embryonic blastocyst period, and it is difficult to integrate global context features to enhance the learning ability of the network.

Method used

Using an axial free attention network method, by constructing an encoder-decoder structure, an axial free attention mechanism network and a soft aggregation operation network are used to capture non-local information and integrate global context features.

Benefits of technology

It realizes non-local feature maps that capture long-distance clues under low computing resources, alleviates local structure errors, and improves the learning ability of the network and the accuracy of biological tissue recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119723577B_ABST
    Figure CN119723577B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for biological tissue recognition based on an axial-free attention network, comprising: constructing an axial-free attention transformer, which adopts an encoder-decoder structure and includes a plurality of convolutional blocks, axial-free attention blocks and upsampling blocks. Each axial-free attention block includes an input network, D axial-free attention mechanism networks, a soft aggregation operation network, a 1×1 convolutional network and an output network; each axial-free attention mechanism network processes the input image feature data in different angular directions; training the axial-free attention transformer with training sample data to obtain a biological tissue recognition processing training model; and using the biological tissue recognition processing training model to recognize and process biological tissue image data. This method captures non-local feature maps with long-range cues with lower computational resources and expands diverse receptive fields.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for biological tissue recognition based on an axial free attention network, and particularly to the construction and application of an axial free attention network for tissue recognition during the embryo blastocyst stage, belonging to the technical field of image pattern recognition based on artificial neural networks and the field of artificial intelligence applied to medicine. Background Art

[0002] Since the implantation potential of in vitro fertilization embryos is related to their morphology, the morphological structure including cell density, size, and expansion level will determine the embryo quality grading, which makes the assessment of embryo quality rely on the segmentation of blastocyst components. However, in the prior art, when selecting blastocysts with high implantation potential for transfer by visually evaluating their morphological characteristics, due to problems such as poor contrast of embryo images, noise, and unclear boundaries of different tissue structures, it becomes a challenging problem to identify different tissue regions and their morphological structures in blastocyst images to achieve automatic blastocyst segmentation.

[0003] In the prior art, there have been some techniques for automatic segmentation and recognition of blastocyst images. For example, traditional computer vision methods (level sets) are used to segment the inner cell mass and trophectoderm; by combining text information (Gabor and DCT features) and the level set method to identify the ICM boundary, the feature expression can be enhanced to a certain extent; by introducing the clustering idea, the features of the same tissue can be better aggregated. With the development of deep learning, there have emerged methods of inputting the extracted traditional features into a two-layer neural network to identify the ZP, TE, and ICM; there have also emerged pure end-to-end deep learning methods, such as using a fully convolutional neural (FCN) network to segment the ICM. However, these methods only increase the receptive field of a single pixel but do not change the attribute relationship between pixels.

[0004] Ronneberger et al. proposed U-Net for the segmentation of biological microscopy images. Subsequently, various U-Net variants have been proposed for medical image analysis, such as Attention U-Net, Unet++, DU-Net, PraNet, nnUnet, etc. These networks are not designed specifically for embryo data and only utilize effective locality at the cost of sacrificing long-range cues.

[0005] Currently, there have also been attempts to use CNN and Transformer-based methods for blastocyst image segmentation, but there are still two main challenges: (1) how to effectively capture non-local dependencies and reduce local inconsistencies, as shown in Figure 1 Figure (a), to solve the complex manifestations of embryo morphology, especially regarding the problems of blurred edges and similar textures. (2) How to integrate global context features and obtain a wide receptive field to enhance the learning ability of the network. Summary of the Invention

[0006] To solve the above technical problems, the present invention provides a method for biological tissue recognition based on an axial free attention network, including the steps of:

[0007] Construct an axial free attention transformer, which adopts an encoder-decoder structure and includes a plurality of convolutional blocks, axial free attention blocks and upsampling blocks. Among them, each axial free attention block includes an input network, D axial free attention mechanism networks, a soft aggregation operation network, a 1×1 convolutional network and an output network; each axial free attention mechanism network processes the input image feature data in different angular directions, and it includes: a grid generator, a direction sampler, a one-dimensional attention module and a learnable sine relative position encoding module; wherein, the grid generator is used to calculate the pixel position relationship between the input feature map and the generated direction feature map Z, the direction sampler uses the sampling grid to generate the sampled direction feature map Z, the one-dimensional attention module is used to perform one-dimensional self-attention mechanism calculation processing on the direction feature map Z, and the learnable sine relative position encoding is used to introduce continuous deviation and adaptability;

[0008] Obtain training sample data, which is biological tissue image data and is labeled data;

[0009] Use the training sample data to train the axial free attention transformer to obtain a biological tissue recognition processing training model;

[0010] Use the biological tissue recognition processing training model to recognize and process biological tissue image data.

[0011] In the above technical solution, for the number D of the axial free attention mechanism networks, D × θ = π, where θ represents the included angle between the attention axes processed by adjacent axial free attention mechanism networks.

[0012] In the above technical solution, the input network of each axial free attention block divides the input image feature data into D paths and inputs them into D axial free attention mechanism networks respectively. The D axial free attention mechanism networks process the input image feature data in one angular direction respectively and output the results to the soft aggregation operation network; the soft aggregation operation network soft-aggregates the output results of the D axial free attention mechanism networks and then outputs them to the 1×1 convolutional network; the output of the 1×1 convolutional network is added to the input image feature data and then output to the output network to obtain the output image feature data.

[0013] In the above technical solution, the pixel position relationship between the input feature map X calculated by the grid generator and the direction feature map Z is as follows:

[0014]

[0015] where is the coordinate of the sample grid in the input feature map, θ is the angle of rotation of the processed attention axis direction relative to the coordinate axis, is the coordinate for generating a new pixel in the direction feature map Z.

[0016] In the above technical solution, the kernel function of the sampling grid used by the direction sampler is:

[0017]

[0018] where and define the interpolation operations in the x and y directions, is the value at the position (n, m) on the feature map, and H and W are the height and width of the image data respectively.

[0019] In the above technical solution, for a given pixel coordinate a = (i, j) and another pixel coordinate b = (l, n), the learnable sine relative position encoding (LSRPE) is defined as:

[0020]

[0021] where the learnable parameter , , is the relative position encoding.

[0022] In the above technical solution, the specific structure of the axial-free attention transformer includes:

[0023] The first layer includes the first convolutional block;

[0024] The second layer includes the second convolutional block and the first axial-free attention block;

[0025] The third layer includes the second axial-free attention block;

[0026] The fourth layer includes the third axial-free attention block;

[0027] The fifth layer includes the first upsampling block;

[0028] The sixth layer includes the second upsampling block;

[0029] The seventh layer includes the third upsampling block;

[0030] The eighth layer includes the fourth upsampling block;

[0031] Among them, the first layer is the input layer, the eighth layer is the output layer. One of the two outputs of the first layer is input to the second layer, and the other is added to the output of the seventh layer and then input to the eighth layer; one of the two outputs of the second layer is input to the third layer, and the other is added to the output of the sixth layer and then input to the seventh layer; one of the two outputs of the third layer is input to the fourth layer, and the other is added to the output of the fifth layer and then input to the sixth layer; the output of the fourth layer is input to the fifth layer.

[0032] The present invention also provides a biological tissue recognition device based on an axial free attention network, which is characterized by comprising:

[0033] A model construction module, used to construct an axial free attention transformer. The axial free attention transformer adopts an encoder-decoder structure and includes a plurality of convolutional blocks, axial free attention blocks and upsampling blocks. Among them, each axial free attention block includes an input network, D axial free attention mechanism networks, a soft aggregation operation network, a 1×1 convolutional network and an output network; each axial free attention mechanism network processes the input image feature data in different angular directions respectively, and it includes: a grid generator, a direction sampler, a one-dimensional attention module and a learnable sine relative position encoding module;

[0034] Among them, the grid generator is used to calculate the pixel position relationship between the input feature map and the generated direction feature map Z, the direction sampler uses the sampling grid to generate the sampled direction feature map Z, the one-dimensional attention module is used to perform one-dimensional self-attention mechanism calculation processing on the direction feature map Z, and the learnable sine relative position encoding is used to introduce continuous deviation and adaptability;

[0035] A training sample acquisition module, used to acquire training sample data. The training sample data is biological tissue image data, and the biological tissue image data is labeled data;

[0036] A model training module, which uses the training sample data to train the axial free attention transformer to obtain a biological tissue recognition processing training model;

[0037] A data processing module, which uses the biological tissue recognition processing training model to recognize and process biological tissue image data.

[0038] The present invention also provides a biological tissue recognition system based on an axial free attention network, including a processor and a storage medium; the program code is stored on the storage medium; the processor is used to call the program code stored in the storage medium to execute the biological tissue recognition method described in the above technical solution based on the axial free attention network.

[0039] The present invention has achieved the following technical effects: First, the present invention proposes an axial-free attention mechanism that uses low computational resources to capture non-local feature maps with long-range cues to mitigate errors in local structures. Second, the present invention proposes an axial-free attention block with a soft aggregation operation that embeds features extracted from different angles with the features of axial-free attention, which collect global cues and expand diverse receptive fields. By validating on typical public datasets, the best segmentation performance is achieved in terms of accuracy, precision, recall, Dice coefficient, and Jaccard index, with specific values of 93.86%, 91.86%, 92.20%, 92.03%, and 85.46% respectively; and extensive qualitative experimental results prove the effectiveness of the method. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 is a detailed feature map of a blastocyst tissue image, where Figure 1 in (a) shows the local inconsistency in the blastocyst tissue image, Figure 1 in (b) shows the rotational consistency of the blastocyst tissue image;

[0041] Figure 2 is the overall architecture diagram of BTFormer, where Figure 2 in (a) shows the specific architecture of BTFormer, Figure 2 in (b) shows the specific structure of the axial-free attention block, where Figure 2 in (c) shows the specific manner of the soft aggregation operation;

[0042] Figure 3 is the structure diagram of axial-free attention. DETAILED DESCRIPTION OF THE INVENTION

[0043] To facilitate the understanding and implementation of the present invention by those of ordinary skill in the art, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0044] The local inconsistency of the blastocyst tissue image, as Figure 1 shown in (a), the blastocyst tissue image is captured using a Hofmann optical sampling system, which is known for its three-dimensional relief effect and transparent tissue properties. The overlapping structures present in the captured image will exacerbate this local inconsistency; in addition, blastocyst tissue images are usually captured under low-light conditions. Although the entire embryo is transparent and has good brightness, under low-light conditions, the low contrast and noise of the captured image will further exacerbate this phenomenon. Figure 1In (a), two small regions of interest (ROI 1 and ROI 2) are extracted from the embryo image, and their positions in the original image and its corresponding label image are marked. It is worth noting that ROI 1 and ROI 2 exhibit similar colors and textures, although they belong to different categories in the label image. This observation indicates that for blastocyst tissue images, the classification of specific regions is affected not only by their colors and textures but also by their spatial relationships with neighboring regions, suggesting the importance of context relationships in embryo images. And the main labels of human blastocysts, such as Figure 1 shown in the label of (a), generally include the inner cell mass (ICM), blastocoel, trophectoderm (TE), zona pellucida (ZP), and background from the inside out. In particular, the TE is usually surrounded by the ZP and the blastocoel on both sides, while the ICM tends to be located within the boundary of the TE and closer to the center of the embryo. Therefore, for Figure 1 a local view like ROI 2 in (a), the combination of local inconsistencies and blurred edges requires the integration of context and non-local information for accurate analysis. Figure 1 In (b), the rotational consistency of the blastocyst tissue image is shown, and its semantic information is not lost when rotated around the center by any angle.

[0045] To address the above technical challenges in the prior art, the present invention proposes a method for biological tissue recognition based on an axially free attention network, a novel Transformer (Blast Transformer, BTFormer) architecture based on axially free attention layers and axially free attention blocks. By introducing the axially free attention mechanism, it can capture non-local information from any direction at a reduced computational cost. Such axially free attention blocks can utilize the principle of rotational consistency of blastocyst images to effectively extract global salient information. For simplicity, the BTFormer in the present invention can be referred to as the "axially free attention Transformer" or "axially free attention transducer".

[0046] The present invention proposes an axially free attention transducer (BTFormer), as Figure 2 shown, which adopts an encoder-decoder structure and incorporates multi-scale fusion. The overall architecture of the BTFormer of the present invention is as Figure 2As shown in (a), it uses a Transformer-based backbone as an encoder to extract global features. Among them, the decoder consists of upsampling blocks, including an upsampling operator and some simple convolutional layers (Conv-BN-ReLU). It retains a new independent Transformer-based encoder by replacing the convolutional blocks in the standard decoder structure with the axial-free attention blocks proposed in the present invention. The typical structure of the BTFormer of the present invention is as Figure 2 shown in (a). The first layer is the input layer, which includes the first convolutional block; the second layer includes the second convolutional block and the first axial-free attention block; the third layer includes the second axial-free attention block; the fourth layer includes the third axial-free attention block; the fifth layer includes the first upsampling block; the sixth layer includes the second upsampling block; the seventh layer includes the third upsampling block; the eighth layer is the output layer, including the fourth upsampling block; one of the two outputs of the first layer is input to the second layer, and the other is added to the output of the seventh layer and then input to the eighth layer; one of the two outputs of the second layer is input to the third layer, and the other is added to the output of the sixth layer and then input to the seventh layer; one of the two outputs of the third layer is input to the fourth layer, and the other is added to the output of the fifth layer and then input to the sixth layer; the output of the fourth layer is input to the fifth layer.

[0047] The specific structure of the axial-free attention block in the BTFormer proposed by the present invention is as Figure 2 shown in (b). It is implemented as an independent convolutional module, rather than the implementation method above the high-level convolutional position (bottleneck position) similar to Unet. It is suitable for capturing long-range dependencies. To promote a wider receptive field, the present invention collects information from a wide range of angles by applying rotational consistency. Its theoretical basis is that the embryonic image retains its physiological meaning when rotated at any angle (as Figure 1 shown in (b)). That is to say, semantic clues in any direction are used for segmentation. Therefore, the present invention evenly arranges multiple axial-free attention layers in the 2D space:

[0048] D × = π (1)

[0049] where D is the number of axial-free attention layers used, and θ represents the angle between two adjacent layers. Preferably, the present invention selects a parallel structure, D = 4.

[0050] As Figure 2 shown in (b), the typical structure of each axial-free attention block includes an input network, D axial-free attention mechanism networks, a soft aggregation operation network (SAO), a 1×1 convolutional network, and an output network; where the input network inputs the input image feature data It is divided into D paths and respectively input into D axial free attention mechanism networks, where are respectively the number of color channels, height, and width of the input image feature data; the D axial free attention mechanism networks respectively process the input image feature data in one angular direction and output the results to the soft aggregation operation network; the soft aggregation operation network performs soft aggregation on the output results of the D axial free attention mechanism networks and then outputs them to the 1×1 convolutional network; the output of the 1×1 convolutional network is added to the input image feature data and then output to the output network to obtain the output image feature data , where are respectively the number of color channels, height, and width of the output image feature data.

[0051] Typically, the present invention applies four axial free attention mechanism networks to obtain semantic information in four directions: According to Equation (1), it can be correspondingly known that among them .

[0052] The soft aggregation operation network (SAO) of the axial free attention block of the present invention, as shown in (c) in Figure 2 , is used to collect semantic information in D directions (0, 45, 90, and 135 degrees shown in the figure). By adjusting the value of D, it is possible to expand or shrink . This design enriches the context information in each direction and promotes the global receptive field. Correspondingly, this method also obtains rotational invariance. Additionally, when the input is rotated by times, the output shares the global connection from the same direction.

[0053] In the above operation combining semantic information, for each point in the feature map, the required semantic information is not always the same; therefore, it is not possible to simply add the results of these four axial free attentions. As shown in (c) in Figure 2 , the present invention assumes that the axial free attention block is composed of the outputs of D axial free attention mechanism networks, that is: . The final result is obtained through the sum of the two branches of the soft aggregation operation network (SAO), that is and , with the same input .

[0054] For Figure 2The left branch of the Soft Aggregation Operation Network (SAO) in (c) includes a two-layer structure connected in sequence. Among them, the first layer includes a 1×1 convolutional network, a BN network (Batch Normalization), and a ReLU network, and the second layer includes a SoftMax operation network; the right branch of the Soft Aggregation Operation Network (SAO) includes a two-layer structure connected in sequence. Among them, the first layer includes a Concat operation network, and the second layer includes a 1×1 convolutional network. The output of each of the D axial free attention mechanism networks is divided into three paths. Among them, the first path of D groups becomes a group of inputs to the left branch for processing after addition operation, the second path of D groups is input to the right branch for processing, and the third path of D groups is multiplied by the D-group outputs of the left branch respectively and then added to become a group added to the output of the right branch to obtain the output result.

[0055] Specifically, the left branch of the Soft Aggregation Operation Network (SAO) processes to obtain D groups of context features, and then uses an aggregation operation to aggregate them into a set; the right branch of the Soft Aggregation Operation Network (SAO) mainly implements the connection operation Cat. Thus, the processing operation of the Soft Aggregation Operation Network (SAO) can be described by the following formula:

[0056]

[0057]

[0058]

[0059]

[0060] Among them, represents the output of the d-th axial free attention mechanism network, is the processing result of the output of the first layer of the left branch, represents the convolution operation, represents the function operation; is the processing result of the left branch, which compresses the biased information into an output, maintains the same size, and generates weighted and diversified features; is the processing result of the right branch, which promotes feature fusion through a simple connection operation; is the processing result of the Soft Aggregation Operation Network.

[0061] In the Soft Aggregation Operation Network (SAO), the weights of each axial free attention layer are learned through convolution, and the sum of the results of each axial free attention layer is . Aggregate through the weighted sum of each axial free attention layer. Through the above processing method, four representative axial features from different angles (the features shown in the figure represent 0, 45, 90, and 135 degrees respectively) are aggregated and combined with the original features to form output features containing global information.

[0062] Due to the correlation between high resolution and the high cost of Transformer, the first layer and the second layer of BTFormer in the present invention include traditional convolutional blocks to extract image features. Generally speaking, in the Transformer architecture, the role of the convolutional block is to reduce the calculation and resolution of the feature map; while in the present invention, there is no pooling operation at the end of the second convolutional block of BTFormer, so it does not change the resolution of the feature map; all attention blocks of BTFormer in the present invention add an additional 2×2 max pooling operation; the resolution of the feature map will change after being processed by each axial free attention block.

[0063] Specifically, the present invention realizes the extraction of non-local correlations in any given direction at low cost through the axial free attention mechanism network. As Figure 3 shown, the axial free attention mechanism network consists of four main components: a grid generator, a direction sampler, a 1D attention module, and a learnable sine relative position encoding module. For simplicity, the following mainly describes the specific implementation method of the axial free attention mechanism network in a single-head attention manner, but it can be generalized to multi-head attention. Combining Figure 2 (b) and Figure 2 (c) shows that the axial free attention block of the present invention uses multi-head attention for each axis, the results of each head are concatenated as the output, and the whole block adopts a residual structure, and the two 1×1 convolutions in the residual branch maintain the mixture of features.

[0064] As Figure 3 shown, the role of the grid generator in the axial free attention mechanism network of the present invention is to calculate the pixel positions for generating the direction feature map Z through the given input feature map , where is the number of color channels of the input image, H is the height of the input image, and W is the width of the input image. In the present invention, the input feature map X is projected along the width or height axis direction, and its size (excluding channels) is limited to the width W×1 or height H×1 of the input feature map X; those skilled in the art can expect that this operation can reduce the calculation amount and improve the robustness.

[0065] Generally speaking, the new pixel a for generating the direction feature map Z is transformed at the original grid of the input feature map X, where In this rotation case, the projection relationship of the pixels is

[0066] (3)

[0067] Among them, defines the coordinates of the sample grid in the input feature map X, and θ represents the rotation angle. According to the above formula, it can be known that the direction information contained in the direction feature map Z can be controlled by setting the original input grid.

[0068] The direction sampler in the present invention uses a sampling grid to generate a sampled direction feature map Z for the input feature map X, which is achieved by applying a sampling kernel to the pixels as follows:

[0069] (4)

[0070] Among them, and define the interpolation operations (such as bilinear interpolation and bicubic interpolation, etc.) in the x and y directions. is the value at the position (n, m) on the feature map.

[0071] Since the sampling in the above formula is performed independently for each channel, the spatial invariance between channels is maintained. Specifically, if bilinear interpolation is used, is:

[0072] (5)

[0073] The 1D attention module in the present invention performs one-dimensional self-attention mechanism calculation on the input direction feature map Z. For each element in the input sequence, the model calculates a query vector, a key vector, and a value vector; when the model processes an element, it generates a query; then, the model compares this query with the keys of all other elements to determine which elements are relevant to the currently processed element. Finally, the model updates the representation of the currently processed element according to the values of these relevant elements. In the multi-head attention mechanism, the input data is first divided into multiple "heads", and each "head" has its own query, key, and value vectors. Each head independently performs self-attention calculation to obtain its own attention output, and finally these outputs are concatenated together to form the final output.

[0074] Specifically, in Figure 3 , the one-dimensional axial free attention along the angle θ of the given parameter is defined as follows:

[0075] (6)

[0076] where θ is the rotation direction, i.e., the angle between the attention axis and the width axis. Here, the query , the key , and the value are linear transformations of the pixels in the direction map, where are learnable weights, the same as the setting of self-attention; , is the element value in the direction feature map Z. In addition, , and are learnable sine relative position encodings. Considering that the input width W is not always the same as the height H, project the direction onto the height axis. When θ is greater than , to obtain multi-range dependencies, the formula is as follows:

[0077] (7)

[0078] where , are the coverage areas of the axial-free attention.

[0079] In particular, when θ = 0, Equation (6) is the same as the width-axis attention, and the axial-free attention turns to the width-axis attention

[0080] (8)

[0081] When θ becomes , the axial-free attention is converted to the height-axis attention,

[0082] (9)

[0083] By performing the above axial-free attention, the computational cost can be reduced to . At the same time, it can capture long-range dependencies instead of focusing on local interactions.

[0084] The present invention proposes learnable sine relative position encoding (LSRPE) to introduce continuous bias and adaptability. In the learnable sine relative position encoding, the position encoding provides information about the spatial relationship. Usually, an additional position bias term is added to the value, which is usually used to make the model sensitive to the position information.

[0085] Given a pixel coordinate a = (i, j) and another pixel coordinate b = (l, n), the learnable sine relative position encoding (LSRPE) is defined as follows:

[0086] (10)

[0087] Among them, the learnable parameters , , are relative position encodings. All position encodings for keys, queries, and values here are random and independent.

[0088] In the above model, the weighted soft Jaccard index approximation (WSJA) loss is adopted. The soft Jaccard index for class i is calculated as follows:

[0089] (11)

[0090] where is a smoothing constant, are the true label and the predicted label of class i, respectively. The loss function is completed by summing SJA over five classes and applying weighted attention to the two smallest SJA classes, i.e.:

[0091] (12)

[0092] Preferably, the weights are set to 1.0 and 0.8, respectively.

[0093] The present invention was tested and evaluated on a public blastocyst dataset, which contains 249 blastocyst images and their masks. There are five classes in the dataset: ZP, TE, ICM, blastocoel, and background. 199 blastocyst images were randomly selected as our training set, and the remaining 50 images constitute the test set. Five evaluation metrics are used: accuracy, precision, recall, Dice coefficient (Dice), and Jaccard index (Jaccard), and their definitions are as follows:

[0094]

[0095]

[0096]

[0097]

[0098]

[0099] These metrics are defined based on four parameters: TP (true positive), FP (false positive), TN (true negative), and FN (false negative).

[0100] In the test, 2000 cycles of training were carried out, and specific hyperparameters were used: the initial learning rate was 0.0008, the momentum was set to 0.9, and the weight decay was 0.0005. To adaptively adjust the learning rate, the Poly learning rate strategy was adopted. The network parameters were updated using the Adam optimizer. The specific implementation used PyTorch (version ≥1.10), and both training and testing were performed on a 32GB NVIDIA Tesla V100 GPU. Before inputting the images into the model for training and testing, their sizes were normalized to 256×256. In addition, during the training phase, data augmentation was performed on each sample, including random scaling, rotation, cropping, and horizontal flipping.

[0101] The baselines for evaluation and comparison include two types: CNN-based methods such as U-Net, AttUNet, Unet++, Pranet, nnUnet, Deeplabv3+, Unet3+, FCN, PSPNet, DANet, CCNet, and Blast-Net; and transformer-based methods such as SETR, TransUNet, Segtran, Swin-Unet, UCTransnet, MedT, PCS-Net, and SPC-Net. The baselines were trained using the same binary cross-entropy (CE) loss between the predictions and the ground truth labels.

[0102] The test results show that the present invention achieved the best performance, with accuracy, precision, recall, Dice, and Jaccard being 93.86%, 91.81%, 92.25%, 92.02%, and 85.45% respectively.

[0103] The test results also show that the selection of the number D in Equation (1) is crucial for forming the axial free attention block. First, the present invention tried to explore the influence of different Ds on the receptive field; when D equals 1, the model only obtained a narrow effective receptive field along the width axis; when D equals 2, the axial free attention layers would cross, and in this case, the BTFormer collected dense attention from both the width axis and the height axis; as the number of Ds increased, the dense attention became scattered and filled the entire image, which means that a larger D can obtain global representations and a broad receptive field. Then, the Dice and Jaccard comparisons of different Ds were also made; the BTFormer with a larger D consumed more computations, and a larger number of Ds did not always ensure good performance. Therefore, considering computational efficiency and performance, D cannot be too large. Specifically, BTFormer (D = 4) is the best choice, with the highest Dice and Jaccard and lower computational costs.

[0104] Through the detailed description of the specific embodiments and examples of the present invention above, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments and examples without departing from the principles and basic concepts of the present invention. The scope of the present invention is defined by the appended claims and their equivalents. The content not described in detail in this specification, such as the specific structures and implementation details of typical neural networks and their functional layers, can be achieved by those of ordinary skill in the art based on the well-known prior art and the technical teachings of the present invention.

Claims

1. A biological tissue recognition method based on axial free attention network, characterized in that Includes steps: An axial free attention transformer is constructed. The axial free attention transformer adopts an encoder-decoder structure, including multiple convolution blocks, axial free attention blocks and upsampling blocks, wherein each axial free attention block includes an input network, D axial free attention mechanism networks, a soft aggregation operation network, a 1×1 convolution network and an output network; each axial free attention mechanism network processes the input image feature data according to different angular directions, and includes: a grid generator, a direction sampler, a one-dimensional attention module and a learnable sinusoidal relative position encoding module; wherein the grid generator is used to calculate the pixel position relationship between the input feature map and the generated direction feature map Z, the direction sampler generates the sampled direction feature map Z using a sampling grid, the one-dimensional attention module is used to perform one-dimensional self-attention mechanism calculation processing on the direction feature map Z, and the learnable sinusoidal relative position encoding is used to introduce continuous bias and adaptability; Acquire training sample data, where the training sample data is biological tissue image data, and the biological tissue image data is labeled data; Using the training sample data to train the axial free attention transformer to obtain a biological tissue recognition processing training model; Use biological tissue recognition and processing training models to recognize and process biological tissue image data; Among them, the input network of each axial free attention block divides the input image feature data into D paths and inputs them into D axial free attention mechanism networks respectively. The D axial free attention mechanism networks process the input image feature data according to an angle direction respectively, and output the results to the soft aggregation operation network; for the number D of the axial free attention mechanism networks, D × θ = π, where θ represents the angle between the attention axes processed by adjacent axial free attention mechanism networks.

2. The biological tissue recognition method based on axial free attention network as claimed in claim 1, characterized in that: The number D of the axial free attention mechanism networks is 4 or more.

3. The biological tissue recognition method based on axial free attention network as claimed in claim 2, characterized in that: The soft aggregation operation network soft-aggregates the output results of the D axis-free attention mechanism networks and outputs them to the 1×1 convolutional network; the output of the 1×1 convolutional network is added to the input image feature data and then output to the output network to obtain the output image feature data.

4. The biological tissue recognition method based on axial free attention network as claimed in claim 3, characterized in that: The pixel position relationship between the input feature map X and the directional feature map Z calculated by the grid generator is: ; in, is the coordinate of the sample grid in the input feature map, is the angle of rotation of the processed attention axis relative to the coordinate axis, is the coordinate of the new pixel used to generate the directional feature map Z.

5. The biological tissue recognition method based on axial free attention network as claimed in claim 4, characterized in that: The kernel function of the sampling grid used by the direction sampler is: ; in, and Defines the interpolation operation in the x and y directions, is the value of the (n, m) position on the feature map, and H and W are the height and width of the image data respectively.

6. The biological tissue recognition method based on axial free attention network as claimed in claim 5, characterized in that: For a given pixel coordinate and another pixel coordinate , the learnable sinusoidal relative position encoding (LSRPE) is defined as: ; Among them, the learnable parameters , , is a relative position encoding.

7. The biological tissue recognition method based on axial free attention network as claimed in claim 6, characterized in that: The specific structure of the axial free attention converter includes: The first layer, including the first convolutional block; The second layer, includes the second convolution block and the first axis-free attention block; The third layer, includes the second axial free attention block; The fourth layer includes the third axis free attention block; The fifth layer, including the first upsampling block; The sixth layer, including the second upsampling block; The seventh layer, including the third upsampling block; The eighth layer, including the fourth upsampling block; Among them, the first layer is the input layer, and the eighth layer is the output layer. One of the two outputs of the first layer is input to the second layer, and the other is added to the output of the seventh layer and then input to the eighth layer; one of the two outputs of the second layer is input to the third layer, and the other is added to the output of the sixth layer and then input to the seventh layer; one of the two outputs of the third layer is input to the fourth layer, and the other is added to the output of the fifth layer and then input to the sixth layer; the output of the fourth layer is input to the fifth layer.

Citation Information

Patent Citations

  • Retinal vessel segmentation method based on gated axial self-attention double-coding convolutional neural network

    CN116309629A

  • Squeeze-enhanced axial transformer, its layer and methods thereof

    US11776240B1