A method for image prediction classification using a context reasoning network based on Graphormer
By adopting a Graphormer-based context reasoning network in small object detection, and using Transformer to aggregate multiple coded information to update node characteristics, the problems of information loss and redundancy in complex environments in the prior art small-object detection are solved, and higher detection accuracy and efficiency are achieved.
Patent Information
- Application Number
- CN202210940698.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-06
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-08-06
AI Technical Summary
Existing small object detection methods are difficult to effectively detect small objects in the real world with complex environments, especially in dense scenarios. The existing context learning methods are limited by the receptive field size, resulting in the loss of important information, and there is information redundancy in the global context information.
A contextual reasoning network based on Graphormer is used to extract image features through the Backbone network, generate regional suggestions, construct graph structure models, and use Transformer to aggregate central encoding, semantic encoding and spatial layout encoding information between nodes to update node features, and obtain classification results through the Softmax function.
It effectively improves the accuracy of small object detection, enhances the characterization ability of the graph, reduces information redundancy, and improves the efficiency of the model.
Smart Images

Figure CN115457345B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image classification methods, and particularly relates to a method for predicting and classifying pictures by using a context reasoning network based on Graphormer. Background Art
[0002] As a key technology in aerial image analysis, small target detection has become a popular research topic in the field of computer vision. However, challenges in small target detection still exist, especially in detecting small targets in the complex real world: (1) Small targets occupy a small proportion of pixels and have few available features; (2) Existing context learning methods are limited by the size of the receptive field, which may lead to the loss of important information.
[0003] When extracting features, most existing context learning methods rely on the design of context windows or are limited by the size of the receptive field, which may lead to the loss of important context information. Therefore, in order to make more full use of context information, some methods attempt to integrate global context information into the target detection model. However, in dense small object scenes, a lot of useless feature information will be obtained in the global context information, resulting in information redundancy and increasing the cost of training the model. Therefore, how to find valuable context information from the global scene is the current research focus. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a method for predicting and classifying pictures by using a context reasoning network based on Graphormer in view of the above problems and requirements.
[0005] To solve the above technical problems, the present invention adopts the following technical solutions:
[0006] A method for predicting and classifying pictures by using a context reasoning network based on Graphormer, comprising the following steps:
[0007] Step 1: Use the Backbone network to extract features from the input picture, and input the extracted features into the RPN network to generate region proposals;
[0008] Step 2: The context relationship modeling module constructs a graph structure model by using the generated region proposals. Each edge in the graph structure model contains centrality encoding, semantic encoding, and spatial layout encoding information. Among them, the centrality encoding is constructed according to the degree of the target node, the semantic encoding is constructed based on the sparse semantic relationship of the initial region features, and the spatial layout encoding is constructed based on the sparse spatial layout relationship of the position and shape information.
[0009] Step 3: The context relationship reasoning module aggregates three types of encoding information between nodes through Transformer, further fuses it with the initial features to update the node information, then uses the Softmax function to obtain the output probability distribution of all nodes, gets the probability of each node belonging to each category, takes the category with the highest probability as the label of the node, and finally outputs the labels of all nodes as the classification result of the input image.
[0010] Further, in step 2, the centrality encoding is a learnable embedding vector assigned to each node according to the degree of each node.
[0011] Further, in step 2, by calculating the semantic similarity of each edge, the semantic similarity is input into a multi-layer perceptron to obtain the semantic encoding of each edge; for edge e ij The formula for calculating the semantic similarity is: where δ(i, j) is an indicator function, which is equal to 0 if the nodes i and j at both ends of edge e ij highly overlap with each other, and equal to 1 otherwise, P O ∈R n×d is the given initial region feature library, d is the dimension of the initial region feature, f(·, ·) is a defined learnable semantic correlation function used to calculate the semantic correlation of each pair of initial node features in the original fully connected graph, and φ(·) is a projection function that projects the initial region feature into a latent representation.
[0012] Further, in step 2, by calculating the spatial layout similarity of each edge, the spatial layout similarity is input into a multi-layer perceptron to obtain the spatial layout encoding of each edge; for edge e ij The formula for calculating the spatial layout similarity s″ ij is:
[0013]
[0014] where and are the regional coordinates of nodes i and j at both ends of edge e ij respectively, and are the spatial similarity and spatial distance weight between nodes i and j respectively:
[0015]
[0016]
[0017] where λ is a scale parameter, is the spatial distance between the centers of nodes i and j.
[0018] Further, step 2 specifically includes the following steps:
[0019] Step 2.1: Update the nodes. The steps for updating the i-th node in the first layer, i.e., node i, are as follows: Let the initial node feature V of node i i be denoted as x i , and add it to the central node feature matrix to obtain which represents the hidden feature h of node i, where d is the hidden dimension. Take the node feature and the features of other nodes in the graph structure as the input of the self-attention module, and then multiply them with three weight matrices and respectively to obtain vectors Q, K, and V;
[0020] Perform a dot product calculation on vector Q and vector K to obtain the correlation weight A of node i. Then normalize the correlation weight A of node i, and perform a matrix addition calculation on the semantic encoding s′ ij and the spatial encoding s″ ij as the bias term of weight A. Convert the weight A of node i into a probability distribution between [0, 1] through the softmax function, and then perform a dot product calculation on the probability distribution of each node and vector V. Finally, accumulate the results of the dot product calculation to obtain the updated node feature of node i;
[0021] Step 2.2: After updating all nodes, the feature sequence H of all updated nodes output by the current Graphormer layer (l+1) ;
[0022] Step 2.3: Input the feature sequence H (l+1) into the Decoder to calculate the self-attention weights, and obtain the updated node feature sequence H (l) ;
[0023] Step 2.4: Input the updated node feature sequence H (l) into the fully connected layer FC for a linear transformation. Each node will be connected to a fully connected layer. Use the Softmax function to obtain the probability distribution of the outputs of all nodes, and obtain the probability of each node belonging to each category. Take the category with the highest probability as the label of the node, and obtain the prediction output results of all nodes in the input image.
[0024] Further, in the testing phase, the context relationship modeling module and the context relationship reasoning module are the final modules that have undergone the training phase. In the training phase, the images input in step 1 are the images in the training set, and step 3 further includes the following steps: calculating the loss function, and adjusting the parameters of the context relationship modeling module and the context relationship reasoning module by minimizing the loss function. After training, the final modules are obtained.
[0025] After the present invention adopts the above technical solutions, compared with the prior art, it has the following advantages:
[0026] Based on the traditional feature extraction method, the present invention innovatively combines the advantages of the two methods by using the Graphormer method. Since the graph neural network (GCN) adopts local perception regions, weight sharing, and downsampling in the spatial domain, it can well extract the spatial features of images. However, due to the local receptive field of GCN, its calculation process is limited to the feature information of neighboring nodes. On the contrary, Transformer can regard each node in the graph as connected to calculate the Attention Bias between pairwise nodes; the present invention uses Transformer for combined feature aggregation / propagation and feature transformation to solve the problem of few challenging scale features in the field of small target detection.
[0027] In the method of representing the graph structure, the traditional method only considers the semantic and spatial similarities between targets, while ignoring the characteristics that the nodes themselves may have. The present invention constructs degree centrality encoding considering that the more degrees each node has, the greater its weight may be, and inputs the centrality encoding as the bias term of attention into Transformer. This degree of centrality can make the model pay more attention to some special nodes, thereby enhancing the representation ability of the graph and effectively improving the accuracy of small target detection.
[0028] The present invention will be described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings
[0029] Figure 1 is the overall flowchart of the method of the present invention;
[0030] Figure 2 is the flowchart for obtaining the semantic encoding;
[0031] Figure 3 is the flowchart for obtaining the spatial layout encoding;
[0032] Figure 4 is the flowchart for calculating the feature weights;
[0033] Figure 5 is the flowchart for feature transfer of the Graphormer layer;
[0034] Figure 6 Schematic diagram of the Encoder-Decoder structure in Graphormer;
[0035] Figure 7 Schematic diagram of the qualitative detection results of Coni-GT on MS COCO;
[0036] Figure 8 Qualitative detection result graph of Coni-GT on the TinyPerson dataset. Detailed implementation manners
[0037] The principles and features of the present invention will be described below in conjunction with the accompanying drawings. The examples given are only used to explain the present invention and are not intended to limit the scope of the present invention.
[0038] I. As Figure 1 shown, the overall technical solution of this project includes the following steps:
[0039] 1. First, the input image will undergo feature extraction through the Backbone network, and then the extracted features will be input into the RPN network to generate region proposals. We use the generated region proposals to construct a graph structure model. The context modeling module is used to extract the node features of the graph structure, and the context reasoning module is used to aggregate the node feature information and update the node feature information. Finally, a fully connected network layer (FC) is used to classify and regress the nodes.
[0040] 2. Context relationship modeling module: Extract the node feature information of the graph. We use Faster R-CNN as the baseline and use the generated region proposals to construct a graph structure model G = <V, E>, where V = {v1, v2,..., v n}, n = |V| is the number of region proposal nodes, and each edge e ij < E contains centrality encoding, semantic encoding, and spatial layout encoding information. Among them, the centrality encoding is constructed according to the degree of the target node, the semantic encoding is constructed based on the sparse semantic relationship of the initial region features, and the spatial layout encoding is constructed based on the sparse spatial layout relationship of the position and shape information.
[0041] Context Inference Module: Aggregate the context information of nodes. Aggregate three types of encoded information between nodes through a Transformer, further fuse it with the initial features to update the node information, then use the Softmax function to obtain the probability distribution of the outputs of all nodes, get the probability of each node belonging to each category, and take the category with the highest probability as the label of the node to obtain the predicted output results of all nodes in the input image.
[0042] II. Implementation Method
[0043] The present invention has invented a new algorithm called Coni-GT, which is mainly divided into two parts: 1. Context Relationship Modeling; 2. Context Inference Module.
[0044] (1) Context Relationship Modeling
[0045] 1. Construct Centrality Encoding
[0046] In daily life, celebrities with a large number of fans usually have greater influence, and this phenomenon is an important factor in predicting social network trends. Similarly, we can associate that in a graph structure, the more neighbor nodes a node has, the greater its influence factor. This degree centrality feature's impact on the prediction result is easily overlooked, but according to our research, it is an essential factor in target prediction. In Coni-GT, we construct degree centrality encoding as an additional signal for the neural network.
[0047] Specifically, we propose a centrality encoding that assigns a real-valued embedding vector to each node according to its degree. Since centrality encoding is applied to each node, we only need to add it as an input to the node features.
[0048]
[0049] where x i is the feature vector V i , represents the feature representation of V i at the initial layer, z ∈ R d is a learnable embedding vector specified by the degree of the node deg(v i ), that is, the centrality encoding. For a directed graph, deg(v i ) can be split into in-degree deg - (v i ) and out-degree deg + (v i) By using centrality encoding in the input, the Softmax attention in the Transformer can capture the node importance signals in the Queries and Keys. Therefore, the model can capture the importance of nodes.
[0050] 2. Construct semantic encoding
[0051] The human visual system can distinguish object categories based on the similarity between objects, which proves that there are the same intrinsic semantic relationships among objects of the same category. We can use this semantic similarity to identify objects that are difficult to recognize in a scene. Specific examples are as Figure 2 shown. In the same scene, the semantic information of easily detectable people is used to identify blurred people. The context information of such easily detectable objects often helps to identify objects that are difficult to detect.
[0052] In the real world, most of the connections of object interactions are invalid. So when we construct semantic encoding, we need to calculate the semantic correlation degree between nodes in the fully connected graph G and retain the relationships with high correlation degree, and prune the relationships with low correlation degree. Therefore, the semantic similarity calculation formula is:
[0053]
[0054] where δ(i, j) is an indicator function that is equal to 0 if the i-th and j-th regions overlap highly with each other, and equal to 1 otherwise. P O ∈R n×d is the given initial region feature library, d is the dimension of the initial region features. f(·, ·) is a defined learnable semantic correlation function used to calculate the semantic correlation degree of each pairwise initial region feature in the original fully connected graph. φ(·) is a projection function that projects the initial region features into a latent representation. Since different regions are parallel and there is no distinction between the main and object, in this paper, it is set as a multi-layer perceptron (MLP) to encode undirected relationships.
[0055] 3. Spatial layout encoding
[0056] One advantage of the Transformer is that it has a global receptive field. In each Transformer layer, each token can attend to information at any position and then process its representation. However, this operation has a side effect that the model must explicitly specify different positions or encode position dependencies (such as positional) in each layer. For sequential data, an embedding (i.e., absolute position encoding) can be given to each position as input, or the relative distance between any two positions can be encoded in the Transformer layer (i.e., relative position encoding).
[0057] However, for graphs, the nodes are not arranged as a sequence. They can be located in a multi-dimensional space and connected by edges. To encode the structural information of graphs in the model, we propose a new spatial encoding based on the characteristics that small objects belonging to the same class in the same scene tend to have similar spatial aspect ratios and small objects of the same category are clustered together in the spatial layout. An example is Figure 3 as shown. Specifically, for a fully connected graph G, we define a spatial layout correlation function g(·, ·) to calculate the correlation in the original fully connected graph. The spatial layout similarity s″ ij is calculated by the formula
[0058]
[0059] where and are the node coordinates corresponding to nodes i and j respectively. and are the spatial similarity and the spatial distance weight respectively.
[0060]
[0061]
[0062] where λ is the scale parameter, which is empirically set to 5e -4 . is the spatial distance between the centers of the two nodes.
[0063] (2) Context relationship reasoning
[0064] 1. Transformer layer
[0065] The attention mechanism of Transformer enables it to have a global receptive field, so that the model can pay attention to the global context information in the scene. Graphormer is a model that applies Transformer to graph structures, and it proves that Transformer has the ability to pass messages in non-Euclidean spaces. Inspired by this, we propose to use Graphormer to perform feature aggregation and feature transformation operations to solve the problem of limited receptive fields of existing context learning methods, enabling the model to enrich the target feature information from a global perspective of the scene.
[0066] The architecture of Graphormer is composed of Graphormer layers. Each Graphormer layer consists of two parts: a self-attention module (MHA) and a position-wise feed-forward network (FFN). In the context reasoning module, we use the Graphormer network to aggregate centrality encoding, semantic encoding, and spatial encoding. These three encodings will be used as the bias terms of the multi-head attention (MHA) in the Encoder, enabling message passing between nodes to update the feature information of the nodes.
[0067] Figure 4 It is the calculation process of the feature weights of the Encoder in the Graphormer layer for nodes. Specifically, we let the initial node feature V i be denoted as x i , and add it to the central node feature matrix to obtain which represents the hidden feature h of the i-th node in the l-th layer, where d is the hidden dimension. We take the node feature and the features of other nodes in the graph structure as the inputs of the self-attention module, and then multiply them with the three weight matrices and respectively, and project them to the corresponding representations Q, K, V. We use the dot product of vector Q and vector K to calculate the correlation weight A between each node, and then normalize the correlation weight A between each node. The purpose of normalization is mainly to make the gradient stable during training. At the same time, we also use the semantic encoding s′ ij and the spatial encoding s″ ij as the bias terms of the weight A for matrix addition calculation. We convert the weight A between each node into a probability distribution between [0, 1] through the softmax function, and then perform a dot product calculation between the probability distribution between each word and the corresponding Values. Finally, by accumulating the results of the dot product calculation, we can obtain the updated node feature of node i The entire calculation process of node update can be expressed by the formula:
[0068]
[0069]
[0070]
[0071] Each node in the graph will go through the above calculation process in parallel. Finally, the output of the current Graphormer layer is the feature sequence H (l+1) of all updated nodes.
[0072] 2. Implementation Details
[0073] Graphormer is implemented based on the basic Transformer network. Figure 5 It is the message passing process of node feature information in the Graphormer layer. First, input the node feature H of the previous network layer (l-1) To the application network layer (Add&Normalize), and perform LN normalization. Then, the normalized data is input into the multi-head attention module (MHA) and combined with the original layer feature H (l-1) Calculate the attention weights between nodes together to get the l-layer node representation matrix H′ (l) Finally, the l-layer node feature information H′ (l) It will be normalized (LN) and then input into the feedforward neural network layer (FFN) and combined with the feature information H′ input in the current layer. (l) Update the current layer node feature information H together (l) The above description can be expressed by the formula:
[0074] H′ (l) =MHA(LN(h (l-1) ))+H l-1 (9)
[0075] H (l) =FFN(LN(h′ (l) ))+H′ (l) (10)
[0076] The node feature sequence in the Graphormer layer updated by the Encoder is input into the Decoder for self-attention weight calculation, and finally the updated node feature sequence H is obtained. (l) The node feature sequence updated by the Decoder will be input into the fully connected layer FC after a linear transformation, where each node will be connected to a fully connected layer, and then the Softmax function will be used to obtain the output probability distribution. Then, through the label of each node, the corresponding node with the highest probability will be output as our prediction output.
[0077] During the training phase, the loss L is calculated while performing linear transformation on each node. G , L G =∑ v∈V y v log(σ(H v θ))+(1-y v )log(1-σ(H v θ)) (11) where V = {v1, v2, ..., v n} is the set of nodes in graph G, yv is a set of node class labels, H v is a set of output node feature embeddings, θ is the weight of the fully connected layer, and σ is the set learning rate.
[0078] In the training phase, the model is trained by minimizing the loss function and the parameters of the model are adjusted. In the prediction phase, the trained model is used for prediction.
[0079] The Encoder-Decode structure of the Graphormer layer is as Figure 6 shown. The structures of the Encoder and Decoder are similar, both being a combination of the multi-head self-attention mechanism and the feed-forward neural network. The way they perform message passing is as Figure 5 shown. The input to the Encoder is the node features obtained by adding the initial node features and the centrality node feature matrix, and the output is the updated node feature sequence for each node. The Decoder has two inputs. The first input is the classification label corresponding to each node, and the second input is the output of the last layer in the Encoder. The Decoder has two self-attention modules. The first Masked Multi-Head Attention is to obtain the information of the previously predicted output, which is equivalent to recording the information between the current inputs. The second Multi-Head Attention is to represent the relationship between the current input and the feature vector extracted by the encoder to predict the output. The final output of the Decoder, like that of the Encoder, is the feature sequence for each node. We input the updated features of each node into the fully connected layer and finally obtain the class and prediction probability of each node through the Softmax function.
[0080] Experimental results of small object detection
[0081] To verify the effectiveness of the proposed Coni-GT, we conducted extensive experiments on the challenging MS COCO dataset and Tinyperson dataset and provided the qualitative results on these two datasets.
[0082] Figure 7 are the qualitative results of Coni-GT on the MS COCO dataset. MS COCO is a common publicly available dataset in the field of object detection. This dataset contains 140,000 images (80,000 for training, 40,000 for validation, and 20,000 for testing) and is richly annotated for 91 classes of objects.
[0083] Figure 8Shows some visual detection results of Coni-GT and Mask R-CNN on the MS COCO dataset. From the figure, we can observe that Mask R-CNN fails to detect some small objects. However, Coni-GT works well on small objects by using appearance and context information. It can also be found that Coni-GT is more robust than Mask R-CNN. Mask R-CNN suffers from appearance changes and background interference. As Figure 4 shown in the third row, Mask R-CNN detects the bridge opening and the building as a ship. However, Coni-GT can handle these interferences well by leveraging the context information in the surrounding scene.
[0084] The TinyPerson dataset is a benchmark for person detection in the context of long distances and large backgrounds, opening up a new promising direction for extremely small object detection. Figure 5 Shows some detection results of our proposed Coni-GT on the TinyPerson dataset. Among them, the first two rows represent the scenes of people at sea, and the last row represents the scenes of people on land. From the first two rows, it can be observed that the postures of the people on the surfboard vary greatly, and even some people only show their heads under backlight conditions, but our Coni-GT can still accurately identify and locate them. In addition, the last row contains scale variations, crowd scenes, and cluttered backgrounds. Even so, most people can still be identified and located. This proves the effectiveness of our detector for small object detection.
[0085] The above are examples of the best implementation modes of the present invention, and the parts not described in detail are the common general knowledge of those skilled in the art. The protection scope of the present invention shall be subject to the content of the claims, and any equivalent transformation based on the technical inspiration of the present invention shall also be within the protection scope of the present invention.
Claims
1. A method for image prediction classification using a context reasoning network based on Graphormer, characterized in that, Including the following steps: Step 1: Extract features from the input image, and input the extracted features into the RPN network to generate region proposals; Step 2: The context relationship modeling module constructs a graph structure model using the generated region proposals. Each edge in the graph structure model contains centrality encoding, semantic encoding, and spatial layout encoding information. Among them, the centrality encoding is constructed according to the degree of the target node, the semantic encoding is constructed based on the sparse semantic relationship of the initial region features, and the spatial layout encoding is constructed based on the sparse spatial layout relationship of the position and shape information; In step 2, by calculating the semantic similarity of each edge and inputting the semantic similarity into a multi-layer perceptron, the semantic encoding of each edge can be obtained; for edge e ij , the formula for calculating the semantic similarity is: where δ(i, j) is an indicator function, which is equal to 0 if the nodes i and j at both ends of edge e ij highly overlap with each other, and equal to 1 otherwise. P O ∈R n×d is a given initial regional feature library, d is the dimension of the initial regional feature, and f(·, ·) is a defined learnable semantic correlation function used to calculate the semantic correlation of each pairwise initial node feature in the original fully connected graph. Φ(·) is a projection function that projects the initial regional feature into a latent representation; In the above-mentioned step 2, by calculating the spatial layout similarity of each edge and inputting the spatial layout similarity into a multi-layer perceptron, the spatial layout code of each edge can be obtained; for edge e ij the spatial layout similarity s″ ij The calculation formula is as follows: wherein and are respectively the regional coordinates of node i and node j at both ends of edge e ij , and and are respectively the spatial similarity and spatial distance weight of node i and node j where λ is the scale parameter, is the spatial distance between the centers of node i and node j. Step 3: The context relationship reasoning module aggregates the three types of encoding information between nodes through Transformer, and further fuses with the initial features to update the node information. Then, the Softmax function is used to obtain the output probability distribution of all nodes, obtain the probability of each node belonging to each category, take the category with the largest probability as the label of the node, and finally output the labels of all nodes as the classification result of the input image.
2. The method for image prediction classification using a context reasoning network based on Graphormer according to claim 1, characterized in that, In step 2, the centrality encoding is a learnable embedding vector assigned to each node according to the degree of each node.
3. The method for image prediction classification using a context reasoning network based on Graphormer according to claim 1, characterized in that, Step 2 specifically includes the following steps: Step 2.1: Update the node. The steps for updating the \(i\)-th node in the \(l\)-th layer, i.e., node \(i\), are as follows: Let the initial node feature \(V\) of node \(i\) i be denoted as \(x\) i , and add it to the central node feature matrix to obtain which represents the hidden feature \(h\) of node \(i\). Here, \(D\) is the hidden dimension. Take the node feature and the features of other nodes in the graph structure as the inputs of the self-attention module, and then multiply them with three weight matrices and respectively to obtain the vector matrices \(Q\), \(K\), and \(V\); Perform a dot product calculation on vector Q and vector K to obtain the relevance weight A of node i, and then normalize the relevance weight A of node i. Take the semantic encoding s′ ij and the spatial encoding s″ ij as the bias terms of weight A for matrix addition calculation. Convert the weight A of node i into a probability distribution between [0, 1] through the softmax function, and then perform a dot product calculation on the probability distribution of each node and vector V. Finally, accumulate the results of the dot product calculation to obtain the updated node feature of node i Step 2.2: After updating all nodes, the feature sequence H of all nodes updated by the current Graphormer layer (l+1) ; Step 2.3: Input the feature sequence H (l+1) into the Decoder to calculate the self-attention weights, and obtain the updated node feature sequence H (l) ; Step 2.4: Input the updated node feature sequence H (l) into the fully connected layer FC for a linear transformation. Each node is connected to a fully connected layer. The Softmax function is used to obtain the probability distribution of the outputs of all nodes, and the probability of each node belonging to each category is obtained. The category with the highest probability is used as the label of the node, and the predicted output results of all nodes in the input image are obtained.
4. The method for image prediction classification using the context reasoning network based on Graphormer according to claim 1, wherein In the test stage, the context relationship modeling module and the context relationship reasoning module are the final modules after the training stage. In the training stage, the image input in step 1 is the image in the training set. Step 3 also includes the following steps: calculate the loss function, adjust the parameters of the context relationship modeling module and the context relationship reasoning module by minimizing the loss function, and obtain the final module after training.