Target detection method based on multi-task model
The multi-task model integrates target detection and semantic segmentation tasks using a deformable convolution network and attention mechanism, reducing computational load and enhancing detection accuracy and scene understanding in monitoring systems.
Patent Information
- Application Number
- CN202510402096.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-15
AI Technical Summary
In the prior art, the independent operation of object detection and semantic segmentation tasks leads to high computational complexity, high resource consumption and insufficient information sharing, which limits the accuracy and reliability of monitoring environment perception.
Using the object detection method based on multi-task model, by building a multi-task model, integrating the object detection decoder and semantic segmentation decoder, using deformable convolutional network, SimAM attention mechanism and Wise-IoU loss function, synchronous processing and resource sharing of object detection and semantic segmentation are realized, and small object detection capabilities are enhanced.
It reduces the computing pressure of the monitoring computing platform, improves the efficiency and accuracy of target detection, enhances the detection ability of small targets, and improves the perceived performance and reliability of the monitoring and alarm system.
Smart Images

Figure CN120318761A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of object detection, and particularly relates to an object detection method based on a multi-task model. Background Art
[0002] In the monitoring and alarm scenario, visual perception plays a crucial role. It mainly performs operations such as object detection and recognition, behavior analysis, and intrusion detection on targets, enabling it to obtain rich image information in complex environments and providing a reliable basis for security monitoring.
[0003] Object detection and semantic segmentation, as the core tasks in the field of visual perception, have irreplaceable application values in the monitoring scenario. Object detection is dedicated to identifying the object categories in an image and determining their positions, providing key object positioning information for monitoring and realizing object intrusion and left-behind detection; semantic segmentation focuses on pixel-level semantic information extraction, which can accurately distinguish different scene elements such as people, vehicles, animals, and buildings, and construct a more refined environmental model for the monitoring scenario.
[0004] However, at the current stage of technological development, the two tasks of object detection and semantic segmentation often run independently. This means that the monitoring computing platform needs to separately process and calculate data for the two tasks, which not only increases the computational complexity but also results in a large consumption of computing resources, bringing huge pressure to the monitoring computing platform. In addition, there is a lack of effective information sharing and cooperation between the two independently running tasks, making it difficult to fully utilize the overall information in the image and restricting the accuracy and reliability of environmental perception. This series of problems restricts the further development of the monitoring and alarm technology, and there is an urgent need for a technical solution that can effectively integrate the functions of object detection and semantic segmentation, reduce the computational pressure, and improve the perception performance. Summary of the Invention
[0005] The purpose of the present invention is to provide an object detection method based on a multi-task model. The present invention can simultaneously perform object detection and semantic recognition, and has the advantage of high efficiency.
[0006] The technical solution of the present invention: The object detection method based on a multi-task model is carried out according to the following steps:
[0007] Step S1: Construct a multi-task model. The multi-task model includes an encoder, which is connected to an object detection decoder and a semantic segmentation decoder. The object detection decoder has a decoupled head; the encoder includes a convolutional network and a feature fusion module. The convolutional network is used to extract features from the input data, and the feature fusion module is used to fuse the features extracted by the convolutional network; after obtaining the fused features of the encoder, the object detection decoder performs object detection. The decoupled head is used to generate a classification branch and a regression branch for the classification task and regression task of object detection, and convergence is achieved after generating the object detection box; after obtaining the fused features of the encoder, the semantic segmentation decoder performs semantic segmentation;
[0008] Step S2: Train the constructed multi-task model, and perform object detection and semantic recognition using the trained multi-task model.
[0009] In the above object detection method based on a multi-task model, the convolutional network is a deformable convolutional network. While obtaining data features, the deformable convolutional layer network obtains the offset in the data through the difference algorithm of the convolutional layer, and adjusts its parameters according to the offset through the backpropagation algorithm; the output calculation of the deformable convolutional network is shown in the following formula:
[0010]
[0011] In the formula, y is the new pixel coordinate, R = {(-1, -1), (-1, 0), …, (0, 0), …, (1, 0), (1, 1)}, P0 is any point on the input feature map, P n is the offset of each point in the convolution kernel relative to the center point, ΔP n is the offset offset, P0 + P n + ΔP n is the new coordinate, X(P0 + P n + ΔP n ) is to take out the pixel value at this position, W(P n ) is the weight value of the pixel at the relative position of the convolution kernel; iteratively add the values of all positions W(P n ) * X(P0 + P n + ΔP n ) to output the result pixel value.
[0012] In the aforementioned object detection method based on a multi-task model, a SimAM attention mechanism is added before the feature fusion module. The SimAM attention mechanism is used to strengthen the feature representation of small objects, and the calculation is shown in the following formula:
[0013]
[0014] Wherein, t is the target neuron, x represents the neighboring neurons, λ is a hyperparameter, M = H × W, H and W represent the height and width of the energy function on each channel, and M represents the number of energy functions on each channel. represents the mean value on each channel in x, and X is the input feature map. is the output feature map. represents the variance on each channel in x. represents the distinguishability between the target neuron and the neighboring neurons.
[0015] In the foregoing object detection method based on a multi-task model, the object detection decoder adopts a Wise-IoU loss function to focus on the prediction regression of ordinary-quality anchor boxes for object detection, and the calculation is as follows:
[0016] L IoU = 1 - IoU
[0017] L WIoUv1 = R WIoU L IoU
[0018]
[0019] Wherein, IoU is used to measure the overlap degree between the predicted box and the ground truth box in the object detection task, and WIoUv1 is a loss function that integrates a dual attention mechanism; r is a non-monotonic focusing coefficient used to adjust the focusing degree of the loss function; R WIoU is a metric for calculating the distance between the predicted box and the ground truth box; W g , H g are the width and height of the minimum bounding box; x and y are the center coordinates of the bounding box; χ gt , y gt are the center points of the ground truth box; β is an outlier parameter for measuring the quality of the anchor box; α and δ are two hyperparameters for adjusting the loss function; is a monotonic focusing coefficient; is a moving average with a momentum of m.
[0020] In the foregoing object detection method based on a multi-task model, the decoupled head includes a 1×1 convolutional layer for dimensionality reduction connected to the object detection decoder, and the output end of the 1×1 convolutional layer is connected to a classification branch and a regression branch; the classification branch is a plurality of sequentially connected 3×3 convolutional layers, and a 1×1 convolutional layer is connected after the 3×3 convolutional layer to complete the classification task; the regression branch includes a localization sub-branch and a confidence sub-branch, and both the localization sub-branch and the confidence sub-branch are single 1×1 convolutional layers, which are respectively used for localization calculation and confidence calculation in the regression task.
[0021] In the aforementioned object detection method based on a multi-task model, a segmentation head branch connected to the semantic segmentation decoder is provided on the encoder.
[0022] In the aforementioned object detection method based on a multi-task model, the semantic segmentation decoder includes a plurality of parallel 1×1 convolutional layers and a segmentation module. The 1×1 convolutional layers perform channel adjustment and sampling on the fused features of the encoder through the segmentation head branch, and then splice the sampled features of each 1×1 convolutional layer through channel splicing and output them to the segmentation module.
[0023] In the aforementioned object detection method based on a multi-task model, the segmentation module includes a semantic reconstruction module, a pyramid pooling module, and a semantic feature fusion module connected in sequence; the semantic reconstruction module is used to perform enlarged receptive field and multi-scale fusion processing on the features; the pyramid pooling module is used to increase the receptive field and perform classification; the semantic feature fusion module is used to calculate the input feature attention weights and obtain weighted features, and then add the weighted features to the input features to obtain output features.
[0024] In the aforementioned object detection method based on a multi-task model, the data used for training the multi-task model is the Cityscapes dataset. After the Cityscapes dataset is proportionally divided into a training set, a validation set, and a test set, the training set and the validation set are converted into the txt format.
[0025] Compared with the prior art, the present invention has the following beneficial effects:
[0026] 1. Multi-task parallelism and resource sharing: The multi-task model of the present invention integrates an object detection decoder and a semantic segmentation decoder. The two can simultaneously obtain the fused features processed by the convolutional network and the feature fusion module, realizing multi-task synchronous processing. This architecture allows the two tasks to share model resources, avoiding repeated calculations, greatly reducing the computational pressure on the monitoring computing platform, and improving the system operation efficiency. When processing complex monitoring images, compared with the traditional independent task processing method, this model can complete object detection and semantic recognition in a shorter time, providing more timely and accurate environmental information for the monitoring and alarm system.
[0027] 2. Fast convergence and efficient detection: During the object detection process, the decoupled head of the multi-task model generates a classification branch and a regression branch, which are respectively responsible for the classification and regression tasks of object detection. This division of labor and cooperation enables the model to quickly converge after generating the object detection box, effectively improving the efficiency and accuracy of object detection, and can more accurately identify and locate monitoring targets, reducing missed detections and false detections.
[0028] 3. Enhance the ability to detect small targets: The model replaces the original standard convolutional network with a deformable convolutional network. When the deformable convolutional network obtains data features, it uses the convolutional layer difference algorithm to obtain the offset, and adjusts the parameters according to the offset through the backpropagation algorithm. This enables it to adaptively sample, expand the receptive field of small targets, reduce the loss of effective feature information, better focus on key features, and significantly improve the ability to detect small targets. In practical applications, for small targets such as people and animals in the distance, the detection effect is significantly better than traditional methods, effectively enhancing the perception ability of the monitoring and alarm system for complex scenarios.
[0029] 4. Improve the accuracy of target detection: The SimAM attention mechanism added before the feature fusion module evaluates the importance of features based on the energy function theory of neuroscience, which can strengthen the feature representation of small targets and reduce the impact of background noise on target detection. Moreover, this mechanism does not require additional learnable parameters. Without increasing the model complexity, it effectively improves the accuracy of target detection and further enhances the reliability of monitoring.
[0030] 5. Optimize the loss function to improve detection accuracy: The target detection decoder uses the Wise-IoU loss function. By introducing various parameters and calculation methods, such as fusing the dual attention mechanism and setting non-monotonic focusing coefficients, it can focus on the prediction regression of ordinary-quality anchor boxes for target detection and more accurately measure the difference between the predicted box and the ground truth box. Compared with the traditional IoU loss function, the Wise-IoU loss function significantly improves the accuracy of target detection and provides more accurate target position information for the monitoring and alarm system. Brief Description of the Drawings
[0031] Figure 1 is the flowchart of the present invention;
[0032] Figure 2 is the overall structure diagram of the multi-task model
[0033] Figure 3 is the structure diagram of the deformable convolution of the present invention;
[0034] Figure 4 is the structure diagram of the decoupled head of the present invention;
[0035] Figure 5 is the structure diagram of the segmentation head branch of the present invention;
[0036] Figure 6 is the structure diagram of the semantic reconstruction module of the present invention;
[0037] Figure 7 is the structure diagram of the semantic feature fusion module of the present invention. Detailed Embodiments
[0038] The present invention will be further described below in conjunction with the accompanying drawings and embodiments, but it shall not be used as a basis for limiting the present invention.
[0039] Embodiment: A target detection method based on a multi-task model, as shown in the accompanying Figure 1 drawings, is carried out according to the following steps:
[0040] Step S1: Construct a multi-task model based on YOLOv5. The structure of the multi-task model is as shown in the accompanying Figure 2 drawings, including an encoder, which is connected to a target detection decoder and a semantic segmentation decoder. The target detection decoder has a decoupled head; the encoder includes a convolutional network and a feature fusion module. The convolutional network is used to extract features from the input data, and the feature fusion module is used to fuse the features extracted by the convolutional network; after obtaining the fused features of the encoder, the target detection decoder performs target detection, and the decoupled head is used to generate a classification branch and a regression branch for the classification task and regression task of target detection, and converges after generating the target detection box; the semantic segmentation decoder performs semantic segmentation after obtaining the fused features of the encoder.
[0041] Extract features from the data through the convolutional network of the multi-task model encoder; the convolutional network is a deformable convolutional network, which replaces the original standard convolutional network. The deformable convolutional network realizes adaptive sampling, expands the receptive field of small targets, and reduces the loss of effective feature information; the difference between the deformable convolution and the ordinary convolution is that an offset is added during the calculation process to adapt to the image features, as shown in the accompanying Figure 3 drawings. The offset is obtained through a convolutional layer conv. The input feature map is input, and the deviation is output. The generated channel dimension is 2N, where the 2 respectively correspond to the 2D offsets of X and Y, and N is calculated from the convolutional kernel size. For example, for a common 3×3 convolution with 9 parameters, then N = 9; during the training process, the convolutional kernels for generating the output features and the offsets are optimized simultaneously. The offset is obtained through interpolation technology, and the parameters of the deformable convolutional network are adjusted using the backpropagation algorithm; the calculation of the output pixel value is shown in the following formula:
[0042]
[0043] In the formula, y is the new pixel coordinate, R = {(-1, -1), (-1, 0), …, (0, 0), …, (1, 0), (1, 1)}, P0 is any point on the input feature map, P n is each point in the convolutional kernel relative to the central point, and ΔP n is the offset, P0 + P n + ΔP n is the new coordinate, X(P0 + P n + ΔP n) is to extract the pixel value at this position, W(P n ) is the weight value of the pixel at the relative position of the convolution kernel; iteratively add the values at all positions W(P n )*X(P0+P n +ΔP n ) to output the result pixel value.
[0044] The feature fusion module of the multi-task model encoder performs feature fusion on the features extracted by the deformable convolutional network; a SimAM attention mechanism is added before the feature fusion module. By strengthening the feature representation of small targets and reducing the influence of background noise on object detection, the accuracy of object detection can be effectively improved without increasing the model complexity. SimAM evaluates the importance of features based on the energy function theory of neuroscience and does not require additional learnable parameters. This mechanism enhances the interpretability of the model by intuitively highlighting key features, and the calculation is as shown in the following equations (2)-(5):
[0045]
[0046] In the formula, t is the target neuron, x represents the neighboring neurons, λ is a hyperparameter, M = H×W, H and W represent the height and width of the energy function on each channel, M represents the number of energy functions on each channel, represents the mean value on each channel in x, X is the input feature map, is the output feature map, represents the variance on each channel in x, represents the distinguishability between the target neuron and the neighboring neurons.
[0047] The object detection decoder adopts the Wise-IoU loss function to focus on the prediction regression of the ordinary quality anchor boxes for object detection, and the calculation formula is as follows:
[0048] L IoU = 1 - IoU
[0049] L WIoUv1 = R WIoU L IoU #(6)
[0050]
[0051] In the formula, IoU is used to measure the overlap degree between the predicted box and the ground truth box in the object detection task, and WIoUv1 is the loss function that integrates the dual attention mechanism; r is the non-monotonic focusing coefficient, which is used to adjust the focusing degree of the loss function; R WIoU is the metric used to calculate the distance between the predicted box and the ground truth box; W g 、H gare the width and height of the minimum bounding box; x and y are the center coordinates of the bounding box; χ gt , y gt are the center points of the ground truth boxes; β is an outlier parameter for measuring the quality of the anchor boxes; α and δ are two hyperparameters for adjusting the loss function; is the monotonic focusing coefficient; is the moving average with momentum m.
[0052] In this embodiment, the original detection head of YOLOv5 is replaced with a decoupled head, and the specific structure is as shown in the appendix Figure 4 . The decoupled head first realizes the dimensionality reduction of features through a 1×1 convolutional layer. Then, the features are fed into two parallel branches: the classification branch and the regression branch. In the classification branch, after continuously applying two 3×3 convolutional layers, a 1×1 convolutional layer is used to complete the classification task. In the regression branch, it is further divided into two sub-branches for localization and confidence. Each sub-branch uses a 1×1 convolutional layer to separately process the calculations of localization and confidence. Finally, the multi-task model uses an Anchor-based method to predict the target boxes, and these prediction results will be compared with the ground truth annotations of the dataset to evaluate the differences between them. In this way, the model can generate multiple detection results, thereby improving the accuracy and efficiency of object detection.
[0053] A segmentation head branch connected to the semantic segmentation decoder is added to the encoder to achieve semantic segmentation, as shown in the appendix Figure 2 . After the picture is subjected to feature extraction and fusion by the encoder, a 1×1 convolution is performed after each of the three segmentation head branches to adjust the number of channels to 128, and different multiples of upsampling are performed according to the size of the input feature map, and the output is 128×80×80. Finally, the outputs of the three branches are concatenated in channels (Concat) and input into the segmentation module (Seg); the segmentation module is as shown in the appendix Figure 5 . After the three channels are concatenated, they are first input into the semantic reconstruction (RM) module. The semantic reconstruction module has the functions of good nonlinearity, expanding the receptive field, and multi-scale fusion. Then, it is input into the pyramid pooling module (PPM) module, and then enters the semantic feature fusion (FFM) module. The FFM mainly plays the role of feature fusion. Finally, the number of channels is adjusted through a 1×1 convolution and the resolution is restored by 8 times of upsampling. The semantic reconstruction module is as shown in the appendix Figure 6 . Through four branches including convolutional layers and BNSILU, the output feature maps are spliced together and then the number of channels is adjusted through a 1×1 convolution. The semantic feature fusion module is as shown in the appendix Figure 7As shown, the input feature map first passes through a 3×3 convolution block. Then, the feature map passes through a channel attention module, which uses global average pooling and two 1×1 convolutional layers to calculate the channel attention weights. Finally, the original feature map is added to the weighted feature map to obtain the final output feature map.
[0054] Step S2: Train the constructed multi-task model. During the training of object detection, classification and regression branches are generated through the decoupled head of the multi-task model for the classification and regression tasks of object detection. After generating the object detection boxes, they are compared with the validation set to achieve fast convergence. After the semantic segmentation training process, the output feature map is compared with the validation set to achieve convergence. In this step, the dataset used for training is converted into the txt format adopted by YOLOv5 by programming the downloaded Cityscapes dataset. The Cityscapes dataset is a pixel-level dataset that contains the semantic information of each pixel. Since the object detection model usually requires annotation information at the rectangular box level to learn how to detect and locate objects, the Cityscapes data format cannot be directly used for the training of the multi-task model. Therefore, the Cityscapes dataset is processed to generate a rectangular box-level dataset with the annotation format for object detection. The downloaded Cityscapes dataset includes 10 object detection dataset categories (traffic light, traffic sign, person, rider, car, truck, bus, train, motorcycle, and bicycle) and 19 semantic segmentation dataset categories (road, sidewalk, building, wall, fence, pole, traffic light, traffic sign, vegetation, terrain, sky, person, rider, car, truck, bus, train, motorcycle, and bicycle), which are divided into a training set, a validation set, and a test set. Among them, the training set (train) has 2975 images, the validation set (val) has 500 images, and the test set (test) has 1525 images. The training set and the validation set are converted into the txt format adopted by YOLOv5. Use the trained multi-task model to perform object detection and semantic recognition in the monitoring scenario.
[0055] In summary, the encoder of the multi-task model of the present invention extracts and fuses features from the data, and then simultaneously obtains the fused features through the object detection decoder and the semantic segmentation decoder of the multi-task model to perform object detection and semantic segmentation respectively, sharing resources while realizing multi-task processing to reduce the computational pressure. Moreover, the deformable convolutional network is sampled, which can better focus on key features and improve the ability of the convolution to globally understand the scene. During the object detection process, the decoupled head of the multi-task model generates a classification branch and a regression branch for the classification task and the regression task of object detection. After generating the object detection box, it is compared with the validation set to achieve convergence. After the semantic segmentation process, the feature map is output, and the feature map is compared with the validation set to achieve convergence. The multi-task parallel processing realizes efficient convergence, with high efficiency. Adding the SimAM attention mechanism improves the accuracy of object detection, and adopting the optimized Wise-IoU loss function further enhances the accuracy of object detection.
Claims
1. A target detection method based on a multi-task model, characterized in that: The steps are as follows: Step S1: Construct a multi-task model. The multi-task model includes an encoder, which is connected to an object detection decoder and a semantic segmentation decoder. The object detection decoder has a decoupled head; the encoder includes a convolutional network and a feature fusion module. The convolutional network is used to extract features from the input data, and the feature fusion module is used to fuse the features extracted by the convolutional network; after obtaining the fused features of the encoder, the object detection decoder performs object detection. The decoupled head is used to generate a classification branch and a regression branch for the classification task and regression task of object detection, and converges after generating the object detection box; after obtaining the fused features of the encoder, the semantic segmentation decoder performs semantic segmentation; Step S2: Train the constructed multi-task model, and perform object detection and semantic recognition with the trained multi-task model.
2. The object detection method based on a multi-task model according to claim 1, wherein: The convolutional network is a deformable convolutional network. While obtaining the data features, the deformable convolutional layer network obtains the offsets in the data through the difference algorithm of the convolutional layer, and adjusts its parameters according to the offsets through the backpropagation algorithm; the output of the deformable convolutional network is calculated as shown in the following formula: Where y is the new pixel coordinate, R = {(-1, -1), (-1, 0), …, (0, 0), …, (1, 0), (1, 1)}, P0 is an arbitrary point on the input feature map, and P n is the offset of each point in the convolution kernel relative to the center point, and ΔP n is the offset value offset, P0 + P n + ΔP n is the new coordinate, X(P0 + P n + ΔP n ) is to extract the pixel value at this position, and W(P n ) is the weight value of the pixel at the relative position of the convolution kernel; iteratively add the values W(P n ) * X(P0 + P n + ΔP n ) to output the result pixel value.
3. The object detection method based on a multi-task model according to claim 1, wherein: The feature fusion module is added with a SimAM attention mechanism, which is used to enhance the feature representation of small targets, and the calculation is as shown in the following formula: Wherein, t is the target neuron, x represents the neighboring neurons, λ is a hyperparameter, M = H × W, H and W represent the height and width of the energy function on each channel, and M represents the number of energy functions on each channel. represents the mean value on each channel in x, and X is the input feature map. is the output feature map. represents the variance on each channel in x. represents the distinguishability between the target neuron and the neighboring neurons.
4. The object detection method based on a multi-task model according to claim 1, characterized in that: The object detection decoder adopts a Wise-IoU loss function to focus on the prediction regression of ordinary-quality anchor boxes for object detection, and the calculation is as shown in the following formula: L IoU = 1 - IoU L WIoUv1 = R WIoU L IoU L WIoUv3 = rL WIoUv1 , In the formula, IoU is used to measure the overlap degree between the predicted box and the ground truth box in the object detection task, and WIoUv1 is a loss function that integrates a dual attention mechanism; r is a non-monotonic focusing coefficient, which is used to adjust the focusing degree of the loss function; R WIoU is a metric for calculating the distance between the predicted bounding box and the ground truth bounding box; W g and H g are the width and height of the minimum bounding box; x and y are the center coordinates of the bounding box; χ gt and y gt are the center points of the ground truth bounding box; β is an outlier parameter for measuring the quality of the anchor box; α and δ are two hyperparameters for adjusting the loss function; is the monotonic focusing coefficient; is the moving average with momentum m.
5. The object detection method based on a multi-task model according to claim 1, characterized in that: The decoupled head includes a 1×1 convolutional layer for dimensionality reduction connected to the object detection decoder. The output end of the 1×1 convolutional layer is connected to a classification branch and a regression branch; the classification branch is multiple 3×3 convolutional layers connected in sequence, and a 1×1 convolutional layer is connected after the 3×3 convolutional layer to complete the classification task; the regression branch includes a localization sub-branch and a confidence sub-branch. Both the localization sub-branch and the confidence sub-branch are single 1×1 convolutional layers, which are respectively used for the localization calculation and confidence calculation of the regression task.
6. The object detection method based on a multi-task model according to claim 1, characterized in that: The encoder has a segmentation head branch connected to the semantic segmentation decoder.
7. The object detection method based on a multi-task model according to claim 6, wherein: The semantic segmentation decoder includes multiple parallel 1×1 convolutional layers and a segmentation module. The 1×1 convolutional layer adjusts the channels of the fused features of the encoder through the segmentation head branch and samples them, and then splices the sampled features of each 1×1 convolutional layer through channel splicing and outputs them to the segmentation module.
8. The object detection method based on a multi-task model according to claim 7, wherein: The segmentation module includes a semantic reconstruction module, a pyramid pooling module, and a semantic feature fusion module connected in sequence; the semantic reconstruction module is used to perform enlarged receptive field and multi-scale fusion processing on the features; the pyramid pooling module is used to increase the receptive field and classify; the semantic feature fusion module is used to calculate the input feature attention weights and obtain weighted features, and then add the weighted features to the input features to obtain the output features.
9. The object detection method based on a multi-task model according to claim 1, wherein: The data used for training the multi-task model is the Cityscapes dataset. After dividing the Cityscapes dataset into a training set, a validation set, and a test set according to a certain ratio, the training set and the validation set are converted into the txt format.