A data enhancement method for sparse turnout images and its parameter optimization method
By performing data augmentation processing on sparse switch images, including splicing transformation, affine transformation and hyperparameter optimization, the problem that data augmentation method in the prior art is not suitable for sparse switch images, and the detection performance of the network is improved.
Patent Information
- Application Number
- CN202211174300.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-26
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-09-26
AI Technical Summary
The existing data enhancement methods are not suitable for sparse switch images, which makes it difficult for deep learning networks to learn effective switch features, learn too much noise features, and network performance cannot be improved.
By performing switch image stitching transformation, affine transformation and specified size interception on the switch image data set with labels, switch image and label information are formed after data augmentation, and a hyperparameter optimization method is designed to optimize hyperparameters using the variation and cross-section ideas of genetic algorithms.
The positive sample ratio of the switch image is improved, the background information of the switch image is enhanced, the overfitting problem is avoided, the rapid convergence of network training is promoted, and better detection performance is obtained.
Smart Images

Figure CN115601587B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of turnout image data processing, and relates to a data enhancement method for sparse turnout images and a parameter optimization method therefor. Background Art
[0002] With the gradual transformation of the turnout image processing technology from the traditional manually designed feature method to the data-driven deep learning method, the performance of the object detection algorithm based on the turnout image has been significantly improved, and a new technical means has been provided for the turnout state detection problem in the track scene. The data-driven turnout state detection method requires a large number of turnout targets to learn the relevant features of different turnout states. However, since the turnout opening appears with a small probability in the railway scene and the target density is very low, in the turnout images collected from the driver's perspective in reality, each frame of turnout image often contains only one turnout target, and the pixel proportion of the turnout target area in the entire turnout image is also relatively low, resulting in the proportion of negative samples in the training samples far exceeding that of positive samples. The above characteristics make it difficult to fully learn effective turnout features during the training process of the deep learning network, but instead learn too many noise features, resulting in the inability to improve the network performance. How to design a data enhancement method suitable for sparse turnout images has become an urgent problem to be solved in the current turnout state detection task.
[0003] The invention patent with the publication number CN202111005581.4 discloses a turnout image automatic data enhancement method, system, medium and terminal, and proposes a parameter optimization and model selection method for the data enhancement method. It obtains the final data enhancement scheme by fixing the existing data enhancement method and parameter range in a set and then testing and screening the data enhancement instances in the set, but does not design a data enhancement method suitable for sparse target turnout images, and only improves the selection strategy of the existing data enhancement method.
[0004] The invention patent with the publication number CN202111103809.3 discloses a data enhancement method, system, device and computer-readable storage medium, and proposes a data enhancement method that comprehensively performs multiple operations such as picture scale scaling, horizontal flipping, chromaticity transformation of the turnout image, and saturation adjustment. This method encodes the above multiple data enhancement operations, then tests them in a self-built RNN controller, calculates the reward value, and then obtains the data enhancement method with the largest reward value as the optimal data enhancement method by continuously optimizing the adjustable parameters. This data enhancement method combines common data enhancement operations and obtains the optimal parameters through repeated tests to improve the training effect of the deep neural network. However, this method still takes the entire picture as the most basic unit of data enhancement, and the data enhancement effect is relatively single and cannot provide rich turnout image features.
[0005] The invention patent with the publication number CN202210158191.9 discloses a data augmentation processing method, device, computer equipment and storage medium, and proposes a selection strategy for data augmentation methods. First, a large dataset is divided into multiple sub-datasets, and then different data augmentation methods are used to train each sub-dataset. The data augmentation method corresponding to the optimal training result is selected, and then the total dataset is tested to finally determine the optimal data augmentation method. This method can reduce the training time and the computational complexity in the process of selecting data augmentation methods by applying different data augmentation methods to each sub-dataset. However, this method does not solve the problem of imbalance between positive and negative samples in sparse turnout images. Summary of the Invention
[0006] The purpose of the embodiments of the present invention is to provide a data augmentation method for sparse turnout images to solve the problem that the existing data augmentation methods are not applicable to the data augmentation of sparse turnout images.
[0007] Another purpose of the embodiments of the present invention is to provide a parameter optimization method for the data augmentation method of sparse turnout images.
[0008] The technical solution adopted in the embodiments of the present invention is as follows: A data augmentation method for sparse turnout images includes the following steps:
[0009] Step 1: Label the turnout categories of each turnout image in the turnout image dataset to form a turnout image dataset with labels. The turnout categories are divided into left-turn turnouts and right-turn turnouts.
[0010] Step 2: Perform turnout image splicing transformation on the turnout images and corresponding labels in the turnout image dataset with labels to form spliced turnout images and corresponding labels.
[0011] Step 3: Perform affine transformation on the spliced turnout images and corresponding labels to obtain affine-transformed turnout images and corresponding labels.
[0012] Step 4: Intercept turnout images of a specified size from the affine-transformed turnout images and corresponding labels, and clear smaller fragmented target labels to obtain the finally data-augmented turnout images and label information.
[0013] Another technical solution adopted in the embodiments of the present invention is as follows: A parameter optimization method for the data augmentation method of the sparse turnout images is carried out according to the following steps:
[0014] Step S1: Determine the parameters to be optimized and the value ranges of the parameters: (1) The left and right boundary expansion coefficients T l 、T r; (2) Affine transformation matrix hyperparameters center1, center2, revo1, revo2, trans1, trans2; (3) Label retention minimum area ratio A min ; (4) The probability value of whether to perform data augmentation, denoted as Prob;
[0015] Step S2: Initialize the optimizable parameters, then multiply the initialized optimizable parameters by 10^2, convert the hyperparameters to integers, and represent them in binary;
[0016] Step S3: Denote the number of parameter optimizations as n. When n = 1, use the selected hyperparameters to perform data augmentation on the turnout image dataset with labels, combine with the deep learning network model for training and testing. Convert the hyperparameters encoded in binary in Step S2 to decimal and input them into the deep learning network model for training and testing. The test results are selected as accuracy P, recall R, and mean average precision mAP, and record the hyperparameters used in each training and testing, store them in the variable passhyp. When n is not equal to 1, first determine whether the selected hyperparameters are in the passhyp variable. If they are in passhyp, discard the hyperparameters and directly execute Step S6. Otherwise, use the selected hyperparameters to perform data augmentation on the dataset and combine with the deep learning network model for training and testing;
[0017] Step S4: According to the test metrics of the deep learning network model, design the objective function J. In order to comprehensively consider the detection ability of the deep learning network model for turnout categories, set the weight coefficients to balance the contribution of each metric to the objective function. The objective function is expressed as J = 0.2 * P + 0.4 * R + 0.4 * mAP;
[0018] Step S5: When n ≤ 10, record the number of parameter optimizations, the objective function result J, and the hyperparameters of this calculation in the Best_result variable in the form of a dictionary. The storage form is as follows: {circle:times,hyp:value1,J:value2}, where times represents which parameter optimization this calculation process belongs to, value1 represents the hyperparameters used in this calculation, saved in decimal form, value2 represents the objective function value corresponding to this calculation, also saved in decimal. The Best_result variable is set with a maximum storage space of 10 dictionary data. When n > 10, compare the objective function value J of this calculation with the value2 of each item in the Best_result variable. When the J value is greater than the value2 of a certain item, replace the information of that item in Best_result with the result of this calculation;
[0019] Step S6: In Best_result, select 5 groups of hyperparameters as the new basic hyperparameters with the J value as the selection probability;
[0020] Step S7: Introduce the mutation and crossover ideas of the genetic algorithm to update the new basic hyperparameters;
[0021] Step S8: Repeat Steps S3 - S7. When the number of iterations reaches the preset value or the objective function reaches the expected goal, extract the optimal result from Best_result as the final hyperparameter value to complete the hyperparameter optimization process.
[0022] The beneficial effects of the embodiments of the present invention are:
[0023] 1. A method for extracting the turnout area of a sparse turnout image is proposed. By determining the area extraction range through the turnout image label information, it can ensure that each extracted turnout area block contains the turnout target;
[0024] 2. A method for splicing area blocks of a sparse turnout image is proposed. By randomly splicing the turnout area blocks of different turnout images, it can enrich the background information of the turnout image, and at the same time achieve the regularization effect of the turnout image, avoiding the overfitting problem in the training process;
[0025] 3. A data augmentation method for sparse turnout images is proposed. By combining and splicing the extracted turnout area blocks, the turnout target density of a single turnout image is increased, the positive sample ratio of the turnout image is improved, which is beneficial to the rapid convergence of the network training process and obtains better detection performance, and solves the problem that the existing data augmentation methods are not applicable to the data augmentation of sparse turnout images;
[0026] 4. A hyperparameter optimization method for the data augmentation method of sparse turnout images is proposed. The encoding representation, mutation, and crossover methods of hyperparameters are designed, and the mutation direction of hyperparameters is guided by the generalized gradient, which can achieve a faster global optimal parameter search. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0028] Figure 1 It is a flowchart of the data augmentation method provided by the embodiments of the present invention.
[0029] Figure 2(a) is a sample of the left - hand turnout annotation example provided by the embodiments of the present invention.
[0030] Figure 2(b) is a sample of the right-turn turnout annotation instance provided by the embodiment of the present invention.
[0031] Figure 3 It is a schematic diagram of the turnout image stitching method provided by the embodiment of the present invention.
[0032] Figure 4 It is a comparison chart of the effects of the data augmentation method provided by the embodiment of the present invention.
[0033] Figure 5 It is a schematic diagram of the hyperparameter coding method provided by the embodiment of the present invention.
[0034] Figure 6 It is a flowchart of the parameter optimization of the data augmentation method provided by the embodiment of the present invention.
[0035] Figure 7 It is a schematic diagram of the mutation operation of hyperparameter coding provided by the embodiment of the present invention.
[0036] Figure 8 It is a schematic diagram of the crossover operation of hyperparameter coding provided by the embodiment of the present invention. Detailed implementation manners
[0037] Next, in combination with the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0038] Embodiment 1
[0039] This embodiment provides a data augmentation method for sparse turnout images. As Figure 1 shown, first, a dataset is collected to obtain the original turnout images, and the original turnout images are annotated to obtain the label files corresponding to the turnout images. Then, the turnout region blocks are extracted using the label information and the turnout region blocks are stitched to obtain the stitched turnout images. Further, an affine transformation is performed on the stitched turnout images, and turnout image blocks of a specified size are randomly intercepted within the region of the affine-transformed turnout images as the data-augmented turnout images. Finally, the labels of the original turnout images are also subjected to corresponding coordinate transformations to obtain the labels of the data-augmented turnout images, completing the data augmentation process.
[0040] The specific process of the data augmentation method for sparse turnout images is as follows:
[0041] Step 1: Obtain a turnout image dataset with labels:
[0042] Step 11: Collect switch images in the train operation scenario. During the collection process, the camera should be located at the center position of the cab, and the optical axis of the camera should be consistent with the train operation direction, so that the track area is imaged in the central area of the switch image as much as possible, avoiding adverse effects on the track area imaging caused by edge distortion of the switch image. Finally, manually screen the collected switch images to select the switch images that can identify the switch category, so as to eliminate the switch images with blurred pixels that cannot be identified due to reasons such as camera jitter and far switch targets. The switch categories are divided into left-turn switches and right-turn switches;
[0043] Step 12: Use the VIA (VGG Image Annotator) annotation tool to annotate the switch categories of the switch images collected in Step 11 to obtain a switch image dataset with labels; the basis for judging the switch category is the geometric relationship between the switch rail and the stock rail. When the gap between the switch rail and the stock rail is on the right side, the switch is considered a right-turn switch. When the gap between the switch rail and the stock rail is on the left side, the switch is considered a left-turn switch. The annotated switch images are shown in Figures 2(a) and 2(b); since the switch targets are relatively sparse, the directly annotated switch images are not suitable for direct input into the deep learning network model for calculation, and further data augmentation is required.
[0044] In some embodiments, the resolution of the switch images is uniformly 1080×1920, and a buffer stabilizing device should be installed between the switch image collection camera and the train cab to minimize the impact of train vibration on camera imaging during train operation.
[0045] In some embodiments, the switch image collection process should be carried out under different railway lines and different weather conditions. The final collected switch images require at least 8000 frames of switch images and at least 10000 instances of each type of switch.
[0046] Step 2: Perform switch image stitching transformation on the switch images and corresponding labels in the switch image dataset with labels, which can reduce the target search space and enrich the data representativeness. Refer to Figure 1 , the specific steps for performing switch image stitching transformation on the switch images and corresponding labels in the switch image dataset with labels are as follows:
[0047] Step 21: Number the switch images and corresponding labels in the switch image dataset with labels in sequence to obtain the numbering information of the switch images to be stitched;
[0048] Step 211: The turnout images and corresponding labels of the labeled turnout image data set are numbered in sequence, and the numbering is performed in units of data pairs. A turnout image and a corresponding label constitute a data pair, that is, the turnout image and label of a data pair use the same number that uniquely identifies the data pair, and the numbering increases in sequence from 00000, and the numbering information is stored in the indexes list.
[0049] Step 212: Randomly divide the batch blocks in indexes. Specifically, the labeled turnout image dataset is divided in indexes in units of batch_size. batch_size represents the batch size, which refers to the number of turnout images and their corresponding labels used in each calculation process. The number of batch blocks obtained by training the labeled turnout image dataset once is recorded as i. Finally, the data that is less than one batch block is also regarded as a batch block. Each completion of the calculation process represents the completion of the calculation of one batch block. i=ceil(len(indexes) / batch_size), wherein len() represents the size of the indexes list, and ceil() is the encapsulation function of the math module in Python, which represents rounding up the calculation result.
[0050] Step 22: Perform a splicing transformation on the turnout images and corresponding labels in the labeled turnout image dataset:
[0051] The calculation of the mth batch block is called the mth calculation process, and the indexth turnout image and the corresponding label in the mth calculation process are spliced and transformed:
[0052] Step 221: sequentially obtain the number information of the turnout images and corresponding labels in the mth calculation process, and record the number of the turnout images and corresponding labels obtained each time as index. Randomly select 7 turnout images and corresponding label numbers from the numbers other than index in indexs, and store them together with index in P, where P is the number list of turnout images to be spliced and their labels;
[0053] Step 222: extract the corresponding turnout images and label numbers from P in sequence, the turnout image number extracted each time is recorded as j, and the value range of j is 0-7. The turnout image and label corresponding to each number are resized, and the height of the resized turnout image is h1=r*h, and the width of the resized turnout image is w1=r*w, where h is the height of the original turnout image before rescaling, w is the width of the original turnout image before rescaling, r is the scale ratio of the turnout image, and r=s / h is taken, s is 1 / 2 of the side length of the resized turnout image canvas; extract the label information [num, x] of the jth turnout image 1 min ,y 1min ,x 1 max ,y 1 max ], where num represents the label category; num has two values: 0 and 1, 0 represents a left turn and 1 represents a right turn; x 1 min Indicates the horizontal coordinate of the upper left corner of the label box, y 1 min Indicates the upper left corner ordinate of the label box, x 1 max Indicates the horizontal coordinate of the lower right corner of the label box, y 1 max Indicates the vertical coordinate of the lower right corner of the label box; the label information of the turnout image is scaled in proportion to the turnout image, that is, the label information becomes [num, x min ,y min ,x max ,y max ], where [x min ,y min ,x max ,y max ]=[x 1 min ,y 1 min ,x 1 max ,y 1 max ]*r,x min Indicates the horizontal coordinate of the upper left corner of the label box after scaling, y min Indicates the upper left corner ordinate of the label box after scaling, x max Indicates the horizontal coordinate of the lower right corner of the label box after scaling, y max Indicates the vertical coordinate of the lower right corner of the label box after scaling;
[0054] Step 223: extract the turnout area of the j-th turnout image in sequence, and the turnout area is represented by [x1, y1, x2, y2], where x1 is the upper left corner abscissa of the turnout area, y1 is the upper left corner ordinate of the turnout area, x2 is the lower right corner abscissa of the turnout area, and y2 is the lower right corner ordinate of the turnout area; in order to ensure the richness of the turnout area background, the turnout area is determined according to the label coordinate information, and the calculation formula is as follows: y1=0, y2=h1, x1=max(x min *(1-T l ),0),x2=min(x max *(1+T r ),w1), where T l represents the expansion coefficient of the left boundary of the turnout area, T r represents the expansion coefficient of the right boundary of the turnout area, Tl and T r The default value of is: generate a random number between 0.3 and 0.8 through the random.uniform() function of Python respectively; max() means taking the maximum value of the two numbers, which is used to constrain the upper left corner point of the turnout area to avoid the extraction area exceeding the left boundary of the turnout image; min() means taking the minimum value of the two numbers, which is used to constrain the lower right corner point of the turnout area to avoid the extraction area exceeding the right boundary of the turnout image; repeat the above process, and extract all turnout areas of all turnout images corresponding to the numbers in P in turn. The extracted turnout area turnout image serial number is represented by [j, k], and [j, k] represents the kth turnout area extracted from the jth turnout image;
[0055] Step 224: Create a turnout image stitching turnout image canvas: Use numpy.zeros() function to generate a square canvas, such as Figure 3 As shown, a (h1*2, h1*2) canvas is generated, and the canvas is divided into four areas, each of which is a square area of h1*h1. From left to right and from top to bottom, they are recorded as area 1, area 2, area 3, and area 4 respectively. The default pixel value of the canvas is set to 0. The square area of h1*h1 is used to pair with the turnout image, which can make the splicing point closer to the center area of the canvas, and it is easier for the turnout image after data enhancement to contain rich turnout image background information;
[0056] Step 225: Using the method of matching the turnout regions of two turnout images with a canvas region, pair the turnout region of j∈[0,7] with regions 1 to 4 of the turnout image canvas: Figure 3 As shown, the turnout area and the canvas area are paired, the turnout area of j=0,1 is paired with canvas area 1; the turnout area of j=2,3 is paired with canvas area 2; the turnout area of j=4,5 is paired with canvas area 3; the turnout area of j=6,7 is paired with canvas area 4, and the method of matching the turnout areas of two turnout images with one canvas area can ensure that the canvas contains as much valid pixel information as possible with less calculation amount; because the height of the turnout area is consistent with the height of the canvas area, the turnout areas of the corresponding turnout images are selected in turn to expand from the center point of the canvas to the boundary. When the turnout area overflows the canvas area, the turnout area in the canvas area is intercepted.
[0057] The turnout area of the turnout image is represented by [x 1j_k ,y 1j_k ,x 2j_k ,y 2j_k ] indicates that x 1j_k Refers to the horizontal coordinate of the upper left corner of the kth turnout area of the jth turnout image in the original turnout image, y 1j_kRefers to the ordinate of the upper left corner of the kth turnout area of the jth turnout image in the original turnout image, x 2j_k Refers to the horizontal coordinate of the lower right corner of the kth turnout area of the jth turnout image in the original turnout image, y 2j_k Refers to the ordinate of the lower right corner of the kth turnout area of the jth turnout image in the original turnout image. The corresponding position in the canvas area is [x 1pj_k ,y 1pj_k ,x 2pj_k ,y 2pj_k ], x 1pj_k Refers to the upper left corner horizontal coordinate of the kth turnout area of the jth turnout image in the corresponding area of the canvas, y 1pj_k Refers to the upper left corner ordinate of the kth turnout area of the jth turnout image in the corresponding area of the canvas, x 2pj_k Refers to the lower right corner horizontal coordinate of the kth turnout area of the jth turnout image in the corresponding area of the canvas, y 2pj_k Refers to the lower right corner ordinate of the kth turnout area of the jth turnout image in the corresponding area of the canvas.
[0058] For region 1:
[0059] Specifically, when j=0, k=0, that is, the position information of the 0th turnout area of the 0th turnout image in the original turnout image is [x 10_0 ,y 10_0 ,x 20_0 ,y 20_0 ], the corresponding position in area 1 of the canvas is [x 1p0_0 ,y 1p0_0 ,x 2p0_0 ,y 2p0_0 ], x 1p0_0 =max(h1-(x 20_0 -x 10_0 ),0),y 1p0_0 =0,x 2p0_0 =h1,y 2p0_0 =h1; and record the leftmost horizontal coordinate of the turnout area on the canvas, denoted as last1, last1=x 1p0_0 ; When j = 0, k = 1, that is, the position information of the first turnout area of the 0th turnout image in the original turnout image is [x 10_1 ,y 10_1 ,x 20_1 ,y 20_1 ], the corresponding position in canvas area 1 is [x 1p0_1 ,y 1p0_1 ,x 2p0_1 ,y 2p0_1 ], x 1p0_1 =max(last1-(x 20_1 -x10_1 ),0),y 1p0_1 =0,x 2p0_1 =max(last1,0),y 2p0_1 =h1; and update the leftmost horizontal coordinate of the turnout area on the canvas, that is, last1=x 1p0_1 ; Similarly, when j∈[0,1], except for the cases of j=0 and k=0, the position information of the kth turnout area of the jth turnout image in the original turnout image is [x 1j_k ,y 1j_k ,x 2j_k ,y 2j_k ], the corresponding position in canvas area 1 is [x 1pj_k ,y 1pj_k ,x 2pj_k ,y 2pj_k ], x 1pj_k =max(last1-(x 2j_k -x 1j_k ),0),y 1pj_k =0,x 2pj_k =max(last1,0),y 2pj_k =h1; and update the leftmost horizontal coordinate of the turnout area on the canvas, that is, last1=x 1pj_k This continues until all the turnout areas of the 0th and 1st turnout images are matched with canvas area 1.
[0060] In particular, when the remaining allocated area of the last canvas is not enough to store the next turnout area, the obtained turnout area needs to be regularized according to the size of the remaining allocated area. j∈[0,1], when the canvas area last1=0 that matches the kth turnout area of the jth turnout image, it means that the last remaining allocated area of canvas area 1 is not enough to store the next turnout area, and the horizontal coordinate of the upper left corner of the position information of the turnout area in the original turnout image is adjusted to x 1j_k =x 2j_k -(x 2pj_k -x 1pj_k ), that is, the position information of the turnout area in the original turnout image is adjusted to [x 2j_k -(x 2pj_k -x 1pj_k ),y 1j_k ,x 2j_k ,y 2j_k ], that is, the turnout area is cropped from the left side so that its size is just the same as the remaining canvas area, and the canvas is filled. The turnout area input later is directly removed and not spliced into the canvas. If the turnout area is insufficient, all the turnout areas are spliced together, and the remaining unspliced areas keep the pixel value of the canvas itself.
[0061] For zone 2:
[0062] Specifically, when j=2 and k=0, that is, the position information of the 0th turnout area of the 2nd turnout image in the original turnout image is [x 12_0 ,y 12_0 ,x 22_0 ,y 22_0 ], the corresponding position in canvas area 2 is [x 1p2_0 ,y 1p2_0 ,x 2p2_0 ,y 2p2_0 ], x 1p2_0 =h1,y 1p2_0 =0,x 2p2_0 =min(h1+(x 22_0 -x 12_0 ),2*h1),y 2p2_0 =h1; and record the rightmost horizontal coordinate of the turnout area on the canvas, denoted as last2, last2=x 2p2_0 ; When j = 2, k = 1, that is, the position information of the first turnout area of the second turnout image in the original turnout image is [x 12_1 ,y 12_1 ,x 22_1 ,y 22_1 ], the corresponding position in canvas area 2 is [x 1p2_1 ,y 1p2_1 ,x 2p2_1 ,y 2p2_1 ], x 1p2_1 =min(last2,2*h1),y 1p2_1 =0,x 2p2_1 =min((last2+(x 22_1 -x 12_1 )),2*h1),y 2p2_0 =h1; and update the rightmost horizontal coordinate of the turnout area on the canvas, last2 = x 2p2_1 Similarly, when j∈[2,3], except for the case of j=2 and k=0, the position information of the kth turnout area of the jth turnout image in the original turnout image is [x 1j_k ,y 1j_k ,x 2j_k ,y 2j_k ], the corresponding position in canvas area 2 is [x 1pj_k ,y 1pj_k ,x 2pj_k ,y 2pj_k ], x 1pj_k =min(last2,2*h1),y 1pj_k =0,x 2pj_k =min((last2+(x 2j_k -x 1j_k)),2*h1),y 2pj_k =h1. And update the rightmost horizontal coordinate of the turnout area on the canvas, last2=x 2pj_k This continues until all the turnout areas of the second and third turnout images are matched with canvas area 2.
[0063] In particular, when the remaining allocated area of the last canvas is not enough to store the next turnout area, the obtained turnout area needs to be regularized according to the size of the remaining allocated area. j∈[2,3], when the kth turnout area of the jth turnout image matches the canvas area, last2=2*h1, which means that the remaining allocated area of the last canvas is not enough to store the next turnout area, and the horizontal coordinate of the lower right corner of the position information of the turnout area in the original turnout image is adjusted to x 2j_k =x 1j_k +(x 2pj_k -x 1pj_k ), that is, the position information of the turnout area in the original turnout image is adjusted to [x 1j_k ,y 1j_k ,x 1j_k +(x 2pj_k -x 1pj_k ),y 2j_k ]; If the turnout area is insufficient, all turnout areas matching area 2 will be spliced together, and the remaining unspliced areas will keep the pixel values of the canvas itself;
[0064] For zone 3:
[0065] Specifically, when j=4 and k=0, that is, the position information of the 0th turnout area of the 4th turnout image in the original turnout image is [x 14_0 ,y 14_0 ,x 24_0 ,y 24_0 ], the corresponding position in canvas area 3 is [x 1p4_0 ,y 1p4_0 ,x 2p4_0 ,y 2p4_0 ], x 1p4_0 =max(h1-(x 24_0 -x 14_0 ),0),y 1p4_0 =h1,x 2p4_0 =h1,y 2p4_0 =2*h1. And record the leftmost horizontal coordinate of the turnout area on the canvas, denoted as last3, last3=x 1p4_0 When j = 4 and k = 1, the position information of the first turnout area of the fourth turnout image in the original turnout image is [x 14_1 ,y 14_1 ,x 24_1 ,y 24_1], the corresponding position in canvas area 3 is [x 1p4_1 ,y 1p4_1 ,x 2p4_1 ,y 2p4_1 ], x 1p4_0 =max(last3-(x 24_1 -x 14_1 ),0),y 1p4_0 =h1,x 2p4_0 =max(last3,0),y 2p4_0 =2*h1. And update the leftmost horizontal coordinate of the turnout area on the canvas, last3=x 1p4_1 Similarly, when j∈[4,5], except for the case of j=4 and k=0, the position information of the kth turnout area of the jth turnout image in the original turnout image is [x 1j_k ,y 1j_k ,x 2j_k ,y 2j_k ], the corresponding position in canvas area 3 is [x 1pj_k ,y 1pj_k ,x 2pj_k ,y 2pj_k ], x 1pj_k =max(last3-(x 2j_k -x 1j_k ),0),y 1pj_k =h1,x 2pj_k =max(last3,0),y 2pj_k =2*h1. And update the leftmost horizontal coordinate of the turnout area on the canvas, last3=x 1pj_k This continues until all the turnout areas of the 4th and 5th turnout images are matched with canvas area 3.
[0066] In particular, when the remaining allocated area of the last canvas is not enough to store the next turnout area, the obtained turnout area needs to be regularized according to the size of the remaining allocated area. j∈[4,5], when the kth turnout area of the jth turnout image matches the canvas area, last3=0 means that the remaining allocated area of the last canvas is not enough to store the next turnout area, and the horizontal coordinate of the upper left corner of the position information of the turnout area in the original turnout image is adjusted to x 1j_k =x 2j_k -(x 2pj_k -x 1pj_k ), that is, the position information of the turnout area in the original turnout image is adjusted to [x 2j_k -(x 2pj_k -x 1pj_k ),y 1j_k ,x 2j_k ,y 2j_k ].
[0067] For zone 4:
[0068] Specifically, when j=6 and k=0, that is, the position information of the 0th turnout area of the 6th turnout image in the original turnout image is [x 16_0 ,y 16_0 ,x 26_0 ,y 26_0 ], the corresponding position in canvas area 4 is [x 1p6_0 ,y 1p6_0 ,x 2p6_0 ,y 2p6_0 ], x 1p6_0 =h1,y 1p6_0 =h1,x 2p6_0 =min(h1+(x 26_0 -x 16_0 ),2*h1),y 2p6_0 =2*h1. And record the rightmost horizontal coordinate of the turnout area on the canvas, denoted as last4, last4=x 2p6_0 When j = 6 and k = 1, the position information of the first turnout area of the sixth turnout image in the original turnout image is [x 16_1 ,y 16_1 ,x 26_1 ,y 26_1 ], the corresponding position in canvas area 4 is [x 1p6_1 ,y 1p6_1 ,x 2p6_1 ,y 2p6_1 ], x 1p6_1 =min(last4,2*h1),y 1p6_1 =h1,x 2p6_1 =min((last4+(x 26_0 -x 16_0 )),2*h1),y 2p6_1 =2*h1. And update the rightmost horizontal coordinate of the turnout area on the canvas, last4=x 2p6_1 Similarly, when j∈[6,7], except for the case of j=6 and k=0, the position information of the kth turnout area of the jth turnout image in the original turnout image is [x 1j_k ,y 1j_k ,x 2j_k ,y 2j_k ], the corresponding position in canvas area 4 is [x 1pj_k ,y 1pj_k ,x 2pj_k ,y 2pj_k ], x 1p6_1 =min(last4,2*h1),y 1pj_k =h1,x 2pj_k=min((last4+(x 2j_k -x 1j_k )),2*h1),y 2pj_k =2*h1. And update the rightmost horizontal coordinate of the turnout area on the canvas, last4=x 2pj_k , until the matching of all turnout areas of the 6th and 7th turnout images and canvas area 3 is completed.
[0069] In particular, when the remaining allocated area of the last canvas is not enough to store the next turnout area, the obtained turnout area needs to be regularized according to the size of the remaining allocated area. j∈[6,7], when the kth turnout area of the jth turnout image matches the canvas area, last4=2*h1, which means that the remaining allocated area of the last canvas is not enough to store the next turnout area, and the horizontal coordinate of the lower right corner of the position information of the turnout area in the original turnout image is adjusted to x 2j_k =x 1j_k +(x 2pj_k -x 1pj_k ), that is, the position information of the turnout area in the original turnout image is adjusted to [x 1j_k ,y 1j_k ,x 1j_k +(x 2pj_k -x 1pj_k ),y 2j_k ]; If the turnout area is insufficient, all turnout areas matching area 4 will be spliced together, and the remaining unspliced areas will keep the pixel values of the canvas itself.
[0070] Step 226: Set the turnout area matched in step 225 as P area , the corresponding canvas area is C area , the canvas area C area The pixel values of the corresponding turnout area P area The pixel values of the canvas area not covered (i.e., the area without matching turnout) are still kept at their default pixel values, and a spliced turnout image is obtained;
[0071] Step 227: coordinate transformation is performed on the label of the original turnout image to obtain the spliced turnout image label and the position information of the spliced turnout image label:
[0072] First, let the offset distance of the intercepted turnout area in the x and y directions relative to the origin of the canvas be P x , P y According to the coordinate information of each paired turnout area and the corresponding canvas area obtained in step 225, P x =x 1pj_k -x 1j_k , P y =y 1pj_k -y 1j_k, since the label information of the original turnout image is [num,x min ,y min ,x max ,y max ], so the label information of the original turnout image in the new spliced turnout image becomes [num,x min +P x ,y min +P y ,x max +P x ,y max +P y ].
[0073] During the turnout image stitching process, the labels of the original turnout image may exceed the area range of the turnout area, so it is necessary to remove the invalid labels of each turnout area block. The removal process uses the numpy clip() function to check the new label information [num, x min +P x ,y min +P y ,x max +P x ,y max +P y ], when x min +P x 、x max +P x All in [x 1j_k ,x 2j_k ] outside the interval, the label is removed; when x min +P x 、x max +P x All in [x 1j_k ,x 2j_k ], keep the label; when x min +P x In [x 1j_k ,x 2j_k ] is within the interval and x max +P x In [x 1j_k ,x 2j_k ], correct the new label information to [num,x min +P x ,y min +P y ,x 2j_k ,y max +P y ]; when x max +P x In [x 1j_k ,x 2j_kwithin the interval and x min +P x at [x 1j_k , x 2j_k outside the interval, correct the new label information to [num, x 1j_k , y min +P y , x max +P x , y max +P y , and obtain the spliced turnout image label and its position information;
[0074] Repeat steps 221 - 227 to complete the splicing transformation of all turnout images and corresponding labels in the turnout image dataset with labels.
[0075] When performing splicing transformation on the i - th turnout image and the corresponding label in the i - th calculation process, in step 221, 17 turnout image numbers and corresponding label numbers other than index can be randomly selected from the numbers in indexs and stored in P together with index; at this time, s in step 221 is 1 / 3 of the side length of the resized turnout image canvas, and in step 224, a square canvas of (h1 * 3, h1 * 3) needs to be generated, and the canvas is divided into 9 square sub - regions from left to right and top to bottom, so that in step 225, the turnout regions of two turnout images are used to match one canvas region method to pair the turnout region with the turnout image canvas region. That is, the number of square sub - regions divided by the generated square canvas is half of the number of turnout images and corresponding label numbers stored in the turnout image and corresponding label number list P to be spliced.
[0076] Step 3: Perform affine transformation on the spliced turnout image and the corresponding label, introduce random noise of the input data to avoid overfitting problems during the training process, and obtain the affine - transformed turnout image and the corresponding label;
[0077] The affine transformation mainly includes translation and size scaling of the spliced turnout image and the corresponding label. The specific operations are as follows:
[0078] Step 31: Determine the transformation matrix parameters. The transformation matrices are the centering matrix C, the size scaling matrix R, and the translation matrix T. The centering matrix represents translating the turnout image coordinate system to the center position of the spliced turnout image. Since in the real - world scenario, there is no rotational position relationship of a certain angle for the turnout target, the R matrix only performs size scaling operations on the turnout image, and the rotation angle parameter is set to 0. Specifically, the above matrix expressions are as follows:
[0079]
[0080] Among them, c1, c2, r1, r2, t1, and t2 are the transformation coefficients of the corresponding matrices. Among them, c1 = -s(1 + center1), c2 = -s(1 + center2); center1 and center2 are both hyperparameters of the centering matrix, and the default value is a random number between -0.1 and 0.1 generated by the random.uniform() function in Python; r1 = 1 + revo1, r2 = 1 + revo2, and revo1 and revo2 are both scale hyperparameters, and the default value is a random number between -0.5 and 0.5 generated by the random.uniform() function in Python; the values of t1 and t2 are determined according to the size of the finally output turnout image, t1 = W last *(0.5 + trans1), t2 = H last *(0.5 + trans2), where, W last is the width of the finally output turnout image, H last is the height of the finally output turnout image, and trans1 and trans2 are hyperparameters of the translation matrix, and a random number between -0.2 and 0.2 is generated by the random.uniform() function in Python;
[0081] Step 32: Affinely transform the spliced turnout image and the spliced turnout image label generated in Step 225 and Step 226 respectively using the transformation matrix obtained in Step 31 to obtain the affinely transformed turnout image and the corresponding label.
[0082] Step 4: Crop the turnout image of the specified size from the affinely transformed turnout image and the corresponding label, and clear the smaller fragmented labels to obtain the finally data-augmented turnout image and label information.
[0083] Specifically, crop the affinely transformed turnout image according to the size of the data-augmented target turnout image. The upper left corner of the cropping area is the coordinate origin of the affinely transformed turnout image, so that the pixel coordinate systems of the cropped turnout image and the affinely transformed turnout image are the same coordinate system, so that the turnout image label does not need to perform coordinate transformation, and only the labels within the area range need to be selected according to the cropping area. The size of the cropping area is (W last , H last ), W last is the width of the finally output turnout image, H last is the height of the finally output turnout image, and this size can be set according to the requirements of the feature map size at the input end of the input deep learning network model, and the default value is set to (640 * 640).
[0084] Filter the turnout image labels after affine transformation and keep only the labels within the cropping area. Set the label information after affine transformation to [num 仿 , x 1仿 ,y 1仿 ,x 2仿 ,y 2仿 ], where num 仿 Indicates the label category after affine transformation, num 仿 The value can be 0 or 1. 0 means left turn and 1 means right turn. 1仿 Indicates the horizontal coordinate of the upper left corner of the label box after affine transformation, y 1仿 Indicates the ordinate of the upper left corner of the label box after affine transformation, x 2仿 Represents the horizontal coordinate of the lower right corner of the label box after affine transformation, y 2仿 Indicates the ordinate of the lower right corner of the label box after affine transformation; when x 1仿 、x 2仿 Both in [0,W last ] outside the interval, y 1仿 ,y 2仿 Both in [0,H last ] outside the interval, the label is removed; when x 1仿 、x 2仿 Both in [0,W last ] interval, y 1仿 ,y 2仿 Both in [0,H last ], keep the label; when x 1仿 、x 2仿 Both in [0,W last ] interval, y 1仿 In [0,H last ] interval, y 2仿 In [0,H last ], correct the label information to [num 仿 ,x 1仿 ,y 1仿 ,x 2仿 ,H last ]; when x 1仿 In [0,W last ] interval, x 2仿 In [0,W last ] outside the interval, y 1仿 ,y 2仿 In [0,H last ], correct the label information to [num 仿 ,x 1仿 ,y 1仿 ,W last ,y 2仿 ]; when x2仿 Outside the interval [0, W last , y 2仿 Outside the interval [0, H last , correct the label information to [num 仿 , x 1仿 , y 1仿 , W last , H last .
[0085] The cropped turnout image will generate smaller fragmented target labels. When extracting the label information corresponding to the final output turnout image, these fragmented target labels need to be removed to avoid interference with subsequent use (such as network model training). The data augmentation effect is as Figure 4 shown. Specifically, the label information of the cropped turnout image is denoted as [x 1lat , y 1lat , x 2lat , y 2lat , where x 1lat represents the abscissa of the upper left corner of the cropped label box, y 1lat represents the ordinate of the upper left corner of the cropped label box, x 2lat represents the abscissa of the lower right corner of the cropped label box, and y 2lat represents the ordinate of the lower right corner of the cropped label box; it can be seen from step 222 that the original turnout image label coordinates of the spliced turnout image are [num, x min , y min , x max , y max . Let the ratio of the label area after the shear transformation to the original turnout image label area of the spliced turnout image be R area , and the calculation formula is as follows:
[0086]
[0087] When R area > A min , retain the label after the shear transformation; otherwise, delete the label. A min represents the minimum area ratio for label retention, and the default value is set to 0.5.
[0088] Example 2
[0089] This example provides a parameter optimization method for the data augmentation method of sparse turnout images. To describe the optimization process more clearly, the following uses a specific example in combination with an existing deep learning network model for illustration.
[0090] A parameter optimization method for the data augmentation method of sparse turnout images, as Figure 5 shown, the specific steps are as follows:
[0091] Step S1: Determine the optimizable parameters and their value ranges: The adjustable hyperparameters of a data augmentation method for sparse turnout images include: (1) The left and right boundary expansion coefficients T of the turnout area l , T r ; (2) The hyperparameters center1, center2, revo1, revo2, trans1, trans2 of the affine transformation matrix; (3) The minimum area ratio A for retaining labels min ; (4) The probability value for whether to perform data augmentation, denoted as Prob. The value ranges of each hyperparameter are shown in Table 1 below.
[0092] Table 1 Value ranges of the parameters of a data augmentation method for the sparse turnout images
[0093] Hyperparameter Value range <![CDATA[T l 、T r > 0.30-0.80 <![CDATA[center1, center2]]> -0.10-0.10 <![CDATA[revo1, revo2]]> -0.50-0.50 <![CDATA[trans1, trans2]]> -0.20-0.20 <![CDATA[A min > 0.10-0.80 Prob 0.00-1.00
[0094] Step S2: As Figure 6 shown, encode the optimizable parameters and initialize them within the value ranges: As known from Step S1, the value ranges of each hyperparameter are all decimals between -1 and 1. To facilitate the subsequent search for the hyperparameter space in the following steps, first multiply the determined optimizable parameters by 10^2 to convert the hyperparameters into integers and represent them in binary, so as to design crossover and mutation rules according to the binary bits during subsequent crossover, mutation and other operation processes. Each parameter is arranged in sequence according to the table order, and each parameter occupies 8 bits. The first bit of each 8 bits is the sign bit, 0 indicates no negative sign, and 1 indicates a negative sign. Each parameter is initialized within the value range using the random function;
[0095] Step S3: Denote the number of parameter optimizations as n. When n = 1, use the selected hyperparameters to perform data augmentation on the turnout image dataset, and combine with the deep learning network model for training and testing. The essence of data augmentation is a processing module before the turnout image dataset is input into the deep network model. The effect of data augmentation needs to be evaluated by the improvement of the performance indicators of the deep network model. Convert the hyperparameters encoded in binary in Step S2 into decimal and input them into the deep learning network model for training and testing. The test results are the accuracy P, recall rate R, and mean average precision mAP, and record the hyperparameters used in each calculation and store them in the variable passhyp. When n is not equal to 1, first determine whether the selected hyperparameters are in the passhyp variable. If they are in passhyp, discard the hyperparameters and directly execute Step S6. Otherwise, use the selected hyperparameters to perform data augmentation on the dataset and combine with the deep learning network model for training and testing;
[0096] Specifically, in order to accurately test the performance of the data augmentation algorithm, the deep learning network model preferably selects the current state-of-the-art detection model, such as the YOLOv5 model, etc.;
[0097] Step S4: According to the algorithm test metrics, design the objective function J. In order to comprehensively consider the detection ability of the deep learning network model for turnout categories, set the weight coefficients to balance the contributions of each metric to the objective function. The objective function is expressed as J = 0.2 * P + 0.4 * R + 0.4 * mAP;
[0098] Step S5: When n ≤ 10, record the number of parameter optimizations times, the objective function result J, and the hyperparameters calculated this time in the form of a dictionary in the Best_result variable. The storage form is as follows: {circle: times, hyp: value1, J: value2}, where times represents which parameter optimization this calculation process belongs to, value1 represents the hyperparameters used in this calculation, saved in decimal form, and value2 represents the objective function value corresponding to this calculation, also saved in decimal. The Best_result variable is set to have a maximum storage space of 10 dictionary data; when n > 10, compare the objective function value J calculated this time with the value2 of each item in the Best_result variable. When the J value is greater than the value2 of a certain item, replace the information of that item in the Best_result with the calculation result of this time;
[0099] Specifically, the specific process of converting value1 from binary to decimal is as follows: First, obtain the binary encoding of the hyperparameters. Then, take 8 bits as a unit. After removing the highest bit of each unit, convert it to a decimal number. The highest bit is the sign bit. Finally, divide the converted decimal number by 10^2 to obtain the decimal value of each hyperparameter;
[0100] Specifically, when n ≥ 2, record the number of parameter optimizations, the objective function calculation result J n-1 and the hyperparameter values in the form of a dictionary in the past variable, that is, the past is used to record the value1 and value2 of the previous parameter optimization; when the calculation result of the nth time is recorded in the Best_result, extract the values in the past variable to Best_past. Best_past is used to store the value1 and value2 of the previous parameter optimization of the optimization result recorded in the Best_result, otherwise, directly clear the past;
[0101] Step S6: Screen high-quality hyperparameters. In Best_result, select 5 groups of hyperparameters as the new basic hyperparameters with the J value as the selection probability. The larger the J value, the more likely this group of data will be selected;
[0102] Step S7: Introduce the mutation and crossover ideas of the genetic algorithm to update the hyperparameter encoding. The specific operations are as follows:
[0103] Perform mutation and partial coding crossover operations on the new basic hyperparameters:
[0104] (1) First, extract the high-quality hyperparameters screened in step S6. Let the corresponding number of optimization times be represented by g. Extract the corresponding hyperparameters and objective function values with the optimization times of g - 1 from the Best_past variable. Set the coding mutation weight, determine the mutation probability according to the mutation weight, and then perform mutation operations on the corresponding coding. The specific process is as follows:
[0105] When g = 1, initialize the weights of all hyperparameter codings to 1;
[0106] When g ≥ 2, calculate the generalized gradient grad of J with respect to each hyperparameter g :
[0107]
[0108] where J g represents the objective function value calculated in the g-th calculation, J g-1 represents the objective function value calculated in the (g - 1)-th calculation, x ig is the i-th hyperparameter in the g-th calculation process, x ig-1 is the i-th hyperparameter in the (g - 1)-th calculation process. i takes values from 1 to 10. The specific meanings and value ranges of the 10 hyperparameters are shown in step S1. ΔB i represents the length of the value range corresponding to the i-th hyperparameter, that is, ΔB i = x maxi - x mini , x maxi and x mini represent the maximum and minimum values of the value range of the i-th hyperparameter respectively;
[0109] For example, Figure 7 as shown, when |grad gWhen |grad| > G, where G is the hyperparameter evolution threshold determined by empirical values, it indicates that the objective function enters the sensitive region with respect to this set of hyperparameters and the hyperparameter evolution speed is too fast. Adjust the mutation probabilities of each encoding, increase the mutation probability of the lower bits of the binary data of each hyperparameter, and guide the hyperparameters to reduce the evolution speed. Specifically, first initialize the weights of all hyperparameter encodings to 1, and then use the random.sample() function in Python to randomly select the number of encoding bits in the lower hyperparameter encodings of each hyperparameter, and add a random number between 0 and 1 to the weights of the selected lower hyperparameter encodings. The first four bits of each hyperparameter are the high bits, and the last four bits are the low bits. When |grad| g | ≤ G, it indicates that the objective function enters the sluggish region with respect to this set of hyperparameters and the hyperparameter evolution is too slow. Adjust the mutation probabilities of each encoding, increase the mutation probability of the high bits of the binary data of each hyperparameter, and guide the hyperparameters to generate more significant mutation characteristics and accelerate the evolution process. Specifically, first initialize the weights of all hyperparameter encodings to 1, and then use the random.sample() function in Python to randomly select the number of encoding bits in the high hyperparameter encodings of each hyperparameter, and add a random number between 0 and 1 to the weights of the selected high hyperparameter encodings.
[0110] Specifically, record the weights of each encoding of all obtained hyperparameters as W R , where R ranges from 1 to 80 representing the binary encoding bits of 10 hyperparameters, and determine the mutation probability P of the R-th bit encoding of all hyperparameters according to the weights of each encoding of all hyperparameters R , and the calculation formula is as follows:
[0111]
[0112] Among them, α is the mutation probability coefficient, and the default value is set to 1. The larger the encoding weight W R , the larger the mutation probability P R .
[0113] The binary encoding of each hyperparameter determines whether to perform bit-flip mutation on this encoding bit with P R as the mutation probability. Bit-flip mutation of the encoding bit according to the determined mutation probability is the content in the genetic algorithm. When performing bit-flip mutation, if the original corresponding encoding bit is 0, it becomes 1 after flipping mutation, and if the original corresponding encoding bit is 1, it becomes 0 after flipping mutation.
[0114] (2) After completing the mutation of the hyperparameter encoding, it is also necessary to perform crossover on the mutated groups of hyperparameter encodings. For example Figure 8As shown, the crossover can perform single-encoding swapping or encoding-block swapping to achieve the crossover of process hyperparameter encoding segments, generating new individuals, which is beneficial to forming high-quality hyperparameter combinations. The crossover of process hyperparameter encoding segments refers to swapping some binary bits of the binary encoding of each group of hyperparameters after mutation, which belongs to the crossover operation of the genetic algorithm;
[0115] (3) After completing the hyperparameter crossover, it is necessary to convert the hyperparameters into decimal fractions and then normalize them within the value range of each hyperparameter. The hyperparameters exceeding the value range take the boundary values;
[0116] Step S8: Repeat steps S3 - S7. When the number of iterations reaches the preset value or the objective function reaches the expected goal, extract the optimal result from Best_result as the final hyperparameter value to complete the hyperparameter optimization process.
[0117] The above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.
Claims
1. A data enhancement method for sparse turnout images, characterized in that Including the following steps: Step 1: Label the turnout categories of each turnout image in the turnout image dataset to form a turnout image dataset with labels. The turnout categories are divided into left-turn turnouts and right-turn turnouts; Step 2: Perform turnout image splicing transformation on the turnout images and corresponding labels in the turnout image dataset with labels to form spliced turnout images and corresponding labels; Step 3: Perform affine transformation on the spliced turnout images and corresponding labels to obtain affine-transformed turnout images and corresponding labels; Step 4: Crop the turnout images of the specified size from the affine-transformed turnout images and corresponding labels, and clear the smaller fragmented target labels to obtain the finally data-augmented turnout images and label information; In Step 1, the VIA annotation tool is used to label the turnout categories of the turnout images. The basis for judging the turnout categories is the geometric relationship between the switch rail and the stock rail. When the gap between the switch rail and the stock rail is on the right side, the turnout is considered a right-turn turnout. When the gap between the switch rail and the stock rail is on the left side, the turnout is considered a left-turn turnout; Before Step 2 performs the turnout image splicing transformation on the turnout images and corresponding labels in the turnout image dataset with labels, first number the turnout images and corresponding labels in the turnout image dataset in order to obtain the numbering information of the turnout images to be spliced. The specific process is as follows: Step 211: Number the turnout images and corresponding labels in the dataset in order. The numbering is carried out in units of data pairs. A turnout image and its corresponding label form a data pair. The numbering starts from 00000 and increases sequentially. The numbering information is stored in the indexs list; Step 212: Divide the dataset in units of batch_size in indexs. batch_size represents the batch processing size, which refers to the number of data pairs used in each calculation process. The last data that is less than one batch processing block is also regarded as a batch processing block. The number of divided batch processing blocks i = ceil(len(indexs) / batch_size), where len() represents obtaining the size of the indexs list, and ceil() is a wrapper function of the math module in Python, which means rounding up the calculation result; 2. The data enhancement method for sparse turnout images according to claim 1, wherein The specific operation of Step 2 for performing the turnout image splicing transformation on the turnout images and corresponding labels in the turnout image dataset with labels is as follows: The calculation of the mth batch processing block is called the mth calculation process. Perform splicing transformation on the indexth turnout image and its corresponding label in the mth calculation process, where m ∈ [1, i], and i is the total number of batch processing blocks: Step 221: Sequentially obtain the numbering information of the turnout image and its corresponding label in the mth calculation process. The numbering of the turnout image and its corresponding label obtained each time is recorded as index. Randomly select 7 more numbers of turnout images and corresponding labels from the numbers in the indexs list other than index, and store them together with index in P. P is the numbering list of the turnout images to be spliced and their labels; Step 222: Sequentially extract the corresponding turnout images and label numbers from P. The turnout image number extracted each time is denoted as j, where the value range of j is 0 - 7. Scale the size of the turnout image and label corresponding to each number: The height h1 of the scaled turnout image = r * h, and the width w1 of the scaled turnout image is taken as w1 = r * w, where h is the height of the original turnout image before scaling, w is the width of the original turnout image before scaling, and r is the turnout image scaling ratio, taking r = s / h, where s is 1 / 2 of the side length of the turnout image canvas for size scaling; Extract the label information of the j-th turnout image [num, x 1 min , y 1 min , x 1 max , y 1 max , where num represents the labeled label category; num can take two values, 0 and 1. 0 represents a left-turn turnout, and 1 represents a right-turn turnout; x 1 min represents the abscissa of the upper left corner of the label box, and y 1 min represents the ordinate of the upper left corner of the label box, and x 1 max represents the abscissa of the lower right corner of the label box, and y 1 max represents the ordinate of the lower right corner of the label box; perform an equal-scale scaling operation on the turnout image label information, that is, the label information becomes [num, x min , y min , x max , y max , where [x min , y min , x max , y max = [x 1 min , y 1 min , x 1 max , y 1 max * r, and x min represents the abscissa of the upper left corner of the scaled label box, and y min represents the ordinate of the upper left corner of the scaled label box, and x max represents the abscissa of the lower right corner of the scaled label box, and y max represents the ordinate of the lower right corner of the scaled label box; Step 223: Sequentially extract the turnout area of the j-th turnout image. The turnout area is represented by [x1, y1, x2, y2], where x1 is the abscissa of the upper left corner of the turnout area, y1 is the ordinate of the upper left corner of the turnout area, x2 is the abscissa of the lower right corner of the turnout area, and y2 is the ordinate of the lower right corner of the turnout area. To ensure the richness of the turnout area background, the turnout area is determined according to the label coordinate information, and the calculation formula is as follows: y1 = 0, y2 = h1, x 1= max(x min* (1 - T l ), 0), x2 = min(x max *(1 + T r ), w1), where T l represents the turnout area left boundary expansion coefficient, and T r represents the turnout area right boundary expansion coefficient, which is a random number between 0.3 and 0.
8. Repeat the above process to sequentially extract all turnout areas of the turnout images corresponding to all numbers in P. The turnout image serial number of the extracted turnout area is represented by [j, k], and [j, k] represents the k-th turnout area extracted from the j-th turnout image; Step 224: Use the numpy.zeros() function to generate a canvas of size (h 1* 2, h 1* 2), divide the canvas into four regions, each region being a square region of size h 1* h1. Denote the regions as Region 1, Region 2, Region 3, and Region 4 from left to right and top to bottom. Set the default pixel value of the canvas to 0; Step 225: Adopt the method of matching the turnout areas of two turnout images to one canvas area. Pair the turnout areas for j ∈ [0, 7] with areas 1 - 4 of the turnout image canvas. The turnout areas for j = 0, 1 are paired with canvas area 1; the turnout areas for j = 2, 3 are paired with canvas area 2; the turnout areas for j = 4, 5 are paired with canvas area 3; the turnout areas for j = 6, 7 are paired with canvas area 4. And during the pairing process of areas 1 - 4, sequentially select the turnout areas of the corresponding turnout images to expand from the center point of the canvas to the boundary. When the turnout area overflows the canvas area, intercept the turnout area within the canvas area; Step 226: Set the paired turnout area in Step 225 as P area , and the corresponding canvas area as C area . Then, use the pixel values of the turnout area P area to cover all the pixel values of the canvas area C area . The uncovered canvas area still retains its default pixel values, obtaining a spliced turnout image; Step 227: Perform coordinate transformation on the labels of the original turnout images to obtain the spliced turnout image labels and the position information of the spliced turnout image labels; Then, repeat steps 221 - 227 to complete the splicing transformation of all turnout images and corresponding labels in the turnout image dataset with labels.
3. A data enhancement method for sparse turnout images according to claim 2, characterized in that In step 225, the process of pairing the turnout areas with areas 1 - 4 of the turnout image canvas by adopting the method of matching the turnout areas of two turnout images is as follows: The turnout area of the turnout image is represented by [x 1j_k , y 1j_k , x 2j_k , y 2j_k . Here, x 1j_k refers to the abscissa of the upper-left corner of the k-th turnout area of the j-th turnout image in the original turnout image, and y 1j_k refers to the ordinate of the upper-left corner of the k-th turnout area of the j-th turnout image in the original turnout image. x 2j_k refers to the abscissa of the lower-right corner of the k-th turnout area of the j-th turnout image in the original turnout image, and y 2j_k refers to the ordinate of the lower-right corner of the k-th turnout area of the j-th turnout image in the original turnout image. The corresponding position in the canvas area is [x 1pj_k , y 1pj_k , x 2pj_k , y 2pj_k . Here, x 1pj_k refers to the abscissa of the upper-left corner of the k-th turnout area of the j-th turnout image in the corresponding area of the canvas, and y 1pj_k refers to the ordinate of the upper-left corner of the k-th turnout area of the j-th turnout image in the corresponding area of the canvas. x 2pj_k refers to the abscissa of the lower-right corner of the k-th turnout area of the j-th turnout image in the corresponding area of the canvas, and y 2pj_k refers to the ordinate of the lower-right corner of the k-th turnout area of the j-th turnout image in the corresponding area of the canvas; For area 1, j ∈ [0, 1]: When j = 0 and k = 0, that is, the position information of the 0th turnout area of the 0th turnout image in the original turnout image is [x 10_0 , y 10_0 , x 20_0 , y 20_0 , and the corresponding position in area 1 of the canvas is [x 1p0_0 , y 1p0_0 , x 2p0_0 , y 2p0_0 , x 1p0_0 = max(h1 - (x 20_0 - x 10_0 ), 0), y 1p0_0 = 0, x 2p0_0 = h1, y 2p0_0 = h1; and denote the leftmost abscissa of this turnout area on the canvas as last1, last1 = x 1p0_0 ; Except when j = 0 and k = 0, x 1pj_k = max(last1 - (x 2j_k - x 1j_k ), 0), y 1pj_k = 0, x 2pj_k = max(last1, 0), y 2pj_k = h1, and update the leftmost abscissa last1 = x of this turnout area on the canvas 1pj_k ; When the last remaining allocation area in Region 1 is not sufficient to store the next turnout area, i.e., last1 = 0, the abscissa of the upper left corner of the position information of the turnout area in the original turnout image is adjusted to x 1j_k = x 2j_k -(x 2pj_k - x 1pj_k ), that is, the position information of the turnout area in the original turnout image is adjusted to [x 2j_k -(x 2pj_k - x 1pj_k ), y 1j_k , x 2j_k , y 2j_k ; if the turnout area is insufficient, all the turnout areas matching Region 1 are spliced completely, and the remaining unspliced areas keep the pixel values of the canvas itself; For area 2, j ∈ [2, 3]: When j = 2 and k = 0, that is, the position information of the 0th turnout area of the 2nd turnout image in the original turnout image is [x 12_0 , y 12_0 , x 22_0 , y 22_0 . The corresponding position in the canvas area 2 is [x 1p2_0 , y 1p2_0 , x 2p2_0 , y 2p2_0 . x 1p2_0 = h1, y 1p2_0 = 0, x 2p2_0 = min(h1+(x 22_0 - x 12_0 ), 2*h1), y 2p2_0 = h1; and record the rightmost abscissa of this turnout area on the canvas as last2, last2 = x 2p2_0 ; Except for the case where j = 2 and k = 0, x 1pj_k = min(last2, 2*h1), y 1pj_k = 0, x 2pj_k = min((last2+(x 2j_k -x 1j_k )), 2*h1), y 2pj_k = h1, and update the rightmost abscissa last2 = x of this turnout area on the canvas 2pj_k ; When the remaining allocation area in Region 2 is finally insufficient to store the next turnout area, i.e., last2 = 2 * h1, the abscissa of the lower right corner of the position information of the turnout area in the original turnout image is adjusted to x 2j_k = x 1j_k +(x 2pj_k -x 1pj_k ), that is, the position information of the turnout area in the original turnout image is adjusted to [x 1j_k , y 1j_k , x 1j_k +(x 2pj_k -x 1pj_k ), y 2j_k ; if the turnout area is insufficient, all turnout areas matching Region 2 are spliced completely, and the remaining unspliced area maintains the pixel value of the canvas itself; For area 3, j ∈ [4, 5]: When j = 4 and k = 0, that is, the position information of the 0th turnout area in the 4th turnout image in the original turnout image is [x 14_0 , y 14_0 , x 24_0 , y 24_0 . The corresponding position in the canvas area 3 is [x 1p4_0 , y 1p4_0 , x 2p4_0 , y 2p4_0 . x 1p4_0 = max(h1 - (x 24_0 - x 14_0 ), 0), y 1p4_0 = h1, x 2p4_0 = h1, y 2p4_0 = 2 * h1, and record the leftmost abscissa of this turnout area on the canvas as last3, last3 = x 1p4_0 ; Except for the case where j = 4 and k = 0, x 1pj_k = max(last3 - (x 2j_k - x 1j_k ), 0), y 1pj_k = h1, x 2pj_k = max(last3, 0), y 2pj_k = 2 * h1, and update the leftmost abscissa last3 = x of this turnout area on the canvas 1pj_k ; When the last remaining allocation area in area 3 is not sufficient to store the next turnout area, i.e., last3 = 0, the abscissa of the upper left corner of the position information of the turnout area in the original turnout image is adjusted to x 1j_k = x 2j_k -(x 2pj_k - x 1pj_k ), that is, the position information of the turnout area in the original turnout image is adjusted to [x 2j_k -(x 2pj_k - x 1pj_k ), y 1j_k , x 2j_k , y 2j_k ; if the turnout area is insufficient, all the turnout areas matched with area 3 are spliced completely, and the remaining unspliced areas keep the pixel values of the canvas itself; For area 4, j ∈ [6, 7]: When j = 6 and k = 0, that is, the position information of the 0th turnout area in the 6th turnout image in the original turnout image is [x 16_0 , y 16_0 , x 26_0 , y 26_0 . The corresponding position in the canvas area 4 is [x 1p6_0 , y 1p6_0 , x 2p6_0 , y 2p6_0 . x 1p6_0 = h1, y 1p6_0 = h1, x 2p6_0 = min(h1 + (x 26_0 - x 16_0 ), 2 * h1), y 2p6_0 = 2 * h1, and record the rightmost abscissa last4 of this turnout area on the canvas, last4 = x 2p6_0 ; Except for the case of j = 6 and k = 0, x 1p6_1 = min(last4, 2 * h1), y 1pj_k = h1, x 2pj_k = min((last4 + (x 2j_k - x 1j_k )), 2 * h1), y 2pj_k = 2 * h1, and update the rightmost abscissa last4 = x of this turnout area on the canvas 2pj_k ; When the last remaining allocation area in area 4 is not sufficient to store the next turnout area, i.e., last4 = 2 * h1, the abscissa of the position information of the turnout area in the lower right corner of the original turnout image is adjusted to x 2j_k = x 1j_k +(x 2pj_k -x 1pj_k ), that is, the position information of the turnout area in the original turnout image is adjusted to [x 1j_k , y 1j_k , x 1j_k +(x 2pj_k -x 1pj_k ), y 2j_k ; if the turnout area is insufficient, all the turnout areas matching area 4 are completely spliced, and the remaining unspliced areas maintain the pixel values of the canvas itself to obtain the spliced turnout image.
4. A data enhancement method for sparse turnout images according to claim 2, characterized in that In step 227, first, set the offset distances of the intercepted turnout area in the x and y directions relative to the origin of the canvas as P x and P y . According to the coordinate information of each paired turnout area and the corresponding canvas area obtained in step 225, we can get P x = x 1pj_k - x 1j_k and P y = y 1pj_k - y 1j_k . Here, x 1j_k refers to the abscissa of the upper left corner of the k-th turnout area of the j-th turnout image in the original turnout image, and y 1j_k refers to the ordinate of the upper left corner of the k-th turnout area of the j-th turnout image in the original turnout image. x 1pj_k refers to the abscissa of the upper left corner of the k-th turnout area of the j-th turnout image in the corresponding area of the canvas, and y 1pj_k refers to the ordinate of the upper left corner of the k-th turnout area of the j-th turnout image in the corresponding area of the canvas. Since the label information of the original turnout image is [num, x min , y min , x max , y max , the label information of the original turnout image in the new spliced turnout image becomes [num, x min + P x , y min + P y , x max + P x , y max + P y , thus obtaining the label and position information of the spliced turnout image.
5. A data enhancement method for sparse turnout images according to claim 2, characterized in that, In step 3, the affine transformation includes translation and size scaling of the spliced turnout images and corresponding labels. The specific operations are as follows: Step 31: Determine the transformation matrix parameters. The transformation matrices are the centering matrix C, the size scaling matrix R, and the translation matrix T. The centering matrix C translates the turnout image coordinate system to the center position of the spliced turnout image. The size scaling matrix R only performs size scaling operations on the spliced turnout image, and the rotation angle parameter is set to 0. The above matrix expressions are as follows: Among them, c1, c2, r1, r2, t1, and t2 are the transformation coefficients of the corresponding matrices. Among them, c1 = -s(1 + center1), c2 = -s(1 + center2); center1 and center2 are both hyperparameters of the centering matrix, and the default value is a random number between -0.1 and 0.1 generated by the random.uniform() function in Python; r1 = 1 + revo1, r2 = 1 + revo2, revo1 and revo2 are both scale hyperparameters, and the default value is a random number between -0.5 and 0.5 generated by the random.uniform() function in Python; the values of t1 and t2 are determined according to the size of the finally output turnout image, t1 = W last* (0.5 + trans1), t2 = H last *(0.5 + trans2), where W last is the width of the finally output turnout image, H last is the height of the finally output turnout image, and trans1 and trans2 are hyperparameters of the translation matrix, which are random numbers between -0.2 and 0.2 generated by the random.uniform() function in Python; Step 32: Respectively perform affine transformation on the generated spliced turnout images and corresponding labels using the transformation matrices obtained in step 31 to obtain the affine transformation turnout images and corresponding labels; In step 4, the operation of intercepting the turnout image with a specified size in the affine transformation turnout images and corresponding labels is as follows: Crop the turnout image after affine transformation according to the size of the data-augmented target turnout image. The upper left corner of the cropping area is the coordinate origin of the turnout image after affine transformation, and the size of the cropping area is (W last , H last ); Filter the turnout image labels after affine transformation and keep only the labels within the cropping area. Set the label information after affine transformation to [num 仿 , x 1仿 ,y 1仿 ,x 2仿 ,y 2仿 ], where num 仿 Indicates the label category after affine transformation, num 仿 The value can be 0 or 1. 0 means left turn and 1 means right turn. 1仿 Indicates the horizontal coordinate of the upper left corner of the label box after affine transformation, y 1仿 Indicates the ordinate of the upper left corner of the label box after affine transformation, x 2仿 Represents the horizontal coordinate of the lower right corner of the label box after affine transformation, y 2仿 Indicates the ordinate of the lower right corner of the label box after affine transformation; when x 1仿 、x 2仿 Both in [0,W last ] outside the interval, y 1仿 ,y 2仿 Both in [0,H last ] is outside the interval, the label is removed; when x 1仿 、x 2仿 Both in [0,W last ] interval, y 1仿 ,y 2仿 Both in [0,H last ], keep the label; when x 1仿 、x 2仿 Both in [0,W last ] interval, y 1仿 In [0,H last ] interval, y 2仿 In [0,H last ], correct the label information to [num 仿 ,x 1仿 ,y 1仿 ,x 2仿 ,H last ]; when x 1仿 In [0,W last ] interval, x 2仿 In [0,W last ] outside the interval, y 1仿 ,y 2仿 In [0,H last ], correct the label information to [num 仿 ,x 1仿 ,y 1仿 ,W last ,y 2仿 ]; when x 2仿 Outside the interval [0, W last , when y 2仿 is outside the interval [0, H last , correct the tag information to [num 仿 , x 1仿 , y 1仿 , W last , H last ; In step 4, the operation of clearing smaller fragmented labels is as follows: The label information of the sheared turnout image is denoted as [x 1lat , y 1lat , x 2lat , y 2lat , x 1lat represents the abscissa of the upper left corner of the label box after shearing, and y 1lat represents the ordinate of the upper left corner of the label box after shearing, x 2lat represents the abscissa of the lower right corner of the label box after shearing, and y 2lat represents the ordinate of the lower right corner of the label box after shearing; it can be seen from step 222 that the original turnout image label coordinates of the spliced turnout image are [num, x min , y min , x max , y max , and let the ratio of the label area after shear transformation to the original turnout image label area of the spliced turnout image be R area , and the calculation formula is as follows: When R area > A min , retain the tag after shear transformation; otherwise, delete the tag. A min represents the minimum area ratio for tag retention, and the default value is set to 0.
5.
6. A parameter optimization method for a data enhancement method of sparse turnout images, characterized in that, Perform according to the following steps: Step S1: Determine the optimizable parameters and their value ranges: (1) The left and right boundary expansion coefficients T of the turnout area l , T r ; (2) The hyperparameters center1, center2, revo1, revo2, trans1, trans2 of the affine transformation matrix; (3) The minimum area ratio A for label retention min ; (4) The probability value of whether to perform data augmentation, denoted as Prob; Step S2: Initialize the optimizable parameters, then multiply all the initialized optimizable parameters by 10^2, convert the hyperparameters to integers, and represent them in binary; Step S3: Denote the number of parameter optimizations as n. When n = 1, use the selected hyperparameters to perform data augmentation on the turnout image dataset with labels, and combine with the deep learning network model for training and testing. Convert the hyperparameters encoded in binary in Step S2 into decimal, and input them into the deep learning network model for training and testing. The test results use accuracy P, recall R, and mean average precision mAP, and record the hyperparameters used in each training and testing, and store them in the variable passhyp. When n is not equal to 1, first determine whether the selected hyperparameters are in the passhyp variable. If they are in passhyp, discard the hyperparameters and directly execute Step S6. Otherwise, use the selected hyperparameters to perform data augmentation on the dataset and combine with the deep learning network model for training and testing; Step S4: According to the test metrics of the deep learning network model, design the objective function J. In order to comprehensively consider the detection ability of the deep learning network model for turnout categories, set the weight coefficient to balance the contribution of each index to the objective function. The objective function is expressed as J = 0.2*P + 0.4*R + 0.4*mAP; Step S5: When n ≤ 10, record the number of parameter optimizations, the objective function result J, and the hyperparameters calculated this time in the form of a dictionary in the Best_result variable. The storage form is as follows: {circle:times,hyp:value1,J:value2}, where times represents which parameter optimization this calculation process belongs to, value1 represents the hyperparameters used in this calculation, saved in decimal form, and value2 represents the objective function value corresponding to this calculation, also saved in decimal. The Best_result variable is set with a maximum storage space of 10 dictionary data; when n > 10, compare the objective function value J calculated this time with the value2 of each item in the Best_result variable. When the J value is greater than the value2 of a certain item, replace the information of this item in Best_result with the calculation result of this time; Step S6: In Best_result, select 5 groups of hyperparameters as the new basic hyperparameters with the J value as the selection probability; Step S7: Introduce the mutation and crossover ideas of the genetic algorithm to update the new basic hyperparameters; Step S8: Repeat Steps S3 - S7. When the number of iterations reaches the preset value or the objective function reaches the expected goal, extract the optimal result from Best_result as the final hyperparameter value to complete the hyperparameter optimization process.
7. The parameter optimization method of a data enhancement method for sparse turnout images according to claim 6, characterized in that In Step S7, perform mutation and partial coding crossover operations on the new basic hyperparameters, specifically as follows: (1) First, extract the 5 groups of high-quality hyperparameters selected in Step S6, and use g to represent their corresponding optimization times. Extract the corresponding hyperparameters and objective function values with the optimization times of g - 1 from the Best_past variable, set the coding mutation weight, determine the mutation probability according to the mutation weight, and then perform mutation operations on the corresponding coding according to the mutation probability. The specific process is as follows: When g = 1, all hyperparameter encoding weights are initialized to 1; When g ≥ 2, calculate the generalized gradient grad of J with respect to each hyperparameter g : Among them, J g represents the objective function value of the g-th calculation, and J g-1 represents the objective function value of the (g - 1)-th calculation. x ig is the i-th hyperparameter in the g-th calculation process, and x ig-1 is the i-th hyperparameter in the (g - 1)-th calculation process. i takes values from 1 to 10. For the specific meanings and values of the 10 hyperparameters, refer to step S1. ΔB i represents the length of the value range corresponding to the i-th hyperparameter, that is, ΔB i = x maxi - x mini , x maxi , x mini respectively represent the maximum and minimum values of the value range of the i-th hyperparameter; When |grad g | > G, where G is the hyperparameter evolution threshold. First, initialize the weights of all hyperparameter encodings to 1. Then, use the random.sample() function in Python to randomly select the encoding bits and the number of encoding bits in the lower-order hyperparameter encodings of each hyperparameter, and add a random number between 0 and 1 to the weights of the selected lower-order hyperparameter encodings; When |grad g | ≤ G, first initialize the weights of all hyperparameter encodings to 1, and then use the random.sample() function in Python to randomly select the encoding bits and the number of encoding bits in the high-order hyperparameter encodings of each hyperparameter, and add a random number between 0 and 1 to the weights of the selected high-order hyperparameter encodings; Denote the respective coding weights of all the obtained hyperparameters as W R , where R ranges from 1 to 80 representing the binary coding bits of 10 hyperparameters, and determine the mutation probability P of the R-th bit coding of all the hyperparameters according to the respective coding weight magnitudes of all the hyperparameters R , and the calculation formula is as follows: Among them, α is the mutation probability coefficient, and the default value is set to 1. The encoding weight W R is larger, and the mutation probability P R is larger; The binary encoding of each hyperparameter is based on P R As the mutation probability, it determines whether bit flipping mutation is performed on this encoding bit; (2) Cross the mutated hyperparameter encodings of each group. The crossing is to swap individual encodings or encoding blocks to achieve the crossing of process hyperparameter encoding segments and generate new individuals; (3) After completing the hyperparameter crossing, convert the hyperparameters into decimal fractions and then normalize them within the value range of each hyperparameter. Hyperparameters outside the value range take the boundary values.
8. A parameter optimization method for the data enhancement method of sparse turnout images according to claim 6, characterized in that, The specific process of converting value1 from binary to decimal is as follows: First, obtain the binary encoding of the hyperparameter. Then, take 8 bits as a unit. After removing the highest bit of each unit, convert it into a decimal number. The highest bit is the sign bit. Finally, divide the converted decimal number by 10^2 to obtain the decimal value of each hyperparameter; When n ≥ 2, record the number of parameter optimizations times, the objective function calculation result value2, and the hyperparameter value value1 in the n - 1th calculation result into the past variable in the form of a dictionary. When the nth calculation result is recorded in Best_result, extract the value in the past variable into Best_past. Otherwise, directly clear the past.
Citation Information
Patent Citations
An automatic image data enhancement method, system, medium, and terminal
CN113537406B
Data enhancement method, system and device and computer readable storage medium
CN113570046A
Data enhancement processing method and device, computer equipment and storage medium
CN114549932A
Turnout and non-turnout rail fastener positioning method based on deep learning
CN113506269A
Localization method and system based on deep learning
WO2020173036A1