Target recognition model training method and device, and terminal equipment

By introducing cascaded convolutional branches and identity jump connection layers into the neural network model, the problem of low recognition accuracy in existing neural network models is solved, and higher recognition accuracy is achieved.

CN115761400BActive Publication Date: 2026-04-10苏州万集车联网技术有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
苏州万集车联网技术有限公司
Filing Date
2022-11-03
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing neural network models only include partial details when extracting image features, resulting in low recognition accuracy.

Method used

It employs a neural network block consisting of N cascaded layers. Each layer includes a convolutional branch, a target convolutional layer, and an identity skip connection layer. The convolutional branch and the target convolutional layer use different convolutional kernels, and the identity skip connection layer preserves feature information to generate target features containing more detailed features.

Benefits of technology

It improves the recognition accuracy of the target recognition model, can extract feature information of different dimensions and retain high-dimensional features, thus improving the recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115761400B_ABST
    Figure CN115761400B_ABST
Patent Text Reader

Abstract

Embodiments of the present application are suitable for the technical field of neural networks, and provide a target recognition model training method and device, terminal equipment and a storage medium. The method comprises: obtaining training data; inputting the training data into a first network structure of an initial recognition model to be trained to obtain target features; the first network structure comprises N cascaded neural network blocks, each neural network block comprising a convolution branch, a target convolution layer connected in parallel with the convolution branch, and an identity skip connection layer; the convolution branch and the target convolution layer use different convolution kernels; the identity skip connection layer is used to retain feature information input into the neural network block; and the initial recognition model is trained by using the target features to obtain a target recognition model. The above method can improve the recognition accuracy of the target recognition model.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of neural networks, and particularly relates to a target recognition model training method and device and a terminal device. BACKGROUND

[0002] At present, neural network learning algorithms are widely used in the field of computer vision, such as target detection, image classification and the like. Specifically, an image containing a target object captured by a camera device is labeled as training data, and then an existing neural network model is trained by using the training data to obtain target detection weights.

[0003] However, in the prior art, when extracting image features, the image is usually subjected to convolution processing to obtain the features of the last convolution layer output, and target recognition is performed according to the features. However, the features obtained by this method only contain part of the detailed information in the image, so that the recognition accuracy of the trained neural network model is low when detecting the image. SUMMARY

[0004] The embodiments of the present application provide a target recognition model training method, device, terminal device and storage medium, which can solve the problem of low recognition accuracy of the existing trained recognition model.

[0005] In a first aspect, the embodiments of the present application provide a target recognition model training method, which comprises:

[0006] obtaining training data;

[0007] inputting the training data into a first network structure of an initial recognition model to be trained to obtain target features; the first network structure comprises N cascaded neural network blocks, each neural network block comprising a convolution branch, a target convolution layer connected in parallel with the convolution branch, and an identity skip connection layer; the convolution branch and the target convolution layer use different convolution kernels; and the identity skip connection layer is used to retain the feature information input into the neural network block;

[0008] training the initial recognition model by using the target features to obtain a target recognition model.

[0009] In a second aspect, the embodiments of the present application provide a target recognition model training device, which comprises:

[0010] a obtaining module configured to obtain training data;

[0011] The input module is configured to input training data into a first network structure of an initial identification model to be trained to obtain target features; the first network structure comprises N cascaded neural network blocks, each neural network block comprising a convolution branch, a target convolution layer connected with the convolution branch, and an identity skip connection layer; the convolution branch and the target convolution layer use different convolution kernels; and the identity skip connection layer is configured to retain feature information input into the neural network block.

[0012] The training module is configured to train the initial identification model based on the target features to obtain a target identification model.

[0013] In a third aspect, an embodiment of the present application provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the method of the first aspect when executing the computer program.

[0014] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program, and the computer program is executable on a processor to implement the method of the first aspect.

[0015] In a fifth aspect, an embodiment of the present application provides a computer program product, which, when executed on a terminal device, causes the terminal device to implement the method of the first aspect.

[0016] Compared with the prior art, the embodiment of the present application has the beneficial effects that: training data is obtained, and the training data is input into a first network structure of an initial identification model to be trained to obtain target features. In the first network structure, a neural network block not only comprises a convolution branch, a target convolution layer connected with the convolution branch, and an identity skip connection, but also the convolution branch and the target convolution layer use different convolution kernels. Therefore, when the training data is processed using different convolution kernels, different dimensional feature information can be extracted, and the original high-dimensional feature information in the training data is also retained through the identity skip connection, so that the neural network block can generate target features containing more detailed feature information of the training data. Further, the recognition accuracy of a target identification model trained based on the target features is improved. BRIEF DESCRIPTION OF DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0018] Figure 1is an implementation flowchart of a target recognition model training method provided by an embodiment of the present application;

[0019] Figure 2 is an implementation manner schematic diagram of generating a training image in a target recognition model training method provided by an embodiment of the present application;

[0020] Figure 3 is an implementation manner schematic diagram of determining a target feature in a target recognition model training method provided by an embodiment of the present application;

[0021] Figure 4 is a structure schematic diagram of a neural network block in a target recognition model training method provided by an embodiment of the present application;

[0022] Figure 5 is an implementation manner schematic diagram of obtaining a target recognition model in a target recognition model training method provided by an embodiment of the present application;

[0023] Figure 6 is a structure schematic diagram of a target recognition model training apparatus provided by an embodiment of the present application;

[0024] Figure 7 is a structure schematic diagram of a terminal device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0025] In the following description, specific details are set forth in order to provide a thorough understanding of embodiments of the present application. However, persons of ordinary skill in the art will readily recognize that embodiments of the present application can be practiced without these specific details. In other instances, well-known structures, circuits, and processes have not been described in detail in order to avoid obscuring the present application.

[0026] It should be understood that, when used in the specification and the appended claims of the present application, the term "comprising" indicates the presence of the described features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0027] In addition, in the description of the specification and the appended claims of the present application, the terms "first", "second", "third", and the like are only used to distinguish descriptions, and cannot be understood as indicating or implying relative importance.

[0028] Currently, in order to alleviate the pressure brought by the sharply increasing number of vehicles to the traffic and improve the driving safety, the vehicle-road cooperative system emerges as the times require. In the vehicle-road cooperative system, the traffic information is sensed by the roadside sensing device in multiple dimensions, and then the traffic information is identified by the roadside computing device. Then, according to the identification result, wireless communication and Internet technology are used to interact with the vehicle terminal on the vehicle to improve the driving safety. The traffic information includes, but is not limited to, one or more of the number of obstacles, the speed, the category and the position of the obstacles. The obstacles include, but are not limited to, vehicles, pedestrians, animals or other static obstacles, and are not limited thereto.

[0029] In the vehicle-road cooperative system, the roadside sensing device usually identifies the collected data according to a pre-set neural network model to output corresponding traffic information. However, when the existing neural network model processes the collected data, the data is usually convoluted to obtain the features of the last convolution layer output, and the target in the image is identified according to the features. However, the features obtained by this method cannot contain more detailed information of the target, so that the identification accuracy of the neural network model is low.

[0030] Therefore, in order to improve the identification accuracy of the neural network model, the embodiment of the present application provides a target identification model training method, which can be applied to a terminal device. The terminal device can be the roadside computing device as described above, or the vehicle terminal on the vehicle, and is not limited thereto.

[0031] It can be understood that the vehicle is also usually provided with multiple sensing devices such as camera devices and radar devices to sense the external environment. Therefore, when the terminal device is the vehicle terminal, the following method can also be executed.

[0032] Referring to Figure 1 , Figure 1 An implementation flowchart of a target identification model training method provided by the embodiment of the present application is shown, and the method includes the following steps:

[0033] S101, acquiring training data.

[0034] In an embodiment, the training data is the data required for training the target identification model. The data type of the training data can be a point cloud or an image, and is not limited thereto.

[0035] It can be understood that when the target identification model is an image recognition model, the data type of the training data should be an image; when the target identification model is a point cloud recognition model, the data type of the training data should be a point cloud. In this embodiment, the data type of the training data is an image.

[0036] Specifically, the terminal device can acquire a plurality of video images containing the to-be-identified target collected by the camera device at a plurality of preset time periods; the weather conditions of different time periods are different; then, a plurality of video images of a preset number of frames are randomly selected from the plurality of video images for splicing to obtain a training image. The training data contains the training image.

[0037] In an embodiment, the camera device can be a camera or a photographing camera in a roadside perception device. The plurality of preset time periods can be set by a staff according to actual conditions, which is not limited. The weather conditions include, but are not limited to, sunny, rainy, snowy, and the like.

[0038] It should be noted that, in actual application, the camera device needs to shoot images at each time and under various weather conditions to dynamically improve traffic information for the roadside computing device or the vehicle-mounted terminal in real time. Therefore, in order to improve the identification accuracy of the target identification model, the camera device also needs to shoot video images at a plurality of time periods and under different weather conditions when generating the training data.

[0039] In addition, in the traffic environment in which vehicles travel, the obstacles have the characteristics of large quantity, small target, and being blocked. Therefore, in order to improve the identification accuracy of the obstacles, after the plurality of video images are acquired, a plurality of video images of a preset number of frames can be randomly selected for splicing to obtain a training image.

[0040] The preset number of frames can be set according to actual conditions, for example, the preset number of frames can be 4 frames. The splicing of the video images can be to splice the obstacles contained in each frame of the video images in one frame of image to obtain a training image containing a plurality of obstacles.

[0041] Specifically, the terminal device can generate the training image according to the S201-S203 steps as shown in FIG. 8. Figure 2 The details are as follows:

[0042] S201, performing a preset data enhancement processing on each frame of video image respectively to obtain a plurality of first images.

[0043] In an embodiment, the data enhancement processing includes, but is not limited to, random flipping, translation, or pixel value change of the video image, which is not described in detail. The purpose of performing the data enhancement on each frame of video image is to expand the number of training data as much as possible so that the target identification model has sufficient training data.

[0044] S202, respectively cutting a second image containing the to-be-identified target from each first image.

[0045] S203, splicing a plurality of second images to obtain a training image.

[0046] In an embodiment, the target to be identified is the obstacle described above, wherein one or more of the following conditions may exist in a frame of video image: the number of targets to be identified is small, the target is large, and the target is not occluded. Therefore, the terminal device can crop a second image containing the target to be identified from the first image. At this time, the image range of the second image is generally smaller than that of the first image. For example, the terminal device can generate a cropping frame for the target to be identified to crop the target to obtain the second image. Finally, the terminal device can stitch the edges of each frame of the second image to form a training image. At this time, the training image can include a plurality of targets to be identified, and the occluded target to be identified can also be removed.

[0047] In an embodiment, when the second image is obtained by cropping the target to be identified, the staff can also count the aspect ratio corresponding to each second image to obtain an optimal aspect ratio from a plurality of aspect ratios, so that the collected image can meet the requirements of the trained target recognition model for processing. For example, the number of second images corresponding to each aspect ratio is counted, and the aspect ratio corresponding to the maximum number is determined as the optimal aspect ratio.

[0048] It should be noted that the terminal device can obtain one frame of training image by executing the above S201-S203 steps once. If the training data needs to include multiple frames of training images, the terminal device also needs to execute the above S201-S203 steps multiple times, which will not be described herein.

[0049] S102, input the training data into a first network structure of an initial recognition model to be trained to obtain target features; the first network structure includes N cascaded neural network blocks, each neural network block includes a convolution branch, a target convolution layer connected with the convolution branch, and an identity skip connection layer; the convolution branch and the target convolution layer use different convolution kernels; and the identity skip connection layer is used to retain the feature information input into the neural network block.

[0050] In an embodiment, the initial recognition model described above can be a neural network model constructed based on a two-stage target detection algorithm. The two-stage target detection algorithm includes but is not limited to a region-based convolutional neural network (R-CNN), a cascade region-based convolutional neural network (Cascade R-CNN), and the like, which will not be limited herein. In this embodiment, the initial recognition model described above can be a Cascade R-CNN model.

[0051] It should be noted that the first network structure can adopt a ResNeXt50 structure to replace the original residual neural network 50 (Residual Neural Network, ResNet 50) structure in the Cascade R-CNN model to improve the feature extraction capability.

[0052] In addition, in order to further improve the feature extraction capability, in each neural network block of the first network structure, not only the convolution branch originally used to process the training data is included, but also a target convolution layer and an identity skip connection layer are added in parallel with the convolution branch.

[0053] The convolution kernel used in the convolution branch needs to be different from the convolution kernel used in the target convolution layer. In this way, the neural network block can use different convolution kernels to process the input features to extract feature information of different dimensions from the input features. In addition, the identity skip connection layer can also be used to retain the feature information input to the neural network block. Based on this, the neural network block can perform multi-dimensional information fusion on the feature information of different dimensions and the originally input feature information, so that the feature finally output by the neural network block retains the feature information of each dimension and is input to the next level neural network block until the target feature is output by the N-level neural network block. Wherein, N is usually an integer greater than 1.

[0054] Specifically, referring to Figure 3 , the terminal device can obtain the target feature according to steps S301-S303 as shown in Figure 3 . Details are as follows:

[0055] S301, input the training data into the first-level neural network block, process the training data through the convolution branch and the target convolution layer in the first-level neural network block, obtain the target output feature of the first-level neural network block, and output the target output feature to the second-level neural network block.

[0056] In an embodiment, since the first-level neural network block is the first module to process the training data among all neural network blocks, the data input into the first-level neural network block is the training data. At the same time, since the identity skip connection layer is used to retain the feature information input to the neural network block, and the target output feature of the first-level neural network block is also the feature generated by directly processing the training data. Therefore, the terminal device can directly determine the feature obtained by processing the training data through the convolution branch and the target convolution layer in the first-level neural network block as the target output feature of the first-level neural network block. Without the need to perform feature fusion processing again on the feature obtained by processing the training data through the convolution branch and the target convolution layer in the first-level neural network block and the training data retained by the identity skip connection layer to obtain the target output feature, the speed of processing the training data by the initial identification model is improved.

[0057] In another embodiment, the terminal device can also perform feature fusion processing on the features obtained by the convolution branch and the target convolution layer in the first-level neural network block after processing the training data in the first-level neural network block, and the training data retained by the identity skip connection layer to obtain the target output features. At this time, although the speed of processing the training data by the initial identification model is reduced, the generated target output features can retain as much high-dimensional feature information in the training data as possible.

[0058] In each neural network block, the structure of the convolution branch and the target convolution layer is the same. The convolution branch can include one convolution layer or multiple convolution layers, which is not limited. When only one convolution layer is included, the convolution kernel of the convolution layer should be different from the convolution kernel of the target convolution layer; and when multiple convolution layers are included, the convolution kernel of at least one convolution layer in the multiple convolution layers is different from the convolution kernel of the target convolution layer, which is not limited.

[0059] Specifically, referring to Figure 4 , Figure 4 A structure diagram of a neural network block in a target identification model training method provided by an embodiment of the present application is shown. Wherein, "256, 1*1, 128" represents a convolution layer in a convolution branch, which means that the input is a feature of 256 dimensions, and after processing by a 1*1 convolution kernel, a feature of 128 dimensions is output to the next convolution layer in the convolution branch. Wherein, Group=C means that the feature maps of the input features to the "128, 3*3, 128" convolution layer are divided into C groups for processing.

[0060] In each convolution branch, the first convolution layer uses a 1*1 convolution kernel to compress the target output feature channel input from the previous neural network block; then a 3*3 convolution kernel is used to reduce the resolution, i.e. the height and width of the feature map corresponding to the target output feature of the previous neural network block are halved, and finally a 1*1 convolution kernel is used to increase the depth, so as to improve the ability to extract multi-dimensional feature information.

[0061] It should be noted that the purpose of adding an identity skip connection layer in each of the N cascaded neural network blocks is to perform multiple identity skip connections on the target input features input to the neural network block, which can improve the probability of learning feature information from the target input features each time, so that the final generated target output features can include more detail information of the to-be-identified image in the training data.

[0062] S302, for each of the second to the Nth neural network block, the target output feature of the previous neural network block is processed by the convolution branch and the target convolution layer in the neural network block to obtain the first output feature and the second output feature of the neural network block, and the first output feature, the second output feature and the target output feature reserved by the identity skip connection are fused to obtain the target output feature of the neural network block, and the target output feature is output to the next neural network block.

[0063] S303, the target output feature output by the Nth neural network block is determined as the target feature.

[0064] In an embodiment, for each of the neural network blocks, the convolution branch and the target convolution layer in the neural network block are used to process the target output feature of the previous neural network block to obtain the first output feature and the second output feature. Then, the neural network block also fuses the first output feature, the second output feature and the target output feature reserved by the identity skip connection layer to obtain the target output feature of the neural network block. Wherein, the fusion method can be splicing or weighting the first output feature, the second output feature and the target output feature reserved by the identity skip connection layer, which is not limited.

[0065] Wherein, the target output feature of each neural network block can be represented by the following formula:

[0066] T(x)=x+g(x)+f(x);

[0067] Wherein, T(x) represents the target output feature of the current neural network block; x represents the target output feature of the previous neural network block; g(x) represents the first output feature; f(x) represents the second output feature.

[0068] It can be understood that the target feature is the target output feature output by the Nth neural network block (i.e. the last neural network block).

[0069] S103, training the initial recognition model by the target feature to obtain a target recognition model.

[0070] In an embodiment, the initial recognition model can predict the to-be-recognized target in the training data according to the target feature, output the predicted position and the predicted category of the to-be-recognized target. Then, according to the real position and the real category of the to-be-recognized target labeled in advance, the training loss value is calculated to train the initial recognition model.

[0071] Specifically, the terminal device can train the initial recognition model according to the S501-S504 steps as shown in Figure 5 The details are as follows:

[0072] S501, obtaining a predicted position and a predicted category of the target to be recognized in the training data by the initial recognition model.

[0073] S502, calculating a training loss value based on the predicted position, the predicted category, the real position and the real category.

[0074] In an embodiment, the initial recognition model described above generally outputs a predicted probability of belonging to the predicted category when outputting the predicted category of the target to be recognized. Based on this, the terminal device can calculate a position loss value according to the predicted position and the real position, and calculate a category loss value according to the real category and the predicted probability of the predicted category. Finally, the training loss value is determined according to the position loss value and the category loss value.

[0075] Wherein, the predicted position is generally a predicted bounding box containing the target to be recognized; the real position is a labeled bounding box of the actual target to be recognized. Based on this, the terminal device can import the predicted position and the real position into the following position loss function to obtain the position loss value:

[0076]

[0077] Wherein, GIOU represents the position loss value; IOU represents the ratio of the intersection area of the predicted bounding box and the labeled bounding box to the merged area; C represents the area of the boundary of the predicted bounding box and the labeled bounding box enclosed by the smallest rectangle; A∪B represents the union set of the predicted bounding box and the labeled bounding box.

[0078] Wherein, the calculation of the position loss value by using the above position loss function can optimize the regression accuracy of the boundary box and improve the accuracy of the target recognition model trained to predict the position of the target to be recognized.

[0079] In addition, the loss function for calculating the category loss value can be a logistic regression loss function, a cross-entropy loss function, which will not be described in detail.

[0080] S503, determining a target learning rate for updating the model parameters in the initial recognition model at the current iteration number.

[0081] In an embodiment, the target learning rate can be a preset value or determined according to the current iteration number, which is not limited. In the embodiment, in order to improve the convergence effect of the initial identification model, the terminal device can first determine the iteration number of the model parameter in the initial identification model and the historical learning rate of the model parameter in the last iteration. Then, the ratio of the iteration number to the preset number is determined as the first weight, and the second weight corresponding to the current iteration number is determined according to the association relationship between the preset iteration number and the learning weight. Finally, the target learning rate under the current iteration number is determined according to the historical learning rate, the first weight and the second weight.

[0082] In the initial model, the model parameter usually needs to be iterated for multiple times during training. Based on this, during the iteration process, the terminal device can record the iteration number of the model parameter and the historical warm-up factor at each iteration. That is, the terminal device can directly determine the iteration number and the historical warm-up factor at the last iteration of the model parameter.

[0083] In an embodiment, the above association relationship can be preset by the staff according to the actual situation. For example, the staff can preset a weight adjustment interval of the learning rate, which records the iteration number interval corresponding to each second weight. The terminal device can determine the second weight corresponding to the current iteration number according to the iteration number interval corresponding to the current iteration number.

[0084] In a specific embodiment, the terminal device can import the historical warm-up factor, the first weight and the second weight into the preset learning rate formula to calculate the target learning rate. The preset learning rate formula is as follows:

[0085] l r =base_lr*S r *β 2 ;

[0086] S r =α+S r-1 *(1-α)*γ;

[0087] Wherein, l r represents the target learning rate, S r-1 represents the historical warm-up factor, a represents the first weight, b represents the second weight, g represents the preset constant, and base_lr represents the preset learning rate.

[0088] Based on the above formula, the historical warm-up factor S r-1 is used to calculate the warm-up factor S r corresponding to the current iteration number. It can be understood that in the r+1 iteration process, S r will be used as the historical warm-up factor to participate in the next iteration.

[0089] And, the target learning rate calculated by the preset learning rate formula can make the initial recognition model start training at a larger rate, and then slow down, ensuring that the learning rate is very low at the beginning and then increases rapidly, ensuring convergence while improving convergence speed.

[0090] S504, update the model parameters according to the training loss value and the target learning rate until the number of iterations reaches the preset number, and then use the model parameters after the preset number of iterations as the model parameters of the target recognition model to obtain the target recognition model.

[0091] In an embodiment, the back propagation based on the training loss value and the target learning rate to update the model parameters is the basic process of neural network model training, which will not be described in detail.

[0092] The preset number is a judgment condition for ending the training of the initial recognition model. The terminal device can use the model parameters after the preset number of iterations as the model parameters of the target recognition model to obtain the target recognition model. The preset number can be set by the staff according to the actual situation, which is not limited.

[0093] In another embodiment, the judgment condition can also be that the training loss value obtained continuously for multiple times is lower than the preset value, or the fluctuation amplitude is less than the preset amplitude, and the judgment condition in this embodiment is not limited.

[0094] In this embodiment, by obtaining the training data and inputting the training data into the first network structure of the initial recognition model to be trained for processing, the target feature is obtained. Because the neural network block in the first network structure not only includes a convolution branch, a target convolution layer connected with the convolution branch, and an identity skip connection, and the convolution kernels used in the convolution branch and the target convolution layer are different. Therefore, when different convolution kernels are used to process the training data, different dimensional feature information can be extracted, and the original high-dimensional feature information in the training data is also retained through the identity skip connection, so that the neural network block can generate a target feature containing more detailed feature information in the training data. Further, the recognition accuracy of the target recognition model trained based on the target feature is improved.

[0095] In another embodiment, in the above S103 step, in order to improve the accuracy of the initial recognition model in predicting the target feature, the terminal device can also perform secondary processing on the target feature to obtain a processed target feature, and then train the initial recognition model according to the processed target feature to obtain the target recognition model. For example, the processed target feature is subjected to the above S501-S504 steps.

[0096] Specifically, the terminal device can input the target feature obtained in the S102 step into a second network structure preset in the initial recognition model for feature processing to obtain a processed target feature. The second network structure is specifically a path aggregation network (PANet), which can be used to further fuse feature information in multiple dimensions and improve the extraction capability of low-dimensional feature information.

[0097] It should be noted that when the model parameters in the initial recognition model are iterated, the model parameters in the second network structure should also be iterated at the same time.

[0098] In another embodiment, in order to obtain an optimal target recognition model, the terminal device can repeat the above S101-S103 steps to obtain multiple target recognition models. Then, a verification set containing the target to be recognized is used to verify each target recognition model to obtain a verification recognition result corresponding to each target recognition model. Then, according to the verification recognition result, a target recognition model to be tested is determined from all target recognition models, and the target recognition model to be tested is used to identify the test set to obtain a test recognition result. Then, it is determined whether the first evaluation index of the target recognition model to be tested meets the preset evaluation index based on the test recognition result. If the first evaluation index meets the preset evaluation index, the target recognition model to be tested is determined as the optimal target recognition model; if the first evaluation index does not meet the preset evaluation index, the steps of determining the target recognition model and the optimal target recognition model are re-executed.

[0099] The verification set is used to verify the recognition result of each target recognition model in identifying the target to be recognized. The verification recognition result can be one or more evaluation indexes such as the accuracy, recall rate, and mean average precision (mAP) of the target recognition model in identifying the target to be recognized. Then, the target recognition model whose evaluation index meets the preset evaluation index is determined as the target recognition model to be tested.

[0100] In an embodiment, the test set is used to test the target recognition model to be tested. The preset evaluation index can include not only the above-mentioned accuracy, recall rate, and mean average precision, but also indexes such as the missed detection rate and false detection rate. The types and number of preset evaluation indexes used for evaluation are not limited in this embodiment.

[0101] In this embodiment, the multiple target recognition models obtained are sequentially verified and tested by using the verification set and the test set to obtain the optimal target recognition model. In this way, the recognition accuracy of the optimal target recognition model obtained finally can be improved.

[0102] It should be noted that the number of data in the training data, the validation set and the test set can be set according to actual conditions, and is not limited. In the embodiment, in order to improve the recognition accuracy of the target recognition model, the number of data contained in the training data can be much larger than the number of data contained in the validation set and the test set. Specifically, the ratio of the number of data in the training data, the validation set and the test set can be 7:1:2.

[0103] Please refer to Figure 6 , Figure 6 is a structural block diagram of a target recognition model training device provided by an embodiment of the present application. The target recognition model training device in the embodiment includes various modules for executing Figures 1 to 3 、 Figure 5 steps in the corresponding embodiments. For details, please refer to the related description in Figures 1 to 3 、 Figure 5 and Figures 1 to 3 、 Figure 5 corresponding embodiments. For ease of illustration, only the parts related to the present embodiment are shown. Referring to Figure 6 , the target recognition model training device 600 can include an acquisition module 610, an input module 620 and a training module 640, wherein:

[0104] The acquisition module 610 is configured to acquire training data.

[0105] The input module 620 is configured to input the training data into a first network structure of an initial recognition model to be trained to obtain target features; the first network structure includes N cascaded neural network blocks, each neural network block includes a convolution branch, a target convolution layer connected with the convolution branch, and an identity skip connection layer; the convolution branch and the target convolution layer use different convolution kernels; and the identity skip connection layer is configured to retain feature information input into the neural network block.

[0106] The training module 630 is configured to train the initial recognition model by using the target features to obtain a target recognition model.

[0107] In an embodiment, the training data includes training images; and the acquisition module 610 is further configured to:

[0108] acquire a plurality of video images containing a target to be recognized collected by a camera device at a plurality of preset time periods; the weather conditions of different time periods are different; and randomly select a preset number of video images from the plurality of video images to splice to obtain the training images.

[0109] In an embodiment, the acquisition module 610 is further configured to:

[0110] The preset data enhancement processing is respectively performed on each frame of video image to obtain a plurality of first images; a second image containing the to-be-recognized target is respectively cut from each first image; and the plurality of second images are spliced to obtain a training image.

[0111] In an embodiment, the input module 620 is further configured to:

[0112] The training data is input into the first neural network block, the training data is processed by the convolution branch and the target convolution layer in the first neural network block, the target output feature of the first neural network block is obtained, and the target output feature is output to the second neural network block; for each neural network block in the second to Nth neural network block, the target output feature of the previous neural network block is processed by the convolution branch and the target convolution layer in the neural network block, the first output feature and the second output feature of the neural network block are obtained, the first output feature, the second output feature and the target output feature reserved by the identity skip connection are fused to obtain the target output feature of the neural network block, and the target output feature is output to the next neural network block; and the target output feature output by the Nth neural network block is determined as the target feature.

[0113] In an embodiment, the training module 630 is further configured to:

[0114] The target feature is input into the second network structure in the initial recognition model for feature processing to obtain a processed target feature; the second network structure is a path aggregation network;

[0115] The initial recognition model is trained according to the processed target feature to obtain a target recognition model.

[0116] In an embodiment, the training module 630 is further configured to:

[0117] The initial recognition model is trained according to the processed target feature to obtain a target recognition model.

[0118] In an embodiment, the training module 630 is further configured to:

[0119] Determine the number of iterations of the model parameters in the initial recognition model, and the historical warm-up factor of the model parameters in the last iteration; determine the ratio of the number of iterations to the preset number of iterations as the first weight; determine the second weight corresponding to the current iteration number based on the correlation between the preset number of iterations and the learning weight; determine the target learning rate for the current iteration number based on the historical warm-up factor, the first weight, and the second weight.

[0120] In one embodiment, the training module 630 is further configured to:

[0121] Import the historical warm-up factor, the first weight, and the second weight into the preset learning rate formula to calculate the target learning rate; the preset learning rate formula is as follows:

[0122] l r =base_lr*S r *β 2 ;

[0123] S r =α+S r-1 *(1-α)*γ;

[0124] Among them, l r S represents the target learning rate. r-1 The term "historical warm-up factor" refers to the historical warm-up factor, where α represents the first weight, β represents the second weight, γ represents a preset constant, and base_lr represents a preset learning rate.

[0125] When it is understood that, Figure 6 The block diagram of the target recognition model training device shown illustrates that each module is used to perform... Figures 1 to 3 , Figure 5 The steps in the corresponding embodiments, and for Figures 1 to 3 , Figure 5 The steps in the corresponding embodiments have been explained in detail in the above embodiments. Please refer to them for details. Figures 1 to 3 , Figure 5 as well as Figures 1 to 3 , Figure 5 The relevant descriptions in the corresponding embodiments will not be repeated here.

[0126] Figure 7 This is a structural block diagram of a terminal device provided in one embodiment of this application. For example... Figure 7 As shown, the terminal device 700 of this embodiment includes: a processor 710, a memory 720, and a computer program 730 stored in the memory 720 and executable on the processor 710, such as a program for a target recognition model training method. When the processor 710 executes the computer program 730, it implements the steps of each embodiment of the target recognition model training method described above, for example... Figure 1S101 to S103. Alternatively, the processor 710 implements the functions of the above-described modules 610 to 630 when executing the computer program 730. Figure 6 The functions of the modules 610 to 630 are described above in the corresponding embodiments, for example, Figure 6 The functions of the modules 610 to 630 are described above in the corresponding embodiments, for example, Figure 6 The functions of the modules 610 to 630 are described above in the corresponding embodiments, for example,

[0127] The computer program 730 can be segmented into one or more modules that are stored in the memory 720 and executed by the processor 710 to implement the target recognition model training method provided by the embodiments of the present application. One or more modules can be a series of computer program instruction segments capable of completing a specific function, which are used to describe the execution process of the computer program 730 in the terminal device 700. For example, the computer program 730 can implement the target recognition model training method provided by the embodiments of the present application.

[0128] The terminal device 700 can include, but is not limited to, the processor 710 and the memory 720. Those skilled in the art can understand that the terminal device 700 can further include other components required for the terminal device 700 to perform the target recognition model training method provided by the embodiments of the present application. Figure 7 The terminal device 700 is only an example and does not constitute a limitation on the terminal device 700, and can include more or fewer components than those shown, or combine certain components, or different components, for example, the terminal device can also include an input / output device, a network access device, a bus, etc.

[0129] The processor 710 can be a central processing unit, and can also be other general-purpose processors, digital signal processors, application-specific integrated circuits, ready programmable gate arrays or other programmable logic devices, discrete gates or transistor logic components, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.

[0130] The memory 720 can be an internal storage unit of the terminal device 700, such as a hard disk or a memory of the terminal device 700. The memory 720 can also be an external storage device of the terminal device 700, such as a plug-in hard disk, a smart memory card, a flash memory card, etc. equipped on the terminal device 700. Further, the memory 720 can include both the internal storage unit and the external storage device of the terminal device 700.

[0131] The embodiments of the present application provide a computer readable storage medium, including a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the target recognition model training method in the above-described various embodiments when executing the computer program.

[0132] The embodiment of the present application provides a computer program product, when the computer program product runs on a terminal device, makes the terminal device execute the target identification model training method in each of the above embodiments.

[0133] The above embodiments are only used to illustrate the technical solutions of the present application, but not limit the same; although the present application is described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.

Claims

1. A method for training a target recognition model, characterized in that, The method includes: Obtain training data; The training data is input into the first network structure of the initial recognition model to be trained to obtain target features. The first network structure includes N cascaded neural network blocks. Each level of the neural network block includes a convolutional branch, a target convolutional layer connected in parallel with the convolutional branch, and an identity skip connection layer. The convolutional branch and the target convolutional layer use different convolutional kernels. The identity skip connection layer is used to retain the feature information input to the neural network block. The initial recognition model is trained using the target features to obtain the target recognition model; The step of inputting the training data into the first network structure of the initial recognition model to be trained to obtain target features includes: The training data is input into the first-level neural network block, and the training data is processed by the convolutional branch and the target convolutional layer in the first-level neural network block to obtain the target output feature of the first-level neural network block, and the target output feature is output to the second-level neural network block; For each level of the neural network block from the second to the Nth level, the target output features of the previous level neural network block are processed by the convolutional branches and the target convolutional layer in the neural network block to obtain the first output feature and the second output feature of the neural network block. The first output feature, the second output feature and the target output feature preserved by the identity jump connection are fused to obtain the target output feature of the neural network block. The target output feature is then output to the next level neural network block. The target output feature of the Nth level neural network block is determined as the target feature.

2. The method according to claim 1, characterized in that, The training data includes training images; acquiring the training data includes: Acquire multiple frames of video images containing the target to be identified, captured by the camera device at multiple preset time periods; the weather conditions vary at different time periods; The training image is obtained by randomly selecting a preset number of video images from the multi-frame video images and stitching them together.

3. The method according to claim 2, characterized in that, The step of randomly selecting a preset number of video images from the multi-frame video images and stitching them together to obtain the training image includes: Each frame of the video image is subjected to a preset data enhancement process to obtain multiple first frames; A second image containing the target to be identified is extracted from each frame of the first image; The training image is obtained by stitching together multiple frames of the second image.

4. The method according to claim 1, characterized in that, The step of training the initial recognition model using the target features to obtain the target recognition model includes: The target features are input into the second network structure of the initial recognition model for feature processing to obtain the processed target features; the second network structure is a path aggregation network. The initial recognition model is trained based on the processed target features to obtain the target recognition model.

5. The method according to any one of claims 1-4, characterized in that, The training data includes the true location and true category of the target to be identified; The step of training the initial recognition model using the target features to obtain the target recognition model includes: Obtain the predicted location and predicted category of the target to be identified in the training data predicted by the initial recognition model; Calculate the training loss value based on the predicted location, the predicted category, the true location, and the true category; Determine the target learning rate for updating the model parameters in the initial recognition model at the current iteration number; The model parameters are updated based on the training loss value and the target learning rate until the number of iterations reaches a preset number. The model parameters after the preset number of iterations are then used as the model parameters of the target recognition model to obtain the target recognition model.

6. The method according to claim 5, characterized in that, Determining the target learning rate for updating the model parameters in the initial recognition model at the current iteration number includes: Determine the number of iterations of the model parameters in the initial identification model, and the historical warm-up factor of the model parameters in the last iteration; The ratio of the number of iterations to the preset number is determined as the first weight; Based on the correlation between the preset number of iterations and the learning weights, determine the second weight corresponding to the current iteration number; The target learning rate for the current iteration number is determined based on the historical warm-up factor, the first weight, and the second weight.

7. The method according to claim 6, characterized in that, The step of determining the target learning rate for the current iteration number based on the first weight and the second weight of the historical warm-up factors includes: The historical warm-up factor, the first weight, and the second weight are imported into a preset learning rate formula to calculate the target learning rate; the preset learning rate formula is as follows: in, This represents the target learning rate. This refers to the historical warm-up factor. This represents the first weight. This represents the second weight. This represents a preset constant. .

8. A target recognition model training device, characterized in that, The device includes: The acquisition module is used to acquire training data; An input module is used to input the training data into the first network structure of the initial recognition model to be trained to obtain target features; the first network structure includes N cascaded neural network blocks, each level of the neural network block includes a convolutional branch, a target convolutional layer connected in parallel with the convolutional branch, and an identity skip connection layer; the convolutional branch and the target convolutional layer use different convolutional kernels; the identity skip connection layer is used to retain the feature information input to the neural network block; The training module is used to train the initial recognition model using the target features to obtain the target recognition model; The input module is also used for: The training data is input into the first-level neural network block, and the training data is processed by the convolutional branch and the target convolutional layer in the first-level neural network block to obtain the target output feature of the first-level neural network block, and the target output feature is output to the second-level neural network block; For each level of the neural network block from the second to the Nth level, the target output features of the previous level neural network block are processed by the convolutional branches and the target convolutional layer in the neural network block to obtain the first output feature and the second output feature of the neural network block. The first output feature, the second output feature and the target output feature preserved by the identity jump connection are fused to obtain the target output feature of the neural network block. The target output feature is then output to the next level neural network block. The target output feature of the Nth level neural network block is determined as the target feature.

9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Image processing method and device, storage medium and electronic device

    CN110991298A

  • Neural network distillation method, target detection method and device

    CN115018039A