Target recognition method, device, electronic device and storage medium

By building a training network with decoupled detection heads and using the target training sample set for model training, the problem of slow operation of existing object detection methods is solved, and fast and accurate object recognition is achieved.

CN114565780BActive Publication Date: 2025-05-16COSMO INSTITUTE OF INDUSTRIAL INTELLIGENCE (QINGDAO) CO LTD +3
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210175862.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-25
Publication Date
2025-05-16
Estimated Expiration
2042-02-25

AI Technical Summary

Technical Problem

The existing target detection methods are slow to run and are difficult to meet the real-time detection requirements.

Method used

By obtaining the image to be detected and determining the target training sample set based on its source information, a training network with decoupled detection heads is built, and initial training is performed through the initial data set, and then further training the model is used to use the target training sample set to obtain the target recognition model.

Benefits of technology

It realizes the rapid identification model based on the region extraction algorithm, which improves the speed and accuracy of target recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114565780B_ABST
    Figure CN114565780B_ABST
Patent Text Reader

Abstract

The present invention discloses a target recognition method, device, electronic device and storage medium. The method includes: obtaining an image to be detected, and determining a target training sample set corresponding to the image to be detected according to the source information of the image to be detected; building a training network according to a decoupled detection head, and training the training network according to an initial data set to obtain an initial recognition model; training the initial recognition model according to a target training sample set to obtain a target recognition model; inputting the image to be detected into the target recognition model to perform target recognition, and obtaining a detection result of the image to be detected. That is, in an embodiment of the present invention, the convergence speed of the model is improved from the network structure and data source by building a training network and determining target training samples; the training network is initially trained by an initial data set to obtain an initial recognition model, and then different targets are trained according to the target training sample set, and the convergence speed of the model is improved from the training step setting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to computer technology, and in particular to a target recognition method, device, electronic device and storage medium. Background Art

[0002] With the rapid development of computer vision research, target detection plays an indispensable role in the fields of intelligent industrial manufacturing and artificial intelligence information technology. Among them, image recognition is mainly used to determine whether an object instance of a predefined type exists in the target image, and locate the spatial position and range of the target to be identified through the region box. When there is a target in the target image, the spatial position and range of the identified target are returned as the recognition result. Safety detection in industrial manufacturing requires manual observation. Because of the high-intensity work, it is easy to cause dangers caused by human fatigue, so vision-based detection is very necessary. Existing target detection methods may generate a large number of prior frames containing the objects to be detected, and then use classifiers to determine whether each prior frame contains the object to be detected and the confidence of the category to which the object belongs. At the same time, post-processing is required to correct the bounding box. Finally, based on some criteria, the bounding boxes with low confidence and high overlap are filtered out to obtain the detection results. Although this method has a relatively high detection accuracy, it runs slowly. Summary of the invention

[0003] The present invention provides a target recognition method, device, electronic equipment and storage medium to realize rapid acquisition of a recognition model based on a region extraction algorithm, so as to facilitate recognition of different targets.

[0004] In a first aspect, an embodiment of the present invention provides a target recognition method, the method comprising:

[0005] Acquire an image to be detected, and determine a target training sample set corresponding to the image to be detected according to source information of the image to be detected;

[0006] Building a training network according to the decoupled detection head, and training the training network according to the initial data set to obtain an initial recognition model;

[0007] Training the initial recognition model according to the target training sample set to obtain a target recognition model;

[0008] The image to be detected is input into the target recognition model for target recognition to obtain the detection result of the image to be detected.

[0009] Further, determining a target training sample set corresponding to the image to be detected according to the source information of the image to be detected includes:

[0010] Determining a detection target according to source information of the image to be detected;

[0011] Determine a target training sample set corresponding to the image to be detected from a training database according to the detection target;

[0012] The target training sample set is sorted to obtain a target label corresponding to the target training sample set.

[0013] Furthermore, a training network is built based on the decoupled detection head, including:

[0014] Adding a data enhancement module for feature enhancement at the input end of the network, and adding the decoupled detection head for object detection in the remaining network;

[0015] A spatial pyramid pooling module for image normalization is added after the backbone network to obtain the training network.

[0016] Furthermore, the decoupled detection head includes a convolutional network, a classification detection head, a regression detection head and a confidence detection head;

[0017] Among them, the convolution network is used to perform feature dimensionality reduction, the classification detection head is used to perform target classification, the regression detection head is used to perform position recognition, and the confidence detection head is used to determine the accuracy of the classification detection head and the regression detection head.

[0018] Furthermore, the training network is trained according to the initial data set to obtain an initial recognition model, including:

[0019] Training the training network according to the target detection data in the initial data set to obtain initial parameters corresponding to the training network;

[0020] The initial parameters are updated to the training network to obtain the initial recognition model.

[0021] Further, the initial recognition model is trained according to the target training sample set to obtain a target recognition model, including:

[0022] Dividing the target training sample set into data of a first training cycle and data of a second training cycle according to a preset training cycle;

[0023] The initial recognition model is trained according to the data of the first training cycle and the data of the second training cycle to obtain the target recognition model.

[0024] Further, the initial recognition model is trained according to the data of the first training cycle and the data of the second training cycle to obtain the target recognition model, including:

[0025] Freezing initial parameters of a backbone network in the initial recognition model, and training the initial recognition model using data from the first training cycle to obtain a first training model;

[0026] Unfreeze the initial parameters of the backbone network in the first training model, and use the data of the second training cycle to train the first training model to obtain the target recognition model.

[0027] In a second aspect, an embodiment of the present invention further provides a target recognition device, the device comprising:

[0028] A sample determination module is used to obtain an image to be detected and determine a target training sample set corresponding to the image to be detected according to source information of the image to be detected;

[0029] A network building module, used to build a training network according to the decoupling detection head, and train the training network according to the initial data set to obtain an initial recognition model;

[0030] A model training module, used to train the initial recognition model according to the target training sample set to obtain a target recognition model;

[0031] The image detection module is used to input the image to be detected into the target recognition model for target recognition to obtain the detection result of the image to be detected.

[0032] In a third aspect, an embodiment of the present invention further provides an electronic device, the electronic device comprising:

[0033] one or more processors;

[0034] a storage device for storing one or more programs,

[0035] When the one or more programs are executed by the one or more processors, the one or more processors implement the target identification method.

[0036] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, wherein the program implements the target identification method when executed by a processor.

[0037] In the embodiment of the present invention, an image to be detected is obtained, and a target training sample set corresponding to the image to be detected is determined according to the source information of the image to be detected; a training network is built according to the decoupled detection head, and the training network is trained according to the initial data set to obtain an initial recognition model; the initial recognition model is trained according to the target training sample set to obtain a target recognition model; the image to be detected is input into the target recognition model for target recognition to obtain a detection result of the image to be detected. That is, in the embodiment of the present invention, the convergence speed of the model is improved from the network structure and data source by building a training network and determining target training samples; the training network is initially trained with the initial data set to obtain an initial recognition model, and then different targets are trained according to the target training sample set, and the convergence speed of the model is improved from the training step setting. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 is a flow chart of a target recognition method provided by an embodiment of the present invention;

[0039] Figure 2 is another flowchart of the target recognition method provided by an embodiment of the present invention;

[0040] Figure 2A A schematic diagram of the structure of a training network provided by an embodiment of the present invention;

[0041] Figure 2B A schematic diagram of the structure of a decoupling detection head provided by an embodiment of the present invention;

[0042] Figure 2C A schematic diagram of the training process of the target recognition model provided by an embodiment of the present invention;

[0043] Figure 3 is a schematic diagram of the structure of a target recognition device provided by an embodiment of the present invention;

[0044] Figure 4 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0045] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the present invention, rather than to limit the present invention. It should also be noted that, for ease of description, only parts related to the present invention, rather than all structures, are shown in the accompanying drawings.

[0046] Figure 1The present invention provides a flow chart of a target recognition method provided by an embodiment of the present invention. The method can be performed by a target recognition device provided by an embodiment of the present invention. The device can be implemented in software and / or hardware. In a specific embodiment, the device can be integrated in an electronic device, such as a server. The following embodiments will be described by taking the device integrated in an electronic device as an example. Figure 1 , the method may specifically include the following steps:

[0047] S110, obtaining an image to be detected, and determining a target training sample set corresponding to the image to be detected according to source information of the image to be detected;

[0048] For example, the image to be detected can come from an image acquisition device, and the image acquisition device can be a device with image acquisition function such as a camera or a video recorder installed in the area to be detected, wherein the area to be detected can be an acquisition area for acquiring the image to be detected. When the detection target is an emergency target, it is necessary to acquire images in the area to be detected in real time, and the image to be detected can be the latest image in the area to be detected. The image to be detected is used to detect whether there is a target in the area to be detected. If the image to be detected has a target, emergency processing is performed; when the detection target is a non-emergency target, it can be an image acquired in the area to be detected within a preset time period. If the image to be detected is used to detect the appearance of the detection target in the area to be detected, the operation can be adjusted according to the appearance time and frequency of the detection target. The source information of the image to be detected can be the information of the enterprise or unit that obtains the image to be detected. According to the different working attributes of different enterprises or units, the detection of the target in the same image may be different. The target training sample set corresponding to the image to be detected can be a target training sample set corresponding to different detection targets, and the target mark position and target positive and negative sample marks in the target training sample set corresponding to different detection targets are different.

[0049] In the specific implementation, after the image acquisition device obtains the image to be detected from the area to be detected, the detection target corresponding to the image to be detected can be determined according to the source information of the image to be detected, and then the target training sample set corresponding to the image to be detected can be determined according to the detection target corresponding to the image to be detected; each image in the target training sample set corresponding to the image to be detected has a marked position of the detection target, and the positive and negative samples can be distinguished according to the positive and negative sample classification method, and the positive and negative labels can be marked, so as to train the model according to the target training sample set to obtain the target recognition model.

[0050] S120, building a training network according to the decoupling detection head, and training the training network according to the initial data set to obtain an initial recognition model;

[0051] For example, the decoupled detection head can be connected to different detection branches and detection heads after the dimensionality reduction operation, and is used to detect the target category, location and confidence in the image. While improving the detection effect, the target detection speed can be improved to avoid the increase in the amount of calculation. The training network is a neural network for training target recognition models built according to demand. The training network includes a backbone network and a residual network, wherein the detection head is decoupled in the residual network. The initial data set can be understood as a standard target training data set, such as: ImageNet data set and COCO data set. The initial recognition model can be a model with certain recognition capabilities trained according to the category labels of the images in the standard target training data set.

[0052] In the specific implementation, the training network includes a backbone network and a residual network, wherein the backbone network is used to extract target features in the image, and the residual network is used to perform feature recognition and position information prediction on target features in the image. The residual network includes a decoupled detection head, which is used to detect the target category, location, and confidence in the image, which can improve the detection effect and the target detection speed. After the training network is built, the training network is trained with the initial data set until the model converges to obtain the initial recognition model, which can recognize more than 80 or more types of targets in the standard type as the basic model of the target recognition model. Based on the basic model, targeted training is performed with the target training samples to obtain the required target recognition model.

[0053] S130, training the initial recognition model according to the target training sample set to obtain a target recognition model;

[0054] In a specific implementation, the images in the target training sample set are input into the initial recognition model for target recognition, and the output of the initial recognition model can be the predicted position information of the detected target in each image and the confidence corresponding to the predicted position information. Among them, the confidence corresponding to the predicted position information can be the confidence determined by detecting the overlap between the predicted position information and the actual position information of the image in the target training sample set. A learning correction function can also be set in the initial recognition model to determine the degree of model training using the confidence corresponding to the predicted position information.

[0055] S140, inputting the image to be detected into the target recognition model for target recognition, and obtaining a detection result of the image to be detected.

[0056] In a specific implementation, the detection result of the image to be detected can be the output result of the target recognition model obtained by inputting each image to be detected into the target recognition model for target recognition, and the result includes the detection position information and confidence of the detected target in the image to be detected. In addition, the target recognition model can be a deep neural network model, which is obtained based on a pre-trained initial recognition model.

[0057] In the embodiment of the present invention, an image to be detected is obtained, and a target training sample set corresponding to the image to be detected is determined according to the source information of the image to be detected; a training network is built according to the decoupled detection head, and the training network is trained according to the initial data set to obtain an initial recognition model; the initial recognition model is trained according to the target training sample set to obtain a target recognition model; the image to be detected is input into the target recognition model for target recognition to obtain a detection result of the image to be detected. That is, in the embodiment of the present invention, the convergence speed of the model is improved from the network structure and data source by building a training network and determining target training samples; the training network is initially trained with the initial data set to obtain an initial recognition model, and then different targets are trained according to the target training sample set, and the convergence speed of the model is improved from the training step setting.

[0058] The target recognition method provided by the embodiment of the present invention is further described below. Figure 2 As shown, the method may specifically include the following steps:

[0059] S210, obtaining an image to be detected, and determining a target training sample set corresponding to the image to be detected according to source information of the image to be detected;

[0060] Furthermore, determining a target training sample set corresponding to the image to be detected according to the source information of the image to be detected includes:

[0061] Determine the detection target according to the source information of the image to be detected;

[0062] Determine a target training sample set corresponding to the image to be detected from a training database according to the detection target;

[0063] The target training sample set is sorted to obtain the target label corresponding to the target training sample set.

[0064] For example, the detection target can be an object that needs to be monitored in the image to be detected, which can be determined based on the source information of the image to be detected, such as: an air gun flame in an industrial environment, and the detection target air gun flame is used to identify the use of the air gun flame. Among them, the source information of the image to be detected can be the detection target corresponding to the manufacturing object of the enterprise in the manufacturing industry, such as: the use of the target in the manufacturing process or whether there are dangerous items in a special area. The training database is a database for storing training samples corresponding to different detection targets. The target label corresponding to the target training sample set can be the label of the positive and negative samples corresponding to the target mark in each image in the target training sample set.

[0065] In a specific implementation, when acquiring an image to be detected, the source information of the image to be detected is also acquired. The detection target corresponding to the image to be detected is determined according to the source information of the image to be detected. According to the detection target, the target training sample set corresponding to the image to be detected is matched from the training database. The target training sample set is sorted by using a positive and negative sample sorting algorithm to obtain a target label corresponding to the target training sample set. Among them, the sample sorting algorithm can be a simOTA algorithm to screen the anchor frame of the target training sample set, and divide the samples in the target training sample set into positive and negative samples, and train the initial recognition model according to the positive samples to obtain the target recognition sample. In addition, the target training sample set can include a training sample set and a test sample set, and the corresponding sample set ratio can be set according to actual needs and experimental data, such as dividing it into a training set and a test set in a ratio of 7:3 to prevent overfitting during the training process.

[0066] S220, building a training network according to the decoupling detection head, and training the training network according to the initial data set to obtain an initial recognition model;

[0067] Furthermore, a training network is built based on the decoupled detection head, including:

[0068] Adding a data enhancement module for feature enhancement at the input end of the network, and adding the decoupled detection head for object detection in the remaining network;

[0069] A spatial pyramid pooling module for image normalization is added after the backbone network to obtain the training network.

[0070] In the specific implementation, the training network includes a backbone network and a residual network. The backbone network is used to extract image features, and the residual network is used to identify image features to determine the predicted position information of the detection target in the image to be detected and the confidence of the predicted position information. A spatial pyramid pooling module for image normalization is added after the backbone network, so that the network does not need to preprocess the image size, and the size of the network output is the same. The same image of different sizes can be used as input to obtain pooling features of the same length. Among them, the spatial pyramid pooling module includes a maximum pooling layer, a connection function for connecting multiple matrices, a convolutional layer, a normalization layer, and a residual component.

[0071] In an embodiment of the present invention, a data enhancement module for feature enhancement is added to the input end of the network, wherein the data enhancement module can be two data enhancement methods, namely, the Mosaic algorithm and the Mixup algorithm. Among them, Mosaic splices images by random scaling, random cropping, and random arrangement to improve the detection effect of small targets; wherein Mixup presents the area between samples as linear, which can reduce the memory of wrong labels and increase robustness. The image data input to the training network is enhanced by the data enhancement module to prevent the occurrence of overfitting in the training process. In the backbone network, anchor-free and simOTA algorithms can also be added to screen the ambiguous prediction boxes that appear in the target recognition process, make accurate predictions, and reduce the amount of calculation in the prediction process. In addition, the backbone network can be a Yolox-Darknet53 structure.

[0072] Figure 2A A schematic diagram of the structure of a training network provided in an embodiment of the present invention, such as Figure 2A As shown in the figure, a data enhancement module is added to the input end of the network to enhance the image features input into the network, and then the image is input into the backbone network for feature extraction to obtain the feature map of the image. The feature map is input into the spatial pyramid pooling module for normalization to obtain pooled features of the same length, and then the pooled features are input into the remaining network for image feature detection to predict the location information and confidence of the detection target.

[0073] Furthermore, the decoupled detection head includes a convolutional network, a classification detection head, a regression detection head, and a confidence detection head;

[0074] Among them, the convolutional network is used for feature dimensionality reduction, the classification detection head is used for target classification, the regression detection head is used for location recognition, and the confidence detection head is used to determine the accuracy of the classification detection head and the regression detection head.

[0075] In the specific implementation, the decoupled detection head in the remaining network of the training network can use the convolutional network to perform feature dimensionality reduction and expand the field of view on the feature map corresponding to the image to be detected. For example: a 1*1 convolutional layer can be used to reduce the dimensionality of feature maps with different numbers of channels, and then multiple 3*3 convolutional layers can be used to expand the field of view. Two branches, one is used to form a classification detection head, and the other is used to form a regression detection head. At the same time, an IOU branch is added to the output of the regression detection head to form a confidence detection head, which is used to determine the accuracy of the classification detection head and the regression detection head.

[0076] Figure 2B A schematic diagram of the structure of a decoupling detection head provided by an embodiment of the present invention is shown in FIG. Figure 2BAs shown, by setting multiple convolution layers of different dimensions in the convolutional network, the image features corresponding to the image to be detected can be reduced in dimension and the field of view can be expanded, providing a clear feature image for the subsequent classification detection head, regression detection head and confidence detection head.

[0077] Furthermore, the training network is trained according to the initial data set to obtain an initial recognition model, including:

[0078] The training network is trained according to the target detection data in the initial data set to obtain the initial parameters corresponding to the training network;

[0079] Update the initial parameters to the training network to obtain the initial recognition model.

[0080] In the specific implementation, an initial data set for image recognition is provided on the Internet, wherein the images in the initial data set are divided into training, verification and test sets. The initial data set can be a public target test set with annotated types, i.e., the COCO data set. After the training network is built, the training network is trained with the initial data set until the model converges to obtain an initial recognition model, which can recognize more than 80 or more types of targets in the standard type, as the basic model of the target recognition model. On the basis of the basic model, targeted training is performed through the target training samples to obtain the required target recognition model.

[0081] S230, dividing the target training sample set into data of a first training cycle and data of a second training cycle according to a preset training cycle;

[0082] In a specific implementation, the preset training cycle can be the number of training sub-cycles preset according to actual needs and experimental data. In the preset training cycle, each image can be used as a training sample, and a training sample is input once for training as a training sub-cycle. It can also be divided into preset training sub-cycles according to the target training sample set, and each image is used as a training sample, and a training sample is input once for training as a training cycle. The data of the first training cycle is used for the data in the target training sample set for parameter training of the remaining network when the backbone network is frozen. The data of the second training cycle is used for the data in the target training sample set for parameter training of the training network when the backbone network is unfrozen.

[0083] S240: Train the initial recognition model according to the data of the first training cycle and the data of the second training cycle to obtain a target recognition model.

[0084] In a specific implementation, the data of the first training cycle is input into the initial recognition model for model training of the first training cycle, and the parameters in the initial recognition model are updated. Then, the data of the second training cycle is input into the initial recognition model after the updated parameters for model training of the second training cycle, and the parameters of the initial recognition model are updated for the second time to obtain the target recognition model.

[0085] Further, the initial recognition model is trained according to the data of the first training cycle and the data of the second training cycle to obtain a target recognition model, including:

[0086] Freeze the initial parameters of the backbone network in the initial recognition model, and train the initial recognition model using the data of the first training cycle to obtain a first training model;

[0087] The initial parameters of the backbone network in the first training model are unfrozen, and the first training model is trained using the data of the second training cycle to obtain a target recognition model.

[0088] In the specific implementation, the initial parameters of the backbone network in the initial recognition model are frozen, the data of the first training cycle is input into the initial recognition model for model training of the first training cycle, the parameters in the initial recognition model are updated, and the first training model is obtained; the initial parameters of the backbone network in the first training model are unfrozen, and the data of the second training cycle is input into the first training model for model training of the second training cycle, and the parameters of the initial recognition model are updated for the second time to obtain the target recognition model. For example: the preset training cycle is set to 300 training sub-cycles, the data of the first training cycle is 50 training sub-cycles, the initial parameters of the backbone network are frozen, the data of the first training cycle is trained, and the first training model is obtained; the initial training parameters of the backbone network are unfrozen, and the remaining 250 sub-cycle data is input into the first training model as the second training cycle data for model training of the second training cycle to obtain the target recognition model. Among them, when the target is detected as an airgun flame, an airgun flame recognition model is obtained.

[0089] Figure 2C A schematic diagram of the training process of the target recognition model provided by the embodiment of the present invention, such as Figure 2C As shown, the image to be detected is obtained, and the target training sample set corresponding to the image to be detected is determined according to the source information of the image to be detected, and the initial data set is obtained through the network; a training network is built, and the initial data set is input into the built training network for initial training to obtain an initial recognition model. Based on the initial recognition model, the target training sample set is input into the initial recognition model for training to obtain a target recognition model corresponding to the image to be detected.

[0090] S250, inputting the image to be detected into the target recognition model for target recognition, and obtaining a detection result of the image to be detected.

[0091] In the embodiment of the present invention, an image to be detected is obtained, and a target training sample set corresponding to the image to be detected is determined according to the source information of the image to be detected; a training network is built according to the decoupled detection head, and the training network is trained according to the initial data set to obtain an initial recognition model; the initial recognition model is trained according to the target training sample set to obtain a target recognition model; the image to be detected is input into the target recognition model for target recognition to obtain a detection result of the image to be detected. That is, in the embodiment of the present invention, the convergence speed of the model is improved from the network structure and data source by building a training network and determining target training samples; the training network is initially trained with the initial data set to obtain an initial recognition model, and then different targets are trained according to the target training sample set, and the convergence speed of the model is improved from the training step setting.

[0092] Figure 3 is a schematic diagram of the structure of a target recognition device provided by an embodiment of the present invention, such as Figure 3 As shown, the target recognition device comprises:

[0093] The sample determination module 310 is used to obtain an image to be detected and determine a target training sample set corresponding to the image to be detected according to source information of the image to be detected;

[0094] A network building module 320 is used to build a training network according to the decoupled detection head, and train the training network according to the initial data set to obtain an initial recognition model;

[0095] A model training module 330 is used to train the initial recognition model according to the target training sample set to obtain a target recognition model;

[0096] The image detection module 340 is used to input the image to be detected into the target recognition model to perform target recognition and obtain the detection result of the image to be detected.

[0097] In one embodiment, the sample determination module 310 determines the target training sample set corresponding to the image to be detected according to the source information of the image to be detected, including:

[0098] Determining a detection target according to source information of the image to be detected;

[0099] Determine a target training sample set corresponding to the image to be detected from a training database according to the detection target;

[0100] The target training sample set is sorted to obtain a target label corresponding to the target training sample set.

[0101] In one embodiment, the network building module 320 builds a training network according to the decoupled detection head, including:

[0102] Adding a data enhancement module for feature enhancement at the input end of the network, and adding the decoupled detection head for object detection in the remaining network;

[0103] A spatial pyramid pooling module for image normalization is added after the backbone network to obtain the training network.

[0104] In one embodiment, the decoupled detection head includes a convolutional network, a classification detection head, a regression detection head and a confidence detection head;

[0105] Among them, the convolution network is used to perform feature dimensionality reduction, the classification detection head is used to perform target classification, the regression detection head is used to perform position recognition, and the confidence detection head is used to determine the accuracy of the classification detection head and the regression detection head.

[0106] In one embodiment, the network building module 320 trains the training network according to the initial data set to obtain an initial recognition model, including:

[0107] Training the training network according to the target detection data in the initial data set to obtain initial parameters corresponding to the training network;

[0108] The initial parameters are updated to the training network to obtain the initial recognition model.

[0109] In one embodiment, the model training module 330 trains the initial recognition model according to the target training sample set to obtain a target recognition model, including:

[0110] Dividing the target training sample set into data of a first training cycle and data of a second training cycle according to a preset training cycle;

[0111] The initial recognition model is trained according to the data of the first training cycle and the data of the second training cycle to obtain the target recognition model.

[0112] In the embodiment, the model training module 330 trains the initial recognition model according to the data of the first training cycle and the data of the second training cycle to obtain the target recognition model, including:

[0113] Freezing initial parameters of a backbone network in the initial recognition model, and training the initial recognition model using data from the first training cycle to obtain a first training model;

[0114] Unfreeze the initial parameters of the backbone network in the first training model, and use the data of the second training cycle to train the first training model to obtain the target recognition model.

[0115] The device of the embodiment of the present invention obtains the image to be detected, and determines the target training sample set corresponding to the image to be detected according to the source information of the image to be detected; builds a training network according to the decoupled detection head, and trains the training network according to the initial data set to obtain an initial recognition model; trains the initial recognition model according to the target training sample set to obtain a target recognition model; inputs the image to be detected into the target recognition model for target recognition to obtain the detection result of the image to be detected. That is, the embodiment of the present invention improves the convergence speed of the model from the network structure and data source by building a training network and determining the target training samples; performs initial training on the training network through the initial data set to obtain an initial recognition model, and then trains different targets according to the target training sample set, and improves the convergence speed of the model from the training step setting.

[0116] Figure 4 A schematic diagram of the structure of an electronic device provided in Example 4 of the present invention. Figure 4 A block diagram of an exemplary electronic device 12 suitable for use in implementing embodiments of the present invention is shown. Figure 4 The electronic device 12 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.

[0117] like Figure 4 As shown, the electronic device 12 is in the form of a general purpose computing device. The components of the electronic device 12 may include, but are not limited to: one or more processors or processing units 16, a system memory 28, and a bus 18 that connects various system components (including the system memory 28 and the processing unit 16).

[0118] Bus 18 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor or a local bus using any of a variety of bus architectures. By way of example, these architectures include, but are not limited to, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MAC) bus, an Enhanced ISA bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus.

[0119] The electronic device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the electronic device 12, including volatile and non-volatile media, removable and non-removable media.

[0120] The system memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The electronic device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 34 may be used to read and write non-removable, non-volatile magnetic media ( Figure 4 not shown, usually called a "hard drive"). Although Figure 4 Not shown in the figure, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk"), and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, a DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to the bus 18 via one or more data medium interfaces. The memory 28 may include at least one program product having a set (e.g., at least one) of program modules that are configured to perform the functions of the various embodiments of the present invention.

[0121] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in the memory 28, such program modules 42 including, but not limited to, an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment. The program modules 42 generally perform the functions and / or methods of the embodiments described herein.

[0122] The electronic device 12 may also communicate with one or more external devices 14 (e.g., keyboards, pointing devices, displays 24, etc.), may communicate with one or more devices that enable a user to interact with the electronic device 12, and / or may communicate with any device that enables the electronic device 12 to communicate with one or more other computing devices (e.g., network cards, modems, etc.). Such communication may be performed via an input / output (I / O) interface 22. Furthermore, the electronic device 12 may also communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via a network adapter 20. As shown, the network adapter 20 communicates with other modules of the electronic device 12 via a bus 18. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0123] The processing unit 16 executes various functional applications and data processing by running the program stored in the system memory 28, for example, implementing the target recognition method provided by the embodiment of the present invention, which includes:

[0124] Acquire an image to be detected, and determine a target training sample set corresponding to the image to be detected according to source information of the image to be detected;

[0125] Building a training network according to the decoupled detection head, and training the training network according to the initial data set to obtain an initial recognition model;

[0126] Training the initial recognition model according to the target training sample set to obtain a target recognition model;

[0127] The image to be detected is input into the target recognition model for target recognition to obtain the detection result of the image to be detected.

[0128] The embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the target recognition method is implemented. The method includes:

[0129] Acquire an image to be detected, and determine a target training sample set corresponding to the image to be detected according to source information of the image to be detected;

[0130] Building a training network according to the decoupled detection head, and training the training network according to the initial data set to obtain an initial recognition model;

[0131] Training the initial recognition model according to the target training sample set to obtain a target recognition model;

[0132] The image to be detected is input into the target recognition model for target recognition to obtain the detection result of the image to be detected.

[0133] The computer storage medium of the embodiment of the present invention can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, a system, device or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples (non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, a computer-readable storage medium can be any tangible medium containing or storing a program, which can be used by an instruction execution system, a device or a device or used in combination with it.

[0134] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, which carry computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0135] The program code embodied on the computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0136] Computer program code for performing the operation of the present invention may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0137] Note that the above are only preferred embodiments of the present invention and the technical principles used. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and that various obvious changes, readjustments and substitutions can be made by those skilled in the art without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments, and may include more other equivalent embodiments without departing from the concept of the present invention, and the scope of the present invention is determined by the scope of the appended claims.

Claims

1. A target recognition method, characterized in that: include: Acquire an image to be detected, and determine a target training sample set corresponding to the image to be detected according to source information of the image to be detected; Building a training network according to the decoupled detection head, and training the training network according to the initial data set to obtain an initial recognition model; Training the initial recognition model according to the target training sample set to obtain a target recognition model; Inputting the image to be detected into the target recognition model for target recognition to obtain a detection result of the image to be detected; Wherein, determining the target training sample set corresponding to the image to be detected according to the source information of the image to be detected includes: Determining a detection target according to source information of the image to be detected; Determine a target training sample set corresponding to the image to be detected from a training database according to the detection target; Sorting the target training sample set to obtain target labels corresponding to the target training sample set; The step of building a training network based on the decoupled detection head includes: Adding a data enhancement module for feature enhancement at the input end of the network, and adding the decoupled detection head for object detection in the remaining network; A spatial pyramid pooling module for image normalization is added after the backbone network to obtain the training network; Wherein, the decoupled detection head includes a convolutional network, a classification detection head, a regression detection head and a confidence detection head; Among them, the convolution network is used to perform feature dimensionality reduction, the classification detection head is used to perform target classification, the regression detection head is used to perform position recognition, and the confidence detection head is used to determine the accuracy of the classification detection head and the regression detection head.

2. The method according to claim 1, characterized in that: The training network is trained according to the initial data set to obtain an initial recognition model, including: Training the training network according to the target detection data in the initial data set to obtain initial parameters corresponding to the training network; The initial parameters are updated to the training network to obtain the initial recognition model.

3. The method according to claim 1, characterized in that The initial recognition model is trained according to the target training sample set to obtain a target recognition model, including: Dividing the target training sample set into data of a first training cycle and data of a second training cycle according to a preset training cycle; The initial recognition model is trained according to the data of the first training cycle and the data of the second training cycle to obtain the target recognition model.

4. The method according to claim 3, characterized in that The initial recognition model is trained according to the data of the first training cycle and the data of the second training cycle to obtain the target recognition model, including: Freezing initial parameters of a backbone network in the initial recognition model, and training the initial recognition model using data from the first training cycle to obtain a first training model; Unfreeze the initial parameters of the backbone network in the first training model, and use the data of the second training cycle to train the first training model to obtain the target recognition model.

5. A target recognition device, characterized in that: include: A sample determination module is used to obtain an image to be detected and determine a target training sample set corresponding to the image to be detected according to source information of the image to be detected; A network building module, used to build a training network according to the decoupling detection head, and train the training network according to the initial data set to obtain an initial recognition model; A model training module, used to train the initial recognition model according to the target training sample set to obtain a target recognition model; An image detection module is used to input the image to be detected into the target recognition model for target recognition, and obtain a detection result of the image to be detected; The sample determination module determines the target training sample set corresponding to the image to be detected according to the source information of the image to be detected, including: Determining a detection target according to source information of the image to be detected; Determine a target training sample set corresponding to the image to be detected from a training database according to the detection target; Sorting the target training sample set to obtain target labels corresponding to the target training sample set; The network building module builds a training network according to the decoupling detection head, including: Adding a data enhancement module for feature enhancement at the input end of the network, and adding the decoupled detection head for object detection in the remaining network; A spatial pyramid pooling module for image normalization is added after the backbone network to obtain the training network; Wherein, the decoupled detection head includes a convolutional network, a classification detection head, a regression detection head and a confidence detection head; Among them, the convolution network is used to perform feature dimensionality reduction, the classification detection head is used to perform target classification, the regression detection head is used to perform position recognition, and the confidence detection head is used to determine the accuracy of the classification detection head and the regression detection head.

6. An electronic device, characterized in that: The electronic device comprises: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the target recognition method as described in any one of claims 1 to 4.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the target recognition method as claimed in any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Image detection method and device for identifying target object, electronic equipment and storage medium

    CN110826476A

  • Image multi-task multi-label identification method, system and device, and medium

    CN112508078A