Directed target detection network based on YOLOV6, training method thereof and directed target detection method
Patent Information
- Application Number
- CN202280100548.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-05
- Publication Date
- 2025-05-06
AI Technical Summary
Existing target detection networks have problems with unfavorable feature extraction and difficulty in identifying small targets when dealing with directed targets. Especially when there are a large number of small targets in the image, it is difficult for existing methods to detect targets quickly and accurately.
A directed target detection network based on YOLOV6 is used. By obtaining training images and label information, the setting parameters are obtained through training, and the combined architecture of the backbone layer, neck layer and detection head layer is used for feature extraction, cross-scale fusion and prediction, including RepVGG, The use of FPN-PAN and decoupling heads to achieve angle prediction and target detection.
It improves the accuracy and speed of target detection, and can effectively process the angle information of directed targets, improving detection performance in fields such as biomedicine and aerial images.
Smart Images

Figure CN119948539A_ABST
Abstract
Description
Directed target detection network based on YOLOV6 and its training method, as well as directed target detection method Technical Field
[0001] The present disclosure generally relates to the field of image vision, and in particular, to a YOLOv6-based directed target detection network and a training method thereof, as well as a directed target detection method using the YOLOv6-based directed target detection network, which can improve the accuracy and detection speed of target detection. Background Art
[0002] Object detection is a core problem in image vision. Because objects can be oriented in any direction, existing detection networks that use horizontal object detection boxes to locate objects can hinder feature extraction and subsequent dark line (also known as "track line") detection. Furthermore, when images contain numerous small objects, quickly and accurately orienting and identifying them becomes a challenge.
[0003] In this context, directed object detection has become an important research direction in the field of image vision and is widely used in fields such as biomedical detection, aerial imagery object detection, scene text detection, and densely packed object detection. Current directed object detection methods are mainly divided into two categories: single-stage directed object detection and two-stage directed object detection. Single-stage directed object detection uses direct regression analysis to detect objects. Compared to two-stage directed object detection based on candidate generation, it has the advantages of fast prediction speed and ease of training, but its accuracy is often limited.
[0004] Summary of the Invention
[0005] The present disclosure aims to solve at least part or all of the above problems.
[0006] One aspect of the present disclosure provides a method for training a directed target detection network based on YOLOV6, the method comprising: obtaining label information of a plurality of training images and labels included in the plurality of training images; and inputting the plurality of training images and the label information into the directed target detection network to obtain setting parameters of the directed target detection network through training, wherein the directed target detection network may comprise: a backbone layer, configured to perform feature extraction on the input image to obtain a plurality of feature maps of different scales; a neck layer, configured to receive the plurality of feature maps obtained from the backbone layer, and perform cross-scale fusion of the plurality of feature maps to obtain a plurality of fused feature maps; and a detection head layer, configured to perform prediction on the plurality of fused images to obtain prediction results about the images, wherein the prediction results comprise: classification information, target frame coordinate prediction information, target prediction information, and angle prediction information.
[0007] In one example, the Backbone layer adopts the RepVGG architecture.
[0008] In another example, the Neck layer adopts FPN-PAN architecture.
[0009] In another example, the Head layer adopts a decoupling head so that category analysis is performed via a classification branch to obtain the classification information, and regression analysis is performed via a regression branch to respectively obtain the target frame coordinate prediction information, the target prediction information, and the angle prediction information.
[0010] In another example, inputting the multiple training images and the label information into the directed target detection network to obtain setting parameters of the directed target detection network through training includes: in each training, obtaining corresponding training images and corresponding label information for the current training; using the directed target detection network to predict targets in the corresponding training images to obtain corresponding prediction results; calculating the loss function of the directed target detection network based on the corresponding label information and the corresponding prediction results; and adjusting the setting parameters of the directed target detection network based on the calculation result of the loss function until the calculation result of the loss function converges.
[0011] In another example, the loss function is a cross entropy loss function.
[0012] In another example, in the cross entropy loss function, the weight of the loss associated with angle prediction is 0.1.
[0013] In another example, the acquired label information in the first data format is preprocessed to process the label information into data in the form of [x, y, w, h, θ, classID], and the processed data is input into the directed target detection network, wherein (x, y) represents the position of the center of the target detection box, w represents the width of the long side of the target detection box, h represents the height of the short side of the target detection box, θ represents the angle of rotation of the long side of the target detection box relative to the horizontal axis, and classID represents the category of the target detection box.
[0014] In another example, the label information may have a first data form, and the first data form may be [x1, y1, x2, y2, x3, y3, x4, y4, classID], where (x1, y1), (x2, y2), (x3, y3) and (x4, y4) respectively define the positions of the four vertices of the target detection box.
[0015] In another example, the method may further include: post-processing the obtained prediction result in the form of [x, y, w, h, θ, conf, classID] to process the prediction result into data in a second data form, and presenting the data in the second data form to the user, wherein (x, y) represents the position of the center of the target detection frame, w represents the width of the long side of the target detection frame, h represents the height of the short side of the target detection frame, θ represents the angle of rotation of the long side of the target detection frame relative to the horizontal axis, conf represents the confidence that the target detection frame has the corresponding target, and classID represents the category of the target detection frame.
[0016] In another example, the second data form may be [x1, y1, x2, y2, x3, y3, x4, y4, conf, classID], where (x1, y1), (x2, y2), (x3, y3) and (x4, y4) respectively define the positions of the four vertices of the target detection box.
[0017] According to another aspect of the present disclosure, a directed target detection method is provided, which includes: acquiring an image to be tested; and inputting the image to be tested into a directed target detection network based on YOLOV6 to obtain a target detection result, wherein the directed target detection network may include: a backbone layer, configured to perform feature extraction on the image to be tested to obtain multiple feature maps of different scales; a neck layer, configured to receive the multiple feature maps obtained from the backbone layer, and perform cross-scale fusion of the multiple feature maps to obtain multiple fused feature maps; and a detection head layer, configured to predict the multiple fused images to obtain a target detection result for the image to be tested, wherein the target detection result includes: classification information, target frame coordinate prediction information, target prediction information and angle prediction information.
[0018] In one example, the Backbone layer adopts the RepVGG architecture, and the Neck layer adopts the FPN-PAN architecture.
[0019] In another example, the Head layer adopts a decoupling head so that category analysis is performed via a classification branch to obtain the classification information, and regression analysis is performed via a regression branch to respectively obtain the target frame coordinate prediction information, the target prediction information, and the angle prediction information.
[0020] In another example, the method further includes: post-processing the target detection result obtained in the form of [x, y, w, h, θ, conf, classID] to process the target detection result into data in a second data form, and presenting the data in the second data form to the user, wherein (x, y) represents the position of the center of the target detection frame, w represents the width of the long side of the target detection frame, h represents the height of the short side of the target detection frame, θ represents the angle of rotation of the long side of the target detection frame relative to the horizontal axis, conf represents the confidence that the target detection frame has the corresponding target, and classID represents the category of the target detection frame.
[0021] In another example, the second data format is [x1, y1, x2, y2, x3, y3, x4, y4, conf, classID], where (x1, y1), (x2, y2), (x3, y3) and (x4, y4) respectively define the positions of the four vertices of the target detection box.
[0022] In another example, the directed object detection network is trained using the training method according to an exemplary embodiment of the present disclosure.
[0023] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method for training a directed target detection network based on YOLOV6 according to an exemplary embodiment of the present disclosure or the directed target detection method according to an exemplary embodiment of the present disclosure.
[0024] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute the method for training a directed target detection network based on YOLOV6 according to an exemplary embodiment of the present disclosure or the directed target detection method according to an exemplary embodiment of the present disclosure.
[0025] According to another aspect of the present disclosure, a computer program product is also provided, including a computer program, which, when executed by a processor, implements the method for training a directed target detection network based on YOLOV6 according to an exemplary embodiment of the present disclosure or the directed target detection method according to an exemplary embodiment of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] FIG1 shows the architecture of a directed object detection network based on YOLOV6 according to an exemplary embodiment of the present disclosure;
[0027] 2A and 2B show a flowchart of a method for training a YOLOv6-based directed object detection network according to an exemplary embodiment of the present disclosure;
[0028] FIG3 is a flow chart of a directed target detection method according to an exemplary embodiment of the present disclosure;
[0029] 4 and 5 show the test results on mouse brain data by using the directed target detection method according to an exemplary embodiment of the present disclosure; and
[0030] FIG6 shows a schematic block diagram of an example electronic device that can be used to implement embodiments of the present disclosure. DETAILED DESCRIPTION
[0031] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the detailed description below, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.
[0032] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise," "include," etc. used herein indicate the presence of the features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0033] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0034] When expressions such as “at least one of A, B, and C, etc. is used, it should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, “a system having at least one of A, B, and C” should include but is not limited to systems having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, and C, etc.). When expressions such as “at least one of A, B, or C, etc. is used, it should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, “a system having at least one of A, B, or C” should include but is not limited to systems having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, and C, etc.).
[0035] In the drawings, the same or similar reference numerals are used to represent the same or similar structures.
[0036] FIG1 shows the architecture of a YOLOV6-based directed object detection network according to an exemplary embodiment of the present disclosure.
[0037] As shown in FIG1 , the YOLOV6-based directed target detection network may include a backbone layer 110 , a neck layer 120 , and a detection head layer 130 .
[0038] The backbone layer 110 can be configured to perform feature extraction on the input image to obtain multiple feature maps of different scales. For example, the backbone layer 110 can adopt the RepVGG architecture. The main advantage of the RepVGG architecture is that it decouples the training architecture and the inference architecture. That is, a multi-branch architecture can be used during training, and the trained multi-branch architecture is converted into an equivalent single-path model (or VGG model) during inference, thereby achieving the advantages of fast inference speed, memory saving, and high flexibility.
[0039] In the RepVGG architecture, a 1x1 convolution branch and an identity mapping branch are added in parallel to each 3x3 convolution layer. This structure can be called a RepVGG Block or Rep operator. In the Backbone layer 110 shown in Figure 1, the stem module can be implemented as a RepVGG Block, and the ERBlock module can be implemented as a combination of a RepVGG Block and a Rep Block, where the Rep Block can be further implemented as multiple RepVGG Blocks. Such a RepVGG architecture can be used as the Backbone layer for training. After training, the RepVGG architecture can be equivalently transformed so that the three branches of the RepVGG Block (including the 3x3 convolution branch, the 1x1 convolution branch, and the identity mapping branch) are merged into one branch, thereby obtaining a simple single-branch architecture for use in inference.
[0040] The Neck layer 120 may be configured to receive the multiple feature maps obtained from the Backbone layer 110 and perform cross-scale fusion on the multiple feature maps to obtain multiple fused feature maps. The Neck layer 120 may adopt an FPN-PAN architecture.
[0041] In the Neck layer 120 shown, upsampling fusion is performed first, followed by downsampling fusion. This top-down approach conveys strong semantic features, while the bottom-up approach conveys strong positioning features. Unlike the traditional FPN-PAN architecture, the CSP Block is replaced by the Rep Block, making inference more efficient and achieving a better balance between accuracy and speed.
[0042] The head layer 130 can be configured to predict the multiple fused images received from the neck layer 120 to obtain a prediction result for the image, wherein the prediction result includes: classification information Cls_pred for indicating the result of the category analysis, target frame coordinate prediction information Reg_pred for determining the position of the target detection frame, target prediction information Obj_pred for determining whether the corresponding target exists in the target detection frame, and angle prediction information Ang_pred for determining the angle of the target detection frame, as shown in Figure 1. Specifically, the head layer 130 adopts a decoupled head, so that the classification branch performs category analysis to obtain the classification information, and the regression branch performs regression analysis to obtain the target frame coordinate prediction information, the target prediction information, and the angle prediction information, respectively, wherein the output dimension of the angle prediction is 180 to represent 180 angles. For the angle prediction branch, for example, by introducing the circular smooth label (CSL) angle classification method, the target angle can be accurately predicted to increase the error tolerance between adjacent angles. By decoupling the classification branch and the regression branch, not only can the convergence speed be improved, but the complexity of the detection head can also be reduced. When the predictions of different fused feature maps are completed, the final prediction results are used as the output of the directed object detection network.
[0043] It can be seen that by adopting the YOLOV6-based directed target detection network as shown in Figure 1, not only the accuracy and prediction speed of target detection can be improved, but also the angle information of the target can be obtained.
[0044] 2A and 2B show a flow chart of a method for training a directed target detection network based on YOLOV6 according to an exemplary embodiment of the present disclosure. As previously mentioned, the directed target detection network based on YOLOV6 can structurally have the architecture shown in FIG1 . However, it should be understood by those skilled in the art that the directed target detection network based on YOLOV6 is not limited to the architecture shown in FIG1 . For example, the number of layers of the feature extraction graph of the directed target detection network and the size of the convolution kernel used are not limited to the form shown in the figure, but various changes, modifications and variations can be made within the scope of protection of this application.
[0045] The training method according to an exemplary embodiment of the present disclosure is described by taking the YOLOV6-based directed target detection network shown in reference Figure 1 as an example, wherein the directed target detection network may include a Backbone layer based on the RepVGG architecture, a Neck layer using the FPN-PAN architecture, and a Head layer using a decoupling head.
[0046] As shown in FIG2A , the training method 200 of the directed object detection network based on YOLOV6 can generally include steps S210 and S220, wherein in step S210, a plurality of training images and label information of labels included in the plurality of training images are obtained; and in step S220, the plurality of training images and label information obtained are input into the directed object detection network to be trained to obtain setting parameters of the directed object detection network through training. The training images can be input in the form of a 4-dimensional matrix, that is, batch size (batch_size)×channel (channel)×image height (Height)×image width (Width).
[0047] In one embodiment, the training method may further include: preprocessing the acquired label information in the first data format to process the label information into data in the form of [x, y, w, h, θ, classID], thereby inputting the directed object detection network, wherein (x, y) represents the position of the center of the target detection box, w represents the width of the long side of the target detection box, h represents the height of the short side of the target detection box, θ represents the angle of rotation of the long side of the target detection box relative to the horizontal axis, and classID represents the category of the target detection box. In one example, the label information is input into the directed object detection network in the form of a two-dimensional matrix, wherein each row of the two-dimensional matrix represents a piece of label information. Therefore, a two-dimensional matrix of size 187 × the total number of labels can be input, wherein the number of rows of the two-dimensional matrix represents the total number of labels, and the 187 columns are used to store the batch batch_id of the corresponding label, the category class_id, the horizontal and vertical coordinates x and y of the center of the target detection box, the width w and height h of the target detection box, the rotation angle θ of the long side of the target detection box, and the circular smoothing label csl_label related to the rotation angle. In addition, the first data format of the label information can be any data format as needed. For example, the first data format can be a Poly format, that is, [x1, y1, x2, y2, x3, y3, x4, y4, classID], where (x1, y1), (x2, y2), (x3, y3), and (x4, y4) respectively define the positions of the four vertices of the target detection box. During the preprocessing of the label information, a long-edge method representation defined by the center of the target detection box and its long side can be calculated based on the positions of the four vertices.
[0048] In addition, in the preprocessing stage, data enhancement can be performed on the training images by performing scaling, cropping, translation and other processing, so that the directed object detection network model can converge faster and the trained directed object detection network model can be more robust.
[0049] FIG2B shows a specific flow chart of an operation for inputting the plurality of training images and the label information into the directed object detection network to obtain setting parameters of the directed object detection network through training.
[0050] As shown in FIG2B , for each training, first, in step S211 , corresponding training images and corresponding label information for the current training are obtained.
[0051] In step S212, the directed object detection network is used to predict the object in the corresponding training image to obtain a corresponding prediction result.
[0052] As mentioned above, the prediction result may include classification information, target frame coordinate prediction information, target prediction information, and angle prediction information. The prediction result may be data in the form of [x, y, w, h, θ, conf, classID]. In one embodiment, the obtained prediction result in the form of [x, y, w, h, θ, conf, classID] may be post-processed to process the prediction result into data having a second data form, and the data having the second data form may be presented to the user, wherein (x, y) represents the position of the center of the target detection frame, w represents the width of the long side of the target detection frame, h represents the height of the short side of the target detection frame, θ represents the angle of rotation of the long side of the target detection frame relative to the horizontal axis, conf represents the confidence that the target detection frame has the corresponding target, and classID represents the category of the target detection frame. The second data form may be data in any form, for example, data in the form of a user interface to improve user convenience. In one example, the second data form is [x1, y1, x2, y2, x3, y3, x4, y4, conf, classID], where (x1, y1), (x2, y2), (x3, y3) and (x4, y4) respectively define the positions of the four vertices of the target detection box. In addition, in another example, the prediction results or a portion of the post-processed prediction results can be selectively presented to the user. For example, the presentation of the confidence conf can be omitted. Specifically, the remaining prediction results or their post-processing results other than the confidence conf can be presented to the user as needed. For example, a default confidence value can be set. In this case, the remaining prediction results for the predicted targets with a confidence higher than the default confidence value can be presented to the user.
[0053] In step S213, a loss function for the oriented object detection network is calculated based on the corresponding label information and the corresponding prediction results. In one embodiment, the loss function may be a cross-entropy loss function. As an example, a cross-entropy loss is calculated for each category and a weighted sum is performed to obtain a total loss. For example, the loss associated with angle prediction may be weighted to 0.1.
[0054] Finally, based on the calculation result of the loss function, the setting parameters of the directed target detection network are adjusted until the calculation result of the loss function converges. Specifically, it can be determined in step S214 whether the calculation result of the loss function converges, wherein in response to determining that the calculation result does not converge (step S214-no), step S215 is executed. In step S215, the setting parameters of the directed target detection network are adjusted. On the contrary, in response to determining that the calculation result converges (step S214-yes), step S216 is executed to end the training and output the trained setting parameters. At this point, the training of the directed target detection network based on YOLOV6 is completed.
[0055] The above describes a training method for a directed target detection network based on YOLOv6. After the directed target detection network is trained, the trained directed target detection network can be used to perform target detection. By performing target detection using a directed target detection network based on YOLOv6 according to an exemplary embodiment of the present disclosure, not only can the accuracy and prediction speed of target detection be improved, but the angle information of the target can also be obtained. Figure 3 is a flow chart of a directed target detection method according to an exemplary embodiment of the present disclosure.
[0056] As shown in FIG3 , a directed target detection method 300 according to an exemplary embodiment of the present disclosure may use a YOLOV6-based directed target detection network having a network architecture as shown in FIG1 , and the YOLOV6-based directed target detection network is trained via the method shown in FIG2A and FIG2B .
[0057] According to an exemplary embodiment of the present disclosure, a directed target detection method 300 may include: obtaining an image to be detected in step S310; and inputting the image to be detected into the directed target detection network in step S320 to obtain a target detection result. Specifically, the image to be detected is input into a trained directed target detection model, and a multi-level convolutional neural network trained in the directed target detection model is used to perform convolution operations and feature map fusion on the image to be detected, outputting feature fusion maps of different scales; for the feature fusion maps of different scales, corresponding predictors are used to predict category confidence, target frame position confidence, target detection confidence, and angle classification confidence.
[0058] The following will describe the test results on mouse brain data by using the directed target detection method according to an example embodiment of the present disclosure in conjunction with Figures 4 and 5, wherein (a) and (b) in Figure 4 respectively show mouse brain image data and labels on the image data, and the above test can be performed in the spatiotemporal omics ImageQC software. (a) and (b) in Figure 5 show the test results for mouse brain image data in the ImageQC software. During the test, 149 images with different fields of view were used, containing 6,850 targets. After multiple tests with input pictures at different angles, it was found that the directed target detection network based on YOLOV6 according to the example embodiment of the present disclosure or the directed target detection method based on YOLOV6 according to the example embodiment of the present disclosure can be applied to images at any angle, and can accurately and quickly detect targets in the image.
[0059] In addition, the present application also tests the detection performance of the directed target detection model or directed target detection method based on YOLOV6 according to the exemplary embodiment of the present disclosure compared with the directed target detection model or directed target detection method based on YOLOV5. The following Tables 1 and 2 respectively show the comparison results of the detection performance and the running performance.
[0060] Table 1
[0061]
[0062] Table 2 (unit: milliseconds)
[0063]
[0064] As can be seen from Table 1, the directed target detection network or directed target detection method based on YOLOV6 according to the exemplary embodiments of the present disclosure is slightly higher in precision, recall, mAP@0.5 (mean average precision mAP across all classes when the loss function is 0.5), and F1 value than the directed target detection network based on YOLOV5, while the mAP at a loss function with a step size of 0.5 and a threshold range of 0.5 to 0.95 (i.e., mAP@.5∶.95) is slightly lower than that of YOLOV5. In terms of operation, it can be seen from Table 2 that, compared with YOLOV5, in the directed target detection network of YOLOV6, the processing speed of the central processing unit (CPU) and the graphics processing unit (GPU) in the preprocessing stage, the prediction speed in the inference stage, and the speed in the non-maximum suppression (NMS) stage are all significantly improved.
[0065] The above describes a directed target detection network based on YOLOv6 and its training method, as well as a directed target detection method according to an exemplary embodiment of the present disclosure, which can improve the accuracy and speed of prediction.
[0066] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0067] FIG6 shows a schematic block diagram of an example electronic device 600 that can be used to implement an embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0068] As shown in Figure 6, device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. Various programs and data required for the operation of device 600 can also be stored in RAM 603. Computing unit 601, ROM 602, and RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to bus 604.
[0069] Various components in device 600 are connected to I / O interface 605, including an input unit 606, such as a keyboard, mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, optical disk, etc.; and a communication unit 609, such as a network card, modem, wireless communication transceiver, etc. The communication unit 609 allows device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0070] The computing unit 601 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and steps described above, such as the methods and steps shown in Figures 2 and 3. For example, in some embodiments, the methods and steps shown in Figures 2 to 3 can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as a storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the method for training a directed object detection network based on YOLO V6 and / or the directed object detection method described above can be performed. Alternatively, in other embodiments, the computing unit 601 can be configured in any other appropriate manner (for example, by means of firmware) to execute the method for training a directed target detection network model based on YOLOV6 as described above and / or the directed target detection method and its steps described above.
[0071] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0072] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0073] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0074] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0075] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0076] Computer systems may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The client and server relationship arises through computer programs running on the respective computers and having a client-server relationship to each other.
[0077] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions of this disclosure can be achieved, and this document is not limited here.
[0078] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A method for training a directed object detection network based on YOLOv6, the method comprising: acquiring a plurality of training images and label information of labels included in the plurality of training images; as well as Inputting the plurality of training images and the label information into the YOLOV6-based directed target detection network to obtain setting parameters of the YOLOV6-based directed target detection network through training; Among them, the YOLOV6-based directed target detection network includes: The backbone layer is configured to perform feature extraction on the training image to obtain multiple feature maps of different scales; A Neck layer is configured to receive the multiple feature maps from the Backbone layer and perform cross-scale fusion on the multiple feature maps to obtain multiple fused feature maps; and The detection head layer is configured to predict the multiple fused images to obtain a prediction result about the training image, The prediction results include: classification information, target frame coordinate prediction information, target prediction information and angle prediction information.
2. The method according to claim 1, wherein The Backbone layer adopts the RepVGG architecture, and the Neck layer adopts the FPN-PAN architecture.
3. The method according to claim 1, wherein The Head layer adopts a decoupling head, so that category analysis is performed through a classification branch to obtain the classification information, and regression analysis is performed through a regression branch to respectively obtain the target frame coordinate prediction information, the target prediction information and the angle prediction information.
4. The method according to claim 1, wherein Inputting the plurality of training images and the label information into the YOLOV6-based directed object detection network to obtain the setting parameters of the YOLOV6-based directed object detection network through training includes: in each training, Obtain the corresponding training images and corresponding label information for the current training; Use the YOLOV6-based directed target detection network to predict the target in the corresponding training image to obtain a corresponding prediction result; Calculating the loss function of the YOLOV6-based directed object detection network according to the corresponding label information and the corresponding prediction results; and Based on the calculation result of the loss function, the setting parameters of the YOLOV6-based directed target detection network are adjusted until the calculation result of the loss function converges.
5. The method according to claim 1, further comprising: Preprocessing the acquired label information in the first data format to process the label information into data in the form of [x, y, w, h, θ, classID], and inputting the processed data into the YOLOV6-based directed target detection network, Wherein, (x, y) represents the position of the center of the target detection frame, w represents the width of the long side of the target detection frame, h represents the height of the short side of the target detection frame, θ represents the angle of rotation of the long side of the target detection frame relative to the horizontal axis, and classID represents the category of the target detection frame.
6. The method according to claim 1, wherein The first data form of the label information is [x1, y1, x2, y2, x3, y3, x4, y4, classID], where (x1, y1), (x2, y2), (x3, y3) and (x4, y4) respectively define the positions of the four vertices of the target detection box.
7. The method according to claim 1, further comprising: Post-processing the obtained prediction result in the form of [x, y, w, h, θ, conf, classID] to process the prediction result into data in a second data form, and presenting the data in the second data form to the user, Wherein, (x, y) represents the position of the center of the target detection box, w represents the width of the long side of the target detection box, h represents the height of the short side of the target detection box, θ represents the angle of rotation of the long side of the target detection box relative to the horizontal axis, conf represents the confidence that the target detection box has the corresponding target, and classID represents the category of the target detection box.
8. The method according to claim 7, wherein: The second data format is [x1, y1, x2, y2, x3, y3, x4, y4, conf, classID], where (x1, y1), (x2, y2), (x3, y3) and (x4, y4) respectively define the positions of the four vertices of the target detection box.
9. A directed target detection method, comprising: Acquire the image to be tested; as well as Input the image to be tested into the YOLOV6-based directed target detection network to obtain the target detection result. The YOLOV6-based directed target detection network includes: The backbone layer is configured to perform feature extraction on the image to be tested to obtain multiple feature maps of different scales; A Neck layer is configured to receive the multiple feature maps from the Backbone layer and perform cross-scale fusion on the multiple feature maps to obtain multiple fused feature maps; and The detection head layer is configured to predict the multiple fused images to obtain the target detection result of the image to be tested. The target detection result includes: classification information, target frame coordinate prediction information, target prediction information and angle prediction information.
10. The method according to claim 9, wherein: The Backbone layer adopts the RepVGG architecture, and the Neck layer adopts the FPN-PAN architecture.
11. The method according to claim 9, wherein The Head layer adopts a decoupling head, so that category analysis is performed through a classification branch to obtain the classification information, and regression analysis is performed through a regression branch to respectively obtain the target frame coordinate prediction information, the target prediction information and the angle prediction information.
12. The method according to claim 9, further comprising: Post-processing the obtained target detection result in the form of [x, y, w, h, θ, conf, classID] to process the target detection result into data in a second data form, and presenting the data in the second data form to the user, Wherein, (x, y) represents the position of the center of the target detection box, w represents the width of the long side of the target detection box, h represents the height of the short side of the target detection box, θ represents the angle of rotation of the long side of the target detection box relative to the horizontal axis, conf represents the confidence that the target detection box has the corresponding target, and classID represents the category of the target detection box.
13. The method according to claim 12, wherein: The second data format is [x1, y1, x2, y2, x3, y3, x4, y4, conf, classID], where (x1, y1), (x2, y2), (x3, y3) and (x4, y4) respectively define the positions of the four vertices of the target detection box.
14. The directed target detection method according to claim 9, wherein: The YOLOV6-based directed target detection network is trained using the method described in any one of claims 1 to 8.
15. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 14.
16. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 14.
17. A computer program product comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method according to any one of claims 1 to 14.